Archive record traceability model construction method and system based on big data analysis

By constructing an archive record traceability model based on big data analysis, the problems of low efficiency and insufficient security of archive record traceability in the existing technology are solved, and efficient and accurate archive record traceability and causal analysis are achieved, which is adapted to the archive management needs in the big data environment.

CN119988317AActive Publication Date: 2025-05-13CHINA NAT INST OF STANDARDIZATION

Patent Information

Application Number
CN202510444331.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-10
Publication Date
2025-05-13
Estimated Expiration
2045-04-10

AI Technical Summary

Technical Problem

The existing archival record traceability methods are inefficient, insufficient security, limited traceability, and difficult to identify abnormal operations or tampering behaviors, which cannot adapt to the archive management needs in the big data environment.

Method used

Build an archive record traceability model based on big data analysis, and realize regular storage, privacy encryption, data verification, efficient traceability and causal prediction through a multi-level technical architecture, including building an archive traceability database, identity verification layer, data verification layer, evaluation prediction layer and optimization model layer.

Benefits of technology

It improves the efficiency and accuracy of archival record traceability, realizes accurate traceability and causal analysis of archival records, enhances data security and management capabilities, and adapts to the archive management needs in the big data environment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119988317A_ABST
    Figure CN119988317A_ABST
Patent Text Reader

Abstract

The invention discloses an archive record traceability model construction method and system based on big data analysis, and the method comprises the steps: obtaining archive information of an archive chain, carrying out the traceability coding, carrying out the data classification and partition storage of the archive information, setting a traceability key, constructing an identity verification layer, and carrying out the traceability coding. Extracting target archive information according to the tracing key to obtain target archive features, determining archive information integrity and a node to be traced, correspondingly determining a time sequence causal correlation vector according to the node to be traced, and constructing an archive record prediction model, and inputting target archive information into the archive record prediction model to output an archive record prediction result, optimizing the archive record traceability model according to the causal vector deviation, and performing record traceability on a target archive according to the optimized archive record traceability model. The method not only can improve the efficiency and accuracy of archive record traceability, but also has good interpretability, and can be directly applied to an archive record traceability system.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of archive traceability, and in particular to a method and system for constructing an archive record traceability model based on big data analysis. Background Art

[0002] With the rapid development of information technology, the number and types of archival records continue to increase. The security, integrity and traceability of archival information have become increasingly important. Effectively tracing the archival records and ensuring the security and integrity of archival information have become an urgent issue to be addressed.

[0003] Existing archival record traceability methods usually rely on manual operations and simple database queries. Their data storage lacks an effective classification and partitioning mechanism. When processing large-scale data, there are problems such as low efficiency, insufficient security and limited traceability capabilities. They are also easily affected by human factors, resulting in inaccurate traceability results. In addition, when faced with complex archival relationships and operations, traditional methods cannot effectively identify abnormal operations or tampering behaviors in archival records, lack effective prediction and evaluation mechanisms, and are difficult to adapt to the needs of archival management in a big data environment. The present invention proposes a method and system for constructing an archival record traceability model based on big data analysis, aiming to achieve regular storage, privacy encryption, data verification, efficient traceability and causal prediction of archival records through a multi-level technical architecture, and proposes a method for constructing an archival record traceability model based on big data analysis, which overcomes the shortcomings of existing archival record traceability methods, realizes effective temporal causal associations between operation records and archival data and accurate tracing of archival records, and improves the efficiency and accuracy of archival record traceability. Summary of the invention

[0004] The purpose of the present invention is to provide a method and system for constructing an archival record traceability model based on big data analysis.

[0005] To achieve the above object, the present invention is implemented according to the following technical solutions: The present invention comprises the following steps: Constructing an archive traceability database: obtaining archive chain archive information, performing traceability coding on the archive information, and classifying and partitioning the archive information; Constructing the identity authentication layer: Setting the traceability key, and verifying the operator's identity based on the traceability key to perform information reading or file recording operations; Constructing a data verification layer: extracting target archive information to obtain target archive features, and determining archive information integrity and nodes to be traced based on the target archive features; Constructing the evaluation prediction layer: determining the temporal causal correlation vector according to the archival features and historical operation vectors corresponding to the node to be traced and constructing an archival record prediction model, inputting the target archival information into the archival record prediction model to output the archival record prediction result; the archival record prediction result includes the archival operation vector and the causal vector; Optimizing the archive record traceability model: optimizing the archive record traceability model according to the causal vector deviation, and tracing the target archive records according to the optimized archive record traceability model.

[0006] Furthermore, the method for constructing the archive tracing database includes: Obtain the archive information of the entire archive chain, extract the starting timestamp corresponding to the start recording operation, extract the starting recording person corresponding to the start recording operation, and perform traceability coding on the archive information according to the starting timestamp, starting recording person and total number of records; The archival information is divided into archival metadata and archival entity data. The archival entity data is classified into entity data to obtain external input entity data, archival received entity data and actual recorded entity data. The archival entity data is partitioned and stored according to the entity data classification results. New archival metadata is generated according to the entity data classification results and partition storage conditions. Metadata classification is performed on all archival metadata to obtain descriptive metadata, management metadata and access metadata. The metadata is partitioned and stored according to the metadata classification results.

[0007] Furthermore, the method for constructing the identity authentication layer comprises the following steps: Set a traceability key, authenticate identity and obtain traceability authority based on the traceability key, and update archive information based on specific operations and corresponding document traceability codes; the traceability key includes a primary traceability key and a secondary traceability key; the traceability authority includes document access authority and document recording authority; Enter the first-level traceability key to obtain document access rights, view the document and generate access records, and find the corresponding archive information storage location through the traceability code to update the access metadata; the first-level traceability key consists of an access code, an authority code and an identity code. The access code is used to verify the access rights, the authority code is used to determine the access scope of the traceability database document, and the identity code is used to generate an access identity; Or enter the secondary traceability key to obtain document recording authority, make document records and find the corresponding archive information storage location through the traceability code to update the archive information; the secondary traceability key is composed of a record code, an authority code and an identity code, the record code is used to verify the record authority, the authority code is used to determine the document record scope of the traceability database, and the identity code is used to generate a record identity.

[0008] Furthermore, the method for constructing the data verification layer includes: Enter the first-level traceability key to obtain the document access permission, find the corresponding management metadata according to the target archive traceability code, extract the storage location of the target archive entity data in the management metadata, and extract the target archive entity data according to the storage location of the target archive entity data; Processing the target archive entity data to obtain target archive features, and using hierarchical clustering to divide the target archive features into target archive format features and target archive content features; the target archive format features are obtained by processing the target archive entity data that is finally recorded; the target archive content features include external input content features, archive received content features, and actual recorded content features; Compare the target archive format features with the standard archive format features, determine the integrity of the target archive and update the description metadata; Define each archive record operation on the archive chain as an archive record node, calculate the similarity of the feature vectors of the actual record content between adjacent archive record nodes, and determine the node to be traced based on the similarity calculation result.

[0009] Furthermore, the method for determining the temporal causal correlation vector includes: The data size and time operation parameters of each archive record node on the archive chain are determined according to the descriptive metadata, the sliding window size and step length are determined according to the data size and time operation parameters, and the sliding window method is used to calculate the cross-correlation of each window to obtain the temporal correlation of the node to be traced; Determine the record vector of the node to be traced according to the operation record in the management metadata, calculate the deviation between the external input content feature vector and the archive received content feature vector to obtain the operation deviation vector of the node to be traced, perform feature fusion on the record vector of the node to be traced and the archive received content feature vector to obtain the first action vector, perform feature crossover on the first action vector of the node to be traced and the operation deviation vector to obtain the second action vector, calculate the deviation between the actual recorded content feature of the node to be traced and the previous archive record node to obtain the content feature change vector, calculate the correlation between the content feature change vector of the node to be traced and the second action vector to obtain the causal correlation of the node to be traced; Define the traceability node i Content feature change vector , For the traceability node i No. j Content change feature elements, n Define the traceability node for the number of characteristic elements i Second motion vector , For the traceability node i No. j The second action element, the calculation expression of the causal correlation degree of the node to be traced is: , , , in For the traceability node i Content feature change vector With the second motion vector The causal relationship of is the relevance weight, For vector and The cosine Pearson correlation coefficient of For vector and The Pearson correlation coefficient, For vector and The mutual information of is the maximum mutual information, for and The joint probability density of for The marginal probability distribution of for The marginal probability distribution of For the traceability node i No. j Operational deviation, For the traceability node i Operational deviation mean; The temporal correlation and causal correlation of the target archive's traceable nodes are combined according to the time series to obtain a temporal correlation vector and a causal correlation vector, and the temporal correlation vector and the causal correlation vector of the target archive are feature fused to obtain a temporal causal correlation vector.

[0010] Furthermore, the method for outputting the archive record prediction result includes: The content matrix is ​​composed of the content features of the target archive, and the record matrix is ​​composed of the record vectors of the nodes to be traced. The deviation of the actual record content features of the initial and final record nodes of the target archive is calculated to obtain the start-end feature deviation vector. The content matrix, record matrix, start-end feature deviation vector and temporal causal correlation vector of the target archive are combined into an archive record traceability set. Constructing an archive record prediction model based on the archive record traceability set, the archive record prediction model includes a random forest algorithm, an elastic network regression algorithm and an Adam optimizer; the archive record prediction model is used to predict the impact of the document record operation of each traceability node on the overall document content, and extract the operation record of the corresponding traceability node; The random forest algorithm divides the archival record traceability set into training data and test data in a ratio of 7:3; The elastic network regression algorithm learns the relationship between the content matrix, the record matrix, the temporal causal correlation and the overall deviation vector, and performs regression analysis to predict the impact of each traceability node on the overall deviation, and outputs the operation causal vector; Adam optimizer is used to adaptively adjust the network learning rate; Input the target archive information into the archive record prediction model to output the archive record prediction result; the archive record prediction result includes an archive operation vector and a cause-effect vector.

[0011] Furthermore, the method for optimizing the archival record traceability model includes: Calculate the deviation between the predicted value and the actual value of the causal vector, and determine the objective function based on the causal vector deviation and the archive operation vector. The expression is: , in The objective function optimized for the archival record traceability model, is the causal vector i element prediction value, is the causal vector i The actual value of the element, m is the causal vector dimension, For the file operation vector i elements, is the mean of the archive operation vector elements, is the element-wise variance of the archive operation vector, is the hyperparameter of the explanatory term weight, For the model k parameters, K is the number of model parameters; Perform chaotic mapping on the search population and perform carrier operation to optimize the chaotic variables: , , in For the t +1 Chaos Variable, For the t A chaotic variable, T A chaotic variable, t is the current search count, is a chaotic random number, , For the i The first optimization variable j The optimized variable after carrier operation in the chaotic variable sequence is For thei The upper bound of the optimization variable, For the i The lower bounds of the optimization variables are For the i The optimization variable corresponds to j Chaotic variables; Calculate the population fitness and introduce the explosion operator to dynamically adjust the particle swarm according to the population fitness: , , in For fireworks i The number of sparks produced by the explosion, For fireworks i The magnitude of the explosion, h is the total number of sparks, is the maximum explosion amplitude, , are the maximum and minimum fitness of the population, respectively. g is the number of individuals currently involved in the calculation in the population, is the fitness of the current position of the fireworks particle, is a non-zero natural number; Perform crossover and mutation operations on the population to update the speed and position of the particles: , , , in For the i The causal deviation particle t +1 iteration speed, d For the dimension, For the i The causal deviation particle t The speed at the iteration, For the t The causal weight at the iteration, , , is a random number from (0,1), , is the acceleration constant, For the i The causal deviation particle t The position at the iteration, For the i The causal deviation particle t The historical best position at the iteration, For the tThe historical optimal position of all causal deviation particles at the iteration, , Divided into t The upper and lower bounds of the positions of all causal deviation particles at iterations, For the i The causal deviation particle t The position at +1 iteration, , is the causal weight at the beginning and end; Continue iterating until the objective function is minimized or the maximum number of iterations is reached; The target archives are recorded and traced according to the optimized archive record traceability model.

[0012] The second aspect is an archive record tracing system based on big data analysis, including: Archives traceability database module: used to obtain archive chain archive information, perform traceability coding on the archive information, and perform data classification and partition storage on the archive information; Identity verification module: used to verify the operator's identity based on the traceability key to perform information reading or file recording operations; Data verification module: used to extract target archive management metadata, extract target archive entity data according to the target archive management metadata, process the target archive entity data to obtain target archive features, and determine the integrity of the target archive and determine the node to be traced according to the target archive features; Evaluation and prediction module: used to determine the temporal causal correlation vector according to the archival features and historical operation vectors corresponding to the node to be traced and construct an archival record prediction model, input the target archival information into the archival record prediction model and output the archival record prediction result; Model optimization module: optimizing the archival record traceability model according to the causal vector deviation; Intelligent management module: used to store, view and manage the archive information and the archive record prediction results, generate materialized views according to the timeline, the archive operation vector and the causal vector, and display the archive record history to the user.

[0013] The beneficial effects of the present invention are: The present invention is a method and system for constructing an archival record traceability model based on big data analysis. Compared with the prior art, the present invention has the following technical effects: The present invention can improve the data preprocessing capabilities and enhance the model adaptability in archival record tracing by constructing an archival record traceability database, identity authentication, data verification, data evaluation, model construction and model optimization steps, thereby improving the efficiency and accuracy of archival record tracing, and optimizing the archival record tracing technology, which can greatly save resources, improve work efficiency, achieve tracing of archival records, and provide support for effective tracing and precise management of technical archives in various industries. It can adapt to different archival record tracing systems and the terminal tracing needs of archival records of different users, and has a certain universality. BRIEF DESCRIPTION OF THE DRAWINGS

[0014] Figure 1 The present invention is a flowchart of the steps of a method for constructing an archival record traceability model based on big data analysis. DETAILED DESCRIPTION

[0015] The present invention is further described below by means of specific embodiments. The illustrative embodiments and descriptions of the present invention are used to explain the present invention but are not intended to limit the present invention.

[0016] The present invention provides a method and system for constructing an archival record traceability model based on big data analysis, comprising the following steps: like Figure 1 As shown, in this embodiment, the following steps are included: Constructing an archive traceability database: obtaining archive chain archive information, performing traceability coding on the archive information, and classifying and partitioning the archive information; Constructing the identity authentication layer: Setting the traceability key, and verifying the operator's identity based on the traceability key to perform information reading or file recording operations; Constructing a data verification layer: extracting target archive information to obtain target archive features, and determining archive information integrity and nodes to be traced based on the target archive features; Constructing the evaluation prediction layer: determining the temporal causal correlation vector according to the archival features and historical operation vectors corresponding to the node to be traced and constructing an archival record prediction model, inputting the target archival information into the archival record prediction model to output the archival record prediction result; the archival record prediction result includes the archival operation vector and the causal vector; Optimizing the archive record traceability model: optimizing the archive record traceability model according to the causal vector deviation, and tracing the target archive records according to the optimized archive record traceability model.

[0017] In this embodiment, the method for constructing the archive tracing database includes: Obtain the archive information of the entire archive chain, extract the starting timestamp corresponding to the start recording operation, extract the starting recording person corresponding to the start recording operation, and perform traceability coding on the archive information according to the starting timestamp, starting recording person and total number of records; The archival information is divided into archival metadata and archival entity data, and the archival entity data is classified into entity data to obtain external input entity data, archival received entity data and actual recorded entity data, and the archival entity data is partitioned and stored according to the entity data classification result, and new archival metadata is generated according to the entity data classification result and the partition storage situation, and all archival metadata are classified into metadata to obtain description metadata, management metadata and access metadata, and the metadata is partitioned and stored according to the metadata classification result; the description metadata includes the operation time, recorder and size of entity data of each record node of the archive; the management metadata includes the operation record of each operation point and the storage location of entity data; In the actual evaluation, taking the work file management of a production task of a production line as an example, the management information including operators and operation time is obtained: OP001-2024-06-01-08:15, OP002-2024-06-01-10:30, OP003-2024-06-01-13:45, OP004-2024-06-02-08:00, OP005-2024- 06-02-11:00, OP001-2024-06-03-14:30, OP006-2024-06-04-09:15, OP007-2024-06-05- 16:00, OP002-2024-06-06-07:45, OP008-2024-06-07-10:00, OP009-2024-06-08-13:30; The traceability code is determined to be OP001-OP009-2024-06-01-08:15-2024-06-08-13:30, and the archive metadata and archive entity data are stored in partitions. The external input entity data is the adjustment parameters entered by the operator in the system, and the archive received entity data is the adjustment parameters actually executed by the system. The actual recorded entity data includes the adjusted machine production efficiency, production line transmission speed, number of workers, final finished product output, finished product sampling quality, defective rate and worker feedback.

[0018] In this embodiment, the method for constructing the identity authentication layer includes the following steps: Set a traceability key, authenticate identity and obtain traceability authority based on the traceability key, and update archive information based on specific operations and corresponding document traceability codes; the traceability key includes a primary traceability key and a secondary traceability key; the traceability authority includes document access authority and document recording authority; Enter the first-level traceability key to obtain document access rights, view the document and generate access records, and find the corresponding archive information storage location through the traceability code to update the access metadata; the first-level traceability key consists of an access code, an authority code and an identity code. The access code is used to verify the access rights, the authority code is used to determine the access scope of the traceability database document, and the identity code is used to generate the access identity; the access record includes the access identity, access time and access file traceability code; Or input the secondary traceability key to obtain the document record authority, record the document and find the corresponding archive information storage location through the traceability code to update the archive information; the secondary traceability key is composed of a record code, an authority code and an identity code, the record code is used to verify the record authority, the authority code is used to determine the document record range of the traceability database, and the identity code is used to generate the record identity; the content of the updated archive information includes the traceability code, archive metadata and archive entity data, retaining the original archive entity data and storing the external input entity data, archive receiving entity data and final record entity data of this record in partitions, updating the description metadata according to the recorder identity, recording time and final entity data size, updating the management metadata according to the specific operation of the document record, obtaining the starting timestamp, starting operator and total number of records according to the updated description metadata and management metadata, and re-updating the traceability code; In the actual evaluation, taking the example of the production line continuing to run without completing the production task, operator OP004 adjusted the transmission speed of the production line at 2024-06-09-08:00 according to the workers' work intensity feedback, entered its traceability key Record1998OP004, verified the record authority through the record code Record, determined the current recordable document range through the authority code 1998, generated the record identity through the identity code OP004, and updated the corresponding file information according to the traceability code, and updated the traceability code to OP001-OP004-2024-06-01-08:15-2024-06-09-08:00; Or taking the traceability of the target document as an example, the visitor enters his traceability key Query0218QP001, verifies the access right through the access code Query, determines the currently accessible document scope through the permission code 0218, generates the access identity through the identity code QP001, and generates the access record according to the specific access time and the traceability code of the accessed document.

[0019] In this embodiment, the method for constructing the data verification layer includes: Enter the first-level traceability key to obtain the document access permission, find the corresponding management metadata according to the target archive traceability code, extract the storage location of the target archive entity data in the management metadata, and extract the target archive entity data according to the storage location of the target archive entity data; Processing the target archive entity data to obtain target archive features; the method for processing the target archive entity data includes using TF-IDF to perform text classification on document data to calculate the text importance of each word to obtain text features, using a deep neural network combined with Mel frequency cepstral coefficients to perform speech recognition and classification on audio data to obtain audio features, using an image processing method to extract edge / color / grayscale / local information of image data to obtain image features, and using a pre-trained convolutional neural network to extract video image features of video data; the image processing method includes linear discriminant analysis, Canny edge detection, SIFT feature description, color histogram and grayscale co-occurrence matrix; hierarchical clustering is used to divide the target archive features into target archive format features and target archive content features; the target archive format features are obtained by processing the target archive entity data that is finally recorded; the target archive content features include external input content features, archive received content features and actual recorded content features; Compare the target archive format features with the standard archive format features, determine the integrity of the target archive and update the description metadata; Define each archive record operation on the archive chain as an archive record node, calculate the similarity of the feature vectors of the actual record content between adjacent archive record nodes, and when the similarity of the feature vectors of the actual record content between two adjacent archive record nodes is less than the similarity threshold, determine that the archive record operation of the archive record node that occurred later is abnormal, and set the archive record node as the node to be traced; In the actual assessment, the work information of a production task on a production line was traced back to the archive records, and the target archive format features were extracted. The target archive format features were compared with the standard archive format features to determine that the data integrity was 0.99, which was greater than the integrity threshold of 0.95, and the target archive data was considered complete. The similarities of the actual recorded content feature vectors between two adjacent archival record nodes are calculated (arranged in time series): 0.92, 0.89, 0.93, 0.96, 0.90, 0.95, 0.97, 0.87, 0.94, and 0.98. The adjacent archival record nodes whose content feature vector similarity is less than the similarity threshold of 0.93 are screened, and the archival record node that occurs later among the four groups of archival record nodes is determined as the node to be traced.

[0020] In this embodiment, the method for determining the temporal causal correlation degree vector includes: The data size and time operation parameters of each archive record node on the archive chain are determined according to the descriptive metadata, the sliding window size and step length are determined according to the data size and time operation parameters, and the sliding window method is used to calculate the cross-correlation of each window to obtain the temporal correlation of the node to be traced; Taking the archival record tracing as an example, the tracing node and the previous and next record nodes are determined as a sliding window. The window size is set according to the size of the archival information corresponding to the three record nodes. The sliding window step is set according to the time interval between the two tracing nodes. The corresponding step lengths are 3h15min, 48h45min, and 65h15min. Determine the record vector of the node to be traced according to the operation record in the management metadata, calculate the deviation between the external input content feature vector and the archive received content feature vector to obtain the operation deviation vector of the node to be traced, perform feature fusion on the record vector of the node to be traced and the archive received content feature vector to obtain the first action vector, perform feature crossover on the first action vector of the node to be traced and the operation deviation vector to obtain the second action vector, calculate the deviation between the actual recorded content feature of the node to be traced and the previous archive record node to obtain the content feature change vector, calculate the correlation between the content feature change vector of the node to be traced and the second action vector to obtain the causal correlation of the node to be traced; Define the traceability node i Content feature change vector , For the traceability node i No. j Content change feature elements, n Define the traceability node as the number of characteristic elements i Second motion vector , For the traceability node i No. j The second action element, the calculation expression of the causal correlation degree of the node to be traced is: , , , in For the traceability node i Content feature change vector With the second motion vector The causal relationship of is the relevance weight, For vector and The cosine Pearson correlation coefficient of For vector and The Pearson correlation coefficient, For vector and The mutual information of is the maximum mutual information, for and The joint probability density of for The marginal probability distribution of for The marginal probability distribution of For the traceability node i No. j Operational deviation, For the traceability node i Operational deviation mean; Combining the temporal correlation and causal correlation of the target archive's traceable nodes according to the time series to obtain a temporal correlation vector and a causal correlation vector, and fusing the temporal correlation vector and the causal correlation vector of the target archive to obtain a temporal causal correlation vector; In the actual evaluation, the temporal causal correlation vector obtained is (0.45, 0.85, 0.80, 0.35, 0.30, 0.88, 0.50, 0.25, 0.90, 0.40, 0.20).

[0021] In this embodiment, the method of outputting the archive record prediction result includes: The content matrix is ​​composed of the content features of the target archive, and the record matrix is ​​composed of the record vectors of the nodes to be traced. The deviation of the actual record content features of the initial and final record nodes of the target archive is calculated to obtain the start-end feature deviation vector. The content matrix, record matrix, start-end feature deviation vector and temporal causal correlation vector of the target archive are combined into an archive record traceability set. Constructing an archive record prediction model based on the archive record traceability set, the archive record prediction model includes a random forest algorithm, an elastic network regression algorithm and an Adam optimizer; the archive record prediction model is used to predict the impact of the document record operation of each traceability node on the overall document content, and extract the operation record of the corresponding traceability node; The random forest algorithm divides the archival record traceability set into training data and test data in a ratio of 7:3; The elastic network regression algorithm learns the relationship between the content matrix, the record matrix, the temporal causal correlation and the overall deviation vector, and performs regression analysis to predict the impact of each traceability node on the overall deviation, and outputs the operation causal vector; Adam optimizer is used to adaptively adjust the network learning rate; Input the target archive information into the archive record prediction model to output the archive record prediction result; the archive record prediction result includes an archive operation vector and a cause-effect vector.

[0022] In this embodiment, the method for optimizing the archival record traceability model includes: Calculate the deviation between the predicted value and the actual value of the causal vector, and determine the objective function based on the causal vector deviation and the archive operation vector. The expression is: , in The objective function optimized for the archival record traceability model, is the causal vector i element prediction value, is the causal vector i The actual value of the element, m is the dimension of the causal vector, For the file operation vector i elements, is the mean of the archive operation vector elements, is the element-wise variance of the archive operation vector, is the hyperparameter of the explanatory term weight, For the model k parameters, K is the number of model parameters; Perform chaotic mapping on the search population and perform carrier operation to optimize the chaotic variables: , , in For the t +1 Chaos Variable, For the t A chaotic variable, T A chaotic variable, t is the current search count, is a chaotic random number, , For the i The first optimization variable j The optimized variable after carrier operation in the chaotic variable sequence is For the i The upper bound of the optimization variable, For the i The lower bounds of the optimization variables are For the i The optimization variable corresponds to j Chaotic variables; Calculate the population fitness and introduce the explosion operator to dynamically adjust the particle swarm according to the population fitness: , , in For fireworks i The number of sparks produced by the explosion, For fireworks i The magnitude of the explosion, h is the total number of sparks, is the maximum explosion amplitude, , are the maximum and minimum fitness of the population, respectively. g is the number of individuals currently involved in the calculation in the population, is the fitness of the current position of the fireworks particle, is a non-zero natural number; Perform crossover and mutation operations on the population to update the speed and position of the particles: , , , in For the i The causal deviation particle t +1 iteration speed, d For the dimension, For the i The causal deviation particle t The speed at the iteration, For the t The causal weight at the iteration, , , is a random number from (0,1), , is the acceleration constant, For the i The causal deviation particle t The position at the iteration, For the i The causal deviation particle t The historical best position at the iteration, For the t The historical optimal position of all causal deviation particles at the iteration, , Divided into t The upper and lower bounds of the positions of all causal deviation particles at iterations, For the i The causal deviation particle t The position at +1 iteration, , is the causal weight at the beginning and end; Continue iterating until the objective function is minimized or the maximum number of iterations is reached; Conduct record tracing of the target archives according to the optimized archive record tracing model; In the actual evaluation, the traceability code of the target archive is obtained, and the target archive is traced according to the optimized archive record traceability model. The corresponding causal vector is (0.357, 0.786, 0.714, 0.268, 0.225, 0.816, 0.408, 0.163, 0.833, 0.3265, 0.1020). The final entity data changes caused by the archive record operation corresponding to the traceability point are as follows: 1. The machine productivity is adjusted at the OP002-2024-06-01-10:30 record point, causing the production efficiency to increase from 90% to 9 2. The defective rate dropped from 2% to 1.8% at the recording point OP003-2024-06-01-13:45, and the temperature of the input content was adjusted from 180°C to 182°C, and the defective rate was further reduced to 1.5%. 3. The recording point OP001-2024-06-03-14:30 optimized the final product, causing the sampling quality of the finished product to be improved to the excellent level, and the defective rate was reduced to 1%. 4. The recording point OP002-2024-06-06-07:45 increased the transmission speed of the production line to 25m / min, and the number of production workers increased by 0.2 people / m.

[0023] The second aspect is an archive record tracing system based on big data analysis, including: Archives traceability database module: used to obtain archive chain archive information, perform traceability coding on the archive information, and perform data classification and partition storage on the archive information; Identity verification module: used to verify the operator's identity based on the traceability key to perform information reading or file recording operations; Data verification module: used to extract target archive management metadata, extract target archive entity data according to the target archive management metadata, process the target archive entity data to obtain target archive features, and determine the integrity of the target archive and determine the node to be traced according to the target archive features; Evaluation and prediction module: used to determine the temporal causal correlation vector according to the archival features and historical operation vectors corresponding to the node to be traced and construct an archival record prediction model, input the target archival information into the archival record prediction model and output the archival record prediction result; Model optimization module: optimizing the archival record traceability model according to the causal vector deviation; Intelligent management module: used to store, view and manage the archive information and the archive record prediction results, generate materialized views according to the timeline, the archive operation vector and the causal vector, and display the archive record history to the user.

[0024] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principle of the present invention should be included in the protection scope of the present invention.

Claims

1. A method for constructing an archival record traceability model based on big data analysis, characterized in that: The following steps are involved: S1. Constructing an archive traceability database: obtaining archive chain archive information, performing traceability coding on the archive information, and classifying and partitioning the archive information; S2. Build an identity authentication layer: set a traceability key, and verify the operator's identity based on the traceability key to perform information reading or file recording operations; S3, building a data verification layer: extracting target archive information to obtain target archive features, and determining archive information integrity and nodes to be traced based on the target archive features; S4, constructing an evaluation prediction layer: determining a temporal causal correlation vector according to the archival features and historical operation vectors corresponding to the node to be traced, and constructing an archival record prediction model, inputting the target archival information into the archival record prediction model to output an archival record prediction result; the archival record prediction result includes an archival operation vector and a causal vector; S5. Optimizing the archive record traceability model: optimizing the archive record traceability model according to the causal vector deviation, and tracing the target archive records according to the optimized archive record traceability model.

2. The method for constructing an archival record traceability model based on big data analysis according to claim 1 is characterized in that: The method for constructing the archive tracing database comprises: Obtain the archive information of the entire archive chain, extract the starting timestamp corresponding to the start recording operation, extract the starting recording person corresponding to the start recording operation, and perform traceability coding on the archive information according to the starting timestamp, starting recording person and total number of records; The archival information is divided into archival metadata and archival entity data. The archival entity data is classified into entity data to obtain external input entity data, archival received entity data and actual recorded entity data. The archival entity data is partitioned and stored according to the entity data classification results. New archival metadata is generated according to the entity data classification results and partition storage conditions. Metadata classification is performed on all archival metadata to obtain descriptive metadata, management metadata and access metadata. The metadata is partitioned and stored according to the metadata classification results.

3. The method for constructing an archival record traceability model based on big data analysis according to claim 1 is characterized in that: The method for constructing the identity verification layer comprises the following steps: Set a traceability key, authenticate identity and obtain traceability authority based on the traceability key, and update archive information based on specific operations and corresponding document traceability codes; the traceability key includes a primary traceability key and a secondary traceability key; the traceability authority includes document access authority and document recording authority; Enter the first-level traceability key to obtain document access rights, view the document and generate access records, and find the corresponding archive information storage location through the traceability code to update the access metadata; the first-level traceability key consists of an access code, an authority code and an identity code. The access code is used to verify the access rights, the authority code is used to determine the access scope of the traceability database document, and the identity code is used to generate an access identity; Or enter the secondary traceability key to obtain document recording authority, make document records and find the corresponding archive information storage location through the traceability code to update the archive information; the secondary traceability key is composed of a record code, an authority code and an identity code, the record code is used to verify the record authority, the authority code is used to determine the document record scope of the traceability database, and the identity code is used to generate a record identity.

4. The method for constructing an archival record traceability model based on big data analysis according to claim 1 is characterized in that: The method for constructing the data verification layer comprises: Enter the first-level traceability key to obtain the document access permission, find the corresponding management metadata according to the target archive traceability code, extract the storage location of the target archive entity data in the management metadata, and extract the target archive entity data according to the storage location of the target archive entity data; Processing the target archive entity data to obtain target archive features, and using hierarchical clustering to divide the target archive features into target archive format features and target archive content features; the target archive format features are obtained by processing the target archive entity data that is finally recorded; the target archive content features include external input content features, archive received content features, and actual recorded content features; Compare the target archive format features with the standard archive format features, determine the integrity of the target archive and update the description metadata; Define each archive record operation on the archive chain as an archive record node, calculate the similarity of the feature vectors of the actual record content between adjacent archive record nodes, and determine the node to be traced based on the similarity calculation result.

5. The method for constructing an archival record traceability model based on big data analysis according to claim 1 is characterized in that: The method for determining the temporal causal correlation degree vector comprises: The data size and time operation parameters of each archive record node on the archive chain are determined according to the descriptive metadata, the sliding window size and step length are determined according to the data size and time operation parameters, and the sliding window method is used to calculate the cross-correlation of each window to obtain the temporal correlation of the node to be traced; Determine the record vector of the node to be traced according to the operation record in the management metadata, calculate the deviation between the external input content feature vector and the archive received content feature vector to obtain the operation deviation vector of the node to be traced, perform feature fusion on the record vector of the node to be traced and the archive received content feature vector to obtain the first action vector, perform feature crossover on the first action vector of the node to be traced and the operation deviation vector to obtain the second action vector, calculate the deviation between the actual recorded content feature of the node to be traced and the previous archive record node to obtain the content feature change vector, calculate the correlation between the content feature change vector of the node to be traced and the second action vector to obtain the causal correlation of the node to be traced; Define the traceability node i Content feature change vector , For the traceability node i No. j Content change feature elements, n Define the traceability node for the number of characteristic elements i Second motion vector , For the traceability node i No. j The second action element, the calculation expression of the causal correlation degree of the node to be traced is: , , , in For the traceability node i Content feature change vector With the second motion vector The causal relationship of is the relevance weight, For vector and The cosine Pearson correlation coefficient of For vector and The Pearson correlation coefficient, For vector and The mutual information of is the maximum mutual information, for and The joint probability density of for The marginal probability distribution of for The marginal probability distribution of For the traceability node i No. j Operational deviation, For the traceability node i Operational deviation mean; The temporal correlation and causal correlation of the target archive's traceable nodes are combined according to the time series to obtain a temporal correlation vector and a causal correlation vector, and the temporal correlation vector and the causal correlation vector of the target archive are feature fused to obtain a temporal causal correlation vector.

6. The method for constructing an archival record traceability model based on big data analysis according to claim 1 is characterized in that: The method for outputting the archive record prediction result comprises: The content matrix is ​​composed of the content features of the target archive, and the record matrix is ​​composed of the record vectors of the nodes to be traced. The deviation of the actual record content features of the initial and final record nodes of the target archive is calculated to obtain the start-end feature deviation vector. The content matrix, record matrix, start-end feature deviation vector and temporal causal correlation vector of the target archive are combined into an archive record traceability set. Constructing an archive record prediction model based on the archive record traceability set, the archive record prediction model includes a random forest algorithm, an elastic network regression algorithm and an Adam optimizer; the archive record prediction model is used to predict the impact of the document record operation of each traceability node on the overall document content, and extract the operation record of the corresponding traceability node; The random forest algorithm divides the archival record traceability set into training data and test data in a ratio of 7:3; The elastic network regression algorithm learns the relationship between the content matrix, the record matrix, the temporal causal correlation and the overall deviation vector, and performs regression analysis to predict the impact of each traceability node on the overall deviation, and outputs the operation causal vector; Adam optimizer is used to adaptively adjust the network learning rate; Input the target archive information into the archive record prediction model to output the archive record prediction result; the archive record prediction result includes an archive operation vector and a cause-effect vector.

7. The method for constructing an archival record traceability model based on big data analysis according to claim 1 is characterized in that: The method for optimizing the archival record traceability model comprises: Calculate the deviation between the predicted value and the actual value of the causal vector, and determine the objective function based on the causal vector deviation and the archive operation vector. The expression is: , in The objective function for optimizing the archival record traceability model, is the causal vector i element prediction value, is the causal vector i The actual value of the element, m is the dimension of the causal vector, For the file operation vector i elements, is the mean of the archive operation vector elements, is the element-wise variance of the archive operation vector, is the hyperparameter of the explanatory term weight, For the model k parameters, K is the number of model parameters; Perform chaotic mapping on the search population and perform carrier operation to optimize the chaotic variables: , , in For the t +1 Chaos Variable, For the t A chaotic variable, T A chaotic variable, t is the current search count, is a chaotic random number, , For the i The first optimization variable j The optimized variable after carrier operation in the chaotic variable sequence is For the i The upper bound of the optimization variable, For the i The lower bounds of the optimization variables are For the i The optimization variable corresponds to j Chaotic variables; Calculate the population fitness and introduce the explosion operator to dynamically adjust the particle swarm according to the population fitness: , , in For fireworks i The number of sparks produced by the explosion, For fireworks i The magnitude of the explosion, h is the total number of sparks, is the maximum explosion amplitude, , are the maximum and minimum fitness of the population, respectively. g is the number of individuals currently involved in the calculation in the population, is the fitness of the current position of the fireworks particle, is a non-zero natural number; Perform crossover and mutation operations on the population to update the speed and position of the particles: , , , in For the i The causal deviation particle t +1 iteration speed, d For the dimension, For the i The causal deviation particle t The speed at the iteration, For the t The causal weight at the iteration, , , is a random number from (0,1), , is the acceleration constant, For the i The causal deviation particle t The position at the iteration, For the i The causal deviation particle t The historical best position at the iteration, For the t The historical optimal position of all causal deviation particles at the iteration, , Divided into t The upper and lower bounds of the positions of all causal deviation particles at iterations, For the i The causal deviation particle t The position at +1 iteration, , is the causal weight at the beginning and end; Continue iterating until the objective function is minimized or the maximum number of iterations is reached; The target archives are recorded and traced according to the optimized archive record traceability model.

8. An archive record tracing system based on big data analysis, used to execute the method described in any one of claims 1 to 7, characterized in that: include: Archives traceability database module: used to obtain archive chain archive information, perform traceability coding on the archive information, and perform data classification and partition storage on the archive information; Identity verification module: used to verify the operator's identity based on the traceability key to perform information reading or file recording operations; Data verification module: used to extract target archive management metadata, extract target archive entity data according to the target archive management metadata, process the target archive entity data to obtain target archive features, and determine the integrity of the target archive and determine the node to be traced according to the target archive features; Evaluation and prediction module: used to determine the temporal causal correlation vector according to the archival features and historical operation vectors corresponding to the node to be traced and construct an archival record prediction model, input the target archival information into the archival record prediction model and output the archival record prediction result; Model optimization module: optimizing the archival record traceability model according to the causal vector deviation; Intelligent management module: used to store, view and manage the archive information and the archive record prediction results, generate materialized views according to the timeline, the archive operation vector and the causal vector, and display the archive record history to the user.

Citation Information

Patent Citations

  • Document traceability device, system and method

    CN108304724A

  • A method for tracing and encrypting employee archives data based on block chain technology

    CN109460666A

  • Tracing method and system for fault root cause of Internet of Things equipment, and storage medium

    CN117743843A

  • APT attack detection and tracing method based on graph attention sequential network

    CN117749437A

  • File data processing method and device, equipment and storage medium

    CN117909556A

Cited By

  • File global sensing control intelligent protection method and system

    CN121093387A