Electronic data evidence chain curing method and system based on block chain technology
By using blockchain technology and text classification model in the solidification of electronic data evidence chains, the problems of easy tampering, untraceable and low storage efficiency in the existing technology are solved, and the security, integrity and evidence effectiveness of data are improved.
Patent Information
- Application Number
- CN202510207531.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-25
- Publication Date
- 2025-06-13
AI Technical Summary
The prior art has problems such as tampering, untraceable and inefficient storage in the solidification of evidence links of electronic data, especially in ensuring data privacy and the applicability of legal evidence.
The electronic data evidence chain solidification method based on blockchain technology is adopted, and real-time target evidence data is classified and mapped through text classification model, combining encryption processing and multi-signature technology to ensure the security and integrity of the data.
It realizes accurate classification of electronic data and solidification of evidence links, ensures the security and integrity of data, enhances the flexibility and reliability of data sharing, and improves the effectiveness and credibility of evidence in the legal and financial fields.
Smart Images

Figure CN120144665A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of blockchain, and particularly to a method and system for solidifying an electronic data evidence chain based on blockchain technology. Background Art
[0002] In modern society, electronic data has become the main form of information exchange and storage, and is widely used in multiple fields such as law, finance, medical care, and education. However, the easy tampering and difficult traceability of electronic data pose many challenges in its application as legal evidence, including easy tampering, difficult verification, and difficult traceability.
[0003] The easy tampering means that electronic data can be easily copied, modified, or deleted, making it difficult to be used as reliable evidence in legal disputes.
[0004] The difficult verification means that traditional electronic data storage methods lack effective verification mechanisms and it is difficult to ensure the authenticity and integrity of the data.
[0005] The difficult traceability means that the generation, transmission, and storage processes of electronic data are complex, lacking effective traceability means, and it is difficult to restore the historical state of the data.
[0006] As a new distributed ledger technology, blockchain provides a new solution for the storage and management of electronic data with its characteristics of decentralization, immutability, and traceability. Through a distributed network structure, blockchain eliminates the dependence on centralized institutions, improves the reliability and security of the system. Once the data in the blockchain is written, it cannot be tampered with, ensuring the integrity and authenticity of the data. Through timestamps and a chain structure, blockchain records all historical operations of the data, providing a complete traceability path.
[0007] Although blockchain technology has achieved remarkable success in the financial field, there are still some deficiencies in the solidification of electronic data evidence chains. The open and transparent characteristics of blockchain may lead to the leakage of sensitive data. How to achieve public verification of data while ensuring data privacy is a challenge. The storage efficiency of blockchain is relatively low. How to improve the storage efficiency while ensuring data security is an urgent problem to be solved. The application of blockchain technology in the legal field is still in the exploration stage. How to ensure its applicability and recognition in legal evidence requires further research.
[0008] CN115048670A discloses an encryption and deposit evidence method, device, equipment and storage medium based on blockchain. The encryption and deposit evidence method is applied to a blockchain system. The encryption and deposit evidence method specifically includes: a first user terminal obtains a second user public key of a second user terminal from a deposit evidence system; encrypts the evidence data to be deposited according to the second user public key and a first user private key of the first user terminal to obtain target deposit evidence data; and stores the target deposit evidence data and a first user public key of the first user terminal in the deposit evidence system. This patent encrypts through different user public keys, which is not similar to the present invention, and this patent does not classify the original data, resulting in poor technical effects.
[0009] CN116052171A discloses an electronic evidence relevance calibration method, device, equipment and storage medium. First, text features of text-state electronic evidence and picture features of picture-state electronic evidence are extracted. Then, for each of the text features and the picture features, the relevance between the constituent elements in the features is calculated based on an attention mechanism, and the features are updated using the relevance to obtain updated text features and updated picture features. The updated text features and the updated picture features are fused to obtain a fused feature. Finally, based on the fused feature, the relevance between the text-state electronic evidence and the picture-state electronic evidence is calibrated. This patent only calibrates the relevance of electronic evidence. Electronic evidence has deficiencies such as being easily tampered with and non-traceable, and this patent does not classify the original data, resulting in poor technical effects. Summary of the Invention
[0010] To solve the deficiencies such as being easily tampered with, non-traceable, and poor technical effects in the prior art, the present invention provides an electronic data evidence chain solidification method and system based on blockchain technology, which realizes the precise classification of electronic data, solidifies various evidence chains, and ensures the security and integrity of data.
[0011] The present invention adopts the following technical solutions.
[0012] On the one hand, the present invention discloses an electronic data evidence chain solidification method, including:
[0013] Step 1: Collect real-time target evidence data from multiple data sources; the format of the target evidence data is text format;
[0014] Step 2: Classify the real-time target evidence data based on a text classification model to obtain various types of data to be mapped; the text classification model is constructed based on a pre-trained language model, a gated recurrent unit, and a one-dimensional convolutional model;
[0015] Step 3: Map various types of data to be mapped based on a preset evidence chain standard, and perform data conversion on the mapped data based on a preset storage format to obtain standard evidence chain data;
[0016] Step 4: Perform encryption processing and blockchain operation on the standard evidence chain data to obtain an evidence blockchain containing the standard evidence chain data;
[0017] Step 5: Use multi-signature and distributed storage technologies to protect the data security in the evidence blockchain and complete the solidification of the electronic data evidence chain.
[0018] Further preferably,
[0019] In step 1, the multiple data sources include file storage systems, databases, emails, social media, sensors, web pages, user inputs, audio, legal documents, and logs;
[0020] The target evidence data includes:
[0021] Text extracted from file storage systems;
[0022] Text extracted from databases;
[0023] Email content and attachments extracted from email systems;
[0024] Posts, comments, and messages extracted from social media platforms;
[0025] Text obtained by sensor acquisition;
[0026] Text content obtained by web scraping;
[0027] Text content input by users;
[0028] Text obtained by audio conversion;
[0029] Legal documents;
[0030] Log text data generated by systems, applications, or Internet of Things devices.
[0031] Further preferably,
[0032] In step 2, the following steps are included:
[0033] 2.1 Construct a text classification model based on a pre-trained language model, a gated recurrent unit, and a one-dimensional convolutional model;
[0034] 2.2 Collect historical target evidence data to be classified, annotate the historical target evidence data to be classified to obtain an annotated data set; train the text classification model based on the annotated data set;
[0035] 2.3 Collect real-time target evidence data and input it into the text classification model to obtain the classification results of the real-time target evidence data, and obtain various types of data to be mapped.
[0036] Further preferably,
[0037] The text classification model includes an input layer, a pre-trained language model layer, a forward gated recurrent unit, a backward gated recurrent unit, a context relationship feature acquisition layer, a first one-dimensional convolutional layer, a second one-dimensional convolutional layer, a third one-dimensional convolutional layer, an activation layer, a first pooling layer, a pooling result integration layer, a dropout layer, a concatenation layer, a second pooling layer, a fully connected layer, and an output layer.
[0038] Further preferably,
[0039] The context relationship feature acquisition layer obtains the overall context relationship feature based on the outputs of the forward gated recurrent unit and the backward gated recurrent unit;
[0040] The calculation formula for the overall context relationship feature is:
[0041]
[0042] where h t represents the context relationship feature of the forward and backward gated recurrent units at time t; G represents the overall context relationship feature; h 1 represents the context relationship feature of the forward and backward gated recurrent units at the first moment; h 2 represents the context relationship feature of the forward and backward gated recurrent units at the second moment; h l the context relationship feature of the forward and backward gated recurrent units at the l-th moment; l is the total number of time steps.
[0043] Further preferably,
[0044] Denote the convolutional kernels of the first one-dimensional convolutional layer, the second one-dimensional convolutional layer, and the third one-dimensional convolutional layer as the first convolutional kernel, the second convolutional kernel, and the third convolutional kernel respectively;
[0045] The selection of the first, second, and third convolutional kernels has the following constraints:
[0046]
[0047] where L 1 represents the length of the first convolutional kernel; L 2 represents the length of the second convolutional kernel; L 3 represents the length of the third convolutional kernel; W 1 represents the width of the first convolutional kernel; W 2 represents the width of the second convolutional kernel; W 3represents the width of the third convolutional kernel; max{·} represents taking the maximum value within the brackets; min{·} represents taking the minimum value within the brackets; N L represents the total number of characters in each line of the text; N W the number of lines in the text; |·| represents taking the absolute value.
[0048] Further preferably,
[0049] In step 2.2, the text classification model is trained based on the labeled dataset, and the loss function L oss is shown as follows:
[0050]
[0051] where i represents the i-th sample; N is the number of samples; c represents the c-th label category, and C is the total number of labeled label categories; y i,c is a binary indicator, which is 1 when the true label category of the i-th sample is category c, and 0 otherwise; is the prediction coefficient that the i-th sample belongs to the c-th label category.
[0052] Further preferably,
[0053] The prediction coefficient that the i-th sample belongs to the c-th label category is shown as follows:
[0054]
[0055] where max{·} represents taking the maximum value within the brackets; e is the natural base; j represents the j-th label category; Z i,c is the confidence that the model has for the i-th sample belonging to category c; Z i,j is the confidence that the model has for the i-th sample belonging to category j; A t is shown as follows:
[0056]
[0057] where y r is the true label category of the i-th sample; y f is the predicted label category of the i-th sample.
[0058] Further preferably,
[0059] In step 3, the preset standard for the evidence chain is as follows: each evidence chain includes fields of "DocumentID", "Title", "Content", and "Timestamp"; where "DocumentID" is the document number, "Title" is the title, "Content" is the content, and "Timestamp" is the timestamp.
[0060] On the other hand, the present invention discloses an electronic data evidence chain solidification system based on an electronic data evidence chain solidification method, including a data acquisition module, a data conversion module, a data encryption module, and a storage module:
[0061] The data acquisition module is used to collect real-time target evidence data from multiple data sources; the format of the target evidence data is text format.
[0062] The data conversion module is used to perform data classification, data mapping, and data conversion on the real-time target evidence data to obtain standard evidence chain data.
[0063] The data encryption module performs encryption processing and on-chain operation on the standard evidence chain data to obtain an evidence blockchain containing the standard evidence chain data.
[0064] The storage module uses multi-signature and distributed storage technologies to protect the secure storage of data in the evidence blockchain and complete the solidification of the electronic data evidence chain.
[0065] On the other hand, the present application discloses an electronic device, including a processor and a storage medium; characterized in that:
[0066] The storage medium is used to store instructions.
[0067] The processor is used to operate according to the instructions to execute the foregoing electronic data evidence chain solidification method.
[0068] The present application also discloses a computer-readable storage medium, on which a computer program is stored, characterized in that when the program is executed by a processor, the electronic data evidence chain solidification method is implemented.
[0069] The beneficial effect of the present invention is that, compared with the prior art,
[0070] The electronic data evidence chain solidification method and system based on blockchain technology provided by the present invention achieve accurate text classification through a text classification model constructed based on a pre-trained language model, a gated recurrent unit, and a one-dimensional convolutional model, as well as the selection constraints of the first, second, and third convolutional kernels proposed by the present invention, and the loss function proposed by the present invention.
[0071] By combining the decentralized, immutable, and traceable characteristics of blockchain technology, the present invention realizes the solidification of the evidence chain of electronic data, ensuring the security and integrity of the data.
[0072] Through the application of data encryption and smart contracts, the present invention not only protects data privacy but also enhances the flexibility and reliability of data sharing.
[0073] The present invention further improves the security of data storage by introducing multi-signature and distributed storage technologies, ensuring the evidentiary effect and credibility in applications in fields such as law and finance;
[0074] The design of the management module of the present invention not only improves the security and reliability of the system but also provides users with convenient means of management and maintenance. BRIEF DESCRIPTION OF THE DRAWINGS
[0075] Figure 1 It is a schematic flowchart of a method for solidifying the evidence chain of electronic data based on blockchain technology according to the present invention;
[0076] Figure 2 It is a schematic structural diagram of a system for solidifying the evidence chain of electronic data based on blockchain technology provided by an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0077] To make the objectives, technical solutions, and advantages of the present invention clearer, the technical solutions of the present invention will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of the present invention. The embodiments described in this application are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the spirit of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts fall within the protection scope of the present invention.
[0078] This application discloses a method for solidifying the evidence chain of electronic data based on blockchain technology, as Figure 1 shown, including:
[0079] Step 1: Collect real-time target evidence data from multiple data sources; the format of the target evidence data is text format;
[0080] The multiple data sources include file storage systems, databases, emails, social media, sensors, web pages, user inputs, audio, legal documents, and logs;
[0081] The target evidence data includes:
[0082] Texts extracted from file storage systems;
[0083] Texts extracted from databases;
[0084] The content and attachments extracted from the email system;
[0085] The posts, comments, and messages extracted from the social media platform;
[0086] The text collected by the sensor;
[0087] The text content obtained by web scraping;
[0088] The text content input by the user;
[0089] The text obtained by audio conversion;
[0090] Legal documents;
[0091] The log text data generated by the system, application, or Internet of Things device.
[0092] Step 2: Classify the real-time target evidence data based on the text classification model to obtain various types of data to be mapped; the text classification model is constructed based on a pre-trained language model, a gated recurrent unit, and a one-dimensional convolutional model;
[0093] In step 2, the following steps are included:
[0094] 2.1 Construct a text classification model based on a pre-trained language model, a gated recurrent unit, and a one-dimensional convolutional model;
[0095] The text classification model includes an input layer, a pre-trained language model layer, a forward gated recurrent unit, a backward gated recurrent unit, a context relationship feature acquisition layer, a first one-dimensional convolutional layer, a second one-dimensional convolutional layer, a third one-dimensional convolutional layer, an activation layer, a first pooling layer, a pooling result integration layer, a dropout layer, a concatenation layer, a second pooling layer, a fully connected layer, and an output layer.
[0096] The context relationship feature acquisition layer obtains the overall context relationship feature based on the outputs of the forward gated recurrent unit and the backward gated recurrent unit;
[0097] The calculation formula for the overall context relationship feature is:
[0098]
[0099] where h t represents the context relationship feature of the forward and backward gated recurrent units at time step t; G represents the overall context relationship feature; h 1 represents the context relationship feature of the forward and backward gated recurrent units at the first time step; h 2 represents the context relationship feature of the forward and backward gated recurrent units at the second time step; h l The context relationship feature of the forward and backward gated recurrent units at the l-th time step; l is the total number of time steps.
[0100] Denote the convolution kernels of the first one-dimensional convolution layer, the second one-dimensional convolution layer, and the third one-dimensional convolution layer as the first convolution kernel, the second convolution kernel, and the third convolution kernel respectively;
[0101] The selection of the first, second, and third convolution kernels has the following constraints:
[0102]
[0103] where L 1 represents the length of the first convolution kernel; L 2 represents the length of the second convolution kernel; L 3 represents the length of the third convolution kernel; W 1 represents the width of the first convolution kernel; W 2 represents the width of the second convolution kernel; W 3 represents the width of the third convolution kernel; max{·} represents taking the maximum value within the brackets; min{·} represents taking the minimum value within the brackets; N L represents the total number of characters in each line of the text; N W represents the number of lines in the text; |·| represents taking the absolute value.
[0104] 2.2 Collect historical target evidence data to be classified, label the historical target evidence data to be classified to obtain a labeled data set; train the text classification model based on the labeled data set;
[0105] When training the text classification model based on the labeled data set, the loss function L oss is as shown in the following formula:
[0106]
[0107] where i represents the i-th sample; N is the number of samples; c represents the c-th label category, and C is the total number of label categories obtained by labeling; y i,c is a binary indicator, which is 1 when the true label category of the i-th sample is category c, and 0 otherwise; is the prediction coefficient of the i-th sample belonging to the c-th label category.
[0108] The prediction coefficient of the i-th sample belonging to the c-th label category is as shown in the following formula:
[0109]
[0110] where max{·} represents taking the maximum value within the brackets; e is the natural base; j represents the j-th label category; Z i,c is the confidence of the model in the i-th sample belonging to category c; Z i,jis the model's confidence that the i-th sample belongs to category j; A t As shown below:
[0111]
[0112] Among them, y r is the true label category of the i-th sample; y f is the predicted label category of the i-th sample.
[0113] 2.3 Collect real-time target evidence data and input it into the text classification model to obtain the classification results of the real-time target evidence data and obtain various types of data to be mapped.
[0114] Step 3: Map various types of data to be mapped based on the preset evidence chain standard, and convert the mapped data based on the preset storage format to obtain standard evidence chain data;
[0115] The preset evidence chain standard is: each evidence chain includes the fields of "DocumentID", "Title", "Content", and "Timestamp"; "DocumentID" is the file number, "Title" is the title, "Content" is the content, and "Timestamp" is the timestamp.
[0116] Step 4: Encrypt and upload the standard evidence chain data to obtain an evidence blockchain containing the standard evidence chain data;
[0117] Step 5: Use multi-signature and distributed storage technology to protect the data security in the evidence blockchain and complete the solidification of the electronic data evidence chain.
[0118] Example 1
[0119] A method for solidifying an electronic data evidence chain based on blockchain technology.
[0120] Step 1: Collect target evidence data from various data sources; the target evidence data is in text format.
[0121] Preferably, the data source includes a file storage system, a database, email, social media, a sensor, an IoT device, a web page, user input, audio, legal documents, a log, etc. The database includes a relational database, a non-relational database, etc.
[0122] Target evidence data include:
[0123] Extract text data from documents, reports, contracts, agreements, etc. in file storage systems (local, cloud storage, etc.). Examples: PDF files, Word documents, spreadsheets, etc.
[0124] Extract the stored text data from the database. Examples: text records in SQL databases and NoSQL databases.
[0125] Extract the email content, attachments, and metadata from the email system.
[0126] Extract the user-generated content from social media platforms, including posts, comments, and messages.
[0127] Extract the text data obtained through sensors; the text data obtained through sensors is the text recorded by logging sensor data, and the recorded text is the text data obtained through sensors.
[0128] Extract the text data generated by Internet of Things devices, which is usually used to record events or states. Examples: log information of smart home devices, industrial sensors, etc.
[0129] Extract the text content obtained through web scraping, that is, the text content scraped from web pages by web crawler technology. Examples: public information on news websites, forums, blogs, etc.
[0130] Extract the text data input by users, that is, the text data directly input through user interfaces (such as forms, applications). Examples: online surveys, feedback forms, registration information, etc.
[0131] Convert audio to text, that is, convert audio files to text data through speech recognition technology. Examples: meeting records, telephone recordings, interview content, etc.
[0132] Legal documents, including the text data extracted from legal documents such as courts and arbitration institutions. Examples: judgments, pleadings, mediation agreements, etc.
[0133] Extract the text data from the log files generated by the system or application. Examples: server logs, application logs, audit logs, etc.
[0134] Step 2: Classify the target evidence data to obtain various types of data to be mapped;
[0135] The data classification is to use a preset text classification model to perform context analysis and classification on the target evidence data, assign the data to different categories, and obtain various types of data to be mapped. The preset text classification model is preferably a machine learning or natural language processing model.
[0136] 2.1 Build a text classification model based on pre-trained language models, gated recurrent units, context relationship feature acquisition layers, one-dimensional convolutional models, etc.; among them, the context relationship feature acquisition layer calculates based on the output results of the gated recurrent unit to obtain context relationship features.
[0137] The text classification model includes an input layer, a pre-trained language model layer, a forward gated recurrent unit, a backward gated recurrent unit, a context relationship feature acquisition layer, a first one-dimensional convolutional layer, a second one-dimensional convolutional layer, a third one-dimensional convolutional layer, an activation layer, a first pooling layer, a pooling result integration layer, a dropout layer, a concatenation layer, a second pooling layer, a fully connected layer, and an output layer;
[0138] The input layer receives the historical target evidence data to be classified and inputs the historical target evidence data into a pre-trained XLNet model, a first one-dimensional convolutional layer, a second one-dimensional convolutional layer, and a third one-dimensional convolutional layer respectively;
[0139] The pre-trained language model layer inputs the historical target evidence data into the pre-trained language model and takes the output vector of the pre-trained language model as a feature, which is input into the forward gated recurrent unit in the forward order (from left to right) and input into the backward gated recurrent unit in the reverse order (from right to left); the pre-trained language model is used to perform preliminary feature extraction on the historical target evidence data.
[0140] Those skilled in the art should know that the pre-trained language model is an artificial intelligence model that obtains knowledge through unsupervised learning on a large-scale text corpus. In practical applications, the pre-trained language model can be a pre-trained XLNet model, a pre-trained BERT model, a pre-trained ELMo model, etc.; those skilled in the art can make a choice according to the actual situation; the pre-trained language model provided in the embodiment of the present invention is only a preferred embodiment and is not an inevitable limitation for implementing the present invention.
[0141] Preferably, the present invention selects a pre-trained XLNet model as the pre-trained language model to perform preliminary feature extraction on the historical target evidence data.
[0142] The XLNet model is a pre-trained language model based on the combination of autoregression and autoencoding. Its core part is the Transformer-XL architecture, which solves the defect of the BERT model, that is, the Bidirectional Encoder Representations from Transformers model, in terms of the upper and lower context input length constraints, so that the pre-trained model can learn more distant upper and lower context information.
[0143] Taking the output vector of the pre-trained XLNet model as a feature, it is input into the forward gated recurrent unit in the forward order (from left to right) and input into the backward gated recurrent unit in the reverse order (from right to left); the pre-trained language model is used to perform preliminary feature extraction on the historical target evidence data.
[0144] The forward gated recurrent unit obtains the forward dependency at time step t based on the input at time step t - 1, that is, the output of the forward gated recurrent unit at time t Output the forward dependencies of each time step to the context relationship feature acquisition layer; where t ∈ [1, l], and l is the total number of time steps; set to 1;
[0145] The backward gated recurrent unit obtains the backward dependency at time step t based on the input at time step t - 1, that is, the output of the backward gated recurrent unit at time t Output the backward dependencies of each time step to the context relationship feature acquisition layer; where t ∈ [1, l], and l is the total number of time steps; set to 0;
[0146] The context relationship feature acquisition layer obtains the overall context relationship feature based on the outputs of the forward gated recurrent unit and the backward gated recurrent unit
[0147] Preferably, the calculation formula for the overall context relationship feature is:
[0148]
[0149] where h t represents the context relationship feature of the forward and backward gated recurrent units at time t; G represents the overall context relationship feature; h 1 represents the context relationship feature of the forward and backward gated recurrent units at the first time step; h 2 represents the context relationship feature of the forward and backward gated recurrent units at the second time step; h l The context relationship feature of the forward and backward gated recurrent units at the l-th time step. Where represents the exclusive OR logical operation, which is 1 when the symbols before and after are different and 0 when they are the same
[0150] In summary, the present invention obtains the context relationship feature of the historical target evidence data through the pre-trained XLNet model, the forward gated recurrent unit, the backward gated recurrent unit, and the context relationship feature acquisition layer
[0151] The first one-dimensional convolutional layer performs a one-dimensional convolutional operation on the input historical target evidence data and transmits the output result to the activation layer
[0152] The second one-dimensional convolutional layer performs a one-dimensional convolutional operation on the input historical target evidence data and transmits the output result to the activation layer
[0153] The third one-dimensional convolutional layer performs a one-dimensional convolutional operation on the input historical target evidence data and transmits the output result to the activation layer
[0154] Denote the convolution kernels of the first one-dimensional convolutional layer, the second one-dimensional convolutional layer, and the third one-dimensional convolutional layer as the first convolution kernel, the second convolution kernel, and the third convolution kernel respectively;
[0155] The selection of the first convolution kernel, the second convolution kernel, and the third convolution kernel has the following constraints:
[0156]
[0157] Among them, L 1 represents the length of the first convolution kernel; L 2 represents the length of the second convolution kernel; L 3 represents the length of the third convolution kernel; W 1 represents the width of the first convolution kernel; W 2 represents the width of the second convolution kernel; W 3 represents the width of the third convolution kernel; ⊙ represents the exclusive NOR logical operation, where the result is 0 if the symbols before and after are different, and 1 if they are the same; ∩ represents the AND logical operation; ∪ represents the OR logical operation; max{·} represents taking the maximum value within the parentheses; min{·} represents taking the minimum value within the parentheses; N L represents the total number of characters in each line of the text; N W represents the number of lines in the text; |·| represents taking the absolute value.
[0158] The activation layer uses a non-linear activation function (such as ReLU) to transform the output results of the first one-dimensional convolutional layer, the second one-dimensional convolutional layer, and the third one-dimensional convolutional layer respectively, so as to introduce non-linearity into the three convolution results, and then outputs the three convolution results with introduced non-linearity to the first pooling layer.
[0159] The first pooling layer adopts max pooling to reduce the size of the result of the first one-dimensional convolutional layer with introduced non-linearity, obtaining the first convolution pooling result reduce the size of the result of the second one-dimensional convolutional layer with introduced non-linearity, obtaining the second convolution pooling result reduce the size of the result of the third one-dimensional convolutional layer with introduced non-linearity, obtaining the third convolution pooling result and transmit it to the pooling result integration layer. This layer also has the function of retaining the most significant features while reducing the size.
[0160] The pooling result integration layer receives and integrates the first convolution pooling result, the second convolution pooling result, and the third convolution pooling result, and outputs the integrated result to the dropout layer;
[0161] The integration of the first convolution pooling result, the second convolution pooling result, and the third convolution pooling result is shown as follows:
[0162]
[0163] Among them, is the integration result obtained by integrating the first convolutional pooling result, the second convolutional pooling result, and the third convolutional pooling result, and the integration result is input into the dropout layer;
[0164] The dropout layer randomly sets the outputs of a part of the neurons in the integration result to 0, that is, "discards" these neurons, so as to reduce the co-adaptability between neurons, make the network more robust, and be able to better generalize to unseen data, thereby extracting the global feature information, and outputting the global feature information to the splicing layer.
[0165] In summary, the present invention obtains global feature information through the first one-dimensional convolutional layer, the second one-dimensional convolutional layer, the third one-dimensional convolutional layer, the activation layer, the first pooling layer, the pooling result integration layer, and the dropout layer.
[0166] The splicing layer fuses the data of the overall context relationship feature and the global feature information to obtain the fused text feature;
[0167] The second pooling layer performs max pooling on the fused text feature output by the splicing layer, while reducing the size of the output result of the splicing layer, and retaining the most significant features of the output result of the splicing layer; the output result of the second pooling layer is output to the fully connected layer.
[0168] The fully connected layer maps all the features of the input data to the output layer, so as to perform the classification task.
[0169] The output layer uses the Softmax activation function to generate the prediction result.
[0170] The loss function adopted by the text classification model is as follows:
[0171] The loss function L adopted by the text classification model of the present invention oss is constructed based on the fused text feature, as shown in the following formula:
[0172]
[0173] Among them, i represents the i-th sample; N is the number of samples; c represents the c-th label category, and C is the total number of label categories; y i,c is a binary index, which is 1 when the true label category of the i-th sample is category c, otherwise it is 0. For each sample, only one y i,c is 1, and the rest are all 0. ln represents the natural logarithm. is the prediction coefficient of the i-th sample belonging to the c-th label category; represents to find the natural logarithm of.
[0174] The prediction coefficient of the i-th sample belonging to the c-th label category is as follows:
[0175]
[0176] where max{·} represents taking the maximum value within the brackets; e is the natural base; j represents the j-th label category; Z i,c is the confidence of the model for the i-th sample belonging to category c; Z i,j is the confidence of the model for the i-th sample belonging to category j; A t is as follows:
[0177]
[0178] where y r is the true label category of the i-th sample; y f is the predicted label category of the i-th sample.
[0179] Those skilled in the art should know that during the prediction process of the deep learning model, there is a confidence for each category, and the one with the largest confidence value is the classification prediction result.
[0180] Iteration ends and the training of the text classification model is completed until the set number of training rounds is reached or the iteration end condition is met.
[0181] The iteration end condition is as follows:
[0182] L oss < 0.05N.
[0183] 2.2 Obtain the historical target evidence data to be classified, annotate the historical target evidence data to be classified to obtain an annotated data set; train the text classification model based on the annotated data set.
[0184] The obtaining of the historical target evidence data to be classified means the historical target evidence data that can be found.
[0185] Those skilled in the art should know that the annotation is classified according to different categories of the historical target evidence data, and those skilled in the art can select the number and categories of classification according to the actual situation for annotation; the annotation method proposed in the embodiments of the present invention is only a preferred embodiment and is not an inevitable limitation for implementing the present invention.
[0186] Preferably, the annotation includes: Label 1 to Label 3; where Label 1 represents legal document category; Label 2 represents abnormal power consumption category; Label 3 represents illegal dissemination category.
[0187] Training the text classification model based on the labeled dataset means using the historical target evidence data in the labeled dataset as the input of the text classification model, and using the labels in the labeled dataset as the output of the text classification model to train the text classification model.
[0188] 2.3 Real-time collect target evidence data and input it into the text classification model to obtain the classification result of the target evidence data. According to the classification result, divide the target evidence data into multiple categories, and use each category of target evidence data as a type of data to be mapped, thereby obtaining various types of data to be mapped.
[0189] Step 3: Perform data mapping on the obtained various types of data to be mapped, and then perform data conversion to convert the target evidence data into a preset evidence chain standard to obtain standard evidence chain data.
[0190] 3.1 The data mapping is to perform data mapping on each type of data to be mapped according to the preset evidence chain standard to obtain various types of data to be converted;
[0191] The attribute alignment and field matching can be completed by using tools such as Excel data matching assistant, Allegro automatic alignment tool, etc., and those skilled in the art can choose by themselves.
[0192] Furthermore, the preset evidence chain standard refers to a set of rules for standardizing the structure and content of evidence data to ensure that data from different sources can be processed and stored consistently. These standards usually include field names, data types, required fields, and optional fields, etc.
[0193] The present invention relates to the legal field. In the legal field, the evidence chain standard may require each evidence data to include fields such as "DocumentID", "Title", "Content", "Timestamp", etc. to ensure the integrity and consistency of the data; where "DocumentID" is the document number, "Title" is the title, "Content" is the content, "Timestamp" is the timestamp, etc.
[0194] The purpose of data mapping in this embodiment is to map the classified data to a predefined standard format of the evidence chain. In a preferred embodiment of the present invention, the classification of the present invention includes legal document category, abnormal power consumption category, and illegal dissemination category. First, according to the evidence chain standard, mapping rules for each data type are defined. The mapping rules are the principles for matching the content in the classified data to each field in the evidence chain standard. For example, "Case Number" in the legal document category is mapped to "DocumentID" in the standard format; in the abnormal power consumption category, the meter number + the number of abnormal times of the meter in the current year (that is, concatenating the meter number string with the number of abnormal times) is mapped to "DocumentID" in the standard format. For example, if the meter number is XX-XXXXXXXXXXX and the number of abnormal times of this meter up to the current power abnormality is 3, then the "DocumentID" of this abnormal power consumption category is set to XX-XXXXXXXXXXX3; in the illegal dissemination category, the first string obtained by concatenating the first pinyin characters of the first 6 characters of the illegal dissemination content + the dissemination year + the number of times this first string appears in that year is mapped to "DocumentID" in the standard format. According to the mapping rules, the original data fields are matched to the standard format fields. During the mapping process, necessary data cleaning is performed, such as removing redundant information and correcting format errors.
[0195] Through the above examples, those skilled in the art should know how to set mapping rules according to the actual situation, so the setting methods of other mapping rules will not be elaborated further.
[0196] The data conversion described in 3.2 is to perform data conversion on each type of data to be converted according to a preset storage format to obtain various types of standard evidence chain data.
[0197] Performing data conversion according to a preset storage format means converting the mapped data into a specific storage format, such as JSON, XML, or other structured formats, for easy storage and transmission in the blockchain. For example, the standard evidence chain data can be formatted as a JSON object, which contains all necessary fields and their corresponding values, such as {"DocumentID":"12345","Title":"Contract","Content":"Agreementdetails...","Timestamp":"2023-10-01T12:00:00Z"}, for subsequent encryption and blockchain operations.
[0198] The evidence chain conversion unit converts the mapped intermediate format data into the final format that conforms to the evidence chain standard according to the preset storage format, generates standard evidence chain data, and prepares for subsequent encryption and blockchain operations. The preset storage formats include JSON, XML, etc.
[0199] The purpose of data conversion in this embodiment is to convert the mapped data into the standard evidence chain format. The implementation steps are as follows:
[0200] Format conversion, converting the data into the structure required for the standard format; the required structure includes JSON, XML, etc.
[0201] Data verification, verifying the converted data to ensure that it meets the requirements of the standard format.
[0202] Exception handling, for the abnormal data that appears during the conversion process, record the log and perform processing, and the processing, for example, notify the administrator or re-collect the data.
[0203] After the above steps, this embodiment generates standard evidence chain data and prepares for encryption and blockchain operation.
[0204] This embodiment adopts the data classification in step 2, and the data mapping and data conversion in step 3 ensure the integrity and traceability of the data, and provide a consistent standardized data basis for subsequent processing.
[0205] Step 4: Perform encryption processing and blockchain operation on the standard evidence chain data to obtain an evidence blockchain containing the standard evidence chain data; and verify the integrity and authenticity of the evidence blockchain to obtain a verified blockchain;
[0206] Those skilled in the art should know how to encrypt the standard evidence chain data. The encryption method proposed by the present invention is only a preferred embodiment and is not an inevitable limitation of the present invention; preferably, the present invention can adopt symmetric encryption or asymmetric encryption to encrypt the standard evidence chain data.
[0207] Each piece of the standard evidence chain data is stored in the form of a block, and the evidence blockchain includes the standard evidence chain data, a storage timestamp, and the hash value of the previous block;
[0208] Those skilled in the art should know that both encryption and blockchain operation are existing technologies, and the present invention will not elaborate.
[0209] The method for verifying the integrity and authenticity of the blockchain to obtain a verified blockchain is as follows:
[0210] Since each block in the blockchain contains the hash value of the previous block, which forms an immutable chain, the verification process first calculates the hash value of the current block and compares it with the hash value stored in the block. If the two are consistent, it means that the block has not been tampered with, ensuring the integrity of the data.
[0211] The blockchain network uses a consensus mechanism to ensure that all nodes agree on the validity of the block. During the verification process, the nodes check whether the transactions in the newly generated block comply with the consensus rules, ensuring that only legitimate transactions are added to the blockchain, thereby maintaining the authenticity of the data.
[0212] Each transaction is digitally signed by the sender with his / her private key before submission. The verification module uses the sender's public key to verify the validity of the signature to ensure that the transaction is indeed initiated by the legitimate sender. This process prevents forgery and replay attacks and further enhances the authenticity of the data.
[0213] By auditing all transaction records on the blockchain, the source and change history of the data can be tracked. The verification module will regularly check the history of the blockchain to ensure that all transactions comply with the predetermined standards and rules. This transparency and traceability allows any abnormal activities to be discovered in a timely manner, thereby maintaining the overall credibility of the blockchain.
[0214] Step 5: Use multi-signature and distributed storage technology to protect the secure storage of data in the evidence blockchain and complete the solidification of the electronic data evidence chain for various types of data.
[0215] The present application also discloses an electronic data evidence chain solidification method based on blockchain technology and an electronic data evidence chain solidification system based on blockchain technology, including a data acquisition module, a data conversion module, a data encryption module and a storage module:
[0216] A data collection module, used to collect real-time target evidence data from multiple data sources; the target evidence data is in a text format;
[0217] A data conversion module is used to classify, map and convert real-time target evidence data to obtain standard evidence chain data;
[0218] A data encryption module encrypts and uploads the standard evidence chain data to obtain an evidence blockchain containing the standard evidence chain data;
[0219] The storage module uses multiple signatures and distributed storage technology to protect the secure storage of data in the evidence blockchain and complete the solidification of the electronic data evidence chain.
[0220] Example 2
[0221] like Figure 2 As shown, the electronic data evidence chain solidification system provided by the present invention includes a data acquisition module, a data conversion module, a data encryption module, a data verification module, a smart contract module and a storage module:
[0222] A data collection module for collecting target evidence data from various data sources; the format of the target evidence data is text format;
[0223] Preferably, the data sources include file storage systems, databases, emails, social media, sensors and Internet of Things devices, web pages, user inputs, audio, legal documents, logs, etc.
[0224] Specifically, the data sources in this embodiment are as follows:
[0225] 1. File storage system, extracting text data from documents, reports, contracts, agreements and other files in local or cloud storage. Examples: PDF files, Word documents, spreadsheets, etc.
[0226] 2. Database, extracting stored text data from relational or non-relational databases. Examples: text records in SQL databases and NoSQL databases.
[0227] 3. Email, extracting email content, attachments and metadata from email systems.
[0228] 4. Social media, extracting user-generated content from social media platforms, including posts, comments and messages.
[0229] 5. Sensors and Internet of Things devices, generating text data through Internet of Things devices, usually used to record events or states. Examples: log information of smart home devices, industrial sensors, etc.
[0230] 6. Web scraping, extracting text content from web pages through web crawler technology. Examples: public information on news websites, forums, blogs, etc.
[0231] 7. User input, directly inputting text data through user interfaces (such as forms, applications). Examples: online surveys, feedback forms, registration information, etc.
[0232] 8. Audio-to-text conversion, converting audio files into text data through speech recognition technology. Examples: meeting records, phone recordings, interview contents, etc.
[0233] 9. Legal documents, extracting text data from legal documents such as courts and arbitration institutions. Examples: judgments, pleadings, mediation agreements, etc.
[0234] 10. Log files, extracting text data from log files generated by systems or applications. Examples: server logs, application logs, audit logs, etc.
[0235] Through these diverse data sources, the system in this embodiment can comprehensively collect and integrate different types of text evidence data, providing rich basic information for subsequent data conversion, encryption, and storage, thereby enhancing the reliability and effectiveness of the electronic data evidence chain.
[0236] A data conversion module, configured to perform data classification, data mapping, and data conversion on target evidence data, so as to convert the target evidence data into a preset evidence chain standard to obtain standard evidence chain data;
[0237] Preferably, the data conversion module includes a data classification unit, a data mapping unit, and an evidence chain conversion unit.
[0238] The data conversion module realizes the standardization of target evidence data by combining text classification, data mapping, and format conversion technologies. Specifically, first, the data classification unit uses a preset text classification model to perform context analysis and classification on the target evidence data, and assigns the data to different categories. The preset text classification model is preferably a machine learning or natural language processing model.
[0239] The data classification unit is configured to classify the target evidence data according to the context relationship by using a preset text classification model to obtain various types of data to be mapped;
[0240] Preferably, the data classification unit includes a historical data acquisition subunit, a context extraction subunit, a global feature extraction subunit, a data fusion subunit, a function optimization subunit, and a classification subunit.
[0241] The historical data acquisition subunit, which is the input layer in the text classification model, is configured to acquire historical target evidence data to be classified and transmit the acquired data to the context extraction subunit and the global feature extraction subunit;
[0242] The context extraction subunit is configured to receive the historical target evidence data to be classified transmitted by the initial feature extraction component to extract context relationship features;
[0243] Preferably, the context extraction subunit includes an initial feature extraction component, a gated recurrent component, and a feature acquisition component:
[0244] Among them, the initial feature extraction component, which is the pre-trained language model layer in the text classification model. It is configured to receive the historical target evidence data to be classified transmitted by the initial feature extraction component and extract the initial feature information of the historical target evidence data to be classified, that is, perform a preliminary feature extraction step on the historical target evidence data to be classified;
[0245] Those skilled in the art should know that in practical applications, the pre-trained language model can adopt pre-trained XLNet model, pre-trained BERT model, pre-trained ELMo model, etc.; those skilled in the art can make selections according to actual situations; the pre-trained language model provided by the embodiments of the present invention is only a preferred embodiment and is not an inevitable limitation for implementing the present invention.
[0246] Preferably, the present invention selects a pre-trained XLNet model as the pre-trained language model for performing preliminary feature extraction on historical target evidence data.
[0247] The XLNet model is a pre-trained language model based on the combination of autoregression and autoencoding. Its core part is the Transformer-XL architecture, which solves the defect of the BERT model, that is, the Bidirectional Encoder Representations from Transformers model, in terms of the upper and lower context input length constraints, so that the pre-trained model can learn more distant upper and lower context information.
[0248] A gated recurrent component for inputting the initial feature information passed in by the initial feature extraction component into a forward gated recurrent unit and a backward gated recurrent unit;
[0249] Among them, the forward gated recurrent unit obtains the forward dependence relationship at time step t according to the input at time step t - 1, that is, the output of the forward gated recurrent unit at time t As shown in the following formula:
[0250]
[0251] Among them, GRU z (·) represents inputting the content in the brackets into the forward gated recurrent unit; H t ={H 1 ,H 2 ,…H l} represents the initial feature information input at time t; represents the output of the forward gated recurrent unit at time t; is the output of the forward gated recurrent unit at time t - 1; t ∈ [1, l], l is the total number of time steps; it is set that is 1.
[0252] Output the forward dependence relationships of each time step to the context relationship feature acquisition layer.
[0253] The backward gated recurrent unit obtains the backward dependence relationship at time step t according to the input at time step t - 1, that is, the output of the backward gated recurrent unit at time t As shown in the following formula:
[0254]
[0255] Among them, GRU f (·) represents inputting the content in the parentheses into the reverse gated recurrent unit; represents the output of the reverse gated recurrent unit at time t; is the output of the forward gated recurrent unit at time t - 1; It is set that is 0
[0256] Output the backward dependencies of each time step to the context relationship feature acquisition layer.
[0257] The feature acquisition component, that is, the context relationship feature acquisition layer in the text classification model, is used to splice the outputs of the forward gated recurrent unit and the reverse gated recurrent unit to obtain context relationship features;
[0258] Preferably, the calculation formula of the context relationship feature is:
[0259]
[0260] where h t represents the context relationship feature of the forward and reverse gated recurrent units at time t; G represents the overall context relationship feature; h 1 represents the context relationship feature of the forward and reverse gated recurrent units at the first time step; h 2 represents the context relationship feature of the forward and reverse gated recurrent units at the second time step; h l The context relationship feature of the forward and reverse gated recurrent units at the l-th time step. Among them represents the exclusive OR logical operation, where the result is 1 if the symbols before and after are different, and 0 if they are the same.
[0261] Preferably, in this embodiment, by splicing the outputs of the forward gated recurrent unit and the reverse gated recurrent unit, the text context relationship can be learned more fully, and the context information can be obtained.
[0262] The global feature extraction subunit is used to receive the historical target evidence data to be classified passed in by the initial feature extraction component to extract global feature information;
[0263] The global feature extraction subunit includes a first one-dimensional convolutional layer, a second one-dimensional convolutional layer, a third one-dimensional convolutional layer, an activation layer, a first pooling layer, a pooling result integration layer, and a dropout layer;
[0264] Among them, the first one-dimensional convolutional layer, the second one-dimensional convolutional layer, and the third one-dimensional convolutional layer all receive the historical target evidence data to be classified passed in by the initial feature extraction component and transmit the output results to the activation layer.
[0265] Denote the convolution kernels of the first, second, and third one-dimensional convolutional layers as the first convolution kernel, the second convolution kernel, and the third convolution kernel respectively;
[0266] The selection of the first convolution kernel, the second convolution kernel, and the third convolution kernel has the following constraints:
[0267]
[0268] where L 1 represents the length of the first convolution kernel; L 2 represents the length of the second convolution kernel; L 3 represents the length of the third convolution kernel; W 1 represents the width of the first convolution kernel; W 2 represents the width of the second convolution kernel; W 3 represents the width of the third convolution kernel; ⊙ represents the exclusive NOR logic operation, where the result is 0 when the symbols before and after are different, and 1 when they are the same; ∩ represents the AND logic operation; ∪ represents the OR logic operation; max{·} represents taking the maximum value within the parentheses; min{·} represents taking the minimum value within the parentheses; N L represents the total number of characters in each line of the text; N W represents the number of lines in the text; |·| represents taking the absolute value.
[0269] The activation layer uses a non-linear activation function (such as ReLU) to transform the output results of the first one-dimensional convolutional layer, the second one-dimensional convolutional layer, and the third one-dimensional convolutional layer respectively, so as to introduce non-linearity into the three convolution results, and then outputs the three convolution results with introduced non-linearity to the first pooling layer.
[0270] The first pooling layer adopts max pooling to reduce the size of the result of the first one-dimensional convolutional layer with introduced non-linearity, obtaining the first convolution pooling result reduce the size of the result of the second one-dimensional convolutional layer with introduced non-linearity, obtaining the second convolution pooling result reduce the size of the result of the third one-dimensional convolutional layer with introduced non-linearity, obtaining the third convolution pooling result and transmit to the pooling result integration layer. This layer also has the function of retaining the most significant features while reducing the size.
[0271] The pooling result integration layer receives and integrates the first convolution pooling result, the second convolution pooling result, and the third convolution pooling result, and outputs the integrated result to the dropout layer;
[0272] The integration of the first convolution pooling result, the second convolution pooling result, and the third convolution pooling result is shown as follows:
[0273]
[0274] Among them, is the integration result obtained by integrating the first convolution pooling result, the second convolution pooling result, and the third convolution pooling result, and the integration result is input into the dropout layer;
[0275] The dropout layer randomly sets the outputs of a part of the neurons in the integration result to 0, that is, "discards" these neurons, so as to reduce the co-adaptability between neurons, make the network more robust, and be able to better generalize to unseen data, thereby extracting the global feature information and outputting the global feature information to the splicing layer.
[0276] The data fusion subunit, that is, the splicing layer in the text classification model, is used to fuse the context relationship feature and the global feature information to obtain the fused text feature;
[0277] The data fusion subunit includes a splicing layer, a second pooling layer, and a fully connected layer;
[0278] Among them, the splicing layer fuses the data of the overall context relationship feature and the global feature information to obtain the fused text feature;
[0279] The second pooling layer performs max pooling on the fused text feature output by the splicing layer, while reducing the size of the output result of the splicing layer, and retains the most significant features of the output result of the splicing layer; the output result of the second pooling layer is output to the fully connected layer.
[0280] The fully connected layer maps all the features of the input data to the output layer, thereby performing the classification task.
[0281] The function optimization subunit is used to continuously optimize the loss function to obtain the text classification model;
[0282] The loss function L adopted by the text classification model of the present invention oss is constructed based on the fused text feature, as shown in the following formula:
[0283]
[0284] Among them, i represents the i-th sample; N is the number of samples; c represents the c-th label category, and C is the total number of label categories; y i,c is a binary indicator, which is 1 when the true label category of the i-th sample is category c, otherwise it is 0. For each sample, only one y i,c is 1, and the rest are all 0. ln represents the natural logarithm. is the prediction coefficient of the i-th sample belonging to the c-th label category.
[0285] The prediction coefficient of the i-th sample belonging to the c-th label category is as follows:
[0286]
[0287] where max{·} represents taking the maximum value within the brackets; e is the natural base; j represents the j-th label category; Z i,c is the confidence of the model that the i-th sample belongs to category c; Z i,j is the confidence of the model that the i-th sample belongs to category j; A t is as follows:
[0288]
[0289] where y r is the true label category of the i-th sample; y f is the predicted label category of the i-th sample.
[0290] Those skilled in the art should know that during the prediction process of the deep learning model, there is a confidence for each category, and the one with the largest confidence value is the classification prediction result;
[0291] The classification subunit, i.e., the output layer in the text classification model, uses the Softmax activation function to perform text classification, generates the text classification result, and then obtains various types of data to be mapped based on the text classification result.
[0292] In practical applications, the classification subunit outputs the classification result of the target evidence data in real time. According to the classification result, the target evidence data is divided into multiple categories, and each category of target evidence data is used as a type of data to be mapped, thereby obtaining various types of data to be mapped.
[0293] The data mapping unit is used to perform data mapping for each type of the data to be mapped according to the preset evidence chain standard to obtain various types of data to be converted;
[0294] The data mapping unit applies the predefined evidence chain standard to each type of data to be mapped for attribute alignment and field matching, and converts it into an intermediate format. The attribute alignment and field matching can be completed using tools such as Excel data matching assistant, Allegro automatic alignment tool, etc., and those skilled in the art can choose by themselves.
[0295] Furthermore, the predefined evidence chain standard refers to a set of rules for standardizing the structure and content of evidence data to ensure that data from different sources can be processed and stored consistently. These standards usually include field names, data types, required fields, and optional fields, etc.
[0296] The present invention relates to the field of law. Preferably, the evidence chain standard may require each evidence data to include fields such as "DocumentID", "Title", "Content", "Timestamp", etc., to ensure the integrity and consistency of the data; where "DocumentID" is the document number, "Title" is the title, "Content" is the content, "Timestamp" is the timestamp, etc.
[0297] Preferably, the purpose of data mapping in this embodiment is to map the classified data to a predefined evidence chain standard format. In the preferred embodiment of the present invention, the classification of the present invention includes legal document category, electricity consumption anomaly category, and illegal dissemination category. First, according to the evidence chain standard, define the mapping rules for each data type. The mapping rules are the principles for matching the content in the classified data to each field in the evidence chain standard. For example, map the "case number" in the legal document category to "DocumentID" in the standard format; for the electricity consumption anomaly category, map the meter number + the number of anomalies of this meter in the current year (that is, concatenate the meter number string with the number of anomalies) to "DocumentID" in the standard format. For example: if the meter number is XX-XXXXXXXXXXX and the number of anomalies of this meter up to this electricity consumption anomaly is 3 in the current year, then the "DocumentID" of this electricity consumption anomaly category is set to XX-XXXXXXXXXXX3; for the illegal dissemination category, map the first character string concatenated by the first characters of the pinyin of the first 6 characters of the illegal dissemination content + the dissemination year + the number of times this first character string appears in that year to "DocumentID" in the standard format. According to the mapping rules, match the original data fields to the standard format fields. During the mapping process, perform necessary data cleaning, such as removing redundant information and correcting format errors.
[0298] Through the above examples, those skilled in the art should know how to set the mapping rules according to the actual situation, so the setting methods of other mapping rules will not be elaborated here.
[0299] An evidence chain conversion unit is used to perform data conversion for each type of the to-be-converted unit according to a preset storage format to obtain various types of the standard evidence chain data.
[0300] Data conversion according to a preset storage format means converting the mapped data into a specific storage format, such as JSON, XML, or other structured formats, for easy storage and transmission in the blockchain. For example, standard evidence chain data can be formatted as a JSON object containing all necessary fields and their corresponding values, such as {"DocumentID":"12345","Title":"Contract","Content":"Agreementdetails...","Timestamp":"2023-10-01T12:00:00Z"}, to facilitate subsequent encryption and blockchain operations.
[0301] The evidence chain conversion unit converts the mapped intermediate format data into the final format that complies with the evidence chain standard according to the preset storage format, generates standard evidence chain data, and prepares for subsequent encryption and blockchain operations.
[0302] The purpose of data conversion in this embodiment is to convert the mapped data into the standard evidence chain format. The implementation steps are as follows:
[0303] Format conversion, converting the data into the structure required for the standard format; the required structure includes JSON, XML, etc.
[0304] Data verification, verifying the converted data to ensure that it meets the requirements of the standard format.
[0305] Exception handling, for abnormal data that appears during the conversion process, recording logs and performing processing, such as notifying the administrator or re-collecting data.
[0306] After the above steps, this embodiment generates standard evidence chain data and prepares for encryption and blockchain operations.
[0307] This embodiment uses the data classification unit, data mapping unit, and evidence chain conversion unit in the data conversion module to ensure the integrity and traceability of the data and provide a consistent standardized data basis for subsequent processing.
[0308] The data encryption module is used to encrypt and perform blockchain operations on the standard evidence chain data to obtain a blockchain including the standard evidence chain data; each standard evidence chain data is stored in the form of a block, and the blockchain includes standard evidence chain data, a storage timestamp, and the hash value of the previous block;
[0309] The data verification module is used to verify the integrity and authenticity of the blockchain including the standard evidence chain data to obtain a verified blockchain;
[0310] Specifically, each block in the blockchain contains the hash value of the previous block, which forms an immutable chain. The verification process first calculates the hash value of the current block and compares it with the hash value stored in the block. If the two are consistent, it means that the block has not been tampered with, ensuring the integrity of the data.
[0311] The blockchain network uses a consensus mechanism to ensure that all nodes agree on the validity of the block. During the verification process, the nodes check whether the transactions in the newly generated block comply with the consensus rules, ensuring that only legitimate transactions are added to the blockchain, thereby maintaining the authenticity of the data.
[0312] Each transaction is digitally signed by the sender with his / her private key before submission. The verification module uses the sender's public key to verify the validity of the signature to ensure that the transaction is indeed initiated by the legitimate sender. This process prevents forgery and replay attacks and further enhances the authenticity of the data.
[0313] By auditing all transaction records on the blockchain, the source and change history of the data can be traced. The verification module regularly checks the history of the blockchain to ensure that all transactions comply with the predetermined standards and rules. This transparency and traceability allows any abnormal activities to be discovered in a timely manner, thereby maintaining the overall credibility of the blockchain.
[0314] Smart contract module, used to build smart contracts containing data sharing rules and logic, and add the smart contracts to the verified blockchain;
[0315] The storage module uses multi-signature and distributed storage technology to protect the secure storage of data in the verified blockchain.
[0316] Preferably, the electronic data evidence chain solidification system based on blockchain technology provided by the present invention also includes:
[0317] The management module is used to verify and manage the user identity corresponding to the blockchain verified in the data verification module, and to monitor and record the query process.
[0318] Preferably, the management module includes:
[0319] User authentication unit, used to verify the user's identity through password and biometrics;
[0320] A transaction monitoring unit for real-time monitoring of query status and recording relevant logs; the relevant logs include user login and logout times, records of successful or failed user authentication, the authentication methods used (such as passwords, fingerprints, facial recognition, etc.), the timestamp of the query request, the user identity ID of the query, the specific content or parameters of the query, the return status of the query result (success or failure), and the generation time of the query result.
[0321] A data backup and recovery unit for assisting users in backing up and restoring the verified blockchain.
[0322] Preferably, the management module of this embodiment realizes the secure management and maintenance of the blockchain through the collaborative work of three units: user identity authentication, transaction monitoring, and data backup and recovery. The specific implementation process is as follows: The user identity authentication unit performs multi-factor identity authentication on users by combining passwords and biometric technologies to ensure the security of access permissions. The transaction monitoring unit tracks user query operations in real time and provides comprehensive monitoring and auditing functions for blockchain activities by recording relevant logs and status information. The data backup and recovery unit ensures the quick restoration of the integrity and availability of the blockchain in case of data loss or system failure by regularly backing up blockchain data and providing data recovery tools. The biometric technologies include fingerprints, facial recognition, etc. The status information includes query status, transaction status, user activity status, system performance status, security status, and audit information.
[0323] Preferably, the electronic data evidence chain solidification system based on blockchain technology provided by the present invention further includes:
[0324] A query signature module for encrypting the private key of the verified blockchain in the data verification module and storing the encrypted private key on the user's device; the private key is used for query signature only on the user's device.
[0325] Specifically, this embodiment enhances the security of the evidence chain and user privacy protection by encrypting and storing the private key on the user device and performing query signature. The specific implementation process is as follows: First, after the blockchain private key is generated by the system, strong encryption algorithms such as the Advanced Encryption Standard AES or the asymmetric encryption algorithm RSA are used to encrypt the private key, and the encrypted private key is securely stored on the user device. When the user performs a query operation, the private key is only decrypted on the device and used to generate a query signature, ensuring that the private key is not transmitted over the network, thus reducing the risk of leakage. This design effectively prevents unauthorized access and use of the private key, ensures the uniqueness and non-repudiation of the user's blockchain query, and at the same time enhances the overall security of the system and user trust. Among them, query signature refers to a digital signature generated by encrypting a query request with the user's private key when the user performs a query operation in the blockchain system.
[0326] Furthermore, based on the traditional Advanced Encryption Standard (AES) encryption method, this embodiment enhances its security by introducing dynamic key generation and multi-layer encryption strategies.
[0327] Specifically, first, the dynamic key generation mechanism generates a unique encryption key based on the user's device characteristics (such as device ID, geographical location, timestamp, etc.):
[0328] The dynamic key generation mechanism collects the characteristic information of the user's device, such as device ID, geographical location, and timestamp, and combines an encryption algorithm (such as SHA-256 or HMAC) for hashing operations to generate a unique encryption key. In specific implementation, first, these characteristic information are concatenated, and then the selected hash algorithm is used to process them to generate a key with a fixed length. Since the device characteristics may change during each encryption operation, the generated key will also change accordingly, thus enhancing the unpredictability and security of the key.
[0329] Since the characteristics used for dynamic key generation change during each encryption operation, the key is difficult to predict. Secondly, a multi-layer encryption strategy is adopted, that is, after the initial AES encryption, the ciphertext is encrypted one or more times with different algorithms (such as ChaCha20 or Twofish) to form a multi-layer encryption structure. In this way, even if an attacker cracks one layer of encryption, they still need to face the complex encryption algorithms of other layers. This embodiment significantly improves the security and anti-attack ability of data by increasing the dynamicity of the key and the encryption levels.
[0330] The beneficial effects of the present invention are as follows:
[0331] By combining the decentralized, immutable, and traceable characteristics of blockchain technology, the present invention realizes the solidification of the evidence chain of electronic data, ensuring the security and integrity of the data. Through the application of data encryption and smart contracts, the system not only protects data privacy but also enhances the flexibility and reliability of data sharing. The introduction of multi-signature and distributed storage technologies further improves the security of data storage, ensuring the evidentiary effect and credibility in applications in the fields of law, finance, etc.
[0332] The present disclosure may be a system, a method, and / or a computer program product. The computer program product may include a computer-readable storage medium having computer-readable program instructions thereon for causing a processor to implement various aspects of the present disclosure.
[0333] A computer-readable storage medium can be a tangible device that can retain and store instructions for use by an instruction execution device. A computer-readable storage medium may be, for example—but not limited to—an electrical storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any suitable combination of the foregoing. More specific examples (a non-exhaustive list) of the computer-readable storage medium include: a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), a static random access memory (SRAM), a portable compact disc read-only memory (CD-ROM), a digital versatile disc (DVD), a memory stick, a floppy disk, a mechanically encoded device such as a punched card or raised structures in grooves having instructions stored thereon, and any suitable combination of the foregoing. The computer-readable storage medium as used herein is not construed as being a transient signal per se, such as a radio wave or other freely propagating electromagnetic wave, an electromagnetic wave propagating through a waveguide or other transmission medium (e.g., an optical pulse through an optical fiber cable), or an electrical signal transmitted through a wire.
[0334] The computer-readable program instructions described herein can be downloaded from a computer-readable storage medium to various computing / processing devices, or downloaded to an external computer or external storage device through a network, such as the Internet, a local area network, a wide area network, and / or a wireless network. The network may include a copper transmission cable, an optical fiber transmission, a wireless transmission, a router, a firewall, a switch, a gateway computer, and / or an edge server. A network adapter card or network interface in each computing / processing device receives the computer-readable program instructions from the network and forwards the computer-readable program instructions for storage in a computer-readable storage medium in each computing / processing device.
[0335] The computer program instructions for performing the operations of the present disclosure may be assembly instructions, instruction set architecture (ISA) instructions, machine instructions, machine - related instructions, microcode, firmware instructions, state - setting data, or source code or object code written in any combination of one or more programming languages, including object - oriented programming languages such as Smalltalk, C++, etc., and conventional procedural programming languages such as the "C" language or similar programming languages. The computer - readable program instructions may be executed entirely on the user's computer, partially on the user's computer, executed as a stand - alone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., via an Internet service provider through the Internet). In some embodiments, by using the state information of the computer - readable program instructions to customize an electronic circuit, such as a programmable logic circuit, a field - programmable gate array (FPGA), or a programmable logic array (PLA), the electronic circuit can execute the computer - readable program instructions to implement various aspects of the present disclosure.
[0336] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit them. Although the present invention has been described in detail with reference to the above embodiments, those of ordinary skill in the art should understand that: modifications or equivalent replacements can still be made to the specific embodiments of the present invention. Any modification or equivalent replacement that does not depart from the spirit and scope of the present invention shall be covered by the protection scope of the claims of the present invention.
Claims
1. A method for solidifying an electronic data evidence chain based on blockchain technology, characterized in that: include: Step 1: Collect real-time target evidence data from multiple data sources; The format of the target evidence data is text format; Step 2: Classify the real-time target evidence data based on a text classification model to obtain various types of data to be mapped; the text classification model is constructed based on a pre-trained language model, a gated recurrent unit and a one-dimensional convolutional model; Step 3: Map various types of data to be mapped based on the preset evidence chain standard, and convert the mapped data based on the preset storage format to obtain standard evidence chain data; Step 4: Encrypt and upload the standard evidence chain data to obtain an evidence blockchain containing the standard evidence chain data; Step 5: Use multi-signature and distributed storage technology to protect the data security in the evidence blockchain and complete the solidification of the electronic data evidence chain.
2. The electronic data evidence chain solidification method according to claim 1 is characterized in that: In step 1, multiple data sources include file storage systems, databases, email, social media, sensors, web pages, user input, audio, legal documents, and logs; Target evidence data include: Text extracted from file storage systems; Text extracted from databases; Message content and attachments extracted from email systems; posts, comments, and messages extracted from social media platforms; Text collected by sensors; Text content obtained through web crawling; The text content entered by the user; The text obtained by audio conversion; Legal documents; Log text data generated by a system, application, or IoT device.
3. The electronic data evidence chain solidification method according to claim 1, characterized in that: In step 2, the following steps are included: 2.1 Build a text classification model based on the pre-trained language model, gated recurrent unit and one-dimensional convolutional model; 2.2 Collect historical target evidence data to be classified, annotate the historical target evidence data to be classified, and obtain annotated data sets; train the text classification model based on the annotated data sets; 2.3 Collect real-time target evidence data and input it into the text classification model to obtain the classification results of the real-time target evidence data and obtain various types of data to be mapped.
4. The electronic data evidence chain solidification method according to claim 1 or 3, characterized in that: The text classification model includes an input layer, a pre-trained language model layer, a forward gated recurrent unit, a reverse gated recurrent unit, a contextual relationship feature acquisition layer, a first one-dimensional convolutional layer, a second one-dimensional convolutional layer, a third one-dimensional convolutional layer, an activation layer, a first pooling layer, a pooling result integration layer, a random inactivation layer, a splicing layer, a second pooling layer, a fully connected layer, and an output layer.
5. The method for solidifying the electronic data evidence chain according to claim 4, characterized in that: The contextual relationship feature acquisition layer obtains the overall contextual relationship feature based on the output of the forward gated recurrent unit and the reverse gated recurrent unit; The calculation formula of the overall contextual relationship feature is: Among them, h t represents the contextual relationship features of the forward and reverse gated recurrent units at time t; G represents the overall contextual relationship features; h1 represents the contextual relationship features of the forward and reverse gated recurrent units at the first moment; h2 represents the contextual relationship features of the forward and reverse gated recurrent units at the second moment; h l The contextual relationship features of the forward and reverse gated recurrent units at the lth moment; l is the total number of time steps.
6. The electronic data evidence chain solidification method according to claim 4, characterized in that: The convolution kernels of the first one-dimensional convolution layer, the second one-dimensional convolution layer, and the third one-dimensional convolution layer are respectively recorded as the first convolution kernel, the second convolution kernel, and the third convolution kernel; The selection of the first, second, and third convolution kernels has the following constraints: Wherein, L1 represents the length of the first convolution kernel; L2 represents the length of the second convolution kernel; L3 represents the length of the third convolution kernel; W1 represents the width of the first convolution kernel; W2 represents the width of the second convolution kernel; W3 represents the width of the third convolution kernel; max{·} represents the maximum value in the brackets; min{·} represents the minimum value in the brackets; N L Indicates the total number of characters in each line of text; N W The number of lines in the text; |·| means taking the absolute value.
7. The method for solidifying the electronic data evidence chain according to claim 3, characterized in that: In step 2.2, the text classification model is trained based on the annotated dataset, and the loss function L is used. oss As shown below: Where i represents the i-th sample; N is the number of samples; c represents the c-th label category, and C is the total number of label categories obtained by annotation; y i,c It is a binary indicator, which is 1 when the true label category of the i-th sample is category c, otherwise it is 0; is the prediction coefficient that the i-th sample belongs to the c-th label category.
8. The method for solidifying the electronic data evidence chain according to claim 7, characterized in that: The prediction coefficient that the i-th sample belongs to the c-th label category As shown below: Among them, max{·} means taking the maximum value in the brackets; e is the natural base; j represents the jth label category; Z i,c is the model's confidence that the i-th sample belongs to category c; Z i,j is the model's confidence that the i-th sample belongs to category j; A t As shown below: Among them, y r is the true label category of the i-th sample; y f is the predicted label category of the i-th sample.
9. The electronic data evidence chain solidification method according to claim 1, characterized in that: In step 3, the preset evidence chain standard is: each evidence chain includes "DocumentID", "Title", "Content", and "Timestamp" fields; "DocumentID" is the file number, "Title" is the title, "Content" is the content, and "Timestamp" is the timestamp.
10. An electronic data evidence chain solidification system using the electronic data evidence chain solidification method according to any one of claims 1 to 9, characterized in that: Including data acquisition module, data conversion module, data encryption module and storage module: A data collection module, used to collect real-time target evidence data from multiple data sources; the target evidence data is in a text format; A data conversion module is used to classify, map and convert real-time target evidence data to obtain standard evidence chain data; A data encryption module encrypts and uploads the standard evidence chain data to obtain an evidence blockchain containing the standard evidence chain data; The storage module uses multiple signatures and distributed storage technology to protect the secure storage of data in the evidence blockchain and complete the solidification of the electronic data evidence chain.
11. A terminal comprising a processor and a storage medium; characterized in that: The storage medium is used to store instructions; The processor is used to operate according to the instructions to execute the steps of the electronic data evidence chain solidification method according to any one of claims 1-9.
12. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, the steps of the method for solidifying the electronic data evidence chain described in any one of claims 1 to 9 are implemented.
Citation Information
Patent Citations
Encryption evidence storage method and device based on block chain, equipment and storage medium
CN115048670A