Electronic archive online encryption circulation transmission method and system
By processing electronic archives in blocks and encrypting their transmission, and utilizing blockchain timestamps and topological relationship diagrams, the problems of low security and low efficiency of traditional electronic archive transmission are solved, achieving efficient and secure archive transmission.
Patent Information
- Application Number
- CN202511044683.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-29
- Publication Date
- 2025-09-23
AI Technical Summary
Traditional electronic file transmission is difficult to ensure security and block efficiency, especially when processing PB-level unstructured data, the classification error rate is high and the time consumption is doubled.
By dividing electronic files into blocks, the privacy information of each independent file block is judged, and only the file blocks containing privacy information are encrypted and transmitted. The receiving end completes the decryption and reconstruction of the file blocks, and uses blockchain timestamps, quantum random numbers to generate dynamic keys and topological relationship diagrams for security management.
It improves computer computing efficiency, reduces computer burden, ensures the security of file transmission, and ensures data integrity and consistency through intelligent error correction and reconstruction mechanisms.
Smart Images

Figure CN120692089A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of electronic file transmission, and in particular relates to an online encrypted flow transmission method and system for electronic files. Background Art
[0002] Electronic archives are electronic documents that have documentary, reference, and preservation value and are archived. Electronic documents are various information records in digital format that are generated, processed, transmitted, and stored by computers and other electronic devices in the process of state agencies, social organizations, or individuals performing their legal duties or handling affairs. Electronic files are composed of content, structure, and context. Electronic archives refer to a collection of general electronic image files that are stored on devices such as computer disks and are related to paper archives. They are usually stored in file units. As the number of electronic files increases exponentially, traditional electronic file transmission cannot guarantee the security of their sharing; and manual block segmentation methods cannot process PB-level unstructured data, resulting in a high classification error rate and doubling the time consumption. Summary of the Invention
[0003] The purpose of the present invention is to provide a method and system for online encrypted transfer and transmission of electronic archives. By dividing the electronic archive into blocks, determining the privacy information of each independent file block, and encrypting and transmitting only the file blocks containing privacy information, the receiving end simultaneously completes the decryption and reassembly of the file blocks. This solves the problems of low security in existing electronic archive transmission, low block efficiency, and encryption occupying computer resources.
[0004] To solve the above technical problems, the present invention is achieved through the following technical solutions: The present invention provides an online encrypted transfer method for electronic files, comprising the following steps: Step S1: Obtain the electronic file to be transmitted for identification, divide it into blocks according to the identification results, and inject the blockchain timestamp into each block as a separate header identifier; Step S2: Determine the privacy information of each independent file block, encrypt the file block containing the privacy information, generate a dynamic key based on the quantum random number, and append the block hash value as the integrity check code; Step S3: The server authenticates the client and the receiving end decrypts the ciphertext; Step S4: The receiving end decrypts and obtains document blocks, which are then spliced together according to the topological relationship diagram; Step S5: Use an intelligent error correction and reorganization mechanism to manage electronic files.
[0005] In step S1, the file type is identified using OCR technology, and the character count threshold is adjusted based on the identification result. The NLP model is used to detect the theme change points of the file to ensure that each file block contains a complete semantic unit. The file is segmented based on the theme change points. During segmentation, character overlap is set between adjacent file blocks to facilitate maintaining contextual coherence during reassembly. The NLP model detects the theme change points of the file, including: Semantic coherence analysis uses the BiLSTM-CRF sequence model to capture the semantic relevance between paragraphs. When the cosine similarity of the semantic vectors of adjacent text blocks is lower than a threshold, it is marked as a potential topic boundary. Identify text structure features: The TF-IDF algorithm verifies differences in file content, detects layout features such as title level changes and paragraph indentation, and identifies transition words in the file. This is combined with dependency analysis to confirm topic boundaries. Statistical mutation detection uses the CUSUM algorithm to monitor the cumulative deviation of word frequency distribution. Segmentation is triggered when the frequency of a specific domain term changes suddenly. KL divergence is calculated based on a dynamic sliding window to detect the probability of word distribution. For example, if a certain drug name appears frequently in the previous section of a medical record, but the number of occurrences decreases significantly in the next section, the file segmentation can be triggered at the obvious sudden change. A multimodal fusion measurement strategy combines OCR layout analysis results in scanned archives to strengthen segmentation decisions when physical separation and semantic mutations of text regions occur simultaneously.
[0006] As an optimal technical solution, after the archive is segmented according to the topic change points, the initial segmentation results can be further subdivided, and the segmentation positions can be optimized in combination with syntax tree analysis. The specific solution is as follows: first, syntax analysis is performed based on the initial segmentation to construct a syntax tree containing sentence structure, clause hierarchy and key punctuation; then, complex sentences and logical connectives are detected through depth-first traversal, and while maintaining semantic integrity, syntax nodes such as coordinating conjunctions and complete clause boundaries are preferentially selected as segmentation points; then, a dynamic recursive strategy is adopted to trigger secondary segmentation when it is detected that the nesting is too deep or the sentence complexity exceeds the standard, until the minimum semantic unit is met or the recursive depth limit is reached; finally, the syntax integrity and contextual coherence of each sub-block are verified to ensure that the segmentation results comply with the grammatical rules and retain the semantic logic of the original document; the entire process uses recursive segmentation guided by the syntax tree to ensure that the segmentation of complex documents maintains both structural rationality and semantic coherence.
[0007] As a preferred technical solution, in step S2, the specific process of determining the privacy information of each independent file block is as follows: Step S21, data matrix construction: converting independent file block information into a real number matrix ; Among them, the real matrix center, number of rows Indicates different record entries, number of rows Indicates different attributes of the archive; Step S22, seed gene generation: randomly select a profile attribute as the initial seed gene, and calculate the correlation with other attributes through mutual information. The calculation formula is as follows: ; Where, are two information elements in the real matrix The mutual information value of Represents two information elements in the real matrix The information entropy of Represents two information elements The joint information entropy of Step S23, biclustering set construction: Calculate the mutual information value between the remaining attributes and the current seed gene, select the attribute with the smallest mutual information value as the new seed, and form two biclustering sets of general information and private information; Step S24, fuzzy partitioning implementation: Calculate the membership probability of non-seed genes to the biclusters, and the calculation formula is as follows: ; Where, Indicates that non-seed genes belong to the bicluster set The probability value of represents the seed gene set, Represents the real matrix Rank Column element biclustering probability, The calculation formula is as follows: ; Where, represents the size of the condition set in the biclustering set; Step S25: Classification determination: setting threshold , the final classification is completed through the discriminant, which is expressed as follows: ; Where, Indicates the double clustering result of archival information, 1 indicates general information, 0 indicates private information, represents the biclustering threshold.
[0008] As a preferred technical solution, in step S2, the encryption strategy for the file block containing private information is as follows: Step YS01: Encode the attributes of the privacy file block; Step YS02: Use an asymmetric encryption algorithm to encrypt the private file block to generate ciphertext to ensure data confidentiality and integrity. The core of the RSA algorithm is to convert the original plaintext data into ciphertext data after calculation using a linear feedback shift register. Step YS03: The encrypted document is sent to the receiving end through the server; Step YS04: After the server verifies the client's identity by parsing the address data and port number of the receiving end, the receiving end receives the ciphertext and decrypts it.
[0009] As a preferred technical solution, in step YS04, when the receiving end requests access, the server uses the socket interface to listen and accept the connection request. The server performs identity authentication by parsing the address data and port number sent by the receiving end. The address data includes the IP address, etc.; if the identity authentication is passed, data transmission is performed; when the point-to-point communication transmission task is completed, the receiving end and the server terminate the transmission control protocol connection by closing the program; data is sent in the form of packets, each packet including the data itself and necessary control information, such as the source address, destination address, port number, etc. The server must also be able to handle multiple concurrent connections. This is typically achieved through technologies such as multithreading or asynchronous I / O to ensure that each client receives a timely response. Based on this communication scenario, a multi-value mapping function specifically for encrypted transmission is designed. This function transmits files through a complex spatial mapping mechanism, thereby achieving decryption of terminal data. The specific calculation formula is as follows: ; Where, Indicates the private information of the archive transmitted to the information receiving end. represents the mapping space function, Indicates the mapping link, Indicates a single-phase mapping link in a communication network; Once the mapped data is extracted, it's immediately sent to the file server for further processing using its built-in decryption tools. If manual operation is required, the user must ensure that the passwords used for the encryption and decryption programs on the file server match the passwords currently in use to ensure smooth operation.
[0010] As a preferred technical solution, in step S4, the topological relationship graph is constructed based on the electronic archive blocks generated in the intelligent blocking stage, and a directed acyclic graph is constructed to represent the topological relationship between archive blocks; the nodes represent the block data, and the edges include the sequence relationship between archive blocks and the double hash check code generated in the dynamic encryption stage; at the same time, a topological coordinate attribute is added to each archive block to record its logical position in the global archive, and an anti-collision hash function is used to calculate the correlation between adjacent archive blocks as the edge weight value.
[0011] As a preferred technical solution, the specific process of dynamic verification of the double hash check code is as follows: Step S41: The user submits the zero-knowledge proof and the list of block IDs to be accessed, and the verifier confirms the validity of the permission; Step S42: The system loads the block topology graph and checks whether the requested block set constitutes a legal connected subgraph; Step S43: Decrypt the blocks in topological order, and verify the hash reference relationship between each decrypted block and the decrypted block; Step S44: Finally verify the topological hash tree root of the entire document and match the blockchain evidence value to complete end-to-end verification.
[0012] As a preferred technical solution, in step S5, the intelligent error correction and reconstruction mechanism automatically adjusts the fault tolerance threshold based on the block characteristics, uses the sliding window algorithm to detect data continuity in real time, and triggers the reconstruction process when the missing area exceeds the preset threshold; a new hash value is generated for the reorganized block through the blockchain verification layer, and is stored on the chain together with the original block hash, and a smart contract is used to verify whether the reconstruction logic complies with the preset rules.
[0013] The present invention is an online encrypted transfer and transmission system for electronic files, comprising a generating end, a server and a receiving end; The generating end includes an electronic archive recognition module, an intelligent block segmentation module, a privacy archive block determination module, a dynamic encryption module, and a secure transmission module; the electronic archive recognition module is used to convert paper archives into electronic archives and recognize them using OCR technology; the intelligent block segmentation module is used to segment electronic archives based on subject change points; the privacy archive block determination module is used to perform privacy determination on the segmented archive blocks; the dynamic encryption module is used to encrypt the privacy archive blocks; and the secure output module is used to transmit the encrypted privacy blocks and common archive blocks to the server; The server includes an identity authentication module and a file block management module; the identity authentication module is used to; the file management module is used to parse the address data and port number of the receiving end to verify the client's identity; the file management module is used to manage electronic files and file blocks; The receiving end includes an archive block decryption module, an archive block splicing module, and a smart contract detection module; the archive block decryption module is used to decrypt the received encrypted archive blocks; the archive block reassembly module is used to reassemble the decrypted archive blocks to complete the electronic archive; and the smart contract detection module is used to verify whether the logic of the reassembled electronic archive complies with preset rules.
[0014] The present invention is an online encrypted transfer method for electronic files, comprising a generating end, a server and a receiving end; The generating end includes an electronic archive recognition module, an intelligent block segmentation module, a privacy archive block determination module, a dynamic encryption module, and a secure transmission module; the electronic archive recognition module is used to convert paper archives into electronic archives and recognize them using OCR technology; the intelligent block segmentation module is used to segment electronic archives based on subject change points; the privacy archive block determination module is used to perform privacy determination on the segmented archive blocks; the dynamic encryption module is used to encrypt the privacy archive blocks; and the secure output module is used to transmit the encrypted privacy blocks and common archive blocks to the server; The server includes an identity authentication module and a file block management module; the identity authentication module is used to; the file management module is used to parse the address data and port number of the receiving end to verify the client's identity; the file management module is used to manage electronic files and file blocks; The receiving end includes an archive block decryption module, an archive block splicing module, and a smart contract detection module; the archive block decryption module is used to decrypt the received encrypted archive blocks; the archive block reassembly module is used to reassemble the decrypted archive blocks to complete the electronic archive; and the smart contract detection module is used to verify whether the logic of the reassembled electronic archive complies with preset rules.
[0015] The present invention has the following beneficial effects: (1) The present invention divides electronic files into blocks, determines the privacy information of each independent file block, and encrypts and transmits only the file blocks containing privacy information, thereby improving computer computing efficiency and reducing computer burden. At the same time, the receiving end completes the decryption and reorganization of the file blocks, thereby ensuring the security of file transmission.
[0016] (2) The present invention constructs a document dependency tree through syntactic analysis, and then uses the LDA topic model to generate paragraph topic distribution. Finally, a change point detection algorithm (such as PELT) is used to locate the optimal segmentation point, quickly determine the file segmentation position, make the file segmentation more reasonable, improve the file segmentation efficiency, and set 10% to 15% character overlap between adjacent file blocks to maintain context coherence.
[0017] (3) The present invention uses a biclustering algorithm to extract private information from archive blocks. It randomly selects an archive attribute as the initial seed gene and iterates the genes of subsequent seeds. It quantifies gene correlation based on mutual information to improve the seed selection accuracy. It introduces the concept of simulated fuzzy partitioning to process moderately overlapping data. Finally, it outputs a collection of private archive block information marked as 0, thereby improving the extraction accuracy of private archive blocks.
[0018] (4) The present invention uses a decentralized identifier to generate a user's unique identity, binds the public key and verifiable statement to the blockchain, constructs a directed acyclic graph to represent the topological relationship between blocks based on the electronic file blocks generated in the intelligent block segmentation stage, reduces the exposure of identity data, and uses the topological relationship graph to optimize the decryption order and reduce computational overhead.
[0019] Of course, any product implementing the present invention does not necessarily need to achieve all of the advantages described above at the same time. BRIEF DESCRIPTION OF THE DRAWINGS
[0020] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following briefly introduces the drawings required for describing the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.
[0021] Figure 1 This is a flow chart of an online encrypted transfer method for electronic files according to the present invention; Figure 2 This is a schematic diagram of the structure of an online encrypted transfer system for electronic files according to the present invention; Figure 3 FIG. 4 is a schematic diagram of file block encryption according to an embodiment of the present invention. DETAILED DESCRIPTION
[0022] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making any creative efforts shall fall within the scope of protection of the present invention.
[0023] In addition, the technical features involved in the various embodiments of the present invention described below can be combined with each other as long as they do not conflict with each other.
[0024] In order to make the purpose, technical solutions and advantages of this application more clear, the following Figure 1-2It should be understood that the specific embodiments described herein are only used to explain the present application and are not intended to limit the present application.
[0025] Example 1 See also Figure 1 As shown, the present invention is a method for online encrypted transfer of electronic files, comprising the following steps: Step S1: Obtain the electronic file to be transmitted for identification, divide it into blocks according to the identification results, and inject the blockchain timestamp into each block as a separate header identifier; Step S2: Determine the privacy information of each independent file block, encrypt the file block containing the privacy information, generate a dynamic key based on the quantum random number, and append the block hash value as the integrity check code; Step S3: The server authenticates the client and the receiving end decrypts the ciphertext; Step S4: The receiving end decrypts and obtains document blocks, which are then spliced together according to the topological relationship diagram; Step S5: Use an intelligent error correction and reorganization mechanism to manage electronic files.
[0026] In step S1, OCR technology is used to identify the file type. The character count threshold is adjusted based on the recognition results (such as contracts, reports, and logs). The NLP model is used to detect the file's topic change points to ensure that each file block contains a complete semantic unit. The file is segmented based on the topic change points. During segmentation, a character overlap of 10% to 15% is set between adjacent file blocks to facilitate the maintenance of contextual coherence during reassembly. The NLP model detects the following topic change points in the file: Semantic coherence analysis uses the BiLSTM-CRF sequence model to capture the semantic relevance between paragraphs. When the cosine similarity of the semantic vectors of adjacent text blocks is lower than a threshold (e.g., 0.6), it is marked as a potential topic boundary. Discourse structure feature recognition: The TF-IDF algorithm verifies differences in file content, detects layout features such as title level changes and paragraph indentation, and identifies transition words (such as "however," "in summary," and "but") in the file. It also uses dependency parsing to confirm topic boundaries. Statistical mutation detection uses the CUSUM algorithm to monitor the cumulative deviation of word frequency distribution. Segmentation is triggered when the frequency of a specific domain term changes suddenly. KL divergence is calculated based on a dynamic sliding window to detect the probability of word distribution. For example, if a certain drug name appears frequently in the previous section of a medical record, but the number of occurrences decreases significantly in the next section, the file segmentation can be triggered at the obvious sudden change. A multimodal fusion measurement strategy combines the results of OCR layout analysis in scanned archives to strengthen segmentation decisions when physical spacing of text areas (such as greater than 1.5 times the line spacing) and semantic mutations occur simultaneously.
[0027] After the archive is segmented according to thematic change points, the initial segmentation results can be further subdivided, and the segmentation positions can be optimized in combination with syntax tree analysis. The specific plan is as follows: first, syntax analysis is performed based on the initial segmentation to construct a syntax tree containing sentence structure, clause hierarchy and key punctuation; then, complex sentences and logical connectives are detected through depth-first traversal, and while maintaining semantic integrity, syntax nodes such as coordinating conjunctions and complete clause boundaries are given priority as segmentation points; then, a dynamic recursive strategy is adopted to trigger secondary segmentation when it is detected that the nesting is too deep or the sentence complexity exceeds the standard, until the minimum semantic unit is met or the recursive depth limit is reached; finally, the grammatical integrity and contextual coherence of each sub-block are verified to ensure that the segmentation results conform to the grammatical rules and retain the semantic logic of the original document; the entire process uses recursive segmentation guided by the syntax tree to ensure that the segmentation of complex documents maintains both structural rationality and semantic coherence.
[0028] In step S2, the specific process of determining the privacy information of each independent file block is as follows: Step S21, data matrix construction: converting independent file block information into a real number matrix ; Among them, the real matrix center, number of rows Indicates different record entries, number of rows Indicates different attributes of the archive; Step S22, seed gene generation: randomly select a profile attribute as the initial seed gene, and calculate the correlation with other attributes through mutual information. The calculation formula is as follows: ; Where, are two information elements in the real matrix The mutual information value of Represents two information elements in the real matrix The information entropy of Represents two information elements The joint information entropy of Calculate the mutual information value between each remaining gene in the data set and the current seed gene to quantify the similarity or correlation between genes. Select the gene with the least similarity to the current seed gene, that is, the gene with the smallest mutual information, as the starting point of the new bicluster. Iterate and select subsequent bicluster seed genes in accordance with the above method, and continuously expand the seed gene set and bicluster set. Step S23, biclustering set construction: Calculate the mutual information value between the remaining attributes and the current seed gene. Mutual information can be used to measure the degree of mutual dependence between two random variables. The mutual information of two discrete random variables is equal to the relative entropy between the product of the joint probability distribution function of the two variables and the marginal probability distribution function of the two variables. Select the attribute with the smallest mutual information value as the new seed to form two biclustering sets of common information and private information. Step S24, fuzzy partitioning implementation: Calculate the membership probability of non-seed genes to the biclusters, and the calculation formula is as follows: ; Where, Indicates that non-seed genes belong to the bicluster set The probability value of represents the seed gene set, Represents the real matrix Rank Column element biclustering probability, The calculation formula is as follows: ; Where, represents the size of the condition set in the biclustering set; Step S25: Classification determination: setting threshold , the final classification is completed through the discriminant, which is expressed as follows: ; Where, Indicates the double clustering result of archival information, 1 indicates general information, 0 indicates private information, represents the biclustering threshold.
[0029] In step S2, the encryption strategy for the file block containing private information is as follows: Step YS01: Encode the attributes of the privacy file block; Step YS02: Use an asymmetric encryption algorithm to encrypt the private file block to generate ciphertext to ensure data confidentiality and integrity. The core of the RSA algorithm is to convert the original plaintext data into ciphertext data after calculation using a linear feedback shift register. Step YS03: The encrypted document is sent to the receiving end through the server; Step YS04: After the server verifies the client's identity by parsing the address data and port number of the receiving end, the receiving end receives the ciphertext and decrypts it.
[0030] In step YS04, when the receiving end requests access, the server uses the socket interface to listen and accept the connection request. The server performs identity authentication by parsing the address data and port number sent by the receiving end. The address data includes the IP address, etc.; if the identity authentication is passed, data transmission will be carried out; when the point-to-point communication transmission task is completed, the receiving end and the server terminate the transmission control protocol connection by closing the program; data is sent in the form of packets, each packet includes the data itself and necessary control information, such as the source address, destination address, port number, etc. The server must also be able to handle multiple concurrent connections. This is typically achieved through technologies such as multithreading or asynchronous I / O to ensure that each client receives a timely response. Based on this communication scenario, a multi-value mapping function specifically for encrypted transmission is designed. This function transmits files through a complex spatial mapping mechanism, thereby achieving decryption of terminal data. The specific calculation formula is as follows: ; Where, Indicates the private information of the archive transmitted to the information receiving end. represents the mapping space function, Indicates the mapping link, Indicates a single-phase mapping link in a communication network; Once the mapped data is extracted, it's immediately sent to the file server for further processing using its built-in decryption tools. If manual operation is required, the user must ensure that the passwords used for the encryption and decryption programs on the file server match the passwords currently in use to ensure smooth operation.
[0031] In step S4, a topological relationship graph is constructed based on the electronic archive blocks generated in the intelligent blocking stage, and a directed acyclic graph is constructed to represent the topological relationship between archive blocks; the nodes represent the block data, and the edges include the order relationship between archive blocks and the double hash check code generated in the dynamic encryption stage; at the same time, a topological coordinate attribute is added to each archive block to record its logical position in the global archive, and an anti-collision hash function is used to calculate the correlation between adjacent archive blocks as the edge weight value.
[0032] The specific process of dynamic verification of double hash check codes is as follows: Step S41: The user submits the zero-knowledge proof and the list of block IDs to be accessed, and the verifier confirms the validity of the permission; Step S42: The system loads the block topology graph and checks whether the requested block set constitutes a legal connected subgraph; Step S43: Decrypt the blocks in topological order, and verify the hash reference relationship between each decrypted block and the decrypted block; Step S44: Finally verify the topological hash tree root of the entire document and match the blockchain evidence value to complete end-to-end verification.
[0033] In step S5, the intelligent error correction and reconstruction mechanism automatically adjusts the fault tolerance threshold based on the block characteristics, uses the sliding window algorithm to detect data continuity in real time, and triggers the reconstruction process when the missing area exceeds the preset threshold; generates a new hash value for the reorganized block through the blockchain verification layer, and stores it on the chain together with the original block hash, and uses the smart contract to verify whether the reconstruction logic complies with the preset rules.
[0034] Example 2 See also Figure 2-3 As shown, the present invention is an online encrypted transfer and transmission system for electronic files, which can be used to execute the method of Example 1 of the present invention, including: a generating end, a server and a receiving end; The generating end includes an electronic archive recognition module, an intelligent block segmentation module, a privacy archive block determination module, a dynamic encryption module, and a secure transmission module; the electronic archive recognition module is used to convert paper archives into electronic archives and identify them using OCR technology; the intelligent block segmentation module is used to segment electronic archives based on the subject change points; the privacy archive block determination module is used to perform privacy determination on the segmented archive blocks; the dynamic encryption module is used to encrypt the privacy archive blocks; and the secure output module is used to transmit the encrypted privacy blocks and common archive blocks to the server; The server includes an identity authentication module and an archive block management module; the identity authentication module is used to; the archive management module is used to parse the address data and port number of the receiving end to verify the client's identity; the archive management module is used to manage electronic archives and archive blocks; The receiving end includes an archive block decryption module, an archive block splicing module and a smart contract detection module; the archive block decryption module is used to decrypt the received encrypted archive blocks; the archive block reassembly module is used to reassemble the decrypted archive blocks to synthesize the electronic archive; the smart contract detection module is used to verify whether the logic of the reassembled electronic archive complies with the preset rules.
[0035] It is worth noting that in the above system embodiment, the various units included are only divided according to functional logic, but are not limited to the above division, as long as the corresponding functions can be achieved; in addition, the specific names of the functional units are only for the convenience of distinguishing each other, and are not used to limit the scope of protection of the present invention.
[0036] In addition, those skilled in the art will appreciate that all or part of the steps in the above-mentioned embodiments can be accomplished by instructing related hardware through a program, and the corresponding program can be stored in a computer-readable storage medium.
[0037] The preferred embodiments of the present invention disclosed above are intended only to help illustrate the present invention. These preferred embodiments do not exhaustively describe all details, nor do they limit the present invention to the specific embodiments described. Obviously, many modifications and variations are possible based on the content of this specification. These embodiments are selected and described in detail in this specification to better explain the principles and practical applications of the present invention, thereby enabling those skilled in the art to better understand and utilize the present invention. The present invention is limited only by the claims and their full scope and equivalents.
Claims
1. A method for online encrypted transfer of electronic files, characterized in that: The steps include: Step S1: Obtain the electronic file to be transmitted for identification, divide it into blocks according to the identification results, and inject the blockchain timestamp into each block as a separate header identifier; Step S2: Determine the privacy information of each independent file block, encrypt the file block containing the privacy information, generate a dynamic key based on the quantum random number, and append the block hash value as the integrity check code; Step S3: The server authenticates the client and the receiving end decrypts the ciphertext; Step S4: The receiving end decrypts and obtains document blocks, which are then spliced together according to the topological relationship diagram; Step S5: Use an intelligent error correction and reorganization mechanism to manage electronic files.
2. The method for online encrypted transfer of electronic files according to claim 1, characterized in that: In step S1, the file type is identified using OCR technology, and the character count threshold is adjusted based on the identification result. The NLP model is used to detect the theme change points of the file, and the file is segmented based on the theme change points. During segmentation, character overlap is set between adjacent file blocks. The NLP model detects topic changes in archives, including: Semantic coherence analysis uses the BiLSTM-CRF sequence model to capture the semantic relevance between paragraphs. When the cosine similarity of the semantic vectors of adjacent text blocks is lower than a threshold, it is marked as a potential topic boundary. Identify text structure features: The TF-IDF algorithm verifies differences in file content, detects layout features such as title level changes and paragraph indentation, and identifies transition words in files. Statistical mutation detection uses the CUSUM algorithm to monitor the cumulative deviation of word frequency distribution. When the frequency of a specific domain term suddenly changes, segmentation is triggered. The KL divergence is calculated based on a dynamic sliding window to detect the vocabulary distribution probability. A multimodal fusion measurement strategy combines OCR layout analysis results in scanned archives to strengthen segmentation decisions when physical separation and semantic mutations of text regions occur simultaneously.
3. The method for online encrypted transfer of electronic files according to claim 2, characterized in that: After the archive is segmented according to the thematic change points, the initial segmentation results can be further subdivided, and the segmentation positions can be optimized in combination with syntax tree analysis. The specific plan is as follows: first, syntax analysis is performed based on the initial segmentation to construct a syntax tree containing sentence structure, clause hierarchy and key punctuation; then, complex sentences and logical connectives are detected through depth-first traversal, and syntax nodes are preferentially selected as segmentation points while maintaining semantic integrity; then, a dynamic recursive strategy is adopted to trigger secondary segmentation when it is detected that the nesting is too deep or the sentence complexity exceeds the standard, until the minimum semantic unit is met or the recursive depth limit is reached; finally, the syntax integrity and contextual coherence of each sub-block are verified to ensure that the segmentation results conform to the grammatical rules and retain the semantic logic of the original document.
4. The method for online encrypted transfer of electronic files according to claim 1, characterized in that: In step S2, the specific process of determining the privacy information of each independent file block is as follows: Step S21, data matrix construction: converting independent file block information into a real number matrix ; Among them, the real matrix center, number of rows Indicates different record entries, number of rows Indicates different attributes of the archive; Step S22, seed gene generation: randomly select a profile attribute as the initial seed gene, and calculate the correlation with other attributes through mutual information. The calculation formula is as follows: ; Where, are two information elements in the real matrix The mutual information value of Represents two information elements in the real matrix The information entropy of Represents two information elements The joint information entropy of Step S23, biclustering set construction: Calculate the mutual information value between the remaining attributes and the current seed gene, select the attribute with the smallest mutual information value as the new seed, and form two biclustering sets of general information and private information; Step S24, fuzzy partitioning implementation: Calculate the membership probability of non-seed genes to the biclusters, and the calculation formula is as follows: ; Where, Indicates that non-seed genes belong to the bicluster set The probability value of represents the seed gene set, Represents the real matrix Rank Column element biclustering probability, The calculation formula is as follows: ; Where, represents the size of the condition set in the biclustering set; Step S25: Classification determination: setting threshold , and the final classification is completed through the discriminant.
5. The method for online encrypted transfer of electronic files according to claim 1, characterized in that: In step S2, the encryption strategy for the file block containing private information is as follows: Step YS01: Encode the attributes of the privacy file block; Step YS02: Use an asymmetric encryption algorithm to encrypt the private file block to generate ciphertext; Step YS03: The encrypted document is sent to the receiving end through the server; Step YS04: After the receiving end verifies the client's identity, it receives the ciphertext and decrypts it.
6. The method for online encrypted transfer of electronic files according to claim 5, characterized in that: In step YS04, when the receiving end requests access, the server uses the socket interface to listen and accept the connection request, and the server performs identity authentication by parsing the address data and port number sent by the receiving end; if the identity authentication is passed, data transmission is performed; when the point-to-point communication transmission task is completed, the receiving end and the server interrupt the transmission control protocol connection by closing the program.
7. The method for online encrypted transfer of electronic files according to claim 1, characterized in that: In step S4, a topological relationship graph is constructed based on the electronic archive blocks generated in the intelligent block segmentation stage, and a directed acyclic graph is constructed to represent the topological relationship between archive blocks; the nodes represent the block data, and the edges include the order relationship between archive blocks and the double hash check code generated in the dynamic encryption stage; at the same time, a topological coordinate attribute is added to each archive block to record its logical position in the global archive, and an anti-collision hash function is used to calculate the association degree of adjacent archive blocks as the edge weight value.
8. The method for online encrypted transfer of electronic files according to claim 1, characterized in that: The specific process of dynamic verification of the double hash check code is as follows: Step S41: The user submits the zero-knowledge proof and the list of block IDs to be accessed, and the verifier confirms the validity of the permission; Step S42: The system loads the block topology graph and checks whether the requested block set constitutes a legal connected subgraph; Step S43: Decrypt the blocks in topological order, and verify the hash reference relationship between each decrypted block and the decrypted block; Step S44: Finally verify the topological hash tree root of the entire document and match the blockchain evidence value to complete end-to-end verification.
9. The method for online encrypted transfer of electronic files according to claim 1, characterized in that: In step S5, the intelligent error correction and reconstruction mechanism automatically adjusts the fault tolerance threshold based on the block characteristics, uses a sliding window algorithm to detect data continuity in real time, and triggers the reconstruction process when the missing area exceeds the preset threshold; The blockchain verification layer generates a new hash value for the reorganized block, which is stored on the chain together with the original block hash. Smart contracts are used to verify whether the reorganization logic complies with the preset rules.
10. An online encrypted transfer system for electronic files, comprising a source, a server, and a receiver, characterized by: The generating end includes an electronic archive recognition module, an intelligent block segmentation module, a privacy archive block determination module, a dynamic encryption module, and a secure transmission module; the electronic archive recognition module is used to convert paper archives into electronic archives and recognize them using OCR technology; the intelligent block segmentation module is used to segment electronic archives based on subject change points; the privacy archive block determination module is used to perform privacy determination on the segmented archive blocks; the dynamic encryption module is used to encrypt the privacy archive blocks; and the secure output module is used to transmit the encrypted privacy blocks and common archive blocks to the server; The server includes an identity authentication module and a file block management module; the identity authentication module is used to; the file management module is used to parse the address data and port number of the receiving end to verify the client's identity; the file management module is used to manage electronic files and file blocks; The receiving end includes an archive block decryption module, an archive block splicing module, and a smart contract detection module; the archive block decryption module is used to decrypt the received encrypted archive blocks; the archive block reassembly module is used to reassemble the decrypted archive blocks to complete the electronic archive; and the smart contract detection module is used to verify whether the logic of the reassembled electronic archive complies with preset rules.
Citation Information
Cited By
Audio-visual work material tracing method and system based on block chain
CN121561880A