File uploading and storing method

By extracting common category blocks in historical files for file division and slitting, combining file encoding sets and dual encryption mechanism, data redundancy and security problems in file storage systems are solved, and the efficiency and security of file upload and storage are improved.

CN120050271APending Publication Date: 2025-05-27WUHAN YUXIN SEMICON CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202411892623.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-20
Publication Date
2025-05-27

AI Technical Summary

Technical Problem

There are data redundancy problems in existing file storage systems, resulting in large storage space occupied and slow uploading and storage speeds, reducing the efficiency of file storage.

Method used

Reduce data redundancy by extracting common category blocks from historical files and dividing and slicing files based on these blocks. At the same time, a file encoding set is built and double encryption is performed to ensure the security of the files during transmission and storage.

Benefits of technology

It effectively reduces database redundancy, optimizes storage space, and improves the efficiency and security of file upload and storage.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120050271A_ABST
    Figure CN120050271A_ABST
Patent Text Reader

Abstract

The invention provides a file uploading and storing method, which comprises the following steps of: extracting public category blocks from a historical file, dividing the file into to-be-divided blocks according to the positions and the number of the public category blocks, constructing a dividing and cutting model by utilizing network uploading information and file information of the to-be-divided blocks, further dividing and cutting each to-be-divided block, and storing the to-be-divided blocks into a file storage module; the method comprises the following steps of: dividing a public key into a plurality of public category blocks to obtain split sub-blocks, sequentially arranging all the public category blocks and the split sub-blocks according to file contents, calculating hash values for association, constructing a file code set, performing dual encryption on file codes, and storing the encrypted public key, the public category blocks, the split sub-blocks and the file code set into a database according to category attributes. According to the method, the public category blocks are used for dividing and cutting the files, data redundancy is reduced, file processing efficiency is improved, quick retrieval and verification of the files are achieved by constructing the file coding set and calculating the hash value, and meanwhile the safety of the files in the transmission and storage process is guaranteed through a dual encryption mechanism.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of file storage, and particularly to a file upload and storage method. Background Art

[0002] With the rapid development of information technology, digital and networked data management methods have become an indispensable part of the daily lives of enterprises, organizations, and individuals. However, this convenient data management method also brings data security problems. Traditional file storage systems often have risks such as data leakage, data tampering, illegal access, and data loss. These risks are huge hidden dangers for enterprises and individuals storing sensitive information.

[0003] A file storage system with the publication number 201811191994.4 includes an application layer, an interface adaptation layer, and a storage layer; the application layer is used to provide a user interface for application programs that can generate network traffic and provide corresponding network application services; the interface adaptation layer is created between the application layer and the storage layer, and encapsulates the corresponding interfaces of the file storage system that match two or more file storage methods into the interface adaptation layer. The interface adaptation layer is used to call the storage interface that matches the current storage mode according to the detected storage mode corresponding to the current storage server; after the interface adaptation layer calls the storage interface that matches the current storage mode according to the detected storage mode corresponding to the current storage server, the storage layer is used to perform storage processing that matches the current storage interface.

[0004] In the existing file storage process, due to a large amount of duplicate content in some files, this has led to an increasing amount of redundant data in the database, which not only occupies a large amount of storage space but also reduces the file upload and storage speed, thereby reducing the efficiency of file storage. Summary of the Invention

[0005] In view of this, the present invention proposes a file upload and storage method, which can reduce database redundancy, improve the efficiency of file upload and storage, and at the same time, ensure the security of files during transmission and storage, thereby improving the efficiency and security of file upload and storage.

[0006] The technical solution of the present invention is realized as follows: The present invention provides a file upload and storage method, including the following steps:

[0007] S1, obtain historical file upload data, identify and classify the content of historical files, extract the content areas containing the same content in historical files, and obtain common category blocks;

[0008] S2. Obtain the location information of the common category blocks in the historical files, divide the historical files according to the location information of the common category blocks and the number of file common category blocks contained in the corresponding historical files, and obtain several to-be-sliced blocks corresponding to the historical files;

[0009] S3. Obtain the network upload information, construct a slicing model according to the network upload information and the file information corresponding to the to-be-sliced blocks, and slice each to-be-sliced block based on the slicing model to obtain sliced sub-blocks;

[0010] S4. Obtain all the common category blocks and sliced sub-blocks in each historical file, arrange all the common category blocks and sliced sub-blocks in the order of file content, calculate the corresponding hash values for association, and construct a file coding set;

[0011] S5. Encrypt each file code, generate an encryption private key, encrypt the encryption private key to obtain an encryption public key, and store the encryption public key, common category blocks, sliced sub-blocks and file coding set in the database according to the category attributes.

[0012] Based on the above technical solutions, preferably, step S1 includes the following sub-steps:

[0013] S11. Obtain the historical file upload information from the file server, convert the historical files into text data, and perform classification and data preprocessing on the text data to obtain a historical standard text data set;

[0014] S12. Divide each text data in the corresponding category into paragraph blocks according to paragraphs, calculate the similarity between any two paragraph blocks, obtain a similarity matrix set of each paragraph block in the corresponding category, and the calculation expression is:

[0015]

[0016] In the formula, w is the similarity, A and B are the vectors of two document blocks, · represents the dot product operation, ||A|| and ||B|| are the Euclidean norms of vectors A and B respectively;

[0017] S14. Preset a similarity threshold, traverse the similarity matrix set of each paragraph block in the corresponding category, obtain the paragraph blocks with similarity greater than the similarity threshold, and extract the same content of the paragraph blocks to obtain common category blocks.

[0018] Based on the above technical solutions, preferably, step S14 includes the following sub-steps:

[0019] Obtain the paragraph blocks with similarity greater than the similarity threshold, and for the characters in the paragraph blocks with similarity greater than the similarity threshold, obtain the longest common string in the paragraph blocks based on the LCS algorithm;

[0020] A preset length threshold is used to determine whether the length of the longest common string in the paragraph block is greater than the preset length threshold. If it is greater, the content of the longest common string is used as the same content of the paragraph block, and the same content of the paragraph block is extracted to obtain a common category block. If it is less, the longest common string is ignored.

[0021] Based on the above technical solutions, preferably, step S2 includes the following sub-steps:

[0022] If the number of common category blocks contained in the document data is 0, the text data is divided into one to-be-split block;

[0023] If the number of common category blocks contained in the document data is 1, the text between the first character of the text data and the start of the common category block and the text between the end of the common category block and the end character of the text data are divided, and the text data is divided into two to-be-split blocks;

[0024] If the number of common category blocks contained in the document data is greater than 1, the text between the first character of the text data and the start of the common category block, the text between the end of the common category block and the end character of the text data, and the text between two adjacent common category blocks are divided, and the text data is divided into multiple to-be-split blocks;

[0025] According to the division result, and in the order of the document content, the text content of each to-be-split block is extracted from the text data to obtain several to-be-split blocks corresponding to the text data.

[0026] Based on the above technical solutions, preferably, step S3 includes the following sub-steps:

[0027] Obtain network upload information and file information corresponding to the to-be-split block. The network upload information includes upload average speed, upload packet loss rate, and upload average delay data. The file information corresponding to the to-be-split block includes file size;

[0028] Construct a splitting model according to the upload average speed, upload packet loss rate, upload average delay data, and the file size corresponding to the to-be-split block. The expression of the splitting model is:

[0029]

[0030] In the formula, Q is the size of the split sub-block, W is the file size of the to-be-split block, V s is the upload average speed, G s is the benchmark upload speed, α is the upload speed adjustment coefficient, which is used to adjust the size of the split sub-block according to the upload average speed relative to the benchmark upload speed; D is the upload packet loss rate, β is the upload packet loss rate adjustment coefficient, which is used to adjust the size of the split sub-block according to the upload packet loss rate, G rFor the upload reference delay, V r For the upload average delay data, γ is the upload delay adjustment parameter used to adjust the size of the sliced sub - blocks according to the upload average delay, and n is the exponential parameter;

[0031] Slice each block to be sliced according to the slicing model to obtain a corresponding number of sliced sub - blocks.

[0032] On the basis of the above technical solution, preferably, step S4 includes the following sub - steps:

[0033] Construct a file encoding set, obtain all common category blocks and sliced sub - blocks in each document information, and arrange all common category blocks and sliced sub - blocks in the order of file content;

[0034] Calculate the hash values of all common category blocks and sliced sub - blocks in each document information, and associate the hash values with the corresponding common category blocks or sliced sub - blocks and document information;

[0035] Separate the hash values with a delimiter in the order of file content to obtain the file encoding of the document information, and add the file encodings of each document information to the file encoding set to obtain the file encoding set.

[0036] On the basis of the above technical solution, preferably, step S5 includes encrypting the file encoding using the AES symmetric encryption algorithm to generate an encryption private key, and encrypting the encryption private key using the RSA - 2048 asymmetric encryption algorithm to obtain an encryption public key.

[0037] On the basis of the above technical solution, preferably, the method further includes:

[0038] Obtain the target upload file, identify and classify the content of the target upload file, obtain the common category blocks in the corresponding category from the database according to the category attributes, and identify whether the content of the target upload file contains the same content area;

[0039] If it contains, divide the target upload file into several blocks to be sliced according to the identified common category blocks and the quantity, and slice each block to be sliced based on the slicing model to obtain sliced sub - blocks;

[0040] If it does not contain common category blocks, slice the target upload file based on the slicing model to obtain sliced sub - blocks;

[0041] Obtain all sliced sub - blocks and / or common category blocks in each historical file, arrange all sliced sub - blocks and / or common category blocks in the order of file content, calculate the corresponding hash values for association, and construct a file encoding set;

[0042] Encrypt each file code to generate an encrypted private key, and encrypt the encrypted private key to obtain an encrypted public key. Store the encrypted public key, public category block, sliced sub-block, and file code set in the database according to the category attributes.

[0043] In a second aspect, the present invention also provides an electronic device, including at least one processor, at least one memory, a communication interface, and a bus; wherein, the processor, memory, and communication interface complete communication with each other through the bus; the memory stores a file upload and storage method program executable by the processor, and a file upload and storage method program is configured to implement a file upload and storage method as described above.

[0044] In a third aspect, the present invention also provides a computer-readable storage medium, on which a file upload and storage method program is stored, and when the file upload and storage method program is executed, it implements a file upload and storage method as described above.

[0045] The file upload and storage method of the present invention has the following beneficial effects compared with the prior art:

[0046] (1) By extracting public category blocks from historical files and partitioning and slicing files according to the public category blocks, data redundancy is reduced, storage space is optimized, storage efficiency is improved, and by constructing a file code set and calculating hash values, fast retrieval and verification of files are achieved. At the same time, the dual encryption mechanism ensures the security of files during transmission and storage, improving the efficiency and security of file upload and storage; (2) By pre-computing the similarity of two paragraph blocks from historical file upload data and combining with the LCS algorithm to obtain the longest common string between similar paragraph blocks, the efficiency and accuracy of obtaining public category blocks are improved. At the same time, by setting a length threshold, those common strings that are similar but too short and have no practical significance can be filtered out, avoiding waste of storage space and complexity of subsequent processing caused by extracting trivial content;

[0047] (3) By constructing a slicing model and considering network upload information, i.e., upload average speed, upload packet loss rate, and upload average delay data, and file information of the to-be-sliced block, i.e., file size, the size of the sliced sub-block can be dynamically adjusted to ensure the efficiency and reliability of the sliced sub-block in network transmission, reduce transmission delay or failure caused by poor network conditions, and improve the success rate of file transmission. BRIEF DESCRIPTION OF THE DRAWINGS

[0048] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the accompanying drawings required in the description of the embodiments or the prior art. Obviously, the accompanying drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other accompanying drawings can be obtained based on these drawings.

[0049] Figure 1 It is a flowchart of the file upload and storage method of the present invention;

[0050] Figure 2 It is a schematic structural diagram of the device of the hardware operating environment involved in the embodiment solution of the present invention. Specific embodiments

[0051] The following will describe clearly and completely the technical solutions in the embodiments of the present invention in combination with the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, rather than all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present invention.

[0052] As Figure 1 shown, a file upload and storage method of the present invention includes the following steps:

[0053] S1. Obtain historical file upload data, identify and classify the content of the historical files, extract the content areas containing the same content in the historical files, and obtain common category blocks.

[0054] Among them, step S1 includes the following sub-steps:

[0055] S11. Obtain historical file upload information from the file server, convert the historical files into text data, and perform classification and data preprocessing on the text data to obtain a historical standard text data set;

[0056] S12. Split each text data in the corresponding category into paragraph blocks according to paragraphs, calculate the similarity between any two paragraph blocks, obtain a similarity matrix set of each paragraph block in the corresponding category, and the calculation expression is:

[0057]

[0058] In the formula, w is the similarity, A and B are the vectors of two document blocks, · represents the dot product operation, and ||A|| and ||B|| are the Euclidean norms of vectors A and B respectively;

[0059] S14. A preset similarity threshold is used to traverse the similarity matrix set of each paragraph block under the corresponding category, obtain the paragraph blocks with similarity greater than the similarity threshold, and extract the same content of the paragraph blocks to obtain a common category block.

[0060] It should be noted that converting historical files into text data, classifying and preprocessing the data can standardize and normalize the data, providing convenience for subsequent processing. Among them, through preprocessing the data, including removing stop words and stemming, etc., the processing accuracy and accuracy of text data can be improved. Splitting the text data into paragraphs helps to analyze the text content more precisely. By calculating the similarity between any two paragraph blocks, the similarity degree of the text content can be quantified, providing a basis for extracting common category blocks in the future. By setting a similarity threshold and traversing the similarity matrix set, the same content in the paragraph blocks with higher similarity can be accurately extracted to form a common category block. The extracted common category block can replace the repeated content in the original text, thereby optimizing the storage space and improving the storage efficiency.

[0061] In this embodiment, common category blocks are mined and extracted from the uploaded data of historical files. These common category blocks are regions that frequently appear and have the same or highly similar content in multiple historical files. By identifying and extracting common category blocks, the storage of duplicate data is avoided, thus significantly saving the storage space. Moreover, in the subsequent file processing, search, or retrieval process, by identifying common category blocks, relevant files can be quickly located and processed, improving the processing efficiency.

[0062] The following sub-steps are included in step S14:

[0063] Obtain the paragraph blocks with similarity greater than the similarity threshold. For the characters in the paragraph blocks with similarity greater than the similarity threshold, obtain the longest common string in the paragraph blocks based on the LCS algorithm.

[0064] A preset length threshold is set. Determine whether the length of the longest common string in the paragraph block is greater than the preset length threshold. If it is greater, take the content of the longest common string as the same content of the paragraph block, extract the same content of the paragraph block to obtain a common category block. If it is less, ignore the longest common string.

[0065] It should be noted that by using the LCS algorithm to obtain the longest common string between similar paragraph blocks, it is ensured that the extracted common content is the longest part of the content that is truly shared between these paragraph blocks. By setting a length threshold, those common strings that are similar but too short and have no practical significance can be filtered out, avoiding the waste of storage space caused by extracting trivial content and the complexity of subsequent processing.

[0066] S2. Obtain the position information of the common category blocks in the historical file, and divide the historical file according to the position information of the common category blocks and the number of file common category blocks contained in the corresponding historical file, so as to obtain several to-be-split chunks corresponding to the historical file.

[0067] Among them, step S2 includes the following sub-steps:

[0068] If the number of common category blocks contained in the document data is 0, then divide the text data into one to-be-split chunk;

[0069] If the number of common category blocks contained in the document data is 1, then divide the text between the first character of the text data and the start of the common category block and the text between the end of the common category block and the last character of the text data, and divide the text data into two to-be-split chunks;

[0070] If the number of common category blocks contained in the document data is greater than 1, then divide the text between the first character of the text data and the start of the common category block, the text between the end of the common category block and the last character of the text data, and the text between two adjacent common category blocks, and divide the text data into multiple to-be-split chunks;

[0071] According to the division result, extract the text content of each to-be-split chunk from the text data in the order of the document content, so as to obtain several to-be-split chunks corresponding to the text data.

[0072] It should be noted that dividing the historical file according to the position information and quantity of the common category blocks in the historical file to obtain several to-be-split chunks, this division method more precisely reflects the structure and characteristics of the file content, which helps to achieve more refined file processing and analysis.

[0073] S3. Obtain the network upload information, construct a splitting model according to the network upload information and the file information corresponding to the to-be-split chunks, and split each to-be-split chunk based on the splitting model to obtain split sub-chunks.

[0074] Among them, step S3 includes the following sub-steps:

[0075] Obtain the network upload information and the file information corresponding to the to-be-split chunks. The network upload information includes the upload average speed, upload packet loss rate, and upload average delay data, and the file information corresponding to the to-be-split chunks includes the file size;

[0076] Construct a splitting model according to the upload average speed, upload packet loss rate, upload average delay data, and the file size corresponding to the to-be-split chunks. The expression of the splitting model is:

[0077]

[0078] In the formula, Q is the size of the sliced sub-block, W is the file size of the block to be sliced, V s is the average upload speed, G s is the reference upload speed, α is the upload speed adjustment coefficient, which is used to adjust the size of the sliced sub-block according to the ratio of the average upload speed to the reference upload speed; D is the upload packet loss rate, β is the upload packet loss rate adjustment coefficient, which is used to adjust the size of the sliced sub-block according to the upload packet loss rate, G r is the reference upload delay, V r is the average upload delay data, γ is the upload delay adjustment parameter, which is used to adjust the size of the sliced sub-block according to the upload average delay, and n is the exponential parameter;

[0079] Slice each block to be sliced according to the slicing model to obtain a corresponding number of sliced sub-blocks.

[0080] It should be noted that by constructing a slicing model and considering network upload information, i.e., average upload speed, upload packet loss rate, and average upload delay data, and file information of the block to be sliced, i.e., file size, the size of the sliced sub-block can be dynamically adjusted to ensure the efficiency and reliability of the sliced sub-block in network transmission, reduce transmission delays or failures caused by poor network conditions, and improve the success rate of file transmission.

[0081] S4. Obtain all common category blocks and sliced sub-blocks in each historical file, arrange all common category blocks and sliced sub-blocks in the order of file content, calculate the corresponding hash values for association, and construct a file encoding set;

[0082] Among them, step S4 includes the following sub-steps:

[0083] Construct a file encoding set, obtain all common category blocks and sliced sub-blocks in each document information, and arrange all common category blocks and sliced sub-blocks in the order of file content;

[0084] Calculate the hash values of all common category blocks and sliced sub-blocks in each document information, and associate the hash values with the corresponding common category blocks or sliced sub-blocks and document information;

[0085] Separate the hash values with delimiters in the order of file content to obtain the file encoding of the document information, and add the file encodings of each document information to the file encoding set to obtain the file encoding set.

[0086] It should be noted that by calculating the hash values of the common category blocks and the sliced sub-blocks and associating them with the corresponding file information, a unique file code can be generated for each file. This code can accurately reflect the content and structure of the file, thereby improving the accuracy of file recognition. The file code set provides an efficient indexing mechanism for files. By comparing the file codes, the required file can be quickly located without checking the content of each file one by one, greatly improving the speed and efficiency of file retrieval. Moreover, by using the hash value as the index value, since the hash value is unique and irreversible, it can be used as an effective means to verify the integrity of the file content and can verify whether the file content has been tampered with or damaged.

[0087] S5. Encrypt each file code to generate an encrypted private key, and then encrypt the encrypted private key to obtain an encrypted public key. Store the encrypted public key, the common category blocks, the sliced sub-blocks, and the file code set in the database according to the category attributes.

[0088] Among them, in step S5, it includes using the AES symmetric encryption algorithm to encrypt the file code to generate an encrypted private key, and using the RSA-2048 asymmetric encryption algorithm to encrypt the encrypted private key to obtain an encrypted public key.

[0089] It should be noted that the common category blocks and the sliced sub-blocks are stored dispersedly on multiple physical nodes, and each node is located in a different geographical location to achieve geographical redundancy. The dispersed storage layer records the location information of each file block for subsequent file reconstruction and recovery. Decrypt the encrypted private key according to the corresponding encrypted public key, obtain the file code of the corresponding file from the file code set through the decrypted encrypted private key, retrieve the corresponding common category blocks and / or sliced sub-blocks of the file from each physical node according to the file code and the corresponding recorded location information, and recombine the retrieved common category blocks and / or sliced sub-blocks into a complete file, and the user can download the decrypted file.

[0090] In addition, when a user requests to access or download a file, first perform identity verification to verify the legitimacy of the user; after passing the identity verification, the access control layer checks the user's permissions to ensure that the user has the right to access the requested file; if the user's permission verification passes, the access control layer will allow the user to continue the next operation; otherwise, the user's access request will be rejected.

[0091] In addition, it also includes backing up the data, regularly selecting important data for backup to ensure the recoverability of the data. The backup data will be encrypted and stored on a remote server to prevent data loss or damage. In the case of data loss or damage, the data backup layer can provide the backup data for recovery to ensure the integrity and availability of the data.

[0092] In this implementation, when performing distributed storage in the public category block and / or the sliced sub-block, a space management module is also included, which is responsible for monitoring the storage space of each physical node. When it detects that a certain node is approaching the capacity limit, it automatically triggers the load balancing mechanism to migrate new data or some existing data to other nodes; and a fault tolerance and recovery mechanism, which, in the event of node loss or failure, quickly reconstructs the lost data through the redundant data blocks stored on other nodes to maintain data integrity and system stability.

[0093] In this embodiment, the method not only includes the steps of processing and storing existing files described previously, but also extends the processing flow for newly uploaded target files. The method further includes:

[0094] Obtain the target uploaded file, identify and classify the content of the target uploaded file, obtain the corresponding public category blocks in the corresponding category from the database according to the category attributes, and identify whether the content of the target uploaded file contains the same content area;

[0095] If it contains, divide the target uploaded file into several to-be-sliced blocks according to the identified public category blocks and the quantity; perform slicing on each to-be-sliced block based on the slicing model to obtain sliced sub-blocks;

[0096] If it does not contain public category blocks, perform slicing on the target uploaded file based on the slicing model to obtain sliced sub-blocks;

[0097] Obtain all the sliced sub-blocks and / or public category blocks in each historical file, arrange all the sliced sub-blocks and / or public category blocks in the order of file content, and calculate the corresponding hash values for association to construct a file encoding set;

[0098] Encrypt each file encoding to generate an encrypted private key, and encrypt the encrypted private key to obtain an encrypted public key. Store the encrypted public key, public category blocks, sliced sub-blocks, and file encoding set in the database according to the category attributes.

[0099] In this embodiment, public category blocks are extracted from historical files, and the files are divided into to-be-sliced blocks according to the positions and quantities of these public category blocks. A slicing model is constructed using network upload information and the file information of the to-be-sliced blocks, and each to-be-sliced block is further sliced to obtain sliced sub-blocks. All the public category blocks and sliced sub-blocks are arranged in the order of file content, and hash values are calculated for association to construct a file encoding set. The file encoding is double-encrypted, and the encrypted public key, public category blocks, sliced sub-blocks, and file encoding set are stored in the database according to the category attributes. Using public category blocks for file division and slicing reduces data redundancy and improves file processing efficiency. And by constructing a file encoding set and calculating hash values, fast file retrieval and verification are achieved. At the same time, the double-encryption mechanism ensures the security of files during transmission and storage.

[0100] Those of ordinary skill in the art can realize that the units and algorithm steps of the examples described in combination with the embodiments disclosed herein can be implemented by electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present invention.

[0101] Those skilled in the art can clearly understand that for the convenience and conciseness of description, the specific working processes of the systems and modules described above can refer to the corresponding processes in the foregoing method embodiments, and will not be elaborated herein.

[0102] In the embodiments provided by the present invention, it should be understood that the disclosed systems and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other can be through some interfaces. The indirect couplings or communication connections of devices or units can be in electrical, mechanical or other forms.

[0103] The units described as separate components may or may not be physically separated. The components shown as units may or may not be physical units, that is, they can be located in one place, or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0104] In addition, in each embodiment of the present invention, the functional units can be integrated into one processing unit, or each unit exists physically alone, or two or more units can be integrated into one unit.

[0105] When the above-mentioned functions are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on such understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art or a part of this technical solution can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention. The aforementioned storage medium includes: various media such as USB flash drives, mobile hard disks, ROM, RAM, magnetic disks, or optical discs that can store program codes.

[0106] In addition, it should be noted that in the systems and methods of the present invention, obviously, each component or each step can be decomposed and / or recombined. These decompositions and / or recombinations should be regarded as equivalent solutions of the present invention. And, the steps of performing the above series of processes can naturally be executed in chronological order according to the described order, but it is not necessary to execute them in chronological order. Some steps can be executed in parallel or independently of each other. For those of ordinary skill in the art, it can be understood that all or any steps or components of the methods and devices of the present invention can be implemented in any computing device (including processors, storage media, etc.) or in a network of computing devices in the form of hardware, firmware, software, or a combination thereof, which can be achieved by those of ordinary skill in the art using their basic programming skills after reading the description of the present invention.

[0107] Therefore, the object of the present invention can also be achieved by running a program or a set of programs on any computing system. The computing system can be a well-known general system. Therefore, the object of the present invention can also be achieved only by providing a program product containing program codes for implementing the method or device. That is to say, such a program product also constitutes the present invention, and the storage medium storing such a program product also constitutes the present invention. Obviously, the storage medium can be any well-known storage medium or any storage medium developed in the future. It should also be noted that in the devices and methods of the present invention, obviously, each component or each step can be decomposed and / or recombined. These decompositions and / or recombinations should be regarded as equivalent solutions of the present invention. And, the steps of performing the above series of processes can naturally be executed in chronological order according to the described order, but it is not necessary to execute them in chronological order. Some steps can be executed in parallel or independently of each other.

[0108] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.

Claims

1. A file upload and storage method, characterized in that: The following steps are involved: S1, obtain the uploaded data of historical files, identify and classify the contents of historical files, extract the same content areas in the historical files, and obtain the common category blocks; S2, obtaining the location information of the common category blocks in the history file, dividing the history file according to the location information of the common category blocks and the number of the common category blocks contained in the corresponding history file, and obtaining a number of to-be-divided blocks corresponding to the history file; S3, obtaining network upload information, building a segmentation model according to the network upload information and the file information corresponding to the blocks to be segmented, and segmenting each block to be segmented based on the segmentation model to obtain segmented sub-blocks; S4, obtaining all common category blocks and split sub-blocks in each historical file, arranging all common category blocks and split sub-blocks in the order of file content, and calculating corresponding hash values ​​for association, and constructing a file encoding set; S5, encrypt each file code, generate an encrypted private key, and encrypt the encrypted private key to obtain an encrypted public key, and store the encrypted public key, public category block, split sub-block and file code set in the database according to category attributes.

2. A file upload storage method as claimed in claim 1, characterized in that: Step S1 includes the following sub-steps: S11, obtaining historical file upload information from the file server, converting the historical files into text data, and classifying and preprocessing the text data to obtain a historical standard text data set; S12, according to the category, each text data under the corresponding category is divided into paragraph blocks according to the paragraph, and similarity is calculated for any two paragraph blocks to obtain a similarity matrix set of each paragraph block under the corresponding category. The calculation expression is: Where w is the similarity, A and B are the vectors of two document blocks, · represents the dot product operation, ||A| and ||B|| are the Euclidean norms of vectors A and B respectively; S14, preset a similarity threshold, traverse the similarity matrix set of each paragraph block under the corresponding category, obtain paragraph blocks with similarity greater than the similarity threshold, and extract the same content of the paragraph blocks to obtain common category blocks.

3. A file upload storage method as claimed in claim 2, characterized in that: Step S14 includes the following sub-steps: Obtain paragraph blocks whose similarity is greater than a similarity threshold, and for characters in the paragraph blocks whose similarity is greater than the similarity threshold, obtain the longest common character string in the paragraph blocks based on the LCS algorithm; A preset length threshold is used to determine whether the length of the longest common string in the paragraph block is greater than the preset length threshold. If it is greater, the content of the longest common string is used as the same content of the paragraph block, the same content of the paragraph block is extracted, and the public category block is obtained. If it is less than the longest common string, the longest common string is ignored.

4. A file upload storage method as claimed in claim 3, characterized in that: Step S2 includes the following sub-steps: If the number of common category blocks contained in the document data is 0, the text data is divided into a block to be segmented; If the number of common category blocks included in the document data is 1, the text between the first character of the text data and the beginning of the common category block and the text between the end of the common category block and the end character of the text data are divided, and the text data is divided into two blocks to be divided; If the number of common category blocks included in the document data is greater than 1, the text between the first character of the text data and the beginning of the common category block, the text between the end of the common category block and the end character of the text data, and the text between two adjacent common category blocks are divided to divide the text data into a plurality of blocks to be divided; According to the result of the division and in the order of the document contents, the text content of each block to be divided is extracted from the text data to obtain a plurality of blocks to be divided corresponding to the text data.

5. A file upload and storage method as claimed in claim 4, characterized in that: Step S3 includes the following sub-steps: Obtain network upload information and file information corresponding to the blocks to be segmented, wherein the network upload information includes average upload speed, upload packet loss rate, and average upload delay data, and the file information corresponding to the blocks to be segmented includes file size; A segmentation model is constructed based on the average upload speed, upload packet loss rate, average upload delay data, and the file size corresponding to the block to be segmented. The expression of the segmentation model is: Where Q is the size of the split sub-block, W is the file size of the block to be split, and V s is the average upload speed, G s is the benchmark upload speed, α is the upload speed adjustment coefficient, which is used to adjust the size of the split sub-block according to the average upload speed being equal to the benchmark upload speed; D is the upload packet loss rate, β is the upload packet loss rate adjustment coefficient, which is used to adjust the size of the split sub-block according to the upload packet loss rate, G r For upload reference delay, V r is the average upload delay data, γ is the upload delay adjustment parameter, which is used to adjust the size of the segmented sub-block according to the average upload delay, and n is the exponential parameter; Each block to be cut is cut according to the cutting model to obtain a corresponding number of cut sub-blocks.

6. A file upload and storage method as claimed in claim 5, characterized in that: Step S4 includes the following sub-steps: Construct a file encoding set, obtain all common category blocks and split sub-blocks in each document information, and arrange all common category blocks and split sub-blocks in the order of file content; Calculate the hash values ​​of all common category blocks and split sub-blocks in each document information, and associate the hash values ​​with the corresponding common category blocks or split sub-blocks and document information; The hash values ​​are separated by separators according to the order of the file contents to obtain the file encoding of the document information, and the file encoding of each document information is added to the file encoding set to obtain the file encoding set.

7. A file upload and storage method as claimed in claim 6, characterized in that: Step S5 includes encrypting the file code using the AES symmetric encryption algorithm to generate an encrypted private key, and encrypting the encrypted private key using the RSA-2048 asymmetric encryption algorithm to obtain an encrypted public key.

8. A file upload and storage method as claimed in claim 7, characterized in that: The method further comprises: Obtaining a target uploaded file, identifying and classifying the content of the target uploaded file, obtaining a common category block in a corresponding category from a database according to the category attribute, and identifying whether the content of the target uploaded file contains the same content area; If it includes dividing according to the identified common category blocks and quantity, the target uploaded file is divided into a number of blocks to be divided; each block to be divided is divided based on the division model to obtain divided sub-blocks; If the common category block is not included, the target uploaded file is segmented based on the segmentation model to obtain segmented sub-blocks; Obtain all the segmented sub-blocks and / or public category blocks in each historical file, arrange all the segmented sub-blocks and / or public category blocks in the order of the file contents, calculate the corresponding hash values ​​for association, and construct a file encoding set; Each file code is encrypted to generate an encrypted private key, and the encrypted private key is encrypted to obtain an encrypted public key. The encrypted public key, public category block, split sub-block and file code set are stored in the database according to category attributes.

9. An electronic device, characterized in that: It includes at least one processor, at least one memory, a communication interface and a bus; wherein the processor, memory and communication interface communicate with each other through the bus; the memory stores a file upload storage method program that can be executed by the processor, and the file upload storage method program is configured to implement a file upload storage method as claimed in any one of claims 1 to 8.

10. A computer-readable storage medium, characterized in that: The storage medium stores a file upload storage method program, which, when executed, implements a file upload storage method as claimed in any one of claims 1 to 8.

Citation Information

Patent Citations

  • A file storage system and a file storage method

    CN109508323A