Storage method and device of authentication toolkit, electronic equipment and storage medium
By generating unique identifiers of the toolkit and checking the distributed hash table, the problem of repeated storage of the authentication toolkit is solved, and the savings and efficiency improvement of storage resources are achieved.
Patent Information
- Application Number
- CN202510832044.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-20
- Publication Date
- 2025-07-22
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
The storage method of the authentication toolkit in the prior art leads to wasting storage space, and there is a problem of repeatedly storing the same content.
By generating unique identifiers of the toolkit, using a distributed hash table to check the existence of the identifier, establish the association relationship between the identifier and the node, logically delete the duplicate toolkit copy, and only physically store the new file content.
Avoid repeated storing of the same content in distributed storage systems, saving storage resources and improving storage efficiency.
Smart Images

Figure CN120354400A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of computers, and in particular, to a storage method, device, electronic device, computer-readable storage medium, and computer program product for an authentication toolkit. Background Art
[0002] In the digital age, operating system authentication is one of the important means of security protection, which can ensure whether the hardware and software configurations of the server are compatible with the operating system. Operating system authentication is usually completed using an authentication toolkit.
[0003] In the related art, the authentication toolkit is downloaded and stored in a storage system, and the authentication toolkit is used for authentication when operating system authentication is required.
[0004] However, the storage method of the toolkit in the related art has the problem of wasting storage space. Summary of the Invention
[0005] This application provides a storage method, device, electronic device, computer-readable storage medium, and computer program product for an authentication toolkit, so as to at least solve the problem of duplicate storage of toolkits with the same file content in the related art and waste of storage space.
[0006] This application provides a storage method for an authentication toolkit, which is applied to a distributed storage system. The distributed storage system includes multiple nodes. The method includes: when receiving a first toolkit uploaded by a first node, generating a first identifier corresponding to the first toolkit according to the file content included in the first toolkit, where the first identifier is used to indicate the file content included in the first toolkit; determining whether the first identifier exists in a preset identifier table, where the identifier table includes multiple identifiers, and the multiple identifiers respectively correspond to the file contents of multiple toolkits already stored in the distributed storage system; when the first identifier exists in the identifier table, establishing an association relationship between the first identifier and the first node, and deleting the first toolkit uploaded by the first node.
[0007] This application also provides a storage device for an authentication toolkit, including: an identifier generation module, configured to generate a first identifier corresponding to a first toolkit according to the file content included in the first toolkit when receiving the first toolkit uploaded by a first node, where the first identifier is used to indicate the file content included in the first toolkit; an identifier query module, configured to determine whether the first identifier exists in a preset identifier table, where the identifier table includes multiple identifiers, and the multiple identifiers respectively correspond to the file contents of multiple toolkits already stored in the distributed storage system; a first association module, configured to establish an association relationship between the first identifier and the first node when the first identifier exists in the identifier table, and delete the first toolkit uploaded by the first node.
[0008] The present application also provides an electronic device, including: a memory for storing a computer program; a processor for implementing the steps of any of the above storage methods of the authentication toolkit when executing the computer program.
[0009] The present application also provides a computer-readable storage medium storing a computer program, wherein the computer program implements the steps of any of the above storage methods of the authentication toolkit when executed by a processor.
[0010] The present application also provides a computer program product including a computer program, which implements the steps of any of the above storage methods of the authentication toolkit when executed by a processor.
[0011] Through the present application, when receiving the first toolkit uploaded by the first node, a first identifier corresponding to the first toolkit is generated according to the file content included in the first toolkit, and the identifier is used to indicate the file content. Even if the file names of multiple toolkits are different but the file contents are the same, the same identifier will still be generated. By introducing the identifier of the file content, the identification of the file content of the toolkit is realized. Then, it is determined whether the first identifier exists in the preset identification table. When the first identifier exists in the identification table, an association relationship is established between the first identifier and the first node, and the first toolkit uploaded by the first node is deleted, ensuring that only new file content will be physically stored, avoiding duplicate storage. When the file content is repeated, an association relationship is logically established between the first identifier and the first node, and then the toolkit uploaded by the first node is deleted, reducing the storage cost. Since the unique identifier of the toolkit is generated and it is checked whether the identifier exists in the identification table, the technical problem of duplicate storage of toolkits with the same content in the distributed storage system is avoided, achieving the technical effect of saving storage resources. BRIEF DESCRIPTION OF THE DRAWINGS
[0012] To more clearly illustrate the embodiments of the present application, the drawings required for use in the embodiments will be briefly introduced below. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0013] Figure 1 is a hardware structure block diagram of a server device for a storage method of an authentication toolkit according to an embodiment of the present application;
[0014] Figure 2 is one of the flowcharts of a storage method of an authentication toolkit according to an embodiment of the present application;
[0015] Figure 3 It is the second flowchart of a storage method for an authentication toolkit according to an embodiment of the present application;
[0016] Figure 4 It is the third flowchart of a storage method for an authentication toolkit according to an embodiment of the present application;
[0017] Figure 5 It is the fourth flowchart of a storage method for an authentication toolkit according to an embodiment of the present application;
[0018] Figure 6 It is the fifth flowchart of a storage method for an authentication toolkit according to an embodiment of the present application;
[0019] Figure 7 It is the sixth flowchart of a storage method for an authentication toolkit according to an embodiment of the present application;
[0020] Figure 8 It is the seventh flowchart of a storage method for an authentication toolkit according to an embodiment of the present application;
[0021] Figure 9 It is the eighth flowchart of a storage method for an authentication toolkit according to an embodiment of the present application;
[0022] Figure 10 It is the ninth flowchart of a storage method for an authentication toolkit according to an embodiment of the present application;
[0023] Figure 11 It is the tenth flowchart of a storage method for an authentication toolkit according to an embodiment of the present application;
[0024] Figure 12 It is the eleventh flowchart of a storage method for an authentication toolkit according to an embodiment of the present application;
[0025] Figure 13 It is the structural block diagram of a storage device for an authentication toolkit according to an embodiment of the present application. Detailed implementation manners
[0026] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the protection scope of the present application.
[0027] It should be noted that in the description of this application, the terms "include", "comprise" or any other variant thereof are intended to cover non-exclusive inclusion, such that a process, method, article or device including a series of elements not only includes those elements but also includes other elements not expressly listed, or further includes elements inherent to such process, method, article or device. The terms "first", "second", etc. in this application are used to distinguish similar objects and not to describe a specific order or sequence.
[0028] To enable those skilled in the art of this technology to better understand the solution of this application, the following further detailed description of this application will be given in conjunction with the accompanying drawings and specific embodiments.
[0029] In combination with the specific application environment architecture or specific hardware architecture on which the execution of the storage method of the authentication toolkit depends, the specific application environment architecture or specific hardware architecture will be described herein.
[0030] The embodiment of the storage method of the authentication toolkit provided in the embodiments of this application can be executed in a server device or a similar computing device. Taking the operation on a server device as an example, Figure 1 is the hardware structure block diagram of the server device of a storage method of an authentication toolkit in the embodiments of this application. As Figure 1 shown, the server device may include one or more ( Figure 1 only one is shown in the figure) processors 102 (the processors 102 may include, but are not limited to, processing devices such as a microprocessor MCU or a programmable logic device FPGA) and a memory 104 for storing data. Among them, the above server device may further include a transmission device 106 for communication functions and an input / output device 108. Those of ordinary skill in the art can understand that Figure 1 the structure shown in the figure is only schematic and does not limit the structure of the above server device. For example, the server device may further include more or fewer components than Figure 1 shown in the figure, or have a different configuration from Figure 1 shown in the figure.
[0031] The memory 104 can be used to store computer programs, such as software programs and modules of application software, such as the computer program corresponding to the storage method of the authentication toolkit in the embodiments of the present application. The processor 102 executes various functional applications and data processing by running the computer program stored in the memory 104, that is, the above-mentioned method is implemented. The memory 104 may include a high-speed random access memory, and may also include a non-volatile memory, such as one or more magnetic storage devices, flash memories, or other non-volatile solid-state memories. In some instances, the memory 104 may further include a memory remotely provided with respect to the processor 102, and these remote memories can be connected to the server device through a network. Examples of the above-mentioned network include but are not limited to the Internet, an enterprise intranet, a local area network, a mobile communication network, and combinations thereof.
[0032] The transmission device 106 is used to receive or send data via a network. Specific examples of the above-mentioned network may include a wireless network provided by a communication provider of the server device. In one instance, the transmission device 106 includes a network adapter (abbreviated as NIC), which can be connected to other network devices through a base station so as to communicate with the Internet. In one instance, the transmission device 106 can be a radio frequency (abbreviated as RF) module, which is used to communicate with the Internet wirelessly.
[0033] Embodiments of the present application provide a storage method for an authentication toolkit, which is applied to a distributed storage system. The distributed storage system includes multiple nodes. In combination with the execution process of the storage method of the authentication toolkit, the method is described in detail. As Figure 2 shown, the method includes the following steps S200-220:
[0034] Step S200, when receiving the first toolkit uploaded by the first node, generate a first identifier corresponding to the first toolkit according to the file content included in the first toolkit.
[0035] Wherein, the first identifier is used to indicate the file content included in the first toolkit.
[0036] Specifically, the distributed storage system contains multiple nodes. When a certain node, such as the first node, uploads a toolkit, the system identifies the file content included in the first toolkit and generates a unique identifier (Content Identifier, CID) to identify the file content of this toolkit. The generation of the identifier is based on the file content of the toolkit and is achieved by calculating the hash value of the binary content. This hash value is irreversible, ensuring that the same file content will generate the same identifier. Even if the file names are different, content deduplication can be achieved in the system. The identifier mechanism is the core of the Inter Planetary File System (IPFS), which allows files to be located and retrieved by their content rather than the traditional method of file name or location.
[0037] Exemplarily, a hash function is used to calculate the file content of the toolkit to generate an identifier. The selection of the hash function should ensure the uniqueness and collision resistance of the identifier. After the first node uploads the toolkit, the system calls the hash function to calculate the hash value of the toolkit content and uses this value as the identifier.
[0038] Step S210, determine whether the first identifier exists in the preset identification table.
[0039] Among them, the identification table includes multiple identifiers, and the multiple identifiers respectively correspond to the file contents of multiple toolkits already stored in the distributed storage system.
[0040] Specifically, the identification table can be a Distributed Hash Table (DHT), which is an index table storing existing identifiers and their corresponding toolkit file contents, and can also record the mapping relationship between the identifier and the storage node. After the identifier is generated, the system needs to check whether the same identifier already exists in the identification table to avoid storing files with the same content repeatedly. The identification table (distributed hash table) is a commonly used technology to achieve this function, which allows distributed storage and query of identifier information without a centralized server.
[0041] Exemplarily, query the identification table to find out whether the identifier already exists in the system. Since the identification table is distributed, the query operation can be performed in parallel on multiple storage nodes to improve the query speed. If the identifier exists, it means that the file content has already been stored, and the system does not need to physically store the first toolkit anymore.
[0042] Step S220, when the first identifier exists in the identification table, establish an association relationship between the first identifier and the first node, and delete the first toolkit uploaded by the first node.
[0043] Specifically, if the identifier already exists in the identification table, the system will consider that the tool package uploaded by the first node is the same as the content of a certain stored tool package. At this time, the system establishes a logical association relationship (instead of physical storage) and records the first node as one of the backup nodes for this identifier. Subsequently, the system can safely delete the copy of the tool package uploaded by the first node because the content already exists on other nodes in the network. This approach reduces the occupancy of storage space and improves storage efficiency.
[0044] Exemplarily, record the information of the first node in the identification table as a potential backup node for the identifier without performing duplicate physical storage. Then, the copy of the tool package on the first node can be deleted, but the logical mapping of the identifier is retained.
[0045] In this embodiment, when receiving the first tool package uploaded by the first node, generate a first identifier corresponding to the first tool package according to the file content included in the first tool package, and use the identifier to indicate the file content. Even if the file names of multiple tool packages are different but the file contents are the same, the same identifier will still be generated. By introducing the identifier of the file content, the identification of the file content of the tool package is achieved. Then determine whether the first identifier exists in the preset identification table. If the first identifier exists in the identification table, establish an association relationship between the first identifier and the first node, and delete the first tool package uploaded by the first node to ensure that only new file content will be physically stored, avoiding duplicate storage. When the file content is repeated, logically establish an association relationship between the first identifier and the first node, and then delete the tool package uploaded by the first node, reducing the storage cost. Since the unique identifier of the tool package is generated and it is checked whether the identifier exists in the identification table, the technical problem of duplicate storage of tool packages with the same content in the distributed storage system is avoided, and the technical effect of saving storage resources is achieved.
[0046] In one embodiment, as Figure 3 shown, in step S200, when receiving the first tool package uploaded by the first node, generate a first identifier corresponding to the first tool package according to the file content included in the first tool package. It includes steps S300 - S320:
[0047] Step S300, parse the first tool package to obtain the file content included in the first tool package.
[0048] Among them, the file content is represented in binary form.
[0049] Specifically, the system processes the first tool package uploaded to the distributed storage system to ensure that it has a unique and accurate identifier for subsequent retrieval and deduplication. Specifically, the system receives a tool package upload request from the first node and then deeply analyzes the content of the tool package and converts it into binary form. The reason for using binary form is that it is the basic form for the computer to recognize and process data at the underlying level, which can accurately reflect the actual content of the file and avoid the influence of differences in file names, extensions, or other metadata. Subsequently, the system processes the binary file content using a preset hash algorithm to generate a hash value, which will be used as the unique identifier of the first tool package. The generation of the identifier is based on the binary content of the file, ensuring that tool packages with the same content can be recognized as the same item regardless of how their metadata changes.
[0050] Exemplarily, first read the original data of the first tool package and convert it into binary format. Next, select a hash algorithm to process the binary content to generate the hash value of the first tool package. Finally, encapsulate the hash value in the identifier format specified by IPFS to obtain the first identifier. The format of the identifier usually includes version information, the identifier of the multi-hash algorithm, and the generated hash value to ensure the uniqueness and resolvability of the identifier.
[0051] Step S310: Process the binary file content using a preset hash algorithm to obtain the hash value corresponding to the first tool package.
[0052] Specifically, in the process of generating the identifier, the calculation of the hash value is crucial. The preset hash algorithm encrypts and transforms the binary file content to produce an output of a fixed length, that is, the hash value. This hash value will be used for the generation of the identifier to ensure the uniqueness of the tool package. The selection of the hash algorithm should consider its collision resistance, that is, it is almost impossible for different file contents to generate the same hash value to avoid duplication and confusion of the identifier.
[0053] Step S320: Encapsulate the hash value corresponding to the first tool package in a preset format to obtain the first identifier.
[0054] Specifically, the identifier not only includes the hash value but may also include other meta-information, such as the hash algorithm identifier, identifier version information, etc. Encapsulating the hash value in the preset identifier format is to ensure the standardization and resolvability of the identifier, making it have a unified representation and processing method in the distributed storage system.
[0055] In this embodiment, the toolkit is converted into a binary form and hashed, ensuring that any minor differences in file content can be accurately captured and avoiding storage redundancy. A hash value is generated using a hash algorithm, which is a core component of the identifier. Its high-strength property provides a unique identifier for the content of the toolkit, enhancing the security of the system. Finally, the hash value is encapsulated as an identifier, ensuring the standardization and resolvability of the identifier format, facilitating communication and processing in the distributed storage system.
[0056] In one embodiment, as Figure 4 shown, after determining whether a first identifier exists in a preset identification table in step S210, the method further includes steps S400 - S430:
[0057] Step S400, in the case where the first identifier does not exist in the identification table, add the first identifier to the identification table to update the identification table, and broadcast the updated identification table to each node of the distributed storage system.
[0058] Specifically, when the system discovers that the newly uploaded first toolkit has no corresponding identifier (first identifier) in the identification table, this step S400 is executed. The identification table is a key data structure in the distributed storage system for recording and retrieving identifiers, which helps the system locate the storage location of a specific toolkit. In the case where the identifier does not exist in the identification table, the system needs to add the new identifier to the identification table to ensure that the file can be quickly searched and retrieved through the identifier in the future. The updated identification table needs to be broadcast to all nodes in the system to achieve data consistency and synchronization, ensuring that any node can query the latest identifier record.
[0059] Exemplarily, when the system detects that the identifier of the first toolkit is not in the identification table, it calls the update interface of the identification table to insert a new record into the identification table, which contains information such as the identifier, the upload node number, and file metadata. After the update is completed, the synchronization mechanism of the identification table is used to broadcast the update to all nodes in the entire network, ensuring that each node's copy of the identification table contains the latest identifier record. This broadcast process may include a data exchange protocol between nodes to minimize the network load caused by the broadcast.
[0060] Step S410, determine the shard byte count of the first toolkit according to the file parameters of the first toolkit and the current network parameters of the distributed storage system.
[0061] Specifically, considering the size of the toolkit and network conditions, the system needs to determine an appropriate shard byte count to achieve efficient storage and distribution. File parameters include the size, type, etc. of the file, while network parameters include the latency, bandwidth, packet loss rate, etc. between nodes. By evaluating these parameters, the system can decide on the shard size to ensure that the distribution of the toolkit is both fast and stable in the current network environment.
[0062] Exemplarily, the system design includes an intelligent sharding strategy module that determines the shard size based on file parameters and network parameters. For example, for large files such as image files, a larger shard byte count may be selected to reduce the number of shards. For an environment with high network latency, a smaller shard byte count may be selected to reduce the risk of timeout for the transmission of a single shard. In practical applications, it can be implemented through a dynamic calculation formula, considering parameters such as file size, network bandwidth, latency, etc., to dynamically adjust the shard byte count.
[0063] Step S420, perform sharding processing on the first toolkit according to the shard byte count to divide the first toolkit into multiple shard files, where the byte count of each shard file is less than or equal to the shard byte count.
[0064] Specifically, after determining the shard byte count, the system needs to split the first toolkit into multiple shard files, and the size of each shard does not exceed the set shard byte count. This process ensures that the toolkit can be efficiently stored and distributed, while reducing the storage pressure on a single node.
[0065] Exemplarily, read the file content of the first toolkit and perform cutting according to the determined shard byte count to generate multiple shard files. A unique identifier will be generated for each shard file during the cutting process for indexing and storage. The file cutting can adopt a sliding window algorithm, allowing the last shard to be smaller than the set shard byte count.
[0066] Step S430, store the multiple shard files.
[0067] Specifically, after sharding is completed, the system needs to select appropriate nodes to store these shard files. The storage strategy may include a redundancy factor, node priority, etc. to ensure the availability and reliability of the data.
[0068] Exemplarily, the system calls a storage scheduler to allocate the storage locations of the shard files according to the redundancy policy and node priority (such as low-latency nodes first). The identifier of each shard file and its storage node information are recorded in an identification table for subsequent query and data recovery.
[0069] In this embodiment, by adding a new identifier record in the identification table and broadcasting the update, the global uniqueness of the identifier and the consistency of data among nodes in the network are ensured. The fragmentation byte count is intelligently determined and dynamically adjusted according to the file size and network conditions, avoiding the storage and distribution efficiency problems caused by a fixed fragmentation size. The toolkit is divided into multiple fragments, reducing the storage pressure on a single node while improving the distribution speed and download efficiency. Finally, the fragmented files are stored on the optimal nodes, and fast positioning and access are achieved through the identification table index, improving the data availability and the overall system performance.
[0070] In one embodiment, as Figure 5 shown, step S410, determine the fragmentation byte count of the first toolkit according to the file parameters of the first toolkit and the current network parameters of the distributed storage system. It includes: steps S500 - S520:
[0071] Step S500, determine the file type to which the first toolkit belongs according to the file byte header of the first toolkit.
[0072] Specifically, the file byte header, usually called the "file signature", is a specific byte sequence at the beginning of the file, used to uniquely identify the file type. By parsing the file byte header, the system can automatically identify the type of the uploaded toolkit, providing a basis for setting subsequent fragmentation parameters.
[0073] Exemplarily, read the first 512 bytes of the first toolkit, use the binary reading method to parse the file header, and search for the specific file byte header.
[0074] Step S510, determine the initial fragmentation byte count according to the file type to which the first toolkit belongs.
[0075] Specifically, different types of files may be suitable for different fragmentation sizes. For example, image files, due to their size usually in the gigabyte (GB) level, are suitable for larger fragmentation byte counts to reduce the number of fragments. While text log files may be suitable for smaller fragmentation byte counts to improve the flexibility and storage efficiency of fragmentation.
[0076] Exemplarily, set a mapping table of file type and initial fragmentation byte count inside the system. For example, the image type corresponds to an initial fragmentation byte count of 512 megabytes (MB), the compressed package type corresponds to 128MB, and the text log type corresponds to 64MB. The system selects the corresponding initial fragmentation byte count from the mapping table according to the identified file type.
[0077] Step S520, adjust the initial fragmentation byte count according to the current network parameters of the distributed storage system to obtain the fragmentation byte count.
[0078] Specifically, network parameters (such as latency, bandwidth, packet loss rate, etc.) have an important impact on the selection of shard size. In a network environment with high latency and low bandwidth, a larger shard byte count may cause shard transmission timeouts. In an environment with low latency and high bandwidth, larger shard byte counts can be used to improve transmission efficiency. The system dynamically adjusts the initial shard byte count by evaluating network parameters in real time to ensure that the sharding strategy adapts to the current network environment and improves storage and distribution efficiency.
[0079] Exemplarily, design a dynamic shard adjustment algorithm. The input includes real-time network parameters (latency, bandwidth, packet loss rate) and the initial shard byte count, and the output is the adjusted shard byte count. The algorithm can include the calculation of a dynamic adjustment factor AF based on the network parameter weight formula: For example, AF = α * (1 / latency) + β * (bandwidth / reference bandwidth) + γ * (1 - packet loss rate), where AF is the dynamic adjustment factor, and α, β, and γ are preset weights, with α + β + γ = 1. The weights are set based on empirical values to determine whether to increase or decrease the shard byte count. For example, if AF ≥ 1.2, it is defined as a good network condition, and the shard byte count can be enlarged to a maximum of 1024 MB. If AF ≤ 0.8, it is defined as a poor network condition, and the shard byte count can be reduced to a minimum of 64 MB. If 0.8 < AF < 1.2, the initial shard byte count remains unchanged.
[0080] In this embodiment, first, the file type is identified by parsing the file byte header, providing basic information for subsequent sharding parameter settings. Then, the initial shard byte count is selected according to the file type, considering the special requirements of different file types for shard size. Finally, the shard size is dynamically adjusted by evaluating network parameters in real time to ensure that under the current network conditions, the sharding strategy can both improve transmission efficiency and avoid the risk of shard transmission timeouts.
[0081] In one embodiment, as Figure 6 shown, before step S520 of adjusting the initial shard byte count according to the current network parameters of the distributed storage system to obtain the shard byte count, the method further includes steps S600 - S610:
[0082] Step S600, in the case where the file type of the first toolkit cannot be identified based on the file byte header of the first toolkit, determine the file extension of the first toolkit.
[0083] Specifically, the file byte header is an important basis for file type recognition. However, some files may not have an obvious byte header flag, or the flag is not unique enough to accurately determine the file type. In this case, the system needs to further rely on the file extension for type recognition. The file extension is the suffix of the file name and usually indicates the format or type of the file. For example, docx indicates a Word document, and zip indicates a compressed file. Through the file extension, the system can obtain additional information about the file type to assist in setting the initial shard byte count.
[0084] Exemplarily, read the file name of the first toolkit and extract its extension. This step can be completed through string operations, such as using regular expressions or string splitting functions to split the file name and obtain the last part as the extension.
[0085] Step S610, determine the initial shard byte count according to the file extension.
[0086] Specifically, a mapping rule between the file extension and the initial shard byte count is preset in the system to handle the case where byte header recognition fails. The mapping rule of the file extension can be used as a supplementary method for file type recognition. Although it is not as accurate as the byte header, it can still provide sufficient information to make a reasonable sharding decision in most cases.
[0087] Exemplarily, design a mapping table to associate common file extensions with their recommended initial shard byte counts. For example, for an image file (with the extension.iso), the initial shard byte count can be set to 512MB in the mapping table. For a compressed file (with the extension.zip), it is set to 128MB. If the extension of the first toolkit is in the mapping table, the system will select the corresponding initial shard byte count according to the extension. If the extension is not in the mapping table, a default value (such as 256MB) can be set as the initial shard byte count.
[0088] In this embodiment, when the file byte header cannot provide sufficient information, the system provides another dimension of information about the file type through file extension recognition. This information can also be used to set the initial shard byte count to ensure more efficient subsequent sharding processing. Specifically, by extracting the file extension, a supplementary information source is provided when type recognition fails, increasing the system's adaptability to different types of toolkits. Determining the initial shard byte count according to the file extension can, even in the absence of byte header information, provide reasonable parameter settings for subsequent sharding processing through the extension mapping rule, avoiding problems such as low storage efficiency or network transmission timeouts caused by improper sharding parameter settings.
[0089] In one embodiment, as Figure 7As shown in the figure, in step S520, adjust the initial shard byte count according to the current network parameters of the distributed storage system to obtain the shard byte count. This includes steps S700 - S730:
[0090] Step S700, obtain the current network parameters of the distributed storage system.
[0091] Among them, the network parameters include network latency, network bandwidth, and network packet loss rate.
[0092] Specifically, the network environment can be monitored in real time to determine the current state of the network. Network latency refers to the time required for data to be transmitted over the network, which affects the immediacy and efficiency of data transmission. Network bandwidth refers to the maximum rate at which the network can transmit data, which determines the speed of data transmission. The network packet loss rate refers to the proportion of data packets lost during network transmission, which reflects the reliability and stability of the network. Obtaining these network parameters can help the system judge the current health status of the network, so as to make more reasonable decisions on sharded data transmission.
[0093] Step S710, determine the weights corresponding to network latency, network bandwidth, and network packet loss rate respectively according to the influence degrees of network latency, network bandwidth, and network packet loss rate on data transmission.
[0094] Specifically, the influence degrees of network parameters on data transmission are different. Therefore, when calculating the network state, different weights need to be assigned to each parameter. For example, high network latency may cause data transmission interruption or timeout, so its weight may be the highest. Network bandwidth directly affects the data transmission speed, and its weight is the second. The network packet loss rate affects the reliability of data transmission, and its weight is relatively low. The setting of weights should be based on the actual contributions of each parameter to data transmission quality to ensure the accuracy of network state evaluation.
[0095] Exemplarily, according to experience or test data analysis, set the weight coefficients of network latency, network bandwidth, and network packet loss rate to be α, β, and γ respectively, and α + β + γ = 1. For example, it can be set that α = 0.5, β = 0.3, γ = 0.2, indicating that network latency has the greatest impact on the network state, followed by network bandwidth, and finally network packet loss rate. These weights need to reflect the importance of various parameters to data transmission quality and efficiency and can be set during system initialization or in the configuration file.
[0096] Step S720, determine the current network state of the distributed storage system according to the values corresponding to network latency, network bandwidth, and network packet loss rate respectively and the weights corresponding to network latency, network bandwidth, and network packet loss rate respectively.
[0097] Among them, the current network state of the distributed storage system includes at least one of the following: good state, normal state, high - latency state.
[0098] Specifically, by combining the numerical values of network parameters with the set weights, the system can comprehensively evaluate the transmission capacity of the current network, thereby obtaining a network status level. These levels, such as good status, normal status, and high-latency status, can help the system select the most suitable shard byte count and other transmission strategies in different network environments.
[0099] Exemplarily, using the weighted summation method, the numerical values of network parameters are multiplied by the weights and then summed to obtain a network status index. If the network status index is higher than the preset threshold, it indicates that the network condition is good and the system is in a good state. If it is between two thresholds, it is regarded as a normal state. If it is lower than another preset threshold, it is determined to be in a high-latency state. The setting of these thresholds should be determined through testing and analysis in the initial stage of system design to ensure the accuracy of network status evaluation.
[0100] Step S730: Adjust the initial shard byte count according to the current network status of the distributed storage system to obtain the shard byte count.
[0101] Specifically, based on the network status, the system can adjust the shard byte count to optimize the data transmission process. In a good state, a larger shard byte count can be used to improve the parallelism and speed of data transmission. In a normal state, a medium shard byte count is used to balance the transmission speed and reliability. In a high-latency state, a smaller shard byte count is selected to ensure that the shard transmission does not time out.
[0102] Exemplarily, design a shard byte count adjustment algorithm with the input being the network status level and the output being the adjusted shard byte count. For example, if the network status level is in a good state, the shard byte count is set to the maximum value preset by the system (such as 1024 MB). If it is in a normal state, it is set to a medium value (such as 256 MB). If it is in a high-latency state, it is set to the minimum value (such as 64 MB).
[0103] In this embodiment, first, by real-time monitoring of network parameters, basic data is provided for subsequent network status evaluation. Then, the weights of each network parameter are set to ensure the scientificity and accuracy of network status evaluation. The network status level is calculated based on the actual values and weights of network parameters, providing an intuitive network health indicator. Finally, the initial shard byte count is adjusted according to the network status level to ensure that data transmission can be carried out efficiently, stably, and reliably in different network environments.
[0104] In one embodiment, as Figure 8 shown, step S730: Adjust the initial shard byte count according to the current network status of the distributed storage system to obtain the shard byte count. It includes steps S800 - S820:
[0105] Step S800, when it is determined that the current network state of the distributed storage system is in a good state, adjust the initial shard byte count to increase and obtain the shard byte count.
[0106] Among them, the shard byte count is less than the upper limit threshold.
[0107] Specifically, when the system detects that the current network environment is in a good state, that is, low latency, high bandwidth, and low packet loss rate, the initial shard byte count can be appropriately increased, but it is necessary to ensure that it does not exceed the preset upper limit threshold. The purpose of doing this is to utilize the high-quality network conditions to improve the efficiency and parallelism of data distribution, while avoiding problems with storage and retrieval efficiency caused by overly large shards.
[0108] Exemplarily, set a shard byte count adjustment function, with the input being the network state level (good state) and the output being the adjusted shard byte count. For example, if the initial shard byte count is 256MB, in a good network state, the shard byte count can be increased to the maximum value preset by the system (such as 1024MB), but not exceeding the upper limit threshold. This adjustment process can be achieved through simple mathematical operations, such as shard byte count = min(initial shard byte count * increase ratio, upper limit threshold).
[0109] Step S810, when it is determined that the current network state of the distributed storage system is in a high-latency state, adjust the initial shard byte count to decrease and obtain the shard byte count.
[0110] Among them, the shard byte count is greater than the lower limit threshold.
[0111] Specifically, in a poor network environment, especially under high latency conditions, the shard byte count should be appropriately reduced, but it needs to remain above the set lower limit threshold to avoid excessive additional transmission overhead caused by overly small shards. A smaller shard byte count helps reduce the risk of data transmission timeouts in the case of high network latency, ensuring data integrity and transmission efficiency.
[0112] Exemplarily, similarly, set a shard byte count adjustment function, with the input being the network state level (high-latency state) and the output being the adjusted shard byte count. For example, in a high-latency state, reduce the shard byte count to the minimum value preset by the system (such as 64MB), ensuring that the shard byte count is not lower than the lower limit threshold. This process can also be achieved through simple mathematical operations, such as shard byte count = max(initial shard byte count / reduction ratio, lower limit threshold).
[0113] Step S820, when it is determined that the current network state of the distributed storage system is in a normal state, use the initial shard byte count as the shard byte count.
[0114] Specifically, when the network status is neither particularly good nor particularly bad and is in a normal state, the system should keep the number of shard bytes unchanged and use the initially set number of shard bytes for data distribution. This can balance the data transmission efficiency and storage efficiency, and avoid over-optimization or conservatism in the dynamic adjustment strategy, resulting in a decrease in the actual transmission efficiency.
[0115] Exemplarily, directly assign the initial number of shard bytes to the number of shard bytes, that is, the number of shard bytes = the initial number of shard bytes. There is no need to adjust the sharding strategy in the normal network state, maintaining the stability and efficiency of data distribution.
[0116] In this embodiment, in the case of good network conditions, by increasing the number of shard bytes, the parallelism and transmission speed of data distribution are improved, while ensuring that the shard size is within a reasonable range, avoiding a decrease in storage and retrieval efficiency. In a high-latency network environment, by reducing the number of shard bytes, the risk of data transmission timeout is reduced, ensuring the integrity and security of data, while avoiding an increase in network overhead caused by too small shards. In the normal network state, keeping the number of shard bytes unchanged achieves a balance between data transmission efficiency and storage efficiency, avoiding performance fluctuations caused by excessive strategy adjustment. An intelligent shard byte adjustment strategy based on real-time network status is implemented to optimize data transmission and storage efficiency.
[0117] In one embodiment, as Figure 9 shown, step S430, store multiple shard files. It includes: steps S900 - S940:
[0118] Step S900, determine the network latency corresponding to each of the multiple nodes of the distributed storage system.
[0119] Specifically, network latency refers to the average time required for a data packet to be transmitted in the network and is a key indicator for measuring data transmission efficiency. For a distributed storage system, obtaining the network latency of each node helps determine which nodes are more suitable for storing data shards, because nodes with low latency can respond to data requests and shard downloads faster, improving the data transmission speed. The determination of network latency is usually achieved by sending test data packets to each node and recording the round-trip time.
[0120] Exemplarily, design a network latency monitoring module that periodically sends test data packets to each node in the system and records the response time of each node, calculating the average round-trip time as the network latency of that node. For example, built-in network measurement tools can be used to send multiple data packets to each node, record the response time of each data packet, and finally calculate the average value to obtain the network latency. This process can be integrated into the system management background as a continuous network health monitoring task.
[0121] Step S910: Determine a first node set from multiple nodes according to the network latency corresponding to each of the multiple nodes.
[0122] Among them, the network latency of the nodes in the first node set is less than the latency threshold.
[0123] Specifically, the latency threshold is a benchmark value set by the system for screening out a node set with lower network latency. The first node set consists of those nodes whose network latency is lower than the preset threshold. These nodes are considered ideal candidates for storing data shards because they can quickly respond to data requests and provide high-speed shard download services.
[0124] Exemplarily, set a latency threshold, such as 100 milliseconds, and then traverse the latency data of all nodes to screen out the nodes with a latency lower than 100 milliseconds to form the first node set.
[0125] Step S920: When the number of nodes included in the first node set is less than the preset number, determine a second node set from multiple nodes in the distributed storage system.
[0126] Among them, the network latency of the nodes in the second node set is less than the latency threshold, and the distance between the nodes in the second node set and the first node is less than the distance threshold.
[0127] Specifically, the preset number is the minimum number of storage nodes set by the system to ensure the high availability and redundancy of data shards. If the number of nodes meeting the conditions in the first node set is insufficient, the system will further expand the search scope and consider the nodes physically closer to the first node (i.e., the nodes with a distance less than the distance threshold) to form the second node set. This approach can improve the geographical diversity of data distribution and reduce the impact of a single regional failure.
[0128] Exemplarily, if the number of nodes in the first node set is less than the preset number (such as 5), then continue to screen out the nodes whose distance from any node in the first node set is less than a specific distance threshold (e.g., 1000 kilometers, depending on the global layout strategy of the system), and the network latency of these nodes still needs to be lower than the latency threshold, and finally form the second node set.
[0129] Step S930: When the number of nodes included in the second node set is less than the preset number, determine a third node set from multiple nodes in the distributed storage system.
[0130] Among them, the network latency of the nodes in the third node set is less than the latency threshold, the distance between the nodes in the third node set and the first node is less than the distance threshold, and the load of the nodes in the third node set is less than the load threshold.
[0131] Specifically, the load threshold is a standard for measuring the current processing capacity of a node. If the resource utilization rate of a node is too high, it may affect the efficiency of data storage and transmission. The third node set is a node set that further considers the node load on the basis of the first two steps of screening, ensuring that data shards are stored not only on nodes with fast response speed and close distance, but also on nodes with sufficient resources and strong processing capabilities.
[0132] Exemplarily, if the second node set still cannot meet the preset quantity, continue to screen among all nodes. In addition to meeting the network latency and distance conditions, it is also necessary to meet that the load indicators such as the CPU utilization rate and memory occupancy rate of the node are lower than the set load threshold, and finally form the third node set. For example, set the load threshold of the node to be that the CPU utilization rate is lower than 70% and the memory occupancy rate is lower than 80%.
[0133] Step S940, distribute and store multiple shard files in at least one first target node.
[0134] Among them, the first target node is a node in the third node set.
[0135] Specifically, after the above screening process, the system stores data shards on the first target nodes. These nodes not only have low network latency and close distance, but also have a load at a safe level, which is the best choice for storing data shards. In this way, the system can ensure fast data access while maintaining data redundancy and high availability, and reducing the risk of data loss and access failure.
[0136] Exemplarily, the storage node priority: low-latency node > same-region node > low-load node.
[0137] Exemplarily, the system sends data shards to the nodes in the third node set for storage. Each node stores a part of the shard file, and through the distributed storage mechanism of IPFS, it is ensured that each shard has multiple copies, meeting the requirements of redundant storage. For example, a distributed hash table (identifier table) can be used to record the information of multiple storage nodes corresponding to each identifier (content identifier), ensuring that even if some nodes fail, the data can still be obtained from other nodes.
[0138] In this embodiment, first, by monitoring network latency in real time, key data is provided for subsequent node screening, ensuring that only nodes with fast response speeds are selected as storage targets, thereby improving the immediacy of data access and transmission efficiency. Second, based on the latency threshold, the first node set is screened out, realizing the preliminary positioning of data shard storage and ensuring the basic fast access ability of data shards. When the number of the first node set is insufficient, geographical distance is further considered to ensure the geographical diversity of data distribution, improve data redundancy and fault tolerance, and at the same time promote the efficient cross-regional distribution of data. After screening again considering node load, the third node set is formed, ensuring that data shards are stored not only on nodes with low latency and reasonable geographical distribution, but also on nodes with sufficient processing capacity, improving the overall efficiency of data storage and transmission. Distributing data shards to the nodes in the third node set realizes the optimal strategy for data storage, ensures the high availability and fast access of data, and at the same time, through node load monitoring, avoids performance bottlenecks in the data distribution process, improving the overall response speed and stability of the system. An intelligent data shard storage strategy based on network latency, geographical distance, and node load is realized, improving data storage efficiency, distribution speed, and system stability.
[0139] In one embodiment, as Figure 10 shown, the method further includes: steps S1000 - S1040:
[0140] Step S1000, receive a tool download request from a target object.
[0141] Among them, the tool download request carries identification information of the requested tool package.
[0142] Specifically, in a distributed storage system, a tool download request from a target object (such as a user or an application) usually contains identification information of a specific tool package, such as an identifier (content identifier). The identifier is a unique identifier generated for the file content through a hash algorithm, and the user requests the file from the system through the identifier, ensuring the accuracy and repeatability of the request.
[0143] Exemplarily, design a request processing module on the server side to listen for and receive network requests. When receiving a tool download request, parse the information in the request and extract identification information such as the identifier.
[0144] Step S1010, according to the identification information, determine the target identifier corresponding to the requested tool package in the identification table.
[0145] Specifically, the identification table is a data structure maintained by the distributed storage system for recording the mapping relationship between identifiers and actual storage nodes. By querying the identification table, the system can quickly locate the node storing the data shard of a specific identifier, thus accelerating the data distribution process.
[0146] Exemplarily, design an identification query function with the input being an identifier to find the information of the data shard storage node corresponding to the identifier from the identification table. The identification table can be a distributed hash table (identification table), and each node stores a part of the table data. Through the identification table query mechanism, the storage node corresponding to the identifier can be quickly found.
[0147] Step S1020, determine the target toolkit corresponding to the target identifier according to the file content indicated by the target identifier.
[0148] Specifically, the target identifier is not only used to locate the storage node but also to uniquely identify the file content. Through the identifier, the system can ensure that the downloaded file content exactly matches the requested toolkit, avoiding problems such as file version errors or content tampering.
[0149] Step S1030, determine at least one second target node among multiple nodes of the distributed storage system that stores the target toolkit and whose network status meets the standard.
[0150] Specifically, the network status meeting the standard means that parameters such as the network latency, bandwidth, and packet loss rate of the node are within the preset threshold range, ensuring the efficiency and stability of data transmission. By screening the nodes with network status meeting the standard, the system can optimize the data distribution path and improve the user download speed and data integrity.
[0151] Exemplarily, a screening function can be used. This function receives the node list and the network status threshold as inputs and returns a list of nodes whose network status meets the standard and store the shards of the target identifier. For example, traverse all storage nodes, check whether the network latency of each node is lower than 100 milliseconds, the bandwidth is higher than 50 Mbps, and the packet loss rate is lower than 1%, and combine the identifier storage information to screen out the qualified nodes.
[0152] Step S1040, download the target toolkit from at least one second target node and feedback it to the target object.
[0153] Specifically, once the nodes with network status meeting the standard and storing the data shards are determined, the system will download all shards from these nodes, recombine them, and deliver the complete toolkit to the target object. This process makes full use of the characteristics of the distributed network, downloads in parallel from multiple source nodes, improves the download speed, and provides additional security through data redundancy.
[0154] In this embodiment, receiving an accurate tool download request through an identifier ensures the accurate positioning and repeatability of the data request. Querying the identification table quickly locates the storage nodes, accelerating the data distribution process. Verifying the file content corresponding to the identifier guarantees the integrity and security of the data. Screening the nodes with qualified network status optimizes the data transmission path, improving the data transmission speed and stability. Finally, downloading and recombining data shards in parallel realizes the high-speed distribution of the tool package, and at the same time ensures the data security through data redundancy. The content addressing based on the identifier and the intelligent data distribution strategy based on the network status are realized, significantly improving the efficiency of data distribution and the user access experience.
[0155] In one embodiment, as Figure 11 shown, step S1040, downloading the target tool package from at least one second target node and feeding it back to the target object. It includes: steps S1100 - S1150:
[0156] Step S1100, downloading multiple shard files of the target tool package from at least one second target node.
[0157] Specifically, in a distributed storage system, a complete tool package is sliced into multiple shard files, and each shard file is stored on different nodes in the network. When the system receives a user request, it will download these shard files in parallel from the storage nodes identified as the second target nodes to improve the download speed.
[0158] Step S1110, recombining the multiple shard files of the target tool package to obtain the complete file of the target tool package.
[0159] Specifically, all the downloaded shard files are associated by the identifier and need to be recombined on the client side or the server side to be restored to the original complete tool package file. The recombination process ensures the integrity of the file and is a key step in data recovery.
[0160] Step S1120, using a preset hash verification algorithm to determine the identifier of the complete file of the target tool package.
[0161] Specifically, the hash verification algorithm is used to verify the integrity and consistency of the file. By calculating the hash value of the file and comparing it with the original identifier, it is ensured that the downloaded file content has not been tampered with or damaged.
[0162] Step S1130, when the identifier of the complete file of the target tool package is the same as the target identifier, determining that the complete file of the target tool package passes the hash verification.
[0163] Specifically, if the identifier of the complete file matches the target identifier carried in the download request, it indicates that there are no errors in the download and recombination processes, and the file content is complete and has not been tampered with.
[0164] Step S1140, verify the digital certificate of the target toolkit.
[0165] Specifically, the digital certificate is used to verify the legality and source of the toolkit. Verifying the digital certificate of the target toolkit can ensure that the downloaded toolkit is not only complete in content but also from an authenticated operating system vendor.
[0166] Step S1150, when the complete file of the target toolkit passes the hash verification and the digital certificate verification of the target toolkit passes, feedback the target toolkit to the target object.
[0167] Specifically, only when the file passes the hash verification and the digital certificate verification will the system feedback the toolkit to the user or application program to ensure the integrity of its content and the legality of its source.
[0168] In this embodiment, by parallel downloading the shards of the toolkit, the advantages of the distributed network are fully utilized, the download speed is increased, and the user waiting time is reduced. Then the shard files are recombined to restore the original complete toolkit, providing a basis for subsequent verification. Through hash verification, the integrity of the file content is verified, the risk of data tampering or damage is avoided, and the data security is ensured. In addition, through digital certificate verification, the legal source of the toolkit is further confirmed, improving the security of the system and the user trust. Finally, after all verifications pass, the complete and secure toolkit is feedback to the target object, realizing the secure transmission and high-efficiency delivery of data. Ensuring the integrity and security of downloading and recombining the toolkit from the distributed storage network, while improving the download efficiency.
[0169] In one embodiment, as Figure 12 shown, the method further includes: steps S1200 - S1240:
[0170] Step S1200, when it is detected that at least one toolkit has a version update, download the updated toolkit from the corresponding official website.
[0171] Specifically, version update detection is a key step to ensure that the toolkits in the system are always in the latest state. When a new version of the toolkit is detected, the system should download the latest toolkit from the official website to replace the old version. This process can be automatically performed regularly or manually started when triggered by the user.
[0172] Step S1210, compare the file content of the updated toolkit with the file content of the corresponding old version toolkit to be updated, and determine the differential file content.
[0173] Specifically, file content comparison is an important part of version update. By comparing the binary content of the old and new version toolkits, the different parts, i.e., the differential file content, are identified. These differential contents will be used to create an incremental package instead of uploading the entire new version toolkit again, saving storage space and transmission bandwidth.
[0174] Step S1220, retain the differential file content in the update toolkit to obtain an incremental package and generate a corresponding incremental identifier based on the differential file content.
[0175] Specifically, the incremental package is composed of the differential file content. Retaining this part of the content and creating an incremental package is to ensure that only the incremental package needs to be transmitted during subsequent toolkit version updates, reducing the burden of data transmission. The incremental identifier is the unique identifier of the incremental package and is used for subsequent version management and data retrieval.
[0176] Step S1230, add the incremental identifier to the identification table and store the incremental package.
[0177] Specifically, the identification table is used to record the identifier information of all toolkits and their versions. Adding the incremental identifier means that the system can track the incremental updates of the new version. Storing the incremental package in a distributed storage system is a prerequisite for achieving data redundancy and efficient distribution.
[0178] Exemplarily, update the identification table database and insert metadata such as the identifier of the incremental package, the identifier of the associated old version, the update time, and the version information. Then, use the distributed storage mechanism to upload the incremental package to multiple nodes to ensure the high availability and fault tolerance of the data.
[0179] Step S1240, establish an association relationship between the incremental identifier and the identifier of the file content of the old version toolkit.
[0180] Specifically, establishing the identifier association is the core of version update. Through the association relationship, the system can track the dependencies and changes between the old and new versions, facilitating subsequent incremental updates and version backtracking. That is, the identifier of a new version is associated with the identifier of the old version and the identifier of the corresponding incremental package, and the old version toolkit plus the incremental package equals the new version toolkit.
[0181] Exemplarily, implement a version chain construction function in the system. The input is the identifier of the old version and the identifier of the incremental package, and the output is the version dependency relationship. For example, in the Merkle Directed Acyclic Graph (Merkle DAG), use the identifier of the incremental package as an edge to connect to the node corresponding to the identifier of the old version to form a version chain, recording the changes and dependencies between versions.
[0182] In this embodiment, by regularly detecting version updates, it is ensured that the toolkit is always in the latest state, meeting the continuous evolution of the system and the changing needs of users. Through meticulous comparison of file contents, the differential file contents are accurately determined, laying the foundation for subsequent incremental package creation, avoiding unnecessary full updates, and saving storage space and network resources. Then, the differential file contents are converted into an incremental package and an identifier is generated, achieving lightweight storage and transmission of the new version and improving the data distribution efficiency. By registering the incremental package identifier in the identification table and storing the incremental package, a complete version update data chain is established, ensuring high availability and redundancy of the data, and improving the stability and reliability of the system. Finally, by constructing the identifier association relationship, the dependency chain between the old and new versions is clarified, enabling the system to quickly locate and download the incremental package for efficient version update when needed. At the same time, it also supports version rollback and historical version management, enhancing the flexibility of the system and the user experience. An efficient and flexible toolkit version update strategy is implemented, significantly improving the efficiency of version update and the utilization rate of storage space.
[0183] In one embodiment, when an identifier is found to exist in the identification table, the physical storage is skipped, and only the association relationship between the first node and the existing first identifier is established at the logical layer.
[0184] Specifically, if the identification table query result indicates that the identifier already exists, it means that the file content has been stored on at least one node. At this time, the system will not repeat the physical storage operation, but only establish the mapping relationship between the current administrator and the identifier at the logical level, indicating that the current node can be used as a "potential backup node" for the file. This can avoid duplicate storage and save storage resources.
[0185] Exemplarily, a logical mapping function is implemented in the system. The input is the identification information of the current administrator and the identifier, and the output is the updated logical mapping table. For example, the identifier information is recorded in the metadata of the administrator, indicating that when the original storage node of the identifier fails, the current node can provide the file content as a backup node, but no new physical file copy will be created on the current node.
[0186] Mark the current node as a "potential backup node" and update the logical mapping information according to the redundancy policy and network requirements.
[0187] Specifically, the concept of "potential backup nodes" enhances the data redundancy and recovery capabilities of the system. By marking the current nodes, the system can quickly restore the file content from these nodes when needed in the future. However, whether to actually store redundant backups depends on the redundancy strategy set by the system and the actual requirements of the current network. Exemplarily, a redundancy strategy determination module is developed, which decides whether to create a physical backup on the current node based on system policies (such as redundancy factor, network latency) and the current network state. For example, if the redundancy factor is set to 5 and there are only 4 physical backups of the current file, the system will automatically create a physical backup on the current node. Conversely, if the redundancy conditions are already met, the system only updates the logical mapping information, marks the current node as a potential backup node, and does not perform actual physical storage.
[0188] When a storage node in the distributed storage system fails (denoted as the failed node), query the corresponding identifier (denoted as the failure identifier) according to the file content of the toolkit of the failed node. If the file content corresponding to the failure identifier is not stored in the distributed storage system, query through the identification table whether there is a potential backup node storing the file content corresponding to the failure identifier. If there is a potential backup node storing the file content corresponding to the failure identifier, call the file content corresponding to the failure identifier from the potential backup node to restore the file content of the toolkit of the failed node.
[0189] In this embodiment, by establishing a mapping relationship only at the logical level and marking the current node as a potential backup node, both the redundancy of the data is enhanced and unnecessary physical storage is avoided, improving the storage efficiency. Then, according to the redundancy strategy and network requirements, it is intelligently decided whether to create a physical backup on the current node, realizing the dynamic adjustment of data redundancy and ensuring that the system can maintain high availability and data robustness in various network environments. An intelligent and efficient file backup and redundancy strategy is proposed, significantly optimizing the data storage and recovery processes in the distributed storage system.
[0190] Through the description of the above embodiments, those skilled in the art can clearly understand that the method according to the above embodiments can be implemented by means of software plus a necessary general hardware platform. Of course, it can also be implemented by hardware, but in many cases the former is a better implementation.
[0191] An embodiment of the present application also provides a storage device for an authentication toolkit. Figure 13 It is a structural block diagram of a storage device for an authentication toolkit according to an embodiment of the present application. The device includes:
[0192] An identifier generation module 1301, configured to generate a first identifier corresponding to the first toolkit according to the file content included in the first toolkit when receiving the first toolkit uploaded by the first node, where the first identifier is used to indicate the file content included in the first toolkit.
[0193] An identifier query module 1302, configured to determine whether the first identifier exists in a preset identification table, where the identification table includes multiple identifiers, and the multiple identifiers respectively correspond to the file contents of multiple toolkits already stored in the distributed storage system.
[0194] An association module 1303, configured to establish an association relationship between the first identifier and the first node and delete the first toolkit uploaded by the first node when the first identifier exists in the identification table.
[0195] In an exemplary embodiment, the identifier generation module 1301 is further configured to parse the first toolkit to obtain the file content included in the first toolkit, where the file content is represented in binary form. Process the file content in binary form using a preset hash algorithm to obtain a hash value corresponding to the first toolkit. Package the hash value corresponding to the first toolkit in a preset format to obtain the first identifier.
[0196] In an exemplary embodiment, the apparatus further includes:
[0197] An update module, configured to add the first identifier to the identification table to update the identification table and broadcast the updated identification table to each node of the distributed storage system when the first identifier does not exist in the identification table.
[0198] And a sharding determination module, configured to determine the sharding byte count of the first toolkit according to the file parameters of the first toolkit and the current network parameters of the distributed storage system.
[0199] A sharding module, configured to perform sharding processing on the first toolkit according to the sharding byte count to divide the first toolkit into multiple sharded files, where the byte count of each sharded file is less than or equal to the sharding byte count.
[0200] A storage module, configured to store the multiple sharded files.
[0201] In an exemplary embodiment, the sharding determination module is further configured to determine the file type to which the first toolkit belongs according to the file byte header of the first toolkit. Determine the initial sharding byte count according to the file type to which the first toolkit belongs. Adjust the initial sharding byte count according to the current network parameters of the distributed storage system to obtain the sharding byte count.
[0202] In an exemplary embodiment, the apparatus further includes:
[0203] An extension name determination module, configured to determine the file extension of the first toolkit in case the file type to which the first toolkit belongs cannot be recognized according to the file byte header of the first toolkit.
[0204] An initial sharding module, configured to determine the initial sharding byte count according to the file extension.
[0205] In an exemplary embodiment, the sharding determination module is further configured to obtain the current network parameters of the distributed storage system, where the network parameters include network latency, network bandwidth, and network packet loss rate. Determine the weights corresponding to the network latency, network bandwidth, and network packet loss rate respectively according to the influence degrees of the network latency, network bandwidth, and network packet loss rate on data transmission respectively. Determine the current network state of the distributed storage system according to the values corresponding to the network latency, network bandwidth, and network packet loss rate respectively and the weights corresponding to the network latency, network bandwidth, and network packet loss rate respectively, where the current network state of the distributed storage system includes at least one of the following: good state, normal state, high latency state. Adjust the initial sharding byte count according to the current network state of the distributed storage system to obtain the sharding byte count.
[0206] In an exemplary embodiment, the sharding determination module is further configured to, in case it is determined that the current network state of the distributed storage system is in a good state, adjust the initial sharding byte count to increase to obtain the sharding byte count, where the sharding byte count is less than the upper limit threshold. In case it is determined that the current network state of the distributed storage system is in a high latency state, adjust the initial sharding byte count to decrease to obtain the sharding byte count, where the sharding byte count is greater than the lower limit threshold. In case it is determined that the current network state of the distributed storage system is in a normal state, use the initial sharding byte count as the sharding byte count.
[0207] In an exemplary embodiment, the storage module is further configured to determine the network delays respectively corresponding to multiple nodes of the distributed storage system. According to the network delays respectively corresponding to the multiple nodes, a first node set is determined among the multiple nodes, wherein the network delays of the nodes in the first node set are less than a delay threshold. In the case where the number of nodes included in the first node set is less than a preset number, a second node set is determined among the multiple nodes of the distributed storage system, wherein the network delays of the nodes in the second node set are less than the delay threshold, and the distance between the nodes in the second node set and the first node is less than a distance threshold. In the case where the number of nodes included in the second node set is less than the preset number, a third node set is determined among the multiple nodes of the distributed storage system, wherein the network delays of the nodes in the third node set are less than the delay threshold, the distance between the nodes in the third node set and the first node is less than the distance threshold, and the load of the nodes in the third node set is less than a load threshold. The multiple sharded files are distributed and stored in at least one first target node, wherein the first target node is a node in the third node set.
[0208] In an exemplary embodiment, the apparatus further includes:
[0209] A request receiving module, configured to receive a tool download request of a target object, wherein the tool download request carries identification information of a requested tool package.
[0210] An identifier determining module, configured to determine a target identifier corresponding to the requested tool package in an identification table according to the identification information.
[0211] A tool package determining module, configured to determine a target tool package corresponding to the target identifier according to the file content indicated by the target identifier.
[0212] A node determining module, configured to determine at least one second target node that stores the target tool package and has a qualified network status among the multiple nodes of the distributed storage system.
[0213] A download module, configured to download the target tool package from at least one second target node and feedback it to the target object.
[0214] In an exemplary embodiment, the download module is further configured to download multiple sharded files of the target tool package from at least one second target node. Recombine the multiple sharded files of the target tool package to obtain a complete file of the target tool package. Use a preset hash verification algorithm to determine an identifier of the complete file of the target tool package. In the case where the identifier of the complete file of the target tool package is the same as the target identifier, it is determined that the complete file of the target tool package passes the hash verification. Verify the digital certificate of the target tool package. In the case where the complete file of the target tool package passes the hash verification and the digital certificate of the target tool package passes the verification, the target tool package is feedback to the target object.
[0215] In an exemplary embodiment, the device further includes:
[0216] An update module, configured to download an updated toolkit from the corresponding official website when it is detected that at least one toolkit has a version update.
[0217] A comparison module, configured to compare the file content of the updated toolkit with the file content of the corresponding old version toolkit to be updated, and determine the differential file content.
[0218] An incremental retention module, configured to retain the differential file content in the updated toolkit, obtain an incremental package, and generate a corresponding incremental identifier according to the differential file content.
[0219] An incremental storage module, configured to add the incremental identifier to an identification table and store the incremental package.
[0220] An establishment module, configured to establish an association relationship between the incremental identifier and the identifier of the file content of the old version toolkit.
[0221] For the description of the features in the corresponding embodiment of the storage device of the authentication toolkit, reference can be made to the relevant description in the corresponding embodiment of the storage method of the authentication toolkit, which will not be elaborated here one by one.
[0222] An embodiment of the present application further provides an electronic device, including a memory and a processor. A computer program is stored in the memory, and the processor is configured to run the computer program to execute the steps in any of the above embodiments of the storage method of the authentication toolkit.
[0223] An embodiment of the present application further provides a computer-readable storage medium, in which a computer program is stored. The computer program is configured to execute the steps in any of the above embodiments of the storage method of the authentication toolkit when running.
[0224] In an exemplary embodiment, the above computer-readable storage medium may include, but is not limited to: USB flash drives, read-only memories (ROMs), random access memories (RAMs), external hard drives, magnetic disks, or optical discs, and other various media that can store computer programs.
[0225] An embodiment of the present application further provides a computer program product. The above computer program product includes a computer program, and when the computer program is executed by a processor, the steps in any of the above embodiments of the storage method of the authentication toolkit are implemented.
[0226] Embodiments of the present application also provide another computer program product, including a non-volatile computer-readable storage medium storing a computer program, where the computer program, when executed by a processor, implements the steps in any of the above-described embodiments of the storage method of the authentication toolkit.
[0227] Those skilled in the art can further realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, computer software, or a combination of the two. To clearly illustrate the interchangeability of hardware and software, the components and steps of each example have been generally described according to functions in the above description. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Skilled professionals can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present application.
[0228] The above has introduced in detail a storage method, apparatus, electronic device, computer-readable storage medium, and computer program product of an authentication toolkit provided by the present application. Specific examples are used herein to elaborate on the principles and implementation manners of the present application. The description of the above embodiments is only used to help understand the method and its core idea of the present application. It should be noted that for those of ordinary skill in the art in the technical field, without departing from the principle of the present application, several improvements and modifications can be made to the present application, and these improvements and modifications also fall within the protection scope of the claims of the present application.
Claims
1. A storage method for an authentication toolkit, characterized in that, Applied to a distributed storage system, the distributed storage system includes multiple nodes, and the method includes: When receiving a first toolkit uploaded by a first node, generating a first identifier corresponding to the first toolkit according to the file content included in the first toolkit, where the first identifier is used to indicate the file content included in the first toolkit; Determining whether the first identifier exists in a preset identification table, where the identification table includes multiple identifiers, and the multiple identifiers respectively correspond to the file contents of multiple toolkits already stored in the distributed storage system; When the first identifier exists in the identification table, establishing an association relationship between the first identifier and the first node, and deleting the first toolkit uploaded by the first node.
2. The storage method of the authentication toolkit according to claim 1, wherein, The step of, when receiving a first toolkit uploaded by a first node, generating a first identifier corresponding to the first toolkit according to the file content included in the first toolkit, includes: Parsing the first toolkit to obtain the file content included in the first toolkit, where the file content is represented in binary form; Processing the file content in binary form using a preset hash algorithm to obtain a hash value corresponding to the first toolkit; Encapsulating the hash value corresponding to the first toolkit in a preset format to obtain the first identifier.
3. The storage method of the authentication toolkit according to claim 1, wherein After determining whether the first identifier exists in the preset identification table, the method further includes: When the first identifier does not exist in the identification table, adding the first identifier to the identification table to update the identification table, and broadcasting the updated identification table to each node of the distributed storage system; And determining the sharding byte count of the first toolkit according to the file parameters of the first toolkit and the current network parameters of the distributed storage system; Performing sharding processing on the first toolkit according to the sharding byte count to divide the first toolkit into multiple sharded files, where the byte count of each sharded file is less than or equal to the sharding byte count; Storing the multiple sharded files.
4. The storage method of the authentication toolkit according to claim 3, characterized in that The step of determining the sharding byte count of the first toolkit according to the file parameters of the first toolkit and the current network parameters of the distributed storage system, includes: Determining the file type to which the first toolkit belongs according to the file byte header of the first toolkit; Determining an initial sharding byte count according to the file type to which the first toolkit belongs; Adjusting the initial sharding byte count according to the current network parameters of the distributed storage system to obtain the sharding byte count.
5. The storage method of the authentication toolkit according to claim 4, characterized in that Before adjusting the initial sharding byte count according to the current network parameters of the distributed storage system to obtain the sharding byte count, the method further includes: When the file type to which the first toolkit belongs cannot be identified according to the file byte header of the first toolkit, determining the file extension of the first toolkit; Determining the initial sharding byte count according to the file extension.
6. The storage method of the authentication toolkit according to claim 4 or 5, characterized in that Adjusting the initial shard byte count according to the current network parameters of the distributed storage system to obtain the shard byte count includes: Obtaining the current network parameters of the distributed storage system, where the network parameters include network latency, network bandwidth, and network packet loss rate; Determining the weights corresponding to the network latency, the network bandwidth, and the network packet loss rate respectively according to the influence degrees of the network latency, the network bandwidth, and the network packet loss rate on data transmission; Determining the current network state of the distributed storage system according to the values corresponding to the network latency, the network bandwidth, and the network packet loss rate respectively and the weights corresponding to the network latency, the network bandwidth, and the network packet loss rate respectively, where the current network state of the distributed storage system includes at least one of the following: good state, normal state, high-latency state; Adjusting the initial shard byte count according to the current network state of the distributed storage system to obtain the shard byte count.
7. The storage method of the authentication toolkit according to claim 6, wherein Adjusting the initial shard byte count according to the current network state of the distributed storage system to obtain the shard byte count includes: In the case where it is determined that the current network state of the distributed storage system is the good state, adjusting the initial shard byte count to increase to obtain the shard byte count, where the shard byte count is less than the upper limit threshold; In the case where it is determined that the current network state of the distributed storage system is the high-latency state, adjusting the initial shard byte count to decrease to obtain the shard byte count, where the shard byte count is greater than the lower limit threshold; In the case where it is determined that the current network state of the distributed storage system is the normal state, using the initial shard byte count as the shard byte count.
8. The storage method of the authentication toolkit according to claim 3, wherein Storing the multiple shard files includes: Determining the network latencies corresponding to the multiple nodes of the distributed storage system respectively; Determining a first node set among the multiple nodes according to the network latencies corresponding to the multiple nodes respectively, where the network latencies of the nodes in the first node set are less than the latency threshold; In the case where the number of nodes included in the first node set is less than the preset number, determining a second node set among the multiple nodes of the distributed storage system, where the network latencies of the nodes in the second node set are less than the latency threshold and the distance between the nodes in the second node set and the first node is less than the distance threshold; In the case where the number of nodes included in the second node set is less than the preset number, determining a third node set among the multiple nodes of the distributed storage system, where the network latencies of the nodes in the third node set are less than the latency threshold, the distance between the nodes in the third node set and the first node is less than the distance threshold, and the load of the nodes in the third node set is less than the load threshold; Distributively storing the multiple shard files on at least one first target node, where the first target node is a node in the third node set.
9. The storage method of the authentication toolkit according to any one of claims 1-5, characterized in that, The method further includes: Receive a tool download request of a target object, where the tool download request carries identification information of a requested tool package; Determine a target identifier corresponding to the requested tool package in the identification table according to the identification information; Determine a target tool package corresponding to the target identifier according to the file content indicated by the target identifier; Determine at least one second target node that stores the target tool package and has a qualified network status among multiple nodes of the distributed storage system; Download the target tool package from the at least one second target node and feedback it to the target object.
10. The storage method of the authentication toolkit according to claim 9, characterized in that, The downloading the target tool package from the at least one second target node and feedbacking it to the target object includes: Download multiple shard files of the target tool package from the at least one second target node; Recombine the multiple shard files of the target tool package to obtain a complete file of the target tool package; Determine an identifier of the complete file of the target tool package by using a preset hash check algorithm; Determine that the complete file of the target tool package passes the hash check when the identifier of the complete file of the target tool package is the same as the target identifier; Verify the digital certificate of the target tool package; Feed back the target tool package to the target object when the complete file of the target tool package passes the hash check and the digital certificate of the target tool package passes the verification.
11. The storage method of the authentication toolkit according to any one of claims 1-5, characterized in that The method further includes: When detecting that there is a version update for at least one tool package, download an updated tool package from a corresponding official website; Compare the file content of the updated tool package with the file content of the corresponding old version tool package to be updated, and determine the differential file content; Retain the differential file content in the updated tool package to obtain an incremental package and generate a corresponding incremental identifier according to the differential file content; Add the incremental identifier to the identification table and store the incremental package; Establish an association relationship between the incremental identifier and the identifier of the file content of the old version tool package.
12. A storage device for an authentication toolkit, characterized in that, Includes: An identifier generation module, configured to generate a first identifier corresponding to the first tool package according to the file content included in the first tool package when receiving the first tool package uploaded by the first node, where the first identifier is used to indicate the file content included in the first tool package; An identifier query module, configured to determine whether the first identifier exists in a preset identification table, where the identification table includes multiple identifiers, and the multiple identifiers respectively correspond to the file content of multiple tool packages already stored in the distributed storage system; An association module, configured to establish an association relationship between the first identifier and the first node and delete the first tool package uploaded by the first node when the first identifier exists in the identification table.
13. An electronic device, characterized in that, Includes: A memory, configured to store a computer program; A processor, configured to implement the steps of the method according to any one of claims 1 to 11 when executing the computer program.
14. A computer-readable storage medium, characterized in that, A computer program is stored in the computer-readable storage medium, wherein when the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 11 are implemented.
15. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 11 are implemented.
Citation Information
Patent Citations
Node configuration method, device and equipment, readable storage medium and server
CN115632944A
File management method and device, electronic equipment and storage medium
CN117992411A
Task processing method, internet of things system and computer program product
CN118170538A
Low-delay communication method, system and device of distributed safety production monitoring platform and medium
CN119676149A
Method, system and equipment for supporting uploading of large files in multiple formats and medium
CN119728666A