Business data storage method based on distributed storage architecture

By slicing business data and making decisions on node heterogeneity, combined with deep learning models and encryption algorithms, the problems of insufficient data security and easy data loss in distributed storage architecture are solved, achieving higher data security and resilience against data loss.

CN121940290APending Publication Date: 2026-04-28CHENGDU BELL COMM INDAL
View PDF 3 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
CHENGDU BELL COMM INDAL
Filing Date
2026-03-05
Publication Date
2026-04-28

AI Technical Summary

Technical Problem

In existing distributed storage architectures, business data security is insufficient and easily lost. Storing data on a single node using encryption technology leads to security issues.

Method used

By slicing the data to be stored and using deep learning models to identify network traffic characteristics, combined with node heterogeneity decision-making methods, the target storage node is determined in the distributed storage architecture. Symmetric encryption and threshold secret sharing algorithms are used to split the key, thereby improving data security and resistance to data loss.

Benefits of technology

Effectively identify user access security, increase the difficulty of data cracking, improve data storage security and resistance to loss, and ensure that data is not attacked by unauthorized users.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121940290A_ABST
    Figure CN121940290A_ABST
Patent Text Reader

Abstract

The invention discloses a service data storage method based on a distributed storage architecture, and relates to the technical field of data storage. Service data to be stored are sliced, data copy slices of the service data to be stored are determined, and network flow characteristics in the process that a user accesses the distributed storage architecture are obtained; and the network flow characteristics are transmitted to the access security control model for identification, and the network security access result is determined, so that the security of the user in the access process can be effectively identified, and then on the basis that the network security access result is security access, the security of the user can be effectively identified. The target storage nodes corresponding to the data copy slices are determined in the distributed storage architecture by adopting a node isomerism decision-making method, it can be guaranteed that stored data are not attacked by illegal users, meanwhile, the cracking difficulty is increased, the data storage safety is improved, and finally the data copy slices are stored in the corresponding target storage nodes, so that the data storage efficiency is improved. And the loss resistance and the security of the data are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of data storage technology, and more specifically, to a business data storage method based on a distributed storage architecture. Background Technology

[0002] Business data is a core asset for enterprise operations, encompassing information across multiple dimensions, including sales, customers, products, and operations. By collecting, organizing, and analyzing this data, precise insights into market trends, customer needs, and business bottlenecks can be gained, providing a scientific basis for decision-making. For example, sales data reflects performance and regional differences, customer data reveals user profiles and consumption preferences, and operational data monitors process efficiency and cost control. Data analytics tools enable data visualization, trend prediction, and risk warnings, helping enterprises optimize resource allocation, improve operational efficiency, and develop precise marketing strategies, ultimately driving business growth and enhancing core competitiveness. Therefore, the storage of business data is crucial. Existing technologies use encryption to store business data on a node within a distributed storage architecture to achieve massive data storage, but this leads to insufficient security and vulnerability to data loss. Summary of the Invention

[0003] This application aims to provide a business data storage method based on a distributed storage architecture, which solves the problem that existing technologies use encryption technology to store business data on a node of a distributed storage architecture, resulting in insufficient business data security and easy loss.

[0004] This application provides a business data storage method based on a distributed storage architecture, including: Obtain the service data to be stored transmitted by the user, and slice the service data to be stored to determine the data copy slice of the service data to be stored; The network traffic characteristics during the user's access to the distributed storage architecture are obtained, and the network traffic characteristics are transmitted to the access security control model for identification to determine the network security access result; wherein, the access security control model is deployed based on a deep learning model; Based on the fact that the network security access result is secure access, the node heterogeneity decision method is used to determine the target storage node corresponding to the data replica slice of the business data to be stored in the distributed storage architecture. Based on the target storage node corresponding to the data replica slice of the business data to be stored, the data replica slice is stored in the corresponding target storage node.

[0005] In some possible implementations, acquiring user-transmitted service data to be stored and slicing the service data to be stored to determine the data copy slices of the service data to be stored includes: Obtain the business data to be stored transmitted by the user; The business data to be stored is encrypted using a symmetric encryption algorithm to determine the encrypted business data to be stored and the symmetric encryption key. The symmetric encryption key is segmented to determine the segmentation key information; The encrypted business data to be stored and the segmentation key information are used together as a data copy slice of the business data to be stored.

[0006] In some possible implementations, the step of segmenting the symmetric encryption key and determining the segmentation key information includes: The symmetric encryption key is segmented using a threshold secret sharing algorithm to divide it into multiple key fragments, resulting in multiple segmented key information; wherein the number of segmented key information is less than or equal to the total number of storage nodes in the distributed storage architecture.

[0007] In some possible implementations, the access security control model is built on a convolutional neural network, and the access security control model is trained using the lion herd optimization algorithm before use.

[0008] In some possible implementations, a node heterogeneity decision-making method is used to determine the target storage node corresponding to the data replica slice of the business data to be stored in the distributed storage architecture, including: For any storage node in the distributed storage architecture, a numerical range is assigned to the storage node, each numerical range has the same length, and the numerical ranges of all storage nodes form a continuous total numerical range. Based on the number of data replica slices corresponding to the business data to be stored, multiple different target codes are generated; wherein, the element dimension in the target code is the same as the number of data replica slices, and each element is randomly generated within the range of the total number value; Obtain the node heterogeneity of multiple storage nodes corresponding to each target code, and determine the target code with the highest node heterogeneity as the optimal target code; Based on the optimal target code, perform optimal information learning on the target code to determine the target code after optimal information learning; Random matching learning is performed on the target code after learning the optimal information to determine the target code after random matching learning; The target code after random matching learning is subjected to group information learning to determine the target code after group information learning; Differential evolutionary learning is performed on the target code after learning the group information to determine the target code after differential evolutionary learning; Based on the target code learned from the optimal information, the target code learned from random matching, the target code learned from the population information, and the target code learned from differential evolution, the optimal target code is maintained to obtain the maintained optimal target code. If the total learning process is greater than or equal to the preset maximum learning process, the optimal target code after maintenance is decoded to obtain the target storage node corresponding to the data copy slice of the business data to be stored. Otherwise, based on the optimal target code after maintenance, the step of learning the optimal information for the target code is returned to enter the next learning process.

[0009] In some possible implementations, based on the optimal target code, optimal information learning is performed on the target code to determine the target code after optimal information learning as follows:

[0010]

[0011] in, Indicates the first t The first learning process j One target code, j =1,2,…,NP, where NP represents the total number of target codes. This represents the optimal target encoding. Indicates the first j The target encoding after learning the optimal information. Indicates inertia weight, Represents the natural constant. Represents pi (π). Indicates the first spiral shape coefficient. l This represents a random variable spiral control factor between (-1, 1). This represents the maximum value of the inertia weight. This represents the minimum value of the inertia weight. This indicates the preset maximum learning process. Represents the hyperbolic tangent function. This represents the first random number between (0, 1).

[0012] In some possible implementations, random matching learning is performed on the target code after learning the optimal information, and the target code after random matching learning is determined as follows:

[0013] in, Indicates the first t The first learning process n The target encoding after learning the optimal information. Indicates the first n The target encoding after learning random matching n =1,2,…,NP Let cos represent the first learning rate, and let cos represent the cosine function. Represented as target encoding Other target codes that are randomly matched, Indicates the second spiral shape coefficient. pm (nn) represents other target encodings. The corresponding node heterogeneity ranking, pm(n) represents the target encoding. The corresponding node heterogeneity ranking is based on the node heterogeneity from smallest to largest.

[0014] In some possible implementations, population information learning is performed on the target code after the random matching learning to determine the target code after population information learning as follows:

[0015]

[0016]

[0017] in, Indicates the first t The first learning process m The target encoding after learning random matching m =1,2,…,NP Indicates the first m The target encoding after learning the information of a group. This represents the second learning rate. This represents the third learning rate. This represents the worst-case target encoding, i.e., the target encoding with the minimum node heterogeneity; Indicates fusion encoding, This represents a random adjustment factor between (0,1). Indicates the first t The first learning process m The fusion coefficients corresponding to the target encoding after random matching learning. Indicates the first t The first learning process m The node heterogeneity corresponding to the target encoding after random matching learning.

[0018] In some possible implementations, differential evolution learning is performed on the target code after learning the group information, and the target code after differential evolution learning is determined as follows:

[0019]

[0020] in, Indicates the first t The first learning process k The target encoding after learning the information of a group. k =1,2,…,NP Indicates the first k The target encoding after differential evolution learning Represents the differential evolution factor. This represents the second random number between (0,1). This represents a third random number between (0,1). This represents the fourth random number between (0,1). This represents the fifth random number between (0,1). This represents the sixth random number between (0,1). This represents the first differential target encoding for randomization. This represents the random second differential target encoding. This represents the sine function.

[0021] In some possible implementations, the optimal target code is maintained based on the target code learned from the optimal information, the target code learned from random matching, the target code learned from the population information, and the target code learned from differential evolution, to obtain the maintained optimal target code, including: Based on the target code after learning the optimal information, the target code after learning the random matching, the target code after learning the population information, and the target code after learning the differential evolution, the target code with the greatest node heterogeneity is determined as the target code to be processed. Determine whether the node heterogeneity of the target code to be processed is greater than that of the optimal target code. If so, the target code to be processed is taken as the new optimal target code to obtain the maintained optimal target code. Otherwise, the original optimal target code remains unchanged to obtain the maintained optimal target code.

[0022] Beneficial effects: This application provides a business data storage method based on a distributed storage architecture. It acquires the business data to be stored transmitted by the user, slices the data, determines the data replica slices, then acquires the network traffic characteristics during the user's access to the distributed storage architecture, and transmits these characteristics to an access security control model for identification, determining the network security access result. This effectively identifies the user's security during the access process. Following a confirmed secure access result, a node heterogeneity decision method is used to determine the target storage node corresponding to the data replica slice in the distributed storage architecture. This ensures that the stored data is not attacked by unauthorized users, increases the difficulty of cracking, and improves data storage security. Finally, based on the target storage node corresponding to the data replica slice, the data replica slice is stored in the corresponding target storage node, improving data resilience and security. Attached Figure Description

[0023] To more clearly illustrate the technical solutions of the embodiments of this application, the drawings used in the description of the embodiments of this application will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0024] Figure 1 This is a flowchart of a business data storage method based on a distributed storage architecture proposed in an embodiment of this application.

[0025] Figure 2 This is a flowchart illustrating the process of obtaining a data copy slice of business data to be stored, as proposed in one embodiment of this application.

[0026] Figure 3 This is a flowchart illustrating how to determine the target storage node corresponding to a data copy slice of business data to be stored, as proposed in one embodiment of this application. Detailed Implementation

[0027] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0028] like Figure 1 As shown in the figure, this application provides a business data storage method based on a distributed storage architecture, including: S101. Obtain the service data to be stored transmitted by the user, and slice the service data to be stored to determine the data copy slice of the service data to be stored.

[0029] By slicing the business data to be stored, the business data to be stored can also be stored on different storage nodes using the distributed storage concept, which improves data security and avoids data loss.

[0030] S102. Obtain network traffic characteristics during the user's access to the distributed storage architecture, and transmit the network traffic characteristics to the access security control model for identification to determine the network security access result; wherein, the access security control model is deployed based on a deep learning model; By identifying network traffic characteristics during user access to a distributed storage architecture, it is possible to effectively determine whether a user is accessing the system normally. The process of identifying network traffic characteristics is relatively conventional and will not be described in detail in this application's embodiments.

[0031] Network security access results can be either secure or insecure. If the network security access result is insecure, the user's access can be terminated or the user can be blacklisted for a preset period of time to prevent data from being attacked by unauthorized users.

[0032] S103. Based on the network security access result being secure access, the target storage node corresponding to the data replica slice of the business data to be stored is determined in the distributed storage architecture using the node heterogeneity decision method. By employing a node heterogeneity decision-making method to determine the target storage node corresponding to the data replica slice of the business data to be stored in a distributed storage architecture, the business data to be stored can be distributed across heterogeneous nodes as much as possible, increasing the difficulty of being hacked and further enhancing data security.

[0033] S104. Based on the target storage node corresponding to the data replica slice of the business data to be stored, store the data replica slice in the corresponding target storage node.

[0034] The present application provides a business data storage method based on a distributed storage architecture, which can effectively avoid data loss and attacks, and comprehensively improve the storage security of business data in the distributed storage architecture.

[0035] like Figure 2 As shown, the process involves acquiring user-transmitted service data to be stored, slicing the service data to be stored, and determining the data copy slices of the service data to be stored, including: S201. Obtain the service data to be stored transmitted by the user; S202. The business data to be stored is encrypted using a symmetric encryption algorithm, and the encrypted business data to be stored and the symmetric encryption key are determined. S203. Divide the symmetric encryption key and determine the segmentation key information; S204. The encrypted business data to be stored and the segmentation key information are used together as a data copy slice of the business data to be stored.

[0036] A data replica slice represents a split key and an encrypted copy of the business data to be stored. Utilizing a distributed storage architecture to store the data replica slice can effectively improve data security.

[0037] Optionally, the encrypted business data to be stored can be randomly stored on N servers, where N is user-defined, and the segmentation key information can be used as a data copy slice of the business data to be stored, thereby reducing data storage pressure.

[0038] In some possible implementations, the step of segmenting the symmetric encryption key and determining the segmentation key information includes: The symmetric encryption key is segmented using a threshold secret sharing algorithm to divide it into multiple key fragments, resulting in multiple segmented key information; wherein the number of segmented key information is less than or equal to the total number of storage nodes in the distributed storage architecture.

[0039] By employing a threshold secret sharing algorithm to segment the symmetric encryption key, data decryption can still be achieved even with the loss of some segmented key information, and decryption will be impossible even with the loss of a small amount of segmented key information, thus improving data storage security. Building upon this, using a node heterogeneity decision-making method in a distributed storage architecture to determine the target storage node corresponding to the data replica slice of the business data to be stored can further increase the difficulty of cracking the business data, exponentially increasing the difficulty.

[0040] In some possible implementations, the access security control model is built on a convolutional neural network, and the access security control model is trained using the lion herd optimization algorithm before use.

[0041] like Figure 3 As shown, the node heterogeneity decision-making method is used to determine the target storage node corresponding to the data replica slice of the business data to be stored in a distributed storage architecture, including: S301. For any storage node in the distributed storage architecture, assign a numerical range to the storage node, where each numerical range has the same length, and the numerical ranges of all storage nodes form a continuous total numerical range.

[0042] For example, assuming there are six storage nodes, the numerical ranges for these six storage nodes can be assigned as (0,1], (1,2], (2,3], (3,4], (4,5], and (5,6], respectively, so the total value range is (0,6).

[0043] Optionally, to reduce the amount of data processing, (0,1] can be divided into 6 intervals.

[0044] S302. Based on the number of data replica slices corresponding to the business data to be stored, generate multiple different target codes; wherein, the element dimension in the target code is the same as the number of data replica slices, and each element is randomly generated within the range of the total number value; For example, if the number of data replica slices is 3 and the total value range is (0, 6], then a 3-dimensional vector needs to be generated, with each element randomly generated within (0, 6]. However, it's worth noting that in practical applications, to implement the threshold secret sharing algorithm, a larger value can be set as the number of data replica slices, for example, 10.

[0045] S303. Obtain the node heterogeneity of multiple storage nodes corresponding to each target code, and determine the target code with the largest node heterogeneity as the optimal target code; Optionally, the formula for calculating node heterogeneity is:

[0046]

[0047]

[0048]

[0049] in, This represents node heterogeneity, and the greater the node heterogeneity, the better. Indicates the first h The richness of each functional component category Indicates the first h The differences between the functional component categories, where H represents the total number of functional component categories. Indicates the first h The first functional component category i The frequency of a functional component appearing in multiple storage nodes corresponding to the target encoding, S represents the frequency of the first functional component.h The total number of functional components in each functional component category, where ln represents the logarithmic function. Represents the coefficient of variation. This represents the number of vulnerabilities in the vulnerability intersection among multiple storage nodes corresponding to the target encoding. The node heterogeneity is maximized when the vulnerability intersection is an empty set.

[0050] Optionally, in addition to node heterogeneity, the sum of the proportions of remaining storage capacity corresponding to the storage nodes associated with the target encoding can also be used as an indicator; a larger sum indicates better storage performance. Since the storage capacity of a distributed storage architecture is massive, this embodiment does not consider the problem of insufficient remaining storage capacity for a single storage node. In practical applications, if a single storage node has remaining storage capacity, a numerical range may not be allocated to that storage node during the allocation of numerical ranges.

[0051] A functional component category can refer to a certain category of software or hardware, such as an operating system or processor model. A functional component can be considered as a specific software or hardware under a functional component category. For example, Windows 10, Windows 10, and CentOS 8 are all functional components under an operating system. Other functional components are similar, and will not be described in detail in the embodiments of this application.

[0052] For any given target code, each target code corresponds to multiple storage nodes, and each storage node has corresponding functional components. Therefore, the first storage node corresponding to the target code can be determined. h The first functional component category i The frequency of various functional components is used to realize the computation of node heterogeneity.

[0053] Assuming that the number of data replica slices corresponding to a certain business data is M (M can be set to an even number for ease of processing), then in the best case, the M data replica slices are stored one by one in the M storage nodes. However, this requires a relatively strict constraint on the target encoding, which is not conducive to the optimization of the target encoding. Due to the characteristics of the threshold secret sharing algorithm, at least 1+M / 2 segmentation key information is required to achieve decryption. Therefore, the constraint condition for constructing the target encoding in this application embodiment is: at least 1+M / 2 elements in the target encoding need to be located in different numerical intervals. If there are multiple elements located in the same numerical interval, resulting in the failure to meet the constraint condition, the elements located in the duplicate numerical intervals can be allocated one by one to the numerical intervals that do not appear in the target encoding until the constraint condition is met. For example, if there are six different storage nodes corresponding to the numerical intervals (0,1], (1,2], (2,3], (3,4], (4,5], and (5,6]), and M is 4, then at least 1 + 4 / 2 elements in the target code need to be located in different numerical intervals. Assuming that the four elements in the target code are located in (0,1], (0,1], (0,1], and (1,2] respectively, we can judge them one by one. There are three elements in (0,1], so we can keep the first element and process the subsequent duplicate elements one by one. For example, the second element in (0,1] can be randomly generated in (2,3], (3,4], (4,5], and (5,6], and then the other duplicate elements can be processed until the constraint condition is met.

[0054] During processing, elements within the same numerical range can be processed sequentially. Assume that elements within the same numerical range in the target encoding constitute one element set. Also assume there are three element sets: a first set, a second set, and a third set. If the number of elements in the first set is greater than the number in the second set, and the number of elements in the second set is greater than the number in the third set, then the elements in the first set can be processed one by one until the constraints are met or the number of remaining elements in the first set is the same as the number of elements in the second set. Then, the first and second sets are processed simultaneously, and so on. However, it's worth noting that if two element sets initially have the same number of elements, and both are at their maximum, then these two sets are processed simultaneously initially, followed by all three sets, and so on.

[0055] S304. Based on the optimal target code, perform optimal information learning on the target code to determine the target code after optimal information learning; S305. Perform random matching learning on the target code after learning the optimal information to determine the target code after random matching learning; S306. Perform group information learning on the target code after random matching learning, and determine the target code after group information learning; S307. Perform differential evolution learning on the target code after learning the group information to determine the target code after differential evolution learning; S308. Based on the target code learned from the optimal information, the target code learned from random matching, the target code learned from the population information, and the target code learned from differential evolution, the optimal target code is maintained to obtain the maintained optimal target code. S309. Determine whether the total learning process is greater than or equal to the preset maximum learning process. If so, decode the optimal target code after maintenance to obtain the target storage node corresponding to the data copy slice of the business data to be stored. Otherwise, based on the optimal target code after maintenance, return to the step of learning the optimal information for the target code to enter the next learning process.

[0056] Decoding the optimal target encoding after maintenance may include: determining the target storage node corresponding to the data replica slice corresponding to the element based on the numerical range corresponding to the element.

[0057] This application's embodiments employ several learning methods to find an optimal position (i.e., the global optimal solution) in a high-dimensional solution space of M dimensions, thereby maximizing node heterogeneity. Compared to existing technologies that use genetic algorithms for encoding optimization, the method provided in this application's embodiments offers more precise optimization.

[0058] In some possible implementations, based on the optimal target code, optimal information learning is performed on the target code to determine the target code after optimal information learning as follows:

[0059]

[0060] in, Indicates the first t The first learning process j One target code, j =1,2,…,NP, where NP represents the total number of target codes. This represents the optimal target encoding. Indicates the first j The target encoding after learning the optimal information. Indicates inertia weight, Represents the natural constant. Represents pi (π). Indicates the first spiral shape coefficient. l This represents a random variable spiral control factor between (-1, 1). This represents the maximum value of the inertia weight. This represents the minimum value of the inertia weight. This indicates the preset maximum learning process. Represents the hyperbolic tangent function. This represents the first random number between (0, 1).

[0061] The optimal information learning provided in this application embodiment can achieve rapid optimization of the target encoding. Furthermore, the introduction of the hyperbolic tangent function not only avoids entering a local optimum state too early in the early stages of the algorithm, but also enhances the algorithm's local search capability in the middle and later stages, enabling it to find the global optimum more accurately and balance global and local search capabilities.

[0062] In some possible implementations, random matching learning is performed on the target code after learning the optimal information, and the target code after random matching learning is determined as follows:

[0063] in, Indicates the first t The first learning process n The target encoding after learning the optimal information. Indicates the first n The target encoding after learning random matching n =1,2,…,NP Let cos represent the first learning rate, and let cos represent the cosine function. Represented as target encoding Other target codes that are randomly matched, Indicates the second spiral shape coefficient. pm (nn) represents other target encodings. The corresponding node heterogeneity ranking, pm(n) represents the target encoding. The corresponding node heterogeneity ranking is based on the node heterogeneity from smallest to largest.

[0064] The random matching learning provided in this application embodiment can search the region between two target codes in the solution space in a spiral manner, so as to maintain the diversity of target codes and improve the convergence accuracy of the algorithm and the possibility of finding the global optimal solution.

[0065] In some possible implementations, population information learning is performed on the target code after the random matching learning to determine the target code after population information learning as follows:

[0066]

[0067]

[0068] in, Indicates the first t The first learning process m The target encoding after learning random matching m =1,2,…,NP Indicates the first m The target encoding after learning the information of a group. This represents the second learning rate. This represents the third learning rate. This represents the worst-case target encoding, i.e., the target encoding with the minimum node heterogeneity; Indicates fusion encoding, This represents a random adjustment factor between (0,1). Indicates the first t The first learning process m The fusion coefficients corresponding to the target encoding after random matching learning. Indicates the first t The first learning process m The node heterogeneity corresponding to the target encoding after random matching learning.

[0069] The group information learning provided in this application embodiment can search more local unfamiliar regions based on the known worst position and the known best position, which can improve the algorithm's ability to escape local optima to a certain extent and help the algorithm find the global optimum.

[0070] In some possible implementations, differential evolution learning is performed on the target code after learning the group information, and the target code after differential evolution learning is determined as follows:

[0071]

[0072] in, Indicates the first t The first learning process k The target encoding after learning the information of a group. k =1,2,…,NP Indicates the first kThe target encoding after differential evolution learning Represents the differential evolution factor. This represents the second random number between (0,1). This represents a third random number between (0,1). This represents the fourth random number between (0,1). This represents the fifth random number between (0,1). This represents the sixth random number between (0,1). This represents the first differential target encoding for randomization. This represents the random second differential target encoding. This represents the sine function.

[0073] The differential evolution learning provided in this application embodiment enables the algorithm to have a larger search range in the early stage, which greatly improves the global search capability. In the later stage of the algorithm, the differential evolution capability is gradually reduced to improve the convergence accuracy of the algorithm and gradually find the global optimal solution.

[0074] In some possible implementations, the optimal target code is maintained based on the target code learned from the optimal information, the target code learned from random matching, the target code learned from the population information, and the target code learned from differential evolution, to obtain the maintained optimal target code, including: Based on the target code after learning the optimal information, the target code after learning the random matching, the target code after learning the population information, and the target code after learning the differential evolution, the target code with the greatest node heterogeneity is determined as the target code to be processed. Determine whether the node heterogeneity of the target code to be processed is greater than that of the optimal target code. If so, the target code to be processed is taken as the new optimal target code to obtain the maintained optimal target code. Otherwise, the original optimal target code remains unchanged to obtain the maintained optimal target code.

[0075] The various embodiments in this specification are described in a progressive manner, with each embodiment focusing on the differences from other embodiments. The same or similar parts between the various embodiments can be referred to each other.

[0076] This application provides a business data storage method based on a distributed storage architecture. It acquires the business data to be stored transmitted by the user, slices the data, determines the data replica slices, then acquires the network traffic characteristics during the user's access to the distributed storage architecture, and transmits these characteristics to an access security control model for identification, determining the network security access result. This effectively identifies the user's security during the access process. Next, based on the network security access result indicating secure access, a node heterogeneity decision method is used to determine the target storage node corresponding to the data replica slice in the distributed storage architecture. This ensures that the stored data is not attacked by unauthorized users, increases the difficulty of cracking, and improves data storage security. Finally, according to the target storage node corresponding to the data replica slice, the data replica slice is stored in the corresponding target storage node, improving data resilience and security.

[0077] This application describes embodiments with reference to flowchart illustrations and / or block diagrams of methods, apparatuses, electronic devices, and computer program products according to embodiments of this application. It should be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing terminal device to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing terminal device, generate instructions for implementing the flowchart... Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.

[0078] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing terminal device to operate in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.

[0079] These computer program instructions can also be loaded onto a computer or other programmable data processing terminal equipment, causing a series of operational steps to be performed on the computer or other programmable terminal equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable terminal equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.

[0080] Although preferred embodiments of the present application have been described, those skilled in the art, upon learning the basic inventive concept, can make other changes and modifications to these embodiments. Therefore, the appended claims are intended to be interpreted as including the preferred embodiments as well as all changes and modifications falling within the scope of the embodiments of the present application.

[0081] Finally, it should be noted that in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or terminal device that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or terminal device. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or terminal device that includes said element.

[0082] This document uses specific examples to illustrate the principles and implementation methods of this application. The descriptions of the above embodiments are only for the purpose of helping to understand the methods and core ideas of this application. At the same time, for those skilled in the art, there will be changes in the specific implementation methods and application scope based on the ideas of this application. Therefore, the content of this specification should not be construed as a limitation of this application.

Claims

1. A business data storage method based on a distributed storage architecture, characterized in that, include: Obtain the service data to be stored transmitted by the user, and slice the service data to be stored to determine the data copy slice of the service data to be stored; The network traffic characteristics during the user's access to the distributed storage architecture are obtained, and the network traffic characteristics are transmitted to the access security control model for identification to determine the network security access result; wherein, the access security control model is deployed based on a deep learning model; Based on the fact that the network security access result is secure access, the node heterogeneity decision method is used to determine the target storage node corresponding to the data replica slice of the business data to be stored in the distributed storage architecture. Based on the target storage node corresponding to the data replica slice of the business data to be stored, the data replica slice is stored in the corresponding target storage node.

2. The business data storage method based on a distributed storage architecture according to claim 1, characterized in that, Acquire user-transmitted service data to be stored, and slice the service data to be stored to determine the data copy slice of the service data to be stored, including: Obtain the business data to be stored transmitted by the user; The business data to be stored is encrypted using a symmetric encryption algorithm to determine the encrypted business data to be stored and the symmetric encryption key. The symmetric encryption key is segmented to determine the segmentation key information; The encrypted business data to be stored and the segmentation key information are used together as a data copy slice of the business data to be stored.

3. The business data storage method based on a distributed storage architecture according to claim 2, characterized in that, The step of segmenting the symmetric encryption key and determining the segmented key information includes: The symmetric encryption key is segmented using a threshold secret sharing algorithm to divide it into multiple key fragments, resulting in multiple segmented key information; wherein the number of segmented key information is less than or equal to the total number of storage nodes in the distributed storage architecture.

4. The business data storage method based on a distributed storage architecture according to claim 1, characterized in that, The access security control model is built on a convolutional neural network, and the access security control model is trained using the lion herd optimization algorithm before use.

5. The business data storage method based on a distributed storage architecture according to claim 1, characterized in that, The node heterogeneity decision-making method is used to determine the target storage node corresponding to the data replica slice of the business data to be stored in the distributed storage architecture, including: For any storage node in the distributed storage architecture, a numerical range is assigned to the storage node, each numerical range has the same length, and the numerical ranges of all storage nodes form a continuous total numerical range. Based on the number of data replica slices corresponding to the business data to be stored, multiple different target codes are generated; wherein, the element dimension in the target code is the same as the number of data replica slices, and each element is randomly generated within the range of the total number value; Obtain the node heterogeneity of multiple storage nodes corresponding to each target code, and determine the target code with the highest node heterogeneity as the optimal target code; Based on the optimal target code, perform optimal information learning on the target code to determine the target code after optimal information learning; Random matching learning is performed on the target code after learning the optimal information to determine the target code after random matching learning; The target code after random matching learning is subjected to group information learning to determine the target code after group information learning; Differential evolutionary learning is performed on the target code after learning the group information to determine the target code after differential evolutionary learning; Based on the target code learned from the optimal information, the target code learned from random matching, the target code learned from the population information, and the target code learned from differential evolution, the optimal target code is maintained to obtain the maintained optimal target code. If the total learning process is greater than or equal to the preset maximum learning process, the optimal target code after maintenance is decoded to obtain the target storage node corresponding to the data copy slice of the business data to be stored. Otherwise, based on the optimal target code after maintenance, the step of learning the optimal information for the target code is returned to enter the next learning process.

6. The business data storage method based on a distributed storage architecture according to claim 5, characterized in that, Based on the optimal target code, optimal information learning is performed on the target code to determine the target code after optimal information learning as follows: in, Indicates the first t The first learning process j One target code, j =1,2,…,NP, where NP represents the total number of target codes. This represents the optimal target encoding. Indicates the first j The target encoding after learning the optimal information. Indicates inertia weight, Represents the natural constant. Represents pi (π). Indicates the first spiral shape coefficient. l This represents a random variable spiral control factor between (-1, 1). This represents the maximum value of the inertia weight. This represents the minimum value of the inertia weight. This indicates the preset maximum learning process. Represents the hyperbolic tangent function. This represents the first random number between (0,1).

7. The business data storage method based on a distributed storage architecture according to claim 6, characterized in that, Random matching learning is performed on the target code after learning the optimal information, and the target code after random matching learning is determined as follows: in, Indicates the first t The first learning process n The target encoding after learning the optimal information. Indicates the first n The target encoding after learning random matching n =1,2,…,NP Let cos represent the first learning rate, and let cos represent the cosine function. Represented as target encoding Other target codes that are randomly matched, Indicates the second spiral shape coefficient. pm (nn) represents other target encodings. The corresponding node heterogeneity ranking, pm(n) represents the target encoding. The corresponding node heterogeneity ranking is based on the node heterogeneity from smallest to largest.

8. The business data storage method based on a distributed storage architecture according to claim 7, characterized in that, After random matching learning, the target code is subjected to group information learning to determine the target code after group information learning as follows: in, Indicates the first t The first learning process m The target encoding after learning random matching m =1,2,…,NP Indicates the first m The target encoding after learning the information of a group. This represents the second learning rate. This represents the third learning rate. This represents the worst-case target encoding, i.e., the target encoding with the minimum node heterogeneity; Indicates fusion encoding, This represents a random adjustment factor between (0,1). Indicates the first t The first learning process m The fusion coefficients corresponding to the target encoding after random matching learning. Indicates the first t The first learning process m The node heterogeneity corresponding to the target encoding after random matching learning.

9. The business data storage method based on a distributed storage architecture according to claim 8, characterized in that, Differential evolutionary learning is performed on the target code after learning the group information, and the target code after differential evolutionary learning is determined as follows: in, Indicates the first t The first learning process k The target encoding after learning the information of a group. k =1,2,…,NP Indicates the first k The target encoding after differential evolution learning Represents the differential evolution factor. This represents the second random number between (0,1). This represents a third random number between (0,1). This represents the fourth random number between (0,1). This represents the fifth random number between (0,1). This represents the sixth random number between (0,1). This represents the first differential target encoding for randomization. This represents the random second difference target encoding. This represents the sine function.

10. The business data storage method based on a distributed storage architecture according to claim 9, characterized in that, Based on the target code learned from the optimal information, the target code learned from random matching, the target code learned from the population information, and the target code learned from differential evolution, the optimal target code is maintained to obtain the maintained optimal target code, including: Based on the target code after learning the optimal information, the target code after learning the random matching, the target code after learning the population information, and the target code after learning the differential evolution, the target code with the greatest node heterogeneity is determined as the target code to be processed. Determine whether the node heterogeneity of the target code to be processed is greater than that of the optimal target code. If so, the target code to be processed is taken as the new optimal target code to obtain the maintained optimal target code. Otherwise, the original optimal target code remains unchanged to obtain the maintained optimal target code.

Citation Information

Patent Citations

  • Network information security comprehensive analysis and monitoring system and method

    CN118413359A

  • Highly distributed, cryptographic-based data storage method

    US20230177198A1

  • Method and system for secure and synchronous storage area network (SAN) infrastructure to san infrastructure data replication

    US20240291806A1