Efficient distributed data storage and retrieval method and system

By evaluating node load abnormality and dynamic migration evaluation of the programming platform and adjusting node load, the storage inaccuracy problem caused by single point failure in the distributed data storage of the programming platform is solved, and the stability and accuracy of storage are improved.

CN120144633AInactive Publication Date: 2025-06-13CHANGZHOU INST OF LIGHT IND TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510182492.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-19
Publication Date
2025-06-13
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

In the prior art, when programming platforms perform efficient storage of distributed data, a single point of failure caused by network attacks, affecting load balancing and causing inaccurate storage.

Method used

By performing node load exception evaluation, obtain node load evaluation scores, and adjust node load exceptions based on the evaluation results, including dynamic migration evaluation and retrieval capability evaluation, to ensure that node load is within the preset range and improve storage accuracy.

Benefits of technology

It has achieved the improvement of the stability and accuracy of the programming platform in the distributed data storage process, and solved the problem of inaccurate storage.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120144633A_ABST
    Figure CN120144633A_ABST
Patent Text Reader

Abstract

The invention discloses an efficient distributed data storage and retrieval method and system, and relates to the technical field of electric digital data processing. The efficient distributed data storage and retrieval method comprises the following steps: evaluating node load abnormity; performing node dynamic migration evaluation; and evaluating the node retrieval capability. According to the method, the node load evaluation score is obtained by performing node load abnormity evaluation, and when the node load evaluation score is not within the preset node load range, node dynamic migration evaluation and node load abnormity adjustment are performed; if yes, performing node retrieval capability evaluation to obtain a node retrieval capability evaluation score, and judging whether to perform node retrieval capability adjustment or not, thereby achieving the effect of improving the storage accuracy of the programming platform in the efficient storage process of the distributed data. The problem that in the prior art, storage is inaccurate in the efficient storage process of distributed data through a programming platform is solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of electronic digital data processing, and particularly to an efficient distributed data storage and retrieval method and system. Background Art

[0002] With the continuous progress of technology and the popularization of the Internet, the amount of data generated in various fields shows an explosive growth trend. Traditional centralized storage and retrieval methods can no longer meet this demand because they face performance bottlenecks, poor scalability, and single-point failures when dealing with large-scale data. Therefore, distributed data storage and retrieval technology has emerged. It stores data dispersedly on multiple nodes and utilizes the parallel processing capabilities of these nodes to improve the efficiency of storage and retrieval. The core of distributed data storage and retrieval algorithms lies in how to distribute data to multiple nodes. Commonly used methods include hash mapping and consistent hashing algorithms. Hash mapping uses the keywords of data for hash calculation and distributes the data to the corresponding nodes. The consistent hashing algorithm can solve the problems of dynamic changes of nodes and load balancing, ensuring the efficient storage and retrieval of data.

[0003] Existing methods mainly store data on multiple physical nodes. Each node can independently process data requests and has one or more copies of the data.

[0004] For example, the distributed data storage method, retrieval method, system, and readable storage medium disclosed in the invention patent announcement with the announcement number CN113254505B include: a target node receives target data sent by a sensor, retrieves whether a target index item corresponding to K is included in the first index table corresponding to the target node. When the target index item is included in the first index table, append V to the target index item, generate a target data source corresponding to the target data, and associate and store the target data source with V; when the target index is not included in the first index table, create an index item K in the first index table, and associate and store the target data source with V in the index item K.

[0005] For example, a distributed storage method, an electronic device, and a medium for design code data announced in the invention patent application with the announcement number of CN117909313B, including: Step S1, obtaining design codes for parsing to generate code information to be stored; Step S2, generating a corresponding node for each A1n, setting an edge Din before the node corresponding to A1n and the node corresponding to Bin, setting the weight corresponding to Din based on Cin, and generating a dependency graph; Step S3, splitting the dependency graph into M groups of sub-dependency graphs such that the sum of the weights of the cut edges is minimized on the premise that the total weight of the edges in each group of sub-dependency graphs is greater than a preset first weight threshold; Step S4, dividing the An corresponding to each group of sub-dependency graphs into a group of code information to be stored, generating M groups of code information to be stored, and performing distributed storage on the M groups of code information to be stored.

[0006] However, in the process of implementing the technical solution of the invention in the embodiments of the present application, it is found that the above technology has at least the following technical problems:

[0007] In the prior art, since the programming platform needs to centrally store code data, a single node failure may be caused due to a network attack, affecting the availability of the data of a single node, and further affecting the load balancing during efficient distributed data dynamic storage, resulting in inaccurate storage during the process of the programming platform performing efficient storage of distributed data. Summary of the Invention

[0008] The embodiments of the present application provide an efficient distributed data storage and retrieval method and system, which solve the problem of inaccurate storage during the process of the programming platform performing efficient storage of distributed data in the prior art, and achieve an improvement in storage accuracy during the process of the programming platform performing efficient storage of distributed data.

[0009] The embodiments of the present application provide an efficient distributed data storage and retrieval method, including the following steps: performing node load anomaly evaluation on the data affected by the obtained node load to obtain a node load evaluation score and determining whether to perform node load anomaly adjustment, where the node load evaluation score is used to evaluate the load condition of the node during the storage of the data stored in the specified node; when the node load evaluation score is not within the preset node load range obtained from the database, performing node dynamic migration evaluation on the basis of the obtained node dynamic migration data to obtain a node migration evaluation score, and performing node load anomaly adjustment based on the node migration evaluation score, where the node migration evaluation score is used to evaluate the migration ability during the node migration of the data stored in the specified node; if no node load anomaly adjustment is performed, performing node retrieval ability evaluation based on the obtained qualified node load evaluation score and the processed node retrieval ability data to obtain a node retrieval ability evaluation score, and determining whether to perform node retrieval ability adjustment based on the node retrieval ability evaluation score, where the node retrieval ability evaluation score is used to evaluate the retrieval ability during the retrieval of the data stored in the specified node.

[0010] Further, the node load impact data includes the total number of bits of stored code and the node compliance; the node compliance includes the first node load compliance, the second node load compliance, the third node load compliance, and the fourth node load compliance; the node dynamic migration data includes the first node migration impact, the second node migration impact, and the third node migration impact; the node retrieval ability data includes the first node retrieval impact, the second node retrieval impact, the third node retrieval impact, and the fourth node retrieval impact; the first node load compliance is used to reflect the compliance between the maximum threshold of the preset task quantity and the number of tasks running on the node; the second node load compliance is used to reflect the compliance between the maximum threshold of the preset number of read / write disks and the number of read / write disks of the node; the third node load compliance is used to reflect the compliance between the maximum threshold of the preset connection quantity and the number of client connections; the fourth node load compliance is used to reflect the compliance between the maximum threshold of the preset code reading speed and the code reading speed of the memory; the node dynamic migration data is used to reflect the influence of the node migration parameters on the load evaluation difference evaluation value; the node migration parameters include the node migration success rate, the node migration data volume, and the node migration speed; the load evaluation difference evaluation value is used to reflect the difference degree between the node load evaluation score outside the preset node load range and the preset node load average threshold; the first node retrieval impact is used to reflect the influence of the node retrieval difference evaluation value by the maximum threshold of the preset retrieval data volume and the difference between the node retrieval data volume; the node retrieval difference evaluation value is used to reflect the difference degree between the node retrieval speed and the maximum threshold of the preset node retrieval speed; the second node retrieval impact is used to reflect the influence of the node retrieval difference evaluation value by the node retrieval success rate; the third node retrieval impact is used to reflect the influence of the node retrieval difference evaluation value by the maximum threshold of the preset concurrent retrieval number and the difference between the node concurrent retrieval numbers; the fourth node retrieval impact is used to reflect the influence of the node retrieval difference evaluation value by the qualified node load evaluation score.

[0011] Further, the specific method for obtaining the node load evaluation score is as follows: Obtain the node load value by processing the node load impact data and the preset node processing weight; the node load value includes the first node load value, the second node load value, the third node load value, and the fourth node load value; Obtain the node load evaluation score by combining the obtained node load value and the preset load weight; the preset load weight includes the first preset load weight, the second preset load weight, and the third preset load weight.

[0012] Further, the limiting expression of the node load evaluation score is as follows:

[0013]

[0014] In the formula, FZPGa denotes the node load evaluation score at the a-th first preset time point, where a = 1, 2,.., h. a represents the serial number of the first preset time point, and h represents the total number of the first preset time points. The first preset time point represents a preset time point during the storage process of node-stored data, FQ a1 denotes the first node load value at the a-th first preset time point, FQ a2 denotes the second node load value at the a-th first preset time point, FQ a3 denotes the third node load value at the a-th first preset time point, FQ a4 denotes the fourth node load value at the a-th first preset time point, e represents the natural constant, a 1 denotes the first weight of the preset load, a 2 denotes the second weight of the preset load, a 3 denotes the third weight of the preset load.

[0015] Further, the specific process for obtaining the node migration evaluation score is as follows: performing a logarithmic function processing on the first influence degree of node migration to obtain the first processed value of node migration; performing an inverse hyperbolic sine function processing on the second influence degree of node migration to obtain the second processed value of node migration; performing an exponential function processing on the third influence degree of node migration to obtain the third processed value of node migration; combining the obtained processed values of node migration and the preset migration weights to obtain the node migration evaluation score; the processed values of node migration include the first processed value of node migration, the second processed value of node migration, and the third processed value of node migration; the preset migration weights include the first preset migration weight and the second preset migration weight.

[0016] Further, the specific process for abnormal adjustment of node load is as follows: when the node migration evaluation score is not lower than the preset migration ability threshold in the database, perform migration ability optimization; when the node migration evaluation score is lower than the preset migration ability threshold, perform migration ability adjustment; when the node migration evaluation score after the migration ability adjustment is still lower than the preset migration ability threshold, send an alarm prompt to the preset personnel.

[0017] Further, the migration ability optimization includes parallel processing and lossless compression; the specific steps for migration ability adjustment are: S21, perform breakpoint resumption processing. When the monitored migration evaluation score of the stored data is not lower than the preset migration ability threshold, stop the operation, otherwise execute S22; S22, perform segmented migration. When the monitored migration evaluation score of the stored data is not lower than the preset migration ability threshold, stop the operation, otherwise send an alarm prompt to the preset personnel.

[0018] Further, the specific process of obtaining the node retrieval ability evaluation score is as follows: performing exponential function processing on the node retrieval ability data to obtain node retrieval processing values; the node retrieval processing values include a first retrieval processing value, a second retrieval processing value, a third retrieval processing value, and a fourth retrieval processing value; combining the node retrieval processing values and preset retrieval weights to obtain a node retrieval ability evaluation score; the preset retrieval weights include a preset first retrieval weight, a preset second retrieval weight, and a preset third retrieval weight.

[0019] Further, the specific process of determining whether to perform node retrieval ability adjustment based on the node retrieval ability evaluation score is as follows: determining whether the node retrieval ability evaluation score is not lower than a preset node retrieval ability average threshold; if so, it indicates that the node retrieval ability is qualified and no node retrieval ability adjustment is performed; otherwise, it indicates that the node retrieval ability is unqualified and node retrieval ability adjustment is performed; when the node retrieval ability is adjusted and the node retrieval ability evaluation score is still lower than the preset node retrieval ability average threshold, an alarm prompt is sent to a preset person; the node retrieval ability adjustment includes performing heuristic search and performing matching optimization.

[0020] The embodiment of the present application provides an efficient distributed data storage and retrieval system, including a node load anomaly evaluation module, a node dynamic migration evaluation module, and a node retrieval ability evaluation module: among them, the node load anomaly evaluation module is used to perform node load anomaly evaluation according to the obtained node load impact data, obtain a node load evaluation score, and determine whether to perform node load anomaly adjustment, and the node load evaluation score is used to evaluate the load situation of the node during the storage of the data stored in the specified node; the node dynamic migration evaluation module is used to perform node dynamic migration evaluation based on the obtained node dynamic migration data when the node load evaluation score is not within the preset node load range obtained from the database, obtain a node migration evaluation score, and perform node load anomaly adjustment based on the node migration evaluation score, and the node migration evaluation score is used to evaluate the migration ability during the node migration of the data stored in the specified node; the node retrieval ability evaluation module is used to, if no node load anomaly adjustment is performed, perform node retrieval ability evaluation according to the obtained qualified node load evaluation score and the processed node retrieval ability data, obtain a node retrieval ability evaluation score, and determine whether to perform node retrieval ability adjustment based on the node retrieval ability evaluation score, and the node retrieval ability evaluation score is used to evaluate the retrieval ability during the retrieval of the data stored in the specified node.

[0021] One or more technical solutions provided in the embodiment of the present application have at least the following technical effects or advantages:

[0022] 1. By performing node load anomaly assessment to obtain a node load assessment score, when the node load assessment score is not within the preset node load range, node dynamic migration assessment and node load anomaly adjustment are carried out. If node load anomaly adjustment is not performed, node retrieval ability assessment is carried out to obtain a node retrieval ability assessment score and it is judged whether to perform node retrieval ability adjustment, thereby improving the storage stability during the efficient storage of distributed data by the programming platform, and further improving the storage accuracy during the efficient storage of distributed data by the programming platform, effectively solving the problem of inaccurate storage during the efficient storage of distributed data by the programming platform in the prior art.

[0023] 2. By processing the node load impact data and the preset node processing weight to obtain a node load value, and then combining the obtained node load value and the preset load weight to obtain a node load assessment score, thereby realizing the accurate assessment of the node load situation, and further improving the assessment accuracy of the load situation of the specified node storing data during the storage process.

[0024] 3. By comparing the obtained node migration assessment score with the preset migration ability threshold, when the node migration assessment score is not lower than the preset migration ability threshold, migration ability optimization is carried out, and when the node migration assessment score is lower than the preset migration ability threshold, migration ability adjustment is carried out, thereby realizing the dynamic adjustment of the migration ability during the node migration process, and further improving the migration ability during the node migration process of the specified node storing data. BRIEF DESCRIPTION OF THE DRAWINGS

[0025] Figure 1 It is a flowchart of the efficient distributed data storage and retrieval method provided by the embodiment of the present application;

[0026] Figure 2 It is a change statistical chart of the first influence degree of node migration - the first processing value of node migration provided by the embodiment of the present application;

[0027] Figure 3 It is a schematic structural diagram of the efficient distributed data storage and retrieval system provided by the embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0028] Embodiments of the present application provide an efficient distributed data storage and retrieval method and system, which solve the problem of inaccurate storage in the process of efficient storage of distributed data by a programming platform. By evaluating node load anomalies, a node load evaluation score is obtained, and then it is determined whether to adjust node load anomalies based on the node load evaluation score. When the node load evaluation score is not within the preset node load range, a node dynamic migration evaluation is performed based on the obtained node dynamic migration data, and a node migration evaluation score is obtained. Then, node load anomaly adjustment is performed based on the node migration evaluation score. If no node load anomaly adjustment is performed, a node retrieval ability evaluation is performed based on the obtained qualified node load evaluation score and the processed node retrieval ability data, and a node retrieval ability evaluation score is obtained. Finally, it is determined whether to adjust the node retrieval ability based on the node retrieval ability evaluation score, achieving an improvement in storage accuracy in the process of efficient storage of distributed data by a programming platform.

[0029] The technical solution in the embodiments of the present application is to solve the problem of inaccurate storage in the process of efficient storage of distributed data by the above programming platform. The general idea is as follows:

[0030] By evaluating node load anomalies to obtain a node load evaluation score, when the node load evaluation score is not within the preset node load range, a node dynamic migration evaluation and node load anomaly adjustment are performed. If no node load anomaly adjustment is performed, a node retrieval ability evaluation is performed to obtain a node retrieval ability evaluation score and it is determined whether to adjust the node retrieval ability, achieving the effect of improving storage accuracy in the process of efficient storage of distributed data by a programming platform.

[0031] To better understand the above technical solution, the above technical solution will be described in detail below in conjunction with the accompanying drawings of the specification and specific embodiments.

[0032] Such as Figure 1As shown in the figure, it is a flowchart of an efficient distributed data storage and retrieval method provided by an embodiment of the present application. The method includes the following steps: Node load anomaly assessment: Based on the obtained node load impact data, perform node load anomaly assessment to obtain a node load assessment score and determine whether to perform node load anomaly adjustment. The node load assessment score is used to evaluate the load condition of the node during the storage process of the specified node's stored data. The specified node's stored data refers to the distributed data of the programming platform stored on the node; Node dynamic migration assessment: When the node load assessment score is not within the preset node load range obtained from the database, based on the obtained node dynamic migration data, perform node dynamic migration assessment to obtain a node migration assessment score, and perform node load anomaly adjustment based on the node migration assessment score. The node migration assessment score is used to evaluate the migration ability during the node migration process of the specified node's stored data. The node load anomaly adjustment includes migration ability optimization and migration ability adjustment. Migration ability optimization means optimizing the storage data migration assessment score that is not lower than the preset migration ability threshold, and migration ability adjustment means adjusting the storage data migration assessment score of the node with a migration ability lower than the preset migration ability threshold; Node retrieval ability assessment: If no node load anomaly adjustment is performed, based on the obtained qualified node load assessment score and the processed node retrieval ability data, perform node retrieval ability assessment to obtain a node retrieval ability assessment score, and determine whether to perform node retrieval ability adjustment based on the node retrieval ability assessment score. The node retrieval ability assessment score is used to evaluate the retrieval ability during the retrieval process of the specified node's stored data. The qualified node load assessment score refers to the node load assessment score within the preset node load range.

[0033] It should be added that the node load impact data includes the total number of storage code bits and the node compliance, and the node compliance is greater than 0; the node compliance includes the first node load compliance, the second node load compliance, the third node load compliance, and the fourth node load compliance; the node dynamic migration data includes the first node migration impact, the second node migration impact, and the third node migration impact; the node retrieval ability data includes the first node retrieval impact, the second node retrieval impact, the third node retrieval impact, and the fourth node retrieval impact.

[0034] The values corresponding to the node load affecting data, the node dynamically migrating data, and the node retrieval ability data are all greater than 0; the total number of storage code bits represents the total number of binary codes stored by the node; the first compliance degree of the node load is used to reflect the compliance between the maximum threshold of the preset task quantity and the number of tasks running on the node; the second compliance degree of the node load is used to reflect the compliance between the maximum threshold of the preset number of read / write disks and the number of read / write disks of the node; the third compliance degree of the node load is used to reflect the compliance between the maximum threshold of the preset connection quantity and the number of client connections; the fourth compliance degree of the node load is used to reflect the compliance between the maximum threshold of the preset code reading speed and the code reading speed of the memory.

[0035] The node dynamically migrating data is used to reflect the influence of the node migration parameter on the load evaluation difference evaluation value; the node migration parameter includes the node migration success rate, the node migration data volume, and the node migration speed; the load evaluation difference evaluation value is used to reflect the difference degree between the node load evaluation score outside the preset node load range and the preset node load average threshold.

[0036] The first influence degree of node retrieval is used to reflect the influence of the node retrieval difference evaluation value by the maximum threshold of the preset retrieval data volume and the difference of the node retrieval data volume; the node retrieval difference evaluation value is used to reflect the difference degree between the node retrieval speed and the maximum threshold of the preset node retrieval speed; the second influence degree of node retrieval is used to reflect the influence of the node retrieval difference evaluation value by the node retrieval success rate; the third influence degree of node retrieval is used to reflect the influence of the node retrieval difference evaluation value by the maximum threshold of the preset concurrent retrieval number and the difference of the node concurrent retrieval number; the fourth influence degree of node retrieval is used to reflect the influence of the node retrieval difference evaluation value by the qualified node load evaluation score.

[0037] In this embodiment, there is a relationship of mutual influence among the node load evaluation score, the node migration evaluation score, and the node retrieval ability evaluation score. The node load evaluation score is the basis for node load anomaly adjustment and node retrieval ability evaluation. When the node load evaluation score is not within the preset node load range, it indicates that there is an anomaly in the node load, and the node load anomaly is adjusted through the obtained node migration evaluation score; when the node load evaluation score is within the preset node load range, the node retrieval ability is evaluated; when the monitored node retrieval ability evaluation score is not lower than the preset average threshold of the node retrieval ability, the node retrieval ability is adjusted. The adjustment of the node retrieval ability directly affects the storage, load processing ability, migration ability, and retrieval ability of the data stored on the node, which helps to improve the stability of storage during the efficient storage of distributed data by the programming platform. For example, when the programming platform needs to centrally store code data on the node, the node load anomaly evaluation and node dynamic migration evaluation are performed to balance the load and optimize the storage performance, and the node retrieval ability evaluation is performed to optimize the retrieval performance, thereby achieving an improvement in the storage accuracy during the efficient storage of distributed data by the programming platform.

[0038] It should be noted that the number of node running tasks, the number of node read / write disks, the number of client connections, and the total number of stored code bits at the first preset time point are obtained through the API (Application Programming Interface) of the preset programming platform. The code reading speed of the memory at the first preset time point is obtained through the read / write operation using the API of the preset programming platform. The node migration success rate, the amount of migrated data, and the migration speed at the second preset time point are obtained by tracking and recording the migration process through a log analysis tool (such as ELK Stack, Prometheus, etc.). The retrieval speed, the amount of retrieved data, and the number of concurrent retrievals at the third preset time point are obtained by querying and obtaining through the API of the preset programming platform.

[0039] It should be noted that the first preset time point represents adjacent time points within the preset first time period set by the preset personnel, the second preset time point represents adjacent time points within the preset second time period set by the preset personnel, and the third preset time point represents adjacent time points within the preset third time period set by the preset personnel; the data of node load impact, node dynamic migration data, and node retrieval ability data are all data obtained at the corresponding adjacent time points.

[0040] Furthermore, the specific method for obtaining the node load evaluation score is as follows: the node load value is obtained by processing the node load impact data and the preset node processing weight; the node load value includes the node first load value, the node second load value, the node third load value, and the node fourth load value, that is, FQ a1 、FQ a2, FQ a3 and FQ a4 ; Obtain the node load evaluation score by combining the obtained node load value and the preset load weight; the preset load weight includes a preset load first weight, a preset load second weight, and a preset load third weight; the preset load first weight is used to reflect the influence degree of the node first load value on the node load evaluation score; the preset load second weight is used to reflect the influence degree of the node second load value on the node load evaluation score; the preset load third weight is used to reflect the influence degree of the node third load value on the node load evaluation score.

[0041] Among them, the limit expression of the node load evaluation score is as follows:

[0042]

[0043] In the formula, FZPG a represents the node load evaluation score at the a-th first preset time point, a = 1, 2,.., h, a represents the number of the first preset time point, h represents the total number of the first preset time points, and the first preset time point represents the preset time point during the storage process of the node stored data, FQ a1 represents the node first load value at the a-th first preset time point, FQ a2 represents the node second load value at the a-th first preset time point, FQ a3 represents the node third load value at the a-th first preset time point, FQ a4 represents the node fourth load value at the a-th first preset time point, e represents the natural constant, a 1 represents the preset load first weight, a 2 represents the preset load second weight, a 3 represents the preset load third weight.

[0044] YX 1a represents the node load first compliance at the a-th first preset time point, YX 2a represents the node load second compliance at the a-th first preset time point, YX 3a represents the node load third compliance at the a-th first preset time point, YX 4a represents the node load fourth compliance at the a-th first preset time point, NR a represents the number of node running tasks at the a-th first preset time point, NR max represents the maximum threshold of the preset task quantity, NC max represents the maximum threshold of the preset number of read / write disks, NC a represents the number of node read / write disks at the a-th first preset time point, LJ max represents the maximum threshold of the preset connection quantity, LJa Indicates the number of client connections at the a-th first preset time point, DC max Indicates the maximum threshold of the preset read code speed, DC a Indicates the memory read code speed at the a-th first preset time point, YXWS a Indicates the total number of stored code bits at the a-th first preset time point, YXWS 0 Indicates the maximum threshold of the preset total storage bits, μ 1 Indicates the first load processing weight, μ 2 Indicates the second load processing weight, μ 3 Indicates the third load processing weight, μ 4 Indicates the fourth load processing weight.

[0045] In this embodiment, the aforementioned database is a database established before the design of the efficient distributed data storage and retrieval method provided in the embodiments of the present application for storing various types of set data. The database includes but is not limited to the maximum threshold of the preset task quantity, the maximum threshold of the preset number of read / write disks, the maximum threshold of the preset connection quantity, etc. Various values therein are directly set by technicians; for example, the maximum threshold of the preset task quantity is represented by the maximum value of the number of node processing code tasks in the historical time period in the database, the maximum threshold of the preset number of read / write disks is represented by the maximum value of the number of node read / write disks in the historical time period in the database, the maximum threshold of the preset connection quantity is represented by the maximum value of the number of node connections in the historical time period in the database, the maximum threshold of the preset connection quantity is represented by the maximum value of the number of node connections in the historical time period in the database, the maximum threshold of the preset read code speed is represented by the maximum value of the node read code speed in the historical time period in the database, and the maximum threshold of the preset total storage bits is represented by the maximum value of the total number of node code storage bits in the historical time period in the database.

[0046] It should be added that a mapping set reflecting the mapping relationship between the first load value, the second load value, and the third load value of the node and the corresponding weights is preset in the database. The mapping relationship in this mapping set can be a one-to-one or many-to-one relationship; for example, in this embodiment, the value range of the preset load weight is 0-1. By inputting the real-time first load value, second load value, and third load value of the node into the mapping set, the corresponding preset load first weight, preset load second weight, and preset load third weight can be obtained.

[0047] The preset node processing weights include a first load processing weight, a second load processing weight, a third load processing weight, and a fourth load processing weight; the first load processing weight is used to reflect the influence degree of the first compliance of the node load on the first load value of the node; the second load processing weight is used to reflect the influence degree of the second compliance of the node load on the second load value of the node; the third load processing weight is used to reflect the influence degree of the third compliance of the node load on the third load value of the node; the fourth load processing weight is used to reflect the influence degree of the fourth compliance of the node load on the fourth load value of the node;

[0048] A mapping set reflecting the mapping relationships between the first compliance of the node load, the second compliance of the node load, the third compliance of the node load, the fourth compliance of the node load and the corresponding weights is preset in the database, and the mapping relationships in the mapping set can be one-to-one or many-to-one relationships; for example, in this embodiment, the value range of the preset node processing weight is 0-1. By inputting the real-time first compliance of the node load, the second compliance of the node load, the third compliance of the node load, and the fourth compliance of the node load into the mapping set, the corresponding first load processing weight, second load processing weight, third load processing weight, and fourth load processing weight can be obtained.

[0049] The algorithm of this embodiment combines the analysis of the node load influence data to obtain the node load evaluation score. The greater the first compliance of the node load, the greater the compliance between the maximum threshold of the preset task quantity and the number of tasks running on the node, and then the greater the influence degree on the total number of stored code bits and the maximum threshold of the preset total storage bits, resulting in a decrease in the node load evaluation score; the greater the second compliance of the node load, the greater the compliance between the maximum threshold of the preset number of read / write disks and the number of read / write disks of the node, and then the greater the influence degree on the total number of stored code bits and the maximum threshold of the preset total storage bits, and the node load evaluation score decreases; the greater the third compliance of the node load, the greater the compliance between the maximum threshold of the preset connection quantity and the number of client connections, and then the greater the influence degree on the total number of stored code bits and the maximum threshold of the preset total storage bits, resulting in a decrease in the node load evaluation score; the greater the fourth compliance of the node load, the smaller the compliance between the maximum threshold of the preset code reading speed and the code reading speed of the memory, and then the greater the influence degree on the total number of stored code bits and the maximum threshold of the preset total storage bits, resulting in a decrease in the node load evaluation score. In summary, the node compliance is negatively correlated with the node load evaluation score.

[0050] In the algorithm of this embodiment, the data affected by node load does not exist independently, and there is a mutual correlation between independent variables, which requires comprehensive analysis. An increase in the first compliance of node load may lead to an increase in task load, resulting in a decrease in the processing capacity of the node, which may further exacerbate the access competition of the disk, leading to an increase in the second compliance of node load; an increase in task load may also lead to an increase in network connection load, and then lead to an increase in the third compliance of node load; an increase in the third compliance of node load may lead to congestion and increased latency of network connections, which in turn affects the remote reading of code files, resulting in an increase in the fourth compliance of node load. By analyzing the comprehensive influence between parameters, the accurate evaluation of the load condition of the node during the storage of the specified node's stored data is realized, and then the improvement of the storage accuracy during the efficient storage of distributed data by the programming platform is realized.

[0051] Further, the specific process of obtaining the node migration evaluation score is as follows: perform a logarithmic function processing on the first influence degree of node migration to obtain the first processing value of node migration, that is, QY 1b ; perform an inverse hyperbolic sine function processing on the second influence degree of node migration to obtain the second processing value of node migration, that is, QY 2b ; perform an exponential function processing on the third influence degree of node migration to obtain the third processing value of node migration, that is, QY 3b ; combine the obtained node migration processing values and the preset migration weights to obtain the node migration evaluation score; the node migration processing values include the first processing value of node migration, the second processing value of node migration, and the third processing value of node migration; the preset migration weights include the preset first migration weight and the preset second migration weight; the preset first migration weight is used to reflect the influence degree of the first processing value of node migration on the node migration evaluation score; the preset second migration weight is used to reflect the influence degree of the second processing value of node migration on the node migration evaluation score.

[0052] Among them, the node migration evaluation score is obtained by the following method:

[0053]

[0054] QY 1b = log 2 (2 + DYQY 1b );

[0055] QY 2b = sinh -1 (DYQY 2b );

[0056]

[0057] In the formula, QYNL bDenote the node migration evaluation score at the b-th second preset time point, where b = 1, 2, ..., r. Here, b represents the serial number of the second preset time point, and r represents the total number of the second preset time points. The second preset time point represents a preset time point during the process of a specified node storing data for node migration, QY 1b Denote the first processing value of node migration at the b-th second preset time point, QY 2b Denote the second processing value of node migration at the b-th second preset time point, QY 3b Denote the third processing value of node migration at the b-th second preset time point, where e represents the natural constant, b 1 Denote the preset migration first weight, b 2 Denote the preset migration second weight.

[0058] DYQY 1b Denote the first influence degree of node migration at the b-th second preset time point, DYQY 2b Denote the second influence degree of node migration at the b-th second preset time point, DYQY 3b Denote the third influence degree of node migration at the b-th second preset time point, FZPG b Denote the node load evaluation score outside the preset node load range at the b-th second preset time point, FZPG 0 Denote the preset node load average threshold, CG b Denote the node migration success rate at the b-th second preset time point, QYSJ max Denote the maximum threshold of the preset migration data volume, QYSJ b Denote the node migration data volume at the b-th second preset time point, QS max Denote the maximum threshold of the preset migration speed, QS b Denote the node migration speed at the b-th second preset time point, δ 1 Denote the preset first migration processing weight, δ 2 Denote the preset second migration processing weight, δ 3 Denote the preset third migration processing weight.

[0059] In this embodiment, the maximum threshold of the preset migration data volume is represented by the maximum value of the node migration data volume in the historical time period in the database. The maximum threshold of the preset migration speed is represented by the maximum value of the node migration speed in the historical time period in the database. The preset node load average threshold is represented by the average value of the node load evaluation scores in the historical time period in the database.

[0060] It should be added that a set of mapping sets are preset in the database. The mapping sets reflect the mapping relationships between the first processing value of node migration, the second processing value of node migration and the corresponding weights. The mapping relationships in the mapping sets can be one-to-one or many-to-one relationships. By inputting the real-time first processing value of node migration and the second processing value of node migration into the mapping sets, the corresponding preset first migration weight and preset second migration weight can be obtained. For example, in this embodiment, the value ranges of the preset first migration weight and the preset second migration weight are 0-1.

[0061] The node dynamic migration data is obtained by processing the node migration parameters and the preset migration processing weights. The preset migration processing weights include the preset first migration processing weight, the preset second migration processing weight and the preset third migration processing weight. The preset first migration processing weight is used to reflect the influence degree of the node migration success rate on the first influence degree of node migration. The preset second migration processing weight is used to reflect the influence degree of the node migration data volume on the second influence degree of node migration. The preset third migration processing weight is used to reflect the influence degree of the node migration speed on the third influence degree of node migration.

[0062] A set of mapping sets are preset in the database. The mapping sets reflect the mapping relationships between the node migration success rate, the node migration data volume and the node migration speed and the corresponding weights. The mapping relationships in the mapping sets can be one-to-one or many-to-one relationships. By inputting the real-time node migration success rate, the node migration data volume and the node migration speed into the mapping sets, the corresponding preset first migration processing weight, preset second migration processing weight and preset third migration processing weight can be obtained. For example, in this embodiment, the value ranges of the preset first migration processing weight, the preset second migration processing weight and the preset third migration processing weight are 0-1.

[0063] Specifically, assume that the range of the first influence degree DYQY of node migration 1b is 0.5-0.9. As Figure 2 shown, it is the statistical chart of the change of the first influence degree of node migration - the first processing value of node migration provided by the embodiment of the present application. It can be seen from Figure 2 that as the first influence degree of node migration gradually increases, the first processing value of node migration gradually increases, indicating that the influence of the node migration success rate increases.

[0064] The algorithm of this embodiment combines the analysis of node dynamic migration data to obtain the node migration evaluation score. The greater the first influence degree of node migration means that the load evaluation difference evaluation value is more affected by the node migration success rate, resulting in a decrease in the node migration evaluation score; the greater the second influence degree of node migration means that the load evaluation difference evaluation value is more affected by the maximum threshold of the preset migration data volume and the difference in the node migration data volume, resulting in a decrease in the node migration evaluation score; the greater the third influence degree of node migration means that the load evaluation difference evaluation value is more affected by the maximum threshold of the preset migration speed and the difference in the node migration speed, resulting in a decrease in the node migration evaluation score. In summary, the node dynamic migration data is negatively correlated with the node migration evaluation score.

[0065] In the algorithm of this embodiment, the node dynamic migration data does not exist independently, and there is a mutual correlation between the independent variables, which requires comprehensive analysis. When the node migration success rate is lower, the greater the first influence degree of node migration means that there may be more problems during the migration process, which may lead to an increase in the difference in the migration data volume and the migration speed, and thus lead to a decrease in the node migration evaluation score; the more the node migration data volume, it may lead to a decrease in the storage and processing capabilities of the node, thereby affecting the node migration speed and resulting in a decrease in the node migration evaluation score. By analyzing the comprehensive influence between parameters, the accurate evaluation of the migration ability during the node migration process of storing data in a specified node is realized, and thus the effect of improving the storage accuracy during the efficient storage of distributed data in the programming platform is achieved.

[0066] Further, the specific process of node load abnormal adjustment is as follows: Compare the obtained node migration evaluation score with the preset migration ability threshold obtained from the database; when the node migration evaluation score is not lower than the preset migration ability threshold in the database, perform migration ability optimization; the migration ability optimization is used to optimize the migration ability of the node; when the node migration evaluation score is lower than the preset migration ability threshold, perform migration ability adjustment; when the node migration evaluation score after the migration ability adjustment is still lower than the preset migration ability threshold, send an alarm prompt to the preset personnel; the migration ability adjustment is used to increase the node migration evaluation score lower than the preset migration ability threshold to not lower than the preset migration ability threshold.

[0067] It should be added that the migration ability optimization includes parallel processing and lossless compression; parallel processing means improving the migration speed through parallel processing methods; lossless compression means reducing the migration time through lossless compression algorithms (such as Huffman algorithm); the specific steps for adjusting the migration ability are as follows: S21, perform resume interrupted transfer processing. When the monitored migration evaluation score of the stored data is not lower than the preset migration ability threshold, stop the operation. Otherwise, execute S22. Resume interrupted transfer means ensuring that the transfer can continue after interruption during the migration process through resume interrupted transfer technology; S22, perform segmented migration. When the monitored migration evaluation score of the stored data is not lower than the preset migration ability threshold, stop the operation. Otherwise, send an alarm prompt to the preset personnel. Segmented migration means dividing the migration of the stored data of the specified node into multiple times.

[0068] In this embodiment, the preset node load range is set by the preset personnel, and the preset migration ability threshold is represented by the average value of the migration evaluation scores of the stored data in the historical time period in the database.

[0069] Through parallel processing methods, such as parallelism within instructions, parallelism between instructions, parallelism of task processing, and parallelism of job processing, etc., multiple data sets or data blocks are migrated simultaneously, and these parts are migrated simultaneously on different threads, thereby improving the migration speed and reducing the total migration time; using a lossless compression algorithm (such as Huffman algorithm) can achieve compression by identifying and eliminating redundancy (such as repeated patterns or characters) in the stored data of the specified node, which can reduce the amount of data to be migrated, thereby shortening the migration time and saving bandwidth. For example, at the underlying hardware or compiler level of the programming platform, by using ILP (Instruction-Level Parallelism) technology, through means such as instruction reordering and instruction merging, parallel execution of instructions is achieved within a single processor core.

[0070] Marking the interruption point during the migration of the stored data of the specified node through resume interrupted transfer technology and continuing the migration from the interruption point helps improve the reliability and efficiency of the migration; dividing the stored data of the specified node into small parts for migration through segmented migration helps ensure the efficiency of the migration of the stored data of the specified node; thereby improving the storage accuracy during the efficient storage of distributed data on the programming platform.

[0071] Furthermore, the specific process of obtaining the node retrieval ability evaluation score is as follows: performing exponential function processing on the node retrieval ability data to obtain the node retrieval processing value; the node retrieval processing value includes the first retrieval processing value, the second retrieval processing value, the third retrieval processing value, and the fourth retrieval processing value, that is, DYJS 1c 、DYJS 2c 、DYJS 3c and DYJS 4cCombining the node retrieval processing value and the preset retrieval weight to obtain the node retrieval ability evaluation score; the preset retrieval weight includes a preset retrieval first weight, a preset retrieval second weight, and a preset retrieval third weight; the preset retrieval first weight is used to reflect the influence degree of the first retrieval processing value on the node retrieval ability evaluation score; the preset retrieval second weight is used to reflect the influence degree of the second retrieval processing value on the node retrieval ability evaluation score; the preset retrieval third weight is used to reflect the influence degree of the third retrieval processing value on the node retrieval ability evaluation score.

[0072] Among them, the node retrieval ability evaluation score is obtained by the following method:

[0073]

[0074] In the formula, JSNL c represents the node retrieval ability evaluation score at the c-th third preset time point, c = 1, 2,..., f, c represents the number of the third preset time point, f represents the total number of the third preset time points, and the third preset time point represents the preset time point during the retrieval process of the data stored in the specified node, DYJS 1c represents the first retrieval processing value at the c-th third preset time point, DYJS 2c represents the second retrieval processing value at the c-th third preset time point, DYJS 3c represents the third retrieval processing value at the c-th third preset time point, DYJS 4c represents the fourth retrieval processing value at the c-th third preset time point, e represents the natural constant, c 1 represents the preset retrieval first weight, c 2 represents the preset retrieval second weight, c 3 represents the preset retrieval third weight.

[0075] JS 1c represents the node retrieval first influence degree at the c-th third preset time point, JS 2c represents the node retrieval second influence degree at the c-th third preset time point, JS 3c represents the node retrieval third influence degree at the c-th third preset time point, JS 4c represents the node retrieval fourth influence degree at the c-th third preset time point, JSSD c represents the node retrieval speed at the c-th third preset time point, JSSD max The maximum threshold of the preset node retrieval speed, JSL max The maximum threshold of the preset retrieval data volume, JSL c represents the node retrieval data volume at the c-th third preset time point, JSCG cDenote the node retrieval success rate at the c-th third preset time point, JSBF max The maximum threshold of the preset concurrent retrieval number, JSBF c Denote the node concurrent retrieval number at the c-th third preset time point, FZPG c Denote the qualified node load evaluation score at the c-th third preset time point, θ 1 Denote the first weight of the preset retrieval processing, θ 2 Denote the second weight of the preset retrieval processing, θ 3 Denote the third weight of the preset retrieval processing, θ 4 Denote the fourth weight of the preset retrieval processing.

[0076] In this embodiment, the maximum threshold of the preset node retrieval speed is represented by the maximum value of the node retrieval speed in the historical time period in the database, the maximum threshold of the preset retrieval data volume is represented by the maximum value of the node retrieval data volume in the historical time period in the database, and the maximum threshold of the preset concurrent retrieval number is represented by the maximum value of the node concurrent retrieval number in the historical time period in the database.

[0077] It should be added that a mapping set reflecting the mapping relationship between the first retrieval processing value, the second retrieval processing value, the third retrieval processing value and the corresponding weights is preset in the database. The mapping relationship in this mapping set can be a one-to-one or many-to-one relationship; for example, in this embodiment, the value range of the preset retrieval weight is 0-1. By inputting the real-time first retrieval processing value, the second retrieval processing value and the third retrieval processing value into the mapping set, the corresponding preset retrieval first weight, preset retrieval second weight and preset retrieval third weight can be obtained.

[0078] The first influence degree of node retrieval is obtained by processing the node retrieval data volume, the first weight of the preset retrieval processing and the node retrieval speed at the third preset time point; the second influence degree of node retrieval is obtained by processing the node retrieval success rate, the second weight of the preset retrieval processing and the node retrieval speed at the third preset time point; the third influence degree of node retrieval is obtained by processing the node concurrent retrieval number, the third weight of the preset retrieval processing and the node retrieval speed at the third preset time point; the fourth influence degree of node retrieval is obtained by processing the qualified node load evaluation score, the fourth weight of the preset retrieval processing and the node retrieval speed at the third preset time point; the first weight of the preset retrieval processing is used to reflect the influence degree of the node retrieval data volume on the first influence degree of node retrieval; the second weight of the preset retrieval processing is used to reflect the influence degree of the node retrieval success rate on the second influence degree of node retrieval; the third weight of the preset retrieval processing is used to reflect the influence degree of the node concurrent retrieval number on the third influence degree of node retrieval; the fourth weight of the preset retrieval processing is used to reflect the influence degree of the qualified node load evaluation score on the fourth influence degree of node retrieval.

[0079] A mapping set reflecting the mapping relationship between a set of node retrieval data volume, node retrieval success rate, node concurrent retrieval number, and qualified node load evaluation score and corresponding weights is preset in the database. The mapping relationship in this mapping set can be one-to-one or many-to-one. For example, in this embodiment, the value ranges of the preset retrieval processing first weight, preset retrieval processing second weight, preset retrieval processing third weight, and preset retrieval processing fourth weight are 0-1. By inputting the real-time node retrieval data volume, node retrieval success rate, node concurrent retrieval number, and qualified node load evaluation score into the mapping set, the corresponding preset retrieval processing first weight, preset retrieval processing second weight, preset retrieval processing third weight, and preset retrieval processing fourth weight can be obtained.

[0080] Specifically, assume the first retrieval processing value DYJS 1c ranges from 1.15 to 1.35, the second retrieval processing value DYJS 2c ranges from 1.15 to 1.35, the third retrieval processing value DYJS 3c ranges from 1.15 to 1.35, the fourth retrieval processing value DYJS 4c ranges from 1.15 to 1.35, and the preset retrieval first weight c 1 、preset retrieval second weight c 2 、preset retrieval third weight c 3 are fixed at 0.3, 0.3, and 0.4 respectively. As shown in Table 1, it is the change statistical table of the node retrieval ability evaluation score provided by the embodiment of the present application:

[0081] Table 1 Change statistical table of node retrieval ability evaluation score

[0082]

[0083]

[0084] As can be seen from the above table, as the node retrieval processing value gradually increases, the node retrieval ability evaluation score gradually decreases, indicating that the retrieval ability during the retrieval process of storing data in the specified node gradually declines.

[0085] The algorithm of this embodiment combines the data analysis of the node retrieval ability to obtain the node retrieval ability evaluation score. The greater the first influence degree of node retrieval means that the node retrieval difference evaluation value is more affected by the maximum threshold of the preset retrieval data volume and the difference in node retrieval data volume, resulting in a decrease in the node retrieval ability evaluation score; the greater the second influence degree of node retrieval means that the node retrieval difference evaluation value is more affected by the node retrieval success rate, resulting in a decrease in the node retrieval ability evaluation score; the greater the third influence degree of node retrieval means that the node retrieval difference evaluation value is more affected by the maximum threshold of the preset concurrent retrieval number and the difference in node concurrent retrieval numbers, resulting in a decrease in the node retrieval ability evaluation score; the greater the fourth influence degree of node retrieval means that the node retrieval difference evaluation value is more affected by the qualified node load evaluation score, resulting in a decrease in the node retrieval ability evaluation score. In summary, the node retrieval ability data is negatively correlated with the node retrieval ability evaluation score.

[0086] In the algorithm of this embodiment, the node retrieval ability data does not exist independently, and there is a mutual correlation between the independent variables, which requires comprehensive analysis. When the node retrieval data volume is larger, more time and resources may be required to process the retrieval request, increasing the possibility of retrieval errors, which in turn leads to an increase in the second influence degree of node retrieval and a decrease in the node retrieval ability evaluation score; when the node concurrent retrieval number is smaller, the efficiency of the node in processing retrieval requests is higher, which may lead to an increase in the node retrieval success rate; the larger the qualified node load evaluation score, the greater the possible node retrieval speed sum, and the higher the possible node retrieval success rate, which in turn leads to an increase in the node retrieval success rate.

[0087] By analyzing the comprehensive influence between parameters, the accurate evaluation of the retrieval ability during the retrieval process of storing data in a specified node is realized, and thus the effect of improving the storage accuracy during the efficient storage of distributed data in the programming platform is achieved.

[0088] Further, the specific process of judging whether to adjust the node retrieval ability based on the node retrieval ability evaluation score is as follows: Judge whether the node retrieval ability evaluation score is not lower than the preset average threshold of node retrieval ability; if so, it indicates that the node retrieval ability is qualified and no adjustment of the node retrieval ability is made; otherwise, it indicates that the node retrieval ability is unqualified and the node retrieval ability is adjusted; when the node retrieval ability is adjusted and the node retrieval ability evaluation score is still lower than the preset average threshold of node retrieval ability, an alarm prompt is sent to the preset personnel; the adjustment of the node retrieval ability includes performing heuristic search and performing matching optimization; heuristic search means improving the retrieval efficiency of the node through the Dijkstra algorithm; matching optimization is used to improve the retrieval matching accuracy and speed.

[0089] In this embodiment, the preset average threshold of node retrieval ability is represented by the average value of the node retrieval ability evaluation scores in the historical time period in the database; through heuristic search, some paths that are not the optimal solutions can be skipped during the retrieval process. For example, the Dijkstra algorithm uses a priority queue to maintain the currently known shortest paths and gradually expands these paths to find the shortest paths from the source node to all other nodes; matching optimization can be achieved through various means, such as improving the matching algorithm, optimizing the index structure, increasing the cache, etc. During the retrieval process, the code information queried by the user on the programming platform is matched with the nodes in the database to find the most suitable nodes. For example, for code queries, a matching algorithm based on syntax trees or lexical analysis, such as the Knuth-Morris-Pratt algorithm, can be used to more accurately match the query with the code in the database, which helps to improve the information matching ability during the retrieval process, and thus achieves the effect of improving the storage accuracy during the efficient storage of distributed data on the programming platform.

[0090] Such as Figure 3As shown in the figure, it is a schematic structural diagram of an efficient distributed data storage and retrieval system provided by an embodiment of the present application. The efficient distributed data storage and retrieval system provided by the embodiment of the present application includes a node load anomaly evaluation module, a node dynamic migration evaluation module, and a node retrieval ability evaluation module: Among them, the node load anomaly evaluation module is used to evaluate the node load anomaly based on the obtained node load impact data, obtain a node load evaluation score, and determine whether to perform node load anomaly adjustment. The node load evaluation score is used to evaluate the load condition of the node during the storage of the specified node's stored data. The specified node's stored data represents the distributed data of the programming platform stored on the node; the node dynamic migration evaluation module is used to perform node dynamic migration evaluation based on the obtained node dynamic migration data when the node load evaluation score is not within the preset node load range obtained from the database, obtain a node migration evaluation score, and perform node load anomaly adjustment based on the node migration evaluation score. The node migration evaluation score is used to evaluate the migration ability during the node migration of the specified node's stored data. The node load anomaly adjustment includes migration ability optimization and migration ability adjustment. Migration ability optimization means optimizing the storage data migration evaluation score that is not lower than the preset migration ability threshold. Migration ability adjustment means adjusting the storage data migration evaluation score of the node lower than the preset migration ability threshold; the node retrieval ability evaluation module is used to, if no node load anomaly adjustment is performed, perform node retrieval ability evaluation based on the obtained qualified node load evaluation score and the processed node retrieval ability data, obtain a node retrieval ability evaluation score, and determine whether to perform node retrieval ability adjustment based on the node retrieval ability evaluation score. The node retrieval ability evaluation score is used to evaluate the retrieval ability during the retrieval of the specified node's stored data. The qualified node load evaluation score represents the node load evaluation score within the preset node load range.

[0091] In this embodiment, when the node load evaluation score obtained by the node load anomaly evaluation module is within the preset node load range, the node retrieval ability evaluation module obtains the node retrieval ability evaluation score. When the node retrieval ability evaluation score is lower than the preset average node retrieval ability threshold, node retrieval ability adjustment is performed to improve the retrieval ability of the node; when the node load evaluation score is not within the preset node load range, node load anomaly adjustment is performed based on the node migration evaluation score obtained by the node dynamic migration evaluation module to improve the migration ability of the node; the node load anomaly evaluation module, the node dynamic migration evaluation module, and the node retrieval ability evaluation module interact with each other, which helps to optimize the performance during the efficient storage of distributed data on the programming platform, and further improves the storage accuracy during the efficient storage of distributed data on the programming platform.

[0092] In summary, by performing node load anomaly assessment to obtain a node load assessment score, when the node load assessment score is not within the preset node load range, node dynamic migration assessment and node load anomaly adjustment are carried out. If node load anomaly adjustment is not performed, node retrieval ability assessment is carried out to obtain a node retrieval ability assessment score and it is judged whether to perform node retrieval ability adjustment, thereby improving the storage stability in the process of the programming platform for efficient storage of distributed data, and further improving the storage accuracy in the process of the programming platform for efficient storage of distributed data, effectively solving the problem of inaccurate storage in the prior art in the process of the programming platform for efficient storage of distributed data.

[0093] Those skilled in the art should understand that the embodiments of the present invention can be provided as methods, systems, or computer program products. Therefore, the present invention can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present invention can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0094] The present invention is described with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to embodiments of the present invention. It should be understood that each flow and / or block in the flowcharts and / or block diagrams, and the combination of flows and / or blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, such that the instructions executed by the processor of the computer or other programmable data processing devices generate means for implementing the functions specified in Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.

[0095] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, such that the instructions stored in the computer-readable memory generate a manufactured article including instruction means, and the instruction means implements the functions specified in Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.

[0096] These computer program instructions can also be loaded onto a computer or other programmable data processing device, such that a series of operation steps are performed on the computer or other programmable device to generate a computer-implemented process, so that the instructions executed on the computer or other programmable device provide means for implementing the functions specified in Figure 1Steps of the functions specified in one or more processes and / or boxes Figure 1 Steps of the functions specified in one or more boxes

[0097] Although the preferred embodiments of the present invention have been described, those skilled in the art can make additional changes and modifications to these embodiments once they learn the basic creative concept. Therefore, the appended claims are intended to be construed to include the preferred embodiments and all changes and modifications that fall within the scope of the present invention.

[0098] Obviously, those skilled in the art can make various changes and modifications to the present invention without departing from the spirit and scope of the present invention. Thus, if these modifications and variations of the present invention fall within the scope of the claims of the present invention and their equivalent technologies, the present invention is also intended to include these modifications and variations.

Claims

1. An efficient distributed data storage and retrieval method, characterized in that: The following steps are involved: Perform node load anomaly assessment based on the acquired node load impact data, obtain a node load assessment score and determine whether to perform node load anomaly adjustment, wherein the node load assessment score is used to assess the node load condition of the designated node during the storage process of the stored data; When the node load evaluation score is not within the preset node load range obtained from the database, a node dynamic migration evaluation is performed based on the obtained node dynamic migration data to obtain a node migration evaluation score, and node load abnormality adjustment is performed based on the node migration evaluation score, wherein the node migration evaluation score is used to evaluate the migration capability of the designated node storing data during the node migration process; If no abnormal node load adjustment is performed, a node retrieval capability assessment is performed based on the obtained qualified node load assessment score and the processed node retrieval capability data to obtain a node retrieval capability assessment score. Based on the node retrieval capability assessment score, it is determined whether to adjust the node retrieval capability. The node retrieval capability assessment score is used to evaluate the retrieval capability of the specified node during the retrieval process of stored data.

2. The efficient distributed data storage and retrieval method according to claim 1, characterized in that: The node load impact data includes the total number of bits of storage code and node compliance; The node compliance includes a first node load compliance, a second node load compliance, a third node load compliance and a fourth node load compliance; The node dynamic migration data includes a first impact degree of node migration, a second impact degree of node migration, and a third impact degree of node migration; The node retrieval capability data includes a first influence degree of node retrieval, a second influence degree of node retrieval, a third influence degree of node retrieval, and a fourth influence degree of node retrieval; The first node load compliance is used to reflect the compliance between the preset maximum threshold of the number of tasks and the number of tasks running on the node; The second node load compliance is used to reflect the compliance between the preset maximum threshold of the number of read-write disks and the number of read-write disks of the node; The third node load compliance is used to reflect the compliance between the preset maximum threshold of the number of connections and the number of client connections; The fourth node load compliance is used to reflect the compliance between the preset maximum threshold of the code reading speed and the memory code reading speed; The node dynamic migration data is used to reflect the impact of the load assessment difference evaluation value on the node migration parameters; The node migration parameters include node migration success rate, node migration data volume and node migration speed; The load evaluation difference evaluation value is used to reflect the difference between the node load evaluation score that is not within the preset node load range and the preset node load average threshold; The node retrieval first influence degree is used to reflect the influence of the node retrieval difference evaluation value on the preset retrieval data volume maximum threshold and the node retrieval data volume difference; The node retrieval difference evaluation value is used to reflect the difference between the node retrieval speed and the preset node retrieval speed maximum threshold; The node retrieval second influence degree is used to reflect the influence of the node retrieval difference evaluation value on the node retrieval success rate; The node search third impact degree is used to reflect the influence of the node search difference evaluation value on the preset concurrent search number maximum threshold and the node concurrent search number difference; The node retrieval fourth influence degree is used to reflect the influence of the node retrieval difference evaluation value on the qualified node load evaluation score.

3. The efficient distributed data storage and retrieval method according to claim 1, characterized in that: The specific method for obtaining the node load evaluation score is as follows: The node load value is obtained by processing the node load impact data and the preset node processing weight; The node load values ​​include a first node load value, a second node load value, a third node load value, and a fourth node load value; The node load evaluation score is obtained by combining the obtained node load value and the preset load weight; The preset load weight includes a preset load first weight, a preset load second weight and a preset load third weight.

4. The efficient distributed data storage and retrieval method according to claim 3, characterized in that: The limiting expression of the node load evaluation score is as follows: In the formula, FZPG a represents the node load evaluation score at the ath first preset time point, a=1, 2, .., h, a represents the number of the first preset time point, h represents the total number of the first preset time points, and the first preset time point represents the preset time point in the storage process of the node storing data, FQ a1 represents the first load value of the node at the ath first preset time point, FQ a2 represents the second load value of the node at the ath first preset time point, FQ a3 represents the third load value of the node at the ath first preset time point, FQ a4 represents the fourth load value of the node at the ath first preset time point, e represents a natural constant, a1 represents the first weight of the preset load, a2 represents the second weight of the preset load, and a3 represents the third weight of the preset load.

5. The efficient distributed data storage and retrieval method according to claim 2, characterized in that: The specific process of obtaining the node migration evaluation score is as follows: Performing logarithmic function processing on the first impact degree of node migration to obtain a first processing value of node migration; Performing inverse hyperbolic sine function processing on the second influence of node migration to obtain a second processed value of node migration; Performing exponential function processing on the third impact degree of node migration to obtain a third processing value of node migration; The node migration evaluation score is obtained by combining the obtained node migration processing value and the preset migration weight; The node migration processing value includes a node migration first processing value, a node migration second processing value and a node migration third processing value; The preset migration weight includes a preset migration first weight and a preset migration second weight.

6. The efficient distributed data storage and retrieval method according to claim 1, characterized in that: The specific process of adjusting the abnormal node load is as follows: When the node migration evaluation score is not lower than the preset migration capacity threshold in the database, the migration capacity is optimized; When the node migration evaluation score is lower than the preset migration capacity threshold, the migration capacity is adjusted; When the node migration assessment score after the migration capacity adjustment is still lower than the preset migration capacity threshold, an alarm is sent to the preset personnel.

7. The efficient distributed data storage and retrieval method according to claim 6, characterized in that: The migration capability optimization includes parallel processing and lossless compression; The specific steps of adjusting the migration capability are: S21, performing breakpoint resume processing, when the monitored storage data migration evaluation score is not lower than the preset migration capability threshold, stop the operation, otherwise execute S22; S22, performing migration in stages, when the monitored storage data migration assessment score is not lower than the preset migration capability threshold, stopping the operation, otherwise sending an alarm prompt to the preset personnel.

8. The efficient distributed data storage and retrieval method according to claim 1, characterized in that: The specific process of obtaining the node retrieval capability evaluation score is as follows: Performing exponential function processing on the node retrieval capability data to obtain the node retrieval processing value; The node retrieval processing value includes a first retrieval processing value, a second retrieval processing value, a third retrieval processing value and a fourth retrieval processing value; The node retrieval capability evaluation score is obtained by combining the node retrieval processing value and the preset retrieval weight; The preset retrieval weight includes a preset retrieval first weight, a preset retrieval second weight and a preset retrieval third weight.

9. The efficient distributed data storage and retrieval method according to claim 8, characterized in that: The specific process of determining whether to adjust the node retrieval capability based on the node retrieval capability evaluation score is as follows: Determine whether the node retrieval capability evaluation score is not lower than the preset node retrieval capability average threshold; If yes, it indicates that the node retrieval capability is qualified, and no node retrieval capability adjustment is performed; On the contrary, it indicates that the node retrieval capability is unqualified and the node retrieval capability should be adjusted; When the node retrieval capability evaluation score is still lower than the preset node retrieval capability average threshold after the node retrieval capability adjustment, an alarm is sent to the preset personnel; The node retrieval capability adjustment includes performing heuristic search and matching optimization.

10. An efficient distributed data storage and retrieval system, characterized in that: It includes node load anomaly assessment module, node dynamic migration assessment module and node retrieval capability assessment module: The node load anomaly assessment module is used to perform node load anomaly assessment according to the acquired node load impact data, obtain a node load assessment score and determine whether to perform node load anomaly adjustment, and the node load assessment score is used to assess the node load of the designated node during the storage process of the stored data; The node dynamic migration evaluation module is used to perform node dynamic migration evaluation based on the acquired node dynamic migration data to obtain a node migration evaluation score when the node load evaluation score is not within the preset node load range obtained from the database, and perform node load abnormality adjustment based on the node migration evaluation score, wherein the node migration evaluation score is used to evaluate the migration capability of the designated node storing data during node migration; The node retrieval capability assessment module is used to perform node retrieval capability assessment based on the obtained qualified node load assessment score and the processed node retrieval capability data if no abnormal node load adjustment is performed to obtain a node retrieval capability assessment score, and determine whether to adjust the node retrieval capability based on the node retrieval capability assessment score. The node retrieval capability assessment score is used to evaluate the retrieval capability of the specified node during the retrieval process of stored data.

Citation Information

Patent Citations

  • Distributed data storage methods, retrieval methods, systems, and readable storage media

    CN113254505B

  • Distributed storage method, electronic device and medium for designing code data

    CN117909313B