Data development method and device, equipment, storage medium and product

By applying hot and cold storage algorithms and preset time and space algorithms in Taichung in the data, segmenting the storage space and performing resource scheduling, the problems of resource preemption and multi-user collaborative development are solved, and the efficiency and quality of data development are improved.

CN120010773AActive Publication Date: 2025-05-16CHINA MOBILE INFORMATION TECHNOLOGY CO LTD +1
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
CN202510082112.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-17
Publication Date
2025-05-16
Estimated Expiration
2045-01-17

AI Technical Summary

Technical Problem

The resources of the existing data middle platform cannot be isolated when utilizing, resulting in resource preemption and scheduling capabilities, and the inability to achieve collaborative development of multiple users, resulting in low development quality and low efficiency.

Method used

The storage space is divided into multiple node spaces through the hot and cold storage algorithm, and the storage path is determined using the preset time and space algorithm, and the corresponding relationship between data and node space is determined based on the path, and resource scheduling is performed; at the same time, the decision data of the data development statement is determined based on the mapping relationship between the virtual space and the node space, and data development is carried out.

Benefits of technology

Avoid resource seizing problems from three aspects: time, space, and resource optimization and scheduling, improve the efficiency and quality of data development, and improve scheduling capabilities.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120010773A_ABST
    Figure CN120010773A_ABST
Patent Text Reader

Abstract

The invention discloses a data development method and device, equipment, a storage medium and a product, and relates to the technical field of cloud computing, and the method comprises the steps: dividing a storage space into a plurality of node spaces, determining a storage path of to-be-stored data according to a time and space algorithm, determining a corresponding relation between the to-be-stored data and the node space according to the storage path, performing resource scheduling according to the corresponding relation and the priority parameter of each storage path, and storing the to-be-stored data into the node space; determining a current virtual space statement value according to the data development statement and a virtual mapping relationship between the virtual space and the node space; and determining decision data of the data development statement according to the current virtual space statement value, and performing data development according to the decision data. The resource preemption problem is avoided from the three aspects of time, space and resource optimization scheduling, it is guaranteed that the data development task of the data-in-the-data platform is executed more efficiently and more efficiently, and the scheduling capacity is improved from the aspect of resource assurance.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of cloud computing technology, and in particular to a data development method, device, equipment, storage medium and product. Background Art

[0002] In the existing data development solutions based on the data middle platform, the independent space required for storage is separated, and the similar parallel node space and virtual node space are separated in the independent space. The node association of the storage resource is determined based on the virtual identification information of the predefined mapping relationship of the storage resource. The pre-development and development of the virtual space are used to determine that the virtual identification predefined mapping relationship is performed again after the data development is completed, and the virtual backup mapping storage generation of the virtual identification of the storage resource is completed.

[0003] When solving the resource preemption problem in the data center, existing technical solutions avoid data disorder by logically processing user tasks submitted in the same time period, or by encapsulating interfaces and isolating data through technical architecture. However, the huge resources of the existing development technology center cannot be isolated when resources are used. Due to resource preemption, they affect each other, resulting in reduced scheduling capabilities. At the same time, it is impossible to carry out collaborative development and application of multiple users at the same time, resulting in low development quality and reduced development efficiency. Summary of the invention

[0004] The main purpose of this application is to provide a data development method, device, equipment, storage medium and product, aiming to solve the technical problems that various resource points in the data center cannot be isolated and multi-user collaborative development cannot be achieved, resulting in low development quality and low efficiency.

[0005] To achieve the above objectives, the present application proposes a data development method, which comprises:

[0006] Dividing the storage space into a plurality of node spaces according to a hot-cold storage algorithm, wherein the node space includes a virtual space;

[0007] Determine a storage path for the data to be stored according to a preset time algorithm and a preset space algorithm, and determine a corresponding relationship between the data to be stored and the node space according to the storage path, wherein the attributes in the storage path include hot and cold storage classification;

[0008] Perform resource scheduling according to the corresponding relationship and the priority parameter corresponding to each storage path, and store the data to be stored in the node space;

[0009] Determine the current virtual space statement value according to the received data development statement and the virtual mapping relationship between the virtual space and the node space;

[0010] The decision data of the data development statement is determined according to the current virtual space statement value, and data development is performed according to the decision data.

[0011] In one embodiment, the attributes in the storage path include resource pool, resource pool region, hot and cold storage classification, compression type, data directory, time period, and file name;

[0012] The step of determining the storage path of the data to be stored according to the preset time algorithm and the preset space algorithm includes:

[0013] Determine the resource pool, resource pool area, and first hot and cold storage classification according to a preset space algorithm;

[0014] Determine the second cold and hot storage classification, compression type and time period according to a preset time algorithm;

[0015] Encrypting and fusing the first cold and hot storage classification and the second cold and hot storage classification according to preset weights to obtain a cold and hot storage classification;

[0016] The resource pool, the resource pool area, the hot and cold storage classifications and the compression type are encrypted according to a multi-level linear congruential algorithm to obtain a plurality of encrypted ciphertexts, and a storage path is generated according to the plurality of encrypted ciphertexts, the data directory, the time period and the file name.

[0017] In one embodiment, the step of determining the resource pool, the resource pool area, and the first hot and cold storage classification according to a preset space algorithm includes:

[0018] Collecting spatial characteristic attributes of the data to be stored;

[0019] Determine the information entropy of each of the spatial feature attributes to obtain an uncertainty measure corresponding to each of the spatial feature attributes;

[0020] Determine the information gain value of each of the spatial feature attributes according to the information entropy;

[0021] Determine the optimal spatial feature attribute and model parameters corresponding to the decision tree algorithm according to each of the spatial feature attributes and the information gain value;

[0022] The resource pool, the resource pool area and the first hot and cold storage classification are determined according to the decision tree algorithm, the model parameters and the optimal spatial feature attributes.

[0023] In one embodiment, the step of determining the second cold and hot storage classification, compression type and time period according to a preset time algorithm includes:

[0024] Collecting time characteristic attributes of the data to be stored;

[0025] Determine the hidden state and output vector of the current time step corresponding to the sequence of each data storage unit in the storage space according to the recurrent neural network algorithm, the time feature attribute, the input vector of the current time step and the hidden state of the previous time step;

[0026] The second hot and cold storage classification, compression type and time period are determined according to the hidden state and output vector of the current time step.

[0027] In one embodiment, the step of encrypting the resource pool, the resource pool area, the hot and cold storage classifications, and the compression type according to a multi-level linear congruential algorithm to obtain a plurality of encrypted ciphertexts includes:

[0028] Respectively converting the resource pool, the resource pool area, the hot and cold storage classifications, and the compression type into digital sequences;

[0029] The encrypted resource pool, the encrypted resource pool area, the encrypted hot and cold storage classifications and the encrypted compression type are determined according to the digital sequence, the preset constant and the encrypted character set size.

[0030] In one embodiment, after determining the decision data of the data development statement according to the current virtual space statement value and performing the data development step according to the decision data, the method further includes:

[0031] Adding resource extraction marking points to the decision data;

[0032] The node resources having the resource extraction mark point in the storage space are removed.

[0033] In addition, to achieve the above purpose, the present application also proposes a data development device, the data development device comprising:

[0034] A storage space segmentation module, used to segment the storage space into a plurality of node spaces according to a hot-cold storage algorithm, wherein the node space includes a virtual space;

[0035] A storage path determination module, used to determine a storage path of the data to be stored according to a preset time algorithm and a preset space algorithm, and determine a corresponding relationship between the data to be stored and the node space according to the storage path, wherein the attributes in the storage path include hot and cold storage classification;

[0036] A resource scheduling module, used to perform resource scheduling according to the corresponding relationship and the priority parameter corresponding to each storage path, and store the data to be stored in the node space;

[0037] A statement value determination module, used to determine the current virtual space statement value according to the received data development statement and the virtual mapping relationship between the virtual space and the node space;

[0038] The data development module is used to determine the decision data of the data development statement according to the current virtual space statement value, and perform data development according to the decision data.

[0039] In addition, to achieve the above objectives, the present application also proposes a data development device, which includes: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the computer program is configured to implement the steps of the data development method described above.

[0040] In addition, to achieve the above-mentioned purpose, the present application also proposes a storage medium, which is a computer-readable storage medium, and a computer program is stored on the storage medium. When the computer program is executed by a processor, the steps of the data development method described above are implemented.

[0041] In addition, to achieve the above-mentioned purpose, the present application also provides a computer program product, which includes a computer program, and when the computer program is executed by a processor, the steps of the data development method described above are implemented.

[0042] The present application provides a data development method, which divides the storage space into multiple node spaces according to the hot and cold storage algorithm, determines the storage path of the data to be stored according to the time and space algorithm, and determines the corresponding relationship between the data to be stored and the node space according to the storage path, performs resource scheduling according to the corresponding relationship and the priority parameters corresponding to each storage path, and stores the data to be stored in the node space; determines the current virtual space statement value according to the received data development statement and the virtual mapping relationship between the virtual space and the node space; determines the decision data of the data development statement according to the current virtual space statement value, and performs data development according to the decision data. The present application avoids the problem of resource preemption from the three aspects of time, space, and resource optimization scheduling, ensures more efficient and high-quality execution of data development tasks in the data middle station, and improves scheduling capabilities from the aspect of resource guarantee. BRIEF DESCRIPTION OF THE DRAWINGS

[0043] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the present application and, together with the description, serve to explain the principles of the present application.

[0044] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, for ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative labor.

[0045] Figure 1A flowchart of the first embodiment of the data development method of this application is provided;

[0046] Figure 2 A schematic diagram of the unit structure in the data development system for this application;

[0047] Figure 3 A schematic diagram of the storage unit structure in the data development method of this application;

[0048] Figure 4 A flow chart of the second embodiment of the data development method of this application is provided;

[0049] Figure 5 A schematic diagram of the decision tree for calculating the resource pool in the data development method for this application;

[0050] Figure 6 A schematic diagram of the calculation decision tree for the resource pool area in the data development method for this application;

[0051] Figure 7 A schematic diagram of a computational decision tree for the first hot and cold storage classification in the data development method of this application;

[0052] Figure 8 Schematic diagram of the structure of a simple recurrent network used in the data development method for this application;

[0053] Fig. 9 A schematic diagram of the update process of a simple recurrent network in the data development method of this application;

[0054] Fig.10 An example diagram of spatiotemporal fusion in the data development method for this application;

[0055] Fig.11 This is a schematic diagram of the module structure of the data development device according to an embodiment of the present application;

[0056] Fig.12 Schematic diagram of the device structure of the hardware operating environment involved in the data development method in the embodiment of the present application.

[0057] The purpose, features and advantages of this application will be further described in conjunction with the embodiments and with reference to the accompanying drawings. DETAILED DESCRIPTION

[0058] It should be understood that the specific embodiments described herein are only used to explain the technical solutions of the present application and are not used to limit the present application.

[0059] In order to better understand the technical solution of the present application, a detailed description will be given below in conjunction with the accompanying drawings and specific implementation methods.

[0060] The main solution of the embodiment of the present application is: dividing the storage space into multiple node spaces according to the hot and cold storage algorithm, and the node space includes a virtual space; determining the storage path of the data to be stored according to the preset time algorithm and the preset space algorithm, and determining the correspondence between the data to be stored and the node space according to the storage path, and the attributes in the storage path include hot and cold storage classifications; performing resource scheduling according to the correspondence and the priority parameters corresponding to each storage path, and storing the data to be stored in the node space; determining the current virtual space statement value according to the received data development statement and the virtual mapping relationship between the virtual space and the node space; determining the decision data of the data development statement according to the current virtual space statement value, and performing data development according to the decision data.

[0061] When solving the resource preemption problem in the data center, the existing technical solutions avoid data disorder by logically processing user tasks submitted in the same time period, or by encapsulating the application programming interface (API) through the technical architecture to isolate data. However, the huge resources of the existing development technology center cannot be isolated when the resources are used. The resource preemption affects each other, resulting in a decrease in scheduling capabilities. At the same time, it is impossible to carry out collaborative development and application of multiple users at the same time, resulting in low development quality and reduced development efficiency.

[0062] The present application provides a solution, by dividing the storage space into multiple node spaces according to the hot and cold storage algorithm, determining the storage path of the data to be stored according to the time and space algorithm, and determining the corresponding relationship between the data to be stored and the node space according to the storage path, performing resource scheduling according to the corresponding relationship and the priority parameters corresponding to each storage path, and storing the data to be stored in the node space; determining the current virtual space statement value according to the received data development statement and the virtual mapping relationship between the virtual space and the node space; determining the decision data of the data development statement according to the current virtual space statement value, and performing data development according to the decision data. The present application avoids the problem of resource preemption from three aspects: time, space, and resource optimization scheduling, ensures more efficient and high-quality execution of data development tasks in the data middle station, and improves scheduling capabilities from the aspect of resource guarantee.

[0063] It should be noted that the execution subject of the method of this embodiment can be a computing service device with data development, network communication and program running functions, such as a tablet computer, a personal computer, a mobile phone, etc.; it can also be a data development device based on a data middle platform with the same or similar functions. This embodiment and the following embodiments will be described by taking a data development device based on a data middle platform as an example.

[0064] Based on this, the present application embodiment provides a data development method, referring to Figure 1 , Figure 1 This is a flowchart of the first embodiment of the data development method of this application.

[0065] In this embodiment, the data development method includes steps S10 to S50:

[0066] Step S10, dividing the storage space into a plurality of node spaces according to the hot-cold storage algorithm, wherein the node space includes a virtual space.

[0067] It should be noted that the data center is a concept based on modern data technology and architecture, which aims to build a unified data platform to solve problems such as data silos, data dispersion and low data quality within the enterprise. Figure 2 The units included in the data development system based on the data center are described, including data acquisition unit, data processing unit, decision model unit, data analysis unit, intelligent control unit, data storage unit, data display unit and data judgment unit. These modules together constitute the basic architecture of the data center. The modules work together to ensure that enterprises can efficiently manage and utilize their data resources. The relationship between the data storage unit and the storage space, virtual space and node space can be referred to in Figure 3 .

[0068] It is understandable that the storage space can be segmented, and the hot and cold storage algorithms are added during the segmentation. The segmentation is performed according to the hot and cold storage, which is superior to the contribution level space, and is divided into storage node space and virtual space, and each space is divided into multiple node spaces. Among them, multiple node spaces are associated according to the assigned tag vector; the virtual space includes the assigned tag vector value of the corresponding storage node space, the corresponding storage node space data synchronization amount and the temporary storage space; based on the pre-allocated space, a certain space can be pre-allocated before data insertion to avoid frequent data reorganization, and reasonable space pre-allocation can be performed according to the prediction of data growth. Data is sharded based on multiple segmented node spaces, and data can be stored in shards, and different data can be stored in different locations, which can reduce the impact of a single data insertion or deletion operation on the overall data. Among them, independent data segmentation is performed based on convolutional neural networks or recurrent neural networks according to data characteristics and required segmentation characteristics, and distributed algorithm association is associated based on the correlation or degree of association between individual data.

[0069] Step S20, determining a storage path of the data to be stored according to a preset time algorithm and a preset space algorithm, and determining a correspondence between the data to be stored and the node space according to the storage path, wherein the attributes in the storage path include hot and cold storage classifications.

[0070] It can be understood that the data is loaded by the data middle platform content, and the independent data set is divided into independent individual data. Each data set or a row of independent data is regarded as an individual. At the same time, the data is classified, analyzed and planned, and the characteristic attributes and associations that need to be grouped are determined. Then, according to the results of the analysis and planning, the data set is sharded according to the consistent double hash algorithm, and the correspondence between the data to be stored and the node space is determined according to the storage path to obtain the best storage strategy for each data. Individual data can be dispersed into different node spaces to avoid the probability of data collision. However, this step is not implemented on the ground. This step is the basic model data for subsequent global scheduling execution and data encryption services.

[0071] It should be understood that, considering the off-site storage of multiple resource pools, data storage uses a combination of time and space algorithms to segment data, determine the storage path of the data to be stored (including the attributes of hot and cold storage classification), the time attribute of the data is modeled using a recurrent neural network algorithm, and the spatial attribute of the data is modeled using a decision tree algorithm.

[0072] Step S30: performing resource scheduling according to the corresponding relationship and the priority parameter corresponding to each storage path, and storing the data to be stored in the node space.

[0073] It is understandable that after obtaining the best storage strategy for each data, when the global scheduling task is executed, unified scheduling control can be performed according to the corresponding relationship, so that the corresponding file is generated in the corresponding storage area. At the same time, the task priority can be controlled according to the priority parameter corresponding to each storage path to ensure the rational use of resources according to business characteristics. The scheduling control rules are controlled according to the following four-bit code to adapt to different business task scenarios.

[0074]

[0075] For example, the following is the first global resource scheduling process flow, which is suitable for task processing on a daily cycle.

[0076] Scheduling control code: 1211;

[0077] Get the storage location key: a=5, b=3, n=26, and perform reverse decryption of the linear congruential equation;

[0078] Parse array 1 "resource pool", where global parallel tasks are set, and the scheduling tasks between resource pools are independent and executed in parallel;

[0079] Parse array 2 "resource pool region" to obtain the resource pool pod name (pod refers to the smallest deployable computing unit). Obtain resource pool resource information;

[0080] Pull up the free resource task in the resource pool;

[0081] Analyze array three "cold and hot storage classification" and give priority to hot storage tasks;

[0082] Parse array 4 "Compression type" and give priority to uncompressed file tasks;

[0083] Daily periodic tasks are executed according to the above scheduling execution principles. Tasks in each resource pool are executed in parallel to ensure the timeliness of daily tasks. At the same time, daily tasks are executed on hot storage to speed up the calculation. However, hot storage and resource pool performance have relatively large resource overheads. This scheduling control scenario is suitable for business scenarios with high timeliness.

[0084] For example, the following is a second global resource scheduling process flow, which is suitable for task processing in daily and monthly cycles.

[0085] Scheduling control code: 2322;

[0086] Get the storage location key: a=5, b=3, n=26, and perform reverse decryption of the linear congruential equation;

[0087] Parse array 1 "resource pool", set the global serial task here, and schedule tasks between resource pools to be executed serially in order;

[0088] Parse array 2 "resource pool area", obtain the resource pool pod name, and obtain resource pool resource information;

[0089] Tasks in each area of ​​the resource pool are executed serially;

[0090] Analyze array three "hot and cold storage classification" and give priority to cold storage tasks;

[0091] Parse array 4 "Compression type", and give priority to compressed files;

[0092] The monthly cycle tasks have low timeliness requirements. They are executed according to the above scheduling execution principles. The tasks of each resource pool are serialized to save costs to the maximum extent. At the same time, each task is executed on cold storage, with low computing speed but without occupying expensive hot storage resources. Compressed files have slow computing speed but low resource overhead. This scenario is suitable for task scenarios with low timeliness, saving resources to the maximum extent and reducing investment.

[0093] Step S40, determining the current virtual space statement value according to the received data development statement and the virtual mapping relationship between the virtual space and the node space.

[0094] It is understandable that a virtual identification predefined mapping relationship can be performed for individual storage resources in each node space, and the target available area of ​​the storage resource can be determined based on the predefined mapping relationship information. The target available area is the remaining available value of the current virtual node space. When the remaining available value is less than 20% of the total available area, the virtual node space no longer receives data development statements. The virtual backup mapping storage of the storage resource virtual identification is completed based on the link between the virtual space and the node space mapping relationship.

[0095] It should be understood that the data development statement transmitted by the development unit is transmitted to the operation unit based on the data processing module, and the operation unit receives the data development statement transmitted by the data processing module and executes it. The development data node is determined based on the data development statement, the virtual space node mapping data is determined by the virtual mapping relationship between the virtual space and the node space, and the virtual arrangement and storage are performed based on the virtual space node mapping data, the data development statement and the statement timestamp, and the current virtual space statement value is fed back.

[0096] Among them, the development unit includes an identity authentication module, wherein the development unit includes allowing simultaneous multi-user online development, multi-user visual synchronous feedback interface and multi-user multi-party confirmation, wherein the identity authentication module includes the following steps: a. Input the identity token to read the user information, and authenticate the identity through the identity information stored in the login system. If the verification fails, refuse to continue the operation. After the verification passes, enter the b operation step; b. When the verification passes, read the username and password information, complete the identity authentication after the reading verification passes, and complete the login.

[0097] Step S50, determining the decision data of the data development statement according to the current virtual space statement value, and performing data development according to the decision data.

[0098] It is understandable that the current virtual space statement value can be fed back to the development user, and various data development statements can be displayed. The decision of the data development statement is finally completed based on the multi-party determination and decision model, and the data development execution is completed by the virtual space mapping based on the decision data. The virtual space mapping will make a decision to run the decision data development statement based on the corresponding storage node space data synchronization amount, and synchronize the individual data in the storage space.

[0099] In a feasible implementation manner, after step S50, steps S60 to S70 may also be included:

[0100] Step S60: adding resource extraction marking points to the decision data.

[0101] It is understandable that after the data development is completed, the virtual identification pre-defined mapping relationship is performed again to complete the virtual backup mapping storage generation of the storage resource virtual identification. The resource collection classification is first used to perform secondary collection and classification on the resource data collected and integrated by the data acquisition processing system, and the resource extraction mark point is used to perform extraction marking of the collected resources.

[0102] Step S70: remove the node resources having the resource extraction mark point in the storage space.

[0103] It is understandable that the node resources of the extracted tags are compared with the database to create new tag points and remove duplicate tags for storage, and finally the resource integration pre-stored is integrated and uploaded for storage. The developed data is re-tagged, based on independent scheduling and orchestration management, while completing the improvement of data isolation and scheduling capabilities, it not only removes the resource usage and occupancy rate of redundant data, improves data resource utilization, but also realizes accurate search of resource data of the entire management system and accurate control of effectiveness.

[0104] This embodiment provides a data development method, which divides the storage space into multiple node spaces according to the hot and cold storage algorithm, determines the storage path of the data to be stored according to the time and space algorithm, and determines the corresponding relationship between the data to be stored and the node space according to the storage path, performs resource scheduling according to the corresponding relationship and the priority parameters corresponding to each storage path, and stores the data to be stored in the node space; determines the current virtual space statement value according to the received data development statement and the virtual mapping relationship between the virtual space and the node space; determines the decision data of the data development statement according to the current virtual space statement value, and performs data development according to the decision data. This embodiment avoids the problem of resource preemption from three aspects: time, space, and resource optimization scheduling, ensures more efficient and quality execution of data development tasks in the data middle station, and improves the scheduling capability from the resource guarantee aspect. While isolating resources, this embodiment proposes a specific solution for optimizing resource distribution, while taking into account security and performance issues. It also proposes rule ideas for scheduling execution, enriching the innovation capabilities in the construction and development of the data middle station.

[0105] Based on the first embodiment of the present application, in the second embodiment of the present application, the same or similar contents as those in the above-mentioned embodiment 1 can refer to the above introduction, and will not be repeated later. Figure 4 , the attributes in the storage path include resource pool, resource pool area, hot and cold storage classification, compression type, data directory, time period and file name; step S20, the data development method also includes steps S201 to S204:

[0106] Step S201, determining a resource pool, a resource pool area, and a first cold and hot storage classification according to a preset space algorithm.

[0107] It should be noted that the core attributes of a file's storage path are mainly composed of the following items: / resource pool / resource pool area / cold and hot storage classification / compression type / data directory / time period / file name, etc. The resource pool, resource pool area, and hot and cold storage classification are determined by spatial mathematical modeling calculations using spatial algorithms (e.g., decision tree algorithms). The data directory and file name are provided by data governance data asset management.

[0108] In a feasible implementation, step S201 may include steps S2011 to S2015:

[0109] Step S2011, collecting spatial feature attributes of the data to be stored.

[0110] It can be understood that data attribute samples are input for the spatial dimension, and the decision tree algorithm is used to model the spatial attributes. The spatial characteristic attributes of the collected data, such as the regional attributes, storage frequency, reading frequency, modification frequency, compression frequency, storage size per cycle, etc. are used as input parameters, and the resource pool, resource pool storage area, and hot and cold storage classification of the output file are output.

[0111] For example, the sample data is in the following table:

[0112]

[0113]

[0114] The positive samples are shown in the following table.

[0115]

[0116] Step S2012: determine the information entropy of each of the spatial feature attributes, and obtain the uncertainty measure corresponding to each of the spatial feature attributes.

[0117] It can be understood that the information entropy of each spatial feature attribute can be calculated according to the following formula:

[0118]

[0119] Among them, pi represents the proportion of the i-th spatial feature attribute in the data to be stored, n represents the number of spatial feature attributes, and log represents the calculated logarithm. Thus, the uncertainty measure corresponding to each spatial feature attribute can be obtained, as shown in the following table.

[0120]

[0121] Step S2013: determining the information gain value of each of the spatial feature attributes according to the information entropy.

[0122] It can be understood that the information gain value of each spatial feature attribute can be calculated by combining information entropy with the following formula:

[0123]

[0124] Among them, k represents the number of subsets of data to be stored, H(D) represents the information entropy of the data to be stored, |Di| represents the number of samples in the i-th subset, |D| represents the total number of samples of data to be stored, and H(Di) represents the information entropy of the i-th subset. The larger the information gain, the better the classification effect can be achieved by selecting the spatial feature attribute for segmentation.

[0125] Step S2014: determining the optimal spatial feature attribute and model parameters corresponding to the decision tree algorithm according to each of the spatial feature attributes and the information gain value.

[0126] It can be understood that after training with data samples, the optimal spatial data attributes and model parameters corresponding to the decision tree algorithm can be determined according to the decision tree rules in the decision tree algorithm.

[0127] Step S2015, determining a resource pool, a resource pool area, and a first hot and cold storage classification according to a decision tree algorithm, the model parameters, and the optimal spatial feature attributes.

[0128] It is understandable that the resource pool, resource pool area and first hot and cold storage classification can be determined according to the decision tree algorithm, model parameters and optimal spatial feature attributes. Figure 5 , the calculation decision tree of the resource pool area can be referred to Figure 6 , the first hot and cold storage classification calculation decision tree can refer to Figure 7 The data is classified into regional spaces, the optimal spatial segmentation attributes are defined, stored in the node space, and the result content is output as shown in the following table.

[0129]

[0130] Step S202, determining the second cold and hot storage classification, compression type and time period according to a preset time algorithm.

[0131] It can be understood that the attributes of the storage path, such as the hot and cold storage classification, compression type, and time period, can be determined by modeling and calculation using a time algorithm (eg, a recurrent neural network algorithm).

[0132] In a feasible implementation, step S202 may include steps S2021 to S2023:

[0133] Step S2021, collecting the time characteristic attributes of the data to be stored.

[0134] It can be understood that for the time dimension, data attribute samples are input, and based on the recurrent neural network algorithm, independent data segmentation is performed according to the characteristics of the data and the characteristics of the required segmented time series, and the time characteristic attributes of the collected data, such as the time period, storage size, reading frequency, compression type, output hot and cold storage classification, compression type and time period, are collected.

[0135] Step S2022, determining the hidden state and output vector of the current time step corresponding to the sequence of each data storage unit in the storage space according to the recurrent neural network algorithm, the time feature attribute, the input vector of the current time step and the hidden state of the previous time step.

[0136] It can be understood that, assuming that the sequence of each data storage unit is recorded as RNN, the recurrent neural network (RNN) is a neural network architecture that performs well in processing sequence data and can be used for time series prediction tasks. Then:

[0137] ht=σ(Wxhxt+Whhht-1+bh),

[0138] yt=σ(Whyht+by),

[0139] Among them, ht is the hidden state at time step t, xt is the input at time step t, yt is the output at time step t, Wxh, Whh and Why are weight matrices, bh and by are bias terms, and σ is an activation function (such as Hyperbolic Tangent Function (Tanh) or Rectified Linear Unit (ReLU)). At each time step, RNN accepts an input vector of the current time step (which can be an element in the sequence data) and a hidden state of the previous time step as input. Using these two inputs, RNN calculates the hidden state and output vector of the current time step. The calculation of the hidden state takes into account the hidden state of the previous time step and the input vector of the current time step, thereby capturing the time dependency in the sequence. The hidden state and output vector are then passed to the next time step as one of the inputs of the next time step, so that the storage node space of the data is distributed according to the characteristics of the time series, and the data of different time periods are output for cold storage or hot storage.

[0140] Step S2023, determining the second cold and hot storage classification, compression type and time period according to the hidden state and output vector of the current time step.

[0141] It can be understood that the sample data is shown in the following table:

[0142]

[0143] The positive samples are shown in the following table.

[0144] file name Time period Hot and cold storage dw_nb_mobile_weather_dm 20241020 hot dw_nb_mobile_weather_dm 20241019 hot dw_nb_mobile_weather_dm 20241018 hot dw_nb_mobile_weather_dm 20241017 hot ... dw_nb_mobile_weather_dm 20240831 cold dw_nb_mobile_weather_dm 20240830 cold dw_nb_mobile_weather_dm 20240829 cold ...

[0145] It is understandable that a simple loop network can be established. The structure diagram can be referred to Figure 8 Define the loss function E to represent the error between the output y^ and the true label y. Use the chain rule to find the partial derivative of E with respect to the network weight from top to bottom. Update the weight value in the opposite direction of the gradient until E converges. The update process can be referred to Fig. 9 According to the sample data, the time series training is performed, and the output results of the second hot and cold storage classification, compression type and time period are shown in the following table.

[0146]

[0147] Step S203: encrypt and fuse the first hot and cold storage classification and the second hot and cold storage classification according to preset weights to obtain hot and cold storage classifications.

[0148] It is understandable that the key attribute "hot and cold storage classification" is related to both time and space. The two models output their own calculation results, and the fusion algorithm determines the final result according to the weight. Regarding the hot and cold storage classification of the key fields output by the time model and the space model, considering that the space model lacks the time dimension, and all data in the data center are stored according to the time period, you can set the weight rules of 0.8 or 0.9 for time and 0.2 or 0.1 for space. The reason why it is not completely judged according to the time model is determined by the business characteristics. In the storage process of the data center, time faults are a common phenomenon. It is common for a data cycle to be discontinuous. If there is a time fault, the running result of the time model will cause the model output to be incomplete. When the output of the time model is missing, the result of the space model will be used to supplement the model output. For example, you can refer to Fig.10 ,Because of data faults in the time model, the output of hot and cold storage is empty, and the storage path supplements the result of the ,space model, which is recorded as cold storage.

[0149] Step S204, encrypting the resource pool, the resource pool area, the hot and cold storage classifications and the compression type according to a multi-level linear congruential algorithm to obtain multiple encrypted ciphertexts, and generating a storage path according to the multiple encrypted ciphertexts, the data directory, the time period and the file name.

[0150] It is understandable that after the calculation results are fused, the full path is output as a full path in the model. Here, it is necessary to encrypt the storage location according to the multi-level linear congruential algorithm. The storage location encryption of this proposal mainly considers the following three points: the data storage location is prevented from being displayed in plain text, and encrypted storage is further improved data security; path labels are set during the storage location encryption process, and the calculated locations are segmented and marked through encrypted storage to facilitate the subsequent scheduled analysis call execution; the scheduling call parsing performance is improved, and the encrypted array format parsing performance is higher than the analysis and judgment of the plain text path. According to the time-space algorithm, the storage path examples of the fusion calculation model are shown in the following table:

[0151] file name Time period Full path dw_nb_mobile_weather_dm 20241020 / huchi / pod10 / ssd / 20241020 / dat / dw_nb_mobile_weather_dm 20241019 / huchi / pod10 / ssd / 20241019 / dat / dw_nb_mobile_weather_dm 20241018 / huchi / pod10 / ssd / 20241018 / dat / dw_nb_mobile_weather_dm 20241017 / huchi / pod10 / ssd / 20241017 / dat / ... dw_nb_mobile_weather_dm 20240831 / huchi / pod10 / ssd / 20240831 / gz / dw_nb_mobile_weather_dm 20240830 / huchi / pod10 / ssd / 20240830 / gz / dw_nb_mobile_weather_dm 20240829 / huchi / pod10 / ssd / 20240829 / gz /

[0152] It is worth noting that the core storage path is: / resource pool / resource pool area / cold and hot storage classification / compression type / data directory / time period / file name. The full path of one period is as follows:

[0153] / huchi / pod10 / ssd / dat / dw / newbusi / 20241020 / dw_nb_mobile_weather_dm.dat

[0154] Resource pool: huchi, resource pool region: pod10, hot and cold storage classification: ssd, compression type: dat, data directory: 'dw / newbusi', time period: 20241020, file name: dw_nb_mobile_weather_dm.dat. The attributes that need to be encrypted are: resource pool, resource pool region, hot and cold storage classification, compression type; the attributes that do not need to be encrypted are: data directory, time period, file name. Because the character length of the data directory and file name is relatively long and has no meaning for global resource scheduling and is a static constant, the time period is displayed as a number, and is also a high-performance computing constant in global resource scheduling and has a low security attribute.

[0155] It can be understood that after encrypting the attributes that need to be encrypted one by one, the encrypted resource pool, encrypted resource pool area, encrypted hot and cold storage classifications and encrypted compression type are obtained. Then, they are concatenated with the attribute that does not need to be encrypted, " / data directory / time period / file name / ", to obtain the encrypted storage path.

[0156] In a feasible implementation, step S204 may include steps S2041 to S2042:

[0157] Step S2041, converting the resource pool, the resource pool area, the hot and cold storage classifications and the compression type into digital sequences respectively.

[0158] It can be understood that a simple conversion method is first assumed: convert az to 0-25 respectively. Therefore, taking the first path node to be encrypted, resource pool "hachi" as an example, "hachi" is converted into a digital sequence of [7, 0, 2, 7, 8] (h = 7, a = 0, c = 2, i = 8).

[0159] Step S2042, determining the encrypted resource pool, the encrypted resource pool area, the encrypted hot and cold storage classifications, and the encrypted compression type according to the digital sequence, the preset constant, and the encrypted character set size.

[0160] It can be understood that only the key position elements are stored and encrypted, and the encryption process can be encrypted according to the linear congruential equation, which is as follows:

[0161] C = (a*m+b) mod n,

[0162] Where C is the encrypted ciphertext, m is the plaintext to be encrypted (usually converted to the corresponding number or numerical representation), a, b and n are known constants (a and n are usually relatively prime), and n is usually the size of the character set used by the encryption system (for example, for the English character set, n = 26).

[0163] For example, taking the first path node that needs to be encrypted, the resource pool "hachi", as an example, each character needs to be converted into a number and encrypted using a linear congruential equation. Next, select a set of encryption parameters a, b, n. Assume a = 5, b = 3, n = 26 (the size of the English character set). Then apply a linear congruential equation to encrypt each number:

[0164] C1=(5*7+3)mod 26=(35+3)mod 26=38mod 26=12,

[0165] C2=(5*0+3)mod 26=(0+3)mod 26=3,

[0166] C3=(5*2+3)mod 26=(10+3)mod 26=13,

[0167] C4 = (5*7+3) mod 26 = 38 mod 26 = 12 (same as C1, this is a property of linear congruential equations),

[0168] C5=(5*8+3)mod 26=(40+3)mod 26=43mod 26=17,

[0169] Therefore, the ciphertext of ‘hachi’ after encryption is [12, 3, 13, 12, 17].

[0170] Similarly, according to the above algorithm, the following path elements are encrypted and the corresponding ciphertext is obtained:

[0171] Resource pool area: pod10, ciphertext: [63, 106, 251, 12, 7],

[0172] Hot and cold storage classification: ssd, ciphertext: [96, 96, 51],

[0173] Compression type: dat, ciphertext: [10, 73, 11].

[0174] Therefore, the storage path can be segmented and stored as follows:

[0175] Encryption part: / resource pool / resource pool area / hot and cold storage classification / compression type / ;

[0176] Encrypted into four arrays for storage:

[0177] [12, 3, 13, 12, 17][63, 106, 251, 12, 7][96, 96, 51][10, 73, 11],

[0178] Plain text part: / data directory / time period / file name / : / dw / newbusi / 20241020 / dw_nb_mobile_weather_dm.dat.

[0179] It is understandable that, combined with the basic model after time training, the encrypted storage data is parsed and the labeling strategy is scheduled, making full use of the respective advantages of hot and cold storage for data processing. This innovation is of great significance to the saving of resource expenditure. According to the current model statistics, 80% of the storage in the data center can be stored on cold storage, such as historical data, monthly, quarterly, and annual data, which have low processing timeliness requirements. At the same time, the migration and tracing of historical data are not very time-sensitive operations. The timeliness requirements of some small daily and weekly data volumes can also be met on cold storage, and the proportion of data with high timeliness in the entire data center is relatively less than 10%. If implemented according to this model strategy, 85% of the storage investment in the data warehouse can be gradually replaced by cold storage, and 15% of the storage can be replaced by hot storage with faster access. Not only will the hardware level be optimized in terms of data warehouse speed-up, but resource investment will also be greatly reduced.

[0180] In this embodiment, the resource pool, resource pool area and first cold and hot storage classification are determined according to a preset space algorithm; the second cold and hot storage classification, compression type and time period are determined according to a preset time algorithm; the first cold and hot storage classification and the second cold and hot storage classification are encrypted and merged according to preset weights, and then the resource pool, resource pool area, cold and hot storage classification and compression type are encrypted according to a multi-level linear congruential algorithm, so that the storage path can be generated more accurately according to multiple encrypted ciphertexts, data directories, time periods and file names.

[0181] It should be noted that the above examples are only used to understand the present application and do not constitute a limitation on the data development method of the present application. More simple transformations based on this technical concept are all within the scope of protection of the present application.

[0182] This application also provides a data development device, please refer to Fig.11 , the data development device comprises:

[0183] A storage space segmentation module 10, for segmenting the storage space into a plurality of node spaces according to a hot-cold storage algorithm, wherein the node space includes a virtual space;

[0184] A storage path determination module 20, configured to determine a storage path of the data to be stored according to a preset time algorithm and a preset space algorithm, and determine a correspondence between the data to be stored and the node space according to the storage path, wherein the attributes in the storage path include hot and cold storage classifications;

[0185] A resource scheduling module 30, configured to perform resource scheduling according to the corresponding relationship and the priority parameter corresponding to each storage path, and store the data to be stored in the node space;

[0186] A statement value determination module 40, for determining a current virtual space statement value according to the received data development statement and the virtual mapping relationship between the virtual space and the node space;

[0187] The data development module 50 is used to determine the decision data of the data development statement according to the current virtual space statement value, and perform data development according to the decision data.

[0188] The data development device provided by the present application adopts the data development method in the above embodiment to solve the technical problem. Compared with the prior art, the beneficial effects of the data development device provided by the present application are the same as the beneficial effects of the data development method provided by the above embodiment, and the other technical features in the data development device are the same as the features disclosed in the above embodiment method, which will not be repeated here.

[0189] The present application provides a data development device, which includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the data development method in the above-mentioned embodiment one.

[0190] Reference below Fig.12 , which shows a schematic diagram of the structure of a data development device suitable for implementing an embodiment of the present application. The data development device in the embodiment of the present application may include, but is not limited to, mobile terminals such as mobile phones, laptop computers, digital broadcast receivers, PDAs (Personal Digital Assistants), PADs (Portable Application Descriptions), PMPs (Portable Media Players), vehicle-mounted terminals (such as vehicle-mounted navigation terminals), etc., and fixed terminals such as digital TVs, desktop computers, etc. Fig.12 The data development device shown is merely an example and should not bring any limitation to the functions and scope of use of the embodiments of the present application.

[0191] like Fig.12 As shown, the data development device may include a processing device 1001 (e.g., a central processing unit, a graphics processor, etc.), which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM: Read Only Memory) 1002 or a program loaded from a storage device 1003 to a random access memory (RAM: Random Access Memory) 1004. In RAM1004, various programs and data required for the operation of the data development device are also stored. The processing device 1001, ROM1002, and RAM1004 are connected to each other through a bus 1005. An input / output (I / O) interface 1006 is also connected to the bus. Generally, the following systems can be connected to the I / O interface 1006: an input device 1007 including, for example, a touch screen, a touch pad, a keyboard, a mouse, an image sensor, a microphone, an accelerometer, a gyroscope, etc.; an output device 1008 including, for example, a liquid crystal display (LCD: Liquid Crystal Display), a speaker, a vibrator, etc.; a storage device 1003 including, for example, a magnetic tape, a hard disk, etc.; and a communication device 1009. The communication device 1009 can allow the data development device to communicate with other devices wirelessly or wired to exchange data. Although the data development device with various systems is shown in the figure, it should be understood that it is not required to implement or have all the systems shown. More or fewer systems can be implemented or have alternatively.

[0192] In particular, according to the embodiments disclosed in the present application, the process described above with reference to the flowchart can be implemented as a computer software program. For example, the embodiments disclosed in the present application include a computer program product, which includes a computer program carried on a computer-readable medium, and the computer program includes a program code for executing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from a network through a communication device, or installed from a storage device 1003, or installed from a ROM 1002. When the computer program is executed by the processing device 1001, the above-mentioned functions defined in the method of the embodiment disclosed in the present application are executed.

[0193] The data development device provided by the present application adopts the data development method in the above embodiment to solve the technical problems of data development. Compared with the prior art, the beneficial effects of the data development device provided by the present application are the same as the beneficial effects of the data development method provided by the above embodiment, and the other technical features in the data development device are the same as the features disclosed in the method of the previous embodiment, which will not be repeated here.

[0194] It should be understood that the various parts disclosed in this application can be implemented by hardware, software, firmware or a combination thereof. In the description of the above embodiments, specific features, structures, materials or characteristics can be combined in any one or more embodiments or examples in a suitable manner.

[0195] The above is only a specific implementation of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art who is familiar with the present technical field can easily think of changes or substitutions within the technical scope disclosed in the present application, which should be included in the protection scope of the present application. Therefore, the protection scope of the present application should be based on the protection scope of the claims.

[0196] The present application provides a computer-readable storage medium having computer-readable program instructions (ie, computer programs) stored thereon, and the computer-readable program instructions are used to execute the data development method in the above-mentioned embodiment.

[0197] The computer-readable storage medium provided in the present application may be, for example, a USB flash drive, but is not limited to electrical, magnetic, optical, electromagnetic, infrared, or semiconductor systems, systems or devices, or any combination of the above. More specific examples of computer-readable storage media may include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In this embodiment, the computer-readable storage medium may be any tangible medium containing or storing a program that can be used by or in combination with an instruction execution system, system or device. The program code contained on the computer-readable storage medium may be transmitted using any appropriate medium, including but not limited to: wires, optical cables, RF (Radio Frequency), etc., or any suitable combination of the above.

[0198] The computer-readable storage medium may be included in the data development device; or may exist independently without being assembled into the data development device.

[0199] Computer program code for performing the operations of the present application may be written in one or more programming languages ​​or a combination thereof, including object-oriented programming languages ​​such as Java, Smalltalk, C++, and conventional procedural programming languages ​​such as "C" or similar programming languages. The program code may be executed entirely on the user's computer, partially on the user's computer, as a separate software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., via the Internet using an Internet service provider).

[0200] The flow chart and block diagram in the accompanying drawings illustrate the possible architecture, function and operation of the system, method and computer program product according to various embodiments of the present application. In this regard, each square box in the flow chart or block diagram can represent a module, a program segment or a part of a code, and the module, the program segment or a part of the code contains one or more executable instructions for realizing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the square box can also occur in a sequence different from that marked in the accompanying drawings. For example, two square boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each square box in the block diagram and / or flow chart, and the combination of the square boxes in the block diagram and / or flow chart can be implemented with a dedicated hardware-based system that performs a specified function or operation, or can be implemented with a combination of dedicated hardware and computer instructions.

[0201] The modules involved in the embodiments described in this application may be implemented by software or hardware, wherein the name of the module does not constitute a limitation on the unit itself in some cases.

[0202] The readable storage medium provided in this application is a computer-readable storage medium, which stores computer-readable program instructions (i.e., computer programs) for executing the above-mentioned data development method, and can solve technical problems. Compared with the prior art, the beneficial effects of the computer-readable storage medium provided in this application are the same as the beneficial effects of the data development method provided in the above-mentioned embodiment, and will not be repeated here.

[0203] The present application also provides a computer program product, including a computer program, which implements the steps of the above-mentioned data development method when executed by a processor.

[0204] The computer program product provided by the present application can solve the technical problem. Compared with the prior art, the beneficial effects of the computer program product provided by the present application are the same as the beneficial effects of the data development method provided by the above embodiment, which will not be described in detail here.

[0205] The above descriptions are only some embodiments of the present application, and are not intended to limit the patent scope of the present application. All equivalent structural changes made using the contents of the present application specification and drawings under the technical concept of the present application, or direct / indirect applications in other related technical fields are included in the patent protection scope of the present application.

Claims

1. A data development method, characterized in that: The method includes: Dividing the storage space into a plurality of node spaces according to a hot-cold storage algorithm, wherein the node space includes a virtual space; Determine a storage path for the data to be stored according to a preset time algorithm and a preset space algorithm, and determine a corresponding relationship between the data to be stored and the node space according to the storage path, wherein the attributes in the storage path include hot and cold storage classification; Perform resource scheduling according to the corresponding relationship and the priority parameter corresponding to each storage path, and store the data to be stored in the node space; Determine the current virtual space statement value according to the received data development statement and the virtual mapping relationship between the virtual space and the node space; The decision data of the data development statement is determined according to the current virtual space statement value, and data development is performed according to the decision data.

2. The method according to claim 1, characterized in that The attributes in the storage path include resource pool, resource pool area, hot and cold storage classification, compression type, data directory, time period and file name; The step of determining the storage path of the data to be stored according to the preset time algorithm and the preset space algorithm includes: Determine the resource pool, resource pool area, and first hot and cold storage classification according to a preset space algorithm; Determine the second cold and hot storage classification, compression type and time period according to a preset time algorithm; Encrypting and fusing the first cold and hot storage classification and the second cold and hot storage classification according to preset weights to obtain a cold and hot storage classification; The resource pool, the resource pool area, the hot and cold storage classifications and the compression type are encrypted according to a multi-level linear congruential algorithm to obtain a plurality of encrypted ciphertexts, and a storage path is generated according to the plurality of encrypted ciphertexts, the data directory, the time period and the file name.

3. The method according to claim 2, characterized in that The step of determining the resource pool, the resource pool area, and the first hot and cold storage classification according to the preset space algorithm includes: Collecting spatial characteristic attributes of the data to be stored; Determine the information entropy of each of the spatial feature attributes to obtain an uncertainty measure corresponding to each of the spatial feature attributes; Determine the information gain value of each of the spatial feature attributes according to the information entropy; Determine the optimal spatial feature attribute and model parameters corresponding to the decision tree algorithm according to each of the spatial feature attributes and the information gain value; The resource pool, the resource pool area and the first hot and cold storage classification are determined according to the decision tree algorithm, the model parameters and the optimal spatial feature attributes.

4. The method according to claim 2, characterized in that The step of determining the second cold and hot storage classification, compression type and time period according to a preset time algorithm comprises: Collecting time characteristic attributes of the data to be stored; Determine the hidden state and output vector of the current time step corresponding to the sequence of each data storage unit in the storage space according to the recurrent neural network algorithm, the time feature attribute, the input vector of the current time step and the hidden state of the previous time step; The second hot and cold storage classification, compression type and time period are determined according to the hidden state and output vector of the current time step.

5. The method according to claim 2, characterized in that The step of encrypting the resource pool, the resource pool area, the hot and cold storage classifications, and the compression type according to a multi-level linear congruential algorithm to obtain a plurality of encrypted ciphertexts includes: Respectively converting the resource pool, the resource pool area, the hot and cold storage classifications, and the compression type into digital sequences; The encrypted resource pool, the encrypted resource pool area, the encrypted hot and cold storage classifications and the encrypted compression type are determined according to the digital sequence, the preset constant and the encrypted character set size.

6. The method according to claim 1, characterized in that After determining the decision data of the data development statement according to the current virtual space statement value and performing the data development step according to the decision data, the method further includes: Adding resource extraction marking points to the decision data; The node resources having the resource extraction mark point in the storage space are removed.

7. A data development device, characterized in that: The data development device comprises: A storage space segmentation module, used to segment the storage space into a plurality of node spaces according to a hot-cold storage algorithm, wherein the node space includes a virtual space; A storage path determination module, used to determine a storage path of the data to be stored according to a preset time algorithm and a preset space algorithm, and determine a corresponding relationship between the data to be stored and the node space according to the storage path, wherein the attributes in the storage path include hot and cold storage classification; A resource scheduling module, used to perform resource scheduling according to the corresponding relationship and the priority parameter corresponding to each storage path, and store the data to be stored in the node space; A statement value determination module, used to determine the current virtual space statement value according to the received data development statement and the virtual mapping relationship between the virtual space and the node space; The data development module is used to determine the decision data of the data development statement according to the current virtual space statement value, and perform data development according to the decision data.

8. A data development device, characterized in that: The device comprises: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the computer program is configured to implement the steps of the data development method according to any one of claims 1 to 6.

9. A storage medium, characterized in that: The storage medium is a computer-readable storage medium, and a computer program is stored on the storage medium. When the computer program is executed by a processor, the steps of the data development method according to any one of claims 1 to 6 are implemented.

10. A computer program product, characterized in that The computer program product comprises a computer program, and when the computer program is executed by a processor, the steps of the data development method according to any one of claims 1 to 6 are implemented.

Citation Information

Patent Citations

  • Data management method, device and equipment and storage medium

    CN112445952A

  • Data storage space dynamic adjustment method, storage subsystem and intelligent computing platform

    CN118069073A

  • Methods and systems for detection in an industrial internet of things data collection environment with large data sets

    US20180284737A1

  • Method, apparatus, and system for creating training task on ai training platform, and medium

    US20240061712A1

  • Photo storage method, storage medium, server, and apparatus

    WO2019218459A1