Distributed big data storage system and method

By determining the feature distance and spatiotemporal polygon retrieval granularity of distributed storage nodes, the data storage location and search process are optimized, the problem of high computing costs in distributed storage systems is solved, and efficient spatiotemporal polygon retrieval is achieved.

CN120256494AActive Publication Date: 2025-07-04GUIZHOU BUSINESS SCHOOL
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202510691586.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-27
Publication Date
2025-07-04
Estimated Expiration
2045-05-27

AI Technical Summary

Technical Problem

In distributed storage systems, the high computing cost of spatiotemporal polygon retrieval algorithms has become a bottleneck that restricts system efficiency and expansion capabilities. Especially in massive data retrieval tasks, it is difficult for the existing technology to effectively improve retrieval efficiency.

Method used

By obtaining the feature distance of distributed storage nodes and the granularity of spatiotemporal polygon retrieval, the storage location of economically collected data is determined, and its spatiotemporal storage dimension features are used as search information, a two-stage search process is constructed to reduce the data traversal range.

Benefits of technology

It improves the spatio-temporal polygon retrieval efficiency of distributed storage nodes, reduces the calculation overhead, and improves the search speed.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120256494A_ABST
    Figure CN120256494A_ABST
Patent Text Reader

Abstract

The invention provides a big data storage system and method based on distribution, and the method comprises the steps: determining a space-time polygon retrieval granularity corresponding to a distributed storage node based on historical economic data in the distributed storage node; performing polygon retrieval judgment on the historical economic data in the distributed storage node according to the space-time polygon retrieval granularity to obtain a space-time association historical data set, and obtaining space-time storage dimension features corresponding to the historical economic data in the space-time association historical data set; the storage position of the economic collection data is determined based on the space-time storage dimension features corresponding to the historical economic data, and the space-time storage dimension features of the economic collection data serve as space-time retrieval information for distributed retrieval, so that storage and retrieval can be performed based on the space-time storage dimension features of the data to be stored; and the retrieval efficiency of performing space-time polygon retrieval on the distributed storage nodes is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of data storage and data processing, and more specifically, to a distributed big data storage system and method. Background Art

[0002] In a distributed storage system, a large amount of information is usually split and stored separately in multiple independent databases, such as order transaction data, user payment data, etc. Such a distributed information storage strategy based on cloud computing can not only efficiently handle the data distribution and computing resource scheduling problems generated during the merging process of a large amount of information files, but also utilize the good scalability advantage of the distributed file system to expand the coverage of data storage and computing capabilities.

[0003] In the prior art, the invocation of distributed storage information usually depends on the spatio-temporal correlation of data, and a retrieval algorithm based on a spatio-temporal polygon is widely used for data query. Since a polygon area usually has characteristics such as a wide boundary and an irregular shape, it often needs to be accurately expressed in the form of a vector containing a large number of vertices. In actual queries, in order to complete a polygon range query, the system often needs to traverse hundreds of millions or even billions of spatio-temporal data points, thus bringing huge computational overhead. Especially when facing the retrieval task of a large amount of distributed storage data, this high computational cost has become a key bottleneck restricting the application efficiency and expansion ability of the distributed storage system. Therefore, how to effectively improve the retrieval efficiency of distributed storage information has become a technical problem that needs to be solved urgently. Summary of the Invention

[0004] This application provides a distributed big data storage system and method, which can perform storage and retrieval based on the spatio-temporal storage dimension characteristics of the data to be stored, and improves the retrieval efficiency of spatio-temporal polygon retrieval for distributed storage nodes.

[0005] In a first aspect, this application provides a distributed big data storage method, which can be executed by a network device, or can also be executed by a chip configured in the network device. This application does not make any limitations in this regard.

[0006] Specifically, the method includes:

[0007] Obtain multiple distributed storage nodes, and determine the characteristic distances between the economic acquisition data to be stored and each distributed storage node according to the characteristic tags corresponding to the economic acquisition data to be stored;

[0008] Allocate the economic acquisition data to the distributed storage node with the smallest characteristic distance, and determine the spatio-temporal polygon retrieval granularity corresponding to the distributed storage node based on the historical economic data in the distributed storage node;

[0009] Perform polygon retrieval determination on the historical economic data in the distributed storage node according to the spatio-temporal polygon retrieval granularity, obtain a spatio-temporal associated historical data set, and acquire the spatio-temporal storage dimension features corresponding to each historical economic data in the spatio-temporal associated historical data set;

[0010] Determine the storage location of the economic acquisition data based on the spatio-temporal storage dimension features corresponding to each historical economic data, and use the spatio-temporal storage dimension features of the economic acquisition data as spatio-temporal retrieval information for distributed retrieval.

[0011] Combined with the first aspect, in some implementation manners of the first aspect, in the process of determining the characteristic distances between the economic acquisition data and the respective distributed storage nodes, convert the characteristic values on each spatio-temporal feature dimension in the characteristic label corresponding to the economic acquisition data to be stored into a distance detection vector, and acquire the characteristic vectors corresponding to the storage center data of each distributed storage node, and use the Euclidean distance values between the distance detection vector and each characteristic vector as the characteristic distances between the economic acquisition data and the respective distributed storage nodes.

[0012] Combined with the first aspect, in some implementation manners of the first aspect, determining the spatio-temporal polygon retrieval granularity corresponding to the distributed storage node based on the historical economic data in the distributed storage node specifically includes:

[0013] Acquire the historical economic data in the distributed storage node and construct a historical economic data space;

[0014] In the spatio-temporal feature dimension space, obtain the historical economic data space, acquire the difference between the maximum characteristic value and the minimum characteristic value corresponding to each spatio-temporal feature dimension in the historical economic data space as the data range corresponding to each spatio-temporal feature dimension, and use the ratio between the data range corresponding to each spatio-temporal feature dimension and the corresponding standard data range as the data density feature corresponding to each spatio-temporal feature dimension;

[0015] Perform weighted fusion on the data density features in each spatio-temporal feature dimension based on the dimension weights corresponding to each spatio-temporal feature dimension to obtain the spatio-temporal polygon retrieval granularity corresponding to the distributed storage node.

[0016] In combination with the first aspect, in some implementations of the first aspect, performing polygon retrieval determination on the historical economic data in the distributed storage node according to the spatio-temporal polygon retrieval granularity to obtain the spatio-temporal associated historical data set specifically includes: performing mapping according to the spatio-temporal polygon retrieval granularity to obtain the spatio-temporal retrieval distances corresponding to each spatio-temporal feature dimension. For any one spatio-temporal feature dimension, taking the farthest historical economic data within the corresponding spatio-temporal retrieval distance from the economic collection data as the polygon retrieval vertex data, and forming a polygon retrieval boundary based on the polygon retrieval vertex data corresponding to each spatio-temporal feature dimension respectively, and extracting all the historical economic data within the boundary of the polygon retrieval boundary to form the spatio-temporal associated historical data set.

[0017] In combination with the first aspect, in some implementations of the first aspect, in the process of obtaining the spatio-temporal storage dimension features corresponding to each historical economic data in the spatio-temporal associated historical data set, taking the spatio-temporal center point distance corresponding to each historical economic data in the spatio-temporal associated historical data set as the spatio-temporal storage dimension feature, and the spatio-temporal center point distance is based on a preset spatio-temporal data center point, and using the characteristic singular value of the collection label between the historical economic data and the spatio-temporal data center point as the spatio-temporal storage dimension feature, where the characteristic singular value is the weighted result of the characteristic difference values of the historical economic data and the spatio-temporal data center point in each spatio-temporal feature dimension.

[0018] In combination with the first aspect, in some implementations of the first aspect, in the process of taking the spatio-temporal storage dimension feature of this economic collection data as the spatio-temporal retrieval information for distributed retrieval, taking the spatio-temporal storage dimension feature of this economic collection data as the eigenvalue of the incremental retrieval dimension, and performing spatio-temporal polygon retrieval according to multiple polygon boundary retrieval conditions and incremental retrieval dimension conditions.

[0019] In combination with the first aspect, in some implementations of the first aspect, the historical economic data in the spatio-temporal associated historical data set are stored sequentially according to the size of the spatio-temporal storage dimension feature.

[0020] In a second aspect, the present application provides a distributed big data storage system, which includes a data storage unit, and the data storage unit includes:

[0021] A storage node selection module, configured to obtain a plurality of distributed storage nodes, and determine the characteristic distances between the economic collection data to be stored and each distributed storage node respectively according to the characteristic labels corresponding to the economic collection data;

[0022] A data processing module, configured to allocate the economic acquisition data to the distributed storage node with the smallest feature distance, and determine the spatio-temporal polygon retrieval granularity corresponding to the distributed storage node based on the historical economic data in the distributed storage node;

[0023] The data processing module is further configured to perform polygon retrieval determination on the historical economic data in the distributed storage node according to the spatio-temporal polygon retrieval granularity, obtain a spatio-temporal associated historical data set, and acquire the spatio-temporal storage dimension features corresponding to each historical economic data in the spatio-temporal associated historical data set;

[0024] A data storage control module determines the storage location of the economic acquisition data based on the spatio-temporal storage dimension features corresponding to each historical economic data, and uses the spatio-temporal storage dimension features of the economic acquisition data as spatio-temporal retrieval information for distributed retrieval.

[0025] In a third aspect, the present application provides a computer terminal device, which includes a memory and a processor. The memory stores code, and the processor is configured to obtain the code and execute the above-mentioned distributed-based big data storage method.

[0026] In a fourth aspect, the present application provides a computer-readable storage medium, which stores at least one computer program, and the computer program is loaded and executed by a processor to implement the operations performed by the above-mentioned distributed-based big data storage method.

[0027] The technical solutions provided by the disclosed embodiments of the present application have the following beneficial effects:

[0028] In a distributed-based big data storage system and method provided by the present application, first, a plurality of distributed storage nodes are obtained, and according to the feature tags corresponding to the economic acquisition data to be stored, the feature distances between the economic acquisition data and each distributed storage node are determined; the economic acquisition data is allocated to the distributed storage node with the smallest feature distance, and based on the historical economic data in the distributed storage node, the spatio-temporal polygon retrieval granularity corresponding to the distributed storage node is determined; polygon retrieval determination is performed on the historical economic data in the distributed storage node according to the spatio-temporal polygon retrieval granularity, a spatio-temporal associated historical data set is obtained, and the spatio-temporal storage dimension features corresponding to each historical economic data in the spatio-temporal associated historical data set are acquired; the storage location of the economic acquisition data is determined based on the spatio-temporal storage dimension features corresponding to each historical economic data, and the spatio-temporal storage dimension features of the economic acquisition data are used as spatio-temporal retrieval information for distributed retrieval.

[0029] Therefore, it can be seen that the present application first determines the corresponding spatio-temporal polygon retrieval granularity according to the spatio-temporal characteristics of the data stored in the distributed storage nodes, thereby effectively adjusting the number of samples in the spatio-temporal associated historical data set. Based on the spatio-temporal storage dimension characteristics corresponding to each adjacent historical data sample in the spatio-temporal associated historical data set, the storage location of the data to be stored is determined, and the spatio-temporal storage dimension characteristics of this economic acquisition data are used as the spatio-temporal retrieval information for distributed retrieval, constructing the spatio-temporal storage dimension characteristics of the distributed storage data. The original polygon retrieval process that relies on high-overhead coordinate calculation is converted into a two-stage process of range judgment and spatial verification of the spatio-temporal storage dimension characteristics, thereby reducing the data traversal range during spatio-temporal polygon retrieval and improving the retrieval speed of the distributed storage nodes.

[0030] In summary, the present application can perform storage and retrieval based on the spatio-temporal storage dimension characteristics of the data to be stored, improving the retrieval efficiency of spatio-temporal polygon retrieval for distributed storage nodes. BRIEF DESCRIPTION OF THE DRAWINGS

[0031] Figure 1 is an exemplary flowchart of a distributed big data storage method according to some embodiments of the present application;

[0032] Figure 2 is a schematic structural diagram of a data storage unit according to some embodiments of the present application;

[0033] Figure 3 is a schematic structural diagram of a computer terminal device for implementing a distributed big data storage method according to some embodiments of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0034] The present application obtains multiple distributed storage nodes, determines the characteristic distances between the economic acquisition data and each distributed storage node respectively according to the characteristic tags corresponding to the economic acquisition data to be stored; distributes the economic acquisition data to the distributed storage node with the smallest characteristic distance, and determines the spatio-temporal polygon retrieval granularity corresponding to the distributed storage node based on the historical economic data in the distributed storage node; performs polygon retrieval determination on the historical economic data in the distributed storage node according to the spatio-temporal polygon retrieval granularity to obtain a spatio-temporal associated historical data set, and obtains the spatio-temporal storage dimension characteristics corresponding to each historical economic data in the spatio-temporal associated historical data set; determines the storage location of this economic acquisition data based on the spatio-temporal storage dimension characteristics corresponding to each historical economic data respectively, and uses the spatio-temporal storage dimension characteristics of this economic acquisition data as the spatio-temporal retrieval information for distributed retrieval, and can perform storage and retrieval based on the spatio-temporal storage dimension characteristics of the data to be stored, improving the retrieval efficiency of spatio-temporal polygon retrieval for distributed storage nodes.

[0035] In order to better understand the above technical solution, the above technical solution will be described in detail below in conjunction with the accompanying drawings and specific implementation methods. Figure 1 , which is an exemplary flow chart of a distributed big data storage method according to some embodiments of the present application. The distributed big data storage method 100 mainly includes the following steps:

[0036] In step S101, a plurality of distributed storage nodes are obtained, and characteristic distances corresponding to the economic collection data to be stored and the respective distributed storage nodes are determined according to characteristic labels corresponding to the economic collection data to be stored.

[0037] Optionally, in some embodiments, the feature label includes feature values ​​on multiple spatiotemporal feature dimensions, wherein the spatiotemporal feature dimensions include: spatial dimension, such as the spatial longitude and latitude of the collection location associated with the economic collection data, H3 / S2 code, administrative area code, etc. as the feature value of the spatial dimension; time dimension: such as the collection timestamp as the feature value of the time dimension; in some other embodiments, other dimensional features that can reflect the economic collection data information, such as the label value corresponding to the business category to which the data belongs and the type of economic activity, can also be used as the spatiotemporal feature dimension, and this application does not limit this.

[0038] Optionally, in some embodiments, in the process of determining the characteristic distances corresponding to the economic collection data and each distributed storage node, the characteristic values ​​on each spatiotemporal characteristic dimension in the characteristic labels corresponding to the economic collection data to be stored can be converted into distance detection vectors, and the characteristic vectors corresponding to the storage center data of each distributed storage node can be obtained, and the Euclidean distance values ​​between the distance detection vectors and each characteristic vector can be used as the characteristic distances corresponding to the economic collection data and each distributed storage node.

[0039] In a specific implementation, the mean data of all the historical economic data in the distributed storage node can be used as the storage center data, and the feature mean values ​​of the historical economic data in each feature dimension can be used as the vector values ​​of the feature vector.

[0040] In step S102, the economic collection data is distributed to the distributed storage node with the smallest feature distance, and based on the historical economic data in the distributed storage node, the spatiotemporal polygon retrieval granularity corresponding to the distributed storage node is determined.

[0041] It should be noted that the spatio-temporal polygon retrieval granularity is used to indicate the distribution density of the spatio-temporal dimensional characteristics of the data stored in the distributed storage node. When the spatio-temporal polygon retrieval granularity is large, it indicates that the spatio-temporal dimensional characteristics of the data stored in the distributed storage node are distributed more densely. The polygon boundary of the spatio-temporal dimensional characteristics of the data is usually more complex and sensitive in the high-density area. Therefore, it is necessary to estimate the position of the data to be stored more precisely. If only a small amount of adjacent historical data is used to estimate the feature distribution or determine the storage position, it is easy to cause errors due to local fluctuations. At this time, it is necessary to increase the sample size of the data adjacent to the spatio-temporal dimensional characteristics of the data to be stored, so as to smooth the local fluctuations, form a more stable feature distribution trend, reduce the retrieval difficulty of the spatio-temporal polygon retrieval, and improve the retrieval efficiency. Optionally, in some embodiments, based on the historical economic data in the distributed storage node, determining the spatio-temporal polygon retrieval granularity corresponding to the distributed storage node specifically includes:

[0042] Obtain the historical economic data in the distributed storage node and construct a historical economic data space;

[0043] In the spatio-temporal feature dimension space, obtain the historical economic data space, and obtain the difference between the maximum eigenvalue and the minimum eigenvalue corresponding to the historical economic data space in each spatio-temporal feature dimension as the data range corresponding to each spatio-temporal feature dimension, and use the ratio between the data range corresponding to each spatio-temporal feature dimension and the corresponding standard data range as the data density feature corresponding to each spatio-temporal feature dimension;

[0044] Based on the dimension weights corresponding to each spatio-temporal feature dimension, perform weighted fusion on the data density features in each spatio-temporal feature dimension to obtain the spatio-temporal polygon retrieval granularity corresponding to the distributed storage node.

[0045] Optionally, in some embodiments, in the process of performing weighted fusion on the data density features in each spatio-temporal feature dimension based on the dimension weights corresponding to each spatio-temporal feature dimension to obtain the spatio-temporal polygon retrieval granularity corresponding to the distributed storage node, the weighted weights corresponding to each spatio-temporal feature dimension are calibrated as constants according to experience.

[0046] In step S103, perform polygon retrieval determination on the historical economic data in the distributed storage node according to the spatio-temporal polygon retrieval granularity to obtain a spatio-temporal associated historical data set, and obtain the spatio-temporal storage dimension features corresponding to each historical economic data in the spatio-temporal associated historical data set.

[0047] Optionally, in some embodiments, polygon retrieval determination is performed on the historical economic data in the distributed storage node according to the spatio-temporal polygon retrieval granularity, and obtaining the spatio-temporal associated historical data set specifically includes: performing mapping according to the spatio-temporal polygon retrieval granularity to obtain the spatio-temporal retrieval distances corresponding to each spatio-temporal feature dimension. For any one spatio-temporal feature dimension, the farthest historical economic data within the corresponding spatio-temporal retrieval distance from the economic acquisition data is used as the polygon retrieval vertex data, and a polygon retrieval boundary is formed based on the polygon retrieval vertex data respectively corresponding to each spatio-temporal feature dimension. All historical economic data within the boundary of the polygon retrieval boundary is extracted to form the spatio-temporal associated historical data set.

[0048] In specific implementation, in the process of performing mapping according to the spatio-temporal polygon retrieval granularity to obtain the spatio-temporal retrieval distances corresponding to each spatio-temporal feature dimension, different spatio-temporal retrieval distance levels can be mapped according to a preset linear mapping table based on the granularity range where the spatio-temporal polygon retrieval granularity is located.

[0049] Preferably, in some embodiments, after obtaining the polygon retrieval vertex data, a convex hull algorithm can be used to generate a closed polygon area as the retrieval boundary.

[0050] Optionally, in some embodiments, in the process of obtaining the spatio-temporal storage dimension features corresponding to each historical economic data in the spatio-temporal associated historical data set, the spatio-temporal center point distance corresponding to each historical economic data in the spatio-temporal associated historical data set is used as the spatio-temporal storage dimension feature. It should be noted that in this application, a spatio-temporal center point is set for each node in the distributed system, and the difference values (such as time, geography, event type) between each historical economic data and the center point on multiple acquisition tag dimensions are weighted and integrated to obtain a one-dimensional spatio-temporal storage dimension feature value. All historical data in each node is sorted and stored in ascending order according to this dimension feature, and the complex polygon is converted into a "maximum vector radius boundary" centered on the center point for approximate filtering, so that the first retrieval condition of the polygon boundary can be set as the maximum feature difference of the center point in each spatio-temporal feature dimension, thereby reducing the data comparison range of the spatio-temporal polygon retrieval and improving the retrieval speed of the spatio-temporal polygon retrieval.

[0051] In specific implementation, the distance of the spatio-temporal center point needs to be based on the preset spatio-temporal data center point, and the characteristic singular value of the acquisition tag between the historical economic data and the spatio-temporal data center point is used as the spatio-temporal storage dimension feature. Among them, the characteristic singular value is the weighted result of the characteristic difference value of the historical economic data and the spatio-temporal data center point in each spatio-temporal feature dimension. Among them, the weighted weights corresponding to each spatio-temporal feature dimension are calibrated as constants according to experience. The spatio-temporal storage dimension feature reflects the deviation distance between this historical economic data and the spatio-temporal data center point. Each historical economic data in the spatio-temporal associated historical data set is stored in order according to the size of the spatio-temporal storage dimension feature, so as to ensure that during the spatio-temporal polygon retrieval process, unilateral retrieval is performed based on the polygon data range. For example, when retrieving a randomly selected first historical economic data, if it is determined that the spatio-temporal storage dimension feature of this historical economic data is higher than the maximum distance value between the corresponding polygon retrieval boundary and the spatio-temporal data center point, there is no need to continue traversing the entire distributed storage node for retrieval and comparison. Only the economic storage data in the previous order needs to be retrieved according to the storage order of the spatio-temporal storage dimension feature, which greatly reduces the retrieval range of the spatio-temporal polygon retrieval and improves the rapidity of the retrieval result.

[0052] In step S104, based on the spatio-temporal storage dimension features respectively corresponding to each historical economic data, determine the storage location of this economic acquisition data, and use the spatio-temporal storage dimension feature of this economic acquisition data as the spatio-temporal retrieval information for distributed retrieval.

[0053] It should be noted that in this application, the distance deviation value between the historical data and the spatio-temporal data center point is compressed into the spatio-temporal storage dimension feature, and stored in order with this feature as the index, so that the complex spatio-temporal polygon retrieval problem is transformed into one-way feature value range filtering, thereby greatly reducing the retrieval complexity and system calculation overhead.

[0054] Optionally, in some embodiments, determining the storage location of this economic acquisition data based on the spatio-temporal storage dimension features respectively corresponding to each historical economic data specifically includes: obtaining the spatio-temporal storage dimension features respectively corresponding to each historical economic data, and based on the spatio-temporal storage dimension feature corresponding to this economic acquisition data and the sequential storage rule of the spatio-temporal storage dimension feature, determine the storage location of this economic acquisition data.

[0055] Optionally, in some embodiments, during the process of using the spatio-temporal storage dimension feature of this economic acquisition data as the spatio-temporal retrieval information for distributed retrieval, use the spatio-temporal storage dimension feature of this economic acquisition data as the feature value of the incremental retrieval dimension, and perform spatio-temporal polygon retrieval according to multiple polygon boundary retrieval conditions and incremental retrieval dimension conditions.

[0056] In specific implementation, for example, set the boundary conditions of the spatio-temporal polygon retrieval algorithm as follows: time range: January 2024 to December 2024, spatial boundary: an irregular area formed by polygon vertices (such as a city, a region), type conditions: event type, industry classification, tags, etc. Set the spatio-temporal storage dimension feature value as a one-dimensional numerical dimension approximately representing the offset distance of the data relative to the spatio-temporal data center point. Among them, based on the boundary conditions of the spatio-temporal polygon retrieval algorithm, determine the maximum allowable feature value of the incremental retrieval dimension (i.e., the maximum feature distance between the center point and the boundary), and construct a joint retrieval rule for the constraint conditions of adding the spatio-temporal storage dimension feature value to the spatio-temporal polygon retrieval algorithm, so as to improve the retrieval speed of the spatio-temporal polygon retrieval. Among them, if the minimum value of the spatio-temporal storage dimension feature value in a certain distributed node is higher than the maximum allowable feature value, it means that all the data in this node is greater than the query feature value, and this distributed storage node can be directly skipped.

[0057] In addition, on the other hand of the present application, in some embodiments, the present application provides a distributed big data storage system, and the device includes a data storage unit. Refer to Figure 2 , this figure is a schematic diagram of the exemplary hardware and / or software structure of the data storage unit shown according to some embodiments of the present application. The data storage unit 200 includes: a storage node selection module 201, a data processing module 202, and a data storage control module 203, which are described as follows:

[0058] The storage node selection module 201 is configured to obtain a plurality of distributed storage nodes, and determine the feature distances corresponding to the economic acquisition data to be stored and each distributed storage node respectively according to the feature tags corresponding to the economic acquisition data to be stored;

[0059] The data processing module 202 is configured to allocate the economic acquisition data to the distributed storage node with the smallest feature distance, and determine the spatio-temporal polygon retrieval granularity corresponding to the distributed storage node based on the historical economic data in the distributed storage node;

[0060] The data processing module 202 is further configured to perform polygon retrieval determination on the historical economic data in the distributed storage node according to the spatio-temporal polygon retrieval granularity, obtain a spatio-temporal associated historical data set, and obtain the spatio-temporal storage dimension features corresponding to each historical economic data in the spatio-temporal associated historical data set;

[0061] The data storage control module 203 determines the storage location of the economic acquisition data based on the spatio-temporal storage dimension features corresponding to each historical economic data, and uses the spatio-temporal storage dimension features of the economic acquisition data as spatio-temporal retrieval information for distributed retrieval.

[0062] The above text has introduced in detail an example of a distributed big data storage system and method provided by the embodiments of the present application. It can be understood that, in order to implement the above functions, the corresponding device includes the corresponding hardware structure and / or software module for executing each function.

[0063] Those skilled in the art should easily realize that, for each example of the units and algorithm steps described in combination with the embodiments disclosed in this article, the present application can be implemented in the form of hardware or a combination of hardware and computer software. Whether a certain function in the application is executed in the way of hardware or computer software driving hardware depends on the specific application and design constraints of the technical solution. Therefore, professionals can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present application.

[0064] In addition, the present application also provides a computer terminal device, which includes a memory and a processor. The memory stores code, and the processor is configured to obtain the code and execute the above-mentioned distributed big data storage method.

[0065] In some embodiments, refer to Figure 3 , this figure is a schematic structural diagram of a computer terminal device for implementing a distributed big data storage method according to some embodiments of the present application. The above-mentioned distributed big data storage method in the embodiments can be implemented by Figure 3 the computer terminal device shown. The computer terminal device 300 includes at least one communication bus 301, a communication interface 302, a processor 303, and a memory 304.

[0066] The processor 303 can be a general-purpose central processing unit (CPU), or an application-specific integrated circuit (ASIC), or one or more for controlling the execution of the distributed big data storage method in the present application.

[0067] The communication bus 301 may include a path for transmitting information between the above components.

[0068] The memory 304 can be a read-only memory (ROM) or other types of static storage devices that can store static information and instructions, a random access memory (RAM) or other types of dynamic storage devices that can store information and instructions, or can also be an electrically erasable programmable read-only memory (EEPROM), a compact disc read-only memory (CD-ROM) or other optical disc storage, optical disc storage (including compact discs, laser discs, optical discs, digital versatile discs, Blu-ray discs, etc.), magnetic disks or other magnetic storage devices, or any other medium that can be used to carry or store the desired program code in the form of instructions or data structures and can be accessed by a computer, but is not limited thereto. The memory 304 can exist independently and be connected to the processor 303 through the communication bus 301. The memory 304 can also be integrated with the processor 303.

[0069] Among them, the memory 304 is used to store the program code for executing the solution of this application and is controlled by the processor 303 for execution. The processor 303 is used to execute the program code stored in the memory 304. The program code can include one or more software modules. The determination of the polygon retrieval granularity in the above embodiments can be implemented by one or more software modules in the processor 303 and the program code in the memory 304.

[0070] The communication interface 302 uses any device such as a transceiver to communicate with other devices or communication networks, such as Ethernet, radio access network (RAN), wireless local area networks (WLAN), etc.

[0071] Optionally, the above computer terminal device 300 can further include a power supply 305 for supplying power to various components or circuits in the real-time computer terminal device.

[0072] In a specific implementation, as an embodiment, the computer terminal device can include multiple processors, and each of these processors can be a single-core (single-CPU) processor or a multi-core (multi-CPU) processor. Here, the processor can refer to one or more devices, circuits, and / or processing cores for processing data (such as computer program instructions).

[0073] The above computer terminal device may be a general computer terminal device or a dedicated computer terminal device. In specific implementations, the computer terminal device may be a desktop computer, a laptop computer, a network server, a personal digital assistant (PDA), a mobile phone, a tablet computer, a wireless terminal device, a communication device, or an embedded device. The embodiments of the present application do not limit the type of the computer terminal device.

[0074] In addition, in other aspects of the present application, a computer-readable storage medium is provided. The computer-readable storage medium stores at least one computer program, and the computer program is loaded and executed by a processor to implement the operations performed by the above-mentioned distributed big data storage method.

[0075] In summary, in a distributed big data storage system and method disclosed in the embodiments of the present application, first, a plurality of distributed storage nodes are obtained, and according to the feature tags corresponding to the economic acquisition data to be stored, the feature distances between the economic acquisition data and each of the distributed storage nodes are determined; the economic acquisition data is allocated to the distributed storage node with the smallest feature distance, and based on the historical economic data in the distributed storage node, the spatio-temporal polygon retrieval granularity corresponding to the distributed storage node is determined; the historical economic data in the distributed storage node is subjected to polygon retrieval determination according to the spatio-temporal polygon retrieval granularity to obtain a spatio-temporal associated historical data set, and the spatio-temporal storage dimension features corresponding to each historical economic data in the spatio-temporal associated historical data set are obtained; based on the spatio-temporal storage dimension features corresponding to each historical economic data, the storage location of the economic acquisition data is determined, and the spatio-temporal storage dimension features of the economic acquisition data are used as spatio-temporal retrieval information for distributed retrieval, so that storage and retrieval can be performed based on the spatio-temporal storage dimension features of the data to be stored, and the retrieval efficiency of spatio-temporal polygon retrieval of the distributed storage node is improved.

[0076] The above are only the embodiments of the present application, and common specific technical solutions or characteristics and the like in the solutions are not described in detail herein. It should be noted that for those skilled in the art, without departing from the technical solutions of the present application, several deformations and improvements can still be made, and these should also be regarded as the protection scope of the present application, and these will not affect the implementation effect of the present application and the practicability of the patent.

[0077] The scope of protection claimed in this application shall be determined by the content of its claims, and the specific implementation manners and the like described in the specification may be used to interpret the content of the claims. Obviously, those skilled in the art can make various changes and modifications to this application without departing from the spirit and scope of the present invention. Thus, if these modifications and variations of this application fall within the scope of the claims of this application and their equivalent technologies, this application is also intended to cover these changes and modifications.

Claims

1. A distributed-based big data storage method, characterized in that, Including: Obtain multiple distributed storage nodes, and determine the characteristic distances between the economic acquisition data to be stored and each distributed storage node according to the characteristic tags corresponding to the economic acquisition data to be stored; Allocate the economic acquisition data to the distributed storage node with the smallest characteristic distance, and determine the spatio-temporal polygon retrieval granularity corresponding to the distributed storage node based on the historical economic data in the distributed storage node; Perform polygon retrieval determination on the historical economic data in the distributed storage node according to the spatio-temporal polygon retrieval granularity to obtain a spatio-temporal associated historical data set, and obtain the spatio-temporal storage dimension characteristics corresponding to each historical economic data in the spatio-temporal associated historical data set; Determine the storage location of the economic acquisition data based on the spatio-temporal storage dimension characteristics corresponding to each historical economic data, and use the spatio-temporal storage dimension characteristics of the economic acquisition data as the spatio-temporal retrieval information for distributed retrieval.

2. The method according to claim 1, characterized in that In the process of determining the characteristic distances between the economic acquisition data and each distributed storage node, convert the characteristic values on each spatio-temporal characteristic dimension in the characteristic tags corresponding to the economic acquisition data to be stored into distance detection vectors, and obtain the characteristic vectors corresponding to the storage center data of each distributed storage node. Take the Euclidean distance values between the distance detection vectors and each characteristic vector as the characteristic distances between the economic acquisition data and each distributed storage node.

3. The method according to claim 1, characterized in that, Based on the historical economic data in the distributed storage node, determining the spatio-temporal polygon retrieval granularity corresponding to the distributed storage node specifically includes: Obtain the historical economic data in the distributed storage node and construct a historical economic data space; In the spatio-temporal characteristic dimension space, obtain the historical economic data space, and obtain the difference between the maximum characteristic value and the minimum characteristic value corresponding to each spatio-temporal characteristic dimension in the historical economic data space as the data range corresponding to each spatio-temporal characteristic dimension, and take the ratio between the data range corresponding to each spatio-temporal characteristic dimension and the corresponding standard data range as the data density characteristic corresponding to each spatio-temporal characteristic dimension; Perform weighted fusion on the data density characteristics in each spatio-temporal characteristic dimension based on the dimension weights corresponding to each spatio-temporal characteristic dimension to obtain the spatio-temporal polygon retrieval granularity corresponding to the distributed storage node.

4. The method according to claim 1, characterized in that, Performing polygon retrieval determination on the historical economic data in the distributed storage node according to the spatio-temporal polygon retrieval granularity to obtain a spatio-temporal associated historical data set specifically includes: performing mapping according to the spatio-temporal polygon retrieval granularity to obtain the spatio-temporal retrieval distances corresponding to each spatio-temporal characteristic dimension. For any one spatio-temporal characteristic dimension, take the farthest historical economic data within the corresponding spatio-temporal retrieval distance from the economic acquisition data as the polygon retrieval vertex data, and form a polygon retrieval boundary based on the polygon retrieval vertex data corresponding to each spatio-temporal characteristic dimension, and extract all the historical economic data within the boundary of the polygon retrieval boundary to form a spatio-temporal associated historical data set.

5. The method according to claim 1, characterized in that, In the process of obtaining the spatio-temporal storage dimension features corresponding to each historical economic data in the spatio-temporal associated historical data set, the spatio-temporal center point distance corresponding to each historical economic data in the spatio-temporal associated historical data set is used as the spatio-temporal storage dimension feature. The spatio-temporal center point distance is based on a preset spatio-temporal data center point, and the characteristic singular value of the acquisition tag between the historical economic data and the spatio-temporal data center point is used as the spatio-temporal storage dimension feature. Among them, the characteristic singular value is the weighted result of the characteristic difference values of the historical economic data and the spatio-temporal data center point in each spatio-temporal feature dimension.

6. The method according to claim 1, wherein In the process of using the spatio-temporal storage dimension feature of this economic acquisition data as the spatio-temporal retrieval information for distributed retrieval, the spatio-temporal storage dimension feature of this economic acquisition data is used as the eigenvalue of the incremental retrieval dimension, and spatio-temporal polygon retrieval is performed according to multiple polygon boundary retrieval conditions and incremental retrieval dimension conditions.

7. The method according to claim 1, characterized in that Each historical economic data in the spatio-temporal associated historical data set is sequentially stored according to the size of the spatio-temporal storage dimension feature.

8. A distributed big data storage system, comprising a data storage unit, characterized in that, The data storage unit includes: A storage node selection module, configured to obtain multiple distributed storage nodes, and determine the characteristic distances corresponding to the economic acquisition data and each distributed storage node according to the characteristic tags corresponding to the economic acquisition data to be stored; A data processing module, configured to allocate the economic acquisition data to the distributed storage node with the smallest characteristic distance, and determine the spatio-temporal polygon retrieval granularity corresponding to the distributed storage node based on the historical economic data in the distributed storage node; The data processing module is further configured to perform polygon retrieval determination on the historical economic data in the distributed storage node according to the spatio-temporal polygon retrieval granularity, obtain a spatio-temporal associated historical data set, and obtain the spatio-temporal storage dimension features corresponding to each historical economic data in the spatio-temporal associated historical data set; A data storage control module, determines the storage location of this economic acquisition data based on the spatio-temporal storage dimension features corresponding to each historical economic data, and uses the spatio-temporal storage dimension feature of this economic acquisition data as the spatio-temporal retrieval information for distributed retrieval.

9. A computer terminal device, characterized in that, The computer terminal device includes a memory and a processor. The memory stores code, and the processor is configured to obtain the code and execute a distributed-based big data storage method according to any one of claims 1 to 7.

10. A computer-readable storage medium storing at least one computer program, characterized in that, The computer program is loaded and executed by the processor to implement the operations performed by a distributed-based big data storage method according to any one of claims 1 to 7.

Citation Information

Patent Citations

  • Data storage method and system based on intelligent detection platform

    CN118331968A

  • Method, system and device for retrieving spatio-temporal data and program product

    CN118839072A

  • Task scheduling method, device and storage medium for spatiotemporal big data

    CN119739745A

  • Metadump Spatial Database System

    US20170161308A1