Data storage method, device and equipment based on differential compression and readable medium

By performing quantitative encoding and segmentation and reconstruction of multi-dimensional data, the problem of the depth and edge storage mode of the differential storage tree in high-dimensional data storage is solved, and efficient data query and storage resources are achieved.

CN120104057APending Publication Date: 2025-06-06SHANGHAI QI ZHI INSTITUTE
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510165853.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-14
Publication Date
2025-06-06

AI Technical Summary

Technical Problem

When using high-dimensional data storage and query technology, the tree structure of differential storage is deep, resulting in long query paths and low efficiency of querying data. At the same time, the edge storage method of the differential graph is not optimized, which wastes a lot of storage resources.

Method used

By obtaining the multi-dimensional data to be stored, quantized encoding is performed according to the preset quantized codebook, a difference tree of the difference graph and the minimum spanning tree is generated, and the difference tree is segmented and reconstruction processed, generating the reconstructed differential tree set, and optimizing the edge storage method of the difference graph.

Benefits of technology

The path to query data is simplified, data query efficiency is improved, and storage resources are avoided.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120104057A_ABST
    Figure CN120104057A_ABST
Patent Text Reader

Abstract

The embodiment of the invention discloses a data storage method and device based on differential compression, equipment and a readable medium. A specific embodiment of the method comprises the steps of obtaining to-be-stored multi-dimensional data; according to a preset quantization codebook, mapping the to-be-stored multi-dimensional data into a quantization code block so as to perform quantization coding; generating a difference graph according to the quantized code block; generating a difference tree corresponding to the quantized code block according to the difference graph; and performing segmentation and reconstruction processing on the difference tree to generate a reconstructed difference tree set, and storing the reconstructed difference tree set in a database and a cache. According to the embodiment, the data query path is simplified, the data query efficiency is improved, and waste of storage resources is avoided.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Embodiments of the present disclosure relate to the field of computer technology, and in particular to a data storage method, apparatus, device, and readable medium based on differential compression. Background Art

[0002] When performing data storage, high-dimensional data storage and query technology is usually used. At present, when using high-dimensional data storage and query technology for data storage, the method usually adopted is: by dividing the high-dimensional vector into multiple low-dimensional subspaces, and independently quantizing each subspace, generating a quantization code, and realizing data compression storage and efficient approximate query through differential storage.

[0003] However, when using the above method for data storage, the following technical problems often occur:

[0004] The tree structure of differential storage is usually deep, resulting in long query paths and low efficiency in querying data. At the same time, the edge storage method of the differential graph is not optimized, which wastes a lot of storage resources.

[0005] The above information disclosed in this Background section is only for enhancement of understanding of the background of the inventive concept and therefore it may contain information that does not form the prior art that is already known in this country to a person of ordinary skill in the art. Summary of the invention

[0006] The content of this disclosure is used to introduce concepts in a brief form, which will be described in detail in the detailed implementation section below. The content of this disclosure is not intended to identify the key features or essential features of the technical solution claimed for protection, nor is it intended to limit the scope of the technical solution claimed for protection.

[0007] Some embodiments of the present disclosure propose a data storage method, device, electronic device, and computer-readable medium based on differential compression to solve one or more of the technical problems mentioned in the above background technology section.

[0008] In a first aspect, some embodiments of the present disclosure provide a data storage method based on differential compression, the method comprising: acquiring multidimensional data to be stored; mapping the multidimensional data to be stored into a quantization code group according to a preset quantization code book for quantization encoding, wherein the number of quantization codes included in the quantization code group is a preset number; generating a differential graph according to the quantization code group; generating a differential tree corresponding to the quantization code group according to the differential graph, wherein the differential tree is a minimum spanning tree in the differential graph; performing segmentation and reconstruction processing on the differential tree to generate a reconstructed differential tree set, and storing the reconstructed differential tree set in a database and a cache.

[0009] In a second aspect, some embodiments of the present disclosure provide a data storage device based on differential compression, the device comprising: an acquisition unit, configured to acquire multidimensional data to be stored; a mapping unit, configured to map the multidimensional data to be stored into a quantization code group according to a preset quantization code book for quantization encoding, wherein the number of quantization codes included in the quantization code group is a preset number; a first generation unit, configured to generate a differential graph according to the quantization code group; a second generation unit, configured to generate a differential tree corresponding to the quantization code group according to the differential graph, wherein the differential tree is a minimum spanning tree in the differential graph; a splitting and reconstruction unit, configured to split and reconstruct the differential tree to generate a reconstructed differential tree set, and store the reconstructed differential tree set in a database and a cache.

[0010] In a third aspect, some embodiments of the present disclosure provide an electronic device comprising: one or more processors; a storage device on which one or more programs are stored, and when the one or more programs are executed by the one or more processors, the one or more processors implement the method described in any implementation manner of the above-mentioned first aspect.

[0011] In a fourth aspect, some embodiments of the present disclosure provide a computer-readable medium having a computer program stored thereon, wherein when the program is executed by a processor, the method described in any implementation manner of the above-mentioned first aspect is implemented.

[0012] The above-mentioned various embodiments of the present disclosure have the following beneficial effects: through the data storage method based on differential compression of some embodiments of the present disclosure, the path of querying data is simplified, the efficiency of data query is improved, and the waste of storage resources is avoided. Specifically, the reasons for the long query path, low efficiency of querying data, and waste of more storage resources are: the tree structure of differential storage is usually deep, resulting in a long query path and low efficiency of querying data. At the same time, the edge storage method of the differential graph is not optimized, thereby wasting more storage resources. Based on this, the data storage method based on differential compression of some embodiments of the present disclosure first obtains the multidimensional data to be stored. Thus, the multidimensional data to be stored can be determined. Secondly, according to the preset quantization codebook, the above-mentioned multidimensional data to be stored is mapped into a quantization code group for quantization encoding. Thus, the multidimensional data can be quantized and encoded, thereby compressing the data storage space. Then, according to the above-mentioned quantization code group, a differential graph is generated. Thus, the relationship between the quantization codes can be accurately displayed and redundant information can be reduced. Finally, according to the above difference graph, a difference tree corresponding to the above quantization code group is generated; the above difference tree is split and reconstructed to generate a reconstructed difference tree set, and the reconstructed difference tree set is stored in the database and cache. Thus, the storage of the reconstructed difference tree is completed, and because the edge storage method of the query difference graph is optimized and the tree structure of the difference storage is simplified, the path of querying data can be simplified, the data query efficiency is improved, and the waste of storage resources is avoided. BRIEF DESCRIPTION OF THE DRAWINGS

[0013] The above and other features, advantages and aspects of the embodiments of the present disclosure will become more apparent with reference to the following detailed description in conjunction with the accompanying drawings. Throughout the accompanying drawings, the same or similar reference numerals represent the same or similar elements. It should be understood that the drawings are schematic and that components and elements are not necessarily drawn to scale.

[0014] Figure 1 is a flow chart of some embodiments of a data storage method based on differential compression according to the present disclosure;

[0015] Figure 2 is a schematic structural diagram of some embodiments of a data storage device based on differential compression according to the present disclosure;

[0016] Figure 3 It is a schematic diagram of the structure of an electronic device suitable for implementing some embodiments of the present disclosure. DETAILED DESCRIPTION

[0017] Embodiments of the present disclosure will be described in more detail below with reference to the accompanying drawings. Although certain embodiments of the present disclosure are shown in the accompanying drawings, it should be understood that the present disclosure can be implemented in various forms and should not be construed as being limited to the embodiments set forth herein. On the contrary, these embodiments are provided to provide a more thorough and complete understanding of the present disclosure. It should be understood that the drawings and embodiments of the present disclosure are only for exemplary purposes and are not intended to limit the scope of protection of the present disclosure.

[0018] It should also be noted that, for ease of description, only the parts related to the invention are shown in the drawings. In the absence of conflict, the embodiments and features in the embodiments of the present disclosure can be combined with each other.

[0019] It should be noted that the concepts such as "first" and "second" mentioned in the present disclosure are only used to distinguish different devices, modules or units, and are not used to limit the order or interdependence of the functions performed by these devices, modules or units.

[0020] It should be noted that the modifications of "one" and "plurality" mentioned in the present disclosure are illustrative rather than restrictive, and those skilled in the art should understand that unless otherwise clearly indicated in the context, it should be understood as "one or more".

[0021] The names of the messages or information exchanged between multiple devices in the embodiments of the present disclosure are only used for illustrative purposes and are not used to limit the scope of these messages or information.

[0022] The present disclosure will be described in detail below with reference to the accompanying drawings and in conjunction with embodiments.

[0023] Figure 1 The process 100 of some embodiments of the data storage method based on differential compression according to the present disclosure is shown. The data storage method based on differential compression includes the following steps:

[0024] Step 101: Acquire multidimensional data to be stored.

[0025] In some embodiments, the execution subject (eg, server) of the data storage method based on differential compression may obtain the multidimensional data to be stored. In practice, the multidimensional data sent by the associated terminal may be received as the multidimensional data to be stored.

[0026] Step 102: Map the multi-dimensional data to be stored into a quantization code group according to a preset quantization code book to perform quantization coding.

[0027] In some embodiments, the execution subject may map the multi-dimensional data to be stored into a quantization code group according to a preset quantization code book for quantization coding, wherein the number of quantization codes included in the quantization code group is a preset number.

[0028] In practice, the above preset quantization codebook may be generated by the following steps:

[0029] The first step is to initialize the original multi-dimensional space, wherein the dimension of the original multi-dimensional space may be consistent with the dimension of the multi-dimensional data to be stored.

[0030] In the second step, the original multi-dimensional space is divided to generate a preset number of low-dimensional subspaces to obtain a low-dimensional subspace group.

[0031] The third step is to generate a plurality of corresponding cluster centers for each low-dimensional subspace in the low-dimensional subspace group based on a pre-trained clustering algorithm to obtain a cluster center group. The clustering algorithm may be a K-means algorithm.

[0032] The fourth step is to combine the generated cluster center groups into a preset quantization codebook.

[0033] Step 103: Generate a difference map according to the quantization code group.

[0034] In some embodiments, the execution entity may generate a difference map according to the quantization code group.

[0035] In practice, the query graph can be generated by the following steps:

[0036] The first step is to sort the quantization codes included in the quantization code group to generate a quantization code sequence. In practice, the quantization codes included in the quantization code group can be sorted according to a predefined rule.

[0037] The second step is to determine the difference between the quantization code and the next quantization code for each quantization code in the quantization code sequence. In practice, the difference between the quantization code and the next quantization code can be determined as the difference between the quantization code and the next quantization code.

[0038] The third step is to generate a difference graph based on the determined difference values, wherein each node in the difference graph is each quantization code included in the quantization code group.

[0039] Step 104: Generate a difference tree corresponding to the quantization code group according to the difference graph.

[0040] In some embodiments, the execution subject may generate a difference tree corresponding to the quantization code group according to the difference graph, wherein the difference tree is a minimum spanning tree in the difference graph.

[0041] In some optional implementations of some embodiments, the above execution subject may generate a differential tree corresponding to the corresponding quantization code group through the following steps:

[0042] The first step is to select a difference tree that meets a preset condition from the difference graph, wherein the preset condition is that the sum of the edge weights corresponding to the difference tree is the smallest.

[0043] The second step is to convert the tree structure corresponding to the above differential tree into an undirected structure to generate an undirected differential tree as the differential tree corresponding to the above quantization code group.

[0044] Step 105, splitting and reconstructing the differential tree to generate a reconstructed differential tree set, and storing the reconstructed differential tree set in a database and a cache.

[0045] In some embodiments, the execution entity may perform a splitting and reconstruction process on the differential tree to generate a reconstructed differential tree set, and store the reconstructed differential tree set in a database and a cache.

[0046] In the process of adopting technical solutions to solve the above technical problems, the following problems often occur: the structure of a single differential tree is relatively complex. When querying data through a single differential tree, the depth of the differential tree to be accessed is relatively large, the data query efficiency is low, and more memory is required to store the differential tree with a large depth, resulting in a waste of storage resources. Considering the above technical problems and combining the current technical status, we can decide to adopt the following solution.

[0047] In practice, the differential tree can be split and reconstructed by the following steps:

[0048] The first step is to partition the differential tree based on a preset partitioning condition to generate at least one sub-differential tree and obtain a sub-differential tree set. The preset partitioning condition may be to maximize the node similarity within each sub-differential tree and minimize the similarity between different sub-differential trees.

[0049] In the second step, for each sub-difference tree in the above sub-difference tree set, perform the following reconstruction steps:

[0050] The first reconstruction step is to determine the root node corresponding to the above sub-difference tree.

[0051] The second reconstruction step is to determine the average distance between each node in the sub-difference tree and other nodes in the sub-difference tree.

[0052] The third reconstruction step is to select a target node from each node included in the above-mentioned sub-difference tree according to each determined average distance as a reconstruction root node. In practice, the node with the smallest variance of each corresponding average distance can be selected as the reconstruction root node.

[0053] The fourth reconstruction step is to reconstruct the sub-difference tree based on the reconstructed root node to generate a reconstructed difference tree.

[0054] The third step is to combine the generated reconstructed differential trees into a reconstructed differential tree set.

[0055] The first step to the third step as an inventive point of the embodiment of the present disclosure solves the technical problem mentioned in the background technology that "the structure of a single differential tree is relatively complex. When a single differential tree is used to query data, the depth of the differential tree to be accessed is large, the data query efficiency is low, and more memory is required to store the differential tree with a large depth, resulting in a waste of storage resources". The reasons for the low data query efficiency and waste of storage resources are as follows: the structure of a single differential tree is relatively complex. When a single differential tree is used to query data, the depth of the differential tree to be accessed is large, the data query efficiency is low, and more memory is required to store the differential tree with a large depth, resulting in a waste of storage resources. If the above factors are solved, the effect of improving data query efficiency and reducing storage resource waste can be achieved. In order to achieve this effect, some embodiments of the present disclosure firstly, based on a preset partitioning condition, the above differential tree is partitioned to generate at least one sub-differential tree to obtain a sub-differential tree set. Therefore, the depth of the differential tree can be reduced by partitioning the differential tree into at least one subtree, thereby improving the data query efficiency of the differential tree and reducing the space for storing the differential tree. Second, for each sub-difference tree in the above sub-difference tree set, perform the following reconstruction steps: determine the root node corresponding to the above sub-difference tree; determine the average distance between each node in the above sub-difference tree and other nodes in the above sub-difference tree; select the target node from each node included in the above sub-difference tree as the reconstructed root node based on the determined average distances; based on the above reconstructed root node, reconstruct the above sub-difference tree to generate a reconstructed differential tree. In this way, by reconstructing each sub-difference tree, it is possible to ensure that the root node of each sub-difference tree is as central as possible, thereby further improving the data query efficiency. Third, combine the generated reconstructed differential trees into a reconstructed differential tree set. In this way, the division and reconstruction of the differential tree are completed, which improves the data query efficiency and reduces the waste of storage resources.

[0056] In the process of adopting technical solutions to solve the above technical problems, the following problems often occur: Traditional data query algorithms usually need to traverse the entire tree, resulting in low efficiency of data query and a long time for data query. Considering the above technical problems and combining the current technical status, we can decide to adopt the following solution.

[0057] Optionally, after step 105, the following steps are further included:

[0058] The first step is to receive the data query request sent by the target terminal.

[0059] In some embodiments, the execution subject may receive a data query request sent by a target terminal. The target terminal may be a terminal with data query authority connected to the execution subject via a wired or wireless connection. The data query request may be request information representing a query for a certain data.

[0060] The second step is to determine a data query vector corresponding to the data query request according to the data query request.

[0061] In some embodiments, the execution entity may determine a data query vector corresponding to the data query request according to the data query request.

[0062] In the third step, for each reconstructed differential tree in the reconstructed differential tree set, the distance between the root node included in the reconstructed differential tree and the data query vector is determined as the root distance.

[0063] In some embodiments, the execution entity may determine, for each reconstructed differential tree in the reconstructed differential tree set, a distance between a root node included in the reconstructed differential tree and the data query vector as a root distance.

[0064] The fourth step is to sort the reconstructed differential trees included in the reconstructed differential tree set according to the determined root distances to generate a reconstructed differential tree sequence.

[0065] In some embodiments, the execution subject may sort the reconstructed differential trees included in the reconstructed differential tree set according to the determined root distances to generate a reconstructed differential tree sequence. The sorting may be to sort the reconstructed differential trees in ascending order according to the corresponding root distances.

[0066] Step 5: Based on the reconstructed differential tree sequence, perform the following determination steps:

[0067] In the first determination step, the first reconstructed differential tree in the sequence of reconstructed differential trees is determined as the target differential tree.

[0068] In some embodiments, the execution entity may determine the first reconstructed differential tree in the sequence of reconstructed differential trees as the target differential tree.

[0069] The second determination step is to sequentially determine the asymmetric distance between the data query vector and each subtree node in the target differential tree.

[0070] In some embodiments, the execution entity may sequentially determine the asymmetric distance between the data query vector and each subtree node in the target differential tree.

[0071] In a third determination step, in response to determining an asymmetric distance that satisfies a preset condition, the asymmetric distance that satisfies the preset condition is determined as a target asymmetric distance, and a subtree node corresponding to the target asymmetric distance is determined as a nearest neighbor node.

[0072] In some embodiments, the execution subject may determine the asymmetric distance that satisfies the preset condition as the target asymmetric distance in response to determining the asymmetric distance that satisfies the preset condition, and determine the subtree node corresponding to the target asymmetric distance as the nearest neighbor node. The preset condition may be that the asymmetric distance is the shortest.

[0073] In step 6, in response to the asymmetric distance that satisfies the above preset condition being determined in the first reconstructed differential tree in the reconstructed differential tree sequence, the reconstructed differential tree sequence without the first reconstructed differential tree is determined as the reconstructed differential tree sequence, so as to perform the above determination step again.

[0074] In some embodiments, the execution entity may, in response to determining an asymmetric distance that satisfies the preset conditions in the first reconstructed differential tree in the reconstructed differential tree sequence, determine the reconstructed differential tree sequence from which the first reconstructed differential tree is removed as the reconstructed differential tree sequence, so as to perform the determination step again.

[0075] The seventh step is to send the data corresponding to the nearest neighbor node to the target terminal for display.

[0076] In some embodiments, the execution entity may send the data corresponding to the nearest neighbor node to the target terminal for display.

[0077] The above-mentioned first step to the seventh step, as an inventive point of an embodiment of the present disclosure, solves the technical problem mentioned in the background technology that "traditional data query algorithms usually need to traverse the entire tree, resulting in low efficiency of data query and a long time for data query". The reasons for taking a long time to query data are as follows: Traditional data query algorithms usually need to traverse the entire tree, resulting in low efficiency of data query and a long time for data query. If the above factors are solved, the effect of reducing the time of data query and improving the efficiency of data query can be achieved. In order to achieve this effect, some embodiments of the present disclosure first receive a data query request sent by the target terminal. Thus, it can be determined to start the data query. Second, according to the above data query request, determine the data query vector corresponding to the above data query request. Thus, the data to be queried can be determined. Third, for each reconstructed differential tree in the above reconstructed differential tree set, determine the distance between the root node included in the above reconstructed differential tree and the above data query vector as the root distance. Thus, the distance between the vector and the root node of each tree can be determined. Fourth, according to the determined root distances, the reconstructed differential trees included in the above-mentioned reconstructed differential tree set are sorted to generate a reconstructed differential tree sequence. Thus, each differential tree can be sorted in ascending order according to the root distance. Fifth, the first reconstructed differential tree in the reconstructed differential tree sequence is determined as the target differential tree; the asymmetric distance between the above-mentioned data query vector and each subtree node in the above-mentioned target differential tree is determined in turn; in response to the determination of the asymmetric distance that meets the preset conditions, the above-mentioned asymmetric distance that meets the preset conditions is determined as the target asymmetric distance, and the subtree node corresponding to the above-mentioned target asymmetric distance is determined as the nearest neighbor node. Thus, the asymmetric distance between the query vector and the target subtree node can be calculated layer by layer. Sixth, in response to the determination of the asymmetric distance that meets the above-mentioned preset conditions in the first reconstructed differential tree in the reconstructed differential tree sequence, the reconstructed differential tree sequence with the first reconstructed differential tree removed is determined as the reconstructed differential tree sequence, so as to perform the above-mentioned determination step again; the data corresponding to the above-mentioned nearest neighbor node is sent to the above-mentioned target terminal for display. In this way, the query on the data is completed, and the asymmetric distance between the query vector and the target subtree node is calculated layer by layer until the nearest neighbor point is found, thereby determining the queried data and avoiding traversing the entire tree, thereby improving the data query efficiency and reducing the data query time.

[0078] Optionally, after step 105, the following steps are further included:

[0079] The first step is to determine the target node corresponding to the data deletion request in response to receiving the data deletion request sent by the target terminal.

[0080] In some embodiments, the execution subject may determine the target node corresponding to the data deletion request in response to receiving the data deletion request sent by the target terminal. In practice, the data query vector corresponding to the data deletion request may be determined, and then the node corresponding to the data query vector may be determined as the target node.

[0081] The second step is to generate a virtual node tag corresponding to the above target node.

[0082] In some embodiments, the execution subject may generate a virtual node tag corresponding to the target node, wherein the virtual node tag may be a node tag used to represent an empty node.

[0083] The third step is to replace the target node with the virtual node tag to delete the target node.

[0084] In some embodiments, the execution entity may replace the target node with the virtual node tag to delete the target node.

[0085] The fourth step is to determine the insertion data corresponding to the data insertion request in response to receiving the data insertion request sent by the target terminal.

[0086] In some embodiments, the execution subject may determine the insertion data corresponding to the data insertion request in response to receiving the data insertion request sent by the target terminal.

[0087] The fifth step is to determine the distance between the above-mentioned inserted data and the root node of each reconstructed differential tree in the above-mentioned reconstructed differential tree set as the insertion distance, and obtain an insertion distance group.

[0088] In some embodiments, the execution entity may determine the distance between the insertion data and the root node of each reconstructed differential tree in the reconstructed differential tree set as the insertion distance to obtain an insertion distance group.

[0089] The sixth step is to select a target reconstructed differential tree from the reconstructed differential tree set according to the insertion distance group, and insert the insertion data into the target reconstructed differential tree.

[0090] In some embodiments, the execution entity may select a target reconstructed differential tree from the set of reconstructed differential trees according to the insertion distance group, and insert the insertion data into the target reconstructed differential tree.

[0091] Optionally, after the sixth step, in response to the sum of the number of received data deletion requests and data insertion requests being greater than or equal to a preset value, the reconstructed differential tree set is reconstructed.

[0092] In some embodiments, the execution subject may reconstruct the reconstructed differential tree set in response to the sum of the number of received data deletion requests and data insertion requests being greater than or equal to a preset value.

[0093] The above-mentioned various embodiments of the present disclosure have the following beneficial effects: through the data storage method based on differential compression of some embodiments of the present disclosure, the path of querying data is simplified, the efficiency of data query is improved, and the waste of storage resources is avoided. Specifically, the reasons for the long query path, low efficiency of querying data, and waste of more storage resources are: the tree structure of differential storage is usually deep, resulting in a long query path and low efficiency of querying data. At the same time, the edge storage method of the differential graph is not optimized, thereby wasting more storage resources. Based on this, the data storage method based on differential compression of some embodiments of the present disclosure first obtains the multidimensional data to be stored. Thus, the multidimensional data to be stored can be determined. Secondly, according to the preset quantization codebook, the above-mentioned multidimensional data to be stored is mapped into a quantization code group for quantization encoding. Thus, the multidimensional data can be quantized and encoded, thereby compressing the data storage space. Then, according to the above-mentioned quantization code group, a differential graph is generated. Thus, the relationship between the quantization codes can be accurately displayed and redundant information can be reduced. Finally, according to the above difference graph, a difference tree corresponding to the above quantization code group is generated; the above difference tree is split and reconstructed to generate a reconstructed difference tree set, and the reconstructed difference tree set is stored in the database and cache. Thus, the storage of the reconstructed difference tree is completed, and because the edge storage method of the query difference graph is optimized and the tree structure of the difference storage is simplified, the path of querying data can be simplified, the data query efficiency is improved, and the waste of storage resources is avoided.

[0094] Further references Figure 2 As an implementation of the methods shown in the above figures, the present disclosure provides some embodiments of a data storage device based on differential compression. These device embodiments are similar to Figure 1 Corresponding to the method embodiments shown, the data storage device based on differential compression can be specifically applied to various electronic devices.

[0095] like Figure 2As shown, the data storage device 200 based on differential compression of some embodiments includes: an acquisition unit 201, a mapping unit 202, a first generation unit 203, a second generation unit 204 and a splitting and reconstructing unit 205. The acquisition unit 201 is configured to acquire multidimensional data to be stored; the mapping unit 202 is configured to map the multidimensional data to be stored into a quantization code group according to a preset quantization code book to perform quantization coding, wherein the number of quantization codes included in the quantization code group is a preset number; the first generation unit 203 is configured to generate a differential graph according to the quantization code group; the second generation unit 204 is configured to generate a differential tree corresponding to the quantization code group according to the differential graph, wherein the differential tree is a minimum spanning tree in the differential graph; the splitting and reconstructing unit 205 is configured to split and reconstruct the differential tree to generate a reconstructed differential tree set, and store the reconstructed differential tree set in a database and a cache.

[0096] It can be understood that the units recorded in the data storage device 200 based on differential compression are similar to the reference Figure 1 Therefore, the operations, features and beneficial effects described above for the method are also applicable to the data storage device 200 based on differential compression and the units contained therein, and will not be described in detail here.

[0097] Reference below Figure 3 , which shows a schematic diagram of the structure of an electronic device 300 suitable for implementing some embodiments of the present disclosure. The electronic devices in some embodiments of the present disclosure may include but are not limited to mobile terminals such as mobile phones, laptop computers, digital broadcast receivers, PDAs (personal digital assistants), PADs (tablet computers), PMPs (portable multimedia players), vehicle terminals (such as vehicle navigation terminals), etc., and fixed terminals such as digital TVs, desktop computers, etc. Figure 3 The electronic device shown is only an example and should not bring any limitation to the functions and scope of use of the embodiments of the present disclosure.

[0098] like Figure 3 As shown, the electronic device 300 may include a processing device 301 (e.g., a central processing unit, a graphics processing unit, etc.), which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 302 or a program loaded from a storage device 308 into a random access memory (RAM) 303. In the RAM 303, various programs and data required for the operation of the electronic device 300 are also stored. The processing device 301, the ROM 302, and the RAM 303 are connected to each other via a bus 304. An input / output (I / O) interface 305 is also connected to the bus 304.

[0099] Typically, the following devices may be connected to the I / O interface 305: input devices 306 including, for example, a touch screen, a touch pad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, etc.; output devices 307 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; storage devices 308 including, for example, a magnetic tape, a hard disk, etc.; and communication devices 309. The communication devices 309 may allow the electronic device 300 to communicate with other devices wirelessly or by wire to exchange data. Although Figure 3 The electronic device 300 is shown with various devices, but it should be understood that it is not required to implement or possess all the devices shown. More or fewer devices may be implemented or possessed instead. Figure 3 Each block shown in the figure may represent one device, or may represent multiple devices as required.

[0100] In particular, according to some embodiments of the present disclosure, the process described above with reference to the flowchart can be implemented as a computer software program. For example, some embodiments of the present disclosure include a computer program product, which includes a computer program carried on a computer-readable medium, and the computer program includes a program code for executing the method shown in the flowchart. In some such embodiments, the computer program can be downloaded and installed from the network through the communication device 309, or installed from the storage device 308, or installed from the ROM 302. When the computer program is executed by the processing device 301, the above-mentioned functions defined in the method of some embodiments of the present disclosure are executed.

[0101] It should be noted that the computer-readable medium recorded in some embodiments of the present disclosure may be a computer-readable signal medium or a computer-readable storage medium or any combination of the above two. The computer-readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or device, or any combination of the above. More specific examples of computer-readable storage media may include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In some embodiments of the present disclosure, a computer-readable storage medium may be any tangible medium containing or storing a program that can be used by or in combination with an instruction execution system, device or device. In some embodiments of the present disclosure, a computer-readable signal medium may include a data signal propagated in a baseband or as part of a carrier wave, which carries a computer-readable program code. This propagated data signal may take a variety of forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. The computer readable signal medium may also be any computer readable medium other than a computer readable storage medium, which may send, propagate or transmit a program for use by or in conjunction with an instruction execution system, apparatus or device. The program code contained on the computer readable medium may be transmitted using any suitable medium, including but not limited to: wires, optical cables, RF (radio frequency), etc., or any suitable combination of the above.

[0102] In some embodiments, the client and the server may communicate using any currently known or future developed network protocol such as HTTP (HyperText Transfer Protocol), and may be interconnected with any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include a local area network ("LAN"), a wide area network ("WAN"), an internet (e.g., the Internet), and a peer-to-peer network (e.g., an ad hoc peer-to-peer network), as well as any currently known or future developed network.

[0103] The computer-readable medium may be included in the electronic device; or it may exist independently without being installed in the electronic device. The computer-readable medium carries one or more programs. When the one or more programs are executed by the electronic device, the electronic device: obtains the multidimensional data to be stored; maps the multidimensional data to be stored into a quantization code group according to a preset quantization code book for quantization encoding; maps the encoded multidimensional data into a quantization code group according to a preset quantization code book, wherein the number of quantization codes included in the quantization code group is a preset number; generates a differential graph according to the quantization code group; generates a differential tree corresponding to the quantization code group according to the differential graph, wherein the differential tree is the minimum spanning tree in the differential graph; performs segmentation and reconstruction processing on the differential tree to generate a reconstructed differential tree set, and stores the reconstructed differential tree set in a database and a cache.

[0104] Computer program code for performing the operations of some embodiments of the present disclosure may be written in one or more programming languages ​​or a combination thereof, including object-oriented programming languages ​​such as Java, Smalltalk, C++, and conventional procedural programming languages ​​such as "C" or similar programming languages. The program code may be executed entirely on the user's computer, partially on the user's computer, as a separate software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving a remote computer, the remote computer may be connected to the user's computer via any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., via the Internet using an Internet service provider).

[0105] The flow chart and block diagram in the accompanying drawings illustrate the possible architecture, function and operation of the system, method and computer program product according to various embodiments of the present disclosure. In this regard, each square box in the flow chart or block diagram can represent a module, a program segment or a part of a code, and the module, the program segment or a part of the code contains one or more executable instructions for realizing the specified logical function. It should also be noted that in some implementations as replacements, the functions marked in the square box can also occur in a sequence different from that marked in the accompanying drawings. For example, two square boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each square box in the block diagram and / or flow chart, and the combination of the square boxes in the block diagram and / or flow chart can be implemented with a dedicated hardware-based system that performs a specified function or operation, or can be implemented with a combination of dedicated hardware and computer instructions.

[0106] The units described in some embodiments of the present disclosure may be implemented by software or hardware. The units described may also be provided in a processor, for example, may be described as: a processor including an acquisition unit, a mapping unit, a first generation unit, a second generation unit, and a segmentation and reconstruction unit. The names of these units do not, in some cases, constitute limitations on the units themselves, for example, the acquisition unit may also be described as a "unit for acquiring multidimensional data to be stored".

[0107] The functions described above herein may be performed at least in part by one or more hardware logic components. For example, without limitation, exemplary types of hardware logic components that may be used include: field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on chips (SOCs), complex programmable logic devices (CPLDs), and the like.

[0108] The above descriptions are only some preferred embodiments of the present disclosure and an explanation of the technical principles used. Those skilled in the art should understand that the scope of the invention involved in the embodiments of the present disclosure is not limited to the technical solutions formed by a specific combination of the above-mentioned technical features, but should also cover other technical solutions formed by any combination of the above-mentioned technical features or their equivalent features without departing from the above-mentioned inventive concept. For example, the above-mentioned features are replaced with the technical features with similar functions disclosed in the embodiments of the present disclosure (but not limited to) and the technical solutions formed.

Claims

1. A data storage method based on differential compression, comprising: Obtaining multidimensional data to be stored; According to a preset quantization codebook, the multidimensional data to be stored is mapped into a quantization code group for quantization encoding, wherein the number of quantization codes included in the quantization code group is a preset number; Generating a difference map according to the quantization code group; Generating a differential tree corresponding to the quantization code group according to the differential graph, wherein the differential tree is a minimum spanning tree in the differential graph; The differential tree is split and reconstructed to generate a reconstructed differential tree set, and the reconstructed differential tree set is stored in a database and a cache.

2. The method according to claim 1, wherein: The preset quantization codebook is generated by the following steps: Initialize the original multidimensional space; The original multidimensional space is divided to generate a preset number of low-dimensional subspaces to obtain a low-dimensional subspace group; For each low-dimensional subspace in the low-dimensional subspace group, generating a corresponding plurality of cluster centers based on a pre-trained clustering algorithm to obtain a cluster center group; The generated cluster center groups are combined into a preset quantization codebook.

3. The method according to claim 1, wherein: The step of generating a difference map according to the quantization code group comprises: Sorting the quantization codes included in the quantization code group to generate a quantization code sequence; For each quantization code in the quantization code sequence, determining a difference value between the quantization code and a subsequent quantization code; Based on the determined difference values, a difference graph is generated, wherein each node in the difference graph is each quantization code included in the quantization code group.

4. The method according to claim 1, wherein: The step of generating a differential tree corresponding to the quantization code group according to the differential graph includes: Selecting a differential tree that meets a preset condition from the differential graph, wherein the preset condition is that the sum of the edge weights corresponding to the differential tree is the smallest; The tree structure corresponding to the differential tree is converted into an undirected structure to generate an undirected differential tree as the differential tree corresponding to the quantization code group.

5. The method according to claim 1, wherein: The method further comprises: In response to receiving a data deletion request sent by a target terminal, determining a target node corresponding to the data deletion request; Generate a virtual node tag corresponding to the target node; Replacing the target node with the virtual node tag to delete the target node; In response to receiving a data insertion request sent by a target terminal, determining insertion data corresponding to the data insertion request; Determine the distance between the insertion data and the root node of each reconstructed differential tree in the reconstructed differential tree set as the insertion distance, and obtain an insertion distance group; According to the insertion distance group, a target reconstructed differential tree is selected from the reconstructed differential tree set, and the insertion data is inserted into the target reconstructed differential tree.

6. The method according to claim 5, wherein: The method further comprises: In response to the sum of the number of received data deletion requests and data insertion requests being greater than or equal to a preset value, the reconstructed differential tree set is reconstructed.

7. A data storage device based on differential compression, comprising: An acquisition unit, configured to acquire multidimensional data to be stored; A mapping unit is configured to map the multidimensional data to be stored into a quantization code group according to a preset quantization code book to perform quantization encoding, wherein the number of quantization codes included in the quantization code group is a preset number; A first generating unit, configured to generate a difference map according to the quantization code group; A second generating unit is configured to generate a differential tree corresponding to the quantization code group according to the differential graph, wherein the differential tree is a minimum spanning tree in the differential graph; The splitting and reconstruction unit is configured to perform splitting and reconstruction processing on the differential tree to generate a reconstructed differential tree set, and store the reconstructed differential tree set in a database and a cache.

8. An electronic device comprising: one or more processors; a storage device having one or more programs stored thereon; When the one or more programs are executed by the one or more processors, the one or more processors implement the method according to any one of claims 1 to 6.

9. A computer readable medium having a computer program stored thereon, wherein: When the program is executed by a processor, the method according to any one of claims 1 to 6 is implemented.