Method and system for referencing data
By breaking down data into vectorized data by bytes and returning references, it solves the scale and security issues of data centers and achieves efficient recycling and secure storage of data.
Patent Information
- Application Number
- CN202480012205.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2023-03-30
- Filing Date
- 2024-03-25
- Publication Date
- 2025-09-19
- Estimated Expiration
- 2044-03-25
AI Technical Summary
The scale and maintenance costs of data centers continue to increase, and the risk of data leakage due to cyber attacks is high.
The data to be stored is decomposed into multiple bytes and represented by key values to form vectorized data. Its existence and repeatability in the data center are determined, and a reference to the data is returned instead of the complete data to achieve data recycling.
Effectively recycle and utilize data, reduce the increase in the physical scale of data centers, improve data security, and reduce storage space requirements and network attack risks.
Smart Images

Figure CN120677471A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure generally relates to the field of data science, and more particularly to a method for locally or remotely referencing data and a system capable of executing the method. Background Art
[0002] Data centers are typically used to store and share applications and data. They are designed based on computer networks and storage resources to transport shared applications and data. They may include servers, data storage systems, network equipment such as routers and switches, and security systems such as firewalls and encryption systems.
[0003] As the demand for storing more information continues to grow and become unmanageable, data centers are also expanding in size. For example, data centers store data and applications. A server in a data center is computer hardware and / or software programs that provide services to other programs or devices, such as clients. Servers are often categorized by their purpose. Some examples of these server categories include network servers, application servers, proxy servers, virtual servers, file servers, database servers, and printer servers. For example, the purpose of a database server is to host one or more databases. Client applications can execute database queries, retrieving data from, or writing data to, the databases hosted by the server. Another example is a file server, which stores and manages data files (such as text, images, sounds, pictures, and videos) so that computers on the same network can access these files when needed.
[0004] File server design is complicated by competing demands for storage space, access speed, recoverability, security, and budget. This complexity is further exacerbated by the ever-changing environment, where new hardware and technologies quickly render older equipment obsolete. Server storage quickly fills up with the constant generation of new data, necessitating the need for more servers to store new or modified data, which in turn requires more infrastructure, resulting in additional cost, space, and energy requirements.
[0005] In addition, the security of data stored and shared in data centers is always a focus issue.
[0006] The disadvantage of current methods and / or systems for storing / sharing data or files in data centers is that, as time goes by, as more and more files or data are shared or stored on the Internet, the data centers will expand indefinitely. This leads to ever-increasing scale and maintenance costs of the data centers.
[0007] Another drawback of current methods and / or systems for storing / sharing data or files in data centers is cyberattacks. If a data center is attacked intentionally or accidentally by hackers, the risk of data leakage will be high, which is undesirable for users. Summary of the Invention
[0008] Therefore, it would be advantageous to implement a method and system that overcomes or at least mitigates the aforementioned disadvantages. In particular, it would be desirable to be able to recover data stored in data centers and to ensure the security of data stored in data centers. To better address one or more of the aforementioned issues, a method, system, and non-transitory computer-readable storage medium having a computer program stored thereon are provided, the computer program having the features defined in the method claims. Preferred embodiments are defined in the dependent claims.
[0009] Therefore, according to a first aspect, a method for referencing data is provided, the method being performed by a processor operatively connected to one or more data centers in a network. Each of the one or more data centers is configured to store files and corresponding vectors of files. The method includes receiving a request comprising data to be stored; decomposing the data to be stored into a plurality of bytes on a byte-by-byte basis. Each byte of the data is represented by a key value. The method also includes vectorizing the decomposed data to obtain a first vector of data to be stored. The first vector includes at least one value pair. Each value pair includes a first value and a second value. The first value represents one or more bytes with a unique key value in the decomposed data. The second value represents an instance of one or more bytes with a unique key value present in the decomposed data. The method also includes determining whether the vectorized data exists in one or more data centers, and returning a reference to the data to be stored based on the result of the determination step.
[0010] The advantage of the present invention is that this method can effectively recycle existing data, whether stored locally or on shared devices over the internet. This allows the size of the data center to remain virtually constant, even as the amount of data to be shared or stored continues to grow. With a data center of nearly fixed size, even if the amount of data to be shared or stored continues to grow—that is, the amount of information stored in the data center continues to increase—the physical size of the data center will not increase as in traditional data centers, because the data already stored in the data center can be recycled. In other words, this method significantly reduces the rate at which the physical size of the data center increases compared to the amount of data added to the data center.
[0011] The term "file" refers to any type of data file, such as text, images, sounds, pictures, or videos. A data file may include a sequence of bytes. Each byte sequence may include one or more bytes.
[0012] The term "key value" is used to represent bytes in data. All bytes in the data can be represented by a corresponding key value. Data can be represented in different number systems, such as decimal, binary, octal, and hexadecimal. For example, if data is represented in hexadecimal, each byte in the data is represented by a hexadecimal value. The hexadecimal value is the key value of the byte in the data.
[0013] The term "vector" is used to represent data using only the unique key value of the data. A vector includes at least one value pair. Each value pair includes a first value and a second value. The first value represents one or more bytes with a unique key value in the data, and the second value represents an instance of one or more bytes with a unique key value present in the data. A vector containing at least one value pair may include all unique key values that may exist in the data as the first value of the at least one value pair. For example, the data may be represented by hexadecimal values. In this case, each value pair in the vector may include a first value with a unique hexadecimal value present in the hexadecimal data, and a second value representing an instance of one or more bytes with a unique hexadecimal value present in the hexadecimal data. A vector containing at least one value pair may include all unique hexadecimal values that may exist in the hexadecimal data as the first value.
[0014] The phrase "one or more data centers" is used to refer to a data center capable of storing a file and its vectors.
[0015] According to some embodiments of the present invention, determining whether the vectorized data exists in one or more data centers and returning a reference to the data to be stored based on the result of the determining step may further include: comparing the first vector with vectors in the one or more data centers. If the first vector exists in the one or more data centers, determining whether the data to be stored is a duplicate of a first file already in the one or more data centers. The first file corresponds to a vector having the same value pair as the first vector. If the data to be stored is a duplicate of the first file already in the one or more data centers, returning a reference to the user. The reference is a reference ID of the first file already in the one or more data centers.
[0016] According to some embodiments of the present invention, determining whether the vectorized data exists in one or more data centers may further include comparing the first vector with vectors in the one or more data centers. If the first vector does not exist in any of the one or more data centers, or if the data to be stored is not a duplicate of a first file already in the one or more data centers, then selecting a vector in the one or more data centers that has the highest similarity value compared to the first vector. Determining whether the vectorized data exists in the one or more data centers may further include determining one or more repeated byte sequences in the data to be stored and in a second file corresponding to the selected vector in the one or more data centers, wherein the length of the repeated byte sequence is equal to or less than a first threshold. Returning a reference to the data to be stored may further include returning a first reference, the first reference including one or more sub-references, each sub-reference corresponding to a repeated byte sequence in the repeated byte sequence, each sub-reference including an index and a range of the repeated byte sequence in the second file.
[0017] According to some embodiments of the present invention, the method may further include injecting one or more byte sequences not found in the second file, wherein the length of the unfound byte sequences is less than a second threshold.
[0018] According to some embodiments of the present invention, the method may further include: vectorizing each byte sequence not found in the second file to obtain a respective second vector, wherein the length of each byte sequence not found in the second file is equal to or greater than a second threshold. The following steps are performed for each byte sequence: (i) Determine whether the second vector exists in one or more data centers. (ii) If the second vector exists in one or more data centers, determine whether the byte sequence is repeated in a third file already in one or more data centers, wherein the third file corresponds to a vector having the same value pair as the second vector. (iii) If the byte sequence is repeated in the third file, return a second reference. The second reference is a reference ID of the third file, or includes an index and range of repeated bytes in the third file. (iv) If the second vector does not exist in one or more data centers, or if the byte sequence is not repeated in the third file, add the byte sequence to the one or more data centers and return a third reference. The third reference is a reference ID of a byte sequence that is not repeated in the third file.
[0019] According to some embodiments of the present invention, the returned reference to the data to be stored may include at least one of the first reference, the injected byte sequence, one or more second references, and one or more third references.
[0020] According to some embodiments of the present invention, the method may further include reassembling the data to be stored according to the references.
[0021] According to some embodiments of the present invention, decomposing the data to be stored into bytes may include expressing each byte in hexadecimal form.
[0022] According to a second aspect, a system for referencing data includes a processor and a memory, wherein instructions are stored in the memory. When the instructions are executed by the processor, the processor performs the above method.
[0023] According to a third aspect, a non-transitory computer-readable storage medium stores instructions, which, when executed by a computer, cause the computer to perform the method of the first aspect.
[0024] The advantage of this method and system is that when data is stored or shared, a reference is returned rather than the entire data being stored or shared. Therefore, only a reference to the data is stored on the user terminal, not the entire data itself; the entire data can be obtained through the reference. This reduces the amount of data storage space required on the user terminal. Furthermore, since all data stored in the data center can be recycled and combined into different data sets, the size of the data center can be kept relatively constant without having to save the entire data each time data is added to the data center.
[0025] The method and system may further benefit from the concept that all data stored in the data center can be recycled and combined into different data, thereby significantly improving the security of the data stored therein. BRIEF DESCRIPTION OF THE DRAWINGS
[0026] These and other aspects will now be described in more detail with reference to the embodiments shown in the accompanying drawings.
[0027] Figure 1 An exemplary flow chart of a method for referencing data, the method being performed by a processor operatively connected to one or more data centers in a network, is shown.
[0028] Figure 1a An exemplary flow chart of a method for referencing data is shown, the method additionally / alternatively being performed by a processor operatively connected to one or more data centers in a network.
[0029] Figure 2 A system for referencing data is schematically shown.
[0030] All figures are schematic diagrams and are not necessarily drawn to scale. Generally, only parts necessary for illustrating the embodiments are shown, and other parts may be omitted or merely provided as hints. Throughout the specification, the same reference numerals refer to the same elements. DETAILED DESCRIPTION
[0031] Various aspects of the present invention will now be described more fully with reference to the accompanying drawings, which illustrate presently preferred embodiments. However, these aspects may be embodied in many different forms and should not be construed as limited to the embodiments set forth herein. Rather, these embodiments are provided to be thorough and complete and to fully convey the scope of the present invention to those skilled in the art.
[0032] See also Figure 1The figure shows a flow chart of a method 100 for referencing data, performed by a processor operatively connected to one or more data centers in a network. Each of the one or more data centers is configured to store files and corresponding vectors of the files. Method 100 includes: step 102, receiving a request including data to be stored. For example, the request to store data can be received from a user via an application on a mobile terminal or software on a computer. In some embodiments, the data to be stored can be any suitable data type, such as text, images, sounds, pictures, videos, or any combination thereof. For example, the data to be stored is a text message: "Hello world." The method also includes step 104, decomposing the data to be stored into multiple bytes. In some embodiments, different number systems can be used to further represent the byte-by-byte decomposition of the data to be stored, such as decimal, binary, octal, or hexadecimal. Each byte of the decomposed data is represented by a corresponding key value. For example, if the decomposed data is represented in hexadecimal, each byte in the data is represented by a hexadecimal value. The hexadecimal value is the key value of the byte in the data. Taking the text message "Hello world" as an example, when it is decomposed, the bytes can be represented in hexadecimal values as [48, 65, 6C, 6C, 6F, 20, 77, 6F, 72, 6C, 64], where the "H" in the text message "Hello world" is represented by the corresponding hexadecimal value "48" in the decomposed data [48, 65, 6C, 6C, 6F, 20, 77, 6F, 72, 6C, 64]. Similarly, the "l" in the text message "Hello world" appears three times, and each time it appears in the text message, it is represented by the corresponding hexadecimal value "6C" in the decomposed data [48, 65, 6C, 6C, 6F, 20, 77, 6F, 72, 6C, 64]. In step 106, the decomposed data is vectorized to obtain a first vector of data to be stored. The first vector includes at least one value pair. Each value pair includes a first value and a second value. The first value represents one or more bytes with a unique key value in the decomposed data. The second value represents an instance of one or more bytes with a unique key value in the decomposed data. Taking the text message "Hello world" as an example, when the decomposed data is vectorized, the first vector {48=1, 65=1, 6C=3, 6F=2,...} is obtained. For example, the first vector {48=1, 65=1, 6C=3, 6F=2,...} includes the first value pair "48=1", where "48" is the first value in the first value pair "48=1" and "1" is the second value.The first value "48" in the first value pair "48=1" represents one or more bytes having a unique hexadecimal value of "48" in the decomposed hexadecimal representation data [48, 65, 6C, 6C, 6F, 20, 77, 6F, 72, 6C, 64]. The second value "1" in the first value pair "48=1" represents a unique instance of one or more bytes having a unique hexadecimal value of "48" (i.e., H) in the decomposed hexadecimal representation data [48, 65, 6C, 6C, 6F, 20, 77, 6F, 72, 6C, 64]. Similarly, the second value "3" in the third value pair "6C=3" in the first vector {48=1, 65=1, 6C=3, 6F=2, ...} represents three instances of one or more bytes (i.e., 1) with a unique hexadecimal value "6C" in the decomposed hexadecimal representation data [48, 65, 6C, 6C, 6F, 20, 77, 6F, 72, 6C, 64]. The method also includes step 108, determining whether the vectorized data exists in one or more data centers. The determination may include checking the duplication of the vectorized data in one or more data centers. Different subsequent steps may be performed based on the result of the determination. In some embodiments, this step may be implemented by, for example, determining whether all value pairs of the first vector exist in the vectors of the files stored in one or more data centers, and if so, checking whether the data to be stored is duplicated in one or more data centers, which will be combined with the present application. Figure 1a Other embodiments, such as determining that the data to be stored does not have duplication in one or more data centers, will be described below in conjunction with the present application. Figure 1a Finally, in step 110, a reference to the data to be stored is returned based on the result of determining step 108. In some embodiments, returning the reference may include returning a reference to a file identified in one or more data centers. According to some embodiments, the file identified in the one or more data centers includes all value pairs in the first vector, and the identified file is a duplicate of the data to be stored. For example, the reference may include a reference ID. The reference may be returned to the same or different user, software on a computer, etc. This process will be explained in detail below.
[0033] refer to Figure 1a , a flow chart of a method 100 for referencing data is shown, the method being additionally / alternatively performed by a processor operatively connected to one or more data centers in a network. Figure 1a The following further details the different situations that may occur when quoting data. Figure 1The execution is as follows. In step 108, the method determines whether the vectorized data exists in one or more data centers. This determination may further include: step 109, comparing the first vector with vectors in the one or more data centers; and step 111, if it is determined in step 109 that the first vector exists in the one or more data centers, determining whether the data to be stored is a duplicate of a first file already in the one or more data centers. The first file corresponds to a vector having the same value pairs as the first vector. Taking the text message "Hello world" as an example, in this example, it is determined whether all value pairs in the first vector {48=1, 65=1, 6C=3, 6F=2,...} exist in the vectors of the files stored in the one or more data centers. If it is determined that all value pairs in the first vector exist in the vectors of the files stored in the one or more data centers, step 111, i.e., determining the duplication, may be performed in any suitable manner. According to some embodiments, duplication is determined by checking the order of bytes in the data to be stored. In other words, the data to be stored is compared byte by byte with the first file to determine whether the byte order in the data to be stored is the same as the byte order in the first file. According to some other embodiments, duplication is determined by using a hash algorithm. In one case, if the order is the same, the data corresponding to the first vector is a duplicate of the first file already in one or more data centers. If the data to be stored is a duplicate of the first file already in one or more data centers, a reference is returned to the user in step 110. The reference may be a reference ID of the first file already stored in one or more data centers. According to the present invention, only the reference is returned and stored on the user side without the user noticing. The method may further include step 122, reassembling the data to be stored based on the reference. The reassembly may include restoring the data to be stored based on the reference to the first file already stored in one or more data centers. According to some embodiments, the user may use the reference to obtain the first file through an application in a mobile terminal or software in a computer without the user noticing.
[0034] Returning to step 108, other different scenarios for referencing data will be further described. Step 109 involves comparing the first vector with vectors in one or more data centers. If the comparison reveals that the first vector does not exist in any of the one or more data centers, or if the determination in step 111 reveals that the data to be stored does not duplicate the first file, step 112 is executed, which involves selecting the vector in the one or more data centers with the highest similarity to the first vector. In some embodiments, cosine similarity can be used to calculate the similarity between the first vector and the selected vectors in the one or more data centers. The following explanation will be based on similarity values calculated using cosine similarity. The calculation returns a scalar value between 0 and 1, indicating the similarity value. If the returned similarity value is 1, it indicates that the file corresponding to the selected vector in the one or more data centers is likely a duplicate of the first vector. Various other suitable methods can also be used to calculate similarity values. Therefore, the present invention is not limited to calculating similarity values based on cosine similarity. In step 113, one or more repeating byte sequences are determined in the data to be stored and in the second file corresponding to the vector in the one or more data centers selected in step 112. The length of the repeated byte sequence is equal to or less than a first threshold. For example, the threshold can be any numerical value. The second file corresponding to the selected vector can be different from or the same as the first file in one or more data centers. The second file is the file corresponding to the vector with the highest similarity selected in one or more data centers. In step 114, a first reference to the second file is returned. The first reference includes one or more sub-references. Each sub-reference corresponds to a repeated byte sequence. Each sub-reference includes an index and a range of the repeated byte sequence in the second file. The index can be a file ID, such as the file ID of the second file, which can be expressed as FileID=i02. The range can be defined using the starting byte number in the file corresponding to the selected vector in one or more data centers (e.g., the second file) and the byte length in the file corresponding to the selected vector (e.g., the second file), and the range repeats one of the one or more repeated byte sequences in the second file. The syntax of each sub-reference can be as follows: FileID_StartByte_length. In addition, when the length of the undiscovered byte sequence is less than the second threshold, step 121 is executed to inject one or more byte sequences not found in the second file. The injection can be performed by injecting bytes or byte sequences when the reference is returned in step 110. Typically, the second threshold is less than the first threshold. For example, if the data is represented in a hexadecimal system, the injection may be performed by injecting a hexadecimal value for each byte in one or more byte sequences not found in the second file. However, when the length of each byte sequence not found in the second file is equal to or greater than a second threshold, step 115 is performed to vectorize each byte sequence not found in the second file to obtain a respective second vector.The following steps 116 to 120 may be performed for each byte sequence. In step 116, a determination is made as to whether the second vector exists in one or more data centers. The determination in step 116 may further include comparing the second vector with vectors in the one or more data centers. If the second vector exists in one or more data centers, step 117 is performed to determine whether the byte sequence is duplicated in a third file already in the one or more data centers. The third file corresponds to a vector that has the same value pairs as the second vector. The comparison and duplication check steps are performed similarly to those described above. If the byte sequence is duplicated in the third file, step 118 is performed to return a second reference. The second reference may be a reference ID of the third file, or may include the index and range of the duplicated bytes in the third file. Specifically, if the byte sequence is duplicated in the entire third file, the second reference may be the reference ID of the third file. If the byte sequence is duplicated in a portion of the third file, the second reference may include the index and range of the duplicated bytes in the third file. However, if the second vector does not exist in one or more data centers, or if the byte sequence is not duplicated in the third file, step 119 is performed to add the byte sequence to the one or more data centers. Consequently, in step 120, a third reference to the byte sequence added to the one or more data centers is returned. The third reference may be a reference ID of a byte sequence that is not repeated in the third file. The returned reference 110 of the data to be stored may include at least one of the first reference, the injected byte sequence, one or more second references, and one or more third references. Method 100 may further include step 122 of reassembling the data to be stored based on the references. Reassembling may include restoring the data to be stored based on the references already stored in one or more data centers.
[0035] See also Figure 2 According to one embodiment, a system 200 for referencing data includes a processor 202 and a memory 204. The processor 202 in the system 200 is operatively connected to one or more data centers 206a, 206b, 206c, and 206d in a network. One or more data centers (such as 206a and 206b) can be located in a remote network, such as a network connected via the Internet or a cloud network. One or more data centers (such as 206c and 206d) can also be located in a local network. Each of the one or more data centers 206a, 206b, 206c, and 206d is configured to store files and corresponding vectors of the files. Instructions are stored on the memory 204. When executed by the processor 202, these instructions cause the processor 202 to perform the method 100 for referencing data described in conjunction with the previous figures. The memory 204 is shown as being independently connected to the processor 202. However, those skilled in the art will appreciate that the memory can be built into the processor 202 or arranged to be configured externally to the system 200.
[0036] The memory 204 may be a non-transitory computer-readable storage medium having instructions stored thereon. When the instructions are executed by a computer, the computer executes the Figure 1 and 1 Any one or combination of the methods 100 for referencing data described in a.
[0037] Those skilled in the art will appreciate that the present invention is by no means limited to the preferred embodiments described above. Rather, numerous modifications and variations are possible within the scope of the appended claims. For example, the processors executing the method for referencing data are shown as a single processor. However, they may also function as a group of processors that collectively execute certain portions of the method. Therefore, the embodiments presented in this disclosure are for illustrative purposes only and should not be construed as limiting the scope.
Claims
1. A method (100) for referencing data, performed by a processor operatively connected to one or more data centers (206a, 206b, 206c, 206d) in a network, each of the one or more data centers being configured to store a file and a corresponding vector for the file, the method comprising: receiving (102) a request including data to be stored; Decomposing the data to be stored into a plurality of bytes by byte (104), each byte of the data being represented by a key value; Vectorizing the decomposed data (106) to obtain a first vector of the data to be stored, wherein the first vector includes at least one value pair, each value pair includes a first value and a second value, wherein the first value represents one or more bytes with a unique key value in the decomposed data, and the second value represents an instance of the one or more bytes with a unique key value present in the decomposed data; determining (108) whether vectorized data exists in the one or more data centers; and According to the result of the determining step, a reference to the data to be stored is returned (110).
2. The method according to claim 1, wherein Determining whether the vectorized data exists in the one or more data centers and returning a reference to the data to be stored according to a result of the determining step further includes: comparing the first vector with vectors in the one or more data centers (109); If the first vector exists in the one or more data centers, determining (111) whether the data to be stored is a duplicate of a first file already in the one or more data centers, wherein the first file corresponds to a vector having the same value pair as the first vector; If the data to be stored is duplicated with the first file already in the one or more data centers, a reference is returned to the user, where the reference is a reference ID of the first file already in the one or more data centers.
3. The method according to claim 1, wherein Determining whether the vectorized data exists in the one or more data centers and returning a reference to the data to be stored according to a result of the determining step further includes: comparing the first vector with vectors in the one or more data centers (109); If the first vector does not exist in any of the one or more data centers, or if the data to be stored does not duplicate a first file already in the one or more data centers, selecting (112) a vector in the one or more data centers having a highest similarity value compared to the first vector; determining (113) one or more repeated byte sequences in the data to be stored and in a second file corresponding to the selected vector in the one or more data centers, wherein a length of the repeated byte sequence is equal to or less than a first threshold; and Return (114) a first reference, the first reference including one or more sub-references, each sub-reference corresponding to a repeated byte sequence in the repeated byte sequence, and each sub-reference including an index and a range of the repeated byte sequence in the second file.
4. The method according to claim 3, further comprising: One or more byte sequences not found in the second file are injected (121), wherein the length of the not found byte sequences is less than a second threshold.
5. The method according to claim 4, further comprising: Each byte sequence not found in the second file is vectorized (115) to obtain a respective second vector, wherein the length of each byte sequence not found in the second file is equal to or greater than the second threshold; performing the following steps (i) to (iv) for each byte sequence: (i) determining (116) whether the second vector exists in the one or more data centers; (ii) if the second vector exists in the one or more data centers, determining (117) whether the byte sequence is duplicated in a third file already in the one or more data centers, wherein the third file corresponds to a vector having the same value pair as the second vector; (iii) if the byte sequence is repeated in the third file, returning (118) a second reference, wherein the second reference is a reference ID of the third file or includes an index and range of the repeated bytes in the third file; (iv) if the second vector does not exist in the one or more data centers, or if the byte sequence is not repeated in the third file, adding (119) the byte sequence to the one or more data centers and returning (120) a third reference, the third reference being a reference ID of the byte sequence that is not repeated in the third file.
6. The method according to claim 5, wherein: The returned reference to the data to be stored includes at least one of the first reference, the injected byte sequence, one or more second references, and one or more third references.
7. The method according to any one of claims 1 to 6, further comprising reassembling (122) the data to be stored according to the reference.
8. The method according to any one of claims 1 to 7, wherein Decomposing the data to be stored into bytes includes expressing each byte in hexadecimal form.
9. A system (200) for referencing data, comprising a processor (202) and a memory (204), wherein the memory stores instructions which, when executed by the processor, cause the processor to perform the method according to any one of claims 1 to 8. 10 . A non-transitory computer-readable storage medium having instructions stored thereon, which, when executed by a computer, cause the computer to perform the method according to claim 1 .
Citation Information
Patent Citations
A fuzzy matching-supporting cloud storage data dereplication method
CN105868305A
Repeated sequence identifying method and device, storage medium and electronic equipment
CN110782946A
Large-scale vector data deconstruction and adaptive transmission method and system
CN114048276A
System and method for performing object relational mapping for a data grid
US20120246190A1
Backup and restoration for a deduplicated file system
US20170083408A1