Attribute Grouping for Change Detection in a Distributed Storage System
By using hash functions to generate group hash values in a distributed storage system, the granularity change detection of document attributes is solved, and the resource waste caused by a package of notifications is improved, and the system efficiency and performance are improved.
Patent Information
- Application Number
- CN202080030959.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2019-04-24
- Filing Date
- 2020-03-30
- Publication Date
- 2025-07-29
- Estimated Expiration
- 2040-03-30
AI Technical Summary
When detecting document changes, the existing distributed storage system uses a package of notifications to cause waste of computing resources, and the computing services in different scenarios have different interest in document attribute changes, resulting in unnecessary computing, network and resource consumption.
By implementing granularity change detection of document attributes in a distributed storage system, using hash functions to associate them with the attributes, generating group hash values, and sending notifications only to the scenario computing service of interest, reducing unnecessary computing and network resource consumption.
It reduces the computing load, network bandwidth and resource consumption of distributed storage systems, improves the efficiency and performance of the system, and avoids the waste of resources for uninterested scenario computing services.
Smart Images

Figure CN113767390B_ABST
Abstract
Description
Background Art
[0001] Distributed storage systems typically include routers, switches, bridges, and other network devices that interconnect a large number of computer servers, network storage devices, and other types of computing devices via wired or wireless network links. Computer servers can host one or more virtual machines, containers, or other types of virtualized components to provide various computing and / or storage services to users. For example, a computer server can be configured to provide data storage and retrieval services that allow users to store, edit, retrieve, or perform other data management tasks. Summary of the Invention
[0002] This summary is provided to introduce some concepts in a simplified form that are further described below in the detailed description. This summary is not intended to identify key features or essential features of the claimed subject matter, nor is it intended to be used to limit the scope of the claimed subject matter.
[0003] Today's distributed storage systems are typically arranged in a hierarchical architecture to provide reliable and scalable storage services to a large number of users. Each layer of the distributed storage system can rely on a corresponding computing system to provide the designed services. For example, a distributed storage system can include a storage layer that is highly scalable and capable of storing a large amount of data. The storage layer typically can include a set of backend servers that are configured to facilitate users' storage, retrieval, and / or other data management tasks. This set of backend servers can also provide computing resources for performing various data analysis or analytics tasks. Examples of such analytics tasks include aggregation of document views, detection of modified signals, calculation of trending documents, etc. An example of such a storage layer is Microsoft Outlook provided by Microsoft Corporation of Redmond, Washington
[0004] Although the storage layer is suitable for performing storage and analysis of large amounts of data, the data structures used in the storage layer may not be suitable for providing random access to individual data items. For example, the storage layer may not be able to access a data item in a list in constant time independent of the position of the data item in the list or the size of the list. Additionally, the scalability of such distributed storage systems is typically achieved by splitting or "sharding" the stored data or indexes to such data into partitions. For example, in some distributed storage systems, the primary index of data items in the storage layer can be split into multiple sub-indexes. Instead of referencing the data item, the primary index can reference a sub-index, which in turn can reference the data item. In such a distributed storage system, when a user requests a stored data item, multiple operations (e.g., fan-out federation) are performed in the storage layer to resolve these references. Performing multiple operations can result in high latency and computational load in servicing user requests.
[0005] One technique for providing fast random access to stored data is to cache a subset of the stored data (sometimes referred to as "high - impact data") in a centralized repository configured to provide cache services for high - impact data. For example, a set of analytics servers and / or computing processes (referred to herein as "ingest processors") running in the storage layer can be configured to push a subset of the stored data as high - impact data to the centralized repository. The high - impact data pushed to the centralized repository can include various categories of stored data that various scenario - computing services may be interested in. In one example, the high - impact data for a document stored in the storage layer cached in the centralized repository can include the document name, document extension (e.g., "txt"), last updated date / time, number of views of the document, number of modifications to the document, the Uniform Resource Locator (URL) where the document can be retrieved, and / or other suitable information related to the document. By caching the high - impact data, users can easily retrieve the document and / or its other related information from the centralized repository.
[0006] Scenario - computing services can be configured to retrieve the cached high - impact data from the centralized repository and react to changes in the underlying data stored in the storage layer to provide a corresponding user experience. For example, in the above - mentioned document example, a search - indexing service may be interested in changes to the title, body, appendix, or other suitable types of document content used to update the search index. On the other hand, a document - metadata service may be interested in the list of viewers of the document and the timestamp of the last modification of the document used to update the number of views / edits of the document.
[0007] In some distributed storage systems, upon detecting any change in a stored document, a blanket notification is sent to all scenario computing services for further processing. Examples of detected changes can include write operations to the properties of the document. However, such blanket notifications can result in waste of computing resources in the distributed storage system. First, the transmission of blanket notifications may involve resource-intensive notification calls (e.g., Hypertext Transfer Protocol or "HTTP" requests or Remote Procedure Calls). Second, different scenario computing services may be interested in changes to different properties of the document. For example, in the document example, for updating the search index, the search index service may not be interested in changes to the list of document viewers or the number of times the document has been modified. On the other hand, the document metadata service may not be interested in changes to the document body or title. Thus, the changes indicated in the blanket notification may be relevant to some scenario computing services but irrelevant to others. Despite the irrelevance of the changes indicated in the blanket notification, all scenario computing services may spend additional computing, network, input / output, and / or other suitable types of resources to determine whether the changes indicated in the blanket notification are relevant to an individual scenario computing service. For example, various types of resources may be spent to perform additional read operations on high-impact data at a centralized repository. Such read operations may result in a high computing load on the server hosting the centralized repository, high network bandwidth consumption in the computer network, and / or other negative impacts on the performance and / or operation of the distributed storage system.
[0008] Several embodiments of the disclosed technology can address at least some aspects of the above-described deficiencies by implementing granular change detection of document properties in a distributed storage system. In some implementations, the distributed storage system can include a storage layer having one or more backend servers, one or more centralized repository servers, and one or more scenario computing servers (which are operably coupled to each other via a computer network). The centralized repository server can be configured to provide a centralized repository and caching services for a subset of the data stored in the storage layer. The scenario computing servers can be configured to provide various types of computing services in response to changes in certain properties of the documents stored in the backend storage servers, such as the search index service and the document metadata service described above.
[0009] The backend server can be configured to provide various data storage services to facilitate storage, retrieval, modification, deletion, and / or other suitable types of data management operations. The backend server can also be configured to provide an ingestion processor that is configured to analyze new and / or updated data received at and to be stored in the storage layer. In some implementations, the ingestion processor can be configured to receive registration or other suitable types of indications from a scenario computing service. The registration individually indicates one or more attributes of a document for which the corresponding scenario computing service wishes to receive notifications of changes to the one or more attributes. Such registration can be saved in a data store as a list, table, or other suitable data structure.
[0010] In some embodiments, a centralized repository can include a change detector that is configured to perform granular change detection on one or more attributes of data items received at the ingestion processor. Although the change detector is described below as part of the centralized repository, in other embodiments, the change detector can also be configured as part of the ingestion processor, an independent computing service in a distributed storage system, or other suitable configurations.
[0011] Using a document as an example, when a new version of a document is received at the ingestion processor, the ingestion processor can extract certain high-impact data from the document and write the extracted high-impact data to the centralized repository. In turn, the change detector at the centralized repository (or its component) can be configured to compare the various values of the attributes of the document in the received new version with the various values of the attributes in the previous version stored in the centralized repository. For example, the change detector can compare the body and / or title of the document to determine whether a change is detected in the body and / or title. In another example, the change detector can also compare the view count and / or modification count of the document in the new version with the view count and / or modification count in the old version.
[0012] In response to detecting a change in at least one of the properties, the change detector may be configured to provide an indication of the detected change to the ingestion processor. In response, the ingestion processor may be configured to generate a notification and send the notification to those scenario computing services that have registered to receive notifications for detected changes in the properties. For example, when a change in the body and / or title of a document is detected, the ingestion processor may be configured to send a notification to the search index service rather than the document metadata service. On the other hand, when a change in the view count or modification count is detected, the ingestion processor may be configured to send a notification to the document metadata service rather than the search index service. Thus, the number of notifications sent by the ingestion processor and subsequent read operations for verifying the relevance of the notifications can be reduced, thereby reducing the computational load, reducing network bandwidth consumption, and / or reducing the consumption of other resources in the distributed storage system.
[0013] In some embodiments, detecting which one or more properties have changed may include associating a hash function with the properties of a document and executing the associated hash function at the ingestion processor to facilitate determining whether the corresponding properties have changed. For example, the ingestion processor may be configured to apply the hash functions associated with the respective properties to the values of the properties in the received new version of the document to derive a hash value (referred to herein as the "new hash value" or H new ). Then, the ingestion processor may push the generated new hash value as part of the high-impact data to a central repository.
[0014] Upon receiving the new hash value, the change detector at the centralized repository may be configured to compare the new hash value with the previous value of the same property of the document stored in the centralized repository (referred to herein as the "previous hash value" or H old ). If the new hash value is different from the previous hash value, the change detector may indicate to the ingestion processor that a change in the document property has been detected. If the new hash value matches or substantially matches (e.g., a match degree greater than 90%) the previous hash value, the change detector may indicate that the corresponding property has not changed.
[0015] Executing the hash function in the ingestion processor can utilize the high scalability of the storage layer to reduce the risk of computational resources in the storage layer becoming a bottleneck. Additionally, by selecting hash functions suitable for each property, as described in more detail below, the consumption of computational resources at the storage layer can be further reduced and the derived hash values can have a small data size. The small data size of the generated values can result in low storage overhead, low network bandwidth consumption in the centralized repository, and efficient hash comparison when determining whether a property has changed.
[0016] Aspects of the disclosed technology also relate to selecting an appropriate hash function to associate with respective attributes of stored data items (such as the aforementioned documents). In some implementations, multiple different hash functions can be statically associated with respective attributes of a document. For example, an identity function that returns the input value as the output can be used for the view count and the document modification count. In another example, the xxHash function can be associated with the body of the document.
[0017] In further implementations, the association of a hash function with an attribute can be dynamic, e.g., based on the data type and / or size of the attribute value. For example, the identity function can be used for attributes whose values are integers or short strings (e.g., fewer than 10 or other suitable number of characters). Thus, the view count and the document modification count can be associated with the identity function. For values that are not integers or short strings, other hash functions are configured to map the data of the attribute value to a fixed-size hash value that is different from the data size of the value. For example, the Fowler-Noll-Vo (FNV) hash function can be associated with attributes having a data size of four to twenty bytes. The xxHash64 function can be associated with attributes having a large data size (e.g., greater than 20 bytes). The Secure Hash Algorithm 256 (SHA256) function can be applied to the values of attributes when conflicts for the same data cannot be tolerated (e.g., when the number of changes to be detected is above a threshold). For example, when nearly every change to an attribute needs to be detected, applying SHA256 to the attribute value can produce a 256-bit hash value (or other threshold-bit hash value) that has a much lower risk of producing conflicts.
[0018] In some implementations, the ingestion processor can be configured to utilize a configuration file that associates hash functions with ranges of attribute value sizes. For example, attributes having values less than or equal to four bytes can be associated with the identity function, while attributes having values between five and twenty bytes can be associated with the FNV hash function. Attributes having values greater than twenty bytes may be associated with xxHash64. Additionally, attributes for which conflicts cannot be tolerated under any circumstances can be associated with SHA256.
[0019] In other implementations, the ingestion processor can be configured to use a cost function to present the dynamic selection of a hash function as an optimization operation. Different hash functions have different storage footprints (e.g., output hash value size), computational resource costs (e.g., computational cost of computing the hash value), conflict rates, and / or other suitable characteristics. Thus, the optimization operation can be defined as: for an attribute, select such a hash function that minimizes the processor cycles and storage overhead, while keeping the number of conflicts below a threshold and minimizing the total amount of attribute data that must be read for a full comparison. An example cost function J( Hi ) can be expressed as follows:
[0020] J( Hi ) = W Collision *CR Hi +W CPU *CPU Hi +W storage *StOrage Hi +W Data *Data Hi ,
[0021] where CR Hi is the conflict rate, CPU Hi is the processor cycle, Storage Hi is the storage size of the hash value, Data Hi ∈ {PropertySize, Storage Hi} and W is the weight associated with each corresponding parameter.
[0022] While the distributed storage system is running, the different weights associated with the above parameters may change. For example, the available system resources (e.g., computing or storage) may change, making the computing resources abundant. Therefore, the weights used for the computing resources can be adjusted to result in a different hash function being selected for the same property than before the change in the available system resources.
[0023] An additional aspect of the disclosed technology involves performing change detection by grouping together associated properties and generating a hash value for the group of properties rather than for each property in the group. Without being bound by theory, it is recognized that certain properties of a data item (e.g., a document) often change together. For example, the number of views often changes with the number of modifications to a document. Therefore, grouping the properties into a group and generating a hash value for the group can allow for a quick determination of whether any of the properties in the group have changed. If no change in the group is detected, a full comparison of each property in the group can be skipped, and metadata can be added to the new version for corresponding indication. If a change in the group is detected, each property in the group can be compared to see which one or more properties have changed.
[0024] According to additional aspects of the disclosed technology, grouping attributes can also utilize optimization operations. If the correlation of changes between different attributes is used to group the attributes, a "group correlation threshold" can be set so that the amount of data to be read for the attribute and hash comparison is below a size threshold and / or minimized. The correlation of the change can include data that indicates the probability (for example, as a percentage) that the first attribute may change when the second attribute changes, and vice versa. The correlation of the change can be developed by statistical analysis of historical data of the attribute or by other suitable means. In some embodiments, if the value of the attribute has been hashed, a combination operation (for example, left shift and performing a bitwise XOR) can be used to derive a group hash value. However, if the attribute has been passed through an identity function, an additional hash operation can be performed on the combination of attribute values (called a group value) to derive the group hash value. In other embodiments, multiple attribute groups can be used, where one attribute can be part of multiple groups.
[0025] One technique for generating attribute groupings may include maintaining a covariance matrix for the attributes. A covariance matrix is a matrix whose element at position i,j is the covariance (i.e., a measure of joint variability) between the i-th element and the j-th element of a random vector. The computation of the covariance matrix can be simplified so that if an attribute has changed, the value of that attribute is 1. If the attribute has not changed, the value of that attribute is 0. Clustering heuristics can then be used to construct groupings by assigning an attribute Pi to an existing cluster (attribute group) if the attribute is related to an attribute group P that is already in a cluster. Cj If some (or all) of the attributes in <Pi,P Cj > values in the covariance matrix of the attribute Pi are above a threshold Tc). If the attribute is not sufficiently correlated (i.e., no attribute Pi has a value above the threshold Tc in the covariance matrix), a new cluster (i.e., a new attribute group) containing only this attribute can be created. Another grouping heuristic can include determining the "closest" cluster based on a cluster similarity function and assigning the attribute to the closest cluster (i.e., attribute group). Regardless of which heuristic is used to create the attribute groupings, the amount of data to be read and compared in order to detect which attributes have changed can be kept below a threshold and / or minimized.
[0026] In some embodiments, the attribute grouping can be computed before the distributed storage system becomes operational, but it can also be computed periodically in an offline batch mode. When the attribute grouping is determined, a bit field (where each bit corresponds to an attribute) can be used to record which attribute has changed. For example, if the corresponding attribute has not changed, the bit can be set to zero. When the corresponding attribute changes, the same bit can be set to one. In further implementations, the attribute grouping can be changed in an online manner to better adapt to the temporal changes of the data. In further implementations, a single attribute can be included in multiple different groups. This inclusion can allow ready change detection. For example, when a change in one group is detected while no change is detected in another group, it can be determined that an attribute in the two groups has not changed. BRIEF DESCRIPTION OF THE DRAWINGS
[0027] Figure 1 is a schematic diagram showing a distributed storage system implementing granularity change detection according to an embodiment of the disclosed technology.
[0028] Figures 2A - 2D is a schematic diagram showing certain hardware / software components of a Figure 1 distributed storage system that performs granularity change detection according to an embodiment of the disclosed technology.
[0029] Figures 3A - 3C is a schematic diagram showing certain hardware / software components of a signature component during signature generation according to an embodiment of the disclosed technology. Figures 2A - 2D of the signature component during signature generation according to an embodiment of the disclosed technology.
[0030] Figure 4 is a schematic diagram showing certain hardware / software components of a Figures 2A - 2D signature component that implements attribute grouping according to an embodiment of the disclosed technology.
[0031] Figure 5 is a Venn diagram showing an example grouping of attributes according to an embodiment of the disclosed technology.
[0032] Figures 6A - 6C is a flowchart showing the process of granularity change detection in a distributed storage system according to an embodiment of the disclosed technology.
[0033] Figure 7 is suitable for Figure 1 certain components of the distributed storage system in DETAILED DESCRIPTION
[0034] Certain embodiments of systems, devices, components, modules, routines, data structures, and processes for granularity change detection in a data center or other suitable distributed storage system are described below. In the following description, specific details of the components are included to provide a thorough understanding of certain embodiments of the disclosed technology. Those skilled in the relevant art will also understand that the technology may have additional embodiments. The technology may also be practiced without several details of the embodiments described below. Figures 1 - 7 Description of the embodiments.
[0035] As used herein, the term "distributed storage system" generally refers to an interconnected computer system having routers, switches, bridges, and other network devices that interconnect a large number of computer servers, network storage devices, and other types of computing devices to each other and / or to an external network (such as the Internet) via wired or wireless network links. Each server or storage node may include one or more permanent storage devices, such as hard disk drives, solid state drives, or other suitable computer-readable storage media.
[0036] As also used herein, a "document" generally refers to a computer file that contains various types of data and metadata that describe and / or identify the computer file. The various types of data and metadata are referred to herein as "attributes" of the document. Example attributes may include a title, a content body, an author, a Uniform Resource Locator (URL), a view count, an edit count, and / or other suitable information related to the document. Each attribute may be defined as a parameter that contains a corresponding value. For example, the author attribute may have a value that contains a name, such as "Nick Jones". The view count attribute may include a value that contains an integer, such as "4".
[0037] As used herein, the "version" of a document generally refers to the state or form of the document. A version may be defined by a version number, an edit date / time, an editor name, a save date / time, and / or other suitable attributes. A "new version" generally refers to a document state that has at least one change to any aspect of the document's data and / or metadata. For example, a new version of a document may include an edit to the document content or a change to one or more attributes of the document (such as the view count). A "previous version" is another state of the document from which a new version may be directly or indirectly derived.
[0038] The "scenario computing service" also used herein generally refers to one or more computing resources provided through a computer network (such as the Internet), which are designed to provide a designed user experience. Example scenario computing services include Software as a Service ("SaaS"), Platform as a Service ("PaaS"), and Infrastructure as a Service ("IaaS"). SaaS is a software distribution technology in which software applications are hosted by a cloud service provider in, for example, a data center and accessed by users through a computer network. PaaS generally refers to the delivery of an operating system and related services through a computer network without the need for downloading or installation. IaaS generally refers to outsourced facilities for supporting storage devices, hardware, servers, network devices, or other components, all of which are accessible through a computer network.
[0039] Various scenario computing services can be configured to provide different user experiences. An example scenario computing service can be configured to compile various attributes of a document into a search index. In another example, another scenario computing service can be configured to receive a signal indicating that a user views a document, and in response to the received signal, provide an electronic message to other users to recommend viewing the same document. In a further example, the scenario computing service can be configured to provide other suitable user experiences.
[0040] In some distributed storage systems, once any change in a document is detected, a blanket notification is sent to all scenario computing services for further processing. However, such a blanket notification may result in waste of computing resources in the distributed storage system. First, the transmission of the blanket notification may involve resource-intensive notification calls (e.g., HTTP requests or remote procedure calls). Second, different scenario computing services may be interested in different attribute changes of the document. For example, for updating the search index, the search index service may not be interested in changes to the list of document viewers or the number of modifications made to the document. On the other hand, the document metadata service may not be interested in changes to the document body or title. Therefore, the changes indicated in the blanket notification may be suitable for some scenario computing services but irrelevant to other scenario computing services.
[0041] Although the changes indicated in the blanket notification are irrelevant, all scenario computing services can expand additional computing, network, input / output, and / or other suitable types of resources to determine whether the changes indicated in the blanket notification are relevant to each scenario computing service. For example, various types of resources may be spent on performing additional read operations. Such read operations can result in a high computing load, high network bandwidth in the computer network, and / or other negative impacts on the performance and / or operation of the distributed storage system.
[0042] Certain embodiments of the disclosed technology can address at least some aspects of the above-described deficiencies by implementing fine-grained change detection of document attributes in a distributed storage system. In some implementations, a change detector can be implemented in the distributed storage system to determine which one or more attributes of a new version of a document have changed compared to a previous version of the same document. Based on the determined one or more attributes, notifications can be selectively sent to scenario computing services that were previously registered to receive notifications of changes to the one or more attributes. Thus, expensive HTTP requests and / or remote procedure calls to scenario computing services that are not interested in the detected attribute changes can be avoided. Additionally, additional computing resources to determine whether the changes indicated in the notification are actually relevant to the scenario computing service can be avoided. Therefore, the computational load, network bandwidth, and / or consumption of other computing resources of the distributed storage system can be reduced compared to using blanket notifications, as described in more detail below with reference to Figures 1 - 7 More detailed description.
[0043] Figure 1 is a schematic diagram illustrating a computing environment 100 that implements fine-grained change detection according to an embodiment of the disclosed technology. As Figure 1 shown, the computing environment 100 can include one or more client devices 102 for a user 101 and a distributed storage system 103 interconnected to the client devices 102 via a computer network such as the Internet (not shown). Although Figure 1 specific components of the computing environment 100 are shown, in other embodiments, the computing environment 100 can also include additional and / or different components or arrangements. For example, in some embodiments, the computing environment 100 can also include additional network storage devices, network devices, hosts, and / or other suitable components (not shown) in other suitable configurations.
[0044] Each of the client devices 102 can include a computing device that facilitates access by the user 101 to cloud storage services provided by the distributed storage system 103 via a computer network. In one embodiment, each of the client devices 102 can include a desktop computer. In other embodiments, the client devices 102 can also include a laptop computer, a tablet computer, a smart phone, or other suitable computing devices. Although Figure 1 two users 101 and 101' are shown for illustrative purposes, in other embodiments, the computing environment 100 can facilitate access by other suitable numbers of users 101 to cloud storage or other types of computing services provided by the distributed storage system 103 in the computing environment 100.
[0045] As Figure 1As shown, the distributed storage system 103 may include one or more front-end servers 104, one or more back-end servers 105, a centralized repository 106, and one or more scenario servers 108 operatively coupled via a computer network 117. The back-end server 105 may be operatively coupled to a network storage device 112 configured to store user data. In the example shown, the user data stored at the network storage device 112 is shown as document 110. In other examples, the user data may also include image files, video files, or other suitable types of computer files. Although Figure 1 certain components of the distributed storage system 103 are shown, in other implementations, the distributed storage system 103 may also include additional and / or different components.
[0046] The computer network 117 may include any suitable type of network. For example, in one embodiment, the computer network 117 may include an Ethernet or Fast Ethernet network having routers, switches, load balancers, firewalls, and / or other suitable network components implementing the RDMA (RoCE) protocol over converged Ethernet. In other embodiments, the computer network 117 may also include an InfiniBand network having corresponding network components. In further embodiments, the computer network 117 may also include a combination of the above and / or other suitable types of computer networks.
[0047] The front-end server 104 may be configured to interact with the user 101 via the client device 102. For example, the front-end server 104 may be configured to provide a user interface (not shown) on the client device 102 to allow the corresponding user 101 to upload, modify, retrieve the document 110 stored in the network storage device 112, or perform other suitable operations on the document 110 stored in the network storage device 112. In another example, the front-end server 104 may also be configured to provide a user portal (not shown) that allows the user 101 to set retention, privacy, access, or other suitable types of policies to be implemented on the document 110 in the network storage device 112. In certain embodiments, the front-end server 104 may include a web server, a security server, and / or other suitable types of servers. In other embodiments, the front-end server 104 may be omitted from the distributed storage system 103 and instead provided by an external computing system, such as a cloud computing system (not shown).
[0048] The backend server 105 can be configured to provide various data storage services to facilitate storage, retrieval, modification, deletion, and / or other suitable types of data management operations on the documents 110 in the network storage device 112. For example, the backend server 105 can be configured to store the documents 110 and / or indexes of the documents 110 in shards. The backend server 105 can also be configured to provide facilities for performing fan-out joins of data from various shards in response to appropriate queries from, for example, the user 101. In a further example, the backend server 105 can also be configured to provide facilities for performing deduplication, data compression, access control, and / or other suitable operations.
[0049] The backend server 105 can also be configured to provide an ingestion processor 107 ( Figures 2A - 2D as shown in), which is configured to analyze new and / or updated data of the documents 110 received and stored at the network storage device 112. In some implementations, the ingestion processor 107 can be configured to receive registrations or other suitable types of indications from the scene server 108 and / or the corresponding scene computing service. A registration can individually indicate one or more attributes of the document 110 for which the corresponding scene server or scene computing service wants to receive notifications of changes. Such registrations can be maintained as a notification list 128 ( Figure 2D as shown in), a table, or other suitable data structures in the network storage device 112.
[0050] Although the backend server 105 may be suitable for performing storage and analysis of large amounts of data, the data structures used in the backend server 105 and / or the network storage device 112 may not be suitable for providing random access to individual data items (such as the documents 110). For example, the backend server 105 may not be able to access a document 110 in a list in constant time independent of the position of the document 110 in the list or the size of the list. Additionally, when a user requests a stored document, the backend server 105 can perform multiple operations (such as fan-out joins) to resolve indirect references. Performing multiple operations can result in high latency and computational load for servicing user requests.
[0051] The centralized repository 106 can be configured to provide a subset of the documents 110 cached in the data store 114 (sometimes referred to as "high impact data", as Figure 1As shown, a caching service is provided for "Data 111" to provide fast random access to the documents 110 stored in the network storage device 112. In some embodiments, the ingestion processor 107 running in the backend server 105 may be configured to push a subset of the stored data of the document 110 as Data 111 to the centralized repository 106 for storage in the data storage area 114.
[0052] The Data 111 pushed to the centralized repository 106 may include various categories of stored data that may be of interest to the various scenario servers 108 and / or scenario computing services. In one example, the Data 111 cached at the centralized repository 106 for the document 110 stored in the network storage device 112 may include the document name, document extension (e.g., "txt"), date / time of last update, number of views of the document, number of modifications to the document, uniform resource locator (URL) where the document can be retrieved, and / or other suitable information related to the document. By caching the Data 111, the user 101 can easily retrieve information about the document and / or other relevant information from the centralized repository 106.
[0053] The scenario server 108 may be configured to retrieve the cached Data 111 from the centralized repository 106 and react to changes in the underlying data of the document 110 stored in the network storage device 112 to provide a corresponding user experience. For example, a search index service may be interested in changes to the title, body, appendix, or other suitable types of content of the document 110 for updating a search index (not shown). On the other hand, a document metadata service may be interested in the number of views or viewer list of the document 110 and the timestamp of the last modification of the document 110 for updating the view / edit count of the document 110. In response to a change in the number of views, the document metadata service may be configured to send a message 118 to, for example, the user 101' to notify the user 101' that another user 101 has viewed and / or edited the document 110. Although the scenario server 108 is shown as a physical computing device in Figure 1 some implementations, the functionality of the scenario server 108 may be provided by a suitable server (e.g., the front-end server 104) executing appropriate instructions as the corresponding scenario computing service.
[0054] In some distributed storage systems, once any change is detected in a stored document 110, a change notification 116 is sent to all scenario servers 108 for further processing. Examples of detected changes can include write operations to the properties of the document 110. However, such blanket notifications can result in waste of computing resources in the distributed storage system 103. First, the transmission of the change notification 116 can involve resource-intensive notification calls (e.g., Hypertext Transfer Protocol or "HTTP" requests or Remote Procedure Calls).
[0055] Second, different scenario servers or scenario computing services may be interested in changes to different properties of the document 110. For example, for updating a search index, a search index service may not be interested in changes to the viewer list of the document 110 or the number of times the document 110 has been modified. On the other hand, a document metadata service may not be interested in changes to the body or title of the document 110. Thus, the changes indicated in the change notification 116 may be relevant to some scenario computing services but not to others. Although the changes indicated in the change notification 116 are not relevant, all scenario computing services may expend additional computing, network, input / output, and / or other suitable types of resources to determine whether the changes indicated in the change notification 116 are relevant to the respective scenario computing services. For example, various types of resources may be expended to perform additional read operations on the data 111 at the centralized repository 106. Such read operations may result in high computing loads on the server hosting the centralized repository 106, high network bandwidth in the computer network 117, and / or other negative impacts on the performance and / or operation of the distributed storage system 103.
[0056] Several embodiments of the disclosed technology can address at least some aspects of the above disadvantages by implementing granular change detection of document properties in the distributed storage system 103. In some implementations, a change detector 122 can be implemented in the centralized repository 106 to determine which one or more properties in the new version of the document 110 have changed compared to a previous version of the same document 110. Based on the determined one or more properties, the change notification 116 can be selectively sent to the servers 108 that were previously registered to receive notifications of changes to the one or more properties. Thus, expensive HTTP requests and / or Remote Procedure Calls for scenario computing services that are not interested in the detected property changes can be avoided. In addition, additional computing resources for determining whether the changes indicated in the change notification 116 are actually relevant to the scenario servers 108 can also be avoided. Thus, the computing load, network bandwidth, and / or consumption of other computing resources of the distributed storage system 103 can be reduced when compared to using blanket notifications, as described in more detail below with reference to Figures 2A - 2D More detailed description.
[0057] Figures 2A - 2D is a diagram illustrating performing granular change detection according to an embodiment of the disclosed technology. Figure 1 A schematic diagram of some hardware / software components of the distributed storage system 103. Figures 2A - 2D and other figures herein, omitted for clarity Figure 1 Certain components of the computing environment 100 and the distributed storage system 103. For example, Figures 2A - 2D In the example, for ease of explanation, the front-end server 104 and the computer network 117 are omitted from the distributed storage system 103. Figures 2A - 2D In other figures herein, various software components, objects, classes, modules, and routines may be computer programs, procedures, or processes written as source code in C, C++, C#, Java, and / or other suitable programming languages. A component may include, but is not limited to, one or more modules, objects, classes, routines, properties, processes, threads, executable files, libraries, or other components. A component may be in source code or binary form. A component may include source code aspects before compilation (e.g., classes, properties, procedures, routines), compiled binary units (e.g., libraries, executable files), or artifacts (e.g., objects, processes, threads) that are instantiated and used at runtime.
[0058] Components within a system can take different forms within the system. As an example, a system comprising a first component, a second component, and a third component can include, but is not limited to, a system comprising the first component as an attribute in source code, the second component as a binary compiled library, and the third component as a thread created at runtime. A computer program, procedure, or process can be compiled into object, intermediate, or machine code and presented for execution by one or more processors of a personal computer, network server, laptop, smartphone, and / or other suitable computing device.
[0059] Likewise, a component may comprise hardware circuitry. One of ordinary skill in the art will recognize that hardware may be considered fossilized software, while software may be considered liquefied hardware. As just one example, the software instructions in a component may be burned into a programmable logic array circuit, or may be designed as a hardware circuit with an appropriate integrated circuit. Likewise, hardware may be emulated by software. Various implementations of source code, intermediate code, and / or object code and associated data may be stored in a computer memory, including read-only memory, random access memory, magnetic disk storage media, optical storage media, flash memory devices, and / or other suitable computer-readable storage media (excluding propagating signals).
[0060] like Figure 2AAs shown, the backend server 105 may include an ingestion processor 107 for processing a new version of a document 110 received from a client device 102 of a user 101 via a front-end server 104 ( Figure 1 ). The document 110 may include a plurality of attributes 120. In the illustrated embodiment, the ingestion processor 107 may include an interface component 130, a signature component 132, and a notification component 134 that are operatively coupled to each other. In other embodiments, the ingestion processor 107 may further include network, security, database, or other suitable types of components.
[0061] The interface component 130 may be configured to receive a new version of the document 110 and provide the received document 110 to the signature component 132 for further processing. The interface component 130 may also be configured to store, retrieve, or perform other suitable file operations on the document 110 in the network storage device 112. For example, in the illustrated embodiment, the interface component 130 may be configured to store a copy of the received new version of the document 110 as a file separate from the previous version in the network storage device 112, or replace the previous version of the document 110.
[0062] The signature component 132 may be configured to generate a signature 121 corresponding to one or more attributes 120. In certain embodiments, the signature 121 of the attribute 120 may include a hash value of the attribute 120. For example, the signature 121 of the content body of the document 110 may include a randomly generated alphanumeric string of a preset length. In other embodiments, the signature of the attribute 120 may include the value of the attribute 120. For example, the value of the view count attribute of the document 110 may include the integer "4". In a further example, the signature 121 may include other suitable values derived from and / or corresponding to various attributes 120 of the document 110.
[0063] As Figure 2A shown, when the signature 121 is generated, the signature component 132 may instruct the interface component 130 to transmit the data 111 extracted from the new version of the document 110 together with the signature 121 to a centralized repository 106 for granularity change detection. Although embodiments of granularity change detection are described herein as being implemented in the centralized repository 106, in other embodiments, such granularity change detection may also be implemented on one or more of the backend servers 105 of the distributed storage system 103 or other suitable servers (not shown).
[0064] As Figure 2AAs shown, the centralized repository 106 may include a change detector 122 and a change indicator 124 that are operably coupled to each other. The change detector 122 may be configured to determine which one or more properties have changed. In some embodiments, the change detector 122 may be configured to compare the respective signatures 121 of a new version of the document 110 with the respective signatures 121 of a previous version stored in the data store 114. In response to determining that the signatures 121 of the new version are the same or substantially the same (e.g., more than 90% similar), the change detector 122 may be configured to indicate that the corresponding property 120 of the document 110 has not changed. Otherwise, the change detector 122 may be configured to indicate that the corresponding property 120 has changed. As Figure 2B shown, when iterating through all of the signatures 121( Figure 2A ), the change indicator 124 may be configured to send a list of property IDs 123 identifying the list of properties 120 that have been changed to the ingestion processor 107 at the backend server 105.
[0065] In response to receiving the property IDs 123, the notification component 134 of the ingestion processor 107 may be configured to retrieve a notification list 128 from the network storage device 112. The notification list 128 may include data indicating which one or more scene servers 108 (and / or corresponding scene computing services) are interested in changes to a particular property 120 of the document 110. The notification list 128 may be generated by receiving registration data (not shown) from the scene servers 108, by assigning default values, or by other suitable techniques. Based on the retrieved notification list 128, the notification component 134 may be configured to determine one or more scene servers 108 for receiving change notifications 125 based on the property IDs 123. In the example shown, it may be indicated in the notification list 128 that the first scene server 108a is interested in changes to the property 120 identified by one of the property IDs 123. Thus, the change notification 116 is sent only to the first scene server 108a but not to the second scene server 108b and the third scene server 108c.
[0066] As Figure 2C shown, upon receiving the change notification 116, the first scene server 108a may be configured to provide a message 118 to the user 101' to, for example, notify the user 101' of the detected change to the property 120 of the document 110 or to provide some other suitable user experience. As Figure 2DAs shown, when the user 101 provides another version of the document 110' with different changes to one or more of the properties 120', the notification component 134 can be configured to, when a change to a different property 120' is detected (as indicated by the property ID 123'), provide a change notification 116' to the third scene server 108c, rather than providing the change notification 116' to the first scene server 108a or the second scene server 108b. Thus, resource-intensive notification calls to one or more scene servers 108 and additional read operations by such one or more scene servers 108 can be avoided.
[0067] Figures 3A - 3C is a diagram illustrating the process during signature generation according to an embodiment of the disclosed technology. Figures 2A - 2D Schematic diagram of some hardware / software components of the signature component 132. Figure 3A As shown, signature component 132 may include a function selector 142 and a signature generator 146 operatively coupled to each other. Function selector 142 may be configured to select a function for generating signature 121 based on the value of attribute 120. As described in more detail below, function selector 142 may be configured to select a function in various ways. Signature generator 146 may be configured to apply the selected function (e.g., a hash function) to the value of attribute 120 to generate signature 121.
[0068] Aspects of the disclosed technology involve selecting appropriate functions associated with various attributes 120 of a document 110 in order to send the backend server 104 ( Figure 1 ) is maintained below a threshold or even minimized. In some embodiments, such as Figure 3A As shown, multiple different hash functions can be statically assigned and associated with corresponding attributes 120 of document 110, as indicated by function list 137 stored in network storage device 112. For example, an identity function that returns an input value as output can be used for the number of times a document has been viewed or modified, as the cost of computing hash values for these values may outweigh the savings in computing resources used to compare these values. In another example, an xxHash function can be associated with the body of document 110, as the body may include a large amount of data.
[0069] In further implementations, the association of the hash function with the attribute 120 may be dynamic, for example based on the data type and / or size of the value of the attribute 120 or other suitable function criteria 138, such as Figure 3BAs shown. For example, the function selector 142 can be configured to determine the data type and / or size of the value of the attribute 120, and assign a value that is an integer or a short string (e.g., less than ten or other appropriate number of characters) to the identity function of the attribute 120 according to the function criteria 138. Thus, the number of views and the number of modifications to the document can be associated with the identity function. In other examples, the function selector 142 can be configured to assign the Fowler-Noll-Vo (FNV) hash function to the attribute 120 with a data size of four to twenty bytes according to the function criteria 138. In a further example, the function selector 142 can also be configured to assign the xxHash64 function to the attribute 120 with a large data size (e.g., greater than 20 bytes) according to the function criteria 138. In additional examples, the function selector 142 can be configured to assign the Secure Hash Algorithm 256 (SHA256) function to the attribute 120 with values that may generate hash value conflicts and when such conflicts cannot be tolerated.
[0070] In a further implementation, as Figure 3C shown, the function selector 142 can be configured to use the cost function 139 to present the dynamic selection of the hash function as an optimization operation. Different hash functions can have different storage footprints (e.g., output hash value size), computational resource costs (e.g., the cost of computing the hash value), conflict rates, and / or other suitable characteristics. Thus, the optimization operation can be defined as: for the attribute 120, select the hash function that minimizes the processor cycles and storage overhead, while keeping the number of conflicts below a threshold and minimizing the total amount of attribute data that must be read for the entire comparison process. An example cost function J( Hi ) can be expressed as follows:
[0071] J( Hi ) = W Collision *CR Hi + W CPU *CPU Hi + W Storage *Storage Hi + W Dat a*Data Hi ,
[0072] where CR Hi is the conflict rate, CPU Hi is the processor cycles, Storage Hi is the storage size of the hash value, Data Hi ∈ {PropertySize, Storagem} and W is the weight associated with each corresponding parameter.
[0073] While the distributed storage system 103 is operable, different weights associated with the above parameters may change. For example, the available system resources may change, making computing resources abundant. As such, the weight for computing resources can be adjusted to obtain a different hash function for the same attribute 120 than before the change in the available system resources.
[0074] Figure 4 is a schematic diagram of certain hardware / software components of a signature component that implements attribute grouping according to an embodiment of the disclosed technology. As Figures 2A - 2D shown, the signature component 132 may further include a grouping component 140 that is configured to group some of the attributes 120 of the document 110 into a set, generate a signature 121 for the set, and allow a change detector 122 ( Figure 4 ) at a centralized repository 106 ( Figure 2A ) to efficiently determine which one or more of the attributes 120 have changed. Figure 2A
[0075] In some embodiments, the grouping component 140 may be configured to group related attributes 120 and generate a signature 121 for a set of attributes 120 rather than for each attribute 120 in the set. Without being bound by theory, it is recognized that some of the attributes 120 of the document 110 often change together. For example, the number of views often changes with the number of modifications to the document 110. Thus, combining the attributes 120 into a group or set and generating a hash value for the set can allow for a quick determination of whether any of the attributes 120 in the set have changed. For example, if no change to the set is detected, the change detector 122 can skip a full comparison of each attribute 120 in the set. If a change to the set is detected, then each attribute 120 in the set can be compared to see which one or more of the attributes 120 have changed based on the function list 137, function criteria 138, or cost function 139 (as described above with respect to Figures 3A - 3C ).
[0076] According to additional aspects of the disclosed technology, grouping the attributes 120 can also utilize optimization operations. If the changed correlations between different attributes 120 are used to group the attributes 120, a "grouping correlation threshold" can be set such that the amount of data to be read and hash comparisons for the attributes 120 are kept below a size threshold and / or minimized. In some embodiments, if the values of the attributes 120 have been hashed, a combination operation (e.g., left shift and perform bitwise exclusive OR) can be used to generate a group hash value as the signature 121 of the group. However, if the attributes 120 have already been passed through an identity function, an additional hash operation can be performed on the combination of the attribute values to generate a group hash value. In other embodiments, multiple groupings of the attributes 120 can be used, where one attribute 120 can be part of multiple groups or sets.
[0077] A technique for generating attribute groupings can include maintaining a covariance matrix (not shown) of the attributes 120. The calculation of the covariance matrix can be simplified such that if an attribute 120 has changed, the value of the attribute 120 is 1. If the attribute 120 has not changed, the value of the attribute 120 is 0. If an attribute 120 is sufficiently correlated with some (or all) of the attributes 120 in an attribute group P Cj in a cluster (i.e., the value in the covariance matrix corresponding to the attribute pair <Pi, P Cj > is higher than a threshold Tc), grouping can be constructed using clustering heuristics by assigning the attribute 120 Pi to an existing cluster (attribute group). If the attribute 120 is not sufficiently correlated, i.e., no value of the attribute Pi in the covariance matrix is higher than the threshold Tc, a new cluster (i.e., a new attribute group) containing only that attribute can be created. Another grouping heuristic can include determining the "closest" cluster based on a cluster similarity function and assigning the attribute 120 to the closest cluster (i.e., the attribute group). Regardless of which heuristic is used to create the attribute groupings, the amount of data to be read and compared to detect which attributes 120 have changed can be kept below a threshold and / or minimized.
[0078] In certain implementations, the attribute groupings can be calculated before the distributed storage system 103 becomes operational, but they can also be calculated periodically in an offline batch or other suitable processing manner. When the attribute groupings are determined, a bit field with one bit corresponding to each attribute 120 can be used to record which attribute 120 in the group has changed. For example, if no change has occurred in the corresponding attribute 120, the bit can be set to zero. When the corresponding attribute has changed, the bit of the system can be set to 1. In further implementations, the attribute groupings can be changed in an online manner to better adapt to temporal changes in the data.
[0079] In a further implementation, a single attribute 120 can be included in multiple different attribute groups. Such inclusion can allow changes to be detected at any time. For example, as Figure 5 shown, the first group 150a can include four attributes 120 represented as circles, and the second group 150b can also include four attributes 120'. The first and second groups 150a and 150b share a common attribute 120". Thus, when a change in the first group 150a is detected but no change in the second group 150b is detected, it can be determined that the shared attribute 120" in both the first group 150a and the second group 150b has not changed.
[0080] Figures 6A - 6C is a flowchart showing a fine-grained change detection process in a distributed storage system according to an embodiment of the disclosed technology. Although embodiments of the process are described below in the context of the Figure 1 computing environment 100, in other embodiments, the process can also be implemented in other computing environments with additional and / or different components.
[0081] As Figure 6A shown, the process 200 can include receiving a new version of a document at stage 202. In some embodiments, the new version of the document can have a corresponding previous version stored in the distributed storage system 103 ( Figure 1 ). In other embodiments, there may be no previous version of the document in the distributed storage system 103. The process 200 can then include comparing the individual attributes of the received new version of the document at stage 204. Various techniques for comparing attributes can be applied. Some example techniques are described above with reference to Figures 2A - 2D and below with reference to Figure 6B . The process 200 can also include identifying at stage 206 the attributes whose values have changed in the received new version of the document. When an attribute whose value has changed is identified, the process 200 can continue to send a notification to a scenario server or scenario computing service that was previously registered to receive notifications about changes to the attribute.
[0082] Figure 6B illustrates an example operation of comparing the individual attributes of a document. As Figure 6B shown, the operation can include comparing the signatures (e.g., hash values) of the attribute values 210, including identifying the values of the individual attributes at stage 212. The operation can optionally include selecting a function for each attribute to generate a signature at stage 214. Various example techniques for selecting functions are described above with reference to Figures 3A - 3C . Then, the operation can include applying the selected function to the value of the corresponding attribute to generate a signature at stage 216. Based on the generated signature, the operation can include detecting a change in one or more of the attributes at stage 218.
[0083] Figure 6C Illustrated is an example operation for detecting a change in an attribute by grouping attributes into sets. As Figure 6C shown, the operation may include grouping attributes into sets at stage 222. Example techniques for grouping attributes were described above with reference to Figure 4 . The operation may then include generating a group signature for each set at stage 224. For each set, the operation may include a decision stage 224 for determining whether a change in the set has been detected based on the group signature. In response to determining that no change in the set has been detected, the operation may include marking all attributes in the set as unchanged. In response to determining that a change in the set has been detected, the operation may include detecting a change in one or more attributes in the set at stage 228 by performing operations such as those described above with reference to Figure 6B .
[0084] Figure 7 is a computing device 300 suitable for certain components of the computing environment 100 in Figure 1 . For example, the computing device 300 may be suitable for the client device 102, the front-end server 104, the back-end server 105, the centralized repository 106, and the scenario server 108 in Figure 1 . In a very basic configuration 302, the computing device 300 may include one or more processors 304 and a system memory 306. A memory bus 308 may be used for communication between the processor 304 and the system memory 306.
[0085] Depending on the desired configuration, the processor 304 can be of any type, including but not limited to a microprocessor (μR), a microcontroller (μC), a digital signal processor (DSP), or any combination thereof. The processor 304 may include one or more levels of cache, such as a level 1 cache 310 and a level 2 cache 312, a processor core 314, and registers 316. An example processor core 314 may include an arithmetic logic unit (ALU), a floating-point unit (FPU), a digital signal processing core (DSP Core), or any combination thereof. An example memory controller 318 may also be used with the processor 304, or in some implementations, the memory controller 318 may be an internal part of the processor 304.
[0086] Depending on the desired configuration, the system memory 306 can be of any type, including but not limited to volatile memory (such as RAM), non-volatile memory (such as ROM, flash memory, etc.), or any combination thereof. The system memory 306 may include an operating system 320, one or more applications 322, and program data 324. In Figure 7 , this basic configuration 302 is illustrated by those components within the inner dashed line.
[0087] The computing device 300 may have additional features or functionality and additional interfaces to facilitate communication between the basic configuration 302 and any other devices and interfaces. For example, the bus / interface controller 330 may be used to facilitate communication between the basic configuration 302 and one or more data storage devices 332 via the storage interface bus 334. The data storage devices 332 may be removable storage devices 336, non-removable storage devices 338, or a combination thereof. Examples of removable and non-removable storage devices include: disk devices (such as floppy disk drives and hard disk drives (HDDs)), optical disc drives (such as compact disc (CD) drives or digital versatile disc (DVD) drives), solid state drives (SSDs), and tape drives, among others. Example computer storage media may include volatile and non-volatile, removable and non-removable media implemented in any method or technology for storing information (such as computer-readable instructions, data structures, program modules, or other data). The terms "computer-readable storage medium" or "computer-readable storage device" do not include propagated signals and communication media.
[0088] System memory 306, removable storage device 336, and non-removable storage device 338 are examples of computer-readable storage media. Computer-readable storage media includes, but is not limited to, RAM, ROM, EEPROM, flash memory or other storage technology, CD-ROM, digital versatile disk (DVD) or other optical storage, magnetic cassettes, tapes, disk storage or other magnetic storage devices, or any other medium that can be used to store the desired information and can be accessed by the computing device 300. Any such computer-readable storage media can be part of the computing device 300. The terms "computer-readable storage medium" do not include propagated signals and communication media.
[0089] The computing device 300 may also include an interface bus 340 for facilitating communication from various interface devices (such as output device 342, peripheral interface 344, and communication device 346) to the basic configuration 302 via the bus / interface controller 330. Example output devices 342 include a graphics processing unit 348 and an audio processing unit 350, which may be configured to communicate with various external devices (such as a display or speakers) via one or more A / V ports 352. Example peripheral interfaces 344 include a serial interface controller 354 or a parallel interface controller 356, which may be configured to communicate with external devices such as input devices (such as a keyboard, mouse, pen, voice input device, touch input device, etc.) or other peripheral devices (such as a printer, scanner, etc.) via one or more I / O ports 358. Example communication devices 346 include a network controller 360, which may be arranged to facilitate communication with one or more other computing devices 362 via one or more communication ports 364 on a network communication link.
[0090] A network communication link can be an example of a communication medium. A communication medium can generally be embodied by computer-readable instructions, data structures, program modules, or other data in a modulated data signal, such as a carrier wave or other transmission mechanism, and can include any information delivery medium. A "modulated data signal" can be a signal whose one or more characteristics are set or changed (in a manner that encodes information in the signal). By way of example and not limitation, a communication medium can include wired media such as a wired network or a direct wired connection, and wireless media such as acoustic, radio frequency (RF), microwave, infrared (IR), and other wireless media. The term computer-readable medium as used herein can include storage media and communication media.
[0091] Computing device 300 can be implemented as part of a small form factor portable (or mobile) electronic device, such as a cellular phone, a personal digital assistant (PDA), a personal media player device, a wireless web watch device, a personal headset device, a specific application device, or a hybrid device incorporating any of the above functions. Computing device 300 can also be implemented as a personal computer including both a laptop computer configuration and a non-laptop computer configuration.
[0092] From the foregoing, it will be appreciated that specific embodiments of the present disclosure have been described herein for purposes of illustration, but various modifications can be made without departing from the present disclosure. Additionally, many elements of one embodiment can be combined with other embodiments, either to add to or replace elements of other embodiments. Accordingly, the technology is not limited by the appended claims.
Claims
1. A method for property grouping for change detection of documents stored in a distributed storage system, the distributed storage system having a plurality of servers interconnected with each other via a computer network, the method comprising: Receiving, on one of the servers, data representative of a new version of a document stored in the distributed storage system, the received new version of the document having a plurality of properties, the plurality of properties each having a value that describes or identifies the document; And In response to receiving the data representative of the new version of the document, grouping the plurality of properties into a plurality of groups, the plurality of groups each including a subset of the plurality of properties; And For each of the plurality of groups, Generating a hash value for the group based on the values of the subset of properties in the group; Determining whether the generated hash value of the group is different from the hash value of the corresponding group in a previous version of the document in the distributed storage system; And In response to determining that the generated hash value of the group is not different from the hash value of the corresponding group in the previous version, inserting metadata indicating that none of the properties in the subset of properties in the group have changed into the new version of the document.
2. The method according to claim 1, further comprising: In response to determining that the generated hash value of the group is different from the hash value of the corresponding group in the previous version, Determining whether each property in the subset of properties in the group has changed; And In response to determining that one property in the subset of properties has changed, sending a notification via the computer network to one or more computing services, without sending the notification to other computing services not registered to receive the notification, the one or more computing services previously being registered to receive notifications regarding changes to the one property in the subset of properties.
3. The method according to claim 1, wherein, Grouping the plurality of properties includes: grouping the plurality of properties into a plurality of groups, the plurality of groups each including a subset of the plurality of properties, the plurality of groups not sharing any common properties.
4. The method according to claim 1, wherein Grouping the plurality of properties includes: grouping the plurality of properties into a plurality of groups, the plurality of groups each including a subset of the plurality of properties, at least two of the plurality of groups sharing one or more properties.
5. The method according to claim 1, wherein, Grouping the plurality of properties includes: Identifying a change correlation between one property in the properties and another property in the properties, the change correlation including data indicating the likelihood that when the other property in the properties is changed, the one property in the properties is changed, and vice versa; Determining whether the change correlation exceeds a grouping correlation threshold; and In response to determining that the change correlation exceeds the grouping correlation threshold, grouping the one property in the properties and the other property in the properties into a single group.
6. The method according to claim 1, wherein: Each property in the subset of properties includes a corresponding hash value; And Generating the hash value for the group includes applying a combination operation to the hash values of the subset of properties to derive the hash value for the group.
7. The method according to claim 1, wherein: At least one attribute in the attribute subset does not include the corresponding hash value; And Generating the hash value of the group includes: Combining the values of the attribute subsets in the group into a group value; And Applying a hash function to the group value to derive the hash value of the group.
8. The method according to claim 1, wherein, Grouping the plurality of attributes includes: Calculating a covariance matrix of the plurality of attributes; and Based on the covariance matrix, when the covariance of one of the attributes with at least one attribute of one of the groups in the group is higher than a threshold, assigning the one of the attributes to one of the groups in the group.
9. The method according to claim 1, wherein Grouping the plurality of attributes includes: Calculating a covariance matrix of the plurality of attributes; and Based on the covariance matrix, when the covariance of one of the attributes with at least one attribute of one of the groups in the group is not higher than a threshold, assigning the one attribute of the attributes to a new group.
10. A computing device for attribute grouping for change detection of data stored in a distributed storage system, the distributed storage system having a plurality of servers interconnected with each other via a computer network, the computing device comprising: A processor; And A memory having instructions executable by the processor to cause the computing device to perform the method according to one of claims 1-9.
Citation Information
Patent Citations
Distributed data system with document management and access control
US20150127607A1
File metadata verification in a distributed file system
US20190026308A1