Document Tracking for Graphs Linked by Version Hashes
By constructing version hash link graphics (VHLG) to connect document versions, the lack of deep context and flexible collaborative analysis in the prior art is solved, and a deep understanding of the collaborative network and the provision of productivity tools are achieved.
Patent Information
- Application Number
- CN202080067907.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2019-10-10
- Filing Date
- 2020-10-12
- Publication Date
- 2025-06-03
- Estimated Expiration
- 2040-10-12
AI Technical Summary
The prior art lacks the ability to provide a deeper context in order to understand the user/role/industry performing the submission when using hash-based links in collaborative scenarios, and the ability to have more than just find submission contexts for hashing.
The graph is constructed by using a unique identifier of the document hash and/or its contents, thereby connecting the document version to form a version hash link graph (VHLG). This graph is used for data analysis processing, infer industry and user type collaborative networks, and provides applications to detect expired version document access, document update notifications, and document version coordination.
Achieving an in-depth understanding of the collaborative network provides productivity enhancement tools that solve users’ challenges of collaborating on shared documents without being limited to specific applications or document repositories.
Smart Images

Figure CN114556317B_ABST
Abstract
Description
[0001] Cross - Reference to Related Applications
[0002] This application relates to the following co - pending and co - assigned patent applications, which are hereby incorporated by reference herein:
[0003] U.S. Patent Application No. 62 / 913,380, filed on October 10, 2019, inventors Robert Maguire and Ravinder Krishnaswamy, titled "Document tracking through Version Hash Linked Graphs", Attorney Docket No. 30566.0585USP1. Background of the Invention 1. Field of the Technology
[0005] The present invention generally relates to tracking document versions and, in particular, to methods, systems, devices, and articles of manufacture for using document hashes to link document versions, where the document hashes are used to construct graphs for further analysis and exploitation.
[0006] 2. Description of the Related Art
[0007] In today's cloud - connected and multi - device world, collaboration by sharing documents and document content is the norm. Such collaboration requires coordinated access to document and document content versions. In addition, analysis of collaborators and the collaboration itself in an efficient and comprehensive manner (e.g., how collaborators interact / save / edit / view files / documents in a collaborative environment, etc.) can provide insights into application product feature development. In this regard, collaboration has made the following increasingly important for businesses that provide document viewing and editing applications:
[0008] 1. Understanding how different segments of users and industries interact to help businesses prioritize investments in product features; and
[0009] 2. Providing productivity - enhancing tools to address the challenges of their users collaborating on shared documents, such as timely notification of document edits and document version coordination.
[0010] The use of hash - based linking in collaborative scenarios as it exists in the prior art is described below. These provide the background for the key new concepts in the present invention and the problem space they address.
[0011] Use of Chained Hashes
[0012] The idea of chaining data versions via hashes is not new. GIT's underlying mechanism uses commit hashes and chains the commit hashes together to generate a DAG (Directed Acyclic Graph) of commits in order to efficiently find commit context using only the hashes. FIG. 1 shows the DAG of hashed commits used in GIT. As shown, each time a change is committed to the code (referred to as a "commit"), a snapshot is taken and the set of changes between two snapshots can be applied or rolled back. Each snapshot is named with a commit / hash ID 102 derived from the content of the snapshot (e.g., the actual content and some metadata such as the presentation time, author information, parent, etc.). The change flow in GIT is an ordered list of change sets as the change flow is applied one after another to move from one snapshot / commit 102 to the next. A pointer to a specific snapshot is called a "branch" and the "head" 104 points to the location of the last branch (e.g., the main branch 106) checked out from the workspace. Different features can be defined in other branches 108 that can eventually be merged into the main branch 106. Other branches 106 to 108 can be created on any valid snapshot version 110 to 112. Additionally, as shown, a commit 102 can have multiple parents 114 and each parent can have multiple children 116.
[0013] However, what is missing from such chained hashes is the ability to provide deeper context for a general understanding of the user / role / industry that performed the commit and the ability to do more with the hashes than just find commit context.
[0014] In the world of cryptocurrency and distributed ledgers, different forms of hash chaining are the basis of the algorithms. Figure 2 shows hash chaining as used in cryptocurrency. A blockchain consists of various blocks 202 that are linked to each other 204. A hash tree (also known as a Merkle tree) encodes the blockchain 200 data. Each transaction 206 that occurs on the blockchain 200 network has an associated hash 208, which is stored in a tree-like structure such that each hash 208 is linked to its parent, thus following a parent-child tree relationship. All the transaction hashes 208 in a block 202 are also hashed, resulting in a Merkle root 210. The Merkle root 210 contains all the information about each individual transaction hash 208 that exists on the corresponding block 202. Each block may also have a nonce 212 ("number used once"). The nonce 212 is a number added to the hashed block 202 that, when rehashed, satisfies the difficulty level limit and is the number that the blockchain miner is solving for. Each block 202 also contains a timestamp 214 (e.g., for the current time), which is used to establish the validity of the block (e.g., if the timestamp is greater than the median timestamp of the previous 11 blocks and less than the network-adjusted time + 2 hours, then the timestamp may be accepted as valid). Additionally, the previous hash 216 of each block 202 is the hash of the previous block 202 (e.g., an identifier / pointer to the hash of the previous block).
[0015] However, while hash functions are the basis of blockchain algorithms in securely representing the linkage of data blocks, hash functions do not provide the advantages of the structure of the present invention (as described below).
[0016] Data Management Solutions
[0017] In addition to using chained hashes as described above, prior art data management solutions can also be used to enable collaboration. However, prior art data management solutions (e.g., product data management - PDM applications) require users to make their data work within the constraints of the application for version management and notifications.
[0018] In view of the above, what is needed is a solution that can achieve the flexibility of collaboration and the analysis of such collaboration, which is effective and not limited to a specific application or document repository. Summary of the Invention
[0019] Embodiments of the present invention provide a data structure for constructing a graph to connect document versions by using document hashes and / or unique identifiers of their content. Such a data structure and graph can be used for:
[0020] · Data analysis and processing to infer a collaboration network of industries and user types through their access to common data. This addresses the ability to understand how different segments of users and industries interact with each other (e.g., helping enterprises prioritize investment in product features).
[0021] · An application to detect access to expired version documents, document update notifications, and document version coordination, and to summarize document usage statistics in an organization or project. This addresses the ability to provide productivity-enhancing tools that address the challenges of users collaborating on shared documents. Brief Description of the Drawings
[0022] Now refer to the drawings, where like reference numerals throughout indicate corresponding parts:
[0023] FIG. 1 shows a DAG of hash commits used in prior art GIT;
[0024] FIG. 2 shows a hash chaining used in prior art cryptocurrencies;
[0025] Figure 3 Shows an exemplary version hash link graph according to one or more embodiments of the present invention;
[0026] Figure 4A Shows an exemplary accessible lineage according to one or more embodiments of the present invention;
[0027] Figure 4B Shows the ability to link users based on the version hash link graph of FIG. 4 according to one or more embodiments of the present invention;
[0028] Figure 5 Shows an example of how a business that collaborates through its access to a common lineage can be inferred using a VHLG according to one or more embodiments of the present invention;
[0029] Figure 6 Shows a version hash link graph / tree, where data accessed from various different locations / devices accesses version information using the same signature / hash / GUID;
[0030] Figure 7 Shows a logical flow for tracking document version control according to one or more embodiments of the present invention;
[0031] Figure 8 is an exemplary hardware and software environment for implementing one or more embodiments of the present invention; and
[0032] Figure 9 Schematically shows a typical distributed / cloud-based computer system according to one or more embodiments of the present invention. Detailed implementation mode
[0033] In the following description, reference is made to the accompanying drawings, which form a part of the present invention and illustrate several embodiments of the present invention by way of illustration. It should be understood that other embodiments may be utilized and structural changes may be made without departing from the scope of the present invention.
[0034] Overview
[0035] A document based on a file (such as a DWG file) has a unique signature associated with it based on applying a hashing algorithm to its content or creating a new GUID each time the document is updated. The hashes of successive document versions linked to additional operations and usage context information are used to generate a dependency graph that can connect users to the company. The dependency graph can also be analyzed to provide useful information about the community, collaborators, and features used (and potential areas for expansion).
[0036] Version Hash Link Graph
[0037] Embodiments of the present invention chain together the successive version information (i.e., the versions created and then saved before and after an operation, such as when editing) of a (document). This chain can be used to generate / obtain a complete history of the document (e.g., without actually linking or storing the document itself). As used herein, this chain is referred to as a Version Hash Link Graph (VHLG). Figure 3 Shows an exemplary Version Hash Link Graph according to one or more embodiments of the present invention. The Version Hash Link Graph 300 includes nodes 302 to 304 (such as user nodes 302A and 302B, collectively referred to as user nodes 302) corresponding to users or applications accessing document nodes 304 (i.e., document nodes 304A and 304B, collectively referred to as document nodes 304), edges 306 indicating operations and containing additional meta-information such as access time or product version, and edges 308 representing version relationships.
[0038] As Figure 3 shown, the document nodes 304 are connected by arrows 308 representing a "version" (or "next-generation version") relationship (while the edges 306 are operations to generate or access versions). The end-user nodes 302 are connected to document versions based on reference / editing operations (e.g., user 302 opens or saves a document version). To construct the graph 300 from an application, a small amount of additional information (the hash or id associated with the operation being recorded) needs to be sent to a service or backend. This will allow the backend to assemble the graph based on data obtained from multiple users. In embodiments of the present invention, the information includes the hashes before and after the operation. For example, the following format can be used to record the information required to construct the graph:
[0039] Log Entry: (Anonymous - User - id, Platform, File - Operation, Hash - Before, Hash - After, Time)
[0040] The following represents log entries that can be logged / sent to Figure 3 the service / backend:
[0041] (u88, "Desktop - win", "Open", "EF8A09D", "EF8A09D", 9310028)
[0042] (u88, "Desktop - win", "Save", "EF8A09D", "D9A22B", 9320031)
[0043] (u89, "Mobile - ios", "Open", "D9A22B", "D9A22B", 10311299)
[0044] These log entries indicate that user u88 has opened document EF8A09D and then saved document D9A22B (e.g., document EF8A09D was modified and saved with the new hash as D9A22B). Subsequently, user u89 opened document D9A22B. The above log entries identify the user (u88 or u89), the platform (Desktop Windows or Mobile ios), the operation (Open or Save), the hashes before and after the operation (i.e., EF8A09D and D9A22B), and the time when the operation was performed (9310028, 9320031, or 10311299). By comparing / linking different hashes (e.g., both users u88 and u89 have the hash D9A22B in the log entries associated with each user), different users on different platforms sharing the same data can be linked together.
[0045] As used herein, any type of hashing algorithm can be used to generate the hash as long as the algorithm produces a unique identifier (e.g., a hash or GUID) representing the document. Additionally, the log entry is information that is stored / provided to the service / backend and is independent of the data itself. In this regard, only the log entry / hash information is stored and can be used to reconstruct / generate the history of the document.
[0046] Figure 3 Exemplary graphics (and other VHLG) will scale across time and multiple users. By accessing documents that are part of the "pedigree" of document versions (related or derived versions), there is a natural concept of collaboration. Figure 4AShows an accessible exemplary lineage according to one or more embodiments of the present invention; as shown, desktop user Bob 402 has opened file 404 and then saved / exported the file, thereby creating version 406. Desktop user Scott 408 has opened file 410 and then saved and exported file 410, thereby generating a new version 412 (as shown, versions 404, 406, 410, and 412 are all hashes in the VHLG). Bob 402 has also referenced version 412. Web users Joe 414 and John 416 have both opened version 412. Desktop user Mary 418 has opened version 406 and saved / exported it, thereby creating version 420 which has also been opened by mobile device user Yan 422. Using the information in the usage logs (e.g., via a graph), the lineages 424 and 426 of the document versions can be easily determined. Lineage 424 consists of different document versions 404, 406, and 420, while lineage 426 consists of document versions 410 and 412. The hashes and operation information (e.g., open, save / export, reference) (as well as additional information) can be used to determine the lineage. Figure 4B Shows the ability to link users 402, 408, 414, 416, 418, and 422 based on the VHLG 400 of FIG. 4 according to one or more embodiments of the present invention.
[0047] Exemplary lineage / history determination can be shown by the following example: A PDF (Portable Document Format) file can be published from a CAD (Computer-Aided Design) drawing. Hashes are created for the CAD drawing (hash A) and the PDF (hash B), and the edges of the graph will connect the two. If a user opens the PDF without knowing the source of the PDF, then the hash of the PDF can be recalculated / obtained again (generating hash B). Then the calculated hash (hash B) can be used to query the graph and retrieve the edge / determine that hash B originated from hash A. In this regard, using the previous hash value and the subsequent hash value, it is possible to determine the history / lineage of the document version without accessing / using the application that created the document (e.g., a CAD application) and the document does not need to reside / exist in the same repository. In other words, only by querying the VHLG can the document history be determined.
[0048] Analysis and Insights on Large-Scale Use of Version Hash Link Graph
[0049] By massively constructing a graph and augmenting the graph with data (including more user context such as industry type, company type, etc., and edges with access times, product versions, etc.), it is possible to infer and obtain insights into the community that collaborates and understands the access patterns over time.
[0050] Figure 5Shows an example of how a VHLG can be used to infer businesses that collaborate through their access to a common lineage, according to one or more embodiments of the present invention. More specifically, Figure 5 Shows an example of industry type collaboration that can be inferred from a VHLG. Each tree icon node 502 represents a lineage (a collection of documents related through their version chains). Each company / industry type 504A through 504L (collectively referred to as industry types 504) is represented by an icon. Examples of different types can include business 504A, industrial machinery 504B, unknown 504C, consumer products 504D, education 504E, mining 504F, construction 504G, construction services 504H, engineering service providers 504I, architectural services 504J, building products and fabrication 504K, civil infrastructure 504L, and so on. The thickness of the edges 506 (i.e., edges 506A through 506C collectively referred to as edges 506) is proportional to the number of times a company has accessed a lineage. For example, compared to the access represented by edge 504B, edge 504A represents less access to lineage 502 (e.g., through the industry type connected to edge 504A), and edge 504B has less access than the access represented by edge 506C.
[0051] Figure 5 The example shown in can be generated using big data analytics (e.g., APACHE SPARK and NEO4J (a graph database)). Such analysis can be used to evaluate features for focused development (e.g., as long as some users interact with a large portion of users from a set of disciplines on certain document types or through specific features or actions on the documents).
[0052] Productivity and Data Management End-User Characteristics of Using Version Hash Link Graph
[0053] Additional embodiments of the present invention utilize a VHLG to drive end-user features. For a group of users who collaborate using data (e.g., users within a company), a VHLG can be constructed to capture access patterns and dependencies between file versions (as described above). Importantly, the actual files can reside anywhere. For example, if data is accessed on a mobile device, a web browser, or a local drive - as long as the content in the files is identical, then the data has the same signature.
[0054] Figure 6A version hash link graph / tree is shown where data accessed from various different locations / devices accesses version information using the same signature / hash / GUID. The VHLG 600 is shown accessing the structure from a local drive 602, a mobile device 604, and the web 606. Using the structure 600 (when a user accesses a file), embodiments of the present invention can determine, by referring to the VHLG 600, whether the user is updating the latest version, whether there is a new version available for reference, and so on. For example, both user 1 608 and user 2 610 accessed a file 612 represented by a hash from their local desktop computers 602. File 612 is the parent of file 614 (represented by hash E7D78D), and file 614 is in turn the parent of file 616 (represented by hash C7890) and file 618 (represented by hash 8977AA). Both user 3 620 and user 4 622 are accessing the same file version 618, as can be determined by the same hash 8977AA. It is noted that even though user 3 620 is on a mobile device 604 and user 4 622 is on the web 606, the file content 618 accessed by both is identical (confirmable by the hash 8977AA).
[0055] The concept and application of VHLG can be extended to operations and different formats. For example, a user can export from DWG to PDF or import DWG into a REVIT or INVENTOR document.
[0056] Based on the analysis of the VHLG 600, a graphical user interface (GUI) can be generated that allows the user to visualize the analysis in an understandable and comprehensive manner. In a first example, the GUI can (partially) consist of the VHLG itself (e.g., as Figure 3 and Figure 4A shown). Alternatively, the GUI can consist of various different formats / representations of the VHLG or others. For example, different icons and edges of the GUI of the VHLG can be distinguished based on color and / or pattern (reflecting different types / devices, etc.). In a specific example, different device types can be represented by different colors (e.g., green for desktops, red for the web, blue for mobile devices). In another example, one color (e.g., purple) can be used to specify the fingerprint of a particular file version, and the chain of nodes of that color (purple) can reflect a particular lineage. Additionally, different arrows of different sizes (e.g., combined with or independent of the edges themselves) can reflect the number of accesses to a particular node / fingerprint version).
[0057] While a description of the VHLG may be provided in one GUI, other GUIs may focus on specific aspects of data / data analysis. For example, a GUI may present a simple graph reflecting the portion of data accessed by different devices (e.g., where the x-axis is for the minimum number of file versions per lineage and the y-axis reflects the percentage of lineages accessed by more than one device). In another exemplary embodiment, a time series (heat) graph reflecting access patterns may be generated (where the y-axis is time slices, where the graph shows access frequencies across different companies / company types and allows a viewer to quickly determine whether the communication is synchronous or asynchronous and on which documents most coordination needs are concentrated).
[0058] Logical Flow
[0059] Figure 7 Shows a logic flow for tracking document version control according to one or more embodiments of the present invention.
[0060] At step 702, a first pre-hash of the first document version is generated before an open (or reference) operation is performed on the first document version.
[0061] At step 704, a user or application performs an open operation.
[0062] At step 706, a first post-hash of the first document version is generated after the open operation is performed.
[0063] At step 708, before a save operation is performed on the first document version, the first pre-hash of the first document version is obtained (e.g., the hash function may be executed again, resulting in the same hash ID, and / or the system may recognize that the open operation did not change the first document version and thus, obtain / acquire the same first pre-hash).
[0064] At step 710, a user or application performs a save operation on the first document version, resulting in a second document version.
[0065] At step 712, a second post-hash of the second document version is generated after the save operation is performed.
[0066] At step 714, a Version Hash Link Graph (VHLG) is generated (or amplified). The VHLG includes: a first document version node (including a first pre-hash); a second document version node (including a second post-hash); a user-application node corresponding to the user and application that perform the open operation and the save operation; an open operation edge that connects the user-application node to the first document version node (where the open operation edge identifies the open operation); a save operation edge that connects the user-application node to the second document version node (where the save operation edge identifies the save operation); and a document edge that connects the first document version node to the second document version node.
[0067] The generation / amplification of the VHLG may include sending information about the open operation and the save operation to a VHLG generation service that generates the VHLG based on the information. The information for each operation includes the pre-hash before the operation, the post-hash after the operation, and the identification of the operation. Such information may be specified / provided in a log entry that includes an anonymous user identifier corresponding to the user and application, the platform used by the user or application, the identification of the operation, the pre-hash, the post-hash, and the time when the operation is performed. For example, the information for the open operation may consist of a first pre-hash, a first post-hash, and the identification of the open operation. The corresponding log entry may include an anonymous user identifier corresponding to the user and application, the platform used by the user or application, the identification of the open operation, the first pre-hash, the first post-hash, and the time when the open operation is performed. Thus, the VHLG can be amplified by individual logged entries, and the latest graphical structure can be generated on demand based on the log entries recorded in a database (e.g., SPARK) or other data aggregation backends.
[0068] Similar to the open operation, the information for the save operation may consist of a first pre-hash, a second post-hash, and the identification of the save operation. The corresponding log entry may consist of an anonymous user identifier corresponding to the user and application, the platform used by the user or application, the identification of the open operation, the first pre-hash, the second post-hash, and the time when the save operation is performed.
[0069] At step 716, a complete history of the document based on the VHLG is provided.
[0070] At optional step 718, insights (e.g., for community detection and collaboration networks) / end-user features (e.g., for sharing and collaborating on the document) may be provided.
[0071] It can be noted that the VHLG scales across time and multiple users or multiple applications. At step 718, once scaled, exemplary insights, lineages can be determined based on the VHLG. In an exemplary embodiment, such a lineage can consist of a collection of documents related by a version chain, where the collection consists of a first document version and a second document version, and the version chain includes document edges.
[0072] At step 718, additional insights / features can include / provide augmenting the first document version node and the second document version node with context data (e.g., industry type, company type, etc.). Based on the lineage and the context data, a collaborative community (e.g., industry type) can be determined. In an alternative embodiment, an access pattern over time can be determined based on the VHLG.
[0073] It can also be noted that the data of the documents can be stored independently of the VHLG (e.g., the actual files can reside anywhere). In this regard, regardless of the device platform (e.g., desktop, mobile device, web) used to access the first document version, the first pre-hash remains the same (i.e., as long as the content in the first document version is identical). Similarly, regardless of the device platform used to access the second document version, the second post-hash remains the same (i.e., as long as the content in the second document version is identical).
[0074] The additional insights / features provided at step 718 can include a graphical user interface (GUI) consisting of a visualization of the VHLG that shows the access pattern and dependencies between the first document version and the second document version. Such a GUI can be provided to collaborating users.
[0075] In addition to the above, the collaborative features in step 718 can include actions performed when a user or application accesses the first document version or the second document version. Specifically, based on the VHLG, a determination can be made as to whether the first document version or the second document version is the latest version of the document (e.g., by checking the VHLG). The result of the determination can then be informed to the accessing user / application.
[0076] Advantages
[0077] Embodiments of the present invention provide advantages over the prior art in various different contexts including desktop-web-mobile device workflows, drawing lifecycles, and collaborative clusters.
[0078] The desktop-web-mobile device workflow involves how files move between applications on a desktop computer (e.g., AUTOCAD DESKTOP), web-based versions of applications on the web (e.g., AUTOCAD WEB), and mobile-device-based versions of applications on a mobile device (e.g., AUTOCAD MOBILE), and where each of these products fits within a cross-platform workflow. In this context, analytics can answer the following questions:
[0079] · How many (per week / month) drawings created and edited on the desktop are opened on the web and on mobile devices? (Or, conversely, how many new drawings are created and then opened again on the web and on mobile devices?)
[0080] · What is the typical range of drawing exchange intervals between desktop-web or desktop-mobile devices? (For example, interval range = a dwg is opened on the desktop and then opened on a mobile device two days later, or opened on the web four hours later, etc.)
[0081] · When drawings are sent to the web and mobile devices, how do the drawings change? What activities do these users perform on each platform? Can we distinguish this through single-user (the same person using 3 different platforms) and team (different users using different platforms) use cases?
[0082] · How many dwgs are created first on a mobile device or the web and then brought back to the desktop?
[0083] The drawing lifecycle involves how a single drawing (e.g., a DWG file) evolves over time. This can include how the drawing is version-controlled, how the drawing is enhanced or marked, and who edits the drawing. Analytics can be used to answer the following questions about the drawing lifecycle:
[0084] · Can we see the number of collaborators on different layers per DWG? For example, 1-2, 2-10, 10-20, etc.
[0085] · Can we relate the stages of a DWG to typical project stages, including evolution over time (concept design, design development, construction documentation, etc.)?
[0086] · Can we identify DWGs as part of the same project?
[0087] · How often does the "same" DWG change its file name? (The customer is using "Save As" for project archiving, DWG records, version control, etc.) How do we filter out templates and standard DWGs (e.g., detailed forms repeated in multiple projects, etc.)?
[0088] Can we infer anything about out-of-band collaborators that happened in other tools? (e.g., Revit, Bluebeam, Paper, etc.)
[0089] Can we identify how often users start projects from existing DWGs?
[0090] How do different collaborators access files at different stages and on which platforms?
[0091] Collaboration clustering (also called collaboration role clustering) involves finding similarities in how people collaborate. Are some people super-sharers and others end-sharers? What do these super-sharers have in common? Analytics can be used to answer the following questions about collaboration clustering:
[0092] User clusters: How many DWGs did a customer process in X period of time?
[0093] User Clusters (Collaboration Tiers): With how many people did the client exchange DWGs? Is it possible to break this frequency down into internal (designers to drafters) and external (contractors, project managers, etc.) collaborators?
[0094] User clusters: Can we define super collaborators? What types of users collaborate the most? What industries? What technologies do they use?
[0095] Can we identify “hot” collaboration periods during the lifecycle when more users interact with each other more frequently?
[0096] Drill down deeper into any of these questions: If so, what storage provider are they on (on-premises / cloud)? Are they in the same environment (e.g. 1 office) or multiple environments / at home? Do they use references to break up the work? How many references are there in the drawing? What file format?
[0097] Which platforms have which pieces of collaborative storytelling?
[0098] Can we map out collaboration trends (across products in the user ecosystem) and which stages have the most back-and-forth collaboration, or which stages take the longest to complete?
[0099] Hardware Implementation
[0100] Figure 8Is an exemplary hardware and software environment 800 (referred to as a computer-implemented system and / or computer-implemented method) for implementing one or more embodiments of the present invention. The hardware and software environment includes a computer 802 and may include peripheral devices. The computer 802 can be a user / client computer, a server computer, or can be a database computer. The computer 802 includes a hardware processor 804A and / or a special-purpose hardware processor 804B (collectively referred to hereinafter as the processor 804) and a memory 806, such as random access memory (RAM). The computer 802 can be coupled to and / or integrated with other devices, including input / output (I / O) devices such as a keyboard 814, a cursor control device 816 (e.g., a mouse, pointing device, pen, and tablet computer, touch screen, multi-touch device, etc.), and a printer 828, etc. In one or more embodiments, the computer 802 can be coupled to or can include a portable or media viewing / listening device 832 (e.g., an MP3 player, IPOD, NOOK, portable digital video player, cellular device, personal digital assistant, etc.). In yet another embodiment, the computer 802 can include a multi-touch device, a mobile phone, a gaming system, an Internet-enabled television, a set-top box, or other Internet-enabled devices that execute on various platforms and operating systems.
[0101] In one embodiment, the computer 802 operates by executing instructions defined by a computer program 810 (e.g., a computer-aided design [CAD] application program) under the control of an operating system 808 by a hardware processor 804A. The computer program 810 and / or the operating system 808 can be stored in the memory 806 and can interface with a user and / or other devices to receive input and commands and provide output and results based on such input and commands and the instructions defined by the computer program 810 and the operating system 808.
[0102] The output / result can be presented on the display 822 or provided to another device for presentation or further processing or action. In one embodiment, the display 822 includes a liquid crystal display (LCD) having a plurality of individually addressable liquid crystals. Alternatively, the display 822 may include a light emitting diode (LED) display having clusters of red, green, and blue diodes driven together to form full-color pixels. In response to data or information generated by the processor 804 in accordance with the instructions of the computer program 810 and / or the operating system 808 for the application of inputs and commands, each liquid crystal or pixel of the display 822 changes to an opaque or translucent state to form part of an image on the display. The image can be provided by the graphical user interface (GUI) module 818. Although the GUI module 818 is depicted as a separate module, the instructions for performing the GUI functions may reside in or be distributed among the operating system 808, the computer program 810, or implemented using special-purpose memory and processors.
[0103] In one or more embodiments, the display 822 is integrated with or into the computer 802 and includes a multi-touch device having a touch-sensing surface (e.g., a track pod or touch screen) capable of recognizing the presence of two or more contact points with the surface. Examples of multi-touch devices include mobile devices (e.g., IPHONE, NEXUS S, DROID devices, etc.), tablet computers (e.g., IPAD, HP TOUCHPAD, SURFACE devices, etc.), portable / handheld game / music / video player / console devices (e.g., IPOD TOUCH, MP3 players, NINTENDO SWITCH, PLAYSTATION PORTABLE, etc.), touch tables, and walls (e.g., where an image is projected through acrylic and / or glass and then illuminated from behind with LEDs).
[0104] Some or all of the operations performed by the computer 802 in accordance with the instructions of the computer program 810 can be implemented in a special-purpose processor 804B. In this embodiment, some or all of the computer program 810 instructions can be implemented via firmware instructions stored in read-only memory (ROM), programmable read-only memory (PROM), or flash memory within the special-purpose processor 804B or in the memory 806. The special-purpose processor 804B can also be hardwired by circuit design to perform some or all of the operations to implement the present invention. Additionally, the special-purpose processor 804B can be a hybrid processor that includes dedicated circuitry for performing a subset of functions and other circuitry for performing more general functions such as in response to the instructions of the computer program 810. In one embodiment, the special-purpose processor 804B is an application-specific integrated circuit (ASIC).
[0105] The computer 802 may also implement a compiler 812 that allows application programs or computer programs 810 written in programming languages such as C, C++, Assembly, SQL, PYTHON, PROLOG, MATLAB, RUBY, RAILS, HASKELL, or other languages to be translated into processor 804-readable code. Alternatively, the compiler 812 may be an interpreter that directly executes instructions / source code, translates source code into an intermediate representation to be executed, or executes stored precompiled code. Such source code may be written in a variety of programming languages, such as JAVA, JAVASCRIPT, PERL, BASIC, etc. After completion, the application program or computer program 810 uses the relationships and logic generated using the compiler 812 to access and manipulate data received from the I / O device and stored in the memory 806 of the computer 802.
[0106] The computer 802 also optionally includes external communication devices, such as modems, satellite links, Ethernet cards, or other devices for receiving input from and providing output to other computers 802.
[0107] In one embodiment, the instructions implementing the operating system 808, computer program 810, and compiler 812 are tangibly embodied in a non-transitory computer-readable medium, such as a data storage device 820, which may include one or more fixed or removable data storage devices, such as zip drives, floppy disk drives 824, hard disk drives, CD-ROM drives, tape drives, etc. Additionally, the operating system 808 and computer program 810 consist of computer program 810 instructions that, when accessed, read, and executed by the computer 802, cause the computer 802 to perform the steps required to implement and / or use the present invention, or load the instruction program into the memory 806, thus creating a special-purpose data structure that causes the computer 802 to operate as a specifically programmed computer for performing the method steps described herein. The computer program 810 and / or operating instructions may also be tangibly embodied in the memory 806 and / or data communication device 830, thereby manufacturing a computer program product or article according to the present invention. Thus, as used herein, the terms "article", "program storage device", and "computer program product" are intended to cover computer programs accessible from any computer-readable device or medium.
[0108] Of course, those skilled in the art will recognize that any combination of the above components or any number of different components, peripherals, and other devices may be used with the computer 802.
[0109] Figure 9Schematically shown is a typical distributed / cloud-based computer system 900 that uses a network 904 to connect a client computer 902 to a server computer 906. A typical combination of resources may include: a network 904, which includes the Internet, a LAN (local area network), a WAN (wide area network), an SNA (systems network architecture) network, etc.; a client 902, which is a personal computer or a workstation (as Figure 8 stated in); and a server 906, which is a personal computer, a workstation, a minicomputer, or a mainframe (as Figure 8 stated in). However, it may be noted that different networks such as cellular networks (e.g., GSM [Global System for Mobile Communications] or others), satellite-based networks, or any other type of network may be used to connect the client 902 to the server 906 according to embodiments of the present invention.
[0110] A network 904 such as the Internet connects the client 902 to the server computer 906. The network 904 may utilize Ethernet, coaxial cable, wireless communication, radio frequency (RF), etc. to connect the client 902 to the server 906 and provide communication between the client 902 and the server 906. Additionally, in a cloud-based computing system, resources (e.g., storage devices, processors, applications, memory, infrastructure, etc.) in the client 902 and the server computer 906 may be shared by the client 902, the server computer 906, and users across one or more networks. The resources may be shared by multiple users and may be dynamically reallocated according to demand. In this regard, cloud computing may be referred to as a model for accessing a shared pool of configurable computing resources.
[0111] The client 902 may execute a client application or a web browser and communicate with the server computer 906 that executes a web server 910. Such a web browser is typically a program such as MICROSOFT INTERNET EXPLORER / EDGE, MOZILLAFIREFOX, OPERA, APPLE SAFARI, GOOGLE CHROME, etc. Additionally, software executed on the client 902 may be downloaded from the server computer 906 to the client computer 902 and installed as a plugin for the web browser or an ACTIVEX control. Thus, the client 902 may utilize ACTIVEX components / component object model (COM) or distributed COM (DCOM) components to provide a user interface on the display of the client 902. The web server 910 is typically a program such as MICROSOFT’S INTERNETINFORMATION SERVER.
[0112] The web server 910 can host Active Server Pages (ASP) or Internet Server Application Programming Interface (ISAPI) applications 912 that can execute scripts. The scripts call objects (referred to as business objects) that execute business logic. The business objects then manipulate data in the database 916 through a database management system (DBMS) 914. Alternatively, the database 916 can be part of the client 902 or directly connected to the client 902, rather than communicating with / obtaining information from the database 916 through the network 904. When a developer encapsulates business functionality into objects, the system can be referred to as a Component Object Model (COM) system. Thus, the scripts executed on the web server 910 (and / or the application 912) call COM objects that implement business logic. In addition, the server 906 can utilize MICROSOFT’S TRANSACTION SERVER (MTS) to access the required data stored in the database 916 via interfaces such as ADO (Active Data Objects), OLE DB (Object Linking and Embedding Database), or ODBC (Open Database Connectivity).
[0113] Generally, all of these components 900 to 916 include logic and / or data that is embodied in / devivable from a device, medium, signal, or carrier, such as a data storage device, a data communication device, a remote computer, or a device coupled to a computer via a network or via another data communication device, etc. In addition, this logic and / or data, when read, executed, and / or interpreted, causes the steps required to implement and / or use the present invention to be performed.
[0114] Although the terms "user computer", "client computer", and / or "server computer" are mentioned herein, it should be understood that such computers 902 and 906 can be interchangeable and can also include thin client devices with limited or full processing capabilities, portable devices such as mobile phones, laptop computers, pocket computers, multi-touch devices, etc., and / or any other devices with suitable processing, communication, and input / output capabilities.
[0115] Of course, those skilled in the art will recognize that any combination of the above components or any number of different components, peripherals, and other devices can be used with the computers 902 and 906. Embodiments of the present invention are implemented as software / CAD applications on the client 902 or the server computer 906. In addition, as described above, the client 902 or the server computer 906 can include a thin client device or a portable device with a multi-touch-based display.
[0116] Conclusion
[0117] The description of the preferred embodiments of the present invention ends here. Some alternative embodiments for implementing the present invention are described below. For example, any type of computer (such as mainframe computers, minicomputers, or personal computers, etc.) or computer configuration (such as time-sharing mainframe computers, local area networks, etc.) or stand-alone personal computers can be used with the present invention. In summary, the embodiments of the present invention provide at least one or more of the following features:
[0118] 1. Based on the concept of a Version History Link Graph (VHLG) that can connect users and companies;
[0119] 2. Using the VHLG for insights into community detection and collaboration networks; and
[0120] 3. Using the VHLG to drive end-user features for document sharing and collaboration.
[0121] For purposes of illustration and description, the foregoing description of the preferred embodiments of the present invention has been presented. The foregoing description is not intended to be exhaustive or to limit the invention to the precise form disclosed. Many modifications and variations are possible in light of the above teachings. It is intended that the scope of the present invention not be limited by this detailed description, but rather by the appended claims.
Claims
1. A computer-implemented method for tracking document version control, which includes: (a) Generating a first pre-hash of the first document version before performing an open operation on the first document version; (b) The user or application performing the open operation; (c) Generating a first post-hash of the first document version after performing the open operation; (d) Obtaining the first pre-hash of the first document version before performing a save operation on the first document version; (e) The user or the application performing a save operation on the first document version, thereby generating a second document version; (f) Generating a second post-hash of the second document version after performing the save operation; (g) Generating a Version Hash Link Graph (VHLG), where the VHLG includes: (i) A first document version node that includes the first pre-hash; (ii) A second document version node that includes the second post-hash; (iii) A user-application node that corresponds to the user or application that performs the open operation and the save operation; (iv) An open operation edge that connects the user-application node to the first document version node, where the open operation edge identifies the open operation; (v) A save operation edge that connects the user-application node to the second document version node, where the save operation edge identifies the save operation; and (vi) A document edge that connects the first document version node to the second document version node; and (h) Providing a complete history of the document based on the VHLG.
2. The computer-implemented method according to claim 1, wherein generating the VHLG includes: Sending information about the open operation and information about the save operation to a VHLG generation service, where: The information about the open operation includes the first pre-hash, the first post-hash, and an identifier of the open operation; and The information about the save operation includes the first pre-hash, the second post-hash, and an identifier of the save operation; and The VHLG generation service generates the VHLG based on the information about the open operation and the information about the save operation.
3. The computer-implemented method according to claim 2, wherein: The information about the open operation includes a first log entry; The first log entry includes: An anonymous user identifier corresponding to the user or the application; The platform used by the user or the application; The identifier of the open operation; The first pre-hash; The first post-hash; and The time when the open operation is performed; The information about the save operation includes a second log entry; The second log entry includes: The anonymous user identifier corresponding to the user or the application; The platform used by the user or the application; The identifier of the save operation; The first pre-hash; The second post-hash; and The time when the save operation is performed.
4. The computer-implemented method according to claim 1, which further includes: Scale the VHLG across time and multiple users or multiple applications; Determine a pedigree based on the VHLG, where: The pedigree includes a set of documents related by a version chain; The set of documents includes the first document version and the second document version; and The version chain includes the document edges.
5. The computer-implemented method according to claim 4, which further includes: Augment the first document version node and the second document version node with context data; Determine a collaborative community based on the pedigree and the context data.
6. The computer-implemented method according to claim 5, which further includes: Determine the industry type of the community based on the context data.
7. The computer-implemented method according to claim 5, which further includes: Determine an access pattern over time based on the VHLG.
8. The computer-implemented method according to claim 1, where: The data of the document is stored independently of the VHLG; The first pre-hash remains unchanged regardless of the device platform used to access the first document version; and The second post-hash remains unchanged regardless of the device platform used to access the second document version.
9. The computer-implemented method according to claim 1, which further includes: Provide a graphical user interface GUI including a visualization of the VHLG to collaborating users, where the visualization shows the access pattern and dependencies between the first document version and the second document version.
10. The computer-implemented method according to claim 1, which further includes: The user or the application accesses the first document version or the second document version; Determine whether the first document version or the second document version is the latest version of the document based on the VHLG, where the determination includes querying the VHLG; and Inform the user or the application of the result of the determination.
11. A computer-implemented system for tracking document version control, which includes: (a) A computer having a memory; (b) A processor executing on the computer; (c) The memory storing a set of instructions, where the set of instructions, when executed by the processor, causes the processor to perform operations including the following: (i) Generate a first pre-hash of the first document version before performing an open operation on the first document version; (ii) The user or the application performs the open operation; (iii) Generate a first post-hash of the first document version after performing the open operation; (iv) Obtain the first pre-hash of the first document version before performing a save operation on the first document version; (v) The user or the application performs a save operation on the first document version, thereby generating a second document version; (vi) Generate a second post-hash of the second document version after performing the save operation; (vii) Generate a version hash link graph VHLG, where the VHLG includes: (A) The first document version node, which includes the first pre-hash; (B) The second document version node, which includes the second post-hash; (C) A user-application node, which corresponds to the user or application that performs the open operation and the save operation; (D) An open operation edge that connects the user-application node to the first document version node, where the open operation edge identifies the open operation; (E) A save operation edge that connects the user-application node to the second document version node, where the save operation edge identifies the save operation; and (F) A document edge that connects the first document version node to the second document version node; and (viii) Provide a complete history of the document based on the VHLG.
12. The computer-implemented system according to claim 11, wherein the operation of generating the VHLG comprises: Sending the information of the open operation and the information of the save operation to a VHLG generation service, where: The information of the open operation includes the first pre-hash, the first post-hash, and the identifier of the open operation; and The information of the save operation includes the first pre-hash, the second post-hash, and the identifier of the save operation; and The VHLG generation service generates the VHLG based on the information of the open operation and the information of the save operation.
13. The computer-implemented system according to claim 12, wherein: The information of the open operation includes a first log entry; The first log entry includes: An anonymous user identifier corresponding to the user or the application; The platform used by the user or the application; The identifier of the open operation; The first pre-hash; The first post-hash; and The time when the open operation is performed; The information of the save operation includes a second log entry; The second log entry includes: The anonymous user identifier corresponding to the user or the application; The platform used by the user or the application; The identifier of the save operation; The first pre-hash; The second post-hash; and The time when the save operation is performed.
14. The computer-implemented system according to claim 11, wherein the operation further comprises: Scaling the VHLG proportionally across time and multiple users or multiple applications; Determining a pedigree based on the VHLG, where: The pedigree includes a set of documents related by a version chain; The set of documents includes the first document version and the second document version; and The version chain includes the document edge.
15. The computer-implemented system according to claim 14, wherein the operation further comprises: Augmenting the first document version node and the second document version node with context data; Determining a collaborative community based on the pedigree and the context data.
16. The computer-implemented system according to claim 15, wherein the operation further comprises: Determining the industry type of the community based on the context data.
17. The computer-implemented system according to claim 15, wherein the operation further comprises: determining an access pattern over time based on the VHLG.
18. The computer-implemented system according to claim 11, wherein: data of the document is stored independently of the VHLG; the first pre-hash remains unchanged regardless of the device platform used to access the first document version; and the second post-hash remains unchanged regardless of the device platform used to access the second document version.
19. The computer-implemented system according to claim 11, wherein the operation further comprises: providing a graphical user interface GUI including a visualization of the VHLG to collaborating users, wherein the visualization shows an access pattern and dependencies between the first document version and the second document version.
20. The computer-implemented system according to claim 11, wherein the operation further comprises: the user or the application accesses the first document version or the second document version; determining, based on the VHLG, whether the first document version or the second document version is the latest version of the document, wherein the determination includes querying the VHLG; and informing the user or the application of the result of the determination.
Citation Information
Patent Citations
Dynamically building file graph
US20200151280A1
Method and system for document lineage tracking
US20200341957A1