Information processing apparatus, storage medium, information processing method, and computer program product
By tracing document versions and editing histories to identify sample documents, the problem of reflecting document element relationships in copied documents was solved, achieving accurate reflection of relationships and reasonable selection of sample documents.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- FUJIFILM BUSINESS INNOVATION CORP
- Filing Date
- 2020-09-01
- Publication Date
- 2026-05-08
AI Technical Summary
Existing technologies struggle to reflect the relationships between document elements in documents copied from sample documents, especially when the same relationships are not explicitly defined in the sample documents.
By tracing the document's version management information and editing history, the correspondence between the sample document and the latest version document is determined, and the corresponding document element relationships are set within the sample document to reflect these relationships.
Even if the relationship is not explicitly defined in the sample document, the corresponding document element relationship can be reflected in the copied document, reducing the possibility of unsuitable documents being selected as sample documents and improving the accuracy of relationship reflection.
Smart Images

Figure CN113139372B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to an information processing device, a storage medium, an information processing method, and a computer program product. Background Technology
[0002] As a method for creating electronic documents such as document files, it is common to add or modify a sample (i.e., a template). In addition to sample data prepared in a sample-specific data format, copies of existing documents are sometimes used as samples.
[0003] Furthermore, it records and utilizes information about the relationships between documents or document elements and other documents or document elements. Additionally, a document element is a component of a document; in other words, a document consists of one or more document elements.
[0004] In the system disclosed in Patent Document 1, when a user selects a document as a source for copying, the system compares the source document group to which the selected document belongs with a subsequent document group created based on that source document group. Furthermore, in addition to an image representing the structure of the source document group, symbols are used to display the structural differences between the source and subsequent document groups. The user selects the desired document and presses the registration button. Thus, the document group is copied, including the associations between documents. Patent Document 1 illustrates, as an example of the associations between documents, the associations between these documents when an ISO document group consists of a plurality of documents forming a certain structure.
[0005] Patent Document 1: Japanese Patent Application Publication No. 2007-323441 Summary of the Invention
[0006] The object of the present invention is to enable the corresponding document elements in a document copied from the sample documents to reflect relationships between each other when document elements in the documents are associated with each other, even if the same relationships are not explicitly defined in the sample documents of these documents.
[0007] The invention involved in Scheme 1 is an information processing device, characterized in that it includes a processor, which performs the following processing: when a relationship is established between a first document element and a second document element, a first sample document corresponding to a first document containing the first document element and a second sample document corresponding to a second document containing the second document element are determined; the relationship is established between the document elements in the first sample document that correspond to the first document element (i.e., the first sample element) and the document elements in the second sample document that correspond to the second document element (i.e., the second sample element); after the relationship is established between the first sample element and the second sample element, a first copy document (copied from the first sample document) and a second copy document (copied from the second sample document) are generated; the relationship is established between the document elements in the first copy document that correspond to the first sample element and the document elements in the second copy document that correspond to the second sample element.
[0008] The invention involved in Scheme 2, in the information processing apparatus described in Scheme 1, is characterized in that, when the relationship is set between the first sample element and the second sample element, the processor sets the relationship between the document element corresponding to the first sample element in an existing third document generated from the first sample document and the document element corresponding to the second sample element in an existing fourth document generated from the second sample document.
[0009] In the invention of Scheme 3, in the information processing apparatus described in Scheme 1 or 2, the first sample document is determined from the document group of the ancestors of the first document according to the number of times it has been copied.
[0010] In the information processing apparatus of any one of Schemes 1 to 3, the invention involved in Scheme 4, when tracing the first document back to its ancestor side one by one, will not select the ancestor side document as the first sample document if the similarity between adjacent documents to each other in content other than the main text is below a threshold.
[0011] The invention involved in Scheme 5 is a storage medium storing a program that causes a computer to perform the following processes: when a relationship is established between a first document element and a second document element, a first sample document corresponding to a first document containing the first document element and a second sample document corresponding to a second document containing the second document element are determined; the relationship is established between the document elements in the first sample document that correspond to the first document element (i.e., the first sample element) and the document elements in the second sample document that correspond to the second document element (i.e., the second sample element); after the relationship is established between the first sample element and the second sample element, when a first copy document (copied from the first sample document) and a second copy document (copied from the second sample document) are generated, the relationship is established between the document elements in the first copy document that correspond to the first sample element and the document elements in the second copy document that correspond to the second sample element.
[0012] The invention involved in Scheme 6 is an information processing method, characterized by the following steps: when a relationship is established between a first document element and a second document element, a first sample document corresponding to a first document containing the first document element and a second sample document corresponding to a second document containing the second document element are determined; the relationship is established between the document elements in the first sample document that correspond to the first document element (i.e., the first sample element) and the document elements in the second sample document that correspond to the second document element (i.e., the second sample element); after the relationship is established between the first sample element and the second sample element, a first copy document (copied from the first sample document) and a second copy document (copied from the second sample document) are generated, the relationship is established between the document elements in the first copy document that correspond to the first sample element and the document elements in the second copy document that correspond to the second sample element.
[0013] Invention Effects
[0014] According to the first, fifth, or sixth aspect of the invention, when a relationship is established between document elements in documents, even if the same relationship is not explicitly established in the sample documents of these documents, the relationship can be reflected between corresponding document elements in documents copied from these sample documents.
[0015] According to the second aspect of the present invention, the document elements of an existing document generated from a sample document can also reflect the relationships set between the document elements of the original document.
[0016] According to the third aspect of the present invention, compared with the case where the number of times it is copied is not considered, the possibility of a document that is copied less often being identified as a sample document can be reduced.
[0017] According to the fourth aspect of the present invention, compared with the case where similarity is not considered, it is possible to reduce the possibility that a document with low similarity to a document containing document elements with established relationships is identified as a sample document. Attached Figure Description
[0018] The embodiments of the present invention will be described in detail with reference to the following figures.
[0019] Figure 1 A diagram illustrating an example of establishing relationships between document elements in a document that has been upgraded through editing;
[0020] Figure 2 A diagram used to illustrate the identification of a sample of documents containing document elements with defined relationships, and the establishment of relationships between document elements in these samples;
[0021] Figure 3 A diagram illustrating examples of samples determined not only by version but also by tracing back to copies of the document;
[0022] Figure 4 A diagram illustrating the hardware structure of a computer is shown below;
[0023] Figure 5 Here is a diagram illustrating the content of version management information;
[0024] Figure 6 The following diagram illustrates the content of the operation history information;
[0025] Figure 7 This is a diagram illustrating the processing flow executed by the processor for a document editing system or its sample management functions;
[0026] Figure 8 A diagram illustrating a UI screen used to set relationships between document elements;
[0027] Figure 9 For the purpose of illustration Figure 7 The diagram shows cases of failures that occurred during the process identified in the sample determination.
[0028] Figure 10 The following diagram illustrates the process of determining samples by considering similarity and the number of copies.
[0029] Figure 11 This is a diagram used to illustrate how, when a combination of sample documents that reflects the relationships between document elements is copied, the relationships between document elements in the resulting document are also reflected.
[0030] Figure 12 For example, this diagram illustrates the process of replicating a combination of sample documents that reflect the relationships between document elements, and showing the workflow of replicating documents that reflect the same relationships between document elements.
[0031] Figure 13 The diagram illustrates the process of creating a sample document that reflects the relationships between document elements, and how copying an existing document reflects the same relationships.
[0032] Symbol Explanation
[0033] 11, 21, 31 - Sample documents; 12, 22, 23, 32 - Documents; 102 - Processor; 104 - Memory; 106 - Auxiliary storage device; 108 - Input / output device; 110 - Network interface; 112 - Bus. Detailed Implementation
[0034] In the following description, electronic documents such as document files will be simply referred to as documents. Furthermore, each document consists of one or more document elements. For example, a document can be described using a markup language such as HTML (Hypertext Markup Language), which can describe the structure composed of groups of document elements, or it can be text data that describes document structures such as chapters, sections, and paragraphs according to prescribed rules.
[0035] Figure 1 The diagram illustrates an example of establishing relationships between document elements in documents generated from samples. In this example, sample document 11 contains a first document element with the title "1. Power Consumption Reference" and a second document element with the title "2. Noise Reference". Document 12 is generated by appending text content such as "Power Consumption..." and "Noise..." to each document element of sample document 11. When the document editing environment has version control functionality, the version number of sample document 11 is 1 (marked as "Ver1" in the diagram), and the version number of document 12 is 2. Similarly, text content is appended to the first document element of another sample document 21, thereby generating version 2, i.e., document 22. Then, text content is appended to version 2, i.e., document 23. Furthermore, text content is appended to the two document elements of yet another sample document 31, thereby generating version 2, i.e., document 32.
[0036] Here, the relationship between the second document element of document 23 "referencing" the second document element of document 12 and the relationship between the second document element of document 32 "referencing" the first document element of document 12 are set by the same or different users. In addition, the data representing the relationship between these document elements can be included in the document as attribute data of the document or the document element, or it can be registered separately from each document in the document management system that manages these documents.
[0037] In this case, for example, when a new document is created by copying document 23, the relationship between the second document element of the new document and the second document element of document 12 can be automatically set using conventional techniques. Furthermore, when a group of documents consisting of documents 12, 23, and 32 is copied together to generate three new documents, the relationship between the document elements of the copy source document group can also be copied and automatically set using conventional techniques.
[0038] Sometimes it is desirable to reflect the relationships established between document elements in individual documents 12, 23, and 32 generated from the sample document back to the relationships between document elements in the original sample document. However, a technique for easily reflecting such relationships to the sample document is currently unknown.
[0039] In the embodiments shown below, a method is illustrated for reflecting the relationships established between document elements of individual documents in relation to the document elements of a sample document that forms the basis of these individual documents.
[0040] refer to Figure 2 A summary of the method is provided. Figure 2 In some cases, with Figure 1 The document elements of the latest versions of the same documents 11-32, namely documents 12, 23, and 32, are set with the same... Figure 1 The same "reference" relationship applies to the case.
[0041] The system of this embodiment determines the sample documents 11, 21, and 31 corresponding to each of these documents 12, 23, and 32 by tracing the version management information or editing history of each document 12, 23, and 32. Furthermore, the system determines that the first and second document elements of the determined sample document 11 correspond to the first and second document elements of the latest version of document 12, respectively. This determination can be performed, for example, based on the identification information of the document elements contained in each document element within these documents 11 and 12. This identification information can be, for example, a string representing the title of a document element. Furthermore, when a document is described using a markup language, the document's identification information can be described in tags (e.g., HTML tags) that represent the scope of document elements. Similarly, the system determines that the second document element of the latest version of document 23 corresponds to the second document element of sample document 21, and the second document element of the latest version of document 32 corresponds to the second document element of sample document 31. Furthermore, based on these determined results, the system sets the "reference" relationship between the second document element of sample document 31 and the first document element of sample document 11 according to the "reference" relationship between the second document element of document 32 and the first document element of document 12. Similarly, the system sets the "reference" relationship between the second document element of sample document 21 and the second document element of sample document 11 according to the "reference" relationship between the second document element of document 23 and the second document element of document 12.
[0042] exist Figure 2 In the example shown, it is assumed that individual documents are generated by sequentially editing sample documents. In this example, the earliest version determined by tracing back the versions of documents containing document elements with defined relationships is the sample document corresponding to that document.
[0043] However, sometimes the earliest version obtained by tracing document versions is not suitable as a sample document for that document. A typical example is when the earliest version was created by copying another document.
[0044] For example, a common method is to create a desired document by copying a document, making a copy, and then editing that copy. When the earliest version of a document containing document elements with defined relationships is found to be a copy of another document, the original document or its earliest version is more suitable as a sample document than the original document itself.
[0045] exist Figure 3In the example shown, version 1 of document A2 is created by copying a sample document, namely document A1, which contains only the titles of each document element. Version 2 is then created by adding content to version 1 that describes each document element. Version 1 of document A3 is created by copying version 2 of document A2. Furthermore, version 2 is created by editing the content of each document element in version 1 of document A3, and version 3 is created by editing version 2. In this example, consider setting a relationship with other document elements for, for example, the first document element in version 3 of document A3, and reflect this relationship in the corresponding document elements in the sample document corresponding to document A3. For this purpose, it is necessary to determine the sample document corresponding to document A3 (in this example, version 1 of document A1).
[0046] In this determination process, the system obtains version 1 of document A3 by tracing back from version 3 of document A3, which contains document elements with established relationships. If a history of document copying operations or copying relationships between documents is recorded, the system copies version 1 of document A3 and determines version 2 of document A2 as the copy source. The system obtains version 1 of document A2 by tracing back the versions of document A2 and determines version 1 as a copy of version 1 of document A1. Furthermore, the system determines that there is no earlier version than version 1 in document A1, and that version 1 was not obtained by copying another document; based on this determination result, version 1 of document A1 is determined to be a sample document. Moreover, relationships are established between the first document element and other document elements within the determined sample document.
[0047] in addition, Figure 3 Documents A1, A2, and A3 shown are different documents with different document IDs. The ID is identification information. The document ID is a unique identifier assigned to each document by the system that manages the documents. Different versions of the same document have the same document ID, but their version numbers are different.
[0048] <Example of hardware architecture>
[0049] The relationships between document elements in individual documents, as described above, are reflected in the functions between document elements in sample documents, for example, and are incorporated into a document editing system that provides a document editing environment.
[0050] This function is achieved by having the computer execute a program that represents the function. For computers with document editing systems installed, the program representing the function can be installed as part of that system.
[0051] Here, as Figure 4As shown, the computer has, for example, the following circuit structure: a controller that controls the processor 102 (which is hardware), memory (main storage device) such as random access memory (RAM) 104, and auxiliary storage devices 106 such as flash memory, SSD (solid-state hard disk drive), and HDD (hard disk drive); an interface to various input / output devices 108; and a network interface 110 for controlling connections to a network such as a local area network (LAN), etc., connected via a data transmission path such as a bus 112. A program describing processing content that reflects the above-mentioned relationships is installed on the computer via a network or the like and stored in the auxiliary storage device 106. The program stored in the auxiliary storage device 106 is executed by the processor 102 and using the memory 104, thereby realizing the processing of reflecting the relationships of the sample document based on the above-mentioned functions.
[0052] Here, processor 102 refers to a processor in a broad sense, including general-purpose processors (such as CPU: Central Processing Unit, etc.) or special-purpose processors (such as GPU: Graphics Processing Unit, ASIC: Application Specific Integrated Circuit, FPGA: Field Programmable Gate Array, programmable logic devices, etc.).
[0053] Furthermore, the operation of processor 102 can be performed not only by a single processor 102, but also collaboratively by a plurality of processors 102 located in physically separate positions. Moreover, the operations of each processor 102 are not limited to the order described in the following embodiments and can be appropriately modified.
[0054] Furthermore, the auxiliary storage device 106 stores various information for performing relational processing on the sample documents. As an example of such information, there exists... Figure 5 The version management information shown in the example and Figure 6 The example shows the operation history information.
[0055] Figure 5The version management information illustrated here is the management information for each version of the document with document ID "Document A2". The version management information includes items such as version number, creation date and time, creator, and document data for each version of the document. The version number is the identifier for that version. In the example, a sequence number of more than one integer is used as the version number. However, this is just one example; for version numbers, any numbering system can be used as long as the sequential relationship between versions within the same document is clear. The creation date and time are the date and time when that version was created. Here, the initial (i.e., earliest) version of the document is created entirely new or by copying an existing document, while other versions are created by revising existing versions. The creator field registers the user ID of the user who performed the operation of creating the document on the document editing system. The document data field registers the document data itself for that version or information identifying the document data stored in auxiliary storage device 106, etc. (e.g., a link to the document data). The same version management information is stored in auxiliary storage device 106 for each document ID.
[0056] Figure 6 The example operation history information shows the history of operations performed by each user on the document editing system. Figure 6 The diagram illustrates some records from a large number of history entries within a specific period. Each history record includes items such as the operation date and time, operator, operation type, and operation parameters. The operation date and time are the date and time when the system executed the operation that is the object of this history record. The operator is the ID of the user who instructed the execution of the operation. The operation type is the type of operation. Operation types include copy, edit, and download (i.e., downloading a document from the system to the user's terminal). The operation parameters are parameters that specify the content of the operation. When it is a copy operation, the operation parameters are a group of the document ID of the source document and the ID of the destination document generated by copying from the source document. In the example diagram, the operation parameters of the copy operation involved in the initial history record indicate that the source document is document A and the destination document is document B. The parameters of the edit operation include the document ID of the document being edited, and may also include the version number of the version generated as a result of the edit.
[0057] <Example of processing flow>
[0058] Next, refer to Figure 7 An example of the processing flow executed by the processor 102 to demonstrate the above-mentioned sample management function will be explained. Figure 7 The process, for example, is executed when the user sets relationships between document elements. This relationship setting, for example, uses... Figure 8 The UI screen shown in the example is 200.
[0059] exist Figure 8 In the example, the content of the document element that will be the source of the relationship being set is displayed in the first display bar 202 on the left side of the UI (user interface) screen 200. For example, if the user selects a desired document element in a document being edited in a document editing system and instructs to perform a relationship setting operation, the UI screen 200 is displayed, and the content of the selected document element is displayed in the first display bar 202. Furthermore, the second display bar 204 on the right side of the UI screen 200 displays an input bar 206 for entering search criteria for document elements that will be the destination of the relationship being set, and a search button 208. If the user enters search criteria into the input bar 206 and presses the search button 208, the processor 102 searches for document elements that satisfy the search criteria from a database containing document groups. The search results display bar 210 below the input bar 206 displays the searched document element groups side-by-side in order of increasing degree of satisfaction of the search criteria. A checkbox is displayed for each searched document element to be selected as the destination of the relationship. To avoid complexity, in Figure 8 The document element for which only one search result is displayed is shown. Below the search results display bar 210, buttons 212 and 214 are displayed, indicating the setting of relationships. Buttons 212 and 214 respectively indicate setting a "reference" relationship and a "reference" relationship from the document element displayed in the first display bar 202 to the selected document element in the search results display bar 210. The distinction between "reference" and "reference" is predefined. In the example shown, the user presses button 212 while selecting the document element "1. Power Consumption Baseline" from the document "Electricity Evaluation Report" in the search results display bar 210. Consequently, a dialog box 220 is displayed on the UI screen 200, asking whether a "reference" relationship can be set for the document element "1. Power Consumption Baseline" in the document "Electricity Evaluation Report". The dialog box 220 displays a "Yes" button indicating to set the relationship and a "No" button indicating not to set the relationship. If the user presses the "Yes" button, the processor 102 sets a "reference" relationship from the document element displayed as the relationship source in the first display bar 202 to the document element "1. Power Consumption Basis" in the selected document "Electricity Evaluation Report" as the relationship destination. The information of the set reference relationship is, for example, registered in a database that maintains the relationships between document elements.
[0060] Additionally, an example is shown here where a user sets relationships between document elements in an individual document, but this is only one example. Instead, some document analysis systems can analyze each document element and set relationships between them based on the analysis results. Furthermore, the process can be triggered by a situation where a relationship has been set between document elements by this document analysis system. Figure 7 The process.
[0061] return Figure 7 The following is an explanation. As described above, when a relationship between document elements is established, the processor 102 determines the document containing the document element of the relationship source and the document containing the document element of the relationship destination (S10). In step S10, the document ID and version of the document containing the document element of the relationship source or relationship destination are determined. Furthermore, the processing of steps S20 and S30 is performed for each determined document. In step S20, the processor 102 determines the sample document corresponding to the determined document. The processing of step S20 will be described in detail later. Next, the processor 102 determines the document elements in the determined sample document that correspond to the document elements of the relationship source or relationship destination (S30). Thus, the document elements in the sample document containing the document element of the relationship source that correspond to the relationship source are determined, and the document elements in the sample document containing the document element of the relationship destination that correspond to the relationship destination are determined. The processor 102 sets the relationship between the document elements corresponding to the relationship source and the document elements corresponding to the relationship destination in each of the determined sample documents (S40). The information about the established relationships is registered in a database, for example, that maintains relationships between document elements.
[0062] Next, the detailed process of step S20 is illustrated. Processor 102 first designates one of the documents determined in S10 as a document of interest (S202). Next, processor 102 determines whether a previous version or a copy source document exists in the document of interest (S204). In step S204, processor 102, for example, refers to version management information corresponding to the document ID of the document of interest (reference...). Figure 5 The processor 102 checks whether a version number preceding the version number of the document in question exists. If no previous version number exists, the processor 102 retrieves the information from the operation history (see reference). Figure 6 The system searches for copy operation history records that include the document ID and version number of the document of interest as the copy destination in the operation parameters, and determines the document ID and version of the copy source document included in the operation parameters. Thus, in step S204, if a previous version of the document of interest or a copy source document of the document of interest exists, the version is determined.
[0063] When the determination result of step S204 is "yes", the processor 102 changes the document of interest to the determined previous version or copies the source document (S206) and repeats the processing of step S204. If the determination result of S204 is "yes", the processor 102 determines the document of interest at this time as the sample document corresponding to the document determined in S10 (S208).
[0064] Thus, through Figure 7The process illustrated in the example determines sample documents containing document elements with established relationships, and the document elements corresponding to these document elements within these sample documents establish the relationships between each other.
[0065] When version (k+1) of a document X is generated by editing version k (k is a positive integer), the relationship between version k and version (k+1) can be considered as a parent-child relationship where the former is the parent and the latter is the child. Similarly, when version 1 of another document Z is generated by copying version m of document Y, the relationship between version m of document Y and version 1 of document Z can be considered as a parent-child relationship where the former is the parent and the latter is the child. In this case, step S204 involves determining whether a parent exists in the document of interest. Furthermore, step S206 sets the parent of the current document of interest as the parent for processing new documents of interest in the following S204 and S206. In step S10, a document containing document elements constituting the relationship set by the user is identified, and this document becomes the initial document of interest. This initial document of interest is traced back to its ancestor, and the document of interest where no earlier document of interest can be traced back, such as the ancestor of the initial document of interest, is identified as the sample document in S208.
[0066] The above is triggered by a situation where the user (or the document analysis system) has established relationships between document elements. Figure 7 This is just one example of the process. Alternatively, a user could, for instance, instruct the document editing system to begin execution at any time. Figure 7 The processing flow is as follows. In this case, when the document editing system receives the instruction to start execution from the user, it retrieves the relationships set between document elements from the database and executes those relationships. Figure 7 The processing flow.
[0067] <Another example of sample determination>
[0068] (1) In Figure 7 Case 1 where a problem occurred during processing
[0069] There exists a document editing method that temporarily saves a blank document and appends the structure or content of document element groups to that blank document.
[0070] For example, in Figure 9In the example shown, a blank document is temporarily created as version 1 of document A1. Version 2 of document A1 is created by inputting the titles of two document elements, "1. Power Consumption Baseline" and "2. Noise Baseline," into version 1. Version 1 of document A2 is created by copying version 2, and version 2 of document A2 is created by inputting the body content of each document element into version 2. When the relationship between document elements of version 2 of document A2 is set, the sample document is determined by tracing the version group and copy history starting from version 2. When this determination uses... Figure 7 During the process, version 1 of the blank document A1 is identified as the sample document. A blank document does not contain document elements, therefore relationships cannot be established even if it is used as a sample document.
[0071] Furthermore, when preparing documents used in a company's business with a generic format that includes the company name or corporate logo in headers or footers, sometimes this generic format is temporarily copied, and document editing begins with this copy. In this method of editing documents using a generic format, the same problems as with the blank document situation described above can occur.
[0072] (2) In Figure 7 Case 2 where a problem occurred during processing
[0073] The reason why the sample document is called a sample is that multiple individual documents were created based on this sample document.
[0074] Version 1 of a document has not been created at all, while version 2 has been created multiple times, resulting in multiple individual documents based on version 2, whereas no individual documents based on version 1 exist. Therefore, in this case, version 2 is more suitable as a sample document than version 1. However, even in this situation, Figure 7 In the processing flow, version 1 was also selected as the sample document.
[0075] Thus, in Figure 7 In the processing flow, documents that are not frequently used as samples of other documents may be identified as sample documents.
[0076] (3) Examples of improved treatment
[0077] The following is an example of a method used to solve the above problems.
[0078] In this improved processing, processor 102 determines the sample document ( Figure 7In step S20), the similarity between the document of interest and its previous version or copy source is calculated. During the process of tracing a document back to its previous version or copy source, at a point in time where the calculated similarity is less than a pre-set similarity threshold, the processor 102 stops further tracing the document and selects a sample document from the documents that have been considered as documents of interest up to this point. In one example, the similarity between the document of interest and its previous version or copy source is the similarity of their document structures. In this example, the structural features of the document are considered in the similarity calculation, such as the arrangement of document elements contained in the document (e.g., columns of IDs for these document elements) or the title strings of these document elements, but the body content of these document elements is not considered. This example is based on the view that the document structure is important as a sample document. The similarity between the document structures of documents can be calculated using conventional methods. Furthermore, as another example, the body content of each document element can be considered in the above similarity calculation.
[0079] For example, in Figure 9 In the example, version 2 of document A2 has a similarity of 0.9 with its predecessor version 1, and version 1 has a similarity of 0.85 with its copy source, namely version 2 of document A1. Furthermore, version 2 of document A1 has a similarity of 0.1 with its predecessor version 1. When the similarity threshold is set to 0.6, it is possible to trace back from version 2 of document A2 (as the starting point) to version 1 of the same document, and further to version 2 of document A1 (the copy source of version 1). However, since version 2 of document A1 has a similarity of 0.1 with its predecessor version 1, it is not possible to trace back to version 1 of document A1. Therefore, the endpoint of the tracing is version 2 of document A1. The processor 102 determines a document from which multiple copies suitable for use as sample documents are created from the documents from this starting point to the endpoint as a sample document. Whether the number of created copies corresponds to multiple copies suitable for use as sample documents can be determined, for example, by whether the number of copies exceeds a copying threshold. Furthermore, as another example, the document that is copied the most times from the starting point to the ending point can be identified as the sample document.
[0080] In addition, the number of times each version of each document was copied is based on the operation history information (see reference). Figure 6 This can be calculated. Furthermore, a table can be prepared to manage the number of copies for each version of each document. In this case, if a user instructs the copying of a specific version of a document, the document editing system simply increments the copy count for that version of the document in the table by 1.
[0081] In this example, the processing flow executed by processor 102 is as follows: Figure 10 As shown. Figure 10 The process is as follows Figure 7 Here is an example of a detailed process for step S20 of the process.
[0082] In this process, processor 102 will Figure 7 In step S10 of the process, the document identified is designated as the document of interest, and the number of copies of the document of interest is calculated and stored in memory 104 (S302). Next, processor 102 determines whether a previous version or copy source document exists in the document of interest (S304). When the result of the determination in step S304 is "yes", processor 102 calculates the similarity between the document of interest and its parent, i.e., the previous version or copy source document (S306). Here, the similarity between parts other than the main text content can be calculated. Processor 102 determines whether the calculated similarity is above a preset similarity threshold (S308). If the result of the determination is "yes", the parent, i.e., the previous version or copy source document determined in S304, is designated as a new document of interest (S310). At this time, processor 102 calculates the number of copies of the new document of interest and stores it in memory 104. Moreover, processor 102 repeats the step group after step S304.
[0083] In this repetition, when either the determination result in step S304 or S308 is "no", the processor 102 determines the document with the most copies stored from the documents that are documents of interest after S302 as the sample document (S312).
[0084] exist Figure 10 In the process, the document with the most copies among the documents of interest is selected as the sample document. However, as another example, documents with a copy count exceeding a pre-set copy count threshold can also be selected as sample documents. In this case, sometimes multiple sample documents may be selected.
[0085] <Setting up a copy of the relationship as a sample>
[0086] As described above, the relationships between document elements in the sample documents are further reflected in the relationships between document elements in the document groups generated by copying these sample documents.
[0087] For example, such as Figure 11The example illustrates a sample combination 300 consisting of three sample documents: "Sample 1", "Sample 2", and "Sample 3". In this example, a reference relationship is established from the first document element of "Sample 3" to the first document element of "Sample 1", and a reference relationship is established from the second document element of "Sample 2" to the second document element of "Sample 1". A sample combination is a set of multiple sample documents that satisfies the condition that document elements within these sample documents can have relationships with document elements within the same set, but no relationships are established with document elements of documents not included in the set. In this sense, the relationships between document elements within a sample combination are "closed" within that combination.
[0088] A document editing system determines sample combinations, for example, as follows. That is, the document editing system enables... Figure 7 and Figure 10 The document editing system determines a group of documents that are considered sample documents based on its process, as these groups possess attributes that indicate they are samples. Furthermore, the system identifies a group of documents from this sample document group (i.e., sample documents) whose relationships with document elements are "closed" as described above, and records this group of documents as sample combinations. The parent group of multiple sample documents held by the document editing system may contain multiple sample combinations. The document editing system may, for example, provide the user with a list of these multiple sample combinations and receive a selection from the user of the sample combination they wish to copy.
[0089] When a user instructs the copying of sample combination 300, the document editing system generates "Document 1" as a copy of "Sample 1", "Document 2" as a copy of "Sample 2", and "Document 3" as a copy of "Sample 3". The document editing system can manage these three copied documents as a single copy document combination 310. The document editing system replicates the same relationships between document elements as in sample combination 300 among "Document 1", "Document 2", and "Document 3". That is, the document editing system establishes a reference relationship from the first document element of "Document 3" to the first document element of "Document 1", and establishes a reference relationship from the second document element of "Document 2" to the second document element of "Document 1". Subsequently, when "Document 1", "Document 2", and "Document 3" are upgraded through editing, the same relationships between document elements are maintained among the new versions of these documents. Furthermore, when the document group contained in the copy document group 310 is copied together, the document elements in the document group resulting from the copy are also copied with the same relationships as the document elements in the copy document group 310.
[0090] Figure 12 The example illustrates the processing flow executed by the processor 102 of the document editing system in this example.
[0091] In this process, processor 102 executes upon receiving an instruction to copy a certain sample combination. During this process, information about the relationships established between document elements of the sample documents included in the combination is retrieved from the database (S50). Next, processor 102 establishes the same relationships between document elements within the document group resulting from the copying of each sample document in the combination as the relationships obtained in S50 (S52). For example, consider the case where a relationship is established between a first document element in the first sample document and a second document element in the second sample document within the combination. In this case, processor 102 establishes the same relationship between the document elements corresponding to the first document element in the copy of the first sample document (i.e., the document elements of the first document) and the document elements corresponding to the second document element in the copy of the second sample document (i.e., the document elements of the second document). Information about the relationships established between document elements within the copy documents is registered in the database.
[0092] Figure 12 The process is an example of instructing the copying of a sample combination, including the copying of sample documents within that combination. However, it is not limited to this; for example, when all sample documents within a sample combination are copied over a specified time period (e.g., 1 hour), the process can be executed in the same way as when the sample combination is copied together. Figure 12 The process is as follows: when all sample documents within a sample set are copied within a specified time period, the sample set is considered to have been copied together.
[0093] The relationships defined in the sample will be reflected in the existing document.
[0094] Can be passed Figure 7 (and Figure 10 The process is reflected in the relationship between the document elements of the sample document and in the documents that already exist in the document database at the time of the reflection (hereinafter, such documents are referred to as existing documents).
[0095] exist Figure 13 The example illustrates this reflective processing flow for relationships between existing documents. Figure 13 The process also involves the use of Figure 11 and Figure 12 The above example, which has been explained, also uses the concept of sample combination.
[0096] Figure 13 The process is underway Figure 7 (and Figure 10The process begins when a sample combination is formed, establishing relationships between document elements, through a processing flow such as [processor 102]. In this flow, processor 102 determines a combination of existing documents copied from this sample combination (S60). In step S60, processor 102, for example, retrieves information from operation history information (reference […]). Figure 6 Find the copy history record of the sample combination, obtain information from the history record to determine the document group generated by the copy, and use the document group shown in the information as a combination of existing documents corresponding to the sample combination. Alternatively, when checking the operation history information for the date and time when the copying of each sample document contained in the sample combination was performed, and all sample documents contained in the sample combination were copied within a specified time period, the resulting document group can be used as a combination of existing documents corresponding to the sample combination.
[0097] Next, processor 102 executes step S62 for each determined combination of existing documents. That is, processor 102 establishes the same relationships (S62) between document elements of each document within the determined combination of existing documents as the relationships between document elements within the sample combination. For example, consider establishing a relationship between a first document element in a first sample document and a second document element in a second sample document within a sample combination, where copies of this sample combination, i.e., the first and second existing documents within the combination of existing documents, are copies of the first and second sample documents. In this case, processor 102 establishes the same relationships between the document elements in the first existing document that correspond to the first document element of the first sample document and the document elements in the second existing document that correspond to the second document element of the second sample document. Information regarding the relationships established between document elements within the documents is registered in a database.
[0098] In the embodiments described above, document elements are the elements that constitute a document. Here, there may be a larger unit of document that uses individual documents managed by the system as constituent elements. In this case, the individual documents of the former are document elements in relation to the larger unit of document of the latter. For example, when a hypertext consisting of multiple documents linked by hyperlinks is considered a larger unit of document, from the perspective of that hypertext, these multiple documents correspond to document elements.
[0099] The embodiments of the present invention described above are provided for illustrative purposes. Furthermore, these embodiments do not exhaustively encompass the present invention, nor do they limit the invention to the disclosed methods. It will be apparent to those skilled in the art that various modifications and variations will be readily understood. These embodiments were chosen and described to most readily explain the principles and applications of the invention. Thus, those skilled in the art can understand the invention through various modifications that are assumed to be optimized for specific uses of various embodiments. The scope of the invention is defined by the foregoing claims and their equivalents.
Claims
1. An information processing device, characterized in that, Including processors, The processor performs the following processing: When a relationship is established between the first document element and the second document element, a first sample document corresponding to the first document containing the first document element and a second sample document corresponding to the second document containing the second document element are determined. The relationship is established between the document elements in the first sample document that correspond to the first document element (i.e., the first sample element) and the document elements in the second sample document that correspond to the second document element (i.e., the second sample element). After establishing the relationship between the first sample element and the second sample element, when generating a first copy document (a copy of the first sample document) and a second copy document (a copy of the second sample document), the relationship is established between the document elements in the first copy document corresponding to the first sample element and the document elements in the second copy document corresponding to the second sample element. The first sample document is determined based on the number of times each document in the document group of the first document's ancestors is copied.
2. The information processing device according to claim 1, characterized in that, When the relationship is established between the first sample element and the second sample element, the processor establishes the relationship between the document element corresponding to the first sample element in an existing third document generated from the first sample document and the document element corresponding to the second sample element in an existing fourth document generated from the second sample document.
3. The information processing apparatus according to claim 1 or 2, wherein, In the process of tracing the first document back to its ancestor side one by one, the ancestor side documents whose similarity to each other's content other than the main text is below a threshold will not be selected as the first sample documents.
4. A storage medium storing a program that causes a computer to perform the following processes: When a relationship is established between the first document element and the second document element, a first sample document corresponding to the first document containing the first document element and a second sample document corresponding to the second document containing the second document element are determined. The relationship is established between the document elements in the first sample document that correspond to the first document element (i.e., the first sample element) and the document elements in the second sample document that correspond to the second document element (i.e., the second sample element). After establishing the relationship between the first sample element and the second sample element, when generating a first copy document (a copy of the first sample document) and a second copy document (a copy of the second sample document), the relationship is established between the document elements in the first copy document corresponding to the first sample element and the document elements in the second copy document corresponding to the second sample element. The first sample document is determined based on the number of times each document in the document group of the first document's ancestors is copied.
5. An information processing method, characterized in that, Includes the following steps: When a relationship is established between the first document element and the second document element, a first sample document corresponding to the first document containing the first document element and a second sample document corresponding to the second document containing the second document element are determined. The relationship is established between the document elements in the first sample document that correspond to the first document element (i.e., the first sample element) and the document elements in the second sample document that correspond to the second document element (i.e., the second sample element). After establishing the relationship between the first sample element and the second sample element, when generating a first copy document (a copy of the first sample document) and a second copy document (a copy of the second sample document), the relationship is established between the document elements in the first copy document corresponding to the first sample element and the document elements in the second copy document corresponding to the second sample element. The first sample document is determined based on the number of times each document in the document group of the first document's ancestors is copied.
6. A computer program product, characterized in that, This includes programs that cause the computer to perform the following processes: When a relationship is established between the first document element and the second document element, a first sample document corresponding to the first document containing the first document element and a second sample document corresponding to the second document containing the second document element are determined. The relationship is established between the document elements in the first sample document that correspond to the first document element (i.e., the first sample element) and the document elements in the second sample document that correspond to the second document element (i.e., the second sample element). After establishing the relationship between the first sample element and the second sample element, when generating a first copy document (a copy of the first sample document) and a second copy document (a copy of the second sample document), the relationship is established between the document elements in the first copy document corresponding to the first sample element and the document elements in the second copy document corresponding to the second sample element. The first sample document is determined based on the number of times each document in the document group of the first document's ancestors is copied.
Citation Information
Patent Citations
Document management device and method of controlling the same
JP2007323441A
Template creating method and device, document creating method and device, and document rendering method and device
CN108170656A
Automatic template generation based on previous documents
CN108369578A
Document processing device and program
JP5246371B1