Unique content determination of structured format document
The method enhances the reliability of hashing and digital sealing for structured documents by selectively processing component parts of OOXML files, thereby preventing false change detections and ensuring document integrity.
Patent Information
- Application Number
- JP2025027731
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2020-11-09
- Filing Date
- 2025-02-25
- Publication Date
- 2025-05-27
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
Conventional hashing and digital sealing methods for structured documents, such as OOXML files, can produce false indications of changes due to non-substantive differences, leading to failed authentication processes even when the document content remains unchanged.
A computer-implemented method that selects a subset of component parts of an OOXML document, excluding files like 'docProps\app.xml', 'docProps\core.xml', and 'docProps\custom.xml', processes relationship files, and generates a digest by converting component hash values into byte arrays and appending them to a canonical byte stream.
This approach improves the accuracy of hashing and digital sealing by avoiding false change indications, ensuring that only substantive changes in the document content trigger hash value changes, thus maintaining the integrity and authenticity of the document.
Smart Images

Figure 2025081618000001_ABST
Abstract
Description
Technical Field
[0001] (Field of the Invention) The present disclosure relates to analyzing the content of structured format documents.
Background Art
[0002] (Background) This background section is generally provided for the purpose of explaining the context of the present disclosure. The works of the inventor(s) currently named are not expressly or implicitly admitted as prior art to the present disclosure to the extent that the works are described in this background section and in aspects of the description that cannot otherwise be qualified as prior art at the time of filing.
[0003] Digital files are used for different purposes across many industries. Some industry regulations such as HIPAA, SOX, GDPR, ISO 9000, etc. impose restrictions on digital files that must be protected for security, authenticated, or audited.
[0004] Generally, hashing is used to assess the content of files. Hash algorithms use one-way functions to map the content of a file to a hash value (also called a digest value). Any change to any data within the file causes a different hash value. Related hash algorithms include, but are not limited to, Message Digest 5 (MD5), Rivest Shamir Adleman (RSA), Secure Hash Algorithm (SHA), Script, Ethereum Hash, and Cyclic Redundancy Check (CRC).
[0005] Hashing can be used to determine whether a file has been changed or updated by calculating the current hash value of the file and comparing it to the previous hash value for that file. If the hash values are different, the contents of the digital file have changed since the last hash value was calculated. Matching hash values confirm a high degree of certainty that the contents of the digital file have not been changed.
[0006] By using hashing, a digital seal can be created for a digital file. The digital seal can include a hash value stored within a blockchain ledger, such as a Bitcoin or Ethereum ledger. The hash value can be read from the digital seal and verified to match the contents of the file at the time the digital seal was created. The digital seal can be managed using blockchain technology to ensure the authenticity of the digital seal. Since data errors and intentional modifications to the file will result in different hash values, a user or system can prove the integrity and authenticity of the data of a digital file using this approach.
[0007] Microsoft Office TM Files (e.g., Word TM , Excel TM , PowerPoint TM etc.) Structured digital files, for example, contain various information including content, formatting, metadata information. These files can be stored in the Open Office XML (OOXML) format. Metadata information such as the last access date or print date of the file can be updated or changed when the file is accessed even though the substantial information within the file has not changed.
[0008] As determined by the inventors of the present application, using the entire content of a structured document that includes information not related to the integrity or authenticity of the data may limit the effectiveness of hashing and digital sealing. In addition, sequential data stored within a digital file can actually cause variations in the hashing results even when the underlying data within the file is identical. More specifically, differences between documents may not indicate differences in the content or presentation of the documents. These non-substantive differences can cause conventional sealing and authentication processes to fail after the document has been opened by an Office document editor, even when the user has not made any changes to the document.
Summary of the Invention
Means for Solving the Problems
[0009] (Overview) An object of the present invention is to improve the usefulness of hashing and digital sealing of structured documents, such as OOXML, by hashing only the content of the structured document in order to avoid false indications of changes between copies of a particular document. This object is solved by the subject matter according to the claims described below.
[0010] Embodiments of the present invention are described in the claims, the following description, and the drawings.
[0011] A computer-implemented method for generating a digest for a structured document is provided. The method is to select a subset of a plurality of component parts of an OOXML document, wherein at least one of the selected subset of component parts is an XML file, and selecting the subset includes excluding files named "docProps\app.xml", "docProps\core.xml", and "docProps\custom.xml", ordering the selected subset of component parts, and processing relationship files from the selected subset of component parts. Processing the relationship files includes removing at least one relationship entry that references a component part not included in the selected subset and sorting the relationships by identifier value. The computer-implemented method further includes generating a component hash value for each selected subset of component parts, converting each component hash value into a byte array, and appending the byte array to a canonical byte stream.
[0012] A computer-implemented method for generating a digest for a structured document is provided to include selecting a subset of a plurality of component parts of the structured document, ordering the subset of component parts, generating a component hash value for each subset of component parts, converting each component hash value into a byte array and appending the byte array to a canonical byte stream. In some embodiments, the structured document is an OOXML file, which is a ZIP archive, and each of the plurality of component parts is a file. In some embodiments, the method includes calculating an overall hash value for the canonical byte stream. In some embodiments, the selected subset of component parts does not include any of the following dates, namely, the date the structured document was created, the date the structured document was last printed, the date the structured document was last edited, and the date the structured document was last opened. In some embodiments, the selected subset of component parts includes a relationship file named "_rels\.rels", and the relationship file includes relationships of type "officeDocument". In some embodiments, the selected subset of component parts does not include a document named "custom.xml". In some embodiments, the computer-implemented method includes storing the overall hash value for the canonical byte stream in a record, selecting a second subset of a second plurality of component parts, ordering the second plurality of component parts of the second structured document, generating a second component hash value for each second subset of component parts, converting each second hash value into a second byte array and appending the second byte array to a second canonical byte stream, calculating a second overall hash value for the second canonical byte stream, and determining that the structured document is substantially the same as the second structured document if the second hash value for the second canonical byte stream matches the overall hash value in the record.
[0013] A non-transitory computer-readable medium containing executable software instructions is provided, and the executable software instructions, when executed, perform selecting a subset of a plurality of component parts of a structured document, ordering the subset of component parts, generating a component hash value for each subset of component parts, converting each component hash value into a byte array, and appending the byte array to a canonical byte stream. In some embodiments, the structured document is an OOXML file, which is a ZIP archive, and each of the plurality of component parts is a file. In some embodiments, the medium further contains instructions that, when executed, calculate an overall hash value for the canonical byte stream. In some embodiments, the selected subset of component parts does not include any of the following dates, namely, the date the structured document was created, the date the structured document was last printed, the date the structured document was last edited, or the date the structured document was last opened. In some embodiments, the selected subset of component parts includes a relationship file named "_rels\.rels", and the relationship file includes relationships of type "officeDocument". In some embodiments, the selected subset of component parts does not include a document named "custom.xml".In some embodiments, the medium includes instructions that, when executed, cause the following operations to be performed: store a global hash value for a canonical byte stream in a record; select a second subset of a second plurality of component parts; order the second plurality of component parts of a second structured document; generate a second component hash value for each second subset of component parts; convert each second hash value to a second byte array and append the second byte array to a second canonical byte stream; calculate a second global hash value for the second canonical byte stream; and determine that the structured document is substantially the same as the second structured document if the second hash value for the second canonical byte stream matches the global hash value in the record. This specification also provides, for example, the following items. (Item 1) A computer-implemented method for generating a digest for a structured document, selecting a subset of a plurality of component parts of an OOXML document, wherein at least one of the selected subset of component parts is an XML file, and selecting the subset excludes files named "docProps\app.xml", "docProps\core.xml", and "docProps\custom.xml", ordering the selected subset of component parts, processing relationship files from the selected subset of component parts, namely, removing at least one relationship entry that references component parts not included in the selected subset, sorting the relationships by identifier value, and generating a component hash value for each of the selected subset of component parts, converting each component hash value into a byte array and appending the byte array to a canonical byte stream, a computer-implemented method including the above. (Item 2) A computer-implemented method for generating a digest for a structured document, selecting a subset of a plurality of component parts of a structured document, ordering the subset of component parts, generating a component hash value for each of the subset of component parts, converting each component hash value into a byte array and appending the byte array to a canonical byte stream, a computer-implemented method including the above. (Item 3) The method according to item 1, wherein the structured document is an OOXML file that is a ZIP archive, and each of the plurality of component parts is a file. (Item 4) The method according to item 1, further including calculating an overall hash value for the canonical byte stream. (Item 5) The selected subset of component parts is the following dates, namely, the date when the structured document was created, the date when the structured document was last printed, the date when the structured document was last edited, and The method according to item 1, not including any of the dates on which the structured document was last opened. (Item 6) The method according to item 2, wherein the selected subset of component parts includes a relationship file named "_rels\.rels", and the relationship file includes relationships of type "officeDocument". (Item 7) The method according to item 1, wherein the selected subset of component parts does not include a document named "custom.xml". (Item 8) Storing the overall hash value for the canonical byte stream in a record; Selecting a second subset of the second plurality of component parts; Ordering the second plurality of component parts of the second structured document; Generating a second component hash value for each of the second subset of component parts; Converting each second hash value into a second byte array and appending the second byte array to a second canonical byte stream; Calculating a second overall hash value for the second canonical byte stream; Determining that the structured document is substantially the same as the second structured document if the second hash value for the second canonical byte stream matches the overall hash value in the record The method according to item 1, further comprising. (Item 9) A non-transitory computer-readable medium including executable software instructions that, when executed, Select a subset of a plurality of component parts of a structured document; Order the subset of component parts; Generate a component hash value for each of the subset of component parts; Convert each component hash value into a byte array and append the byte array to a canonical byte stream A medium for performing. (Item 10) The medium according to item 9, wherein the structured document is an OOXML file in a ZIP archive, and each of the plurality of component parts is a file. (Item 11) The medium according to item 9, further including instructions for calculating an overall hash value for the canonical byte stream when executed. (Item 12) The selected subset of component parts does not include any of the following dates, namely, the date on which the structured document was created, the date on which the structured document was last printed, the date on which the structured document was last edited, and the date on which the structured document was last opened the medium according to item 9. (Item 13) The selected subset of component parts includes a relationship file named "_rels\.rels", and the relationship file includes relationships of type "officeDocument", the medium according to item 10. (Item 14) The selected subset of component parts does not include a document named "custom.xml", the medium according to item 9. (Item 15) The medium further comprises instructions which, when executed, store the overall hash value for the canonical byte stream in a record; select a second subset of the second plurality of component parts; order the second plurality of component parts of the second structured document; generate a second component hash value for each of the second subset of component parts; convert each second hash value into a second byte array and append the second byte array to a second canonical byte stream; calculate a second overall hash value for the second canonical byte stream; and determine that the structured document is substantially the same as the second structured document if the second hash value for the second canonical byte stream matches the overall hash value in the record the medium according to item 9.
Brief Description of the Drawings
[0014]
Figure 1
[0015]
Figure 2
[0016]
Figure 3A
Figure 3B
Figure 3C
[0017]
Figure 4
[0018]
Figure 5
[0019]
Figure 6
DETAILED DESCRIPTION OF THE INVENTION
[0020] (Description) As discussed above, the teachings herein relate to an improved method for hashing and sealing structured digital files. Digital files can be represented in different file formats. In some file formats, a document is a package file that includes a set of files representing the content, formatting, and other aspects of that document. For example, Microsoft Office TM documents are stored in the Office Open XML (OOXML) file format. The OOXML file format packages a number of XML files that include the content of a Microsoft Office TM document. These files can be stored within an Open Packaging Conventions (OPC) package that packages together XML and data files for the content, format, and metadata of a Microsoft Office document, all packaged within a ZIP archive file. Exemplary Microsoft Office TM documents that can be processed using the present disclosure include files with DOCX, XLSX, and PPTX file extensions, stored in the OOXML format. Another standard open document format is the Open Document Format (ODF). An ODF document is a ZIP archive package of XML files, as will be explained in more detail below. Other types of structured digital files may similarly be suitable.
[0021] Many structured documents do not require a specific order of internal files and may allow the content to appear in a different order within the internal files without modifying the appearance or content of the Office document. For example, an Office document may define a list of styles such as headings, body text, Heading 1, Heading 2, etc. The definitions of these styles can be rearranged without changing the appearance or content of the document. In some embodiments, the style definitions may be part of a larger structured file (e.g., an XML file) and may be present at the beginning, middle, or end of that file without changing the appearance or content of the document. One application (e.g., Microsoft Office Word TM ) may save style information in a certain order and within a certain location of a Word processing document. A structured document may be stored within a content management system (CMS) that can modify the document metadata. One such CMS system is SharePoint TM . When a user opens a document within the CMS, the CMS can record the date the document was opened by modifying the metadata, for example, by updating the modification date of a list item, without modifying the document. More specifically, the document remains in the same state for the user who views or prints the document, but the hash of the entire document does not match the hash of the original document.
[0022] According to some embodiments of the present disclosure, one or more components may be used to prepare or process digital files for hashing and digital sealing. For example, in some embodiments, a canonical formatter may be used to array a package of files representing a digital file in canonical form for hashing. The canonical formatter generates a canonical stream of the digital file for hashing. A package part selector is used to select one or more parts of the digital file for hashing. A relationship information processor processes relationship information according to the information contained in the digital file or the information identified by a universal resource identifier in the file.
[0023] FIG. 1 is a system for processing a document according to an embodiment of the present disclosure. System 100 includes a central processing unit (CPU) 102, a random access memory (RAM) 104, and a disk 106. Disk 106 is a non-transitory computer-readable medium that stores document 108 and program 110. Document 108 is a document that will be digitally sealed and is formatted, for example, according to an XML-based open document standard. Program 110 is a software program for implementing the methods described in the present disclosure. During operation, CPU 102 reads program 110 into RAM 104 and executes program 110. Program 110 reads document 108 into RAM, implements the methods described in the present disclosure, and generates a canonical stream 112 of a part of the content of document 108. Subsequent execution of program 110 on document 108' generates canonical stream 112'. A comparison of canonical streams 112 and 112' (e.g., by hashing each and comparing their hash values) determines whether the relevant content of document 108 matches that of document 108'.
[0024] Program 110 includes a canonical conversion 113. The canonical conversion 113 arranges a package of files representing digital files in a canonical order and controls a sequence of hash values that will be combined into the final digest of the document. In some embodiments, the canonical conversion 113 may apply rules for arranging files within a document in an order independent of the order in which the files are stored within the package. The canonical formatter does not modify the contents of the digital file or the package representing it, but instead generates canonical-form data in RAM 104 for hashing and collating the contents of the files in a manner independent of the order of the contents of the package.
[0025] The Microsoft OPC Digital Signature Framework illustrates an exemplary canonical formatter for OPC packages. The digital signature framework standardizes the creation and verification of hash values for OPC packages. The framework defines the canonical form by universal resource identifiers within the OPC package. For example, the OPC Digital Signature Framework identifies a manifest file identified by a "manifest" tag within the OPC package. Exemplary manifest tags are <manifest xmlns:opc=""http: / / schemas.openxmlformats.org / package / 2006 / digital-signature”">It is. The manifest file is a collection of all package parts and package relationships. The content of the manifest file can be used by the OPC digital signature framework to array the OPC package in a canonical form for hashing. In some embodiments, the manifest file is signed using the SignedXml API, and the signature is saved within the OPC package together with the manifest file. The OPC digital signature framework supports the W3C XML digital signature standard (which can be found at https: / / www.w3.org / TR / xmldsig-core1 / ) and the processing of the content of the OPC package.
[0026] In some embodiments, program 110 includes a package part selector 114. The package part selector 114 selects parts of the package (e.g., the documents to be hashed) that are to be included within the hash generation. In some embodiments, the package part selector 114 reads a file external to the document, and that file defines parts of the document that are to be included in and / or excluded from the hash process. In some embodiments, the package part selector 114 determines the information to be used or excluded for hashing based on a predetermined set of criteria including attributes or relationships defined within the document.
[0027] In some embodiments, the selection criteria are explicitly included within the digital file to be digested. One approach is to include a flag within a manifest file that indicates files that should (or should not) be included in the digest process. In some embodiments, the selection criteria are defined by reference to an external source such as a standardized canonical form. One approach is to define the selection criteria in a universal resource identifier embedded within a file in the document. In some embodiments, the selection criteria are defined external to the file to be digested. An exemplary specification for digesting a contract includes only content files ("document.xml" and "image1.svg") because formatting information is not relevant to the gist of the legal contract. An exemplary specification for digesting a published work excludes only metadata because formatting often extends and sometimes defines a published work. An exemplary specification for digesting a computer-aided design (CAD) file excludes simulation inputs and results but includes structural design elements. Another exemplary specification for digesting a CAD design excludes corporate logos, copyright notices, and other non-engineering information. Exemplary specifications for canonical forms include forms for different types of documents including word processing documents, spreadsheets, and presentations.
[0028] In some embodiments, the package part selector 114 is based on the type of reference regarding the information within the package, Microsoft Office TM Select parts of the OPC package content related to the document. In one approach, the package part selector 114 selects all package parts that identify the content, style, and layout of the document as represented within the OPC package. Thus, the package part selector 114 excludes certain metadata such as the date last accessed of a digital file that is considered not related to the integrity or authenticity of the file. As an example, the package part selector 114 excludes certain information within the OPC package such as custom parts and references to custom XML properties within the OPC package. In some embodiments, the package part selector 114 selects parts of the OPC package based on information described within the manifest file. The package part selector 114 selects information where ContentType == "application / vnd.openxmlformats-package.relationships+xml" or "application / vnd.openxmlformats-package.digital-signature-xmlsignature+xml".
[0029] In some embodiments, to further the objectives of the present disclosure, the relationship information processor 116 may rearrange the blocks of information in a canonical form. In some embodiments, the relationship information processor 116 identifies structured data within a document that is equivalent when rearranged. The relationship information processor 116 rearranges the data in a canonical form to maintain a consistent order of information that will be fed into a hash function. In some embodiments, the relationship information processor 116 sorts the contents of a manifest file according to the uniform resource locator values within the file. In some embodiments, the relationship information processor 116 sorts the contents of "styles.xml" according to the name of the style. In some embodiments, the relationship information processor 116 generates a hash value of the contents of each element in an unsorted XML file and sorts the elements based on the generated hash values.
[0030] In some embodiments, a structured document includes relationship information that defines the details of the relationships between files and data within a package representing the document. The relationship information provides a roadmap for a viewer / editor of the document to navigate the files within the document and defines the relationships between them. In some embodiments, the relationship information of a package is defined in a markup language such as XML. For example, an OPC package includes relationship information that defines the relationships between the various parts of the package. The relationship information within an OPC package is <relationships>Stored or identified within the manifest file using tags. Exemplary tags are <relations xmlns=""http: / / schemas.openxmlformats.org / package / 2006 / relationships”">It is as follows.
[0031] In some embodiments, the relationship information is generated based on an ordered file stored within the package. If Document A and Document B contain the same files but are in a different order, the hash on the content of Document A will be different from that of Document B, even if both documents appear the same when viewed. The relationship information processor 116 thus sorts the content of the manifest file before the file is fed into the hash function, providing a consistent hashing of relationship information independent of the ordered file stored within the package.
[0032] FIG. 2 is a Word processing document according to an embodiment of the present disclosure. Document 200 is shown within a word processing program such as Microsoft Word TM in print preview mode. Document 200 may include content such as a title 202, a bulleted text element 204, an image 206, and a table 208. Document 200 may be the first page of a multi-page document. Document 200 may also include other data (e.g., metadata) that is invisible when viewing or editing the document. For example, the metadata may record the time the document was created, the time it was last edited, and the time it was last opened.
[0033] FIG. 3a illustrates the internal structure of a document according to an embodiment of the present disclosure. Document 300 is in the Open Office Extensible Markup Language (OOXML) format such as that created by Microsoft Word TM and is stored within disk 106. Document 300 is stored as a ZIP archive file container containing files organized within subdirectories. Document 300 includes subdirectories 302-312 and file 319. This collection of files can be read into RAM 104 and viewed by a word processing program as shown in FIG. 2.
[0034] A file 319 named "[Content_Types].xml" defines the types of content used within the document 300 and references other files 319 within the document 300. The folder 302 is named "docProps" and contains metadata about the document 300. For example, a file 319 named "app.xml" identifies the application used to create the document as "Microsoft Office Word" version 16.0000. A file 319 named "core.xml" enumerates document properties including the subject, author, creation date, and date last modified of the document regarding the document 300. In some embodiments, the folder 302 may also include a file named "custom.xml", which contains additional metadata in addition to the standard metadata fields. The package part selector excludes the files within the folder 302 that include "app.xml", "core.xml", and "custom.xml". The folder 304 is named "word" and contains six files of text content and formatting that make up the document 300. A file 319 named "document.xml" contains text regarding the document 300 including the title, text of the bulleted list, and text of the table. A file 319 named "fontTable.xml" enumerates the fonts used within the document 300, and a file 319 named "numbering.xml" describes the styles of the bullet symbols used within the bulleted list. A file 319 named "settings.xml" defines some settings regarding the word processing application such as the formatting of hyperlinks as shown in the document 200. A file 319 named "styles.xml" enumerates the defined formatting styles used within the document. A file 319 named "webSettings.xml" references various XML schemas for converting the document to HTML for web viewing.
[0035] Folder 306 contains two image files. The image files 319 are named "image1.png" and "image2.svg" and are two different image files for the same airplane icon 206. Folder 308 contains a single file 319 named "theme1.xml", which defines a document theme used by Microsoft Word to coordinate colors and styles.
[0036] The files 319 named "[Content_Types].xml" and "document.xml.rels" are manifests that enumerate and classify other files within the document 300.
[0037] Figure 3b illustrates the internal structure of a document according to an embodiment of the present disclosure. Document 350 is in disk 106 as an Open Document Format (ODF) file. Document 350 is stored as a ZIP archive file container including files and subfolders. Document 350 includes files 352 - 360 and subfolder 351. This collection of files can be read into and viewed in RAM 104 by a Word processing program as shown in FIG. 2. File 352 is named "meta.xml" and includes metadata about document 350 such as the name of the document's creator, the date last printed, the name of the software used to create the file, etc. File 354 is named "style.xml" and enumerates the defined formatting styles used within the document. File 356 is named "setting.xml" and includes settings related to the Word processing program. File 358 is named "content.xml" and includes the content of document 350 such as the text and formatting shown in FIG. 2. File 360 is named "manifest.xml" and enumerates other defined types of content used within document 350 and references the other files 352 - 360 within document 350. The package part selector excludes file 352.
[0038] Figure 3c illustrates the internal structure of a document according to an embodiment of the present disclosure. Document 370 is stored in disk 106 as a structured document including metadata 372 and content 374. Metadata 372 includes metadata about document 370 such as the date last printed. The package part selector excludes metadata 372.
[0039] Figure 4 is an OOXML document content file according to an embodiment of the present disclosure. The content file 400 is an XML-formatted document including elements 402 to 428. Element 402 is a container for the content type found within document 300. Elements 404 to 412 define default content based on the file extension. For example, element 404 defines a content type 404b named "image / png" for files within a document with a file extension 404a named "png", which is a Portable Network Graphic image file. Elements 412 to 428 define overrides for the default content type. For example, element 412 defines a content type 412b of "document.main+xml" (abbreviated in the figure for clarity) for a specific file identified as part 412a named " / word / document.xml" (the non-abbreviated content type 412a is "application / vnd.openxmlformats-officedocument.wordprocessingml.document.main+xml"). This content type means that part 412a is the main content of a Word processing document. The package part selector excludes the files referenced by 426 and 428.
[0040] Figure 5 is a computer-implemented method for generating a digest for a structured document according to an embodiment of the present disclosure. Method 500 includes six blocks of functionality for performing on a document such as document 108. In block 502, the CPU 102 reads a structured document that is a ZIP archive of files and identifies files to be excluded in the digest process. As discussed above, files within the "docProps" folder, namely app.xml, core.xml, and (if present) custom.xml, are excluded.
[0041] In block 504, the CPU 102 orders the set of non-excluded files in a lexicographic order. In one embodiment, the CPU 102 determines the set of files by reading the manifest file included in the ZIP archive. In another embodiment, the CPU 102 determines the set of files by listing the files in the ZIP archive. The CPU 102 sorts the file references alphabetically and generates a sorted list. In some embodiments, the list of files in the document is shown in FIG. 3a. In some embodiments, the CPU 102 sorts the file references alphabetically by the path (or URL) as follows. · _rels / * / .rels · word\_rels\document.xml.rels · word\document.xml · word\fontTable.xml · word\media\image1.png · word\media\image2.svg · word\numbering.xml · word\settings.xml · word\styles.xml · word\theme\theme1.xml · word\webSettings.xml
[0042] Relationship files (e.g., those with a ".RELS" extension) may also be preprocessed using relationship transformation algorithms such as those defined in section 12.2.4.26 of Office Open XML, Part 2: Open Packaging Conventions (December 2006) (available at https: / / www.ecma-international.org / publications-and-standards / standards / ecma-376 / ). This processing includes steps to remove ignorable content, versioning instructions, references to missing parts, etc. This processing also includes sorting relationships by their Id values. This processing also omits any relationships that reference parts that are omitted by the package part selector from the relationship file. Relationship entries within the relationship file include a relationship identifier and a target file name, as in the following examples. <relationship id=""rId2”" type=""http: / / schemas.openxmlformats.org / officeDocument / 2006 / relationships / styles”Target="styles.xml” / ">
[0043] In block 506, the CPU 102 generates a digest value for hashing the files in the file list. The CPU 102 reads the content of the current file as a byte stream and feeds the data into a defined hash function. The output of the hash function is the digest.
[0044] In block 508, the CPU 102 converts the digest from block 506 into a byte array stored in the RAM 104. This byte array is a suitable input for the hash function. Blocks 506 and 508 are repeated for each file.
[0045] In block 510, the CPU 102 associates the byte array from block 508 with a single canonical stream. If no additional files are digested, the CPU 102 returns to block 506.
[0046] In block 512, the CPU 102 returns the canonical stream to calculate the final hash. The CPU 102 reads the data bytes from the canonical stream and feeds the data into the hash function. The resulting digest is stored in the RAM 104 and can be used for comparison with previously existing digest values or for incorporation into a new digital seal.
[0047] In some embodiments, blocks 502, 504, and 506 are rearranged without affecting the result of method 500.
[0048] FIG. 6 is a computer-implemented method for determining a digest for a structured document according to an embodiment of the present disclosure. Method 600 includes six functional blocks for implementation on a document such as document 370. The XML container need not include a unique identifier or other useful key for sorting. Two containers can have the same name and different content, for example, the containers can identify a bulleted list that is distinguishable from a bulleted list container that differs only by the content and position within the document.
[0049] In block 602, the CPU 102 reads a structured document comprising a container of elements. The CPU 102 generates a digest value for each container and stores that digest in the RAM 104 as a byte array associated with its respective container.
[0050] In block 604, the CPU 102 sorts the sections of the structured document in canonical order. The CPU 102 sorts the containers by their corresponding digest values.
[0051] In block 606, the CPU 102 identifies containers that are to be excluded from the final digest. In one example, the CPU 102 excludes metadata containers that include document properties.
[0052] In block 608, the CPU 102 associates the byte arrays from block 602 into a single canonical stream.
[0053] In block 610, the CPU 102 returns a canonical stream in order to compute a final hash on the relevant information. The CPU 102 reads bytes of data from the canonical stream and feeds that data into a defined hash function (which need not be the same as the one within block 506). The resulting digest is stored in the RAM 104 and can be used for comparison with a previously existing digest value or for incorporation into a new digital seal.
[0054] Those skilled in the art will understand that in the disclosed embodiments, several modifications can be made without departing from the spirit and scope of the present invention. For example, while the disclosed embodiments relate to the Open XML Microsoft Office file format, the methods described above can be used to seal and verify other formats and types of digital files. The methods described above can be implemented in a digital circuitry, instructions for execution by a processor, or any suitable combination thereof.
[0055] Specific references to components, process steps, and other elements are not intended to be limiting. Further, it should be understood that like parts will have the same or similar reference numbers when the figures are referenced alternately. Additionally, note that the figures are schematic and provided as a guide to those skilled in the art and are not necessarily drawn to scale. Rather, the various drawing scales, aspect ratios, and number of components shown in the figures can be intentionally distorted to make a particular feature or relationship more understandable.< / relationship> < / relations> < / relationships> < / manifest>
Claims
[Claim 1] The invention described in this specification.
Citation Information
Patent Citations
Method and apparatus and performing electronic signature to document having structure
JP2002229448A
Authentication method, apparatus to be authenticated, authentication apparatus and program
JP2006101284A
Wireless medical data communication system and method
JP2012011204A