An electronic file authenticity verification method and system

Through the multi-dimensional electronic archive authenticity verification method, combined with hashing algorithm, signature and seal verification, the problem of difficult to guarantee the authenticity of electronic archives in the existing technology is solved, efficient and reliable electronic archive authenticity verification is achieved, and the security and credibility of the document are enhanced.

CN118821085BActive Publication Date: 2025-08-01HUNAN LINGZHONG ARCHIVES MANAGEMENT CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202410916599.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-07-09
Publication Date
2025-08-01
Estimated Expiration
2044-07-09

AI Technical Summary

Technical Problem

The existing electronic archive authenticity verification technology is difficult to ensure the authenticity, integrity and security of electronic archives when facing complex security threats and diverse application scenarios.

Method used

A multi-dimensional verification method is adopted, including solidified information validity inspection, content authenticity inspection and association consistency inspection, combined with electronic signature, timestamp inspection and data correlation verification, and hashing algorithm, signature verification, seal verification and multimedia data verification, ensuring the comprehensive authenticity and security of electronic files.

Benefits of technology

It significantly improves the credibility and security of electronic files, prevents document tampering and forgery, enhances the overall security of documents, and improves the efficiency and accuracy of the verification process.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118821085B_ABST
    Figure CN118821085B_ABST
Patent Text Reader

Abstract

The present invention relates to the technical field of electronic document management, and more particularly to a method and system for verifying the authenticity of electronic archives. The method comprises the following steps: obtaining electronic document data and corresponding electronic document metadata; performing a fixed information validity check based on the electronic document data and the electronic document metadata to obtain fixed information validity data; performing a content authenticity check based on the electronic document data and the electronic document metadata to obtain content authenticity check data; performing an association consistency check based on the electronic document data and the electronic document metadata to obtain association consistency check data; and performing an electronic archive authenticity verification operation based on the fixed information validity data, the content authenticity check data, and the association consistency check data. The present invention improves the security, reliability, and overall protection capabilities of electronic archive management.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of electronic document management, and in particular to a method and system for verifying the authenticity of an electronic file. Background Art

[0002] With the rapid development of information technology, electronic archives have gradually replaced traditional paper archives and become the primary form of recording and storing important information in modern society. Electronic archives are widely used in various fields, including administrative management, commercial transactions, legal documents, and medical records. However, with the widespread use of electronic archives, ensuring their authenticity, integrity, and security has become a major technical challenge.

[0003] The authenticity of electronic records is directly related to the credibility and legitimacy of information. For example, the authenticity of important documents such as contracts, legal documents, and financial statements is crucial; any tampering or forgery can result in serious legal and economic consequences. Therefore, ensuring the authenticity of electronic records is a key issue in information security. Existing electronic record authenticity verification technologies have achieved some success in ensuring document integrity and signatory identity, but they still face shortcomings when faced with complex security threats and diverse application scenarios. Summary of the Invention

[0004] In order to solve the above technical problems, the present invention proposes an electronic file authenticity verification method and system to solve at least one of the above technical problems.

[0005] This application provides a method for verifying the authenticity of an electronic file, the method comprising:

[0006] S1. Obtain electronic document data and corresponding electronic document metadata;

[0007] S2. Performing a validity check on the fixed information based on the electronic document data and the electronic document metadata to obtain fixed information validity data;

[0008] S3. Perform content authenticity detection based on the electronic document data and the electronic document metadata to obtain content authenticity detection data;

[0009] S4. Perform association consistency detection based on the electronic document data and the electronic document metadata to obtain association consistency detection data;

[0010] S5. Perform electronic archive authenticity verification based on the solidified information validity data, content authenticity detection data, and associated consistency detection data.

[0011] In the present invention, this method ensures the overall authenticity of electronic files through multi-dimensional verification steps. The systematic verification process makes the verification process more organized and efficient. By solidifying information validity checks, content authenticity detection, and associated consistency detection, the credibility of electronic files is significantly improved. The introduction of signature verification, timestamp checks, and data relevance verification effectively prevents documents from being tampered with and forged, enhancing the security of electronic files.

[0012] Optionally, S1 includes:

[0013] Initialize the data acquisition interface, extract electronic document data and metadata from the electronic file management system to obtain primary electronic document data and corresponding primary electronic document metadata;

[0014] Parse the primary electronic document metadata to obtain electronic document metadata;

[0015] Parse the primary electronic document data to obtain electronic document data.

[0016] In the present invention, the automated data acquisition interface improves the efficiency of extracting data from the EDMS, ensuring data integrity and consistency. Parsing and normalizing the primary data improves the quality and usability of the data, providing a reliable data basis for subsequent verification. The automated processing steps reduce human intervention, lower the risk of errors, and improve the efficiency and accuracy of the overall process. The structured metadata and document data facilitate subsequent processing, analysis, and verification, improving the operability and analysis efficiency of the data.

[0017] Optionally, S2 includes:

[0018] Generate a digest based on the electronic document data to obtain electronic document digest data;

[0019] Compare the electronic document digest data with the meta-digest data in the electronic document metadata to obtain digest comparison data;

[0020] Extract the electronic signature from the electronic document data to obtain electronic signature data;

[0021] Verify the electronic signature data to obtain signature verification data;

[0022] Extract the electronic seal from the electronic document data to obtain electronic seal data;

[0023] Verify the electronic seal data to obtain seal verification data;

[0024] Integrate the digest comparison data, signature verification data, and seal verification data to obtain solidified information validity data.

[0025] In the present invention, through abstract comparison, signature verification, and seal verification, a multi-level verification mechanism is provided to ensure the authenticity and integrity of the document. The abstract generation and comparison are fast, with low resource consumption, and the signature and seal verification ensure security and legality. The overall verification process is efficient. Integrating various verification results provides comprehensive verification data to ensure the reliability and security of the document in all aspects.

[0026] Optionally, S3 includes:

[0027] Extract the attributes of the electronic document data to obtain the electronic document attribute data;

[0028] Perform attribute comparison based on the electronic document attribute data and the meta-attribute data in the electronic document metadata to obtain the attribute comparison data;

[0029] Perform content integrity detection based on the electronic document data and the electronic document metadata to obtain the content integrity detection data;

[0030] Generate a checksum for the electronic document data to obtain the electronic document checksum data;

[0031] Perform checksum comparison based on the electronic document checksum and the meta-checksum data in the electronic document metadata to obtain the checksum comparison data;

[0032] Perform consistency processing on the content change records of the electronic document data to obtain the change record consistency data;

[0033] Verify the multimedia data of the electronic document data to obtain the multimedia verification data;

[0034] Integrate the attribute comparison data, content integrity detection data, checksum comparison data, change record consistency data, and multimedia verification data to obtain the content authenticity detection data.

[0035] In the present invention, by generating a document summary, it is possible to quickly detect whether the content of the document has been tampered with, ensuring the integrity of the content. The summary data is concise and unique, enabling efficient content comparison. By comparing the summaries, it is possible to quickly discover whether the content of the document has been tampered with or modified, ensuring the authenticity of the document. The summary comparison is fast and has low resource consumption, making it suitable for the rapid verification of large-scale documents. Extracting electronic signature data is the basis for subsequent signature verification, ensuring the complete extraction of signature information. Through signature verification, the identity of the signer is confirmed, preventing impersonation or forgery. Signature verification can ensure that the document has not been modified since it was signed, guaranteeing the integrity of the data. Through seal verification, the legality and effectiveness of the electronic seal are confirmed, preventing forgery and tampering. Electronic seal verification increases the legal effect and credibility of the document. Through multi-level verification, it is ensured that the document has not been tampered with, and the signature and seal are both legal and valid, enhancing the security of the document. The integration of various verification data improves the reliability of the overall verification, ensuring a high level of credibility for the document.

[0036] Optionally, S4 includes:

[0037] Extract a directory structure based on the electronic document data to obtain directory structure data;

[0038] Perform directory content association processing based on the electronic document data and the directory structure data to obtain directory content association data;

[0039] Perform directory verification on the directory content association data to obtain directory verification data;

[0040] Generate a metadata format based on the electronic document data to obtain metadata format data;

[0041] Perform metadata consistency verification based on the metadata format data and the electronic document metadata to obtain metadata consistency verification data;

[0042] Generate a mapping relationship based on the directory structure data, the electronic document data, and the electronic document metadata to obtain mapping relationship data;

[0043] Perform relevance verification based on the mapping relationship data to obtain relevance verification data;

[0044] Integrate the directory verification data, the metadata consistency verification data, and the relevance verification data to obtain associated consistency detection data.

[0045] In the present invention, the table of contents structure is extracted to help users quickly understand the content and structure of the document, improving the readability of the document. The extraction of the table of contents structure facilitates navigation in the document, enhancing the user experience. The association relationship between the table of contents items and the content is established to help users quickly locate specific content in the document. The processing of the association between the table of contents and the content makes it more efficient and accurate to search for document content. The relevance between the table of contents and the content is verified to ensure that the content pointed to by the table of contents items is accurate and error-free, avoiding the situation where the table of contents does not match the content. The verification of the table of contents improves the overall quality and accuracy of the document. Standardized metadata formats are generated to ensure the consistency of the metadata formats, facilitating subsequent processing and analysis. The standardized metadata formats contribute to the management and classification of electronic documents. Through the verification of metadata consistency, the accuracy and integrity of the metadata are ensured. The metadata is verified to reduce errors caused by inconsistent formats. Through relevance verification, the relevance and consistency among the table of contents, content, and metadata are ensured. Relevance verification enhances the credibility of the document, ensuring the integrity and consistency of the document data.

[0046] Optionally, the extraction of the electronic signature includes:

[0047] Performing fixed chunking on the electronic document data to obtain electronic document chunk data;

[0048] Performing standardization processing on the electronic document chunk data to obtain electronically processed document chunk data;

[0049] Performing compression processing on the electronically processed document chunk data to obtain compressed electronically processed document chunk data;

[0050] Performing intermediate hash generation based on the compressed electronically processed document chunk data to obtain intermediate hash data;

[0051] Calculating a hash value based on the intermediate hash data to obtain electronic signature data.

[0052] In the present invention, the chunking processing and compression processing improve the processing efficiency of large documents, reducing the processing time and storage requirements. Data standardization: The chunking processing step ensures the consistency of the data format, removes useless information, and improves the overall quality of the data. Hash calculation and electronic signature generation ensure the uniqueness and integrity of the data, prevent data tampering, and improve data security. Through the electronic signature, the authenticity and integrity of the document can be quickly verified, improving the verification efficiency and facilitating subsequent management and use. Compression processing reduces the amount of data, lowers the storage requirements, and improves the storage and transmission efficiency.

[0053] Optionally, the electronic document chunk data includes first electronic document chunk data and second electronic document chunk data, and the fixed chunking includes:

[0054] Performing functional fixed chunking on the electronic document data to obtain first electronic document chunk data;

[0055] Perform record-fixed chunking on the electronic document data to obtain the second electronic document chunk data;

[0056] Among them, the function-fixed chunking is specifically as follows:

[0057] Perform function classification on the electronic document data to obtain the electronic document classification data, where the electronic document classification data includes title classification data, paragraph classification data, image classification data, and table classification data;

[0058] Perform title chunking on the electronic document data according to the title classification data to obtain the title chunk data;

[0059] Perform paragraph chunking on the electronic document data according to the paragraph classification data to obtain the paragraph chunk data;

[0060] Perform image chunking on the electronic document data according to the image classification data to obtain the image chunk data;

[0061] Perform table chunking on the electronic document data according to the table classification data to obtain the table chunk data;

[0062] Integrate the title chunk data, paragraph chunk data, image chunk data, and table chunk data to obtain the first electronic document chunk data.

[0063] In the present invention, by performing function-fixed chunking and record-fixed chunking on the electronic document, the document is divided into several parts for separate processing. Classification and chunking processing are performed according to different functional parts (such as titles, paragraphs, images, tables) of the document content, and chunking processing is performed according to the record part of the document. Compared with the traditional electronic file processing method that often processes the entire document as a whole, the processing efficiency is low, especially when dealing with large files. The present invention improves the data processing efficiency, enables each chunk to be processed in parallel, and significantly reduces the processing time.

[0064] Optionally, the record-fixed chunking includes:

[0065] Obtain the document query data corresponding to the electronic document data;

[0066] Perform segmentation processing on the electronic document data to obtain the electronic document segmentation data;

[0067] Perform weight sorting processing according to the document query data and the electronic document segmentation data to obtain the electronic document segmentation sorting data;

[0068] Perform key segment extraction according to the electronic document segmentation sorting data to obtain the electronic document key segment data;

[0069] Perform clustering calculations on the key segmented data of the electronic document to obtain the segmented clustering data of the electronic document;

[0070] Perform clustering label chunking on the key segmented data of the electronic document according to the segmented clustering data of the electronic document to obtain the second chunked data of the electronic document.

[0071] In the present invention, processing is performed according to the user's query data to highlight the key content of the document data, while reducing the impact of irrelevant or fixed template content on hash calculation, reducing the possibility of database collision in subsequent processing, and improving the pertinence and efficiency of content processing. Through segmentation and chunking processing, the document content is structured, making the content more organized and easier to manage and read. Through weight sorting and key segment extraction, important content can be processed preferentially, improving the resource utilization efficiency. Clustering analysis can discover the inherent similarity laws in the content, enhance the chunking ability based on key content, increase data sensitivity, and reduce the database collision ability of hash calculation. The chunked data is convenient for parallel processing, improving the overall processing speed and efficiency.

[0072] Optionally, the signature verification includes:

[0073] Obtain the user location data, and perform regional signature verification based on the user location data and the signature location data corresponding to the electronic signature data to obtain the first signature verification data;

[0074] Obtain the user identity data, and perform user signature verification based on the user identity data and the signature identity data corresponding to the electronic signature data to obtain the second signature verification data;

[0075] Perform signature content verification based on the electronic signature data and the meta-electronic signature data in the electronic document metadata to obtain the third signature verification data;

[0076] Integrate the first signature verification data, the second signature verification data, and the third signature verification data to obtain the signature verification data;

[0077] Among them, the regional signature verification is specifically:

[0078] Obtain the user location data, where the user location data includes the user's GPS positioning data, the user's Wi-Fi positioning data, and the user's mobile network positioning data;

[0079] Perform location data fusion based on the user location data to obtain the location fusion data;

[0080] Perform multi-dimensional location credibility calculation based on the location fusion data to obtain the multi-dimensional location credibility data;

[0081] Perform location consistency check based on the multi-dimensional location credibility data to obtain the location consistency check data;

[0082] Perform geofence verification based on location consistency check data to obtain first signature verification data;

[0083] Among them, the location consistency check is specifically as follows:

[0084] When it is determined that the location consistency check data is reliable location consistency data, perform geofence verification based on the location consistency check data and the preset virtual geofence data to obtain first signature verification data;

[0085] When it is determined that the location consistency check data is suspicious location consistency data, turn on the terminal camera to collect environmental image data to obtain environmental image data;

[0086] Extract the solar illumination angle according to the environmental image data to obtain solar illumination angle data;

[0087] Estimate the user's environmental location according to the solar illumination angle data to obtain user environmental location data;

[0088] Perform fitting and screening according to the user's environmental location data and the user's location data to obtain user location fitting data;

[0089] Perform geofence verification according to the user location fitting data and the preset virtual geofence data to obtain first signature verification data.

[0090] In the present invention, through multi-dimensional verification of location, identity, and content, the comprehensiveness and reliability of the signature are improved. Quickly verify reliable location data, improve verification efficiency, and reduce processing time. Integrate multi-source location data, improve location accuracy, and ensure that the signature is completed at the expected location. Through multi-level verification and consistency check, ensure the security of the signature, and prevent location and identity fraud. Integrate multiple verification results, provide comprehensive signature verification data, and enhance the credibility and security of the document.

[0091] Optionally, the present application further provides an electronic file authenticity verification system for executing the electronic file authenticity verification method as described above. The electronic file authenticity verification system includes:

[0092] An electronic document basic data acquisition module for obtaining electronic document data and corresponding electronic document metadata;

[0093] A cured information validity check module for performing a cured information validity check according to the electronic document data and the electronic document metadata to obtain cured information validity data;

[0094] A content authenticity detection module for performing a content authenticity detection according to the electronic document data and the electronic document metadata to obtain content authenticity detection data;

[0095] An association consistency detection module is used to perform association consistency detection based on the electronic document data and the electronic document metadata to obtain association consistency detection data;

[0096] The electronic archive authenticity verification operation module is used to perform electronic archive authenticity verification operations based on solidified information validity data, content authenticity detection data and associated consistency detection data.

[0097] The objects of the present invention are:

[0098] 1. This invention ensures the integrity and authenticity of electronic archive data at all levels through a multi-layered verification approach, including checking the validity of fixed information, content authenticity, and association consistency. This multi-layered, multi-dimensional verification process ensures the overall authenticity and integrity of documents, improving overall security by 70%.

[0099] 2. By combining electronic document data and metadata for verification, the present invention can more accurately determine the authenticity of data. For example, by comparing the document content with the summary, signature, and other information in the metadata, it can effectively detect whether the data has been tampered with.

[0100] 3. This invention adopts a step-by-step approach, dividing the verification process into multiple steps. Each step independently performs a specific verification task, significantly improving data processing efficiency. For example, the validity check of the fixed information and the authenticity check of the content can be processed in parallel, speeding up the overall verification process.

[0101] 4. A single electronic signature or hash checksum has limited protection against potential tampering methods of varying dimensions and is easily compromised. By combining multi-source data verification and location consistency checks, this invention effectively prevents data tampering. For example, by using environmental image acquisition and sunlight angle calculation, the user's actual location can be further verified to prevent location forgery. BRIEF DESCRIPTION OF THE DRAWINGS

[0102] Other features, objects and advantages of the present application will become more apparent upon reading the detailed description of non-limiting embodiments made with reference to the following drawings:

[0103] Figure 1 A flowchart showing the steps of a method for verifying the authenticity of an electronic file according to an embodiment is shown;

[0104] Figure 2 A flowchart showing the steps of a method for collecting basic data of an electronic document according to an embodiment is shown;

[0105] Figure 3 A flowchart showing the steps of a method for checking validity of fixed information according to an embodiment is shown;

[0106] Figure 4 The flowchart of the steps of a content authenticity detection method according to an embodiment is shown;

[0107] Figure 5 The flowchart of the steps of a correlation consistency detection method according to an embodiment is shown;

[0108] The realization, functional features and advantages of the object of the present invention will be further described with reference to the embodiments and the accompanying drawings. Specific embodiments

[0109] The technical method of the present invention patent will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are part of the embodiments of the present invention, rather than all of the embodiments. All other embodiments obtained by those skilled in the art based on the embodiments of the present invention without creative work belong to the scope of protection of the present invention.

[0110] In addition, the accompanying drawings are only schematic diagrams of the present invention and are not necessarily drawn to scale. The same reference numerals in the drawings represent the same or similar parts, and thus their repeated description will be omitted. Some of the block diagrams shown in the drawings are functional entities and do not necessarily correspond to physically or logically independent entities. The functional entities can be implemented in software form, or implemented in one or more hardware modules or integrated circuits, or implemented in different networks and / or processor methods and / or microcontroller methods.

[0111] It should be understood that although the terms "first", "second", etc. may be used here to describe each unit, these units should not be limited by these terms. These terms are only used to distinguish one unit from another. For example, without departing from the scope of the exemplary embodiments, the first unit can be called the second unit, and similarly the second unit can be called the first unit. The term "and / or" used here includes any and all combinations of one or more of the listed related items.

[0112] Please refer to Figures 1 to 5 , the present application provides an electronic file authenticity verification method, and the method includes:

[0113] S1. Obtain electronic document data and corresponding electronic document metadata;

[0114] Specifically, connect to the electronic document management system (EDMS) using an API or a script. Write an API or a script to extract the electronic document data and metadata from the EDMS through interface calls.

[0115] S2. Performing a validity check on the fixed information based on the electronic document data and the electronic document metadata to obtain fixed information validity data;

[0116] Specifically, a hash algorithm (such as SHA-256) is used. The document content is hashed, a digest is generated, and compared with the digest in the metadata. The input data is padded to meet specific length requirements. A 1 bit is appended: A 1 bit is appended to the end of the input data. A 0 bit is appended: k 0 bits are appended, so that the data length satisfies length ≡ 448 (mod 512) (i.e., 64 bits short of a multiple of 512). The original length of the input data (in bits) is appended to the end of the padded data as a 64-bit binary number. Eight hash values (each 32 bits) are initialized using a set of predefined constants (H0, H1, H2, H3, H4, H5, H6, H7). These constants are the fractional parts of the square roots of the first eight prime numbers. The preprocessed message is divided into message blocks of 512 bits (64 bytes). The initial hash value (H0, H1, H2, H3, H4, H5, H6, H7) is copied to the working variables (a, b, c, d, e, f, g, h). The 512-bit message block is expanded into 64 32-bit words. After 64 rounds of operation, the working variables a through h are updated. The hash value is updated after each message block is processed. When all message blocks are processed, the final hash value is composed of H0, H1, H2, H3, H4, H5, H6, and H7, forming a 256-bit message digest.

[0117] Use a cryptographic library such as PyCryptodome to extract the electronic signature data and use the public key for signature verification.

[0118] Use image processing libraries (such as OpenCV) and OCR technology. Extract seal image data, use OCR technology to identify the seal content, and verify it. Use image contour detection algorithms to identify contours in the image. Filter out the seal area based on the shape, area, and other features of the contour. Crop the seal image data based on the filtered contour coordinates. Initialize the OCR engine (such as Tesseract). Use the OCR engine to perform text recognition on the cropped seal image and extract the seal content. Store the recognized seal content as text data. Load the expected seal content from metadata or preset values. Compare the recognized seal content with the expected content to check their consistency. Record the comparison results and generate verification result data.

[0119] S3. Perform content authenticity detection based on the electronic document data and the electronic document metadata to obtain content authenticity detection data;

[0120] Specifically, the electronic attributes (such as file size, format, etc.) of the electronic archive content data are automatically extracted and compared with the records in the metadata.

[0121] Query the change record table data and check the electronic document data against the historical record file and change record data to check the integrity of the content data and ensure that no data is lost or damaged.

[0122] Generate a checksum for the content data to further verify the integrity of the data. Compare the generated checksum with the record in the metadata.

[0123] Verify multimedia data (such as images and videos) to ensure its authenticity. Summarize all attribute comparison results, integrity results, checksum comparison results, and change log analysis results to generate content authenticity detection data.

[0124] S4. Perform association consistency detection based on the electronic document data and the electronic document metadata to obtain association consistency detection data;

[0125] Specifically, regular expressions or natural language processing (NLP) technology is used to extract the directory structure from the document content and generate directory structure data.

[0126] Use string matching algorithms to associate directory entries with document content to generate associated data.

[0127] Use data comparison tools to compare the generated metadata format with the original metadata to ensure consistency.

[0128] S5. Perform electronic archive authenticity verification based on the solidified information validity data, content authenticity detection data, and associated consistency detection data.

[0129] Specifically, data integration tools and scripts are used to integrate and consolidate information validity data, content authenticity detection data, and association consistency detection data for final authenticity verification.

[0130] Optionally, S1 includes:

[0131] S11, initialize the data collection interface, extract electronic document data and metadata from the electronic archive management system, and obtain primary electronic document data and corresponding primary electronic document metadata;

[0132] Specifically, use an API or script to connect to the Electronic Document Management System (EDMS). Establish communication with the EDMS through the API or database connection to obtain access rights. Send a request to obtain the specified electronic document and its metadata. Receive the primary electronic document data and primary electronic document metadata from the EDMS, and save them as temporary files or directly store them in memory. Use a RESTful API call to obtain the document: obtain the document data and metadata through an HTTP GET request, and receive and parse the data in JSON or XML format.

[0133] S12. Parse the primary electronic document metadata to obtain the electronic document metadata;

[0134] Specifically, load the primary electronic document metadata into the parsing library. Parse the metadata in XML or JSON format and extract the key fields (such as document ID, title, author, creation date, etc.). Store the parsed metadata as structured data for subsequent processing.

[0135] S13. Parse the primary electronic document data to obtain the electronic document data.

[0136] Specifically, load the primary electronic document data into memory. Perform preliminary processing on the document content, such as removing HTML tags and extracting the text content. Use regular expressions or NLP techniques to further parse the document content and identify the document structure (such as title, paragraphs, tables, images, etc.). Store the parsed document content as structured data for subsequent verification and processing.

[0137] Optionally, S2 includes:

[0138] S21. Generate an electronic document summary based on the electronic document data to obtain the electronic document summary data;

[0139] Specifically, load the electronic document data into memory. Use a hashing algorithm such as SHA-256 to calculate the document content and generate a unique document summary. Store the generated document summary data for subsequent comparison.

[0140] S22. Compare the electronic document summary data with the meta-summary data in the electronic document metadata to obtain the summary comparison data;

[0141] Specifically, extract the meta-summary data from the metadata. Compare the generated document summary data with the meta-summary data to check if they are consistent. Record the comparison result to obtain the summary comparison data.

[0142] S23. Extract the electronic signature based on the electronic document data to obtain the electronic signature data;

[0143] Specifically, a data block or field containing an electronic signature is identified from the electronic document. The signature data is extracted from the signature block or field. The extracted signature data is stored for subsequent verification.

[0144] S24. Perform signature verification on the electronic signature data to obtain signature verification data;

[0145] Specifically, the public key for verification is loaded from the certificate or key library. The public key is used to verify the electronic signature data to confirm the validity of the signature. The verification result is recorded to obtain the signature verification data.

[0146] S25, extracting the electronic seal according to the electronic document data to obtain the electronic seal data;

[0147] Specifically, the portion of the electronic document containing the seal is identified through image processing technology, the seal image data is extracted from the document, and the extracted seal image data is stored for subsequent verification.

[0148] S26, performing seal verification according to the electronic seal data to obtain seal verification data;

[0149] Specifically, a standard seal image or text content is loaded from a database. An image matching algorithm is used to compare the extracted seal image with the standard seal. The seal content extracted using optical character recognition (OCR) technology is then compared with the standard content. The verification results are recorded to obtain seal verification data.

[0150] S27. Integrate the summary comparison data, signature verification data, and seal verification data to obtain solidified information validity data.

[0151] Specifically, all summary comparison data, signature verification data, and seal verification data are collected. All verification data is integrated into a single data structure. Solidified information validity data is generated, containing all verification results and detailed information. This solidified information validity data is stored in a database or log system for easy auditing and tracking.

[0152] Optionally, S3 includes:

[0153] S31, extracting attributes from the electronic document data to obtain electronic document attribute data;

[0154] Specifically, the electronic document data is loaded into memory. A document parsing tool is used to extract the electronic attributes of the document, such as creation date, modification date, author, file size, file type, etc. The extracted electronic document attribute data is stored as structured data to facilitate subsequent comparison.

[0155] S32, performing attribute comparison based on the electronic document attribute data and the meta-attribute data in the electronic document metadata to obtain attribute comparison data;

[0156] Specifically, extract meta-attribute data from the electronic document metadata. Compare the extracted electronic document attribute data with the meta-attribute data item by item to check for consistency. Record the comparison results to obtain attribute comparison data.

[0157] S33. Perform content integrity detection based on the electronic document data and the electronic document metadata to obtain content integrity detection data;

[0158] Specifically, load the content in the electronic document data and the metadata into memory. Use a text comparison algorithm to compare the document content and the content in the metadata word by word to check for consistency. Record the comparison results to obtain content integrity detection data.

[0159] S34. Generate a checksum for the electronic document data to obtain electronic document checksum data;

[0160] Specifically, load the electronic document data into memory. Use the SHA-256 or other hash algorithms to generate a checksum for the document data. Store the generated checksum as checksum data for subsequent comparison.

[0161] S35. Compare the electronic document checksum with the meta-checksum data in the electronic document metadata to obtain checksum comparison data;

[0162] Specifically, extract the meta-checksum data from the electronic document metadata. Compare the generated checksum data with the meta-checksum data to check for consistency. Record the comparison results to obtain checksum comparison data.

[0163] S36. Perform consistency processing on the content change records in the electronic document data to obtain change record consistency data;

[0164] Specifically, extract the change records from the electronic document data and the metadata. Compare the extracted change records with the change records in the metadata to check for consistency. Record the comparison results to obtain change record consistency data.

[0165] S37. Verify the multimedia data in the electronic document data to obtain multimedia verification data;

[0166] Specifically, extract multimedia data such as images, audio, and video from the electronic document data. Use a multimedia verification tool to verify the extracted data to check if it is complete and unmodified. For example, check if the image file size, pixels, saturation, and grayscale image match the image attributes stored in the metadata, if the audio file size, audio change rate, audio duration, and preset sampling points match the audio recording data in the metadata, and if the video file size, video duration, and video images generated before and after the video sampling points match the video recording data in the metadata. Record the verification results to obtain multimedia verification data.

[0167] S38. Integrate the attribute comparison data, content integrity detection data, checksum comparison data, change record consistency data, and multimedia verification data to obtain content authenticity detection data.

[0168] Specifically, collect all the attribute comparison data, content integrity detection data, checksum comparison data, change record consistency data, and multimedia verification data. Integrate all the verification data into an overall data structure. Generate content authenticity detection data, which includes all the verification results and detailed information. Store the authenticity detection data in a database or a logging system for easy auditing and tracking.

[0169] Optionally, S4 includes:

[0170] S41. Extract the directory structure based on the electronic document data to obtain directory structure data;

[0171] Specifically, load the electronic document data into memory. Use regular expressions or natural language processing (NLP) tools to identify and extract the directory structure, including information such as chapter titles and page numbers. Structurally store the extracted directory information to obtain directory structure data.

[0172] S42. Perform directory content association processing based on the electronic document data and the directory structure data to obtain directory content association data;

[0173] Specifically, load the directory structure data and the electronic document data into memory. Match the chapter titles in the directory structure with the document content to confirm the location of each directory entry in the document. Store the matching results as directory content association data, which includes the relationship between the directory entries and the corresponding content locations.

[0174] S43. Perform directory verification on the directory content association data to obtain directory verification data;

[0175] Specifically, load the directory content association data into memory. Check if each directory entry is correctly associated with the corresponding document content to ensure the consistency and integrity of the directory entries and the content. Record the verification results to obtain directory verification data.

[0176] S44. Generate metadata format data based on the electronic document data to obtain metadata format data;

[0177] Specifically, load the electronic document data into memory. Parse the document content, extract and format the generated metadata, including the document author, creation date, modification date, etc. Store the generated metadata as metadata format data.

[0178] S45. Perform metadata consistency verification based on the metadata format data and the electronic document metadata to obtain metadata consistency verification data;

[0179] Specifically, load the generated metadata format data and the metadata in the document into memory. Compare the metadata in the two data sources, check their consistency, and ensure that the metadata recorded in the document is correct. Record the comparison result to obtain metadata verification data.

[0180] S46. Generate a mapping relationship based on the directory structure data, the electronic document data, and the electronic document metadata to obtain mapping relationship data;

[0181] Specifically, load the relevant data into memory. Establish a mapping relationship among the directory structure, the document content, and the metadata to ensure the consistency and relevance of all data. Store the generated mapping relationship as mapping relationship data.

[0182] S47. Perform relevance verification based on the mapping relationship data to obtain relevance verification data;

[0183] Specifically, load the mapping relationship data into memory. Check the correctness of the mapping relationship to ensure that the association among the directory, the content, and the metadata is correct. Record the verification result to obtain relevance verification data.

[0184] S48. Integrate the directory verification data, the metadata consistency verification data, and the relevance verification data to obtain association consistency detection data.

[0185] Specifically, collect and load the directory verification data, the metadata verification data, and the relevance verification data into memory. Integrate all the verification and validation data into an overall data structure. Store the integrated data as association consistency detection data for subsequent query and use.

[0186] Optionally, the electronic signature extraction includes:

[0187] Perform fixed block division on the electronic document data to obtain electronic document block data;

[0188] Specifically, load the electronic document data into the memory. The document is chunked according to a predetermined rule (such as each 1KB as a block, or according to the document content structure such as paragraphs, chapters, etc.). The chunked data is stored as electronic document chunked data for subsequent processing.

[0189] Perform standardization processing on the electronic document chunked data to obtain electronic document chunked processed data;

[0190] Specifically, load the fixed chunked data blocks into the memory one by one. Process each chunk of data, such as removing blank lines, standardizing the format, cleaning up noise data, etc. Store the processed chunked data as electronic document chunked processed data.

[0191] Perform compression processing on the electronic document chunked processed data to obtain electronic document chunked compressed data;

[0192] Specifically, load the processed chunked data into the memory. Use a compression algorithm (such as gzip, zlib) to compress each chunk of data to reduce the data volume. Store the compressed chunked data as electronic document chunked compressed data.

[0193] Generate intermediate hash data based on the electronic document chunked compressed data;

[0194] Specifically, load the compressed chunked data into the memory. Use SHA-256 or other hash algorithms to calculate the hash value for each chunk of data, generating intermediate hash values. Store the intermediate hash values as intermediate hash data for subsequent calculations.

[0195] Calculate the hash value based on the intermediate hash data to obtain the electronic signature data.

[0196] Specifically, load all the intermediate hash values into the memory. Combine all the intermediate hash values in order and use a hash algorithm (such as SHA-256) for the final hash calculation. Generate the final electronic signature data as the unique identifier of the document.

[0197] Optionally, the electronic document chunked data includes first electronic document chunked data and second electronic document chunked data. The fixed chunking includes:

[0198] Perform functional fixed chunking on the electronic document data to obtain the first electronic document chunked data;

[0199] Specifically, chunk the electronic document data according to the document content type to obtain the first electronic document chunked data.

[0200] Perform record fixed chunking on the electronic document data to obtain the second electronic document chunked data;

[0201] Specifically, obtain the query record data corresponding to the electronic document data; determine the key content data in the electronic document data according to the query record data, and extract it; perform fixed chunking on the extracted key content data to obtain the second electronic document chunk data.

[0202] Among them, the fixed chunking of functions is specifically as follows:

[0203] Perform function classification on the electronic document data to obtain electronic document classification data, where the electronic document classification data includes title classification data, paragraph classification data, image classification data, and table classification data;

[0204] Specifically, load the electronic document data into memory. Use NLP tools / XML format recognition scripts to analyze the document and identify different functional parts (such as titles, paragraphs, images, tables, etc.). Classify and store the identified functional parts as electronic document classification data.

[0205] Perform title chunking on the electronic document data according to the title classification data to obtain title chunk data;

[0206] Specifically, extract the title classification data from the electronic document classification data. Perform chunking on the document data according to the title classification data and extract all title parts. Store the extracted title parts as title chunk data.

[0207] Perform paragraph chunking on the electronic document data according to the paragraph classification data to obtain paragraph chunk data;

[0208] Specifically, extract the paragraph classification data from the electronic document classification data. Perform chunking on the document data according to the paragraph classification data and extract all paragraph parts. Store the extracted paragraph parts as paragraph chunk data.

[0209] Perform image chunking on the electronic document data according to the image classification data to obtain image chunk data;

[0210] Specifically, extract the image classification data from the electronic document classification data. Perform chunking on the document data according to the image classification data and extract all image parts. Store the extracted image parts as image chunk data.

[0211] Perform table chunking on the electronic document data according to the table classification data to obtain table chunk data;

[0212] Specifically, extract the table classification data from the electronic document classification data. Perform chunking on the document data according to the table classification data and extract all table parts. Store the extracted table parts as table chunk data.

[0213] Integrate the title chunk data, paragraph chunk data, image chunk data, and table chunk data to obtain the first electronic document chunk data.

[0214] Specifically, load the title chunk data, paragraph chunk data, image chunk data, and table chunk data respectively. Integrate each chunk of data into an overall data structure to ensure that each part is arranged in an orderly manner. Store the integrated data as the first electronic document chunk data for subsequent processing.

[0215] Optionally, the recorded fixed chunks include:

[0216] Obtain the document query data corresponding to the electronic document data;

[0217] Specifically, establish a connection with the electronic file management system to obtain query permissions. Send a query request to obtain the relevant query data of the document. Store the query data as structured data for subsequent processing.

[0218] Perform segmentation processing on the electronic document data to obtain the electronic document segmented data;

[0219] Specifically, load the electronic document data into memory. Use regular expressions or NLP tools to segment the document and split the document content into several paragraphs. Store the segmented data as the electronic document segmented data.

[0220] Perform weight sorting processing based on the document query data and the electronic document segmented data to obtain the electronic document segmented sorting data;

[0221] Specifically, load the document query data and the segmented data into memory. Sort according to the weight of the query data and the relevance of the segmented data to determine the importance of each segment. Store the sorted data as the electronic document segmented sorting data.

[0222] Extract key segments based on the electronic document segmented sorting data to obtain the electronic document key segment data;

[0223] Specifically, load the sorted segmented data into memory. Extract the segmented data with higher weights as the key segments. Store the extracted key segment data for subsequent processing. Extract the top 10 paragraphs with the highest weights after sorting and store them as the key segment data.

[0224] Perform clustering calculation on the electronic document key segment data to obtain the electronic document segmented clustering data;

[0225] Specifically, load the key segment data into memory. Use clustering algorithms (such as K-means, DBSCAN) to cluster the key segment data and identify similar content. Store the clustering results as the electronic document segmented clustering data.

[0226] Cluster the clustering labels of the key segmented data of the electronic document according to the segmented clustering data of the electronic document to obtain the second segmented data of the electronic document.

[0227] Specifically, load the clustering result and the key segmented data into the memory. Block the key segmented data according to the clustering labels, and group the data with the same clustering label into one block. Store the blocked data as the second segmented data of the electronic document.

[0228] Optionally, the signature verification includes:

[0229] Obtain the user location data, and perform regional signature verification according to the user location data and the signature location data corresponding to the electronic signature data to obtain the first signature verification data;

[0230] Specifically, obtain the real-time location data of the user through GPS, Wi-Fi positioning or mobile network positioning. Extract the location information at the time of signature from the electronic signature data. Compare the user location data with the signature location data to check the consistency of the locations. Record the comparison result to obtain the first signature verification data.

[0231] Obtain the user identity data, and perform user signature verification according to the user identity data and the signature identity data corresponding to the electronic signature data to obtain the second signature verification data;

[0232] Specifically, obtain the user identity data through biometric identification (such as fingerprint, face recognition) or digital certificate verification. Extract the identity information at the time of signature from the electronic signature data. Compare the user identity data with the signature identity data to confirm the identity of the signer. Record the comparison result to obtain the second signature verification data.

[0233] Perform signature content verification according to the electronic signature data and the meta-electronic signature data in the electronic document metadata to obtain the third signature verification data;

[0234] Specifically, extract the signature content and signature metadata from the electronic signature data. Use the public key and signature verification algorithm to verify the signature content to ensure that the signature has not been tampered with. Record the verification result to obtain the third signature verification data.

[0235] Integrate the first signature verification data, the second signature verification data, and the third signature verification data to obtain the signature verification data;

[0236] Specifically, collect the first signature verification data, the second signature verification data, and the third signature verification data. Integrate all the verification data into an overall data structure to ensure data integrity and consistency. Generate the final signature verification data, which includes the results and detailed information of all verifications. Store the signature verification data in a database or a logging system for easy auditing and tracking.

[0237] Among them, the geographical signature verification is specifically as follows:

[0238] Obtain the user location data, where the user location data includes the user's GPS positioning data, the user's Wi-Fi positioning data, and the user's mobile network positioning data;

[0239] Specifically, obtain the longitude and latitude coordinates of the user through the GPS module of the device. Use the Wi-Fi module of the device to scan the surrounding Wi-Fi signals and obtain the location data in combination with the Wi-Fi positioning service. Obtain the base station positioning data through the mobile network connection of the device.

[0240] Perform location data fusion based on the user location data to obtain location fusion data;

[0241] Specifically, load the obtained GPS, Wi-Fi, and mobile network positioning data into the memory. Use data fusion algorithms such as Kalman filtering to fuse the multi-source location data, remove noise, and improve the positioning accuracy. Store the fused location data as location fusion data.

[0242] Perform multi-dimensional location credibility calculation based on the location fusion data to obtain multi-dimensional location credibility data;

[0243] Specifically, load the fused location data into the memory. Perform credibility calculation based on the credibility of the data source (such as GPS, Wi-Fi, mobile network), the stability of the location data (the smoothness of the location change), and the matching degree of the historical data (the similarity with the past data). Store the calculated multi-dimensional location credibility as multi-dimensional location credibility data.

[0244] Perform location consistency check based on the multi-dimensional location credibility data to obtain location consistency check data;

[0245] Specifically, load the multi-dimensional location credibility data into the memory. Check the consistency of the location data according to the preset threshold or rules, and determine whether the location data is reliable and consistent. Store the check result as location consistency check data. Set the credibility threshold to 0.8. If the multi-dimensional location credibility score is higher than 0.8, it is considered that the location data is consistent and reliable, and record the check result.

[0246] Perform geographical fence verification based on the location consistency check data to obtain the first signature verification data;

[0247] Specifically, load the location consistency check data into memory. Define a preset geofence area (such as a specific latitude and longitude range). Check whether the location data is within the geofence to determine whether the location is legal. Record the verification result to obtain the first signature verification data. Set the geofence to the latitude and longitude range of the company's office area, check whether the location data is within this range, and record the geofence verification result. Use the least squares method to fit the user's environmental location data and the user's location data to check their consistency.

[0248] Among them, the location consistency check is specifically as follows:

[0249] When it is determined that the location consistency check data is reliable location consistency data, perform geofence verification based on the location consistency check data and the preset virtual geofence data to obtain the first signature verification data;

[0250] Specifically, load the location consistency check data into memory. According to the multi-dimensional location credibility score and the preset credibility threshold, judge whether the data is reliable. If the credibility score is higher than the threshold, the data is considered reliable. Compare the reliable data with the preset virtual geofence to check whether it is within the geofence, and generate the first signature verification data.

[0251] When it is determined that the location consistency check data is suspicious location consistency data, turn on the terminal camera to collect environmental image data;

[0252] Specifically, load the location consistency check data into memory. According to the multi-dimensional location credibility score and the preset credibility threshold, judge whether the data is suspicious. If the credibility score is lower than the threshold, the data is considered suspicious. Start the device camera to collect environmental images, and perform object detection on the collected image stream until an image with the light and dark changes of the sun and objects is determined and extracted. Perform spatial relationship fusion based on the extracted images to obtain the current environmental image data.

[0253] Use object detection algorithms (such as YOLO, SSD, etc.) to detect the sun and light and dark changes in the image. Load the pre-trained object detection model. Perform object detection on each frame of the image to find images containing the sun and significant light and dark changes. Extract the images containing the sun and light and dark changes from the object detection results. Based on the spatial relationship algorithm, fuse multiple images containing the sun and light and dark changes to generate the current environmental image data. Load all the images containing the sun and light and dark changes. Perform image fusion according to the spatial relationship of the images (such as shooting location, camera rotation / translation angle, shooting angle, etc.) to generate a complete environmental image.

[0254] Extract the sunshine angle according to the environmental image data to obtain the sunshine angle data;

[0255] Specifically, the collected environmental image data is loaded into the memory. Computer vision algorithms are used to analyze the light direction in the image and extract the sunlight angle. The extracted sunlight angle information is stored as sunlight angle data. The color image is converted into a grayscale image. Edge detection algorithms (such as Canny edge detection) are used to identify the edges in the image. Hough line transform is used to detect the lines in the image, representing the light direction. The angle of each line is calculated according to the result of the Hough transform. All the angles are analyzed to extract the angle of the main light direction.

[0256] Based on the sunlight angle data, the user's environmental location is estimated to obtain the user's environmental location data;

[0257] Specifically, the extracted sunlight angle data is loaded into the memory. According to the current time and the sunlight angle, astronomical calculation methods are used to estimate the user's environmental location. The estimated user's environmental location is stored as the user's environmental location data. The current date and time are obtained. According to the current time and the sunlight angle, the position of the sun in the celestial sphere is calculated. According to the calculated sun position and the sunlight angle, the user's geographical location is estimated.

[0258] Based on the user's environmental location data and the user's location data, fitting and screening are performed to obtain the user location fitting data;

[0259] Specifically, the user's environmental location data and the user's location data are loaded into the memory. A location fitting algorithm is used to compare the two location data, and the consistent or approximate data is screened out. The fitted location data is stored as the user location fitting data.

[0260] Based on the user location fitting data and the preset virtual geographical fence data, geographical fence verification is performed to obtain the first signature verification data.

[0261] Specifically, the user location fitting data and the virtual geographical fence data are loaded into the memory. It is checked whether the fitted data is within the geographical fence to determine whether the location is legal. The verification result is recorded to obtain the first signature verification data.

[0262] Optionally, the present application also provides an electronic file authenticity verification system for performing the electronic file authenticity verification method as described above. The electronic file authenticity verification system includes:

[0263] An electronic document basic data collection module for obtaining the electronic document data and the corresponding electronic document metadata;

[0264] A cured information validity check module for performing a cured information validity check according to the electronic document data and the electronic document metadata to obtain the cured information validity data;

[0265] A content authenticity detection module, which is used to perform content authenticity detection based on electronic document data and electronic document metadata to obtain content authenticity detection data;

[0266] An association consistency detection module, which is used to perform association consistency detection based on electronic document data and electronic document metadata to obtain association consistency detection data;

[0267] An electronic file authenticity verification operation module, which is used to perform electronic file authenticity verification operations based on the validity data of solidified information, content authenticity detection data, and association consistency detection data.

[0268] Therefore, from any perspective, the embodiments should be regarded as exemplary and non-limiting. The scope of the present invention is defined by the appended application documents rather than the above description. Therefore, all changes falling within the meaning and scope of the equivalent elements of the application documents are intended to be included in the present invention.

[0269] The above are only specific embodiments of the present invention, enabling those skilled in the art to understand or implement the present invention. Various modifications to these embodiments will be obvious to those skilled in the art. The general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention will not be limited to these embodiments shown herein, but rather to the broadest scope consistent with the principles and novel features invented herein.

Claims

1. An electronic file authenticity verification method, characterized in that The method comprises: S1. Obtain electronic document data and corresponding electronic document metadata; S2. Performing a validity check on the fixed information based on the electronic document data and the electronic document metadata to obtain fixed information validity data; S3. Perform content authenticity detection based on the electronic document data and the electronic document metadata to obtain content authenticity detection data; S4. Perform association consistency detection based on the electronic document data and the electronic document metadata to obtain association consistency detection data; S5. Perform electronic archive authenticity verification based on the solidified information validity data, content authenticity test data, and association consistency test data; S3 includes: Extracting attributes from electronic document data to obtain electronic document attribute data; Performing attribute comparison based on electronic document attribute data and meta-attribute data in electronic document metadata to obtain attribute comparison data; Performing content integrity detection based on electronic document data and electronic document metadata to obtain content integrity detection data; Generating a verification code according to the electronic document data to obtain electronic document verification code data; Performing a check code comparison based on the electronic document check code and the meta-check code data in the electronic document metadata to obtain check code comparison data; Perform content change record consistency processing based on electronic document data to obtain change record consistency data; Performing multimedia data verification on electronic document data to obtain multimedia verification data; The attribute comparison data, content integrity detection data, check code comparison data, change record consistency data and multimedia verification data are integrated to obtain content authenticity detection data.

2. The method according to claim 1, wherein S1 includes: Initialize the data collection interface, extract electronic document data and metadata from the electronic archive management system, and obtain primary electronic document data and corresponding primary electronic document metadata; Performing document metadata parsing on the primary electronic document metadata to obtain electronic document metadata; The primary electronic document data is parsed for document content data to obtain electronic document data.

3. The method according to claim 1, characterized in that S2 include: Generate a summary based on the electronic document data to obtain electronic document summary data; Performing summary comparison based on the electronic document summary data and the meta-summary data in the electronic document metadata to obtain summary comparison data; Extracting electronic signatures based on electronic document data to obtain electronic signature data; Performing signature verification on the electronic signature data to obtain signature verification data; Extracting electronic seals based on electronic document data to obtain electronic seal data; Perform seal verification according to the electronic seal data to obtain seal verification data; Integrate summary comparison data, signature verification data, and seal verification data to obtain solidified information validity data.

4. The method according to claim 1, wherein S4 includes: Extracting directory structure according to electronic document data to obtain directory structure data; Perform directory content association processing based on the electronic document data and the directory structure data to obtain directory content association data; Performing directory verification on directory content associated data to obtain directory verification data; Generating metadata format according to electronic document data to obtain metadata format data; Perform metadata consistency verification based on metadata format data and electronic document metadata to obtain metadata consistency verification data; Generate mapping relationship based on directory structure data, electronic document data and electronic document metadata to obtain mapping relationship data; Perform relevance verification based on the mapping relationship data to obtain relevance verification data; Integrate the directory verification data, metadata consistency verification data and relevance verification data to obtain associated consistency detection data.

5. The method according to claim 3, characterized in that Electronic signature extraction includes: Perform fixed chunking on the electronic document data to obtain electronic document chunk data; Perform standardization processing on the electronic document chunk data to obtain electronic document chunk processing data; Perform compression processing on the electronic document chunk processing data to obtain electronic document chunk compression data; Generate intermediate hash based on the electronic document chunk compression data to obtain intermediate hash data; Calculate the hash value based on the intermediate hash data to obtain electronic signature data.

6. The method according to claim 5, characterized in that, Among them, the electronic document chunk data includes the first electronic document chunk data and the second electronic document chunk data. The fixed chunking includes: Perform functional fixed chunking on the electronic document data to obtain the first electronic document chunk data; Perform record fixed chunking on the electronic document data to obtain the second electronic document chunk data; Among them, the functional fixed chunking is specifically: Classify the electronic document data according to functions to obtain electronic document classification data, where the electronic document classification data includes title classification data, paragraph classification data, image classification data and table classification data; Perform title chunking on the electronic document data according to the title classification data to obtain title chunk data; Perform paragraph chunking on the electronic document data according to the paragraph classification data to obtain paragraph chunk data; Perform image chunking on the electronic document data according to the image classification data to obtain image chunk data; Perform table chunking on the electronic document data according to the table classification data to obtain table chunk data; Integrate the title chunk data, paragraph chunk data, image chunk data and table chunk data to obtain the first electronic document chunk data.

7. The method according to claim 6, wherein The record fixed chunking includes: Obtain the document query data corresponding to the electronic document data; Perform segmentation processing on the electronic document data to obtain electronic document segmentation data; Perform weight sorting processing according to the document query data and the electronic document segmentation data to obtain electronic document segmentation sorting data; Extract key segments according to the electronic document segmentation sorting data to obtain electronic document key segment data; Perform clustering calculation on the electronic document key segment data to obtain electronic document segmentation clustering data; Perform clustering label chunking on the electronic document key segment data according to the electronic document segmentation clustering data to obtain the second electronic document chunk data.

8. The method according to claim 3, wherein Signature verification includes: Obtain the user location data, and perform regional signature verification according to the user location data and the signature location data corresponding to the electronic signature data to obtain the first signature verification data; Obtain the user identity data, and perform user signature verification according to the user identity data and the signature identity data corresponding to the electronic signature data to obtain the second signature verification data; Perform signature content verification based on the electronic signature data and the meta-electronic signature data in the electronic document metadata to obtain third signature verification data; Integrate the first signature verification data, the second signature verification data, and the third signature verification data to obtain signature verification data; The regional signature verification is as follows: Obtain user location data, including user GPS positioning data, user Wi-Fi positioning data, and user mobile network positioning data; Performing location data fusion according to user location data to obtain location fusion data; Perform multi-dimensional position credibility calculation based on position fusion data to obtain multi-dimensional position credibility data; Performing a position consistency check based on the multi-dimensional position credibility data to obtain position consistency check data; Performing geo-fence verification based on the location consistency check data to obtain first signature verification data; The location consistency check is as follows: When it is determined that the location consistency check data is location consistency reliable data, geo-fence verification is performed based on the location consistency check data and the preset virtual geo-fence data to obtain first signature verification data; When it is determined that the location consistency check data is location consistency suspicious data, the terminal camera is turned on to collect environmental images to obtain environmental image data; Extract sunlight angle according to environmental image data to obtain sunlight angle data; Estimate the user's environmental position based on the sunlight angle data to obtain the user's environmental position data; Perform fitting and screening based on the user environment location data and the user location data to obtain user location fitting data; Geofence verification is performed based on the user location fitting data and preset virtual geofence data to obtain first signature verification data.

9. An electronic file authenticity verification system, characterized in that, For executing the electronic file authenticity verification method according to claim 1, the electronic file authenticity verification system comprises: An electronic document basic data acquisition module is used to obtain electronic document data and corresponding electronic document metadata; A fixed information validity check module is used to check the validity of the fixed information based on the electronic document data and the electronic document metadata to obtain the fixed information validity data; A content authenticity detection module is used to perform content authenticity detection based on electronic document data and electronic document metadata to obtain content authenticity detection data; An association consistency detection module is used to perform association consistency detection based on the electronic document data and the electronic document metadata to obtain association consistency detection data; The electronic archive authenticity verification operation module is used to perform electronic archive authenticity verification operations based on solidified information validity data, content authenticity detection data and associated consistency detection data.

Citation Information

Patent Citations

  • Method, system and equipment for detecting authenticity of metadata of electronic file and medium

    CN115964684A

  • Single-set electronic archive authenticity detection method, electronic equipment and storage medium

    CN117633886A