File identification processing method, device and equipment

By combining multimodal data with blockchain, the problem of low accuracy and efficiency in archival authentication has been solved, achieving automation and objectivity in the authentication of archival authenticity and enhancing the ability to detect advanced forgery methods.

CN121963244APending Publication Date: 2026-05-01北京合思信息技术有限公司
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
北京合思信息技术有限公司
Filing Date
2026-01-19
Publication Date
2026-05-01

AI Technical Summary

Technical Problem

Existing methods for authenticating archives have low accuracy and efficiency, are easily bypassed by forgery techniques, and lack full lifecycle traceability.

Method used

An identification method combining multimodal data acquisition and blockchain network is adopted. By comparing hyperspectral image data, visible light image data and identification information and querying life cycle flow records, the identification results are generated by integrating decision-making models.

Benefits of technology

It has improved the accuracy and efficiency of archival authentication, automated and made the authentication process more objective, and enhanced the ability to detect advanced forgery techniques.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121963244A_ABST
    Figure CN121963244A_ABST
Patent Text Reader

Abstract

The invention provides an archive identification processing method, device and equipment, and relates to the technical field of archive identification. The method comprises the following steps: acquiring multi-modal data of a to-be-authenticated file; wherein the multi-modal data comprises at least one of the following items: hyperspectral image data, visible light image data and identification information. Comparing the hyperspectral image data and / or the visible light image data with data in a preset standard feature template library to obtain hyperspectral result information and / or visible light result information; and querying a life cycle circulation record associated with the identification information in a preset block chain network. Fusing the target result information, and generating identification result information of the target result information through a preset decision model; wherein the target result information comprises at least one of the following items: hyperspectral result information, visible light result information and life cycle circulation records. The method is used for achieving the effect of improving the accuracy and efficiency of authenticity identification of the archives.
Need to check novelty before this filing date? Find Prior Art

Description

Methods, devices and equipment for identifying and processing archives Technical Field

[0001] This application relates to the field of document recognition technology, and more specifically, to a document recognition and processing method, apparatus, and equipment. Background Technology

[0002] Currently, in fields such as forensic identification, financial risk control, property rights transactions, art auctions, and the management of important historical archives, the authenticity of key archival documents is the cornerstone of these businesses. These key documents, such as property ownership certificates, loan contracts, judicial judgments, calligraphy and paintings by famous figures, and important approvals, possess extremely high value. Therefore, they have become primary targets for forgery and alteration by criminals. Forgery methods range from simple photocopying and printing to sophisticated handwriting imitation, seal forgery, and paper aging techniques, constantly evolving and posing a serious challenge to traditional authentication methods.

[0003] However, existing methods for authenticating archival documents mainly include: expert authentication: relying on experienced experts to make judgments through sensory means such as sight, touch, and smell, and with the help of simple tools such as magnifying glasses and ultraviolet lamps. However, this method is highly subjective, the results are difficult to quantify and reproduce, it is inefficient, and the training period for experts is long and costly.

[0004] Single-technology detection: For example, using only OCR technology to compare text content with electronic archives cannot detect physical forgeries (such as forged signatures and seals). Or, using only a spectrometer to analyze the ink composition at a single point cannot verify the consistency of the entire content and the integrity of its circulation history. Single-technology detection has obvious "shortcomings" and is easily bypassed by targeted forgery methods.

[0005] Therefore, the accuracy and efficiency of existing methods for authenticating archives are both poor. Summary of the Invention

[0006] The purpose of this application is to provide a method, apparatus, and device for identifying and processing archives, thereby solving the aforementioned problems in the prior art and improving the accuracy and efficiency of authenticating archives.

[0007] Firstly, a method for identifying and processing archives is provided. This method may include: acquiring multimodal data of the archive to be identified; wherein the multimodal data includes at least one of the following: hyperspectral image data, visible light image data, and identification information; comparing the hyperspectral image data and / or the visible light image data with data in a preset standard feature template library to obtain hyperspectral result information and / or visible light result information; querying a lifecycle transfer record associated with the identification information in a preset blockchain network; fusing target result information and generating identification result information of the target result information through a preset decision model; wherein the target result information includes at least one of the following: the hyperspectral result information, the visible light result information, and the lifecycle transfer record.

[0008] Secondly, an identification and processing device for archives is provided. This device may include: an acquisition module for acquiring multimodal data of the archive to be identified; wherein the multimodal data includes at least one of the following: hyperspectral image data, visible light image data, and identification information; a comparison module for comparing the hyperspectral image data and / or the visible light image data with data in a preset standard feature template library, respectively, to obtain hyperspectral result information and / or visible light result information; querying a lifecycle transfer record associated with the identification information in a preset blockchain network; and a generation module for fusing target result information and generating identification result information of the target result information through a preset decision model; wherein the target result information includes at least one of the following: the hyperspectral result information, the visible light result information, and the lifecycle transfer record.

[0009] Thirdly, an electronic device is provided, comprising a processor, a communication interface, a memory, and a communication bus, wherein the processor, the communication interface, and the memory communicate with each other via the communication bus; the memory is used to store computer programs; and the processor, when executing the program stored in the memory, implements any of the steps described in the first aspect above.

[0010] Fourthly, a computer-readable storage medium is provided, wherein a computer program is stored therein, and when executed by a processor, the computer program implements the steps of any of the methods described in the first aspect above.

[0011] This application provides a method, apparatus, and device for identifying and processing archives, which acquires multimodal data of the archive to be identified. The multimodal data includes at least one of the following: hyperspectral image data, visible light image data, and identification information. The hyperspectral image data and / or visible light image data are compared with data in a preset standard feature template library to obtain hyperspectral result information and / or visible light result information. The lifecycle transfer record associated with the identification information is queried in a preset blockchain network. The target result information is fused, and an identification result information is generated through a preset decision model. The target result information includes at least one of the following: hyperspectral result information, visible light result information, and lifecycle transfer record. This solution constructs a unified, cross-verified, and automated authentication system that integrates "physical material (i.e., hyperspectral image data)," "information content (i.e., visible light image data)," and "transfer history (determined based on identification information)." This system fuses evidence from different dimensions and enables intelligent decision-making, improving the comprehensiveness, cross-verification capability, accuracy, and objectivity of the authentication process. It achieves highly efficient automation of the authentication process, addressing the core problems of existing methods for authenticating archives, such as limited authentication dimensions, inability to detect advanced forgery techniques, lack of process evidence chains, and subjective and inefficient conclusions. This solution can improve the accuracy and efficiency of archive authentication. Attached Figure Description

[0012] To more clearly illustrate the technical solutions of the embodiments of this application, the accompanying drawings used in the embodiments of this application will be briefly introduced below. It should be understood that the following drawings only show some embodiments of this application and should not be regarded as a limitation of the scope. For those skilled in the art, other related drawings can be obtained based on these drawings without creative effort.

[0013] Figure 1 is a flowchart illustrating a method for identifying and processing archives according to an embodiment of this application; Figure 2 is a flowchart illustrating a method for identifying and processing archives according to an embodiment of this application; Figure 3 is a structural diagram illustrating an apparatus for identifying and processing archives according to an embodiment of this application; Figure 4 is a structural diagram illustrating an electronic device according to an embodiment of this application. Detailed Implementation

[0014] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of this application, and not all of the embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application. Unless otherwise defined, the technical or scientific terms used in this application should have the ordinary meaning understood by those skilled in the art. The words "first," "second," and similar terms used in this application do not indicate any order, quantity, or importance, but are only used to distinguish different components. The words "comprising" or "including," etc., mean that the element or object preceding the word covers the element or object listed after the word and its equivalents, but do not exclude other elements or objects. The words "connected," "coupled," or "connected," etc., are not limited to physical or mechanical connections, but can include electrical connections, whether direct or indirect. "Up," "down," "left," "right," etc., are only used to indicate relative positional relationships. When the absolute position of the described object changes, the relative positional relationship may also change accordingly.

[0015] Currently, in fields such as forensic identification, financial risk control, property rights transactions, art auctions, and the management of important historical archives, the authenticity of key archival documents is the cornerstone of these businesses. These key documents, such as property ownership certificates, loan contracts, judicial judgments, calligraphy and paintings by famous figures, and important approvals, possess extremely high value. Therefore, they have become primary targets for forgery and alteration by criminals. Forgery methods range from simple photocopying and printing to sophisticated handwriting imitation, seal forgery, and paper aging techniques, constantly evolving and posing a serious challenge to traditional authentication methods.

[0016] In one example, the main methods for authenticating archival documents include: expert authentication: relying on experienced experts to make judgments through sensory means such as sight, touch, and smell, and with the help of simple tools such as magnifying glasses and ultraviolet lamps. However, this method is highly subjective, the results are difficult to quantify and reproduce, it is inefficient, and the training period for experts is long and costly.

[0017] Single-technology detection: For example, using only OCR technology to compare text content with electronic archives cannot detect physical forgeries (such as forged signatures and seals). Or, using only a spectrometer to analyze the ink composition at a single point cannot verify the consistency of the entire content and the integrity of its circulation history. Single-technology detection has obvious "shortcomings" and is easily bypassed by targeted forgery methods.

[0018] Therefore, the accuracy and efficiency of existing methods for authenticating archives are both poor.

[0019] In one example, existing technologies lack full lifecycle tracing. Traditional authentication methods are mostly static and isolated, focusing only on the document itself while ignoring its entire lifecycle history from creation and circulation to archiving. A physically authentic document can be questionable in its authenticity if its circulation history is questionable.

[0020] The file identification and processing method provided in this application embodiment can be applied to electronic devices, terminal devices, file identification and processing devices, or other devices or equipment capable of executing this embodiment, and there are no limitations on this application. In this embodiment, the execution subject is described as an electronic device.

[0021] The terminal can be a user equipment (UE) such as a mobile phone, smartphone, laptop computer, digital broadcast receiver, personal digital assistant (PDA), or tablet computer (PAD), handheld device, in-vehicle device, wearable device, computing device, or other processing device connected to a wireless modem, mobile station (MS), or mobile terminal. This terminal has the ability to communicate with one or more core networks via a radio access network (RAN).

[0022] The preferred embodiments of this application are described below with reference to the accompanying drawings. It should be understood that the preferred embodiments described herein are for illustration and explanation only and are not intended to limit this application. Furthermore, the embodiments and features in the embodiments of this application can be combined with each other without conflict.

[0023] Figure 1 is a flowchart illustrating a document identification processing method provided in an embodiment of this application. As shown in Figure 1, the method may include: step S101, acquiring multimodal data of the document to be identified; wherein, the multimodal data includes at least one of the following: hyperspectral image data, visible light image data, and identification information.

[0024] For example, multimodal data of the document to be identified is collected. For instance, when a user places the document to be identified into the identification device, the device automatically performs a pipeline scan of the document, simultaneously collecting the following three types of information: Hyperspectral image data: The hyperspectral camera images the document in hundreds of different wavelengths (from visible light to near infrared), generating a three-dimensional hyperspectral data cube (Hyper-Spectral Image, HSI).

[0025] Visible light image data: used for human visual observation and subsequent OCR processing. For example, visible light image data is a high-resolution color image.

[0026] Identification information: Extract unique identification information ID from the document to be identified through Optical Character Recognition (OCR) or barcode scanning. For example, identification information includes document number, contract number, etc.

[0027] Optionally, any one or two of the multimodal data can be collected; there is no limitation on this.

[0028] Step S102: Compare the hyperspectral image data and / or visible light image data with the data in the preset standard feature template library to obtain hyperspectral result information and / or visible light result information; query the lifecycle transfer record associated with the identification information in the preset blockchain network.

[0029] For example, firstly, authentic document data is pre-archived during the preparation stage, and a standard feature template library is constructed. Specifically, for critical document types requiring protection, once the document is confirmed to be authentic, a comprehensive "feature collection" is performed using authentication equipment. For instance, hyperspectral image data is collected to extract the "standard spectral fingerprints" of paper, ink in specific locations (such as signatures), and seal ink; simultaneously, the document is scanned at high resolution and OCR is performed to obtain an authoritative standard text file. These physical features (i.e., hyperspectral image data) and content features (i.e., the standard text file), along with the document's unique ID, are encrypted and stored in the "standard feature template library."

[0030] Then, the hyperspectral image data and visible light image data are compared with data in a preset standard feature template library to obtain hyperspectral result information and visible light result information. Alternatively, the hyperspectral image data is compared with data in a preset standard feature template library to obtain hyperspectral result information. Alternatively, the visible light image data is compared with data in a preset standard feature template library to obtain visible light result information.

[0031] For example, three analysis engines are launched in parallel to analyze and process the collected data: Physical layer analysis engine: Extracts the spectral curves from the hyperspectral image data of the archive to be identified, which are at the same location (such as the seal or signature) in the standard feature template library, and uses spectral matching algorithms (such as SAM, SID) to calculate the similarity score between the spectral curves and the "standard spectral fingerprints" in the standard feature template library.

[0032] Content layer analysis engine: Performs OCR on high-resolution color images to extract all text content. Then, the extracted text content is precisely compared with a "standard text file" retrieved based on the file ID to identify any character-level differences, thereby obtaining the text difference rate.

[0033] Historical layer analysis engine: Using the file ID as an index, it queries the file management consortium blockchain associated with the file ID to obtain and verify the complete and tamper-proof flow log of the file corresponding to that file ID from its generation, approval, each handover, and each viewing. It checks for any abnormal or unauthorized operation records and ultimately determines the blockchain record integrity flag, which indicates the completeness of the blockchain record.

[0034] Finally, the system will query the lifecycle transfer records associated with the identification information in the pre-defined blockchain network and determine whether there are any anomalies in the lifecycle transfer records at each stage.

[0035] Step S103: Integrate the target result information and generate the identification result information of the target result information through the preset decision model; wherein, the target result information includes at least one of the following: hyperspectral result information, visible light result information, and life cycle flow record.

[0036] For example, based on the multimodal data obtained in the previous step, at least one of the hyperspectral result information, visible light result information, and life cycle flow record is selected as the target result information, and multimodal information fusion and final decision are performed.

[0037] Specifically, all the quantitative results obtained from the above three layers of analysis (such as spectral similarity scores, text difference rates, blockchain record integrity markers, etc.) are aggregated and input into a pre-trained intelligent fusion decision model (such as a weighted scoring model, a Bayesian network, or a small neural network, etc., without limitation). This model comprehensively evaluates all evidence and finally outputs an authentication result. The authentication result indicates the authenticity of the document to be authenticated, and includes an overall "authenticity" probability score (e.g., 99.5%), as well as a detailed and interpretable authentication report. The report lists the detailed results of each layer of analysis and the evidence supporting the final conclusion.

[0038] The method provided in this application embodiment acquires multimodal data of a file to be identified; wherein the multimodal data includes at least one of the following: hyperspectral image data, visible light image data, and identification information. The hyperspectral image data and / or visible light image data are compared with data in a preset standard feature template library to obtain hyperspectral result information and / or visible light result information; the lifecycle transfer record associated with the identification information is queried in a preset blockchain network. The target result information is fused, and identification result information of the target result information is generated through a preset decision model; wherein the target result information includes at least one of the following: hyperspectral result information, visible light result information, and lifecycle transfer record. This solution constructs a unified, cross-verified, and automated authentication system that integrates "physical material (i.e., hyperspectral image data)," "information content (i.e., visible light image data)," and "transfer history (determined based on identification information)." This system fuses evidence from different dimensions and enables intelligent decision-making, improving the comprehensiveness, cross-verification capability, accuracy, and objectivity of the authentication process. It achieves highly efficient automation of the authentication process, addressing the core problems of existing methods for authenticating archives, such as limited authentication dimensions, inability to detect advanced forgery techniques, lack of process evidence chains, and subjective and inefficient conclusions. This solution can improve the accuracy and efficiency of archive authentication.

[0039] Figure 2 is a flowchart illustrating a method for identifying and processing archives provided in this application. As shown in Figure 2, this embodiment describes the method in detail based on the embodiment in Figure 1. The method includes: step S201, collecting multimodal historical data of archives that have been confirmed as genuine; the multimodal historical data includes historical hyperspectral image data, historical visible light image data, and historical identification information.

[0040] For example, first, a standard acquisition process is executed. Specifically, when an authentic document (such as a newly signed important contract) needs to be archived, the document is placed into the authentication device. The authentication device executes a standard acquisition procedure to generate an "authentic digital document package," which contains: historical hyperspectral image data: namely, the original hyperspectral data cube (HSI), which is a three-dimensional data (x, y, λ) where x, y are spatial pixel coordinates and λ is the wavelength channel.

[0041] Standard Spectral Fingerprint Database: Operators select several key areas on the interface, such as official seals, legal representative signatures, and specific printing areas. The average spectral curve of all pixels within these areas is extracted and stored as the "standard spectral fingerprint" (i.e., preset spectral curve) of the file. Each fingerprint is associated with an area label (e.g., "Official Seal 1") and coordinates.

[0042] Historical visible light image data: that is, authoritative standard text files. After high-resolution images are finely calibrated by OCR, an accurate TXT or JSON file is generated, storing all text content.

[0043] Historical identification information: namely, the initial blockchain evidence, calculating the hash value of the entire "genuine digital archive package", along with the archive ID, creation time, operator and other information, is recorded on the archive consortium chain as a genesis transaction.

[0044] Step S202: From historical hyperspectral image data, label and extract the spectral features of at least one key region as standard spectral features and store them.

[0045] For example, spectral features include an average spectral curve. An operator selects several key areas on the interface, such as an official seal, a legal representative's signature, or a specific printed area. Based on the operator's actions, the electronic device extracts the average spectral curve of all pixels within these key areas and stores it as the "standard spectral fingerprint" of the document. Each fingerprint is associated with a region label (e.g., "Official Seal 1") and coordinates, thus establishing a standard spectral fingerprint database.

[0046] Step S203: Perform text recognition and correction on historical visible light image data, and generate standard text files for storage.

[0047] For example, text recognition and correction are performed on historical visible light image data to generate standard text files for storage.

[0048] Step S204: Link and store the standard spectral features, standard text files, and historical identification information to form a genuine product feature template library.

[0049] For example, the hyperspectral image data, authoritative standard text files, and initial blockchain evidence are all ultimately linked and stored in the genuine product feature template library.

[0050] Step S205: Obtain multimodal data of the file to be identified; wherein, the multimodal data includes at least one of the following: hyperspectral image data, visible light image data, and identification information.

[0051] For example, the identification device is a dark box that integrates a light source, a transmission device, and a sensor.

[0052] Among them, the light source provides broadband uniform illumination from the visible light to the near-infrared band.

[0053] Hyperspectral cameras: Push-broom or snapshot hyperspectral cameras are typically used. As the files move at a constant speed on a precision conveyor belt, the push-broom camera scans line by line, acquiring complete spectral information at each spatial point, and finally stitching them together to form a complete hyperspectral data cube.

[0054] Auxiliary HD camera: Deployed alongside or overlapping the optical path of the hyperspectral camera to simultaneously capture high-resolution RGB color images.

[0055] Control system: An industrial PC controls the conveyor belt speed, light source switching, and synchronous triggering of all cameras to ensure that the acquired multimodal data is spatially aligned.

[0056] In this step, at the start of the identification process, multimodal data of the file to be identified is acquired based on the user's identification actions. This multimodal data includes at least one of the following: hyperspectral image data, visible light image data, and identification information.

[0057] Step S206: Extract the actual spectral features of at least one key region from the hyperspectral image data, and calculate the matching degree with the standard spectral features of the same region in the preset standard feature template library to obtain hyperspectral result information.

[0058] In one example, S206 includes: spatially aligning hyperspectral image data with standard template images in a preset standard feature template library based on an image registration algorithm; extracting average spectral curves of the same key regions in the hyperspectral image data for each preset key region in the standard feature template library; calculating the similarity between the preset spectral curves of the key regions and the average spectral curves of the same key regions; and obtaining a similarity score in the hyperspectral result information based on the similarity.

[0059] For example, an image registration algorithm (such as SIFT feature matching) is first used to precisely spatially align the hyperspectral image data of the document to be authenticated with the retrieved standard template image (i.e., the template image of the genuine document) to ensure that the same positions can be matched. The standard template image can be an image identical to the document to be authenticated; or an image of the same type as the document to be authenticated; or, for example, if the document to be authenticated is a contract, and different contracting parties have different contract formats, the standard template image can be an image corresponding to the contract format, etc. Here, the standard template image is merely an example and is not limited thereto.

[0060] For each key region defined in the standard feature template library (e.g., "Official Seal 1"), the corresponding region is found in the aligned HSI to be tested, and the average spectral curve is extracted. Then, algorithms such as Spectral Angle Mapper (SAM) or Spectral Information Divergence (SID) are used to calculate the "distance" or "similarity" between the average spectral curve to be tested and the standard preset spectral curve, outputting hyperspectral result information, which includes a similarity score between 0 and 1. For example, for red inkpads from different manufacturers and batches, the spectral curves will have slight but measurable differences in specific wavelengths. Counterfeit seals are difficult to match the spectral fingerprint of genuine products, and thus the differences in spectral curves can be distinguished based on the output similarity score.

[0061] And / or, in step S207, perform text recognition on the visible light image data to obtain the text content to be identified; search for the standard text file corresponding to the identification information in the preset standard feature template library, and compare the difference between the text content to be identified and the standard text file to obtain the visible light result information.

[0062] In one example, “comparing the text content to be identified with a standard text file to obtain visible light result information” includes: based on a string difference comparison algorithm, comparing the text content to be identified with a standard text file to obtain visible light result information; the visible light result information includes any one or more of the following: total difference rate, difference list.

[0063] For example, the OCR engine is run to perform text recognition on the acquired high-resolution color image to obtain the text content to be identified. From a preset standard feature template library, authoritative standard text files corresponding to the identification information are called. A string difference comparison algorithm (such as the Diff algorithm) is used to compare the text content to be identified with the standard text file word by word and sentence by sentence, outputting visible light result information. The visible light result information includes: total difference rate, and a difference list (including location, original text, and text to be tested). Therefore, any addition, deletion, or modification of content can be accurately detected, such as tampering with contract amount figures or modifying key clauses.

[0064] Step S208: Query the lifecycle transfer records associated with the identification information in the preset blockchain network.

[0065] In one example, S208 includes: querying the historical transaction records of each stage corresponding to the identification information in the blockchain network, sorting the historical transaction records of each stage by time, and generating a lifecycle log flow record of the file to be identified; verifying whether the lifecycle log flow record has any one or more of the following anomalies: continuity of blockchain records, compliance of operation permissions at each stage, or state anomaly; state anomaly refers to the matching between the current state of the file to be identified and the actual stage state.

[0066] For example, using the unique identifier ID obtained from the document to be authenticated, the smart contract query interface of the document consortium blockchain is called through the Software Development Kit (SDK). All historical transaction records associated with this ID are retrieved and sorted by time to form a complete lifecycle log, i.e., the lifecycle log flow record. Then, the completeness of the lifecycle log flow record and the continuity of the hash chain are automatically verified, and the following anomalies are checked: Blockchain record continuity: Are there any discontinuous time periods?

[0067] Compliance of operational permissions at each stage: Whether there are any permission abnormalities. Permission abnormalities include unauthorized IP addresses, user IDs that have been viewed, or handover operations.

[0068] Status Abnormal: Whether the current status of the file to be identified matches the actual stage status (e.g., under identification).

[0069] Finally, it outputs a Boolean "Historically Innocent" flag, along with a detailed source tracing report.

[0070] Step S209: Perform spectral clustering analysis on the overall background area of ​​the hyperspectral image data to generate clustering result information; the clustering result information indicates whether there are material abnormal areas in the hyperspectral image data.

[0071] For example, this involves detecting tampering with hyperspectral image data. Specifically, spectral clustering analysis is performed on the background area of ​​the entire sheet of paper in the hyperspectral image data to generate clustering results. If the paper has been patched or pasted, the material of the tampered area will differ from the original paper, and its spectral characteristics will also differ, forming a distinct anomalous cluster in the clustering results. Therefore, it is possible to accurately detect whether there are areas of material abnormality in the hyperspectral image data.

[0072] Step S210: Integrate the target result information and generate the identification result information of the target result information through a preset decision model; wherein, the target result information includes at least one of the following: hyperspectral result information, visible light result information, and life cycle flow record.

[0073] In one example, S210 includes: if the target result information includes hyperspectral result information, visible light result information and life cycle flow record, then construct the target result information into a fusion feature vector; input the fusion feature vector into a preset decision model to generate the identification result information and confidence level of the target result information.

[0074] For example, this step involves summarizing all evidence and making a final judgment. First, if the target result information includes hyperspectral result information, visible light result information, and lifecycle flow records, a fusion feature vector is constructed. All the quantitative results obtained from the three-layer analysis are used to construct a fusion feature vector V_fusion, for example: V_fusion=[spectral similarity_official seal, spectral similarity_signature, tampered area area ratio, total text difference rate, key field difference number, historical innocence marker,...]; Second, the intelligent decision-making model can be a weighted scoring model or a machine learning model, etc., without limitation. For a weighted scoring model, the simplest way is to assign an expert weight to each feature in V_fusion and then perform a weighted sum to obtain a total score. For a machine learning model, a more advanced approach is to use a pre-trained classifier (such as a support vector machine (SVM), random forest, or a small multilayer perceptron (MLP)) that is trained under supervision using a large number of known true and false archives and the corresponding fusion feature vectors of these known true and false archives. During real-time identification, the new V_fusion is input into the classifier, and the output is the probability of "true". This method can learn the complex non-linear relationships between features and is more accurate than simple weighted summation.

[0075] The fused feature vectors are input into a pre-defined decision model to generate identification results and confidence levels. The identification results include an overall "true" probability score (e.g., 99.5%) and a detailed, interpretable identification report. The report lists the detailed results of each level of analysis and the evidence supporting the final conclusion. Regardless of the decision model used, a graphically presented identification report will be generated, for example, in PDF format. The report includes: Overall Conclusion: A clear conclusion of "true," "false," or "highly questionable," along with a quantified confidence score.

[0076] Detailed Evidence: The report clearly outlines the analysis results across three dimensions: physical layer, content layer, and historical layer. For example, it uses charts to illustrate the comparison of spectral curves, highlighted text to highlight content differences, and a timeline to display the complete blockchain circulation history. This report provides comprehensive and objective evidentiary support for manual review and subsequent legal proceedings.

[0077] The method provided in this application embodiment collects multimodal historical data of archives whose authenticity has been confirmed. This multimodal historical data includes historical hyperspectral image data, historical visible light image data, and historical identification information. From the historical hyperspectral image data, spectral features of at least one key region are labeled and extracted as standard spectral features and stored. Text recognition and correction are performed on the historical visible light image data to generate and store standard text files. The standard spectral features, standard text files, and historical identification information are associated and stored to form a genuine artifact feature template library. Multimodal data of the archive to be authenticated is obtained; wherein, the multimodal data includes at least one of the following: hyperspectral image data, visible light image data, and identification information. The actual spectral features of at least one key region are extracted from the hyperspectral image data, and the matching degree is calculated with the standard spectral features of the same region in the preset standard feature template library to obtain hyperspectral result information. Text recognition is performed on the visible light image data to obtain the text content to be authenticated; the standard text file corresponding to the identification information is searched in the preset standard feature template library, and the difference between the text content to be authenticated and the standard text file is compared to obtain visible light result information. The system queries the lifecycle transfer records associated with the identification information in a pre-defined blockchain network. Spectral clustering analysis is performed on the overall background area of ​​the hyperspectral image data to generate clustering results; these results indicate whether the hyperspectral image data contains areas of material abnormality. The target result information is then integrated, and an identification result is generated through a pre-defined decision model. The target result information includes at least one of the following: hyperspectral result information, visible light result information, and lifecycle transfer records. This solution constructs a unified, cross-validated, and automated identification system that integrates "physical material (i.e., hyperspectral image data)," "information content (i.e., visible light image data)," and "transfer history (determined based on identification information)." By fusing evidence from different dimensions and enabling intelligent decision-making, it improves the comprehensiveness, cross-validation capability, accuracy, and objectivity of the identification dimensions. This achieves highly efficient automation of the identification process, addressing the core problems of existing methods for authenticating archives: limited identification dimensions, inability to detect advanced forgery techniques, lack of process evidence chains, and subjective and inefficient conclusions. This improves the accuracy and efficiency of archive authentication.

[0078] In one example, this application embodiment also provides a document identification and processing system, which includes at least: a hyperspectral imaging module; a high-resolution image scanning and OCR module; a communication module for connecting to a blockchain network; and a processor configured with a multimodal information fusion decision model. This system is a deeply integrated hardware and software automated system, a complete system integrating precision optical instruments, automated control, and complex AI algorithms, with a focus on the seamless integration and collaborative operation of its components.

[0079] Optionally, based on an authentication device integrating multiple sensors and a powerful back-end analysis platform, evidence from different dimensions is fused and intelligently decided: 1) At the physical level, sensing technologies exceeding the limits of the human eye (such as hyperspectral imaging) are used to capture the inherent "spectral fingerprints" of materials such as paper, ink, and printing pads to identify the authenticity and consistency of the materials; 2) At the content level, AI technology is used to accurately extract text and image information from the document and compare it with an authoritative digital copy; 3) At the historical level, blockchain technology is used to securely query and verify every transfer record of the document since its creation to ensure its "pure origin." Ultimately, the system needs to integrate these three independent authentication results through an intelligent fusion decision model to arrive at a comprehensive, quantitative, and interpretable conclusion on authenticity.

[0080] Therefore, hyperspectral imaging technology, primarily used in remote sensing and agriculture, is creatively applied to the material identification and tamper detection of paper archives. The novel archival authentication concept, combining "physical fingerprint (hyperspectral) + content fingerprint (OCR) + historical fingerprint (blockchain)," offers the following benefits: Comprehensiveness of authentication dimensions and cross-verification capability: This invention, for the first time, integrates three independent authentication dimensions—physical material, textual content, and circulation history—into a unified automated system. The three-layered evidence chain can corroborate each other, greatly increasing the difficulty of forgery; any single-layer forgery is unlikely to pass multimodal cross-verification.

[0081] Powerful detection capabilities against advanced counterfeiting techniques: Hyperspectral imaging technology can capture subtle differences in materials that are indistinguishable to the human eye, effectively identifying sophisticated physical counterfeiting techniques such as seal forgery, signature imitation, and paper replacement, with extremely high technical barriers.

[0082] Extremely high accuracy and objectivity: The system's identification process is entirely based on objective data collection and quantitative algorithmic analysis, eliminating the interference of human factors. The machine learning-based fusion decision-making model ensures that the final identification conclusions are more accurate and reliable than any single technology or human judgment.

[0083] The authentication process has been highly automated, reducing the expert authentication process from several hours or even days to a few minutes of automated machine work. This makes it possible to conduct large-scale, routine spot checks or full inspections of key documents, greatly improving the efficiency of risk control.

[0084] Corresponding to the above method, this application embodiment also provides an archive identification and processing device, as shown in FIG3. The device includes: an acquisition module 41, used to acquire multimodal data of the archive to be identified; wherein, the multimodal data includes at least one of the following: hyperspectral image data, visible light image data, and identification information; a comparison module 42, used to compare the hyperspectral image data and / or the visible light image data with data in a preset standard feature template library, respectively, to obtain hyperspectral result information and / or visible light result information; querying the lifecycle transfer record associated with the identification information in a preset blockchain network; and a generation module 43, used to fuse target result information and generate identification result information of the target result information through a preset decision model; wherein, the target result information includes at least one of the following: the hyperspectral result information, the visible light result information, and the lifecycle transfer record.

[0085] The functions of each functional unit of the file identification and processing device provided in the above embodiments of this application can be implemented through the above method steps. Therefore, the specific working process and beneficial effects of each unit in the file identification and processing device provided in the embodiments of this application will not be repeated here.

[0086] This application embodiment also provides an electronic device, as shown in FIG4, including a processor 510, a communication interface 520, a memory 530 and a communication bus 540, wherein the processor 510, the communication interface 520 and the memory 530 communicate with each other through the communication bus 540.

[0087] The memory 530 is used to store computer programs; the processor 510 is used to execute the above steps when executing the program stored in the memory 530.

[0088] The communication bus mentioned above can be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, etc. This communication bus can be divided into address bus, data bus, control bus, etc. For ease of illustration, only one thick line is used to represent it in the diagram, but this does not mean that there is only one bus or one type of bus.

[0089] The communication interface is used for communication between the aforementioned electronic devices and other devices.

[0090] The memory may include random access memory (RAM) or non-volatile memory (NVM), such as at least one disk storage device. Optionally, the memory may also be at least one storage device located remotely from the aforementioned processor.

[0091] The processors mentioned above can be general-purpose processors, including central processing units (CPUs), network processors (NPs), etc.; they can also be digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components.

[0092] Since the implementation methods and beneficial effects of the various devices in the above embodiments of the electronic device can be achieved by referring to the steps in the embodiment shown in Figure 1, the specific working process and beneficial effects of the electronic device provided in this application embodiment will not be repeated here.

[0093] In another embodiment provided in this application, a computer-readable storage medium is also provided, which stores instructions that, when executed on a computer, cause the computer to perform any of the file identification and processing methods described in the above embodiments.

[0094] In another embodiment provided in this application, a computer program product containing instructions is also provided, which, when run on a computer, causes the computer to execute any of the file identification and processing methods described in the above embodiments.

[0095] Those skilled in the art will understand that the embodiments in this application can be provided as methods, systems, or computer program products. Therefore, the embodiments in this application can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, the embodiments in this application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0096] This application describes embodiments of methods, apparatus (systems), and computer program products according to embodiments of this application with reference to flowchart illustrations and / or block diagrams. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, create means for implementing the functions specified in one or more blocks of the flowchart illustrations and / or one or more blocks of the block diagrams.

[0097] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means that implement the functions specified in one or more flowcharts and / or one or more block diagrams.

[0098] These computer program instructions may also be loaded onto a computer or other programmable data processing apparatus to cause a series of operational steps to be performed on the computer or other programmable apparatus to produce a computer-implemented process, such that the instructions, which execute on the computer or other programmable apparatus, provide steps for implementing the functions specified in one or more flowcharts and / or one or more block diagrams.

[0099] Although preferred embodiments have been described in this application, those skilled in the art, upon learning the basic inventive concept, can make other changes and modifications to these embodiments. Therefore, the appended claims are intended to be interpreted as including the preferred embodiments as well as all changes and modifications falling within the scope of the embodiments of this application.

[0100] Obviously, those skilled in the art can make various modifications and variations to the embodiments of this application without departing from the spirit and scope of the embodiments of this application. Therefore, if these modifications and variations to the embodiments of this application fall within the scope of the claims in this application and their equivalents, then this application also intends to include these modifications and variations.

Claims

1. A method for identifying and processing archives, characterized in that, The method includes: acquiring multimodal data of the file to be identified; wherein the multimodal data includes at least one of the following: hyperspectral image data, visible light image data, and identification information; comparing the hyperspectral image data and / or the visible light image data with data in a preset standard feature template library to obtain hyperspectral result information and / or visible light result information; querying the lifecycle transfer record associated with the identification information in a preset blockchain network; fusing the target result information and generating identification result information of the target result information through a preset decision model; wherein the target result information includes at least one of the following: the hyperspectral result information, the visible light result information, and the lifecycle transfer record.

2. The method as described in claim 1, characterized in that, The hyperspectral image data and / or the visible light image data are compared with data in a preset standard feature template library to obtain hyperspectral result information and / or visible light result information. This includes: extracting the actual spectral features of at least one key region from the hyperspectral image data, calculating the matching degree with the standard spectral features of the same region in the preset standard feature template library to obtain hyperspectral result information; and / or performing text recognition on the visible light image data to obtain the text content to be identified; searching for a standard text file corresponding to the identification information in the preset standard feature template library, and comparing the text content to be identified with the standard text file to obtain visible light result information.

3. The method as described in claim 2, characterized in that, Extracting actual spectral features of at least one key region from the hyperspectral image data and calculating the matching degree with the standard spectral features of the same region in a preset standard feature template library to obtain hyperspectral result information includes: spatially aligning the hyperspectral image data with standard template images in the preset standard feature template library based on an image registration algorithm; extracting the average spectral curve of the same key region in the hyperspectral image data for each preset key region in the standard feature template library; calculating the similarity between the preset spectral curve of the key region and the average spectral curve of the same key region; and obtaining a similarity score in the hyperspectral result information based on the similarity.

4. The method as described in claim 2, characterized in that, The visible light result information is obtained by comparing the text content to be identified with the standard text file, based on a string difference comparison algorithm. The visible light result information includes any one or more of the following: total difference rate, difference list.

5. The method as described in claim 1, characterized in that, The method further includes: performing spectral clustering analysis on the overall background region of the hyperspectral image data to generate clustering result information; the clustering result information characterizes whether there are material abnormal regions in the hyperspectral image data.

6. The method as described in claim 1, characterized in that, Querying the lifecycle flow records associated with the identification information in a preset blockchain network includes: querying the historical transaction records of each stage corresponding to the identification information in the blockchain network, sorting the historical transaction records of each stage by time, and generating the lifecycle log flow record of the file to be identified; verifying whether the lifecycle log flow record has any one or more of the following anomalies: continuity of blockchain records, compliance of operation permissions at each stage, or abnormal status; the abnormal status refers to the matching between the current status of the file to be identified and the actual stage status.

7. The method as described in claim 1, characterized in that, The method involves fusing target result information and generating identification result information for the target result information through a preset decision model. This includes: if the target result information includes the hyperspectral result information, the visible light result information, and the life cycle flow record, then constructing the target result information into a fused feature vector; inputting the fused feature vector into the preset decision model to generate identification result information and confidence level for the target result information.

8. The method according to any one of claims 2-4, characterized in that, Before acquiring the multimodal data of the archive to be authenticated, the process further includes: collecting historical multimodal data of archives that have been confirmed as authentic; the historical multimodal data includes historical hyperspectral image data, historical visible light image data, and historical identification information; from the historical hyperspectral image data, marking and extracting the spectral features of at least one key region as the standard spectral features for storage; performing text recognition and correction on the historical visible light image data to generate a standard text file for storage; and associating and storing the standard spectral features, the standard text file, and the historical identification information to form the authentic feature template library.

9. A document identification and processing device, characterized in that, The device includes: an acquisition module for acquiring multimodal data of the file to be identified; wherein the multimodal data includes at least one of the following: hyperspectral image data, visible light image data, and identification information; a comparison module for comparing the hyperspectral image data and / or the visible light image data with data in a preset standard feature template library to obtain hyperspectral result information and / or visible light result information; querying the lifecycle transfer record associated with the identification information in a preset blockchain network; and a generation module for fusing target result information and generating identification result information of the target result information through a preset decision model; wherein the target result information includes at least one of the following: the hyperspectral result information, the visible light result information, and the lifecycle transfer record.

10. An electronic device, characterized in that, The electronic device includes a processor, a communication interface, a memory, and a communication bus, wherein the processor, the communication interface, and the memory communicate with each other through the communication bus; the memory is used to store computer programs; and the processor, when executing the program stored in the memory, implements the method described in any one of claims 1-8.