Training methods for document authenticity verification models and methods and equipment for verifying document authenticity.

By injecting minute artifacts into document images and constructing multi-level question text to train a multimodal model, the problems of local forgery trace recognition and step-by-step identification output in document image authentication are solved, achieving more stable document authenticity authentication and result verifiability.

CN122493470APending Publication Date: 2026-07-31HANGZHOU ANT KUAI TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
HANGZHOU ANT KUAI TECHNOLOGY CO LTD
Filing Date
2026-04-27
Publication Date
2026-07-31

AI Technical Summary

Technical Problem

Existing technologies struggle to simultaneously ensure stable identification of fine-grained forgery traces and step-by-step identification output during automatic identification of document images. In particular, when forgery traces manifest as local high-frequency textures, edge detail variations, or abnormal printing features in specific areas, the model output is not stable enough and lacks step-by-step analysis results.

Method used

By generating a base document image and injecting tiny artifacts, a multi-layered question text is constructed. A multimodal model is trained to identify local forgery traces and output step-by-step identification information, including overall authenticity judgment, suspicious area location, and abnormal detail description.

Benefits of technology

It improves the ability to identify local forgery traces and outputs step-by-step identification information corresponding to the identification process, thereby enhancing the verifiability of the results and the stability of the model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122493470A_ABST
    Figure CN122493470A_ABST
Patent Text Reader

Abstract

This invention relates to the field of document image authentication technology, and discloses a training method for a document authenticity authentication model, a method for authenticating documents, and an apparatus. The training method includes: generating a base document image based on standard document information of a preset type of document; injecting minute artifacts into the base document image according to an artifact generation strategy corresponding to the forgery process to obtain a sample forged document image; constructing a multi-level question text corresponding to the sample forged document image and descriptive information of the minute artifacts; inputting the sample forged document image and the multi-level question text into a multimodal model to be trained, and obtaining the sample authenticity authentication result; adjusting the model parameters based on the deviation between the sample authenticity authentication result and the descriptive information to obtain the document authenticity authentication model. This method can improve the ability to identify local forgery traces in document images and output step-by-step authentication information.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of document image identification technology, specifically to a training method for a document authenticity identification model and a method and device for identifying document authenticity. Background Technology

[0002] In the field of document image authentication technology, document authenticity verification typically refers to the process of determining whether a document, such as a resident ID card, passport, driver's license, business license, or other legally valid document, has been forged, altered, or exhibits abnormal manufacturing marks. This type of verification often includes not only the text and graphic content of the document but also information such as printing texture, local edges, surface reflection, the presentation of anti-counterfeiting features, and the content relationships between fields. Therefore, document authenticity verification is not simply a character recognition or image classification problem, but a comprehensive processing procedure involving image detail analysis and content consistency analysis.

[0003] With the widespread adoption of online identity verification, remote account opening, remote business approval, and online real-name services, more and more document verification processes are shifting from manual counter reviews to automated review methods based on image capture. In these scenarios, users typically upload document images via mobile terminals or other image capture devices, and servers or local devices then determine authenticity based on the captured images. Because the capture process often involves variations in shooting angle, compression distortion, lighting changes, and fluctuations in image clarity, forgery traces in document images may not appear as significant defects, but rather as subtle, localized anomalies. For forged documents created through color printing, scanned and reprinted text, partial replacement, or low-quality copying, these anomalies typically manifest as fine-grained visual features such as printing dots, scan moiré patterns, localized font distortion, weakened anti-counterfeiting markings, or abnormal material textures.

[0004] To address the aforementioned document image authentication requirements, existing technologies include a type of implementation based on traditional computer vision models. This type of approach typically involves first preprocessing and region localization of the document image, then using convolutional neural networks, detection networks, character recognition modules, or rule comparison modules to determine the authenticity of the document. This approach can achieve a certain degree of automated recognition when the document layout is relatively stable, image acquisition conditions are relatively consistent, and the type of forgery is relatively fixed. However, this type of approach often relies on pre-defined local features, fixed templates, or a limited distribution of forged samples. When the document layout changes, or the forgery method shifts from obvious alteration to a more realistic partial imitation, the existing features relied upon by the model can easily deviate from the actual abnormal patterns, making it difficult to maintain stable recognition capabilities across different scenarios.

[0005] Furthermore, the aforementioned traditional computer vision solutions typically focus on binary classification results (true / false) or local anomaly scores as output, rarely providing hierarchical analysis information directly corresponding to forgery traces. For business processes requiring manual review, if the model only outputs true / false conclusions or local suspicious areas without further explanation of the anomaly type and specific basis for the suspicious areas, reviewers still need to rely on human experience to make judgments. While this improves the system's automation level, there is still room for improvement in evidence organization, result verification, and subsequent auditing.

[0006] Besides traditional computer vision solutions, existing technologies also include a type of approach that directly utilizes general multimodal models to perform question-and-answer judgments on document images. This type of approach typically inputs both the document image and the query text into the model, and the model outputs a textual result regarding the authenticity of the document. Since multimodal models inherently possess image-text joint understanding capabilities, this approach has certain advantages in open-ended question answering and semantic description. However, the pre-training targets of general multimodal models usually cover a wide range of natural images and general text tasks, and they may not necessarily develop stable specialized recognition capabilities for local fine-grained anomalies introduced by document manufacturing processes. Especially when document forgery traces mainly manifest as local high-frequency textures, edge detail variations, or abnormal printing features in specific areas, relying solely on general pre-training capabilities often makes it difficult to reliably distinguish between genuine document features and local anomalies caused by forgery processes.

[0007] Meanwhile, general multimodal models typically favor natural language responses in their output format. Without specialized training for document authentication tasks, their output tends to remain at the level of overall judgment or general description, rarely generating step-by-step analysis results that match the authentication process. For document authentication, if the model cannot progressively provide an overall judgment, locate suspicious areas, and explain anomalies in detail, even if the model possesses some recognition ability, it will be difficult to meet the requirements for verifiable results and a stable output format.

[0008] Therefore, existing technologies for automatically verifying the authenticity of documents using images often struggle to simultaneously address the following aspects: first, ensuring the model has a stable ability to identify localized, fine-grained technological anomalies in counterfeit documents; and second, enabling the model to output step-by-step analysis information corresponding to the verification process. Regarding these issues, there is still room for improvement in how to construct training samples suitable for document counterfeiting scenarios and train a document authentication model that can both identify localized forgery traces and output step-by-step verification information. Summary of the Invention

[0009] In view of this, embodiments of the present invention provide a training method for a document authenticity identification model and a method and device for identifying document authenticity, which at least alleviates the problem in the prior art that it is difficult to simultaneously take into account the ability to identify local fine-grained forgery traces and the ability to identify step-by-step identification output during the automatic identification of document images.

[0010] To achieve the above objectives, the training method for the document authenticity identification model provided in this embodiment of the invention includes: generating a base document image based on standard document information of a preset type of document, and injecting micro-artifacts into the base document image according to an artifact generation strategy corresponding to the forgery process to obtain a sample forged document image; constructing multi-level question text corresponding to the sample forged document image and descriptive information of the micro-artifacts, wherein the multi-level question text is used to periodically inquire about the authenticity of the document; inputting the sample forged document image and the multi-level question text into a multimodal model to be trained, and obtaining the sample authenticity identification result for the sample forged document image output by the multimodal model; adjusting the model parameters of the multimodal model based on the deviation between the sample authenticity identification result and the descriptive information of the micro-artifacts, and obtaining the document authenticity identification model when a preset training stopping condition is met.

[0011] In some embodiments, generating a base document image based on standard document information of a preset type of document includes: obtaining a blank document template image and anti-counterfeiting process information of a preset type of document, generating corresponding standard document content according to the anti-counterfeiting process information, and filling the standard document content into the corresponding position in the blank document template image to obtain the base document image.

[0012] In some embodiments, injecting micro-artifacts into a base document image according to an artifact generation strategy corresponding to the forgery process includes: adjusting preset parameters of the base document image based on forgery process information to inject micro-artifacts with at least one of preset spatial location, preset morphological features, and preset physical properties into the base document image. The micro-artifacts may include at least one of the following: overlaying printed halftone dots onto a portrait area and / or text area in the base document image; overlaying scanned moiré patterns onto the base document image; deforming the font in the base document image; removing or weakening anti-counterfeiting marks in the base document image; and replacing local textures in the base document image with preset material textures.

[0013] In some embodiments, the multi-layered question text used for periodically verifying the authenticity of a document includes: a first-layer text for verifying the overall authenticity of the document, a second-layer text for verifying the type of forgery features, and a third-layer text for verifying detailed forgery information. By decomposing the document authentication task into hierarchically defined sub-tasks, the multimodal model can be guided to progressively learn the output methods corresponding to the document authentication process.

[0014] This invention also provides a method for identifying the authenticity of a document, comprising: in response to an authentication request initiated for the document to be authenticated, acquiring a document image collected for the document to be authenticated; acquiring multi-level question text for periodically inquiring about the authenticity of the document; inputting the document image and the multi-level question text into a document authentication model, and acquiring the authentication result output by the document authentication model after inference; if the authentication result indicates that the document to be authenticated is a counterfeit document, outputting multiple step-by-step authentication information corresponding to the multi-level question text.

[0015] This invention also provides an electronic device, including a processor and a memory. The memory stores computer instructions, and the processor invokes the computer instructions to execute the steps in the aforementioned training method for the document authenticity verification model or the method for verifying the authenticity of documents.

[0016] Compared with the prior art, the embodiments of the present invention have at least the following beneficial effects: This invention generates a base document image based on standard document information and further injects minute artifacts into the base document image according to an artifact generation strategy corresponding to the forgery process. This enables the construction of sample forged document images containing local fine-grained anomaly features. After training a multimodal model based on such samples, it can more effectively learn local anomaly patterns in document forgery scenarios, thereby improving the ability to identify local forgery traces.

[0017] This invention employs a combination of sample forged document images and multi-level question text for model training. This allows the model to learn not only how to determine authenticity but also the step-by-step output methods corresponding to overall judgment, feature localization, and detailed explanations. Therefore, in the actual authentication phase, the document authentication model can output not only the authenticity conclusion but also step-by-step authentication information corresponding to the authentication process, thereby improving the verifiability of the results.

[0018] In some embodiments, by introducing logical consistency analysis related to document content, the model's ability to identify logical anomalies in content can be further enhanced, enabling the model to identify not only local visual anomalies related to manufacturing processes, but also content anomalies related to field relationships.

[0019] In some embodiments, by updating model parameters using an efficient fine-tuning method during model training, the existing text and image understanding capabilities of the baseline model can be preserved while reducing training resource consumption, thereby improving training efficiency and deployment feasibility. Attached Figure Description

[0020] To more clearly illustrate the technical solutions of the embodiments in this specification, the drawings used in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments recorded in this specification. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0021] Figure 1 This is a schematic diagram of the overall architecture of a document authenticity verification system according to one embodiment of the present invention.

[0022] Figure 2 This is the main flowchart of the training method for a document authenticity identification model in one embodiment of the present invention.

[0023] Figure 3 This is a schematic diagram of the generation of the base document image and the injection of micro-artifacts in one embodiment of the present invention.

[0024] Figure 4 This is a schematic diagram illustrating the correspondence between multi-level question text and descriptive information in one embodiment of the present invention.

[0025] Figure 5 This is a flowchart of a method for identifying the authenticity of documents in one embodiment of the present invention.

[0026] Figure 6 This is a schematic diagram of the data flow during the training phase in one embodiment of the present invention.

[0027] Figure 7 This is an input / output timing diagram of the identification stage in one embodiment of the present invention.

[0028] Figure 8 This is a schematic diagram of the training extension structure principle in one embodiment of the present invention.

[0029] Figure 9 This is a schematic diagram of an electronic device structure according to one embodiment of the present invention.

[0030] Figure 10 This is a schematic diagram of the input method of multi-level question text in the identification stage in one embodiment of the present invention. Detailed Implementation

[0031] To enable those skilled in the art to better understand the technical solutions in this specification, the technical solutions in the embodiments of this specification will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this specification, and not all embodiments. Based on the embodiments in this specification, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of this specification.

[0032] The technical solution of this application will be further described below with reference to the accompanying drawings. In the drawings, the same or similar reference numerals indicate the same or similar objects. The connection relationship between the objects can be a direct connection relationship or an indirect connection relationship established through an intermediate object. The process sequence in the drawings is used to illustrate a feasible processing order. Without affecting the data dependency relationship, some processes can also be executed in parallel or in other orders.

[0033] Figure 1 This diagram illustrates the overall structure of a document authentication system 100. System 100 may include a training server 110, an authentication server 120, a terminal device 130, a training data storage unit 140, and a model storage unit 150. The terminal device 130 may be a mobile terminal, a desktop terminal, a self-service verification terminal, or other electronic devices with image acquisition capabilities. The training server 110 is primarily used to generate training samples, train the document authentication model, and output the trained model file. The authentication server 120 is primarily used to receive images of documents to be authenticated, invoke the document authentication model to perform inference, and return the authentication result. The training data storage unit 140 is used to store standard document information, base document images, sample forged document images, multi-level question text, and descriptive information. The model storage unit 150 is used to store intermediate models during training and the completed document authentication model.

[0034] Figure 2 This illustrates the main flow of the training method. Figure 3 This illustrates the main process of the identification method. Figure 4 This illustrates the correspondence between the base document image and the sample forged document image. Figure 5 This demonstrates the organizational relationship between multi-level question texts and corresponding descriptive information. Figure 6 This shows the input-output relationship of the document authenticity verification model during the training phase. Figure 7 This illustrates the input organization method during the identification phase. Figure 8 This shows a schematic diagram of the electronic device. Figure 9 An alternative software logic structure is shown. The following description will unfold step by step around these figures.

[0035] In one embodiment of this application, the training server 110 can first read standard document information of a preset type of document from the training data storage unit 140. The preset type of document may include a resident ID card, passport, driver's license, business license, or other legal documents requiring image authentication. The standard document information is not limited to a fixed format and typically includes a blank document template image, field layout information, and anti-counterfeiting technology information used to characterize the visual features of a genuine document. If it is inconvenient to directly save complete genuine document samples in the actual deployment environment, only the desensitized and abstracted template information can be retained to reduce reliance on genuine document data.

[0036] exist Figure 2 In the illustrated process, the training does not begin by directly feeding real document images into the model. Instead, a stable and controllable training base is first formed. Correspondingly, the training server 110 can execute step S110: generating the base document image. The object processed in this step is not an arbitrary image, but standard document information that corresponds one-to-one with a preset type of document. The blank document template image 210 in the standard document information can provide a background pattern, fixed layout, area boundaries, and basic text framework. The field layout information 220 can define the positional relationships of key areas such as the name area, document number area, portrait area, and validity period area. The anti-counterfeiting technology information 230 is used to characterize the visual feature generation rules in real documents. The training server 110 can generate standard document content 240 based on this information, and then fill the standard document content 240 into the corresponding positions in the blank document template image 210 to obtain the base document image 250.

[0037] The standard document content 240 mentioned here is not required to be completely identical to the information of a real document holder. Its generation can be achieved through rule-based generation, dictionary sampling, template replacement, or other methods that can form a legal combination of fields. For example, in a resident ID card scenario, the name field can be extracted from a preset name database, the date of birth field can be sampled from a preset date range, the ID number field can be generated based on administrative division code, date of birth, and check digit rules, and the portrait area can use an authorized portrait sample or a uniformly styled simulated portrait. The resulting base document image 250 is structurally close to a real document image, but its content does not necessarily correspond to a specific real object. This approach has two advantages: first, it facilitates the batch generation of training samples; second, it reduces the privacy exposure risks associated with directly using real document images.

[0038] Figure 4The diagram illustrates the correspondence between the base document image 250 and the sample forged document image 260. In this diagram, 250a represents a normal human portrait area, 250b represents a normal text area, and 250c represents a normal anti-counterfeiting mark area. After generating the base document image 250, the training server 110 can continue to execute step S120: constructing the sample forged document image. In this step, the training server 110 does not simply superimpose random noise into the image, but rather rewrites the base document image 250 locally according to the artifact generation strategy corresponding to the forgery process, thereby generating the sample forged document image 260 with local anomalies while maintaining the overall document structure stability.

[0039] Micro-artifacts are not small images that exist independently of the document image; rather, they are anomalous visual patterns introduced into specific areas of the document by forgery techniques. These anomalous visual patterns often exhibit locational locality, interpretable causes, and observability. Locational locality means that micro-artifacts are typically concentrated at the edges of the portrait, text, security features, or localized textures, rather than being uniformly distributed across the entire document image. Interpretable causes mean that different types of micro-artifacts can be traced back to specific forgery techniques, such as color printing, scan-and-reprint, low-quality copying, or localized alteration. Observability means that these anomalies, even after being photographed, scanned, or compressed, can still be retained in the document image with certain visual characteristics, thus serving as a basis for training and identification.

[0040] To ensure the controllability of the generation process of minor artifacts, the training server 110 can adjust the preset parameters of the base document image 250 according to the artifact generation strategy. These preset parameters may include, but are not limited to, spatial location parameters, morphological parameters, texture parameters, intensity parameters, and region range parameters. Spatial location parameters define the region where the artifact injection occurs; morphological parameters define the geometric or textural shape of the artifact; texture parameters define local repetitive structures, frequency structures, or color distributions; intensity parameters define the contrast of the artifact relative to the background; and region range parameters define the area and boundaries of the affected region. By combining and controlling these parameters, the training server 110 can generate multiple different forgery samples on the same type of document template without causing all samples to fall into the exact same pattern.

[0041] In a common implementation, minor artifacts in the sample forged document image 260 can be injected using an image transformation function. Let the base document image be I. b (x,y), the sample forged document image after artifact generation is I. f Given an image of a forged document with region (x,y) as the region, an injection mask of region M(x,y), and an artifact transformation result of region A(x,y), the forged document image can be written as: If (x,y)=(1-M(x,y))×I b (x,y)+M(x,y)×A(x,y)(1) in, and Represents image coordinates; The value of can be between 0 and 1, and is used to characterize whether a pixel location belongs to the artifact injection region; when When =0, the corresponding position retains the original value of the base document image; when When = 1, the corresponding position uses the artifact transformation result; when Taking the median value can represent the boundary transition region. In this way, artifact injection can be limited to a local area, thus avoiding the distortion caused by uniform perturbation of the entire image.

[0042] If the artifact belongs to the category of localized printing dot anomalies, then A(x,y) can be obtained through local halftone transformation. If the artifact belongs to the scan moiré anomaly category, A(x,y) can be obtained through periodic texture overlay. If the artifact belongs to the font deformation anomaly category, A(x,y) can be obtained through local character transformation. If the artifact belongs to the anti-counterfeiting mark weakening anomaly category, A(x,y) can be obtained through local contrast reduction, edge blurring, or highlight suppression. By uniformly writing the image implementation of the artifact into A(x,y), the training server 110 can process multiple types of artifacts in the same main process without redefining the overall training framework for each type of artifact.

[0043] In documents like resident identity cards, the portrait area is often prone to printing artifacts and partial replacement artifacts, while the text area is more likely to expose problems such as insufficient printing clarity, abnormal stroke edges, or scaling distortion. The anti-counterfeiting mark area is more likely to expose problems such as abnormal brightness distribution, weakened light-changing effects, or missing boundaries. The training server 110 can set different artifact injection probabilities for these areas based on the document type and forgery technique. This probability is not required to be consistent across all embodiments. For areas where forgery is more common, the injection ratio can be increased; for areas where it is necessary to retain the real background, the perturbation frequency can be reduced to prevent the training samples from deviating too far from the real scene.

[0044] Figure 4In the diagram, 260a represents the portrait region after injecting local printed dots, 260b represents the text region after overlaying scanned moiré patterns, and 260c represents the local region after the anti-counterfeiting mark is weakened. Although several different types of artifacts are shown in the figure, a specific sample may contain only one type, or it may contain two or more types simultaneously. If multiple artifacts are overlaid in the same image, the training server 110 can record the region, order, and parameters of each artifact so that it can accurately indicate the source and manifestation of the forgery features when generating corresponding descriptive information later.

[0045] After performing local artifact injection on the base document image 250, a sample forged document image 260 can be obtained. This image retains the overall structure of the standard document template while also containing local anomalies related to the forgery process. The resulting sample is no longer a simple category label sample, but a structured training sample with known anomaly locations, types, and details. The training server 110 can write these samples into the training data storage unit 140, providing a foundation for subsequently constructing multi-level question text and descriptive information.

[0046] After the sample forged document image 260 is formed, the training process does not immediately enter the model training, but continues to organize text-based supervision information around the image. Figure 5 This demonstrates a multi-layered approach to organizing question text and descriptive information. Figure 5 In the diagram, 510 represents the sample forged document image, 520 represents the first-level question text, 530 represents the second-level question text, 540 represents the third-level question text, 550 represents the first-level descriptive information, 560 represents the second-level descriptive information, and 570 represents the third-level descriptive information. 520, 530, and 540 together constitute the multi-level question text corresponding to the sample forged document image 510; 550, 560, and 570 together constitute the descriptive information corresponding to this image.

[0047] exist Figure 2 In the illustrated process, the training server 110 can execute step S130: constructing multi-layered question text and descriptive information. The key to this step is not simply writing a few questions, but rather organizing the anomalies in the image into a phased representational structure that can be learned by the model. The first-layer question text 520 can be used to inquire about the overall authenticity of the document, such as organizing questions around whether the entire document image is suspected of being forged; the second-layer question text 530 can be used to inquire about suspicious areas and forgery feature types; the third-layer question text 540 can be used to inquire about the details of specific anomalies. Correspondingly, the first-layer descriptive information 550 can provide an overall judgment, the second-layer descriptive information 560 can provide the region and type, and the third-layer descriptive information 570 can provide detailed descriptions of local artifacts.

[0048] Multi-layered question texts do not require manual, sentence-by-sentence writing each time. The training server 110 can automatically generate question texts using preset templates. The variable parts in the templates can be filled with artifact annotation information corresponding to the sample forged document images 260. For example, if the known abnormal region is located in the portrait region and the abnormality type is printing artifact, the second-layer question text can be organized around this region; if the known abnormality is located at the edge of the expiration date text, the third-layer descriptive information can specifically describe the abnormality of stroke boundaries and texture continuity. Thus, a stable mapping relationship is formed between the multi-layered question text and the descriptive information, and this mapping can be directly used as a supervision signal during subsequent model training.

[0049] exist Figure 5 In the structure shown, the training server 110 typically does not stop at a static question list when organizing multi-level question texts. Instead, it organizes the sample forged document image 510, the first-level question text 520, the second-level question text 530, the third-level question text 540, and the corresponding descriptive information 550, 560, and 570 into a complete training record. This training record can be represented as a combination of image objects and multi-round text objects, or as a combination of image objects and hierarchical text fields. Different models differ in input format, but the core semantics used for training remain consistent: the model needs to form an overall judgment after observing the document image, then provide the region and type, and finally provide detailed explanations.

[0050] In one embodiment, the training server 110 can organize training data in a sequential input manner. Sequential input means that on the same sample forged document image 510, a set of training samples corresponding to the first-layer question text 520 and the first-layer description information 550 is constructed first; then a set of training samples corresponding to the second-layer question text 530 and the second-layer description information 560 is constructed; and finally, a set of training samples corresponding to the third-layer question text 540 and the third-layer description information 570 is constructed. The advantage of this organization is that the supervision objective of each set of training samples is relatively clear, the task pressure on the model at each layer is relatively small, and it is easier to converge layer by layer. For multimodal models with long inference paths and large parameter scales, this method helps to reduce the mutual interference between semantics at different levels.

[0051] In another embodiment, the training server 110 can also organize training data in a conversational input manner. Figure 6 This demonstrates the input-output relationship for training a multimodal model. Figure 6In this diagram, 610 represents the input image object, 620 represents the dialogue template, 630 represents the multimodal model, and 640 represents the model output. In the dialogue input mode, the training server 110 can write the first-layer question text 520, the second-layer question text 530, and the third-layer question text 540 into the dialogue template 620 according to the questioner's role message, and then input the sample forged document image 510 and the dialogue template 620 together into the multimodal model 630. At this time, the model output 640 can be multiple responses corresponding to the questions at each layer, or it can be a uniformly organized structured text. Since multimodal models typically already have the ability to process image and dialogue input during the pre-training stage, this organization method is closer to the model's native inference form and is more convenient to reuse directly in the subsequent identification stage.

[0052] Regardless of whether sequential or dialogic input is used, the training server 110 needs to convert known artifact information in the image into descriptive information. This descriptive information is not simply writing down the artifact names; it needs to correspond to the question level at a specific granularity. For the first layer of descriptive information 550, the training server 110 can provide an overall authenticity judgment based on whether at least one valid artifact exists in the sample forged document image 510. For example, when there are obvious local anomalies in the image, the first layer of descriptive information 550 can indicate a suspected forgery; if normal base image samples are introduced during training, the first layer of descriptive information 550 can also indicate a high overall authenticity. For the second layer of descriptive information 560, the training server 110 can write the suspicious region and anomaly type. For the third layer of descriptive information 570, the training server 110 can further write the manifestation of artifacts in local regions, such as edge discontinuities, local discrete dots, periodic interference textures, or character stroke distortion.

[0053] To reduce the cost of relying entirely on manual generation of descriptive information, the training server 110 can automatically generate basic descriptions based on artifact parameters and region annotations, which are then corrected by a rule engine or manual verification. This is because minute artifacts in an image originate from a controllable injection process; their location, type, and some morphological parameters are recorded during sample generation. Therefore, initial descriptive text can be generated based on these known quantities. For example, if an artifact injection record includes information such as region type, portrait edge location, local texture frequency enhancement, and discrete color point overlay, the system can first generate a draft description corresponding to the printing artifact. This draft can then be further adjusted based on the document type to make the description more closely resemble the actual business context.

[0054] To write this process of generating descriptive information from known artifact parameters into an executable form, constraints can be imposed on the parameter mappings in the descriptive information. Let a certain artifact record be P={r,t,v}, where Indicates the area where the artifact is located. Indicates the type of artifact. Let the artifact parameter vector be represented, then the information selection function can be written as: = g{r,t,v}(2) Where D represents the generated description information; g(·) represents the description generation function; It can represent an image area, a text area, an anti-counterfeiting area, or other local areas; It can represent printing artifacts, scanning artifacts, texture anomalies, or character anomalies; This can represent the set of parameters associated with the artifact. Through the function... The function of the training server 110 is to generate text descriptions corresponding to the sample forged document image 510 based on the region, type, and parameter combination. The function here does not need to have a uniform form in all embodiments. Template matching, rule mapping, text retrieval, or small model generation can all be used, as long as the output can establish a stable correspondence with the artifact record.

[0055] During the training phase, descriptive information is used not only to construct supervised text but also to measure the deviation between the model output and the target output. Figure 6 After receiving the input image object 610 and the dialogue template 620, the multimodal model 630 can output a sample authenticity identification result 640. This result 640 can be a single-turn response or a collection of multiple-turn responses. After obtaining this result, the training server 110 does not use whether the true or false label is matched as the sole optimization objective. Instead, it also checks whether the output reflects regions, types, and details consistent with the descriptive information. This is because one embodiment of this application addresses not only the binary classification accuracy, but also whether the model has truly learned the correspondence between local forgery traces and the step-by-step analysis process.

[0056] exist Figure 2 In the corresponding process, the training server 110 continues to execute step S140: adjusting model parameters based on output bias. This bias includes, but is not limited to, overall true / false judgment bias, suspicious region identification bias, forged feature type bias, and detailed description bias. If the overall judgment output by the multimodal model 630 is correct, but the correct region is not identified, the training server 110 can still consider that there is an unconverged part; if the region judgment is correct, but the anomaly type and detailed description are incorrect, training can still continue. The advantage of defining bias in this way is that the training objective is expanded from a single true / false classification to a multi-level task, giving the model stronger interpretability constraints at the output end.

[0057] In a feasible implementation, the training server 110 can divide the bias into two parts: image discrimination bias and text description bias. Let the target of sample authenticity discrimination be... The true / false output corresponding to the model is The target sequence of the description information is The output description sequence corresponding to the model is The overall training loss can then be expressed as: (3) in, Indicates the overall training loss; cls Indicates the loss from determining authenticity; txt This represents the loss in generating descriptive information; and These represent the weight parameters of the two loss terms. The true / false discrimination loss can take the form of a classification loss, and the descriptive information generation loss can take the form of a sequence generation loss. By jointly optimizing the two loss components, the training server 110 can simultaneously adjust the model's mapping ability to visual anomalies and textual representations during a unified training process.

[0058] If a multi-level supervision approach is adopted, the training server 110 can be further... It is broken down into multiple parts corresponding to the three layers of question text. Let the output of the first layer be... The second layer output is The third layer output is The corresponding targets are respectively , , Then it can be written as: (4) in, , and These represent the descriptive deviation items for the first, second, and third layers, respectively. , and This represents the weight parameter for the corresponding level. If a particular implementation focuses more on overall judgment, the weight can be appropriately increased. If a particular implementation method places greater emphasis on detailed attribution, then the appropriate increase can be made. These weights are not a prerequisite for the validity of this application, but rather an adjustable implementation detail during the training phase.

[0059] The model parameters can be updated by the training server 110 performing a regular backpropagation process. The specific optimizer, batch processing method, or training framework used is not a core limitation of any embodiment of this application. In actual deployment, the training server 110 can select a suitable update strategy based on training resources, model size, and sample quantity. If the training data is limited in the early stages, a smaller batch size and a more gradual learning rate can be used to allow the model to first establish a basic correspondence between images and text; if the sample coverage is more sufficient in later stages, the training intensity can be appropriately increased to reduce the identification bias of local artifacts.

[0060] Figure 6 The multimodal model shown in Figure 630 can be a model capable of simultaneously processing image and text inputs and generating text output. Internally, it generally includes a visual encoding part, a visual-to-language feature mapping part, and a language inference part. The visual encoding part extracts visual features from the input image object 610; the visual-to-language feature mapping part converts the visual features into a representation that can be jointly processed with the text input; and the language inference part combines multi-level question text to complete output generation. In one embodiment of this application, the specific network names of these parts are not strictly limited, but in engineering, they typically include two key stages: image feature extraction and image-text feature fusion. If the former stage lacks sufficient response to local high-frequency textures, even with improved subsequent language output, it will be difficult to compensate. Therefore, the small artifacts introduced in the training samples can directly drive visual feature learning. If the latter stage cannot correctly associate image anomalies with text levels, the output is still likely to remain at a generalized judgment level. Introducing sample forged document images and multi-level question text during the training phase is precisely to apply constraints simultaneously to these two stages.

[0061] When training is nearing convergence, the training server 110 can decide whether to end the current round based on preset training stopping conditions. Training stopping conditions can be set in various ways. If a validation set is used, the training server 110 can send a set of validation images that did not participate in the current round of training into the multimodal model 630 to observe the overall discrimination accuracy and the consistency of step-by-step descriptions. If the overall authenticity judgment has reached a preset level, but there are still significant deviations in the step-by-step descriptions, the training server 110 can continue to optimize the description-side loss; if both the overall and description sides tend to stabilize, the trained document authenticity identification model can be output and written to the model storage unit 150.

[0062] The document authenticity verification model obtained after training is not limited to a specific document type. In one embodiment, the training server 110 can train a dedicated model for a single document type, such as training only for resident ID cards; in another embodiment, the training set can be composed of base document images of multiple preset document types and sample counterfeit document images, enabling the model to identify multiple document types. When using the former approach, the model's local anomaly expressions in a single document scenario are usually more concentrated; when using the latter approach, the model has a wider range of applicability, but the training sample coverage also needs to be more comprehensive.

[0063] when Figure 2 After the training process shown is completed, the document authenticity identification model can be saved in the model storage unit 150. Figure 3 The main process of the identification method is shown. Figure 3 In this context, 310 represents the process of acquiring the image of the document to be authenticated, 320 represents the process of acquiring the multi-level question text, 330 represents the model reasoning process, and 340 represents the process of outputting the authenticity authentication result. Figure 7 This further illustrates the input organization method during the identification phase. Figure 7 In this context, 710 represents the image of the document to be authenticated, 720 represents the preset multi-level question text, 730 represents the joint input object, 740 represents the document authenticity authentication model, 750 represents the authenticity authentication result, and 760 represents multiple step-by-step authentication information.

[0064] In one embodiment of this application, after receiving a verification request, the authentication server 120 can execute step S210: acquiring the image of the document to be verified. The image 710 of the document to be verified can be directly uploaded from the terminal device 130, or it can come from a local shooting module, scanning module, or third-party acquisition module. Since the actual acquisition conditions on the business side may fluctuate greatly, the image 710 of the document to be verified usually needs to be processed before being input into the model. The purpose of processing is not to change the content of the document image, but to unify the image to an input format that the model can stably process, such as border cropping, image compression, format conversion, or resolution normalization. If these processes are not performed, local artifacts may be buried by irrelevant backgrounds, or the inference side noise may be increased due to inconsistent input formats.

[0065] In one embodiment, the authentication server 120 may perform step S220: acquiring a multi-level question text 720 for periodically verifying the authenticity of documents. "Acquiring" here does not require manual, sentence-by-sentence input for each authentication; a more common approach is for the system to read the text from a pre-configured prompt text library. In other words, the multi-level question text 720 is typically a pre-defined, standardized set of questions. This is because if external operators freely input the question text each time, fluctuations in expression style, question order, and terminology across different requests can easily occur, leading to instability in the model's output format. By pre-setting a unified text, the authentication server 120 can maintain a consistent inference input structure across different authentication requests, thereby improving the comparability and verifiability of the model's output results.

[0066] In another embodiment, the multi-level question text 720 can also be dynamically selected by the authentication server 120 based on the document type. For example, when the document image 710 to be authenticated is determined to be a resident ID card after preliminary type recognition, the authentication server 120 can call a set of multi-level question texts matching the resident ID card; if it is determined to be a passport or driver's license, then the question text template for the corresponding document type is called. The advantage of doing so is that different documents have different layout structures, field distributions, and typical forgery methods, and dynamically selecting the question text helps to improve the fit between the subsequent output results and the document scenario.

[0067] Figure 7 The multi-level question text 720 shown can adopt the hierarchical structure used in the training phase. In one embodiment, the multi-level question text 720 includes at least a first-level text, a second-level text, and a third-level text. The first-level text is used to inquire about the overall authenticity of the document; the second-level text is used to inquire about the type of forgery features; and the third-level text is used to inquire about detailed forgery information. The reason for adopting this three-level structure is that the identification phase requires the model to reason along the cognitive path learned in the training phase. If the model always learns in the order of overall judgment, region and type analysis, and detailed attribution during training, but suddenly adopts a completely different input structure during identification, the model's efficiency in utilizing the output level will decrease. Therefore, the correspondence between the training phase and the identification phase in terms of text hierarchical structure is one of the important conditions for ensuring the stability of reasoning behavior.

[0068] Figure 10It can be divided into two parts, left and right. The left side shows the sequential input method, in which the same image of the document to be identified is input into the document authenticity identification model three times, along with the first layer of text, the second layer of text, and the third layer of text, respectively, and the overall judgment, suspicious area and type, and detailed description are obtained respectively. The right side shows the dialogic input method, in which the image of the document to be identified and the three layers of question text are organized into a joint input object according to the dialog template, and the document authenticity identification model is input at once, and the authenticity identification result containing multiple steps of identification information is output. Figure 10 This is used to illustrate that during the identification stage, the multi-level question text can be input into the document authenticity identification model 740 in a sequential manner or in a dialogic manner, thereby realizing the authenticity identification of the document image 710 to be identified and the step-by-step information output.

[0069] In one embodiment, the authentication server 120 can use multi-level query text 720 in a sequential invocation manner. Sequential invocation means that the authentication server 120 first inputs the document image 710 to be authenticated along with the first-level text into the document authenticity authentication model 740 to obtain an overall authenticity judgment. If the returned result indicates suspected forgery, the authentication server 120 continues to input the same document image 710 along with the second-level text into the document authenticity authentication model 740 to obtain suspicious areas and forgery feature types. Subsequently, the authentication server 120 inputs the document image 710 along with the third-level text into the document authenticity authentication model 740 to obtain more detailed forgery information. Using this method, there is an explicit order of reasoning between each layer, facilitating judgment of the results at intermediate layers. For example, if the first layer clearly outputs a high overall authenticity, the system can directly end the process without proceeding to more fine-grained analysis.

[0070] In another embodiment, the authentication server 120 can also use multi-level question text 720 in a conversational invocation manner. In this case, the authentication server 120 will write the first-level text, the second-level text, and the third-level text into the joint input object 730 according to a unified dialogue template, and send the document image 710 to be authenticated and the joint input object 730 into the document authenticity authentication model 740 at once. In this way, the model can output the authenticity authentication results 750 and multiple step-by-step authentication information 760 corresponding to each level of question in a single round of reasoning. This implementation method is closer to the common interface form of multimodal dialogue models, which helps to reduce the communication and scheduling overhead caused by repeated calls. When resources permit and the model has the ability to expand semantics in multiple rounds, the conversational invocation method can obtain a relatively complete hierarchical output at the same time.

[0071] It's important to note that in the conversational invocation method, while the questions in the joint input object 730 can be organized into a format similar to dialogue roles, these questions are typically automatically filled in by the system, rather than requiring manual input from individuals in real-world business scenarios. In other words, the questioner role identifier in the joint input object 730 is essentially part of the input format; its purpose is to adapt to the model's dialogue input interface, not to restrict the participation of a specific type of user. This approach preserves the true form of the technical implementation while avoiding the misinterpretation of the model's reasoning process as a multi-turn manual interaction.

[0072] In step S230, the authentication server 120 inputs the document image 710 to be authenticated and the multi-level question text 720 into the document authenticity authentication model 740, and obtains the authenticity authentication result 750 output by the model after inference. The "authenticity authentication result 750" here may include an overall authenticity conclusion, and may further include at least a portion of multiple step-by-step authentication information 760. Regarding the question of whether there is suspicion of forgery, the model can provide an authenticity judgment result; for suspicious areas and feature types, the model can return area location information and type classification information; for detailed attribution, the model can provide explanatory text corresponding to local artifacts. As long as these contents revolve around the same document image 710 to be authenticated, they can be considered as the authenticity authentication result 750 in one embodiment of this application.

[0073] Multiple step-by-step authentication information 760 typically correspond to multi-level question text 720 in structure. To facilitate system processing and manual review, the authentication server 120 can organize the multiple step-by-step authentication information 760 into a structured result object containing multiple fields. For example, the result object can include fields such as "Overall Authenticity Judgment," "Suspicious Area," "Forgery Feature Type," and "Forgery Details." When the model output itself already has a clear structural hierarchy, the authentication server 120 can directly read the corresponding content and write it into these fields; when the model output is continuous text, the authentication server 120 can also parse the output text based on preset rules and map it to the corresponding fields. In this way, the model output results are not only readable by humans but also easy to call by subsequent business modules.

[0074] In one embodiment, if the authentication result 750 indicates that the document to be authenticated is suspected of being counterfeit, the authentication server 120 can execute step S240: output multiple step-by-step authentication information 760. The output method can be varied. For example, the authentication server 120 can send a structured result containing overall judgment, suspicious areas, anomaly type, and forgery details back to the terminal device 130 for review by the auditor; or the result can be directly written into a risk control system, audit system, or case handling system. In some scenarios, the authentication server 120 can also archive the multiple step-by-step authentication information 760 together with the image 710 of the document to be authenticated for subsequent manual review, sampling inspection, or dispute resolution.

[0075] To illustrate the use of the multi-level question text 720 in the authentication phase, consider the following specific example. Assume that terminal device 130 uploads an image of a resident ID card as the document image 710 to be authenticated. The portrait area in the image contains local discrete dots caused by ordinary color printing. Authentication server 120 first reads the preset first-level text, such as "Please determine if the document as a whole is suspected of being counterfeit"; simultaneously, it reads the second-level text, such as "If suspected of being counterfeit, please point out the suspicious area and describe the type of counterfeit feature"; and then reads the third-level text, such as "Please describe the specific details of the counterfeit feature." If a sequential calling method is used, the document authenticity authentication model 740 can output "The document as a whole is suspected of being counterfeit" in the first round of reasoning; in the second round of reasoning, it can output "The suspicious area is located in the portrait area, and the counterfeit feature type is printing artifact"; and in the third round of reasoning, it can output "Discrete printing dots exist at the edge of the portrait, and the local color transition is discontinuous, which does not match the continuous color tone characteristics of a genuine document." If a conversational approach is used, the document authentication model 740 can also output answers corresponding to these three levels of questions in the same round of reasoning. Regardless of the method used, the multiple step-by-step authentication information 760 remains consistent with the hierarchical structure learned during the training phase, thus improving the stability of the authentication results.

[0076] To further improve the adaptability to image quality fluctuations during the identification stage, the identification server 120 can also perform image preprocessing before inputting the data into the model. Preprocessing operations may include at least one of cropping borders, compressing images, and modifying image format. Cropping borders aims to remove background areas unrelated to the document's content, preventing the model from focusing on irrelevant areas; compressing images aims to ensure the input data meets the size limitations required by the model's inference interface; modifying image format aims to unify the encoding of data uploaded from different terminals. It should be noted that these preprocessing operations, as long as they do not change the authenticity of the document's content, can be considered as standardizing the input format and do not affect the core identification logic of one embodiment of this application.

[0077] In one embodiment, the document authenticity verification model 740 can not only output authenticity judgments but also auxiliary results related to credibility. For example, the model can add an authenticity score to the first-level judgment result, a confidence level to the second-level anomaly type, and a set of evidence references associated with local areas to the third-level detailed description. Although such auxiliary results are not entirely necessary features, adding this information to a practical system can help the business side set review thresholds and manual review strategies more flexibly.

[0078] Figure 3 and Figure 7 The illustrated authentication process does not require the authentication server 120 to independently complete all processing steps. In another embodiment, the terminal device 130 may first perform some preprocessing operations before sending the processed document image 710 to the authentication server 120. Alternatively, the terminal device 130 may deploy a lightweight version of the document authenticity authentication model 740 locally, perform a preliminary overall authenticity judgment locally, and then have the authentication server 120 perform a more detailed step-by-step analysis. This kind of end-side and cloud-side division of labor does not change the basic idea of ​​one embodiment of this application. As long as the final authentication result 750 and multiple step-by-step authentication information 760 are formed based on the document image 710, the multi-level question text 720, and the document authenticity authentication model 740, it falls within the basic scope of this application's solution.

[0079] In some preferred embodiments, the training and identification phases can be further enhanced with optional modules to improve overall performance. For example, the method of updating model parameters can be restricted during the training phase to reduce training resource consumption; content logic consistency analysis can also be introduced during the training phase to improve the model's ability to identify problems such as field forgery. Figure 8 This illustrates the training extension structure that includes optional functional modules. Figure 8 In this diagram, 810 represents the baseline multimodal model, 820 represents the incremental parameter module, 830 represents the logic verification module, 840 represents the hard sample feedback module, and 850 represents the training control module.

[0080] In one embodiment, the training server 110 can update the multimodal model parameters using a parameter-efficient fine-tuning method. Here, parameter-efficient fine-tuning refers to training only a newly added or selected subset of parameters without updating all model parameters. Figure 8The baseline multimodal model 810 provides general image and text understanding capabilities, while the incremental parameter module 820 is used to supplement local anomaly recognition capabilities in document forgery scenarios. This approach is adopted because general multimodal models often have a large parameter scale; if all parameters are used in training, training resources are consumed significantly, and the original general capabilities may be negatively impacted. By training only the incremental parameter module 820, the training server 110 can achieve targeted adaptation for document authenticity verification tasks at a lower cost.

[0081] In one embodiment, the incremental parameter module 820 can be associated with the visual feature mapping part, the image-text feature interaction part, or other parts related to local detail representation in the baseline multimodal model 810. Since the minute artifacts that this application focuses on in one embodiment often manifest as local textures, edge details, and local region anomalies, placing the incremental parameter module 820 closer to the visual representation and image-text correspondence helps improve the model's response to fine-grained anomalies. Of course, the specific placement of the incremental parameter module 820 does not constitute a necessary prerequisite for the success of this application's solution, as long as it can receive the bias signal from step S140 and complete parameter updates during training.

[0082] In another embodiment, the training server 110 may also incorporate a logic verification module 830. The logic verification module 830 does not replace image-side artifact recognition, but rather complements it. Specifically, while some counterfeit documents may simulate genuine documents as closely as possible in their image production process, logical inconsistencies may still exist between their fields. For example, the issuance date may be later than the start of the validity period, the birth date may not match the corresponding segment in the document number, or the facial features may not match the gender description in the text field. The training server 110 can generate additional logical consistency annotations for sample counterfeit document images and use the logic verification module 830 to perform consistency analysis on key areas and text content of the samples.

[0083] To better illustrate the function of the logic verification module 830, the set of fields extracted from the sample forged document image can be defined as follows: The visual semantic set extracted from the key regions is The preset content logic rule set is as follows The logical verification result can be expressed as: (5) in, Indicates the result of the logical verification; This represents a logical consistency judgment function. This function can output a binary result or a result with a consistency score. Through... By optimizing the differences between the logical consistency annotation and the training server 110, the logical verification module 830 can gradually acquire the ability to detect anomalies in field relationships. In this way, when the model encounters document images that have both visual and technical anomalies as well as possible logical anomalies in the subsequent identification stage, it can obtain a more comprehensive basis for judgment.

[0084] In addition to efficient parameter fine-tuning and logic verification, the training server 110 can also incorporate a difficult sample feedback module 840. The difficult sample feedback module 840 analyzes samples prone to misclassification during training and optimizes the sample generation process in reverse. Let the set of misclassified samples in the current training round be... The artifact parameter vector corresponding to each sample in this set is Then the training server 110 can collect these statistics. The distribution characteristics under different regions and different anomaly types are analyzed, and the sampling strategy for minor artifacts in subsequent samples is adjusted accordingly. The aim is to ensure that the next round of samples more concentratedly covers the parameter ranges and combination patterns where the current model has weaker recognition capabilities.

[0085] If we express this process in a generalized way, the update of the artifact generation strategy can be written as: (6) in, Indicates the first The artifact generation strategy used during round training; This represents the set of difficult samples in the current round; This represents the generation policy update function; This indicates the artifact generation strategy to be used in the next round of training. By analyzing the strategy... Through iterative updates, training server 110 can make subsequently generated samples closer to the model's recognition blind spots, thereby gradually improving the model's robustness.

[0086] In one embodiment, the training control module 850 is used for coordination. Figure 2 , Figure 5 and Figure 8 The execution order of each functional unit is as follows: The training control module 850 can read the blank document template image and anti-counterfeiting process information, call the sample generation logic to generate the base document image, call the artifact generation logic to obtain the sample forged document image, call the text generation logic to construct multi-level question text and descriptive information, and then send these data to the baseline multimodal model 810 and the incremental parameter module 820 for training. When the logic verification module 830 and the hard sample feedback module 840 are enabled, the training control module 850 can also be responsible for coordinating the use of logical consistency annotation and the feedback of hard sample statistical results.

[0087] Figure 9 A schematic diagram of the electronic device is shown. Figure 9 In this document, 910 represents a processor, 920 represents a memory, 930 represents an optional communication interface, and 940 represents a bus or other connection structure. In one embodiment of this application, the electronic device may be a training server 110, an authentication server 120, a terminal device 130, or other devices capable of executing methods related to the document authenticity verification model. The processor 910 is used to call computer instructions stored in the memory 920 to execute the steps in the aforementioned training or authentication methods. The memory 920 may store data such as blank document template images, anti-counterfeiting process information, base document images, sample counterfeit document images, multi-level question text, descriptive information, document authenticity verification models, and output results. The optional communication interface 930 is used to transmit images, text, and model results between the training server 110, the authentication server 120, and the terminal device 130.

[0088] In one embodiment, when the electronic device is used as a training device, the instructions executed by the processor 910 enable the electronic device to perform the following operations: generate a base document image based on standard document information of a preset type of document; inject tiny artifacts into the base document image to obtain a sample forged document image; construct multi-level question text and descriptive information corresponding to the sample forged document image; input the image and text into the multimodal model to be trained; update the model parameters according to the deviation between the sample authenticity identification result and the descriptive information; and output the trained document authenticity identification model after the training stopping condition is met.

[0089] In another embodiment, when the electronic device operates as an authentication device, the instructions executed by the processor 910 enable the electronic device to perform the following operations: respond to a verification request to acquire an image of the document to be verified; acquire multi-level question text for periodically verifying the document's authenticity; input the image of the document to be verified and the multi-level question text into the document verification model; acquire the verification result; and if the verification result indicates that the document to be verified is a counterfeit document, output multiple step-by-step verification information. Thus, the electronic device claims can form a clear correspondence with the method claims.

[0090] It should be noted that, Figure 9 The illustrated electronic device structure is merely for illustrating the program execution carrier of one embodiment of this application and does not limit the electronic device to having completely identical hardware components. In some scenarios, training-related data and model parameters can be deployed in a remote server cluster, with the terminal device only responsible for acquiring images and displaying results; in other scenarios, lightweight models can also be directly deployed on a local terminal, with the terminal device independently completing part or all of the identification process. As long as the device can execute the aforementioned method steps, it can be considered to fall within the device implementation scope of one embodiment of this application.

[0091] In summary, one embodiment of this application constructs sample forged document images, multi-level question texts, and descriptive information corresponding to the forgery process during the training phase. This enables the multimodal model to learn local fine-grained forgery traces in the document images and the corresponding step-by-step analysis output methods. During the authentication phase, by inputting the document image to be authenticated and the multi-level question text into the trained document authenticity authentication model, the system can not only output the document authenticity judgment but also output multiple step-by-step authentication information corresponding to the overall judgment, region analysis, and detail attribution. Therefore, one embodiment of this application can improve the ability to identify local abnormal patterns and enhance the verifiability of authentication results in automatic document authenticity authentication scenarios.

[0092] In some embodiments, to make the base document image closer to the real appearance of the target document type, the training server 110 can distinguish between general fields and personalized fields when generating standard document content. General fields typically refer to parts of the same type of document with fixed layout, fixed content templates, or minimal variation, such as document name, fixed background texture, fixed logo, issuing authority format area, and standard border structure. Personalized fields typically refer to parts related to the specific document holder and that vary between different document instances, such as name, document number, date of birth, issuance date, validity period, portrait image, or address information. When constructing the base document image, the training server 110 can first write the general fields into a blank document template image, and then generate personalized fields according to preset rules and write them into the corresponding positions. In this way, the generated base document image retains the overall layout consistency of the document type while also having sufficient individual differences to support model training.

[0093] There are multiple ways to generate personalized fields. For example, the training server 110 can extract corresponding field values ​​from preset character, number, and date libraries, or it can automatically generate them according to rules. For instance, if the document type is a resident ID card, the document number can be automatically generated according to administrative division codes, birth date codes, and sequence code rules; the birth date field can be randomly generated by a date rule engine from a legal date within a preset range; the validity period field can be calculated backward based on the birth date or issuance date according to rules. The portrait image can come from a de-identified image library, a legally synthesized image, or other image sets that do not contain sensitive real-identity mapping relationships. In this way, the training server 110 can generate a large number of base document images for training without exposing real sensitive document content.

[0094] In some embodiments, after generating the base document image, the training server 110 can also perform basic imaging processing on the image to simulate non-forgery factors in a real acquisition process. For example, the training server 110 can apply slight changes in resolution, contrast, local illumination, edge cropping offset, or slight compression coding to make the base document image during training closer to the input format in real business scenarios. The purpose of this is to allow the model to learn forgery traces while gradually adapting to imaging disturbances brought about by normal acquisition, avoiding misjudging ordinary shooting deviations as forgery traces. It should be noted that this type of basic imaging processing is different from micro-artifact injection. The former mainly simulates general disturbances in a normal acquisition process, while the latter explicitly corresponds to abnormal features introduced by forgery techniques.

[0095] In some embodiments, the training server 110 can establish structured sample records for sample forged document images. These records, in addition to the image itself, may include sample identifiers, document type identifiers, artifact type identifiers, artifact region identifiers, artifact parameter records, logical consistency status, corresponding query text identifiers, and corresponding descriptive information identifiers. By establishing structured sample records, the training server 110 can directly access sample features by index during subsequent training, validation, and hard sample analysis stages, without needing to repeatedly parse the image generation process. While this data organization method is not a necessary prerequisite for the success of this application, it helps improve sample management efficiency in engineering and provides a foundation for subsequent closed-loop optimization of hard samples.

[0096] In a preferred embodiment, the artifact region identifiers in the sample records can be represented using normalized region coordinates, predefined region numbers, or by binding to document fields. For example, the training server 110 can mark the portrait region as a first-class key region, the main text field region as a second-class key region, and the anti-counterfeiting mark region as a third-class key region. Thus, when generating the second-layer description information, the system can directly generate descriptions such as "abnormality exists in the portrait region," "abnormality exists in the text region," or "abnormality exists in the anti-counterfeiting region" based on the region identifiers, without needing to re-identify the region's location from the image. If normalized region coordinates are used, the training server 110 can also further use these coordinates for subsequent visualization, such as overlaying selected regions in a manual review interface.

[0097] In some embodiments, multi-level question text can employ both fixed and semi-dynamic templates. A semi-dynamic template refers to a text structure that remains unchanged, but whose internal terms or pronouns automatically change based on the sample context. For example, the first-level text remains an overall authenticity question, while the second-level text can dynamically replace "document" with "ID card," "passport," or "driver's license" based on the document type; the third-level text can then point objects such as "anomaly," "region," and "forgery feature" to specific preceding results based on the second-level annotation results. The advantage of using semi-dynamic templates is that they maintain the consistency of training semantics while reducing overfitting of the model to a single fixed question, making it easier to generalize to standardized question variations in the actual identification stage.

[0098] In one embodiment, the training server 110 can organize multi-level question texts and descriptive information into paired supervisory units. Assume a sample image of a forged ID card corresponds to the first-level question text. Second-level question text Third-level question text The corresponding target description information are as follows: , , Then the training unit can be represented as ,in This indicates that the sample is a forged document image. The training server 110 can train these supervised units individually or in combination into a multi-round question-and-answer format. The significance of this approach is that it enables the model to establish semantic mappings of different granularities around the same image at different levels.

[0099] In some embodiments, to reduce the model's exposure to overly complex detail attribution tasks early in training, the training server 110 can employ a phased sample feeding approach. Specifically, in the initial training phase, training units primarily containing obvious artifacts and relatively simple overall judgment tasks can be used; in the middle training phase, training units containing region localization and type judgment tasks are gradually added; and in the later training phase, training units with higher detail attribution requirements, finer regions, and longer descriptions are further added. This approach is essentially a course-based training arrangement.

[0100] In some embodiments, the granularity of descriptive information can be organized into two categories: evidentiary and conclusive. Conclusive descriptions typically emphasize "what," such as "this area belongs to a printing artifact"; evidentiary descriptions further emphasize "why," such as "the image edge has discrete color dots and discontinuous local color transitions, therefore it is judged to be a printing artifact." During the training phase, the training server 110 can either use only conclusive descriptions as supervision text or combine conclusive and evidentiary descriptions. The advantage of combining them is that the model not only learns the anomaly category but also learns the local evidence expression methods that support the judgment of that category. In this way, during the identification phase, the multiple step-by-step identification information output by the model is more likely to include evidence that can be manually verified.

[0101] To illustrate the correspondence between descriptive information and image regions, a region selection mapping can be introduced. Let the set of candidate key regions in the sample forged document image be denoted as . The system selects abnormal areas based on artifact records. Then the selection of abnormal regions can be formalized as follows: (7) in, This indicates a sample image of a forged document; Indicates artifact recording; This represents a function for selecting anomalous regions based on images and records. Since the artifact records are known during the training phase, This can be directly determined by the sample generation process. When generating the second and third layer description information, the training server 110 can identify the abnormal region. This mapping is used to define regions such as "avatar region," "text region," "anti-counterfeiting region," or other more granular region names. Through this mapping, the text supervision received by the model during the training phase will form a stable correspondence with local regions in the image.

[0102] In one embodiment, if multiple anomalous regions exist in the sample forged document image, the training server 110 can employ at least one of the following strategies. The first strategy is to prioritize selecting the most significant anomalous region as the primary descriptive object for the second and third layers, thereby reducing the complexity of the model's output branches. The second strategy is to generate descriptive information sequentially for multiple anomalous regions, enabling the model to learn how to organize its output under multiple anomalous conditions. The third strategy is to group a certain type of region among multiple anomalous regions into the same descriptive object; for example, similar font anomalies in multiple text fields can be combined to describe character deformation in the text region. The choice of strategy in practical engineering depends on the scale of the training data, the model capacity, and the business requirements for output complexity.

[0103] In some embodiments, the authenticity verification result and multiple step-by-step verification information can be organized using an inclusion relationship. That is, the authenticity verification result serves as the overall result object, which includes both the authenticity judgment and the corresponding multiple step-by-step verification information. Using this approach, the verification server 120 can internally maintain a unified result data structure that stores both the overall authenticity conclusion and the field values ​​corresponding to the query text at each level. The advantage of this is that the system's external interface only needs to return a unified result object, and the business side can read different fields as needed. Another feasible approach is to maintain the authenticity verification result and multiple step-by-step verification information separately, but then combine and output them at the business interface layer. Neither approach changes the essence of the solution in this application.

[0104] In some embodiments, before outputting multiple steps of authentication information, the authentication server 120 can also execute post-processing rules based on the authenticity authentication results. For example, if the overall authenticity judgment indicates that the document is basically credible, the system can choose to return only the overall authenticity judgment information without displaying more detailed local descriptions; if the overall authenticity judgment indicates that the document is suspected of being counterfeit, the system returns all or part of the steps of authentication information. The reason for this arrangement is that different business processes have different requirements for the redundancy of output information. In large-scale real-time audit scenarios, it is generally desirable to output as concisely as possible for normal documents, while outputting more explanatory information for suspicious documents.

[0105] In some embodiments, the authentication server 120 can also template multiple steps of authentication information. For example, the output can be organized into four fixed fields: "Overall Judgment," "Suspicious Location," "Anomaly Type," and "Basis Explanation." This not only benefits the front-end display but also facilitates subsequent record keeping and auditing. If the front-end uses a card-style review interface, the overall judgment can be displayed separately at the top, the suspicious area can be overlaid on the document image, and the anomaly type and basis explanation can be placed in the text area below.

[0106] In some embodiments, there can be multiple ways to deliver the model between the training server 110 and the authentication server 120. If training and inference are deployed in the same server cluster, the trained document authentication model can be directly stored in the same model storage unit 150 for the authentication server 120 to call. If training and inference are deployed on different physical nodes, the training server 110 can package the model parameters, model configuration file, and inference interface description and send them to the authentication server 120. In this case, the authentication server 120 completes model loading locally and receives the document image to be authenticated from the terminal device 130. As long as the model called in the authentication stage corresponds logically to the model obtained in the training stage, the requirements of the embodiments of this application can be met.

[0107] In some embodiments, electronic devices can be deployed as integrated devices, in addition to serving as training or authentication carriers. For example, some local verification terminals both acquire document images and locally execute the trained lightweight model, displaying the authentication results locally. Such devices may include an image acquisition module, a processor 910, a memory 920, and a display component. The image acquisition module is responsible for acquiring images of the document to be authenticated, the processor 910 is responsible for calling the locally stored document authenticity authentication model to perform inference, and the display component is used to display overall authenticity judgment information and multiple step-by-step authentication information. This integrated implementation method has certain practical value for scenarios with insufficient network conditions or requiring rapid local verification.

[0108] In some embodiments, when the processor 910 executes the training method, it can also invoke auxiliary modules such as image enhancement, region extraction, text encoding, and logging. The image enhancement module can be used to generate richer normal acquisition perturbations; the region extraction module can be used to read key region definitions from the document template; the text encoding module can be used to convert multi-level question text into a text representation acceptable to the model; and the logging module can be used to record the number of training epochs, loss changes, and sample usage. While these auxiliary modules are not necessary prerequisites for the implementation of this application, they can further enhance the completeness of the system implementation.

[0109] In some embodiments, to facilitate subsequent verification of the effectiveness of the proposed solution, the training server 110 can also statistically analyze the overall judgment accuracy, region identification consistency, type judgment consistency, and detail description consistency during the verification phase. Overall judgment accuracy measures the model's performance at the true / false level; region identification consistency measures whether the model can identify suspicious regions corresponding to artifact records; type judgment consistency measures whether the model can correctly classify anomalies; and detail description consistency measures whether the output text covers key anomaly evidence. By observing these metrics separately, the training server 110 can more accurately determine at which task level the model still has shortcomings. These verification metrics help support the completeness and feasibility of the training process.

[0110] In one embodiment, for content logic anomalies detected by the logic verification module 830, the training server 110 can also write them into the second-layer description information and the third-layer description information. For example, when the issuance date in the sample document image is later than the start date of the validity period, the second-layer description information can be written as a content logic anomaly, and the third-layer description information can further specify that the date relationship does not conform to the preset content logic. Similarly, when the facial semantic features corresponding to the portrait area are inconsistent with the gender description in the text field, a corresponding anomaly description can also be formed. In this way, after training, the document authenticity identification model can not only identify image craftsmanship anomalies, but also content logic anomalies.

[0111] To further illustrate the complementary relationship between logical verification and image-side recognition, the overall judgment result can be represented as a combination of image anomaly branches and logical anomaly branches. Let the output of the image anomaly branch be... The logical exception branch output is Then the total judgment score can be written as: (8) in, This indicates the overall score. This represents the fusion weights. In this way, training server 110 or authentication server 120 can establish collaborative judgment between the image side and the logic side. If an image of a document to be authenticated shows no obvious signs of forgery at the image processing level, but the logic verification result is abnormally significant, the comprehensive judgment can still output a high suspicion of forgery. Conversely, if the logic fields appear consistent, but the image side shows obvious printing dots, moiré patterns, or weakened anti-counterfeiting marks, a suspicious judgment can also be formed. The fusion method here only needs to ultimately utilize both types of information to form a comprehensive judgment.

[0112] For the difficult sample feedback module 840, in some embodiments, the training server 110 can not only analyze the artifact parameter distribution of misjudged samples, but also analyze the failure locations of these samples in multi-level tasks. For example, some samples fail in the first-level overall judgment, indicating that the anomaly significance may be insufficient; some samples are correctly judged in the first level but incorrectly judged in the second-level region type, indicating that the model's mapping between regions and types is not yet stable; and some samples have large deviations in the third-level detailed description stage, indicating that the model has detected the anomaly but has not yet learned how to accurately describe it. For different levels of failure types, the training server 110 can adjust the subsequent sample generation strategy and text supervision strategy respectively, thereby making the closed-loop optimization more targeted.

[0113] The above embodiments illustrate that the solution of this application can not only establish a complete closed loop from four levels—sample generation, text supervision, model training, and inference output—but also further enhance training efficiency and discrimination ability through several optional modules. Through the detailed description of these implementation principles and methods, those skilled in the art can implement the document authenticity authentication model training method, document authenticity authentication method, and corresponding electronic equipment described in this application.

[0114] In some embodiments, the training method and the identification method can be described in detail from the perspective of functional units.

[0115] For example, during the training phase, the electronic device or training server 110 can logically include a sample generation unit, a text construction unit, a model training unit, and a parameter update unit. The sample generation unit generates a base document image based on standard document information of a preset type of document, and further injects minute artifacts into the base document image according to an artifact generation strategy corresponding to the forgery process to obtain a sample forged document image. The text construction unit constructs multi-level question text corresponding to the sample forged document image and descriptive information of the minute artifacts. The model training unit inputs the sample forged document image and multi-level question text into the multimodal model to be trained and obtains the sample authenticity identification result output by the model. The parameter update unit adjusts the model parameters based on the deviation between the sample authenticity identification result and the descriptive information, and obtains the document authenticity identification model when a preset training stopping condition is met.

[0116] The internal implementation of the sample generation unit is not limited to a single structure. In one embodiment, the sample generation unit may include a template acquisition module, a content filling module, and an artifact injection module. The template acquisition module is used to acquire a blank document template image and anti-counterfeiting process information. The content filling module is used to generate standard document content based on the anti-counterfeiting process information and fill it into the corresponding positions in the template. The artifact injection module is used to set local anomalies according to the forgery process information, so that the base document image is converted into a sample forged document image. If the system also needs to take into account normal document samples, the sample generation unit can also output a normal base image without injected artifacts for use as a control input during the training phase.

[0117] In one embodiment, the text construction unit may include a question template generation module and a description text generation module. The question template generation module is used to construct first-layer text, second-layer text, and third-layer text. The description text generation module is used to generate corresponding descriptive information based on the artifact regions, artifact types, and artifact parameters in the sample forged document image. If the system adopts a dialogic organization method, the text construction unit may also include a dialog template assembly module, used to write multi-level question texts and corresponding historical context relationships into a unified input sequence. If the system adopts a sequential organization method, the text construction unit can output the image-text-description triplet required for training layer by layer.

[0118] In one embodiment, the model training unit may further include an input preparation module and a training execution module. The input preparation module is responsible for preparing the image, question text, and optional annotations into the input format required by the multimodal model. The training execution module is responsible for driving the multimodal model to be trained to perform forward inference, obtaining the sample authenticity identification results, and iteratively adjusting the model parameters after receiving the gradient or update signal from the parameter update unit. If the multimodal model to be trained adopts a decoupled implementation of the image encoding part and the language inference part, the input preparation module can also pass the corresponding inputs to the visual side and the text side respectively.

[0119] In one embodiment, the parameter update unit may include a deviation calculation module and an update control module. The deviation calculation module compares the differences between the sample authenticity identification results and the descriptive information, and calculates the deviation items corresponding to the overall authenticity judgment, abnormal regions, abnormal types, and abnormal details. The update control module controls the model update process based on the deviation items. If the system introduces an optional logic verification module 830, the parameter update unit can further update the logic verification module parameters based on the differences between the logic verification results and the logic consistency annotations. If the system introduces a hard sample feedback module 840, the parameter update unit can also output the misjudged sample records in the current training round to the hard sample analysis link.

[0120] During the authentication phase, the electronic device or authentication server 120 may logically include an image acquisition unit, a text acquisition unit, a model invocation unit, and a result output unit. The image acquisition unit acquires an image of the document to be authenticated in response to an authentication request. The text acquisition unit acquires multi-level question text for periodically verifying the document's authenticity. The model invocation unit inputs the image of the document to be authenticated and the multi-level question text into the document authentication model and acquires the authentication result output by the model. The result output unit outputs multiple step-by-step authentication information corresponding to the multi-level question text when the authentication result indicates that the document to be authenticated is counterfeit.

[0121] In one embodiment, the image acquisition unit may further include a preprocessing module. The preprocessing module performs at least one of the following processing steps on the image of the document to be authenticated: cropping borders, compressing the image, and modifying the image format. The preprocessing module may also normalize the image size, correct the image orientation, or reduce non-document background areas in the image. As long as these operations do not substantially change the main content of the document, they can be considered as the input processing steps in this application.

[0122] In one embodiment, the text acquisition unit can read preset multi-level question texts from local storage or retrieve a set of question texts matching the current document type from the service configuration center. If the document image to be identified has been classified into a certain document type in a previous stage, the text acquisition unit can call the corresponding first-level text, second-level text, and third-level text based on the classification result. If the document type is not yet clear, the text acquisition unit can also first output a general question template applicable to multiple document scenarios, and then the document authenticity identification model can automatically adapt to different format features during the inference process.

[0123] In one embodiment, the model invocation unit supports both sequential and dialogic invocation mechanisms. For the sequential mechanism, the model invocation unit can invoke the document authenticity verification model layer by layer, first obtaining the overall authenticity judgment, then the abnormal areas and types, and finally the detailed descriptions. For the dialogic mechanism, the model invocation unit can write multi-level question text into the input sequence at once, and the model can directly output multi-level results. The system can select the appropriate implementation method between the two mechanisms based on inference latency requirements, model interface type, and deployment resources.

[0124] In one embodiment, the output unit may include a structured processing module and a display interface module. The structured processing module maps the model output to overall judgment fields, suspicious area fields, forgery feature type fields, and forgery detail fields. The display interface module writes these fields into the review interface, log system, risk engine, or external interface return messages. If the system needs to be linked with the manual review process, the display interface module can also overlay suspicious areas onto the preview view of the document image, so that reviewers can establish a direct correspondence between the image and the text.

[0125] In some embodiments, although the training method and the identification method are described as two stages, they can form a longer-term closed loop through a continuous learning mechanism. Specifically, high-risk samples, misjudged samples confirmed by manual review, or inexplicable abnormal samples encountered by the identification server 120 during actual operation can be fed back to the training server 110 after desensitization and compliance processing. The training server 110 then adjusts the artifact generation strategy, multi-level question text, or descriptive information template based on the image performance, business tags, and manual analysis results of these samples, and retrains the document authenticity identification model. In this way, the model can continuously adapt to new forgery methods as business scenarios change. It should be noted that this continuous learning mechanism may exist as a preferred embodiment, but it is not a necessary prerequisite for the validity of this application.

[0126] In some embodiments, the multi-level question text can also be tailored according to the deployment scenario. If the deployment scenario has high latency requirements, such as needing to quickly complete judgments in turnstiles, border inspection terminals, or mobile law enforcement equipment, the system can only call the first and second level texts, while using the output of the third level detailed information as an on-demand trigger function. If the deployment scenario focuses more on record keeping, auditing, or case description, the system can retain the complete three-level structure, so that the output results contain more sufficient detailed evidence. This shows that although the multi-level question text is preferably three levels, it can also be appropriately combined according to application requirements without deviating from the core idea of ​​this application.

[0127] In some embodiments, the detailed information in the third-layer text can point not only to image process anomalies but also to image-text consistency anomalies. For example, when the model, based on the logic verification module 830, deems the gender field semantically inconsistent with the avatar region, the output of the third-layer text can directly explain the inconsistency, without being limited to a specific type of visual artifact. This implementation indicates that the "forgery details" in this application can more broadly cover both local visual anomaly details and content logic anomaly details. Thus, the "forgery feature type" in the second-layer text can also correspondingly cover both image process anomalies and content logic anomalies.

[0128] In some embodiments, if the image of the document to be authenticated is blurry, obscured by strong reflections, or severely tilted, the authentication server 120 can first output an image quality deficiency warning before deciding whether to continue the complete authentication process. For example, when the main area of ​​the image is covered by a large area of ​​highlight, making key fields invisible, the first-level overall authenticity judgment may not be reliably provided. In this case, the system can first return a "image quality deficiency, it is recommended to re-capture" warning, or route the image to the manual review process. This quality judgment branch is a valuable supplementary logic in actual product implementation and also helps explain why the system sometimes does not directly output a true or false conclusion.

[0129] In some embodiments, the memory 920 in the electronic device, in addition to storing program instructions, can also store configuration parameters related to model invocation, such as the mapping relationship between document types and question text templates, preprocessing parameters, output field definitions, result display rules, and model version identifiers. Before invoking the model, the processor 910 can first read these configuration parameters from the memory 920, and then complete the assembly and inference invocation of the document image to be identified and the multi-level question text. Through this design, the system can adapt to new document types and new output format requirements by modifying the configuration parameters while keeping the model itself unchanged.

[0130] In some embodiments, the electronic device can also store the trained document authentication model in two parts: an inference model file and an auxiliary configuration file. The inference model file is used to perform forward inference under image and text input; the auxiliary configuration file is used to define the question text template, result field structure, and optional post-processing rules. The advantage of this split storage method is that when the business side only needs to adjust the question text or output format, partial adaptation can be completed without retraining the model itself. However, if the question text hierarchy changes significantly, it is still recommended to re-execute training on the training server 110 to ensure that the input format in the inference stage is consistent with that in the training stage.

[0131] In summary, further analysis from the perspectives of deviceization, continuous learning, scene customization, and deployment configuration reveals that the proposed solution does not rely on a single narrow implementation. Instead, it revolves around the core idea of ​​"training a document authenticity verification model through a sample generation mechanism corresponding to the counterfeiting process and a multi-level questioning supervision mechanism, and outputting multiple step-by-step authentication information during the authentication phase." This approach is feasible across multiple implementation paths. Those skilled in the art can complete the corresponding implementation based on the disclosure in this specification without any inventive effort.

[0132] In some embodiments, to further clarify the applicability of this application, the scope of "preset type documents" can be further explained. Preset type documents can be personal identification documents issued by state organs, public management institutions, or other authorized entities, such as resident identity cards, passports, driver's licenses, residence permits, student IDs, work permits, and social security cards; or institutional documents related to unit qualifications, business licenses, or specific qualifications, such as business licenses, permits, filing certificates, qualification certificates, or industry access certificates. This application does not require all documents to have completely identical layout structures. As long as the document type can provide relatively stable template information, field layout information, or anti-counterfeiting technology information, the corresponding base document image can be generated using this application, and further sample counterfeit document images can be constructed.

[0133] In some embodiments, although the foregoing examples mainly focus on forgery traces such as color printing, scan-and-reprint, character distortion, weakening of anti-counterfeiting marks, and local texture replacement, the present application is not limited to these specific artifact types. In other words, as long as a certain local visual anomaly originates from an unofficial manufacturing process and can form a locally significant anomaly pattern in the document image, it can be regarded as a "minor artifact" in the present application. For example, some counterfeit documents may exhibit edge feathering, significant inconsistency between the sharpness of local areas and adjacent areas, inconsistency between local noise distribution and the overall background, and abnormal local color channel alignment in local areas. As long as these anomalies can be introduced, recorded, and described parametrically, they can also be included in the construction scope of the sample counterfeit document image.

[0134] In some embodiments, the aforementioned "descriptive information" is primarily given in text form, as current multimodal models typically output directly to text. However, without altering the essence of the scheme described in this application, the descriptive information may also include structured fields, label sequences, region numbers, anomaly type code values, or field logical state values. For example, the second-layer descriptive information may not be directly written as a natural language sentence, but rather a structured record constructed using region numbers and type labels, which is then converted into a text format suitable for model training by the text generation module before training. Alternatively, the model training itself may simultaneously receive text supervision and structured supervision. As long as this information is still organized around the anomaly location, anomaly type, and anomaly details in the sample forged document image, it can be considered to fall within the broad scope of the descriptive information described in this application.

[0135] In some embodiments, the "multi-level question text" in this application does not necessarily require the use of natural language sentences ending with a question mark. Its essence lies in using a set of hierarchical text inputs to guide the model to complete a layer-by-layer analysis from overall judgment to local attribution. Therefore, the text can be expressed as interrogative sentences, instruction sentences, task sentences, or structured prompts. For example, "Please judge the overall authenticity of this document," "Explain the suspicious areas and anomaly types," and "Give detailed evidence of the anomaly" can all be considered as multi-level question texts used for phased inquiries about the authenticity of documents. As long as these texts play a hierarchical guiding role during the training and identification phases, they do not affect their inclusion within the scope of protection of this application.

[0136] In some embodiments, the "sample authenticity identification results" and "authenticity identification results" mentioned above can also adopt different organizational forms depending on the system implementation. For example, in one implementation, they can be continuous text directly output by the model; in another implementation, they can be composite structure objects containing true / false labels, region coordinates, type code values, and detailed text; in yet another implementation, they can also be the final result obtained by first outputting intermediate markers by the model and then reassembling them by the post-processing module. Since one embodiment of this application focuses on the technical content carried by these results and their formation method, rather than a certain fixed result encapsulation format, as long as the output can reflect the overall authenticity judgment and the step-by-step information corresponding to the multi-level question text, it should be considered consistent with the solution of this application.

[0137] In some embodiments, the order of steps in the training and identification phases can be locally adjusted, as long as the data dependencies are not disrupted. For example, in some implementations, the training server 110 can generate multi-level question text templates in batches after generating the base document image, and then fill in the corresponding descriptive information for each sample; in other implementations, descriptive information can be generated first based on the artifact injection record, and then the corresponding question template can be selected based on the descriptive information. Similarly, in the identification phase, the identification server 120 can first read the multi-level question text and then receive the document image to be identified; or it can first receive the document image to be identified and complete preprocessing, and then retrieve the corresponding question text according to the document type. As long as a closed loop of "image object + multi-level question text + model inference + result output" is ultimately formed, the validity of the proposed solution remains unaffected.

[0138] In some embodiments, certain functional modules in this application can be implemented in combination or separately. For example, the text construction unit and description information generation module in the training phase can be merged into a unified annotation construction unit; the text acquisition unit and model invocation unit in the identification phase can also be unified into an inference scheduling unit in program implementation. As another example, if the model is deployed locally on the electronic device, the result output unit and the display interface module can also be implemented by the same application. For those skilled in the art, these module merging or splitting are merely changes in engineering implementation methods and do not alter the core technical ideas of this application.

[0139] In some embodiments, some technical actions in this application can be implemented by software or by a combination of software and hardware. For example, steps such as image preprocessing, text template reading, model inference scheduling, and output result processing can typically be executed by a processor driven by program instructions; while steps such as image acquisition, communication transmission, and display output can be completed by corresponding hardware modules. Large-scale sample generation and parameter updates required during the training phase can also be completed in parallel by a server cluster. As long as the relevant hardware and software jointly support the execution of the aforementioned training and discrimination methods, it should be considered one of the implementation forms of this application.

[0140] To further summarize the technical significance of the aforementioned embodiments, we can conclude from the relationship between problem, means, and result. In the task of automatic document authentication, existing technologies often struggle to simultaneously address both the ability to identify fine-grained forgery traces and the step-by-step output capability of the authentication process. One embodiment of this application does not simply rely on a general model to directly answer the true / false question. Instead, it first generates a base document image based on standard document information, injects minute artifacts according to forgery techniques to construct training samples, and then decomposes the true / false judgment task into a hierarchical training task using multi-level question text and descriptive information. Thus, the model learns not only "true / false labels" during the training phase but also the correspondence between "overall judgment—suspicious area—anomaly type—detailed evidence." Furthermore, during the authentication phase, by inputting the document image to be authenticated along with the multi-level question text into the trained document authentication model, the model can output the authentication result and multiple step-by-step authentication information along the path learned during training. Therefore, one embodiment of this application can improve the ability to identify local anomaly patterns and enhance the verifiability of the results in document image authentication scenarios.

[0141] It should also be noted that the specific numerical values, document names, artifact examples, text examples, and model organization methods given in the foregoing embodiments are only for helping to understand the implementation principles and methods of this application, and should not be construed as limiting the scope of protection of this application. Those skilled in the art, without departing from the core ideas of this application, can substitute, adjust, or combine the base document image generation method, artifact injection method, multi-level question text organization method, model training method, result output method, and electronic device deployment method according to different document types, different acquisition scenarios, different model interfaces, and different business requirements. These substitutions, adjustments, or combinations, as long as they do not deviate from the technical logic established by this application regarding the training and identification methods for document authenticity identification models, can all be considered equivalent implementations of the technical solutions disclosed in this application.

[0142] Thus far, the details have been fully explored from multiple perspectives, including system environment, main training method flow, base document image construction, artifact injection, multi-level question text construction, descriptive information generation, model training, main authentication method flow, multiple step-by-step authentication information outputs, electronic device execution mode, efficient parameter fine-tuning, logic verification module, hard sample feedback, and device-based implementation. Based on this disclosure, those skilled in the art can implement the document authenticity authentication model training method and document authentication method described in this application, and accordingly complete the construction and deployment of corresponding electronic devices or program systems.

[0143] It should be understood that the method steps in the foregoing embodiments can be implemented by program instructions controlling related hardware, or by dedicated circuits, programmable logic devices, or a combination thereof. Correspondingly, the systems, devices, modules, units, or components in the foregoing embodiments can be implemented in software, hardware, or a combination of both. The division of modules, units, or components is merely a logical division for the purpose of illustrating the technical solution; in actual implementation, they can be combined, split, or integrated as needed.

[0144] In one embodiment, the electronic device may include a processor, a memory, and a communication interface, wherein the memory is used to store program instructions, and the processor is used to call and execute the program instructions to implement all or part of the steps in the foregoing method embodiments. The electronic device may be a server, a terminal device, an edge computing node, a cloud computing device, or other device with data processing capabilities.

[0145] In one embodiment, this application may also be implemented in the form of a computer-readable storage medium. The computer-readable storage medium stores a computer program, which, when executed by a processor, causes the processor to implement all or part of the steps in the foregoing method embodiments. The computer-readable storage medium may be a read-only memory, random access memory, flash memory, hard disk, solid-state drive, optical disk, or other non-transitory storage medium.

[0146] The various embodiments in this specification are described in a progressive manner. Similar or identical parts between embodiments can be referred to mutually. Each embodiment focuses on describing the differences from other embodiments. In particular, system embodiments are basically similar to method embodiments, so the description is relatively simple; relevant parts can be referred to the descriptions in the method embodiments. In the description of this specification, the terms "one embodiment," "some embodiments," "example," "specific example," or "some examples," etc., refer to specific features, structures, materials, or characteristics described in connection with that embodiment or example, which are included in at least one embodiment or example of this specification. In this specification, the illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described can be combined in any suitable manner in one or more embodiments or examples. Moreover, without contradiction, those skilled in the art can combine and integrate the different embodiments or examples described in this specification and the features of different embodiments or examples.

[0147] Furthermore, the terms "including," "comprising," and "having" used in the specification are all non-exclusive inclusions; the terms "first," "second," etc., are only used to distinguish technical features and do not indicate limitations on order, quantity, or importance. The execution order of each step in the method embodiments is also not absolutely limited. Without departing from the technical concept of this application, the steps can be adjusted in order, executed in parallel, combined, or split for execution.

Claims

1. A training method for a document authenticity verification model, comprising: A base document image is generated based on the standard document information of a preset type of document, and tiny artifacts are injected into the base document image according to the artifact generation strategy corresponding to the forgery process to obtain a sample forged document image. Construct a multi-level question text corresponding to the sample forged document image and descriptive information of the micro-artifacts, wherein the multi-level question text is used to inquire about the authenticity of the document in stages; The sample forged document image and the multi-level question text are input into the multimodal model to be trained, and the authenticity identification result of the sample forged document image is obtained from the output of the multimodal model. The model parameters of the multimodal model are adjusted based on the deviation between the sample authenticity identification results and the description information of the micro-artifacts, and the document authenticity identification model is obtained when the preset training stopping condition is met.

2. The training method of claim 1, wherein, The process of generating a base document image based on standard document information of a preset type of document includes: Obtain the blank document template image and anti-counterfeiting process information of the preset type of document; Generate corresponding standard certificate content according to the anti-counterfeiting process information; The standard document content is filled into the corresponding position in the blank document template image to obtain the base document image.

3. The training method of claim 1 or 2, wherein, The process of injecting minute artifacts into the base document image according to an artifact generation strategy corresponding to the forgery process includes: Based on the forgery process information, the preset parameters of the base document image are adjusted to inject the micro artifact into the base document image, which has at least one of preset spatial position, preset morphological features and preset physical properties.

4. The training method of claim 3, wherein, The micro-artifacts include at least one of the following: Printed dots are superimposed on the portrait area and / or text area in the base document image; A scanned moiré pattern is superimposed on the base document image; The font in the base document image is deformed; Remove or weaken the anti-counterfeiting features in the base document image; Replace the local texture in the base document image with a preset material texture.

5. The training method of claim 1, wherein, The multi-level question text includes: The first layer of text used to inquire about the overall authenticity of the document; The second layer of text used to query forged feature types; The third layer of text is used to inquire about the details of the forgery.

6. A method for verifying the authenticity of a document, comprising: In response to a verification request initiated for a document to be verified, an image of the document to be verified is acquired. Obtain multi-level question texts for periodic verification of document authenticity; Input the document image and the multi-level question text into the document authenticity identification model, and obtain the authenticity identification result output by the document authenticity identification model after reasoning; If the authentication result indicates that the document to be authenticated is a counterfeit document, multiple step-by-step authentication information corresponding to the multi-level question text will be output.

7. The method of authenticating an identification document of claim 6, wherein, The multiple step-by-step authentication information includes: Information for assessing the overall authenticity of the document; Information on suspicious areas; Forged feature type information; At least one of the falsified details.

8. A method of authenticating a security document as claimed in claim 6 or 7 wherein, The multi-level question text includes: The first layer of text used to inquire about the overall authenticity of the document; The second layer of text used to query forged feature types; The third layer of text is used to inquire about the details of the forgery.

9. The method of authenticating an identification document of claim 6, wherein, Before inputting the document image and the multi-level question text used for periodically verifying the document's authenticity into the document authentication model, the following steps are also included: The document image is preprocessed, and the preprocessing includes at least one of cropping the border, compressing the image, and modifying the image format.

10. An electronic device, comprising: processor; Memory, the memory being used to store computer instructions; The processor is configured to invoke the computer instructions to execute the steps of the method as described in any one of claims 1 to 9.