A training method and an authentication method of an offline signature authentication model and related equipment

CN122780975APending Publication Date: 2026-09-18SOUTH CHINA UNIV OF TECH
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202610945648.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-06-29
Publication Date
2026-09-18

AI Technical Summary

Technical Problem

常规PPO(近端策略优化)或GRPO(群组相对策略优化)类方法通常通过较强约束限制策略更新,这虽有助于稳定训练,但也容易将模型限制在初始次优分布附近,难以充分探索更有效的证据比较方式和报告生成策略

Benefits of technology

[0023] The embodiments of this application include at least the following beneficial effects: This application provides a training method, authentication method, electronic device, storage medium, and program product for an offline signature authentication model. This scheme pre-establishes the signature authentication chain inference format and report generation capability through cold-start supervised fine-tuning, providing a stable initial strategy for subsequent reinforcement learning. On this basis, unconstrained group policy optimization (UGPO) is adopted, using the total number of tokens of all candidate outputs as a unified normalized denominator, fundamentally eliminating the systematic gradient bias introduced by the existing GRPO algorithm due to output length differences, thus making the length differences in signature authentication reports more manageable. In the significantly open-ended generation task, the update contribution weight of each token is strictly consistent. Simultaneously, UGPO removes the probability ratio pruning constraint and the reference policy KL penalty term, expanding the feasible domain of policy updates from the local neighborhood near the base model to the entire parameter space. This fully unleashes the model's ability to explore sparse, high-value inference paths, thereby overcoming the suboptimal distribution limitations of the base model and effectively improving the accuracy of true/false discrimination and the evidentiary sufficiency of the identification report. Furthermore, by calculating the dominance value through within-group mean-variance standardization, the number of positive and negative dominance samples is balanced, maintaining numerical stability during training even under unconstrained conditions. In summary, the technical solution of claim 1 can significantly improve the model's exploration efficiency for high-quality inference chains and the credibility and interpretability of the final output report without requiring additional training of the reward model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122780975A_ABST
    Figure CN122780975A_ABST
Patent Text Reader

Abstract

The application provides a training method and an authentication method of an offline signature authentication model and related equipment, and relates to the technical field of computer vision. The method comprises the following steps: obtaining a training data set; generating RGB original image data and grayscale normalized data for a signature image respectively, extracting global visual tokens through a visual Transformer encoder, extracting expert visual tokens through an expert visual encoder, and inputting the spliced data into a large language model to construct a basic model; adopting cold start supervision fine tuning to establish chain reasoning and identification report generation capability; performing reinforcement learning training through unconstrained group strategy optimization, and uniformly normalizing the total token number of all candidate outputs; introducing a signature identification report optimization reward, evaluating report quality from key inspection point coverage, detail description similarity and format compliance, and adding the reward to the total reward through a gradual trigger criterion. The application enables the model to generate an identification report containing a traceable reasoning chain and detailed explanation, and is suitable for scenarios such as bill verification and contract signing verification.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the fields of computer science, biometric recognition, pattern recognition, authenticity discrimination and multimodal large models, and in particular to a training method, authentication method and related equipment for an offline signature authentication model. Background Technology

[0002] Handwritten signatures have long been widely used in scenarios such as financial instruments, contract signing, forensic identification, identity verification, and document management. Offline signature authentication refers to determining whether a signature belongs to the claimant based solely on a static image of the signature obtained through scanning or photography. Compared to online signature authentication, offline signature authentication does not rely on dedicated acquisition equipment and has a wider range of applications, but it is also more difficult to utilize dynamic information such as writing speed, pressure, and trajectory time series.

[0003] Existing offline signature authentication methods mainly include handcrafted feature extraction methods and deep learning methods. Handcrafted feature extraction methods typically extract features such as shape, texture, contour, and local structure, and then output true / false results through distance metrics or classifiers. Deep learning methods typically use convolutional neural networks, Siamese networks, dual-channel networks, or visual Transformers to learn signature representations. These methods have achieved good binary classification performance on several datasets, but most methods can only output similarity scores, true / false labels, or acceptance / rejection results.

[0004] In high-risk scenarios such as judicial, banking, and contract review, binary classification results alone are insufficient to support review, appeals, and accountability. In practical applications, signature authentication experts typically need to formulate structured authentication opinions based on evidence such as stroke trends, starting and ending strokes, character structure, character spacing, overall layout, and proportional relationships. Existing automated signature authentication systems lack such human-readable analysis reports, making it difficult to explain "why it was determined to be genuine / forged," thus limiting the system's credibility and feasibility.

[0005] General-purpose multimodal large models can process images and text simultaneously and generate natural language interpretations, thus possessing the potential to be applied to interpretable signature authentication. However, these models typically lack specialized training for fine-grained visual differences in signatures, document authentication terminology, and signature authenticity reasoning data. When directly used for signature authentication, they are prone to overlooking local stroke differences, generating category bias, or producing reports that do not conform to authentication logic.

[0006] The applicant previously proposed an interpretable offline signature authentication method based on a multimodal large model (CN120510619A), which achieved interpretable signature authentication report generation by combining a visual Transformer with a local signature visual representation enhancement sub-model. However, this method mainly relies on supervised fine-tuning, and there is still room for improvement in terms of report quality, reasoning sufficiency, and format consistency.

[0007] The key to interpretable signature authentication is not merely outputting a true or false conclusion, but rather forming a traceable chain of reasoning based on visual evidence. The model needs to compare the consistency and differences between the reference signature and the signature under investigation in terms of local strokes, stroke beginnings and endings, connection methods, character structure, and overall layout, and generate an authentication opinion accordingly. Furthermore, compared to tasks that only output standard answers, authentication reports have greater openness and expressive diversity. Therefore, the optimization objective of this task cannot be simply reduced to a reward design based on answer accuracy or fixed format constraints. Simply relying on accuracy rewards or format rewards is insufficient to effectively characterize the evidentiary sufficiency, reasoning consistency, and report quality of the authentication conclusion.

[0008] Furthermore, directly applying general reinforcement learning-based post-training algorithms to signature authentication still has shortcomings. Conventional PPO (proximal policy optimization) or GRPO (group relative policy optimization) methods typically restrict policy updates through strong constraints. While this helps stabilize training, it also tends to limit the model to the vicinity of the initial suboptimal distribution, making it difficult to fully explore more effective evidence comparison methods and report generation strategies. Summary of the Invention

[0009] The main objective of this application is to propose a training method, authentication method, electronic device, storage medium, and program product for an offline signature authentication model based on reinforcement learning. This model can receive a reference signature image and a signature image to be inspected, determine the authenticity of the signature to be inspected through step-by-step chain reasoning, and output a natural language analysis report that conforms to the document authentication standard.

[0010] To achieve the above objectives, one aspect of this application proposes a training method for an offline signature authentication model, the method comprising: Obtain a training dataset, wherein each sample in the training dataset contains a reference signature image, a signature image to be tested, an authentication instruction, a chained reasoning process, and an authentication report containing the authenticity verification conclusion; The reference signature image and the signature image to be inspected are preprocessed to generate RGB original image data and grayscale normalized data respectively; A basic model is constructed, which includes a visual Transformer encoder, an expert visual encoder, a multilayer perceptron mapping layer, and a large language model. The original RGB image data is input into the visual Transformer encoder to extract global visual terms, and the grayscale normalized data is input into the expert visual encoder to extract expert visual terms. The global visual terms and expert visual terms are mapped to the hidden dimensions of the large language model through the multilayer perceptron mapping layer and then concatenated and input into the large language model. The training dataset is used to perform cold-start supervised fine-tuning on the base model, so that the model outputs a chain inference process and identification report in a preset format; The model after cold start is used as the initial policy model, and reinforcement learning training is carried out using unconstrained group policy optimization. For each input cue, multiple candidate outputs are sampled and generated. The reward value of each candidate output is calculated separately. The advantage value is calculated using within-group mean-variance standardization. The total number of tokens of all candidate outputs is uniformly normalized. The policy is updated with the goal of maximizing the sum of the products of the normalized policy ratio and the advantage value. No pruning constraint is imposed on the policy ratio and no KL divergence penalty term of the reference policy is introduced. The reward value includes a format reward, an accuracy reward, and a signature verification report quality reward. The signature verification report quality reward is calculated based on the coverage of key verification points, the similarity of detailed descriptions, and the format compliance between the generated report and the labeled report. The signature verification report quality reward is added to the total reward and participates in the strategy update after the triggering criteria are met.

[0011] In some embodiments, the objective function for optimizing the unconstrained group policy is:

[0012] in, The number of sampled candidate outputs for each input prompt. For the first 1 candidate output, This indicates the number of candidate output tokens. and The current strategy and the old strategy are respectively in the 1st and 2nd phases. The output of the first The probability at each token This is the dominant value; This indicates an input cue sampled from the training data distribution. and from the old strategy model The sampled input prompts are in the middle. candidate outputs Find the expected value.

[0013] In some embodiments, the advantage value is calculated such that each token of the G candidate outputs within the same group shares the same advantage value, which is equal to the reward value of the candidate output standardized relative to the mean and standard deviation of the G reward values ​​within the group.

[0014] In some embodiments, the similarity of the detailed descriptions is calculated in the following manner: For each key inspection point covered in the generated report, extract the sentence vector representations of the generated description and the labeled description respectively, and calculate the cosine similarity. The cosine similarity is subjected to exponential sharpening, and the p-th power of the cosine similarity is used as the sharpened similarity, where p>1; The average of the sharpened similarities of each key point is used as the detail description similarity.

[0015] In some embodiments, the triggering criterion is: within a consecutive n-step window ending at the current training step k, if the sum of the format reward and the accuracy reward at each step exceeds the threshold t, then starting from the (k+1)th step, the signature authentication report excellence reward is included in the total reward; otherwise, the total reward consists only of the format reward and the accuracy reward.

[0016] In some embodiments, the weighted calculation formula for the signature authentication report excellence reward is as follows:

[0017] in, To ensure coverage of key inspection points, To describe the similarity in detail, For format compliance; , , As weight, and ; The coverage of key inspection points ,in This is a collection of key inspection points in the annotation report. To generate a subset of key points already covered in the report; The format compliance ,in To report the total number of structural elements required, Indicates the first Does the item structure element exist?

[0018] In some embodiments, during the cold start supervised fine-tuning, the chained inference process is placed in... <think>and< / think> The identification report is placed between the labels. <answer> and< / answer> Between tags; the chain reasoning process is organized in the order of observing local strokes, comparing character structures, analyzing global layout, and making a comprehensive judgment on authenticity.

[0019] To achieve the above objectives, another aspect of this application proposes an offline signature authentication method, comprising the following steps: Obtain the reference signature image to be authenticated and the signature image to be inspected; The reference signature image and the signature image to be tested are input into the offline signature authentication model trained using the training method described above. Obtain the chain-like reasoning process and authentication report output by the model. The authentication report includes a judgment on the authenticity of the signature to be examined relative to a reference signature and an interpretability analysis based on visual evidence.

[0020] To achieve the above objectives, another aspect of this application provides an electronic device, which includes a memory and a processor. The memory stores a computer program, and the processor executes the computer program to implement the method described above.

[0021] To achieve the above objectives, another aspect of the embodiments of this application proposes a computer-readable storage medium storing a computer program that, when executed by a processor, implements the method described above.

[0022] To achieve the above objectives, another aspect of the embodiments of this application proposes a computer program product, including a computer program that, when executed by a processor, implements the method described above.

[0023] The embodiments of this application include at least the following beneficial effects: This application provides a training method, authentication method, electronic device, storage medium, and program product for an offline signature authentication model. This scheme pre-establishes the signature authentication chain inference format and report generation capability through cold-start supervised fine-tuning, providing a stable initial strategy for subsequent reinforcement learning. On this basis, unconstrained group policy optimization (UGPO) is adopted, using the total number of tokens of all candidate outputs as a unified normalized denominator, fundamentally eliminating the systematic gradient bias introduced by the existing GRPO algorithm due to output length differences, thus making the length differences in signature authentication reports more manageable. In the significantly open-ended generation task, the update contribution weight of each token is strictly consistent. Simultaneously, UGPO removes the probability ratio pruning constraint and the reference policy KL penalty term, expanding the feasible domain of policy updates from the local neighborhood near the base model to the entire parameter space. This fully unleashes the model's ability to explore sparse, high-value inference paths, thereby overcoming the suboptimal distribution limitations of the base model and effectively improving the accuracy of true / false discrimination and the evidentiary sufficiency of the identification report. Furthermore, by calculating the dominance value through within-group mean-variance standardization, the number of positive and negative dominance samples is balanced, maintaining numerical stability during training even under unconstrained conditions. In summary, the technical solution of claim 1 can significantly improve the model's exploration efficiency for high-quality inference chains and the credibility and interpretability of the final output report without requiring additional training of the reward model. Attached Figure Description

[0024] Figure 1This is a flowchart of a training method for an offline signature authentication model provided in an embodiment of this application; Figure 2 This is a flowchart illustrating a training method for an interpretable offline signature authentication multimodal inference model provided in the application embodiment; Figure 3 This is a schematic diagram of the hardware structure of the electronic device provided in the embodiments of this application. Detailed Implementation

[0025] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative of this application and are not intended to limit it. In the following description, when referring to the accompanying drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with those of this application; they are merely examples of apparatuses and methods consistent with some aspects of the embodiments of this application as detailed in the appended claims.

[0026] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application belongs. The terminology used herein is for the purpose of describing embodiments of this application only and is not intended to limit this application.

[0027] like Figure 1 As shown, this embodiment provides a training method for an offline signature authentication model, specifically including the following steps: Step S1: Obtain the training dataset. Each sample in the training dataset contains a reference signature image, a signature image to be tested, an authentication instruction, a chain reasoning process, and an authentication report containing the authenticity verification conclusion. Step S2: Preprocess the reference signature image and the signature image to be inspected respectively to generate RGB original image data and grayscale normalized data; Step S3: Construct a basic model, which includes a visual Transformer encoder, an expert visual encoder, a multilayer perceptron mapping layer, and a large language model; input the RGB original image data into the visual Transformer encoder to extract global visual terms, input the grayscale normalized data into the expert visual encoder to extract expert visual terms, map the global visual terms and expert visual terms to the hidden dimensions of the large language model through the multilayer perceptron mapping layer, concatenate them, and input them into the large language model; Step S4: Use the training dataset to perform cold start supervised fine-tuning on the base model, so that the model outputs the chain inference process and identification report in a preset format; Step S5: Using the cold-started model as the initial policy model, reinforcement learning training is performed using unconstrained group policy optimization; for each input cue, multiple candidate outputs are sampled and generated, the reward value of each candidate output is calculated, the advantage value is calculated using within-group mean-variance standardization, and uniform normalization is performed based on the total number of tokens of all candidate outputs. The policy is updated with the goal of maximizing the sum of the products of the normalized policy ratio and the advantage value, where no pruning constraint is imposed on the policy ratio and no KL divergence penalty term of the reference policy is introduced; The reward value includes a format reward, an accuracy reward, and a signature verification report quality reward. The signature verification report quality reward is calculated based on the coverage of key verification points, the similarity of detailed descriptions, and the format compliance between the generated report and the labeled report. The signature verification report quality reward is added to the total reward and participates in the strategy update after the triggering criteria are met.

[0028] The solutions of the embodiments of this application will be described in detail and explained below with reference to specific application examples.

[0029] Reference Figure 2 This embodiment provides a training method for an interpretable offline signature authentication multimodal inference model, including the following steps: (1) Data preparation: Obtain the offline signature instruction fine-tuning dataset, which should include signature images, authentication instructions input to the large model, chained reasoning process, and a detailed analysis report containing the authenticity verification conclusions.

[0030] (2) Data preprocessing: Obtain a reference signature image and a signature to be inspected, wherein the reference signature is a known real signature and the signature to be inspected is a signature that needs to be authenticated; the two signatures are respectively formed into RGB original image input and grayscale normalized input.

[0031] (3) Basic model construction: The RGB signature image is input into the visual Transformer (ViT) encoder of the multimodal large model to extract global visual lexical units; the grayscale signature image is input into the expert visual encoder to extract expert visual lexical units related to fine-grained strokes, character structure and overall shape. The global visual lexical units and the signature expert visual lexical units are mapped to the same dimension and then concatenated and input into the large language model to form the basic model of interpretable signature authentication.

[0032] (4) Cold start training: Supervised fine-tuning of the interpretable signature authentication base model using labeled signature authentication samples, so that the model outputs the chain reasoning process (thought chain) and the final authentication report in a specified format. The thought chain is placed in... <think> and< / think> Between the tags, the final report is placed <answer> and< / answer> Between tags.

[0033] (5) Reinforcement learning training: Based on the base model after cold start, reinforcement learning training is carried out by unconstrained group strategy optimization; multiple candidate outputs are sampled for each input prompt, and the format reward, accuracy reward and open report quality reward are calculated respectively, and the model strategy is updated by taking advantage of in-group normalization.

[0034] (6) Stability-aware reward modeling for open generation: Set the signature identification report excellence reward to evaluate the coverage of key inspection points, similarity of detailed description and format compliance between the generated report and the manual report. After the format reward and accuracy reward are stable, the signature identification report excellence will be gradually added to the total reward to reduce the impact of reward noise on policy update in the early stage of training.

[0035] (7) Model reasoning: Input the reference signature and the signature to be tested into the trained interpretable signature authentication model. The model outputs the authenticity / forgery judgment of the signature to be tested relative to the reference signature, as well as an interpretable authentication report based on key points such as local strokes, character shape, structural proportion, and spatial layout.

[0036] In one implementation, the dataset used to train the model in step (1) consists of a reference signature image, a signature image to be inspected, an authentication instruction, a chained reasoning process, and a final report. The reference signature is a known genuine signature, while the signature to be inspected may be a genuine signature or a forged signature. The authentication instruction can be set to the Chinese "Please verify whether Signature 2 is genuine or forged by comparing it with Signature 1." or the English "Please determine whether Signature 2 is genuine or forged by comparing it against Signature 1."

[0037] The chain-like reasoning process is organized in the order of local to global and then to a comprehensive conclusion. Local reasoning includes observing the curvature, starting and ending points, and turns of strokes character by character or letter by letter; global reasoning includes the overall style, proportion, slant, spacing, and overall rhythm of the signature; comprehensive reasoning combines local and global evidence to give a judgment of truth or falsehood. The final report provides the verification results, key points of verification, and explanations for each item.

[0038] The Chinese report format can be: "Verification result: Signature 2 is genuine / forged compared to signature 1."

[0039] Inspection points: Inspection point 1, Inspection point 2, Inspection point 3,... 1. Inspection Point 1: [Detailed Description] 2. Inspection Point 2: [Detailed Description] 3. Inspection Point 3: [Detailed Description] ..." The English report format can be: “Verification Result: Compared to Signature 1, Signature 2 isgenuine / forged. Verification Aspects: Verification Aspect 1, Verification Aspect 2, Verification Aspect 3, ... Verification Aspect 1: [Detailed description] Verification Aspect 2: [Detailed description] Verification Aspect 3: [Detailed description] ..." In one implementation, step (2) involves two types of processing for both the reference signature and the signature to be inspected: one type retains the original RGB image, which is then input into the ViT visual encoder to extract the overall structure and layout features of the signature; the other type converts it to a grayscale image, removes background interference, and unifies it to a preset size (96×336; height×width), which is then input into the signature expert visual encoder to extract fine-grained stroke features. Text instructions, chained reasoning, and the identification report are processed by word segmentation according to the vocabulary of the basic large language model.

[0040] As one implementation method, in step (3), the base model can be a multimodal large model that supports multiple image inputs, such as Qwen2-VL-2B-Instruct. The native multimodal ViT visual encoder is mainly used to extract global visual information from the two signature images, while the expert visual encoder is mainly used to supplement more fine-grained signature features such as stroke trends, local textures, character structures, and overall signature shape.

[0041] As an alternative implementation, the basic multimodal large model can be replaced with other multimodal large models that support multiple image inputs and Chinese / English language generation, such as Qwen3-VL, InternVL, MiniCPM-o, or other visual language models, as long as they can receive reference signatures and signatures to be checked and output text.

[0042] As one implementation method, the expert visual encoder can employ a lightweight convolutional network structure (such as ResNet-18) and be pre-trained using a signature authenticity binary classification task to enable it to perceive signature differences. The input, main operations, and output functions of each component module of the model are shown in the table below.

[0043] As another implementation, the expert visual encoder can be replaced with ResNet, DenseNet, SwinTransformer, Vision Transformer, Feature Pyramid Network, Siamese Network, or other feature extraction networks suitable for signature images; the number of expert visual terms, input size, and fusion method can also be adjusted according to computing resources.

[0044] As one implementation method, the global visual lexical units output by the ViT visual encoder and the expert visual lexical units output by the expert visual encoder are mapped to the hidden dimensions that the large language model can accept using a multilayer perceptron (MLP), and then concatenated to serve as the visual context of the large language model (LLM). The large language model combines the visual context and textual instructions to generate a chain-like reasoning process and a final identification report.

[0045] The inputs, main operations, and output functions of each component module of the model are shown in Table 1: Table 1

[0046] As one implementation method, in step (4), the cold start training adopts a supervised fine-tuning approach. The input consists of a reference signature, a signature to be inspected, and an authentication command, with the label being the corresponding chained reasoning and report concatenation text. Through this training stage, the model learns signature authentication terminology, structured comparison procedures, and specified output formats.

[0047] The parameter configurations for cold start training are shown in Table 2 below: Table 2

[0048] In one implementation, in step (5), the reinforcement learning phase uses the cold-started model as the initial policy model. For the same input cue, the model generates multiple candidate outputs and calculates scores for each based on the reward function; then, relative comparisons are performed within the same group of candidate outputs, so that high-quality outputs are positively updated and low-quality outputs are suppressed.

[0049] The formula for unconstrained group strategy optimization proposed in this embodiment is:

[0050] in, The number of sampled candidate outputs for each input prompt. For the first 1 candidate output, For its token quantity, and The current strategy and the old strategy are respectively in the 1st and 2nd phases. The output of the first The probability at each token This represents the corresponding advantage value.

[0051] Compared to the existing GRPO algorithm, the UGPO algorithm in this embodiment differs fundamentally in its normalization method. GRPO employs a double averaging approach, first averaging within each candidate output and then averaging across candidate outputs. The normalization coefficients result in shorter candidates receiving larger single-token weights, while longer candidates have their single-token weights compressed, introducing systematic bias in tasks with significant differences in inference length. This embodiment uses the total number of tokens from all candidate outputs instead. As a unified normalized denominator, it fundamentally eliminates the systematic gradient bias introduced by length differences, ensuring that the contribution weight of each token in the group to the policy update is strictly consistent, making it more suitable for open generation tasks such as signature authentication reports with significant differences in output length.

[0052] Furthermore, UGPO also removed the probability ratio pruning constraint and the reference policy KL penalty term. The KL penalty term is formally equivalent to applying a soft attracting field centered on the reference policy to the policy update, the strength of which is proportional to the penalty coefficient. When the reference policy is a base model with insufficient domain capabilities, this attraction field continuously pulls the policy optimization trajectory toward the suboptimal attractor, causing the policy to fail to converge to that region even if a high-reward inference path exists due to excessive penalty costs—essentially a design choice that sacrifices training stability and final performance for path conservatism. The danger of probability ratio pruning constraints lies in their asymmetry: for positively favored samples, the pruning operation will... Cut off to Within this range, the gradient magnitude will be clipped to... The upper bound of the dominance value, and the first-order term of the dominance value. In most high-dominance situations, the number of pruning options is significantly greater than that of pruning options, causing pruning to systematically suppress the learning rate of high-dominance actions, forming an implicit barrier that prevents the strategy from crossing the suboptimal distribution.

[0053] This embodiment removes both of the above constraints, expanding the feasible domain for policy updates from a local neighborhood near the reference policy to the entire parameter space. In highly specialized and heterogeneous tasks such as signature authentication, high-quality inference paths occupy only a very small proportion of the solution space and are often far from the base model; fully exploring this sparse, high-value region is a necessary condition for overcoming domain performance bottlenecks.

[0054] Each candidate output Corresponding advantage value The within-group mean-variance standardization was used for calculation, as follows:

[0055] That is, each candidate output within the group All tokens share the same advantage value. This value is determined by its reward. Compared to the same group The mean and standard deviation of each candidate reward are standardized. This design ensures that the number of positive and negative advantage samples is nearly balanced in each update (approximately half each), maintaining numerical stability of training even after removing KL and pruning constraints.

[0056] In one implementation, the reward function in step (6) consists of a format reward, an accuracy reward, and a signature verification report quality reward. The format reward is used to determine whether the generated text contains... <think> 、< / think> , <answer> and< / answer> The essential structure of a signature authentication report includes: an accuracy bonus to determine whether the report's conclusions on authenticity / forgery are consistent with the label; and a signature authentication report quality bonus to evaluate the quality of open-ended authentication reports. The signature authentication report quality bonus quantifies the quality gap between generated reports and manually annotated reports across three dimensions: (a) Coverage of key inspection points

[0057] set up This is a set of key inspection points involved in the manually annotated report. To generate a subset of key points already covered in the report, coverage is defined as: This metric measures the completeness of the generated report's coverage of key identification points that experts focus on (such as stroke trends, character structure, and overall proportions).

[0058] (b) Similarity of detailed descriptions

[0059] For each matched point in the generated report Extract their generation descriptions respectively Compared with manual annotation description Sentence vector representation and After calculating the cosine similarity, exponential sharpening is performed:

[0060] in:

[0061] The original cosine similarity in the high similarity interval ( The problem exhibits gradient flattening: although the absolute increment of similarity from 0.60 to 0.85 is the same (0.25 for both), the reward signal gain within this interval is weak for generated descriptions already at a high similarity level, making it difficult to provide effective gradient guidance. This problem is particularly prominent in open-ended text evaluation—the model tends to stagnate at local optima that are "generally correct but lack detail," because the reward incentive for further optimization from semantic similarity of 0.75 to 0.90 is much lower than the incentive intensity for increasing from 0.30 to 0.60. This invention introduces a 5th power exponential sharpening transformation to reshape the similarity incentive distribution from a uniform type to a head-sensitive type: After sharpening, the reward gradient in highly similar regions is significantly amplified, forcing the model to continuously pursue more precise detail descriptions after achieving basic semantic matching, thus avoiding getting trapped in local optima based on superficial semantic similarity. (Exponential frequency) It balances sharpening intensity and numerical stability, and can be regarded as a hyperparameter of the bonus signal resolution, which is adjusted according to the task description style and language characteristics.

[0062] (c) Format compliance

[0063] Suppose the report requires a total of structural elements (verification conclusions, key points list, item-by-item numbered explanations, etc.). item, Indicates the first If the term exists, then:

[0064] (d) Overall score The final score for the signature verification report is obtained by weighting three dimensions:

[0065]

[0066] Among them, weight , , It can be configured according to the application scenario, data language, and reporting specifications. In this embodiment, 1.

[0067] (e) Stability-aware progressive triggering mechanism Directly combining open-ended report quality rewards with rule-based rewards is a mismatch in training dynamics: rule-based rewards (format, accuracy) have deterministic, low-variance signal characteristics, while open-ended rewards rely on sentence vector similarity and structural matching. In the early stages of training, when the policy is still unstable, their calculation results are characterized by high noise and high variance. Mixing these two types of signals directly in the early stages of training causes the high-noise gradient of the report quality reward to cover the effective update direction of the low-variance rule-based reward. This leads to the policy being pulled onto an unstable report optimization path before format standardization and basic discriminative ability are stable, ultimately causing training instability phenomena such as format collapse or premature drop in accuracy.

[0068] To avoid fluctuations in signature verification report superiority rewards during the initial training phase interfering with strategy updates, this embodiment employs a trigger criterion based on historical windows to determine the activation timing of signature verification report superiority rewards. The specific criterion is defined as follows:

[0069] In the current training step Continuous with endpoint Within the step window, if the format of each step is rewarded... With accuracy bonus The sum of all of them exceeds the threshold. Then from the first Step 1 will incorporate the signature verification report quality reward into the overall reward participation strategy update; otherwise, the overall reward will consist only of format rewards and accuracy rewards. Parameters The stringency of the stability judgment should be controlled. Control the activation threshold; in this implementation, it is set during training. The core idea of ​​this criterion is to treat the stability of rule-based rewards as an externally observable indicator that the strategy has formed a stable internal representation. Based on this, introducing a high-noise reward signal can both preserve the effective guidance of the SIGNER on report quality and avoid training oscillations caused by early introduction. (Continuous window length) With threshold Together, they form a two-dimensional stability gating mechanism, providing an adjustable trade-off between training conservatism and convergence speed.

[0070] In one implementation, in step (7), the model inference process inputs two signature images and authentication instructions in the corresponding language. The model first... <think>The process of reasoning from local to global is generated in the middle, and then in <answer>The system outputs a report that can be displayed to users. Users can extract keywords such as "genuine / forged" or "genuine / forged" from the report as binary classification results, while retaining the report text for manual review, or using the merit reward of the signature verification report as an indicator to compare with the tagged report to calculate the quality of the generated report.

[0071] In summary, compared with the prior art, this application has at least the following advantages and beneficial effects: 1) This application enables interpretable offline signature authentication. The signature authentication system is no longer limited to a binary output of genuine / forged characters, but can explicitly generate a complete chain of reasoning processes, from observing local strokes, comparing character structures, analyzing global layouts to making a comprehensive judgment of authenticity. Based on this, it generates an evidence-based authentication report, significantly improving the credibility, verifiability, and business applicability of the authentication results.

[0072] 2) This application proposes a cold-start supervised training mechanism for interpretable signature authentication, which enables the model to learn professional signature comparison procedures, authentication terminology and formatted output capabilities under the constraints of chained reasoning and authentication reports. This capability ensures that the model can avoid format collapse during subsequent post-training.

[0073] 3) This application proposes a stability-aware reward model for open-source generation, including two original designs: First, the signature authentication report quality reward evaluates the generated report from three dimensions: coverage of key verification points, similarity of detailed descriptions, and format compliance. The detailed similarity is enhanced by exponential sharpening (replacing the original cosine similarity with the fifth power of cosine similarity) to improve the sensitivity of the reward signal to differences in description quality. Second, a progressive triggering criterion based on historical windows is adopted, only applying the reward when consecutive... The sum of the in-step format reward and accuracy reward exceeds the threshold for each step. When the signature verification report quality reward is added to the total reward, the training stability and open report quality optimization can be achieved without additional training of the reward model.

[0074] 4) This application proposes an unconstrained group strategy optimization, which has three differences compared to the existing GRPO algorithm: (a) based on the total number of tokens in the entire group. (a) It replaces the double averaging approach of GRPO with a unified normalized denominator, eliminating gradient bias introduced by uneven output length; (b) It removes the probability ratio pruning constraint, allowing full-amplitude gradient updates for high-dominance actions; (c) It removes the reference policy KL penalty term, releasing the policy's ability to explore high-quality inference paths beyond the distribution of the base model; and it adopts group relative advantage with within-group mean-variance normalization. This balances the ratio of positive to negative signals, maintaining the stability of training values ​​under unconstrained conditions.

[0075] 5) This application reveals and utilizes the inherent synergistic relationship between stability-aware reward modeling for open-ended generation and unconstrained group policy optimization: Unconstrained group policy optimization, by eliminating the conservative constraints of policy updates, expands the explorable inference path space to high-value sparse regions beyond the distribution of the base model; stability-aware reward modeling for open-ended generation, through asymptotic triggering criteria, ensures that the expanded exploration process is guided by open-ended rewards only after the regularity of the capability has stabilized, thereby avoiding training instability caused by the superposition of free exploration and high-noise rewards. The two approaches form a complementary tension balance in their design, jointly achieving the training objective of "sufficient exploration and stable convergence".

[0076] Experimental results show that the offline signature authentication model trained using the method provided in this application can achieve an accuracy rate of over 80% in identifying authenticity on Chinese signature datasets and over 90% in identifying authenticity on English signature datasets. Furthermore, the generated authentication report is significantly better than the baseline model without reinforcement learning optimization in terms of complete coverage of key points, accuracy of detailed description, and standardization of format.

[0077] This embodiment also provides an offline signature authentication method, the specific implementation of which is as follows: Obtain the reference signature image to be authenticated and the signature image to be inspected. The reference signature is a genuine signature known to belong to the claimant, while the signature to be inspected is the signature whose authenticity needs to be verified.

[0078] Following the above method, the two signature images were subjected to RGB original image extraction and grayscale normalization processing, respectively.

[0079] The processed image is then input into the offline signature authentication model trained using the method described above.

[0080] The model performs the following inference process: a) The visual Transformer encoder extracts global visual terms from the two images; b) An expert visual encoder extracts fine-grained stroke feature words from the two images; c) The two types of lexical units are concatenated after being mapped by MLP and then input into the large language model; d) Large language models combine visual context and textual instructions for autoregressive generation. <think>The chain-like reasoning process within the tag; e) Based on evidence accumulated during the reasoning process, large language models are generated through autoregression. <answer>The identification report inside the label.

[0081] Output from the model <answer>Within the tag authentication report, keywords such as "authentic" / "fake" or "genuine" / "forged" can be extracted as binary classification authentication results. Simultaneously, the complete authentication report text can be output to users for manual review, audit traceability, or archiving.

[0082] Different verification effects and analysis results can be obtained by adjusting the inference parameters of the model for different signature inputs.

[0083] This application also provides an electronic device, which includes a memory and a processor. The memory stores a computer program, and the processor executes the computer program to implement the method described above. This electronic device can be any smart terminal, including tablet computers, in-vehicle computers, etc.

[0084] It is understood that the content of the above method embodiments is applicable to this device embodiment. The specific functions implemented by this device embodiment are the same as those of the above method embodiments, and the beneficial effects achieved are also the same as those achieved by the above method embodiments.

[0085] Please see Figure 3 , Figure 3 The hardware structure of an electronic device according to another embodiment is illustrated. The electronic device includes: The processor 301 can be implemented using a general-purpose CPU (Central Processing Unit), microprocessor, application-specific integrated circuit (ASIC), or one or more integrated circuits, and is used to execute relevant programs to implement the technical solutions provided in the embodiments of this application. The memory 302 can be implemented as a read-only memory (ROM), static storage device, dynamic storage device, or random access memory (RAM). The memory 302 can store the operating system and other applications. When the technical solutions provided in the embodiments of this specification are implemented through software or firmware, the relevant program code is stored in the memory 302 and is called and executed by the processor 301 using the methods described above in the embodiments of this application. Input / output interface 303 is used to implement information input and output; The communication interface 304 is used to enable communication and interaction between this device and other devices. Communication can be achieved through wired means (such as USB, network cable, etc.) or wireless means (such as mobile network, WIFI, Bluetooth, etc.). Bus 305 transmits information between various components of the device (e.g., processor 301, memory 302, input / output interface 303, and communication interface 304); The processor 301, memory 302, input / output interface 303, and communication interface 304 are connected to each other within the device via bus 305.

[0086] This application also provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the above-described method.

[0087] It is understood that the content of the above method embodiments is applicable to this storage medium embodiment. The specific functions implemented in this storage medium embodiment are the same as those in the above method embodiments, and the beneficial effects achieved are also the same as those achieved in the above method embodiments.

[0088] Memory, as a non-transitory computer-readable storage medium, can be used to store non-transitory software programs and non-transitory computer-executable programs. Furthermore, memory may include high-speed random access memory, and may also include non-transitory memory, such as at least one disk storage device, flash memory device, or other non-transitory solid-state storage device. In some embodiments, memory may optionally include memory remotely located relative to the processor, and these remote memories can be connected to the processor via a network. Examples of such networks include, but are not limited to, the Internet, intranets, local area networks, mobile communication networks, and combinations thereof.

[0089] This application also provides a computer program product, including a computer program that, when executed by a processor, implements the above-described method.

[0090] It is understood that the content of the above method embodiments is applicable to the embodiments of this program product. The specific functions implemented by the embodiments of this program product are the same as those of the above method embodiments, and the beneficial effects achieved are also the same as those achieved by the above method embodiments. The executable computer program code or "code" used to perform the various embodiments can be written in high-level programming languages ​​such as C, C++, Python, Smalltalk, Java, JavaScript, Visual Basic, Structured Query Language (e.g., Transact-SQL), Perl, or in various other programming languages.

[0091] The embodiments described in this application are for the purpose of more clearly illustrating the technical solutions of the embodiments of this application, and do not constitute a limitation on the technical solutions provided by the embodiments of this application. As those skilled in the art will know, with the evolution of technology and the emergence of new application scenarios, the technical solutions provided by the embodiments of this application are also applicable to similar technical problems.

[0092] Those skilled in the art will understand that the technical solutions shown in the figures do not constitute a limitation on the embodiments of this application, and may include more or fewer steps than shown, or combine certain steps, or different steps.

[0093] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs.

[0094] Those skilled in the art will understand that all or some of the steps in the methods disclosed above, as well as the functional modules / units in the systems and devices, can be implemented as software, firmware, hardware, or suitable combinations thereof.

[0095] The terms "first," "second," "third," "fourth," etc. (if present) in the specification and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this application described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.

[0096] It should be understood that in this application, "at least one (item)" means one or more, and "more than" means two or more. "And / or" is used to describe the relationship between related objects, indicating that three relationships can exist. For example, "A and / or B" can represent: only A exists, only B exists, and both A and B exist simultaneously, where A and B can be singular or plural. The character " / " generally indicates that the preceding and following related objects are in an "or" relationship. "At least one (item) of the following" or similar expressions refer to any combination of these items, including any combination of single or plural items. For example, at least one (item) of a, b, or c can represent: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, and c can be single or multiple.

[0097] In the several embodiments provided in this application, it should be understood that the disclosed apparatus and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of the units described above is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between apparatuses or units may be electrical, mechanical, or other forms.

[0098] The units described above as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.

[0099] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.

[0100] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes multiple instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods of the various embodiments of this application. The aforementioned storage medium includes various media capable of storing programs, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0101] The preferred embodiments of the present application have been described above with reference to the accompanying drawings, but this does not limit the scope of the claims of the present application. Any modifications, equivalent substitutions, and improvements made by those skilled in the art without departing from the scope and substance of the embodiments of the present application shall be within the scope of the claims of the present application.< / answer> < / answer> < / think> < / answer> < / think>

Claims

1. A training method for an offline signature authentication model, characterized in that, The method includes the following steps: Obtain a training dataset, wherein each sample in the training dataset contains a reference signature image, a signature image to be tested, an authentication instruction, a chained reasoning process, and an authentication report containing the authenticity verification conclusion; The reference signature image and the signature image to be inspected are preprocessed to generate RGB original image data and grayscale normalized data respectively; A basic model is constructed, which includes a visual Transformer encoder, an expert visual encoder, a multilayer perceptron mapping layer, and a large language model. The original RGB image data is input into the visual Transformer encoder to extract global visual terms, and the grayscale normalized data is input into the expert visual encoder to extract expert visual terms. The global visual terms and expert visual terms are mapped to the hidden dimensions of the large language model through the multilayer perceptron mapping layer and then concatenated and input into the large language model. The training dataset is used to perform cold-start supervised fine-tuning of the base model, so that the model outputs a chain inference process and identification report in a preset format; The model after cold start is used as the initial policy model, and reinforcement learning training is carried out using unconstrained group policy optimization. For each input cue, multiple candidate outputs are sampled and generated. The reward value of each candidate output is calculated separately. The advantage value is calculated using within-group mean-variance standardization. The total number of tokens of all candidate outputs is uniformly normalized. The policy is updated with the goal of maximizing the sum of the products of the normalized policy ratio and the advantage value. The reward value includes a format reward, an accuracy reward, and a signature verification report quality reward. The signature verification report quality reward is calculated based on the coverage of key verification points, the similarity of detailed descriptions, and the format compliance between the generated report and the labeled report. The signature verification report quality reward is added to the total reward and participates in the strategy update after the triggering criteria are met.

2. The method according to claim 1, characterized in that, The objective function for optimizing the unconstrained group strategy is: in, The number of sampled candidate outputs for each input prompt. For the first 1 candidate output, This indicates the number of tokens in the candidate output. and The current strategy and the old strategy are respectively in the th... The output of the first The probability at each token This is the dominant value; This indicates an input cue sampled from the training data distribution. and from the old strategy model The sampled input prompts are in the middle. candidate outputs Find the expected value.

3. The method according to claim 1, characterized in that, The advantage value is calculated as follows: each token of the G candidate outputs in the same group shares the same advantage value, which is equal to the reward value of the candidate output after standardization with respect to the mean and standard deviation of the G reward values ​​in the group.

4. The method according to claim 1, characterized in that, The similarity of the detailed descriptions is calculated in the following way: For each key inspection point covered in the generated report, extract the sentence vector representations of the generated description and the labeled description respectively, and calculate the cosine similarity. The cosine similarity is subjected to exponential sharpening, and the p-th power of the cosine similarity is used as the sharpened similarity, where p>1; The average of the sharpened similarities of each key point is used as the detail description similarity.

5. The method according to claim 1, characterized in that, The triggering criterion is as follows: within a continuous n-step window ending at the current training step k, if the sum of the format reward and accuracy reward for each step exceeds the threshold t, then starting from the (k+1)th step, the signature authentication report excellence reward will be included in the total reward; otherwise, the total reward consists only of the format reward and the accuracy reward.

6. The method according to claim 1, characterized in that, The weighted calculation formula for the signature authentication report excellence reward is as follows: in, To ensure coverage of key inspection points, To describe the similarity in detail, For format compliance; , , As weight, and ; The coverage of key inspection points ,in This is a collection of key inspection points in the annotation report. To generate a subset of key points already covered in the report; The format compliance ,in To report the total number of structural elements required, Indicates the first Does the item structure element exist? 7. The method according to claim 1, characterized in that, In the cold start supervised fine-tuning, the chain reasoning process is placed in... <think> and< / think> The identification report is placed between the labels. <answer> and< / answer> Between tags; the chain reasoning process is organized in the order of observing local strokes, comparing character structures, analyzing global layout, and making a comprehensive judgment on authenticity.

8. An offline signature authentication method, characterized in that, Includes the following steps: Obtain the reference signature image to be authenticated and the signature image to be inspected; The reference signature image and the signature image to be inspected are input into the offline signature authentication model trained using the training method described in any one of claims 1 to 7; Obtain the chain-like reasoning process and authentication report output by the model. The authentication report includes a judgment on the authenticity of the signature to be examined relative to a reference signature and an interpretability analysis based on visual evidence.

9. An electronic device, characterized in that, The electronic device includes a memory and a processor, the memory storing a computer program, and the processor executing the computer program to implement the method according to any one of claims 1 to 8.

10. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by a processor, it implements the method of any one of claims 1 to 8.

Citation Information

Patent Citations

  • Method for realizing interpretable offline signature authentication based on multi-modal large model

    CN120510619A