Method and device for intelligent recognition of receipt signature position and angle based on deep learning
Through deep learning methods combined with data augmentation and multimodal fusion technology, the accuracy and efficiency of signature position and angle recognition in return processing in logistics and finance fields is solved, and high-precision and real-time automated processing is achieved.
Patent Information
- Application Number
- CN202510921079.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-04
- Publication Date
- 2025-08-29
- Estimated Expiration
- 2045-07-04
AI Technical Summary
In the prior art, in the logistics and financial fields, the identification of signature positions and angles has problems such as low efficiency, weak anti-interference ability, insufficient angle detection accuracy and insufficient multimodal fusion, which is difficult to meet the needs of high-concurrency scenarios.
A deep learning-based method is adopted to generate diversified training samples through data augmentation techniques (such as brightness transformation, color gamut transformation and MixCut image mixing), combining light-weight diffusion models and low-rank decomposition techniques to deal with noise and low-light problems; using hierarchical angle detection (coarse classification and fine adjustment) and multi-scale sliding window combined with template matching algorithm, combined with OCR text extraction and NLP semantic analysis, to achieve accurate positioning of signature areas.
The recognition accuracy of signature position and angle is improved to 98.7%, significantly improving the noise and complex background robustness of the model, and supporting high concurrency scenarios through model acceleration technology, reducing customized development costs.
Smart Images

Figure CN120411978B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of image processing technology, and in particular to a method and device for intelligently identifying the position and angle of a receipt signature based on deep learning. Background Art
[0002] In receipt processing in logistics, finance, and other fields, accurate identification of signature position and angle is crucial for automated processing. Traditional methods rely on manual visual judgment or simple rule matching, resulting in low efficiency (processing time for a single receipt exceeds 300ms) and weak anti-interference capabilities (accuracy below 85% under complex lighting or noise). With the development of deep learning technology, existing solutions attempt to use CNNs for image classification and object detection, but they still suffer from the following shortcomings:
[0003] Insufficient angle detection accuracy: Existing methods can only achieve coarse classification of fixed angles (0° / 90° / 180° / 270°) and lack the ability to correct subtle angle deviations within the ±10° range. This can cause positioning offsets in the signature area due to image tilt.
[0004] Insufficient multimodal fusion: Signature location relies solely on image visual features, failing to effectively incorporate textual semantic information such as "shipper's signature," leading to missed detections in complex backgrounds or with changing layouts.
[0005] The contradiction between data enhancement and model efficiency: Traditional data enhancement techniques (such as random rotation and cropping) have limited improvement on problems such as uneven lighting and low resolution. While directly using large models can improve accuracy, the inference speed is slow (single-image processing exceeds 200ms), making it difficult to meet the needs of high-concurrency scenarios.
[0006] Therefore, a method is urgently needed to solve at least one of the above problems. Summary of the Invention
[0007] This application provides a method and device for intelligently identifying the position and angle of receipt signatures based on deep learning, which is used to fill the gap in the existing technology in high-precision and real-time receipt processing.
[0008] In a first aspect, the present application provides a method for intelligently identifying the position and angle of a receipt signature based on deep learning, the method comprising:
[0009] Obtain an original receipt image and perform data enhancement processing on the original receipt image. The data enhancement processing includes brightness conversion, color gamut conversion, and image blending to obtain a sample receipt image, wherein the image blending is performed by randomly blending two images corresponding to the receipt image data and annotating the label ratio;
[0010] Input the sample return image into a preset deep learning classification model for coarse classification, perform fine angle adjustment within a range of ±10° on the classification result corresponding to the coarse classification, output the precise angle correction parameter corresponding to the sample return image, and complete the angle correction of the sample return image;
[0011] Perform text extraction on the angle-corrected sample receipt image, match the text extraction result with the preset signature area keywords, achieve coarse positioning of the signature area in the sample receipt image, and output the coarse positioning area;
[0012] A multi-scale sliding window combined with a template matching algorithm is used to search the coarse positioning area to generate a candidate signature area frame. The overlapping candidate frames corresponding to multiple candidate signature area frames are merged through non-maximum suppression, and the coordinates of the candidate signature area frames are optimized using boundary regression. The signature position coordinates and waybill angle corresponding to the sample receipt image are output.
[0013] In some embodiments, before performing data enhancement processing on the original return single image, it also includes: performing image denoising on the original return single image using a median filtering algorithm to remove white noise and salt and pepper noise; for original return single images whose clarity or lighting does not meet preset clarity conditions or preset lighting conditions, using a lightweight diffusion model combined with low-rank decomposition technology to perform image restoration.
[0014] In some embodiments, the sample receipt image is input into a preset deep learning classification model for coarse classification, the classification result corresponding to the coarse classification is fine-tuned within the range of ±10°, and the precise angle correction parameters corresponding to the sample receipt image are output, including: screening the main area of the sample receipt image according to the edge detection algorithm to obtain the main contour area of the waybill; inputting the main contour area into the preset deep learning classification model for coarse classification to obtain four classification angle labels of 0°, 90°, 180°, and 270°; based on the coarse classification result, the text line direction of the sample receipt image is detected, and the text box tilt angle parameter output by the network is used to perform angle offset correction within the range of ±10° on the basis of the coarse classification angle to generate precise angle correction parameters including a rotation matrix.
[0015] In some embodiments, the text extraction result is matched with the preset signature area keywords to achieve coarse positioning of the signature area in the sample receipt image and output the coarse positioning area, including: performing text extraction and coordinate positioning on the sample receipt image after angle correction to generate a text box set containing text content and position coordinates; performing semantic analysis on the text content in the text box set through NLP technology to match the preset signature area keywords, and generating an initial signature area candidate box as the coarse positioning area based on the coordinates of the successfully matched text box and the preset keyword-associated area offset.
[0016] In some embodiments, the use of a multi-scale sliding window combined with a template matching algorithm to search the coarse positioning area and generate a candidate signature area frame includes: traversing the coarse positioning area using a multi-scale sliding window, with the window scale covering a preset common signature size range and the step size set to 1 / 4 of the window size; within each sliding window, using a preset signature area template for template matching and calculating a matching similarity score; when the similarity score exceeds a preset threshold, generating a candidate signature area frame using the current window coordinates; introducing a spatial attention mechanism to weight the features of the sliding window area, enhancing the response to key features such as signature handwriting edges and textures, and suppressing background noise interference.
[0017] In some embodiments, the common size range of the preset signature includes 10×10 pixels to 200×200 pixels.
[0018] In some embodiments, the merging of overlapping candidate frames corresponding to multiple candidate signature region frames by non-maximum suppression includes: calculating the intersection-over-union ratio between all candidate signature region frames; for candidate signature region frames whose intersection-over-union ratio is higher than a preset threshold, retaining the candidate signature region frame with the highest matching similarity score, and deleting the remaining candidate signature region frames as overlapping candidate frames; repeating the above process until the intersection-over-union ratios between all candidate signature region frames are lower than the preset threshold, thereby completing the deletion of overlapping candidate frames.
[0019] In some embodiments, the use of boundary regression to optimize the coordinates of the candidate signature area box includes: inputting the candidate signature area box into a preset boundary regression model, and based on a preset anchor frame mechanism, performing regression adjustment on the coordinates of the candidate signature area box to optimize the position and size of the candidate signature area box.
[0020] In some embodiments, the outputting of the signature position coordinates and waybill angle corresponding to the sample receipt image includes: fusing the image visual features and text semantic features corresponding to the candidate signature area box, jointly optimizing the boundary regression parameters through multimodal alignment technology, and outputting the signature position coordinates and waybill angle detection results including the upper left corner coordinates and the lower right corner coordinates.
[0021] In a second aspect, the present application provides a device for intelligently identifying the position and angle of a receipt signature, which is applied to a computer device, and the device includes:
[0022] An image acquisition unit is configured to acquire an original receipt image and perform data enhancement processing on the original receipt image, wherein the data enhancement processing includes brightness conversion, color gamut conversion, and image blending to obtain a sample receipt image, wherein the image blending is performed by randomly blending two images corresponding to the receipt image data and annotating the label ratio;
[0023] An image input unit is used to input the sample return image into a preset deep learning classification model for coarse classification, perform fine angle adjustment within a range of ±10° on the classification result corresponding to the coarse classification, output precise angle correction parameters corresponding to the sample return image, and complete angle correction of the sample return image;
[0024] a text extraction unit, configured to extract text from the angle-corrected sample receipt image, match the text extraction result with a preset signature area keyword, achieve coarse positioning of the signature area in the sample receipt image, and output the coarse positioning area;
[0025] The angle output unit is used to search the coarse positioning area using a multi-scale sliding window combined with a template matching algorithm to generate a candidate signature area frame, merge the overlapping candidate frames corresponding to multiple candidate signature area frames through non-maximum suppression, and optimize the coordinates of the candidate signature area frames using boundary regression to output the signature position coordinates and waybill angle corresponding to the sample receipt image.
[0026] The present application discloses a method and device for intelligently identifying the position and angle of a waybill signature based on deep learning. The provided method generates diversified training samples through brightness transformation, color gamut transformation, and MixCut image mixing technology (randomly mixing two images and annotating the label ratio), thereby improving the model's robustness against illumination changes and complex backgrounds. The deep learning model is first used to roughly classify the waybill angle (0° / 90° / 180° / 270°), and then the angle is fine-tuned within the range of ±10° through the Dbnet network to output precise correction parameters. Based on the OCR text extraction results, keywords such as "consignor's signature" are matched through NLP, and the coarse positioning area of the signature is determined in combination with the offset of the keyword-associated area. A multi-scale sliding window is used in combination with template matching to generate candidate boxes, which are deduplicated through non-maximum suppression (NMS). Finally, boundary regression is used to optimize the coordinates and output the signature position and angle results.
[0027] Through the "coarse classification + fine adjustment" angle detection mechanism, the angle error is controlled within ±10°. Combined with semantic guidance and multimodal fusion, the signature position positioning accuracy is ≥98.7%, significantly better than the traditional single visual feature model; data enhancement technology (MixCut, lightweight diffusion model repair) effectively deals with low-quality images such as noise, low light, and blur, and the model generalization ability is improved; model quantization and CUDA acceleration technology compress the processing time of a single image, support high-concurrency scenarios, and improve the inference speed compared to traditional deep learning solutions; the multimodal fusion architecture supports different formats of receipts (electronic / paper, multi-language), and can flexibly adapt to the personalized needs of enterprises through keyword configuration, reducing customized development costs.
[0028] In summary, this application systematically solves the shortcomings of existing methods in accuracy, efficiency, and robustness through innovative designs such as hierarchical detection, semantic guidance, and multimodal fusion, providing a breakthrough technical solution for the automated processing of receipts. BRIEF DESCRIPTION OF THE DRAWINGS
[0029] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the following is a brief introduction to the drawings required for use in the description of the embodiments. Obviously, the drawings described below are some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.
[0030] Figure 1 This is a flowchart illustrating the steps of a method for intelligently identifying receipt signature position and angle based on deep learning provided by an embodiment of the present application;
[0031] Figure 2 This is a schematic block diagram of a device for intelligently identifying the position and angle of a receipt signature provided in an embodiment of the present application;
[0032] Figure 3 This is a schematic block diagram of the structure of a computer device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0033] The following will be combined with the drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are part of the embodiments of this application, not all of them. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.
[0034] The flowcharts shown in the accompanying drawings are for illustrative purposes only and do not necessarily include all contents and operations / steps, nor must they be executed in the order described. For example, some operations / steps may be decomposed, combined, or partially merged, so the actual execution order may vary depending on the actual situation.
[0035] It should be understood that the terms used in this specification are only for the purpose of describing specific embodiments and are not intended to limit the present application. As used in this specification and the appended claims, the singular forms "a", "an", and "the" are intended to include the plural forms unless the context clearly indicates otherwise.
[0036] It will also be understood that the term "and / or" as used in this specification and the appended claims refers to and includes any and all possible combinations of one or more of the associated listed items.
[0037] The following describes some embodiments of the present application in detail with reference to the accompanying drawings. In the absence of conflict, the following embodiments and features therein may be combined with each other.
[0038] It should be noted that any data involved in this application is obtained with the permission of the relevant users, complies with relevant policy regulations, and will not infringe on user privacy.
[0039] In receipt processing in logistics, finance, and other fields, accurate identification of signature position and angle is crucial for automated processing. Traditional methods rely on manual visual judgment or simple rule matching, resulting in low efficiency (processing time for a single receipt exceeds 300ms) and weak anti-interference capabilities (accuracy below 85% under complex lighting or noise). With the development of deep learning technology, existing solutions attempt to use CNNs for image classification and object detection, but they still suffer from the following shortcomings:
[0040] Insufficient angle detection accuracy: Existing methods can only achieve coarse classification of fixed angles (0° / 90° / 180° / 270°) and lack the ability to correct subtle angle deviations within the ±10° range. This can cause positioning offsets in the signature area due to image tilt.
[0041] Insufficient multimodal fusion: Signature location relies solely on image visual features, failing to effectively incorporate textual semantic information such as "shipper's signature," leading to missed detections in complex backgrounds or with changing layouts.
[0042] The contradiction between data enhancement and model efficiency: Traditional data enhancement techniques (such as random rotation and cropping) have limited improvement on problems such as uneven lighting and low resolution. While directly using large models can improve accuracy, the inference speed is slow (single-image processing exceeds 200ms), making it difficult to meet the needs of high-concurrency scenarios.
[0043] Therefore, a method is urgently needed to solve at least one of the above problems.
[0044] To address the above problems, the present invention proposes an intelligent recognition method that integrates data enhancement, multimodal features, and model acceleration. Through an innovative angle fine-tuning mechanism, semantically guided coarse positioning, and multi-scale search strategy, it significantly improves the recognition accuracy and processing efficiency in complex scenarios, filling the gap in existing technologies in high-precision, real-time receipt processing.
[0045] See also Figure 1 , Figure 1 This is a schematic flow chart of a method for intelligently identifying receipt signature position and angle based on deep learning, provided in an embodiment of the present application. The method is applied to a computer device.
[0046] like Figure 1 As shown, the specific steps of the receipt signature position and angle intelligent recognition method include: step S101 to step S104.
[0047] S101. Obtain an original receipt image and perform data enhancement processing on the original receipt image. The data enhancement processing includes brightness transformation, color gamut transformation, and image blending to obtain a sample receipt image, wherein the image blending is performed by randomly blending two images corresponding to the receipt image data and marking the label ratio.
[0048] Specifically, the original single-reply images are repaired for noise, enhanced for quality, and expanded for samples in sequence, and multi-dimensional data processing is used to improve image quality and training sample diversity, which specifically includes three steps: denoising, repair, and data enhancement.
[0049] Image denoising first uses the median filtering algorithm to perform global denoising on the original image. The median of the neighborhood pixels is calculated through a sliding window to effectively filter out white noise and salt and pepper noise, improving the image appearance (for example, background noise on waybills is reduced by more than 80%).
[0050] Low-quality image restoration targets images with insufficient clarity or abnormal lighting (such as excessive darkness or reflections). This restoration uses a lightweight diffusion model combined with low-rank decomposition technology. This model gradually restores image details through a diffusion process, while low-rank decomposition reduces the model's parameter size (by 60% compared to traditional diffusion models). This maintains the restoration effect while keeping the processing time per image under 50ms.
[0051] Data augmentation includes: Brightness Transformation: By adjusting the pixel brightness value (±20% dynamic range), image readability in low-light or bright light scenes is enhanced; Color Gamut Transformation: Converting the image from RGB space to HSV / Lab space, separating hue, brightness, and saturation, facilitating subsequent feature extraction and analysis in the color space; MixCut Image Mixing: Randomly select two single images, mix the pixels at a ratio of 0.3-0.7, and generate new samples based on the label ratio (such as weighted average of signature area coordinates) to enhance the model's generalization ability for complex layouts.
[0052] The combination of median filtering and a lightweight diffusion model improves the image signal-to-noise ratio by 40%, effectively addressing common problems in logistics receipts such as scanning blur and printing defects; MixCut technology generates cross-format mixed samples to avoid model overfitting and improves the diversity index of the measured training set; low-rank decomposition technology speeds up the inference of the repair model to meet the needs of large-scale data preprocessing.
[0053] S102: Input the sample return single image into a preset deep learning classification model for coarse classification, perform fine adjustment on the classification result corresponding to the coarse classification within the range of ±10°, output the precise angle correction parameters corresponding to the sample return single image, and complete the angle correction of the sample return single image.
[0054] Specifically, through the two-level angle detection mechanism of "coarse classification + fine adjustment", the overall direction of the waybill is first determined (0° / 90° / 180° / 270°), and then the slight tilt angle (±10°) is accurately corrected to solve the positioning offset problem caused by angle deviation in traditional methods.
[0055] The main area of the image is filtered using the Canny / Sobel edge detection algorithm to extract the outline of the waybill, filtering out irrelevant background (such as the blank area around the express delivery label), and retaining the main area of the waybill as input. Coarse classification positioning inputs the main area into a lightweight ResNet classification model, outputting four-category angle labels (0° / 90° / 180° / 270°). Angle fine-tuning uses the OCR detection network Dbnet based on the coarse classification results to detect the direction of the text line. By analyzing the tilt angle of the text box (e.g., outputting θ∈[-10°, 10°]), an offset is added to the coarse classification angle (e.g., coarse classification 90° + fine adjustment + 5°, resulting in a final correction of 95°), generating precise correction parameters including a rotation matrix.
[0056] Coarse classification quickly determines the main direction, and fine adjustment corrects slight tilt. Compared with the traditional single classification model, the angle detection error is reduced to, ensuring that the coordinate offset caused by tilt in the signature area is greatly reduced.
[0057] S103 , performing text extraction on the angle-corrected sample receipt image, matching the text extraction result with preset signature area keywords, achieving coarse positioning of the signature area in the sample receipt image, and outputting the coarse positioning area.
[0058] Specifically, by combining OCR text extraction and NLP semantic analysis, the signature area is located by associating keywords such as "consignor's signature", which makes up for the problem of missed detection of simple visual features in complex layouts and realizes the guidance of semantic information on visual positioning.
[0059] Text extraction and coordinate positioning use an OCR model based on CNN+RNN+CTC (such as CRNN) to perform text detection and recognition on the rectified image, generating a set of text boxes containing text content and location coordinates (x1, y1, x2, y2). The text detection speed for a single image is ≤ 20ms.
[0060] Semantic keyword matching uses NLP technology to build a keyword library (including "shipper's signature", "consignee's signature", "customer's seal", etc.), performs semantic matching on the text box content, and locates the text box coordinates corresponding to the keyword.
[0061] The coarse positioning area is generated based on the position of the keyword text box and the associated area offset preset by the business rules (such as within the range of 50-200 pixels below the keyword) to generate the initial signature area candidate box (such as expanding the keyword text box height to 1.5 times as the search range).
[0062] After introducing text semantic information, the missed detection rate is reduced in scenarios with changing layouts (such as the signature area without a fixed border and the keyword position being random), which solves the defect of the simple visual model relying on fixed patterns. By narrowing the search range by keywords, the computational complexity of subsequent candidate area searches is reduced, focusing on the effective area.
[0063] S104: Use a multi-scale sliding window combined with a template matching algorithm to search the coarse positioning area to generate a candidate signature area frame. Use non-maximum suppression to merge overlapping candidate frames corresponding to multiple candidate signature area frames, and use boundary regression to optimize the coordinates of the candidate signature area frames. Output the signature position coordinates and waybill angle corresponding to the sample receipt image.
[0064] Specifically, a multi-scale sliding window combined with template matching is used to generate candidate signature areas. After deduplication by non-maximum suppression (NMS), the coordinates are optimized using boundary regression, and finally high-precision signature position and angle results are output.
[0065] The multi-scale sliding window search uses a multi-scale window (10×10 to 200×200 pixels, covering common signature sizes) to traverse the coarse positioning area, with a step size set to 1 / 4 of the window size to ensure no overlap or omissions; the matching score is calculated within the window using a preset template (including prior features such as handwritten stroke density and blank area ratio), and a candidate box is generated when the score is greater than 0.75.
[0066] Spatial attention enhancement introduces a spatial attention mechanism (such as SE-Net channel weighting) during the sliding window process, assigning higher weights to key features such as signature edges and textures, suppressing interference information such as table lines and seals, and improving the accuracy of candidate boxes.
[0067] Candidate box optimization includes: NMS deduplication: calculating the intersection over union (IoU) of candidate boxes (merging when IoU>0.5), retaining high-scoring boxes, and reducing redundant candidates (the average number of candidate boxes in a single image is reduced from 80 to 15); boundary regression: based on the RPN mechanism of Faster R-CNN, regressing and adjusting the coordinates (x, y, w, h) of the candidate boxes, fusing image visual features (edge distribution) with text semantic features (keyword position) for joint optimization, and outputting the final signature position rectangle (upper left corner / lower right corner coordinates) and angle detection results.
[0068] It covers signatures of different sizes (from small characters to large handwritten text), and its positioning recall rate is improved compared to a fixed-size window. After combined multimodal feature regression, the coordinate error of the signature area meets the pixel-level accuracy requirements for automated processing of electronic receipts. The combination of NMS and the attention mechanism maintains accuracy while controlling the single-image positioning time to within 80ms, supporting high-concurrency real-time processing.
[0069] In summary, through four levels of innovation (data enhancement and anti-interference, hierarchical angle detection, semantic-guided positioning, and multimodal optimization), this method systematically solves the pain points of traditional solutions in terms of accuracy, efficiency, and robustness, and provides key technical support for the engineering implementation of automated processing of logistics and financial receipts.
[0070] In some embodiments, before performing data enhancement processing on the original return single image, it also includes: performing image denoising on the original return single image using a median filtering algorithm to remove white noise and salt and pepper noise; for original return single images whose clarity or lighting does not meet preset clarity conditions or preset lighting conditions, using a lightweight diffusion model combined with low-rank decomposition technology to perform image restoration.
[0071] Median filtering denoising constructs a 3×3 or 5×5 sliding window on the original receipt image, traverses each pixel, sorts the grayscale values of all pixels within the window, takes the median, and replaces the current pixel value. This operation effectively removes discrete white noise (pixel value abrupt changes) and salt and pepper noise (black and white spots) through nonlinear filtering properties. It is particularly effective for scanning noise commonly found in receipts, such as ink spots on express delivery slips.
[0072] The processing flow includes inputting the original image → grayscale conversion (if it is a color image) → calculating the median value in the sliding window → outputting the denoised image. The processing time for a single image is ≤10ms (based on a 512×512 resolution image).
[0073] The lightweight diffusion model combined with low-rank decomposition restoration includes: Image quality assessment: By calculating image clarity indicators (such as gradient amplitude mean) and lighting uniformity (such as brightness standard deviation), low-quality images with clarity less than a threshold (such as gradient mean <50) or uneven lighting (brightness standard deviation >80) are screened out.
[0074] The lightweight diffusion model uses a U-Net architecture as its backbone network. During the diffusion process, only critical diffusion steps (e.g., 10) are retained. Low-rank decomposition techniques are used to reduce the dimensionality of the convolutional layer weight matrix (from 1024 to 256), reducing the number of parameters by 70% while maintaining the restoration effect. The model takes a low-quality image as input and outputs a restored image with high contrast and clear details (e.g., recognizing the strokes of a blurred signature). Histogram equalization is performed on the restored image to further improve lighting uniformity.
[0075] In some embodiments, the sample receipt image is input into a preset deep learning classification model for coarse classification, the classification result corresponding to the coarse classification is fine-tuned within the range of ±10°, and the precise angle correction parameters corresponding to the sample receipt image are output, including: screening the main area of the sample receipt image according to the edge detection algorithm to obtain the main contour area of the waybill; inputting the main contour area into the preset deep learning classification model for coarse classification to obtain four classification angle labels of 0°, 90°, 180°, and 270°; based on the coarse classification result, the text line direction of the sample receipt image is detected, and the text box tilt angle parameter output by the network is used to perform angle offset correction within the range of ±10° on the basis of the coarse classification angle to generate precise angle correction parameters including a rotation matrix.
[0076] The main area is screened using the Canny edge detection algorithm (dual thresholds of 100 / 200) to extract image edges. Contour detection is then used to identify the largest closed contour (the main body of the delivery order). The smallest rectangular area containing this contour is then cropped, filtering out irrelevant background (such as blank areas around the delivery order or traces of pasting). Morphological dilation (3×3 kernel) is performed on the contour area after edge detection to ensure the integrity of the delivery order boundary and avoid loss of the main area due to broken edges.
[0077] The coarse classification model uses a lightweight ResNet-18 classification model. The input size is adjusted to 224×224, and the output layer is configured with 4 neurons (corresponding to 0°, 90°, 180°, and 270°). Training is performed using the cross-entropy loss function, achieving a four-category classification accuracy of ≥99.2% on the validation set. The coarse classification results are used to determine the overall orientation of the waybill. For example, when outputting a 90° label, the image is initially determined to require a 90° clockwise rotation to correct for the forward orientation.
[0078] Angle fine-tuning (DBNet detection) uses the DBNet text detection network to detect the text line orientation in the coarsely classified image and output the tilt angle θ (range [-10°, 10°]) for each text box. Precise angle calculation: If the coarse classification angle is α (0° / 90° / 180° / 270°), the final correction angle is α + θ. The corresponding rotation matrix is generated (for example, if θ = +5°, the rotation matrix is [[cos5°, -sin5°], [sin5°, cos5°]]), and affine transformation correction is performed on the image.
[0079] The traditional single model can only output fixed four-category angles. This embodiment reduces the angle error through "coarse classification to determine the main direction + fine adjustment to correct slight tilt", solving the problem of signature area misalignment caused by angle deviation in traditional methods.
[0080] In some embodiments, the text extraction result is matched with the preset signature area keywords to achieve coarse positioning of the signature area in the sample receipt image and output the coarse positioning area, including: performing text extraction and coordinate positioning on the sample receipt image after angle correction to generate a text box set containing text content and position coordinates; performing semantic analysis on the text content in the text box set through NLP technology to match the preset signature area keywords, and generating an initial signature area candidate box as the coarse positioning area based on the coordinates of the successfully matched text box and the preset keyword-associated area offset.
[0081] OCR text extraction and coordinate positioning use a CRNN (CNN+RNN+CTC) model for text detection and recognition. The CNN uses ResNet-34 to extract features, the RNN uses a bidirectional LSTM to process sequences, and the CTC loss function solves the text label alignment problem. The output is a collection of text boxes, each containing content (such as "Shipper's Signature"), coordinates (upper left corner (x1, y1), lower right corner (x2, y2)), and confidence (retained if > 0.9). The average number of text boxes detected per image is 50-100, with a detection speed of ≤ 20ms.
[0082] The preset keyword library includes more than 10 business-related terms, including "shipper's signature," "consignee's signature," "customer's seal," and "signature area," and supports user-defined extensions. Using NLP technology, the text box content is precisely matched (completely matched) and fuzzy matched (including key words, such as "signature"). After a successful match, a coarse positioning area is generated based on the preset offset rules:
[0083] The vertical offset is 50-200 pixels below the keyword text box (to adapt to the layout where the signature is usually located below the keyword); the horizontal offset is 20% of the width of the keyword text box to the left and right (to cover the possible signature range); the area size is a fixed height of 100 pixels (empirical value, covering the common signature height), and the width is the same as the expanded width.
[0084] In some embodiments, the use of a multi-scale sliding window combined with a template matching algorithm to search the coarse positioning area and generate a candidate signature area frame includes: traversing the coarse positioning area using a multi-scale sliding window, with the window scale covering a preset common signature size range and the step size set to 1 / 4 of the window size; within each sliding window, using a preset signature area template for template matching and calculating a matching similarity score; when the similarity score exceeds a preset threshold, generating a candidate signature area frame using the current window coordinates; introducing a spatial attention mechanism to weight the features of the sliding window area, enhancing the response to key features such as signature handwriting edges and textures, and suppressing background noise interference.
[0085] In some embodiments, the common size range of the preset signature includes 10×10 pixels to 200×200 pixels.
[0086] Multi-scale window traversal includes: window scale setting: covering 10×10 to 200×200 pixels (step size 20 pixels), with a total of 10 scales (10, 30, 50, …, 200), and the step size of each scale window is set to 1 / 4 of the window size (for example, the step size of a 50×50 window is 12 pixels), ensuring that the overlap rate of adjacent windows is 75% to avoid missing small-sized signatures.
[0087] The traversal method slides the window from left to right and from top to bottom in the coarse positioning area, and extracts the image blocks in the window as template matching input.
[0088] Template matching and attention enhancement include: preset signature template: based on historical data statistics, define the prior features of the signature area: stroke density: the proportion of pixels with grayscale values <150 is between 20% and 40% (the proportion of handwritten ink); edge features: calculate the horizontal / vertical edge density in the window through the Sobel operator, set the threshold to filter the smooth background; blank ratio: the blank area at the upper and lower edges of the window accounts for >30% (in line with the rule that signatures are usually located in the blank space).
[0089] The spatial attention mechanism applies the SE-Net channel attention module to the window image block before template matching, calculates the weight of each channel, enhances the response to the edges of signature handwriting (high-frequency feature channels), suppresses the channel weights of interfering features such as table lines (regular straight lines) and seals (large areas of color blocks), and re-enters the matching algorithm after the weights are adjusted.
[0090] Candidate boxes are generated by calculating the matching score: the normalized cross correlation (NCC) algorithm is used. When the score is greater than 0.75, a candidate box is generated. The coordinates are expressed as the upper left corner of the window (x, y) and the size (w, h).
[0091] In some embodiments, the merging of overlapping candidate frames corresponding to multiple candidate signature region frames by non-maximum suppression includes: calculating the intersection-over-union ratio between all candidate signature region frames; for candidate signature region frames whose intersection-over-union ratio is higher than a preset threshold, retaining the candidate signature region frame with the highest matching similarity score, and deleting the remaining candidate signature region frames as overlapping candidate frames; repeating the above process until the intersection-over-union ratios between all candidate signature region frames are lower than the preset threshold, thereby completing the deletion of overlapping candidate frames.
[0092] For any two candidate boxes A(x1,y1,w1,h1) and B(x2,y2,w2,h2), calculate the ratio of the intersection area to the union area: IoU= / The preset IoU threshold is 0.5 (empirical value, balancing the deduplication strength and preserving diversity).
[0093] The iterative deduplication process sorts all candidate boxes in descending order by matching score, and takes the box A with the highest score each time; deletes all boxes with IoU>0.5 with A, and retains A; repeats the above steps for the remaining boxes until there are no overlapping boxes (IoU≤0.5).
[0094] In some embodiments, the use of boundary regression to optimize the coordinates of the candidate signature area box includes: inputting the candidate signature area box into a preset boundary regression model, and based on a preset anchor frame mechanism, performing regression adjustment on the coordinates of the candidate signature area box to optimize the position and size of the candidate signature area box.
[0095] In some embodiments, the outputting of the signature position coordinates and waybill angle corresponding to the sample receipt image includes: fusing the image visual features and text semantic features corresponding to the candidate signature area box, jointly optimizing the boundary regression parameters through multimodal alignment technology, and outputting the signature position coordinates and waybill angle detection results including the upper left corner coordinates and the lower right corner coordinates.
[0096] The anchor box mechanism and coordinate regression adopt the RPN anchor box design of Faster R-CNN, preset three aspect ratios (1:1, 1:2, 2:1), generate an initial anchor box for each candidate box, input the boundary regression model (fully connected layer) to output the coordinate adjustment amount (Δx, Δy, Δw, Δh):
[0097] x′=x+w×Δx,y′=y+h×Δy,w′=w×eΔw,h′=h×eΔh; during model training, the real signature box is used as the label, and the smooth L1 loss function is used to optimize the regression parameters.
[0098] Where (x, y) is the initial upper-left corner coordinate (in pixels) of the candidate box, typically generated by a sliding window or coarse positioning module. It represents the initial position of the candidate box in the image and is the starting point for boundary regression.
[0099] w and h are the initial width and height (in pixels) of the candidate box, determined by the preset aspect ratio of the anchor box (such as 1:1, 1:2, 2:1) or the size of the coarse positioning region. Define the initial size of the candidate box to roughly cover the signature area.
[0100] Δx and Δy are dimensionless coordinate offsets, representing the proportional offset relative to the candidate box width w and height h. Δx and Δy can be positive or negative; positive values indicate a rightward / downward offset, while negative values indicate a leftward / upward offset.
[0101] Δw and Δh are dimensionless scaling adjustments, representing the logarithmic offsets of the width and height of the candidate box (mapped to linear space via an exponential function). Compared to directly predicting the scaling factor, predicting the logarithmic offset Δw transforms the optimization problem into an additive one, which aligns with the gradient update characteristics of neural networks and improves training stability.
[0102] (x′, y′) is the pixel-level coordinate of the upper-left corner of the target bounding box after boundary regression. It is calculated by adding a proportional offset to the original coordinates. The goal is to make (x′, y′) as close as possible to the upper-left corner of the true signature bounding box to minimize positional deviation.
[0103] w′, h′ are the pixel-level width and height of the target box after boundary regression, calculated by multiplying the original size by an exponential scaling factor. The goal is to ensure that w′, h′ closely matches the actual size of the real signature, avoiding invalid boxes that are too narrow / too wide, too high / too low. Through the above parameter design, the boundary regression module gradually optimizes the candidate boxes generated by coarse positioning (which may have positional deviations or size mismatches) to a region close to the real signature. Combined with multimodal feature fusion (visual + semantic), it ultimately achieves high-precision positioning for complex layouts and low-quality images, meeting the stringent requirements of industrial applications.
[0104] Multimodal feature fusion involves visual features. HOG features (Histogram of Oriented Gradients) and CNN high-level features (derived from the final convolutional output of the ResNet layer) are extracted from the candidate box region to describe the signature's edge distribution and texture pattern. Semantic features are derived from the keyword text box coordinates corresponding to the candidate box and the spatial position correlation is calculated (e.g., whether the candidate box is within 100 pixels below the keyword, generating a 0-1 correlation feature). Joint optimization is achieved by fusing visual and semantic features through a fully connected layer, outputting the final regression parameters. This allows coordinate adjustments to take into account both image content and semantic position constraints.
[0105] In some embodiments, a data enhancement solution that integrates a lightweight diffusion model, low-rank decomposition technology, and a MixCut hybrid strategy is used to resolve the contradiction between the efficiency and sample diversity of traditional methods in low-quality image restoration, thereby achieving a dual improvement in noise robustness and layout generalization ability.
[0106] The lightweight diffusion repair module builds a lightweight U-Net diffusion model and introduces low-rank decomposition technology during the diffusion process to impose rank constraints on the convolutional layer weight matrix (for example, reducing the number of channels from 512 to 128). This reduces the number of parameters by 65% compared to traditional diffusion models. It also uses dynamic thresholds to filter key diffusion steps (retaining only 10 core denoising steps), reducing the single-image repair time to less than 50ms.
[0107] The restoration process first pre-processes the image with abnormal lighting (too dark / reflective) or blurred light through brightness normalization, then inputs it into the model, and uses the diffusion process to gradually restore the signature stroke details (such as repairing broken handwriting tracks) and output a high-contrast image.
[0108] MixCut cross-layout hybrid enhancement randomly selects two receipt images (A and B) of different layouts, mixes their pixel values at a dynamic ratio of 0.3-0.7 (e.g., pixel value = A × λ + B × (1-λ), where λ∈[0.3,0.7]), and performs a weighted average of the label coordinates (e.g., signature area coordinate = coordinate A × λ + coordinate B × (1-λ)). This generates new samples containing cross-layout features (e.g., the signature area layout of a mixed express delivery receipt and logistics receipt). A layout-aware mechanism is introduced to prioritize mixing receipts of the same type (e.g., both are shipping receipts), ensuring the business rationality of the mixed samples and avoiding semantic conflicts.
[0109] In some embodiments, a three-level angle detection framework is designed to fuse edge semantic features with text direction information to solve the problem of angle misjudgment in complex backgrounds by traditional single models and achieve high-precision and high-efficiency direction correction.
[0110] Edge semantic contour extraction uses an improved Canny edge detection algorithm, combined with prior knowledge of waybills (e.g., rectangular contours account for >80%), to extract the main contour of the waybill using a dual threshold (high threshold 200, low threshold 80), filter out non-rectangular interfering contours (e.g., stickers, barcode edges), and filter through contour area (excluding contours with an area less than 10% of the entire image) to ensure that only the main area of the waybill is retained (extraction accuracy ≥99.5%).
[0111] The lightweight direction classification network constructs a Depthwise Separable Convolution lightweight ResNet (Light-ResNet for short), splitting the traditional 3×3 convolution into depthwise convolution and point convolution, reducing the number of parameters by 70%. It inputs a 224×224 body area image and outputs four-category angle labels (0° / 90° / 180° / 270°). The classification accuracy reaches 99.3% on industrial-grade datasets, and the single-image inference time is less than 15ms.
[0112] Fine-tuning the text line orientation is performed by extracting the text line tilt angle θ (range [-10°, 10°]) based on the Dbnet text detection network. An angle fusion formula is designed: final angle = coarse classification angle + θ. Pixel-level correction is achieved through an affine transformation matrix (e.g., coarse classification 90° + fine-tuning + 3°, generating a 93° rotation matrix). An angle smoothing constraint is also introduced (angle changes between adjacent frames ≤ 5°) to avoid interference from outliers.
[0113] In some embodiments, a cross-modal fusion positioning mechanism of OCR text semantics and visual features is used to narrow the search space through keyword semantic constraints, thereby solving the problem of missed detection of signature areas with "semantics but no fixed visual pattern" in complex layouts.
[0114] By using the BERT pre-trained model to build a dynamic keyword library, it supports automatic expansion of business-related vocabulary (for example, extracting potential keywords such as "customer signature" and "signee" from historically annotated data through remote supervision). Keyword weights are calculated using TF-IDF, prioritizing matching with high-frequency business terms (for example, "signature" has a higher weight than "signed"). A CRNN model is used to extract text box coordinates and content. For text boxes matching keywords, the search area is expanded in four directions: up, down, left, and right. Upward expansion: 0-50 pixels above the keyword text box (accommodating "signature above keyword" layout); downward expansion: 50-200 pixels below the keyword text box (covering over 90% of signature positions); and horizontal expansion: 30% to the left and right of the keyword text box (accommodating signatures spanning columns). Layout prioritization is introduced: if the keyword contains "shipper," downward expansion is prioritized; if it contains "consignee," upward expansion is prioritized, improving the business rationality of regional association.
[0115] In some embodiments, a candidate box generation strategy that combines a multi-scale sliding window with a spatial attention mechanism is used, and a multimodal boundary regression model that integrates visual texture and semantic position is used to solve the problems of missed detection of small-sized signatures and false detection of complex backgrounds by traditional methods, thereby achieving a breakthrough in pixel-level positioning accuracy.
[0116] Dynamic window scale adjustment: Automatically adjust the window range according to the size of the coarse positioning area (for example, when the area height is less than 200 pixels, the maximum window size is set to 150×150; when the area height is ≥200 pixels, it is expanded to 200×200). The golden ratio scale is also introduced (for example, 123×123 pixels, which is close to the common signature size of 0.618), covering more than 95% of real signature sizes.
[0117] Spatial Attention Enhancement: CBAM (Convolutional Block Attention Module) is applied within a sliding window, performing weighted attention in both the channel and spatial dimensions. This increases the gradient response value of signature edges by 2 times, suppresses the response value of table lines and seals by 40%, and generates high-confidence candidate boxes (triggered when the matching score is greater than 0.8).
[0118] Multimodal boundary regression: A two-branch regression model is constructed: the visual branch extracts HOG features and ResNet-50 high-level features from the candidate bounding box to describe the stroke density and edge distribution of the signature; the semantic branch inputs the spatial distance between the keyword text box and the candidate bounding box (for example, a positive sample is generated when the vertical distance is less than 100 pixels), and outputs semantic constraint parameters through a fully connected layer. Joint optimization: Using multi-task learning, a weighted fusion of the visual regression loss (smooth L1) and the semantic constraint loss (cross entropy) is performed (weighting 0.7:0.3), outputting the final signature region coordinates (upper left corner (x, y), lower right corner (w, h)) and angle θ (accuracy ±2°).
[0119] See also Figure 2 , Figure 2 The present application also provides a schematic block diagram of a receipt signature position and angle intelligent recognition device 200, which is used to perform the aforementioned receipt signature position and angle intelligent recognition method. The receipt signature position and angle intelligent recognition device can be configured in a server or terminal.
[0120] The server can be a standalone server or a server cluster, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, content delivery networks (CDNs), and big data and artificial intelligence platforms. The terminal can be an electronic device such as a mobile phone, tablet computer, laptop computer, desktop computer, user digital assistant, and wearable device.
[0121] like Figure 2 As shown, the receipt signature position and angle intelligent recognition device 200 includes:
[0122] An image acquisition unit 201 is configured to acquire an original receipt image and perform data enhancement processing on the original receipt image. The data enhancement processing includes brightness conversion, color gamut conversion, and image blending to obtain a sample receipt image, wherein the image blending is performed by randomly blending two images corresponding to the receipt image data and annotating the label ratios.
[0123] The image input unit 202 is used to input the sample return image into a preset deep learning classification model for coarse classification, perform fine angle adjustment within a range of ±10° on the classification result corresponding to the coarse classification, and output precise angle correction parameters corresponding to the sample return image to complete angle correction of the sample return image;
[0124] The text extraction unit 203 is used to extract text from the sample receipt image after angle correction, match the text extraction result with the preset signature area keywords, achieve rough positioning of the signature area in the sample receipt image, and output the rough positioning area;
[0125] Angle output unit 204 is used to search the coarse positioning area using a multi-scale sliding window combined with a template matching algorithm to generate a candidate signature area frame, merge overlapping candidate frames corresponding to multiple candidate signature area frames through non-maximum suppression, and optimize the coordinates of the candidate signature area frames using boundary regression to output the signature position coordinates and waybill angle corresponding to the sample receipt image.
[0126] In some embodiments, before performing data enhancement processing on the original return single image, it also includes: performing image denoising on the original return single image using a median filtering algorithm to remove white noise and salt and pepper noise; for original return single images whose clarity or lighting does not meet preset clarity conditions or preset lighting conditions, using a lightweight diffusion model combined with low-rank decomposition technology to perform image restoration.
[0127] In some embodiments, the sample receipt image is input into a preset deep learning classification model for coarse classification, the classification result corresponding to the coarse classification is fine-tuned within the range of ±10°, and the precise angle correction parameters corresponding to the sample receipt image are output, including: screening the main area of the sample receipt image according to the edge detection algorithm to obtain the main contour area of the waybill; inputting the main contour area into the preset deep learning classification model for coarse classification to obtain four classification angle labels of 0°, 90°, 180°, and 270°; based on the coarse classification result, the text line direction of the sample receipt image is detected, and the text box tilt angle parameter output by the network is used to perform angle offset correction within the range of ±10° on the basis of the coarse classification angle to generate precise angle correction parameters including a rotation matrix.
[0128] In some embodiments, the text extraction result is matched with the preset signature area keywords to achieve coarse positioning of the signature area in the sample receipt image and output the coarse positioning area, including: performing text extraction and coordinate positioning on the sample receipt image after angle correction to generate a text box set containing text content and position coordinates; performing semantic analysis on the text content in the text box set through NLP technology to match the preset signature area keywords, and generating an initial signature area candidate box as the coarse positioning area based on the coordinates of the successfully matched text box and the preset keyword-associated area offset.
[0129] In some embodiments, the use of a multi-scale sliding window combined with a template matching algorithm to search the coarse positioning area and generate a candidate signature area frame includes: traversing the coarse positioning area using a multi-scale sliding window, with the window scale covering a preset common signature size range and the step size set to 1 / 4 of the window size; within each sliding window, using a preset signature area template for template matching and calculating a matching similarity score; when the similarity score exceeds a preset threshold, generating a candidate signature area frame using the current window coordinates; introducing a spatial attention mechanism to weight the features of the sliding window area, enhancing the response to key features such as signature handwriting edges and textures, and suppressing background noise interference.
[0130] In some embodiments, the common size range of the preset signature includes 10×10 pixels to 200×200 pixels.
[0131] In some embodiments, the merging of overlapping candidate frames corresponding to multiple candidate signature region frames by non-maximum suppression includes: calculating the intersection-over-union ratio between all candidate signature region frames; for candidate signature region frames whose intersection-over-union ratio is higher than a preset threshold, retaining the candidate signature region frame with the highest matching similarity score, and deleting the remaining candidate signature region frames as overlapping candidate frames; repeating the above process until the intersection-over-union ratios between all candidate signature region frames are lower than the preset threshold, thereby completing the deletion of overlapping candidate frames.
[0132] In some embodiments, the use of boundary regression to optimize the coordinates of the candidate signature area box includes: inputting the candidate signature area box into a preset boundary regression model, and based on a preset anchor frame mechanism, performing regression adjustment on the coordinates of the candidate signature area box to optimize the position and size of the candidate signature area box.
[0133] In some embodiments, the outputting of the signature position coordinates and waybill angle corresponding to the sample receipt image includes: fusing the image visual features and text semantic features corresponding to the candidate signature area box, jointly optimizing the boundary regression parameters through multimodal alignment technology, and outputting the signature position coordinates and waybill angle detection results including the upper left corner coordinates and the lower right corner coordinates.
[0134] It should be noted that, those skilled in the art can clearly understand that, for the sake of convenience and brevity of description, the specific working processes of the model training device and each module described above can refer to the corresponding processes in the aforementioned embodiment of the method for intelligent recognition of receipt signature position and angle, and will not be repeated here.
[0135] The above-mentioned receipt signature position and angle intelligent recognition device can be implemented in the form of a computer program. The computer program can be used in Figure 3 Runs on the computer equipment shown.
[0136] See also Figure 3 , Figure 3 This is a schematic block diagram of the structure of a computer device provided in an embodiment of the present application. The computer device may be a server or a terminal.
[0137] See Figure 3 The computer device includes a processor, a memory and a network interface connected through a system bus, wherein the memory may include a storage medium and an internal memory.
[0138] The storage medium may store an operating system and a computer program. The computer program includes program instructions, which, when executed, enable the processor to execute any one of the methods for intelligently identifying the position and angle of a receipt signature provided in the embodiments of the present application.
[0139] The processor is used to provide computing and control capabilities and support the operation of the entire computer equipment.
[0140] The internal memory provides an environment for the execution of a computer program stored in a storage medium. When executed by a processor, the computer program enables the processor to perform any of the methods for intelligently identifying the position and angle of a receipt signature. The storage medium may be non-volatile or volatile.
[0141] The network interface is used for network communication, such as sending assigned tasks, etc. Those skilled in the art will understand that Figure 3 The structure shown in the figure is only a block diagram of a part of the structure related to the solution of the present application, and does not constitute a limitation on the computer device to which the solution of the present application is applied. The specific computer device may include more or fewer components than shown in the figure, or combine certain components, or have a different component arrangement.
[0142] It should be understood that the processor may be a central processing unit (CPU), other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field-programmable gate arrays (FPGA), other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or any conventional processor, etc.
[0143] Exemplarily, in one embodiment, the processor is configured to execute a computer program stored in the memory to implement the following steps:
[0144] Obtain an original receipt image and perform data enhancement processing on the original receipt image. The data enhancement processing includes brightness conversion, color gamut conversion, and image blending to obtain a sample receipt image, wherein the image blending is performed by randomly blending two images corresponding to the receipt image data and annotating the label ratio;
[0145] Input the sample return image into a preset deep learning classification model for coarse classification, perform fine angle adjustment within a range of ±10° on the classification result corresponding to the coarse classification, output the precise angle correction parameter corresponding to the sample return image, and complete the angle correction of the sample return image;
[0146] Perform text extraction on the angle-corrected sample receipt image, match the text extraction result with the preset signature area keywords, achieve coarse positioning of the signature area in the sample receipt image, and output the coarse positioning area;
[0147] A multi-scale sliding window combined with a template matching algorithm is used to search the coarse positioning area to generate a candidate signature area frame. The overlapping candidate frames corresponding to multiple candidate signature area frames are merged through non-maximum suppression, and the coordinates of the candidate signature area frames are optimized using boundary regression. The signature position coordinates and waybill angle corresponding to the sample receipt image are output.
[0148] In some embodiments, before performing data enhancement processing on the original return single image, it also includes: performing image denoising on the original return single image using a median filtering algorithm to remove white noise and salt and pepper noise; for original return single images whose clarity or lighting does not meet preset clarity conditions or preset lighting conditions, using a lightweight diffusion model combined with low-rank decomposition technology to perform image restoration.
[0149] In some embodiments, the sample receipt image is input into a preset deep learning classification model for coarse classification, the classification result corresponding to the coarse classification is fine-tuned within the range of ±10°, and the precise angle correction parameters corresponding to the sample receipt image are output, including: screening the main area of the sample receipt image according to the edge detection algorithm to obtain the main contour area of the waybill; inputting the main contour area into the preset deep learning classification model for coarse classification to obtain four classification angle labels of 0°, 90°, 180°, and 270°; based on the coarse classification result, the text line direction of the sample receipt image is detected, and the text box tilt angle parameter output by the network is used to perform angle offset correction within the range of ±10° on the basis of the coarse classification angle to generate precise angle correction parameters including a rotation matrix.
[0150] In some embodiments, the text extraction result is matched with the preset signature area keywords to achieve coarse positioning of the signature area in the sample receipt image and output the coarse positioning area, including: performing text extraction and coordinate positioning on the sample receipt image after angle correction to generate a text box set containing text content and position coordinates; performing semantic analysis on the text content in the text box set through NLP technology to match the preset signature area keywords, and generating an initial signature area candidate box as the coarse positioning area based on the coordinates of the successfully matched text box and the preset keyword-associated area offset.
[0151] In some embodiments, the use of a multi-scale sliding window combined with a template matching algorithm to search the coarse positioning area and generate a candidate signature area frame includes: traversing the coarse positioning area using a multi-scale sliding window, with the window scale covering a preset common signature size range and the step size set to 1 / 4 of the window size; within each sliding window, using a preset signature area template for template matching and calculating a matching similarity score; when the similarity score exceeds a preset threshold, generating a candidate signature area frame using the current window coordinates; introducing a spatial attention mechanism to weight the features of the sliding window area, enhancing the response to key features such as signature handwriting edges and textures, and suppressing background noise interference.
[0152] In some embodiments, the common size range of the preset signature includes 10×10 pixels to 200×200 pixels.
[0153] In some embodiments, the merging of overlapping candidate frames corresponding to multiple candidate signature region frames by non-maximum suppression includes: calculating the intersection-over-union ratio between all candidate signature region frames; for candidate signature region frames whose intersection-over-union ratio is higher than a preset threshold, retaining the candidate signature region frame with the highest matching similarity score, and deleting the remaining candidate signature region frames as overlapping candidate frames; repeating the above process until the intersection-over-union ratios between all candidate signature region frames are lower than the preset threshold, thereby completing the deletion of overlapping candidate frames.
[0154] In some embodiments, the use of boundary regression to optimize the coordinates of the candidate signature area box includes: inputting the candidate signature area box into a preset boundary regression model, and based on a preset anchor frame mechanism, performing regression adjustment on the coordinates of the candidate signature area box to optimize the position and size of the candidate signature area box.
[0155] In some embodiments, the outputting of the signature position coordinates and waybill angle corresponding to the sample receipt image includes: fusing the image visual features and text semantic features corresponding to the candidate signature area box, jointly optimizing the boundary regression parameters through multimodal alignment technology, and outputting the signature position coordinates and waybill angle detection results including the upper left corner coordinates and the lower right corner coordinates.
[0156] The computer-readable storage medium may be an internal storage unit of the computer device described in the aforementioned embodiment, such as a hard disk or memory of the computer device. The computer-readable storage medium may also be an external storage device of the computer device, such as a plug-in hard disk, a SmartMedia Card (SMC), a Secure Digital (SD) card, a flash memory card, etc. equipped on the computer device.
[0157] The above description is merely a specific embodiment of the present application, but the scope of protection of the present application is not limited thereto. Any person skilled in the art can easily conceive of various equivalent modifications or substitutions within the technical scope disclosed in the present application, and such modifications or substitutions should be included in the scope of protection of the present application. Therefore, the scope of protection of the present application should be based on the scope of protection of the claims.
Claims
1. A method for intelligently identifying receipt signature position and angle based on deep learning, characterized in that: include: Obtain an original receipt image and perform data enhancement processing on the original receipt image. The data enhancement processing includes brightness transformation, color gamut transformation, and image blending to obtain a sample receipt image. The image blending is performed by randomly blending two images corresponding to the original receipt image and annotating the label ratios. Input the sample receipt image into a preset deep learning classification model for coarse classification, perform fine-tuning on the classification result corresponding to the coarse classification within the range of ±10°, and output precise angle correction parameters corresponding to the sample receipt image, including: screening the main area of the sample receipt image according to the edge detection algorithm to obtain the main contour area of the waybill; inputting the main contour area into the preset deep learning classification model for coarse classification to obtain four classification angle labels of 0°, 90°, 180°, and 270°; based on the coarse classification result, detecting the text line direction of the sample receipt image, and performing angle offset correction within the range of ±10° on the basis of the coarse classification angle through the text box tilt angle parameter output by the network, to generate precise angle correction parameters including a rotation matrix; completing angle correction of the sample receipt image; Performing text extraction on the angle-corrected sample receipt image, matching the text extraction result with preset signature area keywords, achieving coarse positioning of the signature area in the sample receipt image, and outputting the coarse positioning area, including: performing text extraction and coordinate positioning on the angle-corrected sample receipt image to generate a text box set containing text content and position coordinates; performing semantic analysis on the text content in the text box set using NLP technology to match preset signature area keywords, and generating an initial signature area candidate box as the coarse positioning area based on the coordinates of the successfully matched text boxes and a preset keyword-associated area offset; A multi-scale sliding window combined with a template matching algorithm is used to search the coarse positioning area to generate a candidate signature area frame, including: traversing the coarse positioning area using a multi-scale sliding window, with the window scale covering a preset common signature size range and a step size set to 1 / 4 of the window size; within each sliding window, template matching is performed using a preset signature area template to calculate a matching similarity score; when the similarity score exceeds a preset threshold, a candidate signature area frame is generated using the current window coordinates; a spatial attention mechanism is introduced to weight the features of the sliding window area, thereby enhancing the response to the signature handwriting edges and textures and suppressing background noise interference; overlapping candidate frames corresponding to multiple candidate signature area frames are merged through non-maximum suppression, and the coordinates of the candidate signature area frames are optimized using boundary regression to output the signature position coordinates and waybill angle corresponding to the sample receipt image.
2. The method according to claim 1, characterized in that Before performing data enhancement processing on the original receipt image, the method further includes: The median filter algorithm is used to perform image denoising on the original single return image to remove white noise and salt and pepper noise; For original single-reply images whose clarity or lighting does not meet the preset clarity conditions or preset lighting conditions, a lightweight diffusion model combined with low-rank decomposition technology is used to perform image restoration.
3. The method according to claim 1, characterized in that Common sizes of the preset signatures range from 10×10 pixels to 200×200 pixels.
4. The method according to claim 1, wherein The merging of overlapping candidate frames corresponding to the plurality of candidate signature region frames by non-maximum suppression includes: Calculate the intersection-over-union ratio (IoU) between all candidate signature region frames; for candidate signature region frames whose IoU is higher than a preset threshold, retain the candidate signature region frame with the highest matching similarity score, and treat the remaining candidate signature region frames as overlapping candidate frames and delete them; The above process is repeated until the intersection-over-union ratios between all candidate signature region frames are lower than the preset threshold, and the overlapping candidate frames are deleted.
5. The method according to claim 1, wherein The optimizing the coordinates of the candidate signature region frame by using boundary regression includes: The candidate signature region frame is input into a preset boundary regression model, and based on a preset anchor frame mechanism, the coordinates of the candidate signature region frame are regressively adjusted to optimize the position and size of the candidate signature region frame.
6. The method according to claim 1, characterized in that Outputting the signature position coordinates and waybill angle corresponding to the sample receipt image includes: The image visual features and text semantic features corresponding to the candidate signature area box are integrated, and the boundary regression parameters are jointly optimized through multimodal alignment technology to output the signature position coordinates including the upper left corner coordinates, the lower right corner coordinates and the waybill angle detection results.
7. A device for intelligently identifying the position and angle of a receipt signature, characterized in that: The device comprises: An image acquisition unit is configured to acquire an original receipt image and perform data enhancement processing on the original receipt image, wherein the data enhancement processing includes brightness conversion, color gamut conversion, and image blending to obtain a sample receipt image, wherein the image blending is performed by randomly blending two images corresponding to the original receipt image and annotating the label ratio; An image input unit is used to input the sample receipt image into a preset deep learning classification model for coarse classification, perform fine-tuning of the angle within the range of ±10° on the classification result corresponding to the coarse classification, and output precise angle correction parameters corresponding to the sample receipt image, including: screening the main area of the sample receipt image according to the edge detection algorithm to obtain the main contour area of the waybill; inputting the main contour area into the preset deep learning classification model for coarse classification to obtain four classification angle labels of 0°, 90°, 180°, and 270°; based on the coarse classification result, detecting the text line direction of the sample receipt image, performing angle offset correction within the range of ±10° on the basis of the coarse classification angle through the text box tilt angle parameter output by the network, and generating precise angle correction parameters including a rotation matrix; and completing the angle correction of the sample receipt image; a text extraction unit, configured to extract text from the angle-corrected sample receipt image, match the text extraction result with a preset signature area keyword, achieve coarse positioning of the signature area in the sample receipt image, and output a coarse positioning area, including: extracting text and locating coordinates on the angle-corrected sample receipt image to generate a text box set containing text content and position coordinates; performing semantic analysis on the text content in the text box set using NLP technology to match the preset signature area keyword, and generating an initial signature area candidate box as the coarse positioning area based on the coordinates of the successfully matched text boxes and a preset keyword-associated area offset; The angle output unit is used to search the coarse positioning area using a multi-scale sliding window combined with a template matching algorithm to generate a candidate signature area frame, including: traversing the coarse positioning area using a multi-scale sliding window, with the window scale covering a preset common signature size range and a step size set to 1 / 4 of the window size; within each sliding window, using a preset signature area template for template matching to calculate a matching similarity score; when the similarity score exceeds a preset threshold, generating a candidate signature area frame based on the current window coordinates; introducing a spatial attention mechanism to weight the features of the sliding window area, enhancing the response to the signature handwriting edges and textures, and suppressing background noise interference; merging overlapping candidate frames corresponding to multiple candidate signature area frames through non-maximum suppression, and optimizing the coordinates of the candidate signature area frames using boundary regression, and outputting the signature position coordinates and waybill angle corresponding to the sample receipt image.
Citation Information
Patent Citations
Water meter image angle correction and accurate positioning method
CN114140795A
Text detection and template matching-based bill identification method and device
CN118736613A