Transaction voucher target detection method and device, equipment and storage medium
Patent Information
- Application Number
- CN202611053821.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-07-15
- Publication Date
- 2026-09-25
AI Technical Summary
[0003]然而,现有目标检测方案在交易凭证图像中面对多尺度目标、复杂背景及形变干扰时,容易出现小尺寸关键信息特征丢失,导致漏检和定位精度下降;同时,为提高检测精度而增加网络规模,又会带来计算量和参数量上升,难以兼顾实时处理需求与部署约束
[0016]本申请实施例提供的交易凭证的目标检测方法、装置、设备及存储介质,通过对待检测的交易凭证图像进行多尺度特征提取,并对浅层细节特征、中层过渡特征和深层语义特征执行双向跨尺度融合,能够强化不同层级特征之间的信息交互与优势互补,提升复杂背景下目标表征的完整性;小目标本身像素占比低、纹理信息稀缺,经多层下采样后深层融合特征极易出现分辨率衰减、边缘纹理丢失,为此,基于高分辨率的浅层细节特征对深层融合特征进行分辨率还原与细节补充,修复小目标模糊轮廓、补充缺失的微细纹理;并结合常规尺寸目标检测与小尺寸目标检测的分层检测策略,专门适配小目标微弱特征提取需求,从特征修复与分层检测双重维度缓解小目标关键信息损耗引发的漏检、定位偏移问题,最终在约束模型参数量与计算开销的前提下,同步提升交易凭证整体检测精度、目标定位准确度与实时推理性能。
Smart Images

Figure CN122821301A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of computer vision and object detection, and in particular to a method, apparatus, device and storage medium for object detection of transaction credentials. Background Technology
[0002] In the remote transaction certificate recognition in the financial field, deep learning-based target detection models are typically used to detect and recognize certificate images collected by mobile devices or self-service terminals in real time.
[0003] However, existing target detection schemes are prone to losing small-sized key information features when faced with multi-scale targets, complex backgrounds, and deformation interference in transaction certificate images, leading to missed detections and decreased positioning accuracy. At the same time, increasing the network size to improve detection accuracy will increase the amount of computation and parameters, making it difficult to balance the needs of real-time processing with deployment constraints.
[0004] Therefore, how to improve the detection accuracy and real-time processing capability of small targets in complex voucher scenarios under limited computing resources has become an urgent technical problem to be solved. Summary of the Invention
[0005] This application provides a target detection method, apparatus, device, and storage medium for transaction vouchers to solve the aforementioned technical problems. This solution addresses the issues of insufficient detection capability for small targets caused by the coexistence of multi-scale targets, complex backgrounds, and local deformation interference in transaction voucher images. It constructs a technical approach that considers feature representation at different levels, collaborative utilization of feature information, and differentiated target detection. Under limited computing resources, it improves the detection and localization performance of various targets, especially small targets, in transaction voucher images, while also considering processing efficiency and deployment adaptability.
[0006] In a first aspect, embodiments of this application provide a target detection method for transaction vouchers, comprising: extracting multi-scale features from the transaction voucher image to be detected to obtain shallow detail features, mid-level transition features, and deep semantic features; performing bidirectional cross-scale fusion on the shallow detail features, mid-level transition features, and deep semantic features to obtain shallow fusion features, mid-level fusion features, and deep fusion features; performing resolution restoration and detail supplementation on the deep fusion features based on the shallow detail features to obtain high-resolution fusion features; performing target detection of transaction vouchers of normal size based on the shallow fusion features, mid-level fusion features, and deep fusion features, and performing target detection of transaction vouchers of small size based on the high-resolution fusion features to obtain detection results.
[0007] In one possible embodiment, resolution restoration and detail supplementation are performed on deep fusion features based on shallow detail features to obtain high-resolution fusion features. This includes: processing the deep fusion features based on processing units composed of convolutional layers, normalization layers, and activation layers; upsampling the processed deep fusion features to obtain high-resolution features with spatial dimensions consistent with the shallow detail features; performing multiple convolution operations on the high-resolution features to obtain enhanced high-resolution features; and fusing the enhanced high-resolution features with the shallow detail features to obtain high-resolution fusion features.
[0008] In one possible embodiment, performing multiple convolution operations on high-resolution features to obtain enhanced high-resolution features includes: inputting the high-resolution features into a first processing branch and a second processing branch respectively; performing multiple convolution operations on the high-resolution features through the first processing branch, and performing a single convolution operation on the high-resolution features through the second processing branch; and performing a single convolution operation after concatenating the outputs of the first processing branch and the second processing branch to obtain the enhanced high-resolution features.
[0009] In one possible embodiment, before performing bidirectional cross-scale fusion on shallow detail features, mid-level transition features, and deep semantic features to obtain shallow fused features, mid-level fused features, and deep fused features, the method further includes: performing global average pooling and global max pooling on the feature maps of each level in the shallow detail features, mid-level transition features, and deep semantic features respectively to obtain two types of global feature descriptors; generating channel attention weight vectors based on the two types of global feature descriptors; multiplying the channel attention weight vectors with the feature maps of the corresponding levels channel by channel to obtain channel-calibrated feature maps; performing multi-scale depthwise separable convolution processing on the channel-calibrated feature maps to generate spatial attention weight maps; and multiplying the spatial attention weight maps with the channel-calibrated feature maps element by element to obtain spatially calibrated feature maps, so as to perform bidirectional cross-scale fusion on the spatially calibrated feature maps of each level.
[0010] In one possible embodiment, the detection of regular-sized targets on transaction vouchers is performed based on shallow fusion features, medium-layer fusion features, and deep fusion features, and the detection of small-sized targets on transaction vouchers is performed based on high-resolution fusion features to obtain detection results. This includes: inputting the shallow fusion features, medium-layer fusion features, deep fusion features, and high-resolution fusion features into the corresponding detection branches to obtain the target detection results of the corresponding sizes; and integrating the target detection results of the corresponding sizes output by each detection branch to obtain the detection results.
[0011] In one possible embodiment, multi-scale feature extraction is performed on the transaction voucher image to be detected to obtain shallow detail features, mid-level transition features, and deep semantic features. This includes: sequentially performing dimensionality reconstruction, downsampling, convolutional transformation, channel normalization, and nonlinear activation processing on the transaction voucher image to obtain shallow detail features; performing multiple sets of cross-stage local convolutional iterations on the shallow detail features, and performing channel integration operations after each set of convolutional iterations to obtain mid-level transition features; performing multiple sets of cross-stage local convolutional iterations on the mid-level transition features, and performing channel integration operations after each set of convolutional iterations to obtain initial deep features; performing spatial pyramid multi-scale pooling operations on the initial deep features, and concatenating the results of pooling operations at different scales to obtain deep semantic features.
[0012] Secondly, embodiments of this application provide a target detection device for transaction vouchers, comprising: a multi-level feature extraction module for extracting multi-scale features from an image of the transaction voucher to be detected, obtaining shallow detail features, mid-level transition features, and deep semantic features; a feature fusion module for performing bidirectional cross-scale fusion on the shallow detail features, mid-level transition features, and deep semantic features, obtaining shallow fused features, mid-level fused features, and deep fused features; a feature enhancement module for performing resolution restoration and detail supplementation on the deep fused features based on the shallow detail features, obtaining high-resolution fused features; and a target detection module for detecting targets of normal size on the transaction voucher based on the shallow fused features, mid-level fused features, and deep fused features, and detecting small-sized targets on the transaction voucher based on the high-resolution fused features, obtaining detection results.
[0013] Thirdly, embodiments of this application provide an electronic device, including: a processor, and a memory communicatively connected to the processor; the memory stores computer-executable instructions; the processor executes the computer-executable instructions stored in the memory to implement the method provided above.
[0014] Fourthly, embodiments of this application provide a computer-readable storage medium storing computer-executable instructions, which, when executed by a processor, are used to implement the methods provided above.
[0015] Fifthly, embodiments of this application provide a computer program product, including a computer program that, when executed by a processor, implements the method described above.
[0016] The target detection method, apparatus, device, and storage medium for transaction vouchers provided in this application extract multi-scale features from the transaction voucher image to be detected, and perform bidirectional cross-scale fusion of shallow detail features, mid-level transition features, and deep semantic features. This strengthens the information interaction and complementary advantages between features at different levels, and improves the integrity of target representation in complex backgrounds. Small targets themselves have low pixel ratios and scarce texture information. After multi-layer downsampling, deep fusion features are prone to resolution decay and edge texture loss. Therefore, based on high-resolution shallow detail features, resolution restoration and detail supplementation are performed on deep fusion features to repair the blurred outline of small targets and supplement missing micro-textures. Furthermore, by combining the layered detection strategy of conventional size target detection and small size target detection, it is specifically adapted to the weak feature extraction requirements of small targets. From the dual dimensions of feature repair and layered detection, it alleviates the problems of missed detection and positioning offset caused by the loss of key information of small targets. Finally, under the premise of constraining the number of model parameters and computational overhead, it simultaneously improves the overall detection accuracy of transaction vouchers, target positioning accuracy, and real-time inference performance. Attached Figure Description
[0017] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this application and, together with the description, serve to explain the principles of this application.
[0018] Figure 1 A flowchart illustrating the target detection method for the transaction certificate provided in this application;
[0019] Figure 2 The structural block diagram of the target detection model provided in this application;
[0020] Figure 3 A schematic diagram of the target detection device for a transaction certificate provided in this application;
[0021] Figure 4 A schematic diagram of the structure of the electronic device provided in this application.
[0022] The accompanying drawings illustrate specific embodiments of this application, which will be described in more detail below. These drawings and descriptions are not intended to limit the scope of the concept in any way, but rather to illustrate the concept of this application to those skilled in the art through reference to particular embodiments. Detailed Implementation
[0023] Exemplary embodiments will now be described in detail, examples of which are illustrated in the accompanying drawings. When the following description relates to the drawings, unless otherwise indicated, the same numbers in different drawings denote the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with this application. Rather, they are merely examples of apparatuses and methods consistent with some aspects of this application.
[0024] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, data stored, data displayed, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties. Furthermore, the collection, storage, use, processing, transmission, provision, disclosure, and application of the relevant data all comply with relevant laws, regulations, and standards, necessary confidentiality measures have been taken, they do not violate public order and good morals, and corresponding operation portals are provided for users to choose to authorize or refuse.
[0025] Furthermore, the technical solution involved in this application, which involves big data analysis of user information (including but not limited to personal biometrics, identity data, consumption data, asset data, electronic terminal operation data, etc.) and the use of artificial intelligence technology for automated decision-making, and makes decisions that have a significant impact on personal rights based on the results of automated decision-making, provides users with corresponding operation entry points for users to choose to agree to or reject the results of automated decision-making; if the user chooses to reject, the process will proceed to the expert decision-making process.
[0026] It should be noted that the target detection method, apparatus, equipment and storage medium provided in this application can be used in the financial field, or in any field other than finance. It is only necessary to replace the transaction certificate image with an image of the corresponding field. This application does not limit the specific application field.
[0027] Transaction document target detection technology is mainly applied in remote transaction document recognition scenarios in the financial sector, typically completed collaboratively by mobile terminals, self-service terminals, or server-side recognition systems. After users take or upload images of transaction documents such as ID cards, bank cards, business documents, and contracts, the system needs to quickly detect the target documents and their key areas in the images, providing basic input for subsequent information recognition, compliance verification, and business review.
[0028] Existing solutions typically employ deep learning-based target detection models to extract features from the collected document images and perform target classification and location regression on feature maps at different scales, in order to adapt to targets of different sizes, such as complete documents, local regions, and text information, in the same image.
[0029] However, in actual transaction document images, complete documents and tiny key areas often coexist, and the background is often interspersed with interference factors such as desktop textures, shadows, reflections, occlusions, and shooting distortions. Limited by existing feature extraction and detection structures, the detailed texture and edge information that small targets rely on are easily weakened during downsampling, leading to missed detections of targets such as document photos, signature areas, and serial number areas, or false detections in complex backgrounds.
[0030] On the other hand, simply increasing network depth or width to improve detection accuracy, while enhancing semantic representation, leads to a significant increase in the number of parameters and computational load, which is detrimental to real-time deployment on mobile devices or self-service terminals. Especially in remote transaction scenarios requiring rapid response, if the detection process is time-consuming or the positioning deviation is large, it will directly affect subsequent recognition results, review efficiency, and user interaction experience.
[0031] In view of this, how to balance the detection performance of both regular-sized and small-sized targets under limited computing resources, and improve the localization accuracy and processing efficiency in complex voucher scenarios, has become an urgent technical problem to be solved. To address the above issues, this application provides a target detection method for transaction vouchers. By extracting multi-scale features from the image to be detected, performing bidirectional cross-scale fusion of shallow detail features, mid-level transition features, and deep semantic features, and then restoring resolution and supplementing details based on shallow detail features to the deep fused features, the method can respectively complete the detection of regular-sized targets and small-sized targets, thereby improving the accuracy and efficiency of transaction voucher detection in complex scenarios.
[0032] The technical solution of this application and how the technical solution of this application solves the above-mentioned technical problems are described in detail below with specific embodiments. These specific embodiments can be combined with each other, and the same or similar concepts or processes may not be described again in some embodiments. The embodiments of this application will now be described with reference to the accompanying drawings.
[0033] Figure 1 A flowchart illustrating the target detection method for the transaction certificate provided in this application is shown below. Figure 1 As shown, the target detection method includes:
[0034] S101: Perform multi-scale feature extraction on the transaction voucher image to be detected to obtain shallow detail features, mid-level transition features, and deep semantic features.
[0035] The transaction voucher image can be a single-frame image captured by a mobile terminal camera, a static image obtained by scanning at a self-service terminal, or a voucher photo uploaded by the business system. The transaction voucher image may include an ID card, bank card, business document, contract page, signature area, document photo area, serial number area, or other target areas that need to be detected.
[0036] Shallow detail features are used to preserve fine-grained information such as image edges and textures, and their size can be 80×80; mid-level transition features are used to connect the semantic transition between shallow and deep features, and their size can be 40×40; deep semantic features are used to represent deeper and more abstract target semantic information, and their size can be 20×20.
[0037] In its implementation, this object detection method can be executed by a detection module in a mobile terminal processor, edge computing device, or server. It first receives the transaction voucher image to be detected and optionally preprocesses it to meet network inference requirements. Multi-scale feature extraction can be accomplished by extracting features at different levels through a backbone network.
[0038] For example, the backbone network can be a lightweight backbone network built on CSPDarknet, a lightweight backbone network of the MobileNet series, a lightweight backbone network of EfficientNet, or a backbone network of the ShuffleNet series.
[0039] After inputting the transaction voucher image or a preprocessed transaction voucher image into the backbone network, the shallow layers of the backbone network undergo a small amount of sampling to retain a large amount of pixel-level texture, text edges, seal outlines, and other fine information, corresponding to the output of shallow detail features. The middle layers of the backbone network undergo moderate downsampling to take into account both local texture details and local target contextual relationships, forming middle-layer transition features that connect the information of the shallow and deep layers. The deep layers of the backbone network undergo multiple downsampling operations, significantly compressing the image spatial resolution, filtering out redundant and fragmented noise, and focusing on high-dimensional abstract information such as the overall layout of the voucher and the target category, outputting deep semantic features with strong semantic discriminative capabilities. By progressively extracting three types of multi-scale features through the backbone network, the network fully covers multi-level information from fine texture to global semantics, providing a multi-dimensional feature foundation for subsequent cross-scale feature fusion.
[0040] Based on the above analysis, it can be seen that multi-scale feature extraction enables the edges, textures, local structures, and overall semantics of transaction voucher images to be represented at different levels. Among them, shallow detail features retain signature strokes, number character edges, and document outlines; mid-level transition features represent the local block structure and region combination relationships of the document; and deep semantic features represent the overall semantics of the complete document, form areas, and page layout. This provides a complementary feature base from different sources for subsequent object detection at different scales.
[0041] S102: Perform bidirectional cross-scale fusion on shallow detail features, mid-level transition features and deep semantic features to obtain shallow fusion features, mid-level fusion features and deep fusion features.
[0042] Among them, bidirectional cross-scale fusion is used to conduct bidirectional information interaction between features at different levels, such as between shallow detail features and mid-level transition features, between shallow detail features and deep semantic features, or between mid-level transition features and deep semantic features, to obtain fused features at the corresponding levels.
[0043] Bottom-up fusion starts with shallow, detailed features and proceeds upwards step by step. Specifically, shallow, detailed features are downsampled to the same spatial scale as intermediate-level transition features and then fused with these features to obtain a preliminary intermediate-level fused feature. This preliminary intermediate-level fused feature is then further downsampled to the spatial scale of deep semantic features and fused with them to obtain a deep-level fused feature. Top-down fusion proceeds downwards step by step. Specifically, deep-level fused features are upsampled to the spatial size of intermediate-level transition features and then fused with the aforementioned preliminary intermediate-level fused feature to obtain a intermediate-level fused feature. This final intermediate-level fused feature is then upsampled to the spatial scale of shallow, detailed features and fused with them to obtain a shallow-level fused feature. The methods for fusing the two features include, but are not limited to, element-wise addition, concatenation, and channel-wise weighting.
[0044] If a concatenation method is used, feature reorganization can be performed after concatenation; if an addition method is used, channel consistency can be achieved for each input feature before fusion.
[0045] In one possible implementation, the shallow, medium, and deep fusion features are all corrected before being output to the subsequent detection module. Correction can be normalization correction, such as using Batch Normalization (BN) layers or layer normalization to standardize feature channels and spatial dimensions, balancing the feature amplitudes of each channel and eliminating numerical imbalances caused by scale differences. Correction can also be attention-weighted correction, introducing lightweight channel attention or spatial attention to enhance the features of effective target areas such as document text, seals, and tables, while suppressing invalid background noise such as watermarks and messy background patterns. Residual correction can also be used, such as fine-tuning the fusion features through 1×1 convolutions and small residual modules to correct feature distortion and edge blurring caused by upsampling and downsampling, further preserving the fine textures of small targets. Correction can also be performed through convolutional dimensionality reduction and feature filtering to eliminate redundant channel features generated during the fusion process, compressing feature computation without losing effective information.
[0046] Based on the above analysis, it can be seen that bidirectional cross-scale fusion is not simply about passing deep fusion features to lower layers, but rather about supplementing deep semantics downwards while preserving shallow layer details, and feeding back shallow layer edges and textures upwards, so that both intermediate and deep layers receive local detail compensation. For complete documents, contract layouts, and table areas in transaction voucher scenarios, deep semantics can provide stable overall target discrimination criteria; for boundary-sensitive areas such as document photo areas, signature areas, and number areas, shallow details can supplement contour and texture information. Therefore, the fused multi-layer features can simultaneously adapt to category judgment and position regression of targets of conventional sizes, alleviating the problem of weakened details in small areas or insufficient deep semantics at a single scale.
[0047] S103: Based on shallow detail features, resolution restoration and detail supplementation are performed on deep fusion features to obtain high-resolution fusion features.
[0048] Resolution restoration and detail enhancement are used to restore deep fusion features to a higher spatial resolution and supplement detailed information to suit small target detection. Resolution restoration specifically involves upsampling low-resolution deep fusion features to restore their spatial size to a high-resolution scale consistent with or similar to shallow detail features. Detail enhancement involves fusing the resolution-restored deep fusion features with shallow detail features, such as through stitching, element-wise addition, or weighted summation. This injects the subtle texture and edge contour information carried by the shallow detail features into the restored deep fusion features, compensating for the key details of small targets lost during downsampling and resulting in high-resolution fusion features.
[0049] For example, the size of the shallow detail feature can be 80×80. When supplementing details, the shallow detail feature can be stitched together with the deep fusion feature after resolution restoration, so that the size of the high-resolution fusion feature can be 160×160.
[0050] In practical implementation, the resolution of deep fusion features can be upsampled as described above, increasing their spatial size to be consistent with or close to that of shallow detail features. After size alignment, details are supplemented by combining the shallow detail features to obtain high-resolution fusion features. Detail supplementation can be achieved through enhancement processing after feature fusion, or by restoring spatial resolution in stages and supplementing edge, texture, and local structural information.
[0051] In this application, the spatial resolution of the high-resolution fusion features can be consistent with that of the shallow detail features to match the small-size target detection branch.
[0052] Based on the above analysis, while deep fusion features possess strong target semantic discrimination capabilities, their low resolution makes it difficult to directly locate minute, critical areas within transaction documents. By restoring resolution and supplementing details, deep semantics can recover spatial location representation capabilities while incorporating shallow edge and texture information, thus forming high-resolution fusion features that combine semantic integrity with fine spatial granularity. When used for detecting small targets such as document photo areas, signature areas, serial number areas, and seal edge areas, this feature provides more stable contour and local texture responses, thereby supporting subsequent detection branches to complete more accurate location regression and category discrimination.
[0053] S104: Detection of regular-sized targets on transaction vouchers based on shallow, medium, and deep fusion features, and detection of small-sized targets on transaction vouchers based on high-resolution fusion features, and obtain the detection results.
[0054] The standard size can include multiple sizes, with the smaller size being smaller than the smallest of the standard sizes.
[0055] For example, common sizes may include 80×80, 40×40, 20×20, etc.
[0056] The standard-sized target detection function is used to detect standard-sized targets in transaction documents. Specifically, it inputs shallow, medium, and deep fusion features into the corresponding detection branches and outputs the corresponding standard-sized target detection results. The small-sized target detection function is used to detect small-sized targets in transaction documents. Specifically, it inputs high-resolution fusion features into the corresponding detection branch and uses features with higher resolution and added details to perform small-sized target detection, obtaining small target detection results. The small target detection results are then integrated with the target detection results of each standard size to obtain the final detection result.
[0057] In the specific implementation, each detection branch can include a classification subnetwork and a regression subnetwork. The classification subnetwork outputs classification information, such as the score indicating whether the current target belongs to the target category. The regression subnetwork outputs bounding box position parameters, and may also include confidence scores, which output the probability of the corresponding type of target existing in the bounding box. For shallow, medium, and deep fusion features, detection heads with different receptive fields can be used to cover regular-sized targets such as complete documents, bank cards, contract pages, business document blocks, and text region boxes in transaction voucher images. For high-resolution fusion features, a detection head specifically designed for small-sized targets is set up, focusing on outputting the detection results of tiny targets such as document photo areas, signature areas, serial number areas, and partial stamp areas.
[0058] Bounding box parameters can be expressed as center point coordinates, width, and height, or as top-left and bottom-right corner coordinates; classification information can be expressed as credential category, region category, or key region category; confidence score reflects the degree of matching between the candidate box and the true target.
[0059] After each detection branch outputs candidate bounding box information, cross-branch result integration is performed, and confidence screening and duplicate box suppression are applied to candidate bounding boxes at different scales. Duplicate box suppression can employ non-maximum suppression or weighted box fusion, selectively retaining or merging candidate bounding boxes with overlap exceeding a threshold. This threshold can be set between 0.4 and 0.7 depending on the scenario. If the same target is detected by both the regular size branch and the small size branch, the final bounding box is retained based on confidence, scale adaptability, or a fusion strategy. The final output detection results include the target category information, bounding box parameters, and confidence scores for each retained bounding box.
[0060] In one possible embodiment, the detection result may also include the corresponding target's document page number, rotation angle information, or key area identifier to meet the subsequent processing requirements in multi-page contracts and multi-region document scenarios.
[0061] The standard-sized target detection branch utilizes multi-layer fusion features to cover main targets and regional targets of different standard sizes, while the small-sized target detection branch uses high-resolution fusion features to supplement the detection of tiny targets and tiny key regions. The results are then integrated to form a complete detection result. Therefore, under limited computational resources, this application achieves joint detection of standard-sized and small-sized targets in complex transaction document scenarios without simply increasing network depth or width. Instead, it achieves this through multi-scale feature extraction, bidirectional cross-scale fusion, resolution restoration and detail supplementation, and collaborative processing of scale-based detection branches. The output results can directly serve subsequent document recognition and verification processes.
[0062] The target detection method for transaction vouchers provided in this application extracts shallow detail information, mid-level transition information, and deep semantic information from the transaction voucher image in layers. It then completes the bidirectional transfer of semantics and details through bidirectional cross-scale fusion. Furthermore, it utilizes shallow detail features to restore resolution and supplement details in the deep fusion features, enabling high-resolution fusion features to handle the detection of small key regions. Simultaneously, multi-layer fusion features handle the detection of regular-sized targets, thus forming a collaborative detection structure for targets of different scales. Based on this implementation, regular-sized targets such as complete documents, document bodies, and contract pages, as well as small-sized targets such as document photo areas, signature areas, and serial number areas, can obtain appropriate feature inputs and detection branches in a unified detection process. Localization errors and missed detections in complex backgrounds are suppressed, and the overall structure does not rely on significantly increasing network depth and width, making it suitable for real-time transaction voucher detection needs in mobile terminals, self-service terminals, and server-side scenarios.
[0063] In one possible implementation, resolution restoration and detail supplementation are performed on deep fusion features based on shallow detail features to obtain high-resolution fusion features. This includes: processing deep fusion features based on processing units composed of convolutional layers, normalization layers, and activation layers; upsampling the processed deep fusion features to obtain high-resolution features with spatial dimensions consistent with shallow detail features; performing multiple convolution operations on the high-resolution features to obtain enhanced high-resolution features; and fusing the enhanced high-resolution features with shallow detail features to obtain high-resolution fusion features.
[0064] After obtaining the deep fusion features, they are processed by the CBS module to obtain the processed deep fusion features. The CBS module consists of convolutional layers, batch normalization layers, and sigmoid linear unit (SiLU) activation layers. The convolutional layers adjust the channel dimensions of the deep fusion features and extract local spatial features, the batch normalization layers normalize the feature distribution, and the activation layers introduce nonlinear transformations to enhance the expressive power of features related to small targets.
[0065] The processed deep fusion features are upsampled to obtain high-resolution features with the same size as the shallow detail features, such as 80×80 high-resolution features. These high-resolution features then require feature enhancement to fully extract texture and edge detail information, filter noise, and enhance the representation of features related to small targets. Feature enhancement can be achieved through methods such as stacking single-branch serial convolutional units, cross-branch residual convolution fusion, and lightweight bottleneck convolution.
[0066] For example, high-resolution features can be enhanced using the CSP3_1 (Cross Stage Partial Bottleneck with 3 Convolutions and 1 Bottleneck Unit) module. The CSP3_1 module is a two-branch cross-stage local convolutional structure, including a main branch and a residual branch. The main branch performs basic transformations on the input high-resolution features using only a single CBS module. The residual branch sequentially connects multiple CBS modules, each containing one convolution operation, extracting refined detail features layer by layer. The output features of the main branch and the output features of the residual branch are concatenated, and the concatenated result is then integrated by the CBS module to output the enhanced high-resolution features.
[0067] For example, high-resolution features can be enhanced by serially stacking multiple CBS modules to obtain enhanced high-resolution features.
[0068] After obtaining the enhanced high-resolution features, the enhanced high-resolution features can be stitched together with the shallow detail features to obtain high-resolution fused features.
[0069] In this process, the deep fusion features first undergo a basic transformation followed by resolution restoration to align their spatial dimensions with the shallow detail features. Then, multiple convolutions enhance the restored local representation capabilities, incorporating fine-grained information from the shallow detail features into the fusion result. This results in a high-resolution fusion feature that combines high spatial resolution with strong semantic representation capabilities. This high-resolution fusion feature can more fully express the edge and local texture information of small-sized targets and complements features of conventional sizes.
[0070] By adopting the above method, deep fusion features not only restore spatial resolution but also supplement shallow detail information, thus preserving and enhancing the local texture, edge, and contour features on which small-sized transaction voucher targets depend. This improves the target representation capability of high-resolution fusion features and enhances the positioning accuracy and detection stability in complex backgrounds.
[0071] In one possible implementation, multiple convolution operations are performed on the high-resolution features to obtain enhanced high-resolution features, including: inputting the high-resolution features into a first processing branch and a second processing branch respectively; performing multiple convolution operations on the high-resolution features through the first processing branch, and performing a single convolution operation on the high-resolution features through the second processing branch; and performing a single convolution operation after concatenating the outputs of the first processing branch and the second processing branch to obtain enhanced high-resolution features.
[0072] The first processing branch performs deeper convolutional transformations on the high-resolution features. It contains multiple convolutional units, each consisting of a Convolutional layer (Conv), a Batch Normalization (BN) layer, and a SiLU activation layer, to continuously map and refine the input features. The second processing branch performs lighter convolutional transformations on the high-resolution features, containing a single convolutional unit to preserve more original spatial information and form complementary representations. The concatenation process merges the features output from the two processing branches along the channel dimension. Then, a single convolution operation is used to perform channel fusion and feature reshaping on the merged features, resulting in enhanced high-resolution features.
[0073] The spatial resolution or size of the high-resolution feature is consistent with that of the shallow detail feature. After the high-resolution feature is fed into two processing branches, the first processing branch gradually enhances the edge, texture, and local structural information through continuous convolution, while the second processing branch retains the original information response through fewer convolutions. The outputs of the two are concatenated and then integrated by subsequent convolutions, thus forming an enhanced feature that combines semantic and detail representations, i.e., the enhanced high-resolution feature.
[0074] By using two branches with different convolution depths to process the same high-resolution feature in parallel, the deep-level refined information and the shallow-level fidelity information are simultaneously fused. After another convolution to complete compression and recombination, the output enhanced high-resolution feature can be better adapted to the subsequent small-size transaction voucher target detection branch.
[0075] After adopting the above processing method, the detailed information in the high-resolution features can be more fully preserved and enhanced. The convolutional fusion after splicing realizes the unified expression of multi-source features, making the edges, corners and local textures of small-sized targets easier to be captured by the detection network. It can also improve the detection accuracy and feature utilization efficiency in complex voucher scenarios without significantly increasing the computational complexity.
[0076] In one possible implementation, before performing bidirectional cross-scale fusion on shallow detail features, mid-level transition features, and deep semantic features to obtain shallow fused features, mid-level fused features, and deep fused features, the method further includes: performing global average pooling and global max pooling on the feature maps of each level in the shallow detail features, mid-level transition features, and deep semantic features respectively to obtain two types of global feature descriptors; generating channel attention weight vectors based on the two types of global feature descriptors; multiplying the channel attention weight vectors with the feature maps of the corresponding levels channel by channel to obtain channel-calibrated feature maps; performing multi-scale depthwise separable convolution processing on the channel-calibrated feature maps to generate spatial attention weight maps; multiplying the spatial attention weight maps with the channel-calibrated feature maps element by element to obtain spatially calibrated feature maps, and then performing bidirectional cross-scale fusion on the spatially calibrated feature maps of each level. Here, shallow detail features, mid-level transition features, and deep semantic features are feature maps of three levels, with the levels increasing sequentially. Global Average Pooling (GAP) is used to extract the mean response of feature maps at each level, while Global Max Pooling (GMP) is used to extract salient responses. Together, they form a global description of each channel. Channel attention weight vectors are used to adaptively weight different channels, spatial attention weight maps are used to recalibrate different spatial locations, and multi-scale depthwise separable convolutions are used to introduce spatial correlation information of different receptive fields with lower computational cost.
[0077] In practical implementation, global average pooling and global max pooling can be performed on each layer's feature map separately, and the concatenated weights are then input into a shared weight generation network. This weight generation network can consist of fully connected layers, activation layers, and normalization layers to output a channel attention weight vector consistent with the number of channels in the feature map. This weight vector is then multiplied channel-by-channel with the corresponding layer's feature map to enhance the effective channel responses. The channel-calibrated feature map is then input into a multi-scale depthwise separable convolutional layer composed of different kernel sizes to generate a spatial attention weight map, which is then multiplied element-by-element with the channel-calibrated feature map to obtain the spatially calibrated feature map. The feature maps from each layer, processed in the above way, then participate in bidirectional cross-scale fusion to complete cross-layer information exchange.
[0078] After obtaining the three scale feature maps output by the backbone network—shallow detail features, mid-level transition features, and deep semantic features—before inputting them into the neck network for bidirectional scale fusion, a SENetV2 (Squeeze-and-Excitation Network Version 2) attention module can be cascaded between the backbone and neck networks. The SENetV2 attention module is a lightweight channel-spatial joint attention unit that integrates dual-channel global pooling, multi-layer perceptrons, and multi-scale depthwise separable convolutions. The module consists of a four-stage sequential processing flow: compression, activation, channel weighting, and spatial attention enhancement. This flow corresponds entirely to the calibration operations of the aforementioned feature maps at each level. The specific execution logic is as follows:
[0079] 1. Squeeze Stage: The single-layer feature map (shallow detail features, mid-layer transition features, and deep semantic features) output by the backbone network is used as module input. Global average pooling (GAP) and global max pooling (GMP) are performed on the feature map simultaneously to extract two types of global features: global mean feature descriptors and global extreme value feature descriptors. This takes into account both global contextual information and target salient texture information, and makes up for the one-sided feature extraction of a single pooling method.
[0080] 2. Excitation Phase: Channel dependency modeling is completed using an Aggregate Multi-Layer Perceptron (AMP) structure. The two types of global feature descriptors output from the compression phase are concatenated and fused along the channel dimension. The input is then used to perform multi-level nonlinear transformations on multiple parallel fully connected branches to fully explore the correlation weights between different channels. After the output results from the multiple branches are aggregated, they are normalized by the Sigmoid activation function to generate a one-dimensional channel attention weight vector with values in the range [0,1]. The magnitude of the weight values represents the effectiveness of the corresponding channel features for the credential recognition task.
[0081] 3. Channel Weighting (Scale) Stage: The channel attention weight vector output from the excitation stage is multiplied channel by channel with the original hierarchical feature map input to the module, and each channel of the feature map is adaptively scaled: the feature response of the effective channels related to the credential target is enhanced, the invalid channel features such as background and noise are suppressed, the feature calibration of the channel dimension is completed, and the channel-calibrated feature map is output.
[0082] 4. Spatial Attention Enhancement Stage: Following the channel calibration process, a multi-scale depth-separable convolutional layer is followed to perform multi-receptive-field lightweight convolution operations on the channel-calibrated feature map. This convolutional layer is configured with multiple sets of depth-separable convolutional branches with different kernel sizes in parallel, simultaneously extracting spatial texture and edge position information of the target object of different sizes, and aggregating them to generate a two-dimensional spatial attention weight map. The spatial attention weight map is multiplied element-wise with the channel-calibrated feature map to adaptively enhance the pixels of key regions of the target object and weaken irrelevant background pixels, resulting in the final spatially calibrated feature map, which is the output feature of the SENetV2 module.
[0083] After performing channel and spatial calibration on the feature maps of the three levels, namely shallow detail features, mid-level transition features and deep semantic features, the calibrated shallow detail features, calibrated mid-level transition features and calibrated deep semantic features are input into the subsequent neck network to perform bidirectional cross-scale fusion to obtain shallow fusion features, mid-level fusion features and deep fusion features.
[0084] By performing channel calibration first and then spatial calibration, the important channels and key regions of shallow, medium and deep features can be reweighted before entering cross-scale fusion, thereby reducing the impact of background interference on the fusion results and allowing subsequent fused features to retain both detailed and semantic information, thus improving the detection accuracy and feature expression stability of transaction voucher targets in complex scenarios.
[0085] In one possible implementation, target detection of regular-sized transaction documents is performed based on shallow, medium, and deep fusion features, while small-sized target detection of transaction documents is performed based on high-resolution fusion features, yielding detection results including:
[0086] The shallow fusion features, medium fusion features, deep fusion features, and high-resolution fusion features are input into the corresponding detection branches to obtain the target detection results of the corresponding size.
[0087] The detection results of the corresponding size targets output from each detection branch are integrated to obtain the detection result.
[0088] The detection branch is used to classify and predict the input features and regress the bounding box, outputting the category confidence, location coordinates and target score of the target of the corresponding size.
[0089] In practical implementation, shallow, medium, and deep fusion features are input into three regular-sized target detection branches. Each detection branch detects candidate regions on the corresponding scale feature map to output detection results for complete documents, larger area identifiers, and regular-sized text blocks in the transaction certificate. High-resolution fusion features are input into the small-sized target detection branch to locate and classify document corners, number areas, signature areas, and other small key areas. The detection boxes, category labels, and confidence scores output by each branch are uniformly organized by the result aggregation module and integrated into the final detection result after threshold filtering, duplicate box suppression, and coordinate mapping.
[0090] The detection branches corresponding to shallow fusion features, medium fusion features, deep fusion features, and high-resolution fusion features are composed of multiple decoupled detection heads, which are responsible for the target detection of the corresponding features and obtain the detection results of the corresponding detection branches.
[0091] Each detection head can use the same network architecture, for example, it can include a 1×1 convolutional layer for channel dimensionality reduction. After dimensionality reduction, the output is divided into two parallel branches, with two layers of 3×3 convolutional units stacked inside each branch. One branch is a classification branch, which is used to predict the target category and target confidence, and the other branch is a regression branch, which is used to predict the coordinate offset of the target bounding box.
[0092] For example, the detection branch can decouple the detection head from YOLOX (You Only Look Once X-version).
[0093] This method feeds different levels of fused features into their respective detection branches, enabling both regular-sized and small-sized targets to be detected at their appropriate feature scales. The results from each branch are then integrated to ensure that the output covers both large and small targets while maintaining consistency and completeness. This improves detection coverage and localization accuracy, and enhances response to small, critical regions, even in complex backgrounds and with significant differences in target size in transaction document images.
[0094] In one possible implementation, multi-scale feature extraction is performed on the transaction voucher image to be detected, yielding shallow detail features, mid-level transition features, and deep semantic features, including:
[0095] The transaction voucher image is sequentially processed with dimensionality reconstruction, downsampling, convolution transformation, channel normalization, and nonlinear activation to obtain shallow detail features;
[0096] Multiple sets of cross-stage local convolution iterations are performed on the shallow detailed features, and channel integration is performed after each set of convolution iterations to obtain the mid-level transition features;
[0097] Multiple sets of cross-stage local convolution iterations are performed on the intermediate transition features, and channel integration is performed after each set of convolution iterations to obtain the initial deep features;
[0098] Spatial pyramid multi-scale pooling is performed on the initial deep features, and the results of pooling operations at different scales are concatenated to obtain deep semantic features.
[0099] Among these, dimensionality reconstruction is used to adjust the input transaction voucher image into a tensor format suitable for subsequent convolution calculations. For example, if the original input is a two-dimensional RGB image, the dimensionality reconstruction operation converts the image data into a standard four-dimensional tensor [Batch, Channel, Height, Width], unifying the batch, channel, and spatial dimension arrangement rules, matching the input data specifications of the convolutional layer, and ensuring that subsequent convolution and pooling operations can be executed normally. Downsampling is used to compress spatial resolution and expand the feature receptive range. Downsampling can be achieved using CBS with a stride of 2. Convolutional transformation is used to extract local texture, edge, and structural information. Convolutional transformation can be performed using a 3×3 convolutional kernel. The sliding convolutional kernel traverses the feature map to capture local spatial features such as text lines, borders, and patterns in the transaction voucher image, thereby achieving basic visual information extraction. Channel normalization is used to stabilize the feature distribution of each channel, which can be achieved through a batch normalization layer (BN). Nonlinear activation processing is used to enhance feature representation capabilities, which can be achieved through an activation layer (SiLU).
[0100] Cross-stage local convolutional iterative operations are used to repeatedly refine responses between adjacent feature stages. A dual-branch architecture can be adopted, with one branch being a shortcut branch and the other an iterative branch. The shortcut branch only performs a single CBS convolution to complete the basic channel transformation, without performing multi-level convolutional iteration operations, directly passing the original feature information across stages, preserving shallow textures and minute document details to the greatest extent, while providing a direct path for gradients and alleviating the gradient decay problem caused by multi-level iterative convolutions. The iterative branch first compresses the channels through CBS convolution, and then connects multiple sets of Bottleneck units to complete multi-level convolutional iterative operations, mining deep feature responses such as document outlines and text layer by layer to achieve refined feature extraction. After the iterative branch completes multiple rounds of convolutional iterations, the output features of the iterative branch and the output features of the shortcut branch are concatenated and fused along the channel dimension, and then channel integration is completed through a layer of CBS convolution to output the feature map processed in this stage.
[0101] Spatial pyramid multi-scale pooling operation is used to aggregate contextual information at different spatial scales. By splicing the results of pooling operations at different scales, deeper semantic features with more complete semantic expression can be obtained.
[0102] In the specific implementation, the transaction certificate image can first be input into the shallow encoding unit of the feature extraction network. The shallow encoding unit can be composed of convolutional layers, normalization layers, and activation layers in sequence. The convolutional kernel size can be set to 3×3 or 5×5 to balance local detail preservation and feature compression, thus obtaining shallow detail features. After the shallow detail features are formed, they are fed into the cross-stage local convolution module. The module can contain multiple cascaded local convolutional sub-units. After each set of convolutional iterations, channel integration is completed by channel-wise concatenation or channel-wise weighted fusion, and finally, the intermediate transition features are output. The initial deep features are obtained in a similar manner. After the initial deep features are generated, spatial pyramid multi-scale pooling can be processed in parallel using multiple pooling windows of different scales. The pooling windows can include adaptive pooling of different scales or fixed window pooling. Then, the outputs of each scale are concatenated along the channel dimension to form deep semantic features.
[0103] This approach extracts local details layer by layer, enhances semantic expression progressively, and introduces multi-scale context aggregation at a deeper level. This allows shallow, mid-level, and deep features to correspond to texture edges, structural transitions, and overall semantic information, respectively, thus providing a clear and complementary feature foundation for subsequent transaction document target detection. This enables more complete preservation of features in the representation of complete documents, local regions, and small-sized key targets, and improves detection stability and localization accuracy in complex backgrounds.
[0104] In this application, a remote transaction voucher recognition system in the financial field is deployed between a mobile terminal and a back-end recognition service. Users take or upload images of transaction vouchers via mobile phones or tablets. Transaction vouchers include ID cards, bank cards, and business documents. The system performs voucher detection and recognition in real time and returns results within milliseconds. The input image can be a mixed image of an ID card and bank card uploaded by the user through a mobile banking app. This image contains both large and small targets. The main body of the ID card is 200×200 pixels, and the photo is 20×20 pixels. The background of the image consists of desktop textures and lighting shadows, resulting in complex background interference. The ID card exhibits trapezoidal distortion due to the shooting angle and may also be accompanied by edge curvature, creating a situation where multiple scale targets such as the complete document, the photo, and the signature area coexist in the same image.
[0105] In one possible implementation, the transaction certificate image to be detected is input into a feature extraction network. This network employs a lightweight backbone network, CSPDarknet, which sequentially performs dimensionality reconstruction, downsampling, convolutional transformation, channel normalization, and nonlinear activation on the mixed image of the ID card and bank card. The nonlinear activation uses SiLU, outputting shallow detail features. Subsequently, multiple sets of cross-stage local convolutional iterations are performed on these shallow detail features. After each set of convolutional iterations, feature filtering and channel integration operations are performed to obtain intermediate transition features. Then, multiple sets of cross-stage local convolutional iterations are performed on these intermediate transition features, and feature filtering and channel integration operations are performed again after each set of convolutional iterations to form initial deep features. Finally, spatial pyramid multi-scale pooling operations are performed on these initial deep features, and the results of pooling operations at different scales are concatenated to obtain deep semantic features. The feature maps output by the backbone network may include feature maps of different scales, such as 80×80, 40×40, and 20×20, for subsequent upsampling, downsampling, and concatenation.
[0106] Before shallow detail features, mid-level transition features, and deep semantic features are fused across scales, global average pooling (GAP) and global max pooling (GMP) are performed on the feature maps of each level to generate two types of global feature descriptors. These are then combined with an AMP structure (aggregate multilayer perceptron) to learn the dependencies between channels. Channel attention weight vectors are calculated based on the two types of global feature descriptors, and then multiplied channel-wise with the corresponding level feature map to obtain the channel-calibrated feature map. In the mixed image of an ID card and a bank card, the high-frequency feature channels of the ID card edge are enhanced in this process. Subsequently, multi-scale depthwise separable convolution is performed on the channel-calibrated feature map to capture important spatial relationships such as document edges and text regions, generating a spatial attention weight map. This spatial attention weight map is then multiplied element-wise with the channel-calibrated feature map to obtain the spatially calibrated feature map, in which the spatial location of the ID photo area is highlighted, and interference from desktop textures and lighting shadows is suppressed.
[0107] After obtaining the spatially calibrated shallow, mid, and deep feature maps, bidirectional cross-scale fusion is performed on each layer's feature maps. This bidirectional cross-scale fusion uses an FPN+PAN structure for upsampling and stitching, performing multi-scale fusion on the 80×80 and 40×40 feature maps output by the backbone network to output shallow, mid, and deep fused features, thereby enhancing the detection capability for large targets such as ID card subjects and bank card subjects. After obtaining the deep fused features, they are first processed using a processing unit composed of convolutional layers, normalization layers, and activation layers. Then, upsampling is performed on the processed deep fused features to obtain high-resolution features with spatial dimensions consistent with the shallow detail features. The high-resolution features are simultaneously input into the first and second processing branches. The first processing branch performs multiple convolution operations on the high-resolution features, while the second processing branch performs a single convolution operation. After the outputs of the two processing branches are stitched together, a final convolution operation is performed to obtain the enhanced high-resolution features.
[0108] The enhanced high-resolution features are fused with shallow detail features to form a high-resolution fused feature. This high-resolution fused feature utilizes the high-resolution feature map output from the CSP1_1 layer in the shallow feature map of the backbone network. Through the CBL module (Conv+BN+SiLU), the upsampling module, and the stitching module, a 160×160 resolution high-resolution feature map is generated and used as a dedicated detection path for the TOD-Net small object detection enhancement network. This provides a dedicated detection head for the 20×20 pixel small target in ID card photos, preserving edge and detail information in the high-resolution path. Correspondingly, the shallow, mid-level, and deep fused features are input into the corresponding regular-size detection branches to detect regular-size targets in transaction documents. Simultaneously, the high-resolution fused feature is input into the small-size detection branch to detect small-size targets in transaction documents, thus forming a parallel detection structure with four detection heads: 160×160, 80×80, 40×40, and 20×20.
[0109] After each detection branch outputs the target detection results of the corresponding size, the predicted bounding boxes output by the detection heads are optimized by bounding box regression and the results are integrated. During the training phase, the predicted bounding boxes are optimized using CIoU Loss, adding aspect ratio consistency constraints on top of DIoU Loss to constrain the localization accuracy of the bounding boxes of ID cards and bank cards. For ID cards with trapezoidal distortion caused by the shooting angle, the aspect ratio deviation is corrected under the CIoU Loss constraint, making the predicted bounding boxes more closely match the shape of the real target. Sample allocation adopts the SimOTA strategy, dynamically allocating positive and negative samples based on the comprehensive cost matrix of the predicted bounding boxes and the ground truth bounding boxes. During the inference phase, the prediction matrices output by the four detection heads are concatenated in dimensions. The prediction matrices include 160×160×45, 80×80×45, and the prediction matrices corresponding to the other scale detection heads, forming a unified tensor containing information of the predicted bounding boxes at all scales. The output results of each detection branch are then integrated, and overlapping predicted bounding boxes are eliminated by non-maximum suppression (NMS). For overlapping predictions in adjacent areas of ID cards and bank cards, the detection result with the highest confidence is retained to obtain the final detection result of the transaction voucher image.
[0110] In one possible implementation, the system also dynamically adjusts the computational resource allocation strategy based on the complexity of the input image, with the complexity determined by the number of targets and the degree of background interference. When the image background is simple and the number of targets is small, the model reduces the feature map resolution, for example, from 80×80 to 40×40, to reduce computation. When the image contains only large targets, such as only bank cards, the TOD-Net small target detection head is turned off to avoid redundant computation. When the image noise level is low, the number of AMP branches in the SENetV2 module is reduced to lower computational overhead. For mixed images of ID cards and bank cards uploaded by mobile banking apps, since there are simultaneously 200×200 pixel ID card subjects, 20×20 pixel document photos, desktop textures, lighting shadows, and trapezoidal distortion, the system keeps four detection heads (160×160, 80×80, 40×40, and 20×20) in working state and retains channel attention and spatial attention enhancement processing to complete real-time detection and recognition in the remote transaction certificate recognition process.
[0111] Figure 2 The diagram below shows the structural block diagram of the object detection model provided in this application. This object detection model is based on the YOLOX-S (lightweight small YOLOX model) network architecture, integrates the SENetV2 attention mechanism module and the TOD-Net small object detection enhancement network, and optimizes the loss function.
[0112] like Figure 2As shown, the target detection model includes a backbone network, a neck network, and a detection head network. The backbone network adopts a lightweight backbone network based on CSPDarknet, which includes a Focus module, a CBS module (Conv+BN+SiLU), a CSP1_X module, and an SPP bottleneck structure, used to extract multi-level feature maps from the input image, including shallow detail features, mid-level transition features, and deep semantic features.
[0113] The neck network can use FPN (Feature Pyramid Network) and PAN (Path Aggregation Network) structures for multi-scale fusion, and add feature map extraction paths for the detection head corresponding to small-sized targets.
[0114] Based on YOLOX's original three decoupled detection heads, the detection head network adds a fourth detection head specifically for small targets, in order to achieve the detection of small-sized targets.
[0115] See also Figure 2 Between the backbone network and the neck network, there are three SENetV2 attention modules to perform channel and spatial calibration on the extracted shallow detail features, mid-level transition features and deep semantic features, respectively.
[0116] In the neck network, in addition to the modules that perform bidirectional fusion of shallow detail features, mid-level transition features and deep semantic features through FPN and PAN structures, a new processing path TOD-Net consisting of CBL module, Upsample module, CSP3_1 module and Concat module is added. This path performs multi-level fusion of shallow detail features and deep fusion features to generate high-resolution fusion features, such as feature maps with a size of 160×160.
[0117] In the detection head network, to adapt to high-resolution fusion features, a detection branch similar to but independent of the YOLOX decoupled detection head is designed. It consists of a 1×1 convolutional dimensionality reduction followed by two parallel 3×3 convolutional branches, which predict the target category, confidence score, and bounding box coordinates, respectively. This detection branch or detection head is specifically responsible for detecting small targets with a size less than 20×20.
[0118] After obtaining the outputs from the four detection heads, the data are spliced and integrated to form a unified tensor containing the detection results at each scale, thus obtaining the final detection result.
[0119] During the training phase, the final detection result can be used for subsequent loss calculations. During the inference phase, the final detection result can be used for non-maximum suppression (NMS).
[0120] Non-maximum suppression is a post-detection processing algorithm used to eliminate numerous overlapping and duplicate predicted boxes generated for the same target, retaining only the single valid detection box with the highest confidence. During model training, the loss function used is bounding box regression loss, with CIoU (Complete Intersection over Union) loss replacing the original YOLOX-s loss function. CIoU loss, based on DIoU (Distance Intersection over Union) loss, adds consideration to the aspect ratio consistency between the predicted box (the bounding box containing the target output by the model) and the ground truth labeled box. Its calculation formula is as follows:
[0121]
[0122]
[0123]
[0124] in, Represents the prediction box With the actual annotation box The ratio of the overlapping area to the total area of the two is [0,1]. It represents the square of the Euclidean distance between the center coordinates of the predicted bounding box and the center coordinates of the actual labeled bounding box, and characterizes the degree of offset between the two box centers. It is the diagonal length of the smallest bounding rectangle that can simultaneously enclose the predicted bounding box and the ground truth bounding box. and These represent the width and height of the actual annotation box, respectively. and These represent the width and height of the prediction box, respectively.
[0125] CIoU loss can simultaneously minimize the center point distance and maximize the overlapping area while optimizing the aspect ratio, making bounding box regression more accurate and convergent more widely.
[0126] During model training, data preprocessing and enhancement can be performed first. For example, image samples can be uniformly scaled to 640×640 and enhancement strategies such as Mosaic, MixUp, random cropping, and color dithering can be used to improve the model's generalization ability.
[0127] The AdamW optimizer can be used to update the gradient of the model network parameters, and the initial learning rate can be set to 0.001. The AdamW optimizer implements weight decay regularization separately on the basis of the traditional Adam optimizer, which can effectively suppress the explosion of network parameter weights and is suitable for the complex network structure of this model, including multi-layer CSP cross-stage convolution, SENetV2 attention, and multi-branch detection heads.
[0128] During the training phase, a cosine annealing learning rate scheduling strategy is employed. After each complete training epoch, the learning rate is scheduled according to the formula. Dynamically update the learning rate ,in, Indicates the number of training rounds and the decay coefficient. The value can be 0.97; the learning slope is gradually reduced as the number of training rounds increases. In the early stage of training, a large learning rate quickly converges the basic features, and in the later stage of training, a small learning rate finely adjusts the bounding boxes and classification weights to ensure smooth convergence of the model training process and avoid gradient oscillation in the later stage of training.
[0129] The model is built and trained based on the PyTorch deep learning framework, and performs 150 rounds of training on a self-built transaction certificate annotation dataset. The training batch size is adaptively adjusted according to the local GPU memory capacity to make full use of hardware computing power to improve training efficiency. The training process introduces an early stopping method to monitor the recognition accuracy of the validation set. When the validation set metric does not improve for several consecutive rounds, the training is automatically terminated to limit the network from overfitting invalid features such as blank backgrounds and noise on the training set, thus ensuring the model's recognition performance in real business scenarios.
[0130] By introducing the TOD-Net small object detection enhancement network, a dedicated high-resolution detection path is provided for small objects. Experiments show that this improvement improves the model's mean accuracy (mAP) for small objects by approximately 6.0% compared to the baseline YOLOX-S model. In real-world mixed transaction document image detection tasks, the improved model significantly reduces the false negative rate for small-sized objects (such as bank cards and ID photos) that are overlapping or occluded.
[0131] After integrating the SENetV2 attention mechanism, the model can more accurately focus on the key feature regions and channels of the credential target, effectively suppressing interference from complex backgrounds. Ablation experiments show that adding only the SENetV2 module can improve precision (P) and recall (R) while reducing the number of model parameters from 8.86M to 7.21M and the computational cost (GFLOPs) from 27.29 to 22.14. This indicates that the attention mechanism not only improves feature quality but also achieves model lightweighting through feature selection. This application significantly improves detection accuracy while increasing inference speed. The M-YOLOX model, which integrates SENetV2 and TOD-Net, achieves an FPS (frames per second) of 42.33, compared to the baseline YOLOX-S's 26.83, representing a speed improvement of approximately 59%. This is mainly due to SENetV2's efficient feature selection, which reduces redundant computation, and the careful design of the overall network structure, which avoids complexity explosion. This characteristic makes this recognition algorithm very suitable for deployment in remote transaction credential recognition scenarios.
[0132] Using CIoU Loss as the regression loss function results in faster model convergence during training, a smoother decrease in the validation set loss curve, and the achievement of a lower value. Simultaneously, CIoU's constraint on the aspect ratio of the bounding box ensures a higher degree of alignment between the predicted bounding box and the actual transaction voucher target, reducing localization bias.
[0133] Specific optimizations were made to address the characteristics of transaction vouchers, which are characterized by "multiple types, diverse forms, and frequent stacking and occlusion." In real-world testing, the M-YOLOX model demonstrated stronger robustness and higher recognition confidence compared to models such as YOLOv5 and the original YOLOX. It effectively copes with challenges such as changes in angle and distance, and target occlusion, providing reliable core algorithmic support for improving the initial screening efficiency of remote transaction vouchers in the financial sector.
[0134] In summary, this application provides a high-precision, fast-speed, and robust target detection scheme for transaction certificates through innovative network structure improvements and loss function design.
[0135] Figure 3 This application provides a schematic diagram of the structure of a target detection device for transaction certificates, as shown below. Figure 3 As shown, the target detection device for the transaction certificate includes:
[0136] The multi-level feature extraction module extracts multi-scale features from the transaction voucher image to be detected, obtaining shallow detail features, mid-level transition features, and deep semantic features. The feature fusion module performs bidirectional cross-scale fusion of shallow detail features, mid-level transition features, and deep semantic features, obtaining shallow fused features, mid-level fused features, and deep fused features. The feature enhancement module restores resolution and supplements details of deep fused features based on shallow detail features, obtaining high-resolution fused features. The target detection module performs regular-sized target detection on the transaction voucher based on shallow, mid-level, and deep fused features, and performs small-sized target detection on the transaction voucher based on high-resolution fused features, obtaining detection results.
[0137] The multi-level feature extraction module simultaneously preserves edge texture, structural contours, and deep semantic information in the transaction document image, ensuring that both the complete document and key local areas receive corresponding scale representations. A bidirectional cross-scale fusion module then performs bidirectional fusion, achieving mutual complementarity between shallow details and deep semantics, reducing interference from complex backgrounds, shadows, reflections, and shooting distortions on localization. Furthermore, the feature enhancement module utilizes shallow detail features to restore resolution and supplement details in the deep fused features, recovering the texture and boundary information required for small targets, thereby improving the detection capability for document photos, signature areas, and serial number areas. Therefore, the target detection module can detect both regular-sized and small-sized targets, balancing detection accuracy and processing efficiency while reducing missed and false detections, improving real-time adaptability in remote transaction document scenarios.
[0138] In one possible embodiment, the feature fusion module is specifically used for: processing deep fusion features based on a processing unit composed of convolutional layers, normalization layers, and activation layers; upsampling the processed deep fusion features to obtain high-resolution features with spatial dimensions consistent with shallow detail features; performing multiple convolution operations on the high-resolution features to obtain enhanced high-resolution features; and fusing the enhanced high-resolution features with shallow detail features to obtain high-resolution fusion features.
[0139] In one possible embodiment, the feature enhancement module is specifically configured to: input high-resolution features into a first processing branch and a second processing branch respectively; perform multiple convolution operations on the high-resolution features through the first processing branch and perform a single convolution operation on the high-resolution features through the second processing branch; and after concatenating the outputs of the first processing branch and the second processing branch, perform a single convolution operation to obtain the enhanced high-resolution features.
[0140] In one possible embodiment, the device further includes a two-dimensional calibration module, configured to: before performing bidirectional cross-scale fusion on shallow detail features, mid-level transition features, and deep semantic features to obtain shallow fused features, mid-level fused features, and deep fused features, perform global average pooling and global max pooling on the feature maps of each level in the shallow detail features, mid-level transition features, and deep semantic features respectively to obtain two types of global feature descriptors; generate channel attention weight vectors based on the two types of global feature descriptors; multiply the channel attention weight vectors with the feature maps of the corresponding levels channel by channel to obtain channel-calibrated feature maps; perform multi-scale depthwise separable convolution processing on the channel-calibrated feature maps to generate spatial attention weight maps; multiply the spatial attention weight maps with the channel-calibrated feature maps element by element to obtain spatially calibrated feature maps, so as to perform bidirectional cross-scale fusion on the spatially calibrated feature maps of each level.
[0141] In one possible embodiment, the target detection module is specifically used to: input shallow fusion features, mid-layer fusion features, deep fusion features and high-resolution fusion features into the corresponding detection branches respectively to obtain target detection results of corresponding sizes; and integrate the target detection results of corresponding sizes output by each detection branch to obtain the detection result.
[0142] In one possible embodiment, the multi-level feature extraction module is specifically used to: sequentially perform dimensional reconstruction, downsampling, convolutional transformation, channel normalization, and nonlinear activation processing on the transaction certificate image to obtain shallow detail features; perform multiple sets of cross-stage local convolutional iteration operations on the shallow detail features, and perform channel integration operations after each set of convolutional iteration operations to obtain intermediate transition features; perform multiple sets of cross-stage local convolutional iteration operations on the intermediate transition features, and perform channel integration operations after each set of convolutional iteration operations to obtain initial deep features; perform spatial pyramid multi-scale pooling operations on the initial deep features, and concatenate the results of pooling operations at different scales to obtain deep semantic features.
[0143] The target detection device for transaction certificates provided in this embodiment can execute the target detection method for transaction certificates provided in the above method embodiment. Its implementation principle and technical effect are similar, and will not be described in detail here.
[0144] Figure 4 A schematic diagram of the structure of the electronic device provided in this application. Figure 4 As shown, the electronic device 40 provided in this embodiment includes at least one processor 401 and a memory 402. Optionally, the electronic device 40 further includes a communication component 403. The processor 401, memory 402, and communication component 403 are connected via a bus 404.
[0145] In a specific implementation, at least one processor 401 executes computer execution instructions stored in memory 402, causing at least one processor 401 to perform the above-described method.
[0146] The specific implementation process of processor 401 can be found in the above method embodiments, and its implementation principle and technical effect are similar. It will not be repeated here.
[0147] In the above embodiments, it should be understood that the processor can be a Central Processing Unit (CPU), or other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), etc. The general-purpose processor can be a microprocessor or any conventional processor. The steps of the method disclosed in this invention can be directly implemented by a hardware processor, or implemented by a combination of hardware and software modules within the processor.
[0148] The memory may include random access memory (RAM) and may also include non-volatile memory (NVM), such as at least one disk storage device.
[0149] The bus can be an Industry Standard Architecture (ISA) bus, a Peripheral Component Interconnect (PCI) bus, or an Extended Industry Standard Architecture (EISA) bus, etc. Buses can be categorized as address buses, data buses, control buses, etc. For ease of illustration, the buses shown in the accompanying drawings are not limited to a single bus or a single type of bus.
[0150] This application also provides a computer program product, including a computer program that, when executed by a processor, implements the above-described method.
[0151] This application also provides a computer-readable storage medium storing computer-executable instructions, which, when executed by a processor, implement the above-described method.
[0152] The aforementioned readable storage medium can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic storage, flash memory, magnetic disk, or optical disk. The readable storage medium can be any available medium accessible to a general-purpose or special-purpose computer.
[0153] An exemplary readable storage medium is coupled to a processor, enabling the processor to read information from and write information to the readable storage medium. Of course, the readable storage medium can also be a component of the processor. The processor and the readable storage medium can reside in an Application Specific Integrated Circuit (ASIC). Alternatively, the processor and the readable storage medium can exist as discrete components in the device.
[0154] The division of units is merely a logical functional division; in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be indirect coupling or communication connection through some interfaces, devices, or units, and may be electrical, mechanical, or other forms.
[0155] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
[0156] In addition, the functional units in the various embodiments of the present invention can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit.
[0157] If a function is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this invention, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods of the various embodiments of this invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0158] Those skilled in the art will understand that all or part of the steps of the above-described method embodiments can be implemented by hardware related to program instructions. The aforementioned program can be stored in a computer-readable storage medium. When executed, the program performs the steps of the above-described method embodiments; and the aforementioned storage medium includes various media capable of storing program code, such as ROM, RAM, magnetic disks, or optical disks.
[0159] It should be noted that, for the sake of simplicity, the foregoing method embodiments are all described as a series of actions. However, those skilled in the art should understand that this application is not limited to the described order of actions, as some steps may be performed in other orders or simultaneously according to this application. Furthermore, those skilled in the art should also understand that the embodiments described in the specification are all optional embodiments, and the actions and modules involved are not necessarily essential to this application.
[0160] It should be further noted that although the steps in the flowchart are shown sequentially according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless explicitly stated herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least some steps in the flowchart may include multiple sub-steps or multiple stages. These sub-steps or stages are not necessarily completed at the same time, but can be executed at different times. The execution order of these sub-steps or stages is not necessarily sequential, but can be performed alternately or in turn with other steps or at least some of the sub-steps or stages of other steps.
[0161] It should be understood that the above-described device embodiments are merely illustrative, and the device of this application can also be implemented in other ways. For example, the division of units / modules in the above embodiments is only a logical functional division, and there may be other division methods in actual implementation. For example, multiple units, modules, or components may be combined, or integrated into another system, or some features may be ignored or not executed.
[0162] Furthermore, unless otherwise specified, the functional units / modules in the various embodiments of this application can be integrated into one unit / module, or each unit / module can exist physically separately, or two or more units / modules can be integrated together. The integrated units / modules described above can be implemented in hardware or as software program modules.
[0163] When an integrated unit / module is implemented in hardware, the hardware can be digital circuits, analog circuits, etc. The physical implementation of the hardware structure includes, but is not limited to, transistors, memristors, etc. Unless otherwise specified, the processor can be any suitable hardware processor, such as a CPU, GPU, FPGA, DSP, and ASIC, etc. Unless otherwise specified, the storage unit can be any suitable magnetic or magneto-optical storage medium, such as Resistive Random Access Memory (RRAM), Dynamic Random Access Memory (DRAM), Static Random-Access Memory (SRAM), Enhanced Dynamic Random Access Memory (EDRAM), High-Bandwidth Memory (HBM), Hybrid Memory Cube (HMC), etc.
[0164] If the integrated unit / module is implemented as a software program module and sold or used as an independent product, it can be stored in a computer-readable storage device (CMD). Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a memory and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods of the various embodiments of this application. The aforementioned memory includes various media capable of storing program code, such as a USB flash drive, read-only memory (ROM), random access memory (RAM), portable hard drive, magnetic disk, or optical disk.
[0165] In the above embodiments, the descriptions of each embodiment have their own emphasis. For parts not described in detail in a certain embodiment, please refer to the relevant descriptions of other embodiments. The technical features of the above embodiments can be combined arbitrarily. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as the combination of these technical features does not contradict each other, it should be considered within the scope of this specification.
[0166] Other embodiments of this application will readily occur to those skilled in the art upon consideration of the specification and practice of the invention disclosed herein. This application is intended to cover any variations, uses, or adaptations of this application that follow the general principles of this application and include common knowledge or customary techniques in the art not disclosed herein.
[0167] It should be understood that this application is not limited to the precise structure described above and shown in the accompanying drawings, and various modifications and changes can be made without departing from its scope.
Claims
1. A target detection method for transaction certificates, characterized in that, include: Multi-scale feature extraction is performed on the transaction voucher image to be detected to obtain shallow detail features, mid-level transition features, and deep semantic features; A bidirectional cross-scale fusion is performed on the shallow detail features, mid-level transition features, and deep semantic features to obtain shallow fused features, mid-level fused features, and deep fused features; Based on the shallow detail features, the deep fusion features are restored in resolution and supplemented in detail to obtain high-resolution fusion features; Based on the shallow, medium, and deep fusion features, the detection of regular-sized targets on transaction vouchers is performed, and based on the high-resolution fusion features, the detection of small-sized targets on transaction vouchers is performed, resulting in detection results.
2. The method according to claim 1, characterized in that, The process of restoring resolution and supplementing details on the deep fused features based on the shallow detailed features to obtain high-resolution fused features includes: The deep fusion features are processed using a processing unit composed of convolutional layers, normalization layers, and activation layers. Upsampling is performed on the processed deep fusion features to obtain high-resolution features with spatial dimensions consistent with the shallow detail features; Perform multiple convolution operations on the high-resolution features to obtain enhanced high-resolution features; The enhanced high-resolution features are fused with the shallow detail features to obtain the high-resolution fused features.
3. The method according to claim 2, characterized in that, The process of performing multiple convolution operations on the high-resolution features to obtain enhanced high-resolution features includes: The high-resolution features are input into the first processing branch and the second processing branch, respectively; The high-resolution features are subjected to multiple convolution operations through the first processing branch, and a single convolution operation is performed on the high-resolution features through the second processing branch. After concatenating the outputs of the first processing branch and the second processing branch, a convolution operation is performed to obtain the enhanced high-resolution features.
4. The method according to claim 1, characterized in that, Before performing bidirectional cross-scale fusion on the shallow detail features, mid-level transition features, and deep semantic features to obtain shallow fused features, mid-level fused features, and deep fused features, the method further includes: For the feature maps of each level in the shallow detail features, mid-level transition features and deep semantic features, global average pooling and global max pooling are performed respectively to obtain two types of global feature descriptors; Based on two types of global feature descriptors, channel attention weight vectors are generated; The channel attention weight vector is multiplied channel by channel with the feature map of the corresponding level to obtain the channel-calibrated feature map; Multi-scale depthwise separable convolution is performed on the channel-calibrated feature map to generate a spatial attention weight map. The spatial attention weight map is multiplied element-wise with the channel-calibrated feature map to obtain the spatially calibrated feature map, and then bidirectional cross-scale fusion is performed on the spatially calibrated feature maps at each level.
5. The method according to any one of claims 1-4, characterized in that, The detection results obtained by performing regular-sized target detection on transaction documents based on the shallow, medium, and deep fusion features, and by performing small-sized target detection on transaction documents based on the high-resolution fusion features, include: The shallow fusion feature, the medium fusion feature, the deep fusion feature, and the high-resolution fusion feature are input into the corresponding detection branches to obtain the target detection results of the corresponding size. The detection results are obtained by integrating the target detection results of the corresponding size output from each of the detection branches.
6. The method according to any one of claims 1-4, characterized in that, The process involves multi-scale feature extraction of the transaction voucher image to be detected, yielding shallow detail features, mid-level transition features, and deep semantic features, including: The transaction voucher image is sequentially subjected to dimensional reconstruction, downsampling, convolution transformation, channel normalization, and nonlinear activation processing to obtain the shallow detail features; Multiple sets of cross-stage local convolution iterations are performed on the shallow detail features, and channel integration is performed after each set of convolution iterations to obtain the middle transition features; Multiple sets of cross-stage local convolution iterations are performed on the intermediate transition features, and channel integration is performed after each set of convolution iterations to obtain the initial deep features. The initial deep features are subjected to spatial pyramid multi-scale pooling operations, and the results of pooling operations at different scales are concatenated to obtain the deep semantic features.
7. A target detection device for transaction vouchers, characterized in that, include: The multi-level feature extraction module is used to extract multi-scale features from the transaction voucher image to be detected, obtaining shallow detail features, mid-level transition features, and deep semantic features; The feature fusion module is used to perform bidirectional cross-scale fusion on the shallow detail features, mid-level transition features and deep semantic features to obtain shallow fused features, mid-level fused features and deep fused features; The feature enhancement module is used to restore the resolution and supplement the details of the deep fusion features based on the shallow detail features, so as to obtain high-resolution fusion features; The target detection module is used to perform regular-sized target detection on transaction vouchers based on the shallow fusion features, medium-sized fusion features and deep fusion features, and to perform small-sized target detection on transaction vouchers based on the high-resolution fusion features, so as to obtain the detection results.
8. An electronic device, characterized in that, include: A processor, and a memory communicatively connected to the processor; The memory stores computer-executed instructions; The processor executes computer execution instructions stored in the memory to implement the method as described in any one of claims 1 to 6.
9. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer-executable instructions, which, when executed by a processor, are used to implement the method as described in any one of claims 1 to 6.
10. A computer program product, characterized in that, Includes a computer program that, when executed by a processor, implements the method of any one of claims 1 to 6.