Bill anti-counterfeiting detection method, device and equipment and storage medium
By employing a target detection model that combines resolution feature enhancement, multimodal feature fusion, and feature pooling kernel processing, the problems of missed detection and false detection in invoice anti-counterfeiting detection have been solved, achieving efficient and accurate identification of genuine and counterfeit marks.
Patent Information
- Application Number
- CN202511067871.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-31
- Publication Date
- 2025-11-07
AI Technical Summary
Existing target detection algorithms suffer from problems such as missing small-scale authenticity marker areas, high false detection rate, and computational redundancy in anti-counterfeiting detection of invoices, making it difficult to efficiently and accurately identify authenticity markers of invoices in complex backgrounds.
By using a trained target detection model, resolution feature enhancement, multimodal feature fusion, and feature pooling kernel adjustment are performed to identify the authenticity marking regions in the ticket image and compare them with the benchmark marking content to optimize feature capture capability and computational efficiency.
It improves the ability to capture features of small-scale genuine and fake marker regions, reduces the false detection rate, and improves recognition efficiency and accuracy, meeting the high precision and high efficiency requirements in complex scenarios.
Smart Images

Figure CN120913045A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of image processing, and in particular to a bill anti-counterfeiting detection method, device, equipment and storage medium. BACKGROUND
[0002] Bill anti-counterfeiting detection is a core technology in the field of financial security, directly affecting the efficiency of anti-fraud, risk control and automated business processes.
[0003] Traditional manual identification relies on experience, and has problems such as low efficiency, high cost, and being easily affected by subjective factors to reduce the accuracy of anti-counterfeiting detection. In recent years, although deep learning-based target detection algorithms have made progress in general scenarios, they still face severe challenges in bill detection in complex scenarios: first, the size of the bill anti-counterfeiting mark is small, and existing target detection algorithms are prone to missed detection due to low feature map resolution; second, the complex background of the bill, such as the background, decorative patterns, etc. is easy to be confused with the anti-counterfeiting features, especially under the interference of dynamic optical features such as optically variable ink and holographic patterns, resulting in a significant increase in the false detection rate of existing target detection algorithms; finally, the pooling process in existing target detection algorithms in high-resolution image processing is prone to redundant calculation, reducing the efficiency of bill anti-counterfeiting detection. SUMMARY
[0004] The present application provides a bill anti-counterfeiting detection method, device, equipment and storage medium to improve the accuracy and efficiency of bill anti-counterfeiting detection.
[0005] In a first aspect, an embodiment of the present application provides a bill anti-counterfeiting detection method, which comprises:
[0006] obtaining a bill image of a target bill;
[0007] identifying a true-false mark region in the bill image based on resolution feature enhancement, multi-modal feature fusion and adjustment of feature pooling kernels on the bill image through a trained target detection model;
[0008] comparing the region content of the true-false mark region with the set reference mark content, and determining the true-false detection result of the target bill according to the comparison result.
[0009] In a second aspect, an embodiment of the present application provides a bill anti-counterfeiting detection device, which comprises:
[0010] an image acquisition module for acquiring a bill image of a target bill;
[0011] a region identification module for identifying a true-false mark region in the bill image based on resolution feature enhancement, multi-modal feature fusion and adjustment of feature pooling kernels on the bill image through a trained target detection model;
[0012] The authenticity detection module is configured to compare the region content of the authenticity mark region with a set reference mark content, and determine an authenticity detection result of the target bill according to a comparison result.
[0013] In a third aspect, an electronic device is provided, and the electronic device comprises:
[0014] at least one processor;
[0015] and a memory in communication with the at least one processor;
[0016] The memory stores a computer program executable by the at least one processor, and the computer program is executed by the at least one processor to enable the at least one processor to perform the bill anti-counterfeiting detection method according to any one of the embodiments of the present application.
[0017] In a fourth aspect, a computer readable storage medium is provided, and the computer readable storage medium stores computer instructions for enabling a processor to perform the bill anti-counterfeiting detection method according to any one of the embodiments of the present application.
[0018] The technical scheme of the embodiments of the present application comprises the following steps: obtaining a bill image of a target bill; identifying an authenticity mark region in the bill image based on resolution feature enhancement, multi-modal feature fusion and adjustment of a feature pooling kernel through a trained target detection model; comparing the region content of the authenticity mark region with a set reference mark content, and determining an authenticity detection result of the target bill according to a comparison result. By using this method, the resolution feature enhancement processing of the bill image based on the target detection model improves the feature capturing capability for small-scale authenticity mark regions, and the multi-modal feature fusion processing of the bill image fuses the features related to optics, enhances the sensitivity of the model to dynamic anti-counterfeiting features, and reduces the false detection rate of the authenticity mark region. In addition, based on the adjustment of the feature pooling kernel, a suitable pooling kernel can be dynamically selected based on the size difference of the authenticity mark region in the target bill, so as to reduce the calculation redundancy while maintaining the multi-scale detection capability of the bill image, improve the efficiency of identifying the authenticity mark region, ensure the real-time detection, and meet the high-precision and high-efficiency requirements of bill anti-counterfeiting detection in complex scenarios.
[0019] It should be understood that the content described in this part is not intended to identify key or important features of the embodiments of the present application, nor is it used to limit the scope of the present application. Other features of the present application will become apparent from the following description. BRIEF DESCRIPTION OF DRAWINGS
[0020] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings needed in the embodiment description. Obviously, the drawings in the following description are only some embodiments of the present application, and for those skilled in the art, other drawings can be obtained without creative effort based on these drawings.
[0021] Figure 1 A flow chart of a bill anti-counterfeiting detection method provided by the embodiment of the present application;
[0022] Figure 2 A structural schematic diagram of a bill anti-counterfeiting detection device provided by the embodiment of the present application;
[0023] Figure 3 A structural schematic diagram of an electronic device that can be used to implement the embodiment of the present application is shown. DETAILED DESCRIPTION
[0024] In order for those skilled in the art to better understand the present application, the following will clearly and completely describe the technical solutions in the embodiments of the present application with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only some embodiments of the present application, not all. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative effort should be within the scope of the present application.
[0025] It should be noted that the terms "first", "second", etc. in the specification and claims of the present application and the above-mentioned drawings are used to distinguish similar objects, and do not necessarily describe a specific order or sequence. It should be understood that the data thus used can be interchanged under appropriate circumstances, so that the embodiments of the present application described herein can be implemented in an order other than that illustrated or described herein. In addition, the terms "include" and "have" and any variations thereof are intended to cover non-exclusive inclusion, for example, a process, method, system, product or device that includes a series of steps or units does not necessarily limit to those steps or units clearly listed, but can include other steps or units not clearly listed or inherent to these processes, methods, products or devices.
[0026] The embodiment of the present application provides a bill anti-counterfeiting detection method, Figure 1 A flow chart of a bill anti-counterfeiting detection method provided by the embodiment of the present application, the embodiment of the present application can be applied to the scene of detecting the anti-counterfeiting mark on the bill to identify the authenticity of the bill. The method can be executed by a bill anti-counterfeiting detection device, which can be realized in the form of software and / or hardware, and can be realized by an electronic device, which is preferably a mobile terminal, a desktop computer, a notebook computer, a server, etc.
[0027] As Figure 1 shown, the bill anti-counterfeiting detection method provided by the embodiment of the application can specifically include:
[0028] S101, obtaining a bill image of a target bill.
[0029] The target bill can be understood as a bill that has anti-counterfeiting detection needs, and the target bill includes a true-false mark region with an anti-counterfeiting mark. For example, the anti-counterfeiting mark can be microtext, watermark, seal, and anti-counterfeiting barcode, and the microtext can be printed text or handwritten signature, and the embodiment is not limited specifically. The bill image can be considered as an image containing the bill, and the bill can be a paper bill or an electronic bill, and the electronic bill can be in PDF format, JPG format, or png format, etc.
[0030] In this embodiment, the bill image of the target bill can be obtained through an image acquisition device, or the bill image of the target bill uploaded by the user through the editing interface provided by the client through clicking, dragging, etc. through the editing interface, or the storage path input by the user on the editing interface can be obtained through the editing interface to obtain the bill image of the target bill under the storage path.
[0031] S102, through the trained target detection model, based on the resolution feature enhancement, multi-modal feature fusion and adjustment of the feature pooling kernel processing on the bill image, the true-false mark region in the bill image is identified.
[0032] The target detection model can be understood as a model for detecting and determining the true-false mark region in the bill image, and the true-false mark region can be considered as a region for true-false judgment, which can be the region where the anti-counterfeiting mark is located. For example, the rectangular bounding box region of the anti-counterfeiting mark can be determined as the true-false mark region.
[0033] It can be understood that the true-false mark region in the target bill usually occupies a very small region of the bill image, such as 0.1%-1% of the pixels. The existing target detection network has a too high down-sampling rate in the deep layer, resulting in insufficient resolution of the feature map, and the small true-false mark region gradually loses the detailed information in the feature extraction process, which increases the small target true-false mark region missing detection rate. Therefore, in this embodiment, the bill image can be input into the trained target detection model, and the bill image can be processed by the target detection model for resolution feature enhancement, so as to improve the feature capture ability of the small scale anti-counterfeiting mark and avoid the missing detection problem caused by insufficient feature map resolution in the feature extraction process of the bill image.
[0034] Secondly, the anti-counterfeit marks in the real and fake mark region, such as optically variable ink and holographic pattern, will dynamically change with the change of the observation angle. The traditional visual attention mechanism only relies on visual texture information, and it is difficult to distinguish the real anti-counterfeit features from the background reflection noise points. Therefore, in the embodiment, the target detection model can be used to extract the optical correlation features of the bill image, and a multi-modal feature fusion processing is performed. For example, the texture features, color features and optical reflection features of the bill image can be fused to enhance the sensitivity of the target detection model to the dynamic anti-counterfeit features.
[0035] In addition, the feature scales of different regions of the target bill are significantly different. For example, the security line is a long strip-shaped large target, and the microtext is a dense small target. It is difficult for a fixed size of a pooling kernel to balance efficiency and accuracy. Therefore, in the embodiment, the target detection model can also be used to adjust the feature pooling kernel of the bill image, to adaptively adjust the size of the pooling kernel according to the local features of the bill image, and to determine the number of the pooling kernels and the connection relationship between the pooling kernels, so as to reduce the calculation redundancy while maintaining the multi-scale detection capability of the bill image, thereby identifying the real and fake mark regions in the bill image.
[0036] For example, the identified real and fake mark regions can be marked in the bill image in the form of different colors or bounding boxes; or the position information of the real and fake mark regions in the bill image can be output. For example, if the identified real and fake mark region is a rectangular region, the pixel coordinates of the upper left corner of the rectangular region, and the width and height of the rectangular region can be output as the position information.
[0037] S103, comparing the region content of the real and fake mark region with the set reference mark content, and determining the authenticity detection result of the target bill according to the comparison result.
[0038] The region content can be understood as the content contained in the real and fake mark region for serving as the basis for authenticity judgment. For example, the region content can be the content of the anti-counterfeit mark in the real and fake mark region. The reference mark content can be understood as the correct mark content that should be possessed in the target bill, which is used as a reference content for the region content to determine the authenticity of the region content.
[0039] In the embodiment, the reference mark content can be set by the bill issuer or the relevant institution in advance. By comparing the region content of the real and fake mark region with the set reference mark content, the comparison result is obtained. If the comparison result represents that the region content is consistent with the reference mark content, it can be determined that the authenticity detection result of the target bill is true, that is, the target bill is a real bill; if the comparison result represents that the region content is inconsistent with the reference mark content, it can be determined that the authenticity detection result of the target bill is false, that is, the target bill is a counterfeit bill.
[0040] The technical scheme of the embodiment is described as follows: a bill image of a target bill is acquired; a trained target detection model is used to identify a true-false mark region in the bill image based on resolution feature enhancement, multi-modal feature fusion and adjustment of a feature pooling kernel; and the region content of the true-false mark region is compared with set reference mark content, and a true-false detection result of the target bill is determined according to the comparison result. With this method, the resolution feature enhancement processing of the target detection model on the bill image improves the feature capturing capability for small-scale true-false mark regions, and the multi-modal feature fusion processing of the bill image fuses the features related to optics, enhances the sensitivity of the model to dynamic anti-false features, and reduces the false detection rate of the true-false mark region. In addition, based on the adjustment of the feature pooling kernel, a suitable pooling kernel can be dynamically selected based on the size difference of the true-false mark region in the target bill, so as to reduce the calculation redundancy while maintaining the multi-scale detection capability of the bill image, improve the efficiency of identifying the true-false mark region, ensure the real-time detection, and meet the high-precision and high-efficiency requirements of bill anti-false detection in complex scenes.
[0041] As a first optional embodiment of the embodiment of the application, on the basis of the above-mentioned embodiment, the target detection model comprises a feature extraction sub-model, an attention perception sub-model, a feature pooling sub-model and a region prediction sub-model.
[0042] Correspondingly, the step of identifying the true-false mark region in the bill image based on the resolution feature enhancement, multi-modal feature fusion and adjustment of the feature pooling kernel of the trained target detection model can be further optimized as follows:
[0043] a1) a first feature map is obtained by performing feature extraction on the bill image according to a set resolution granularity through the feature extraction sub-model including a semantic feature extraction module and a resolution feature enhancement module.
[0044] The feature extraction sub-model can be understood as a model for extracting features of the bill image. The feature extraction sub-model at least includes a semantic feature extraction module and a resolution feature enhancement module. The semantic feature extraction module can be used to extract semantic features in the bill image, and the resolution feature enhancement module can be used to enhance the features of small targets in the bill image, which can include the true-false mark region. The first feature map can be understood as a feature map obtained by performing feature extraction on the bill image by the feature extraction sub-model, which retains high-resolution features.
[0045] In this embodiment, the semantic feature extraction module included in the feature extraction sub-model extracts the semantic features of the bill image at the resolution granularity set therefor, and the resolution feature enhancement module included in the feature extraction sub-model extracts the high-resolution features of the bill image in a larger range at the resolution granularity set therefor, and the high-resolution features can include the context features of the bill image. The first feature map is obtained by fusing the feature map output by the resolution feature enhancement module and the feature map output by the semantic feature extraction module.
[0046] For example, the size of the feature map output by the resolution feature enhancement module can be HxWxC1, and the size of the feature map output by the semantic feature extraction module can be HxWxC2. Wherein, H is the height, W is the width, C1 is the number of channels of the feature map output by the resolution feature enhancement module, and C2 is the number of channels of the feature map output by the semantic feature extraction module.
[0047] As an implementation manner, the feature extraction sub-model at least includes a semantic feature extraction module and a resolution feature enhancement module; a cross-layer skip connection is adopted between the resolution feature enhancement module and the semantic feature extraction module, the cross-layer skip connection includes channel dimension splicing of the resolution feature enhancement module and the semantic feature extraction module and channel adjustment by one convolution; the resolution feature enhancement module includes at least one hollow convolution layer, each hollow convolution layer has a set hollow rate and a convolution kernel, and is used to expand the resolution receptive field to a set range.
[0048] In this embodiment, a cross-layer skip connection is adopted between the resolution feature enhancement module and the semantic feature extraction module. It can be understood that when the target detection model extracts features of the bill image, the resolution of the feature map will decrease with the increase of the number of downsampling, therefore, the resolution feature enhancement module can be arranged in a relatively shallow network layer to effectively ensure the resolution of the feature map. At the same time, with the increase of the number of downsampling, more global and higher-level semantic information can be captured, therefore, the semantic feature extraction module can be arranged in a relatively deep network layer.
[0049] For example, if the target detection model is based on the traditional YOLOv7 model, the resolution feature enhancement module can be arranged after the 3rd layer CBS module of the backbone network, and the semantic feature extraction module can be arranged before the 10th layer ELAN module of the backbone network.
[0050] In this embodiment, the cross-layer skip connection includes channel dimension splicing of the feature map output by the resolution feature enhancement module and the feature map output by the semantic feature extraction module, and channel adjustment by one convolution to adjust the number of channels to a preset value to obtain an adjustment result.
[0051] Exemplarily, the adjustment result F is obtained through the cross-layer jump connection strategy fuse Specifically, it can be represented as:
[0052] F fuse = Conv 1×1 ([F shallow , F deep ]);
[0053] wherein, Conv 1×1 is a convolution; F shallow is a feature map output by the resolution feature enhancement module; F deep is a feature map output by the semantic feature extraction module; and [F shallow , F deep ] represents splicing two feature maps along the channel dimension.
[0054] In the embodiment, the resolution feature enhancement module comprises at least one dilated convolution layer, each dilated convolution layer has a set dilated rate and a convolution kernel, and is used to expand the resolution receptive field to a set range, so as to capture a larger range of context information without reducing the resolution.
[0055] The above technical solution of the embodiment fuses the high-resolution features of the shallow network layer and the semantic features of the deep network layer through the cross-layer jump connection strategy, so that the edge texture of the regional content (such as micro-text) in the true or false mark region and the contour information of the overall bill image are deeply fused, and the detection accuracy of the regional content is improved; the resolution receptive field is expanded to a set range through the dilated convolution layer in the resolution feature enhancement module, so that a larger range of context information can be captured without reducing the resolution, thereby avoiding the resolution loss caused by down-sampling.
[0056] Optionally, the semantic feature extraction module and the resolution feature enhancement module according to the application can be used to extract features of the bill image according to a set resolution granularity, to obtain a first feature map, and the specific optimization is as follows:
[0057] a11) inputting the bill image into a feature extraction sub-model, expanding the resolution receptive field of the bill image to a set range according to each dilated convolution layer in the resolution feature enhancement module, to form a resolution enhanced image with a resolution range expanded to the set resolution granularity.
[0058] wherein, the resolution enhanced image can be understood as a feature map with a resolution range of the set resolution granularity.
[0059] Exemplarily, the resolution receptive field is set to be 7x7, the dilation rate of the dilated convolution layer can be set to be 2, and the convolution kernel can adopt a 3x3 convolution kernel. Under the set range of the resolution receptive field, the resolution enhanced image F dilated One way of representing (x, y) in the resolution enhanced image F
[0060]
[0061] wherein I is the feature map input into the resolution feature enhancement module; x is the position of the pixel in the width direction in the feature map, y is the position of the pixel in the height direction in the feature map; d is the dilation rate, d = 2; k is the step, k = 1; and W is the convolution kernel weight.
[0062] a12) performing semantic feature extraction on the resolution enhanced image according to the semantic feature extraction module to obtain a first feature map.
[0063] In the embodiment, the semantic feature extraction can be directly performed on the resolution enhanced image by the semantic extraction module; or the resolution enhanced image can be down-sampled at least once, and then the semantic feature extraction is performed on the down-sampled resolution enhanced image by the semantic feature extraction module to obtain more global semantic features in a deeper network layer. The first feature map is obtained by fusing and splicing the resolution enhanced image and the feature map composed of the extracted semantic features.
[0064] In the embodiment, the resolution receptive field of the bill image is expanded to a set range by the resolution feature enhancement module in the feature extraction sub-model to form a resolution enhanced image, which retains high-resolution features and captures more context information. The semantic feature extraction is performed on the resolution enhanced image by the semantic feature extraction module, the resolution enhanced image is spliced with the feature map composed of the semantic features to obtain a first feature map, so that the edge texture of the region content in the authenticity mark region and the contour information of the overall bill image are deeply fused, the detection accuracy of the region content is improved, and the missed detection caused by feature blur is reduced.
[0065] b1) determining the optical reflection feature in the bill image according to the multi-modal attention mechanism by the attention perception sub-model, and fusing the optical reflection feature and the first feature map to generate a second feature map.
[0066] The second feature map can be understood as a feature map that at least fuses the optical reflection feature and the texture feature of the bill image.
[0067] It can be understood that the optical reflection characteristics of the anti-counterfeiting marks such as optically variable ink and holographic patterns dynamically change with the observation angle, and the traditional visual attention mechanism is difficult to distinguish the real features from the background reflection noise points, and thus the accuracy of identifying the real mark area may be reduced. The complex background such as the bill background and decorative patterns is easy to be confused with the anti-counterfeiting features, and thus the accuracy of identifying the real mark area is further affected.
[0068] Therefore, in the embodiment, the optical reflection features in the bill image are determined through the multi-modal attention mechanism in the attention perception sub-model, and the optical reflection features are fused with the first feature map to generate the second feature map, so as to realize the fusion of the optical reflection features and the texture features, and improve the accuracy of identifying the real mark area.
[0069] As an implementation manner, the step of determining the optical reflection features in the bill image according to the multi-modal attention mechanism through the attention perception sub-model and generating the second feature map by fusing the optical reflection features with the first feature map can be further optimized as follows:
[0070] b11) inputting the bill image into the attention perception sub-model to extract the optical reflection features from the bill image and generate a gray gradient map.
[0071] The gray gradient map can be understood as an image for describing the gray change intensity of each pixel point in the image. The optical reflection features can be understood as the characteristics of the reflection of light variable ink and holographic patterns and other objects under the light condition and the features related to the light reflection.
[0072] In the embodiment, the feature map of the bill image processed by the feature extraction sub-model can be input into the attention perception sub-model, or the bill image obtained by the feature extraction sub-model can be transmitted to the attention perception sub-model. The gray gradient map is generated by the attention perception sub-model to highlight the reflection intensity difference of the variable area.
[0073] As an optional implementation manner, the attention perception sub-model can first generate a gray image of the input image, calculate the horizontal and vertical gradients using the Sobel operator, then obtain the gradient amplitude of each pixel point in the gray image through a preset formula, and finally, the gradient amplitudes of the pixel points are summarized to generate the gray gradient map.
[0074] For example, the way of obtaining the gradient amplitude G(x', y') of the pixel point (x', y') in the gray image through the preset formula can be represented as:
[0075]
[0076] wherein, I gray is the gray image of the input image. is a Sobel gradient operator in a horizontal direction; is a Sobel gradient operator in a vertical direction.
[0077] b12) determining channel weights of the first feature map through the channel attention mechanism possessed.
[0078] In the embodiment, the manner of determining the channel weights of the first feature map through the channel attention mechanism possessed can be: performing pooling on the first feature map through a global average pooling manner, and processing the pooled first feature map through a multi-layer perception, mapping the processing result into a preset weight interval by using a preset function such as a Sigmoid function, and obtaining the channel weights of the first feature map.
[0079] Exemplarily, the manner of determining the channel weights α c of the first feature map can be represented as:
[0080] α c = σ(MLP(GAP(F RGB ))) ;
[0081] wherein GAP(·) is global average pooling; MLP(·) is a multi-layer perception; σ is a Sigmoid function; F RGB is the first feature map, F RGB ∈ R H×w×3 ; the channel weights α c ∈ R 3 .
[0082] b13) determining spatial weights of the gray-scale gradient map through the spatial attention mechanism possessed.
[0083] In the embodiment, the manner of determining the spatial weights of the gray-scale gradient map through the spatial attention mechanism can be: processing the gray-scale gradient map through a convolution kernel of a preset size, and mapping the processing result into a preset weight interval by using a preset function such as a Sigmoid function, which is the same as the channel attention mechanism, to obtain the spatial weights of the gray-scale gradient map.
[0084] Exemplarily, the manner of determining the spatial weights β s of the gray-scale gradient map can be represented as:
[0085] β s = σ(Conv 3×3 (G)) ;
[0086] wherein G is the gray-scale gradient map; β s ∈ R H×W .
[0087] b14) performing feature concatenation on the first feature map and the gray gradient map according to the channel weight and the spatial weight, to obtain a second feature map.
[0088] In this embodiment, the first feature map is weighted by the channel weight, and the gray gradient map is weighted by the spatial weight. The weighted first feature map and the weighted second feature map are subjected to feature concatenation to obtain the second feature map.
[0089] For example, the second feature map F MMA may be obtained in the following manner:
[0090] F MMA = Concat(a c ⊙F RGB , b s ⊙G);
[0091] wherein Concat(·) is a feature concatenation function.
[0092] The technical solution described above in this embodiment extracts optical reflection features from a bill image to generate a gray gradient map. A dual-modal attention mechanism of a channel attention mechanism and a spatial attention mechanism is used to assign a channel weight to the first feature map and a spatial weight to the gray gradient map. The first feature map and the gray gradient map are then subjected to feature concatenation, which realizes the fusion of texture features and optical reflection features, enhances the sensitivity to dynamic anti-counterfeiting features, reduces the interference of dynamic optical features such as optically variable ink and holographic patterns, and improves the accuracy of identifying the true-false mark region.
[0093] c1) determining a target pooling kernel according to the feature map size of the second feature map by using the feature pooling sub-model, and performing feature concatenation on the second feature map according to the target pooling kernel to obtain a third feature map.
[0094] The feature pooling sub-model can be understood as a model for feature pooling, which is used to reduce the spatial dimension of the feature map while retaining important feature information. The target pooling kernel can be understood as a pooling kernel with a target window size, which is used for pooling processing of the second feature map. The third feature map can be understood as a feature map obtained by concatenating multi-scale features in the second feature map.
[0095] Optionally, the feature pooling sub-model can be built based on a spatial pyramid pooling structure, and a dynamic receptive field adjustment strategy is introduced to adaptively adjust the size of the pooling kernel according to the feature map size of the second feature map, so as to determine the target pooling kernel.
[0096] In this embodiment, one way of determining the target pooling kernel according to the feature map size of the second feature map can be to determine the corresponding target pooling kernel based on the interval in which the product value of the width and height of the second feature map is located.
[0097] As an implementation manner, the serial cascade is adopted for each pooling layer in the feature pooling sub-model.
[0098] It can be understood that each pooling layer can correspond to a target pooling kernel, and in the prior art, three parallel pooling layers are usually designed to correspond to each target pooling kernel.
[0099] In order to reduce the amount of calculation and improve the efficiency of identifying the true and false mark region in the bill image, in the embodiment, the parallel pooling branch is changed to a serial cascade.
[0100] For example, two consecutive target pooling kernels of 5x5 pooling layers can be equivalent to a target pooling kernel of 9x9 pooling layer. Thus, if the target pooling kernel includes 5x5, 7x7 and 9x9, only two selectable pooling kernels of 5x5 and 7x7 can be set, that is, the adaptive selection of the target pooling kernel can be met, and the calculation efficiency can be greatly improved while maintaining the multi-scale detection capability. Specifically, when the target pooling kernel is 5x5, only one pooling layer of the target pooling kernel of 5x5 needs to be called; if the target pooling kernel is 7x7, only one pooling layer of the target pooling kernel of 7x7 needs to be called; if the target pooling kernel is 9x9, two pooling layers of the target pooling kernel of 5x5 can be called, and the two pooling layers are serially cascaded.
[0101] In the experiment, the inference speed of the target detection model on the bill image of 1080p resolution is improved from 30FPS to 45FPS, while the multi-scale detection capability is maintained.
[0102] Correspondingly, the step of determining the target pooling kernel according to the feature map size of the second feature map and performing feature splicing on the second feature map according to the target pooling kernel to obtain a third feature map through the feature pooling sub-model can be further optimized as follows:
[0103] c11) inputting the second feature map into the feature pooling sub-model to determine the feature map size of the second feature map.
[0104] In the embodiment, the feature map size of the second feature map is determined through the feature pooling sub-model, and the feature map size includes width and height.
[0105] c12) searching for preset pooling kernel determination information to determine the target pooling kernel size matched with the feature map size.
[0106] For example, the pooling kernel determination information can be expressed as:
[0107]
[0108] wherein the kernel Size is a target pooling kernel size.
[0109] In the embodiment, the target pooling kernel size matching the feature map size is determined from the pooling kernel determination information, and a pooling kernel satisfying the target pooling kernel size is determined as the target pooling kernel.
[0110] c13) Splicing the feature vectors in the second feature map by each pooling layer in series connection according to the target pooling kernel size to obtain a spliced third feature map.
[0111] The technical solution described above in the embodiment can adaptively determine the target pooling kernel size according to the feature map size of the second feature map, so as to well cope with the feature scale difference of the authenticity mark region in the bill image, such as a long strip-shaped large target of a security line and a dense small target of microtext; by changing each pooling layer from a traditional parallel pooling branch to a serial connection, the multi-scale detection capability is maintained while the calculation efficiency is greatly improved, and the efficiency of bill anti-forgery detection is improved.
[0112] d1) Dividing the bill image into regions according to the third feature map by the region prediction sub-model to obtain authenticity mark confidence corresponding to each region formed by the division, and determining the region corresponding to the highest authenticity mark confidence as the authenticity mark region.
[0113] wherein the region prediction sub-model can be considered as a model for image region division and confidence evaluation. The authenticity mark confidence can be understood as a numerical value representing the possibility of the region being an authenticity mark region.
[0114] In the embodiment, the region prediction sub-model can divide the bill image into regions according to the third feature map by means such as a sliding window, an anchor frame, and a prediction network, and generate an authenticity mark confidence for each divided region by means such as a classifier. The region corresponding to the highest authenticity mark confidence is determined as the authenticity mark region, ensuring the accuracy and reliability of identifying the authenticity mark region.
[0115] In an optional embodiment, the authenticity mark confidence higher than the preset threshold value can also be determined as each authenticity mark region.
[0116] The technical scheme of the embodiment improves the feature capturing capability for small-scale true or false mark regions by performing feature extraction on the bill image according to the included semantic feature extraction module and resolution feature enhancement module in the feature extraction submodel at a set resolution granularity; the attention perception submodel determines the optical reflection features in the bill image according to the multi-modal attention mechanism, and fuses the texture features and the optical reflection features, thereby enhancing the distinguishability of dynamic anti-fake features and reducing the false detection caused by background interference; the feature pooling submodel flexibly determines the target pooling kernel according to the feature map size of the second feature map, thereby reducing the calculation redundancy while maintaining the multi-scale detection capability, so as to improve the accuracy, reliability and real-time performance of the target detection model in bill anti-fake detection.
[0117] As an implementation manner of any one of the above-mentioned embodiments, the training step of the target detection model can be further optimized as follows:
[0118] a2) obtaining an initial detection model and a sample training set, the sample training set including at least one sample bill image and a corresponding label true or false region.
[0119] The initial detection model can be understood as a detection model with training requirements, which can be constructed based on a neural network. The sample bill image can be regarded as a bill image as a sample, and the sample bill image contains a true or false mark for verifying the true or false of the bill. The real region where the true or false mark is located is labeled by artificial or labeling software to form a label true or false region. Optionally, the region surrounded by the minimum rectangular bounding box of the true or false mark can be uniformly labeled as the label true or false region.
[0120] In the embodiment, a pre-constructed initial detection model and a sample training set are obtained.
[0121] b2) inputting the sample bill image into the initial detection model to obtain a current true or false mark region.
[0122] The current true or false mark region can be understood as a predicted region of the label true or false region in the sample bill image output by the initial detection model.
[0123] It can be understood that, in order to ensure the accuracy of the current true or false mark region, the formation rule of the current true or false mark region can be consistent with the formation rule of the label true or false region. Optionally, the region surrounded by the minimum rectangular bounding box of the predicted true or false mark can be regarded as the current true or false mark region.
[0124] In the embodiment, the sample bill image is input into the initial detection model, the initial detection model performs resolution feature enhancement processing on the bill image based on a feature extraction sub-model, performs multi-modal feature fusion processing on the bill image through an attention perception sub-model, performs adjustment feature pooling core processing on the bill image through a feature pooling sub-model, and finally determines the current authenticity mark region through a region prediction sub-model.
[0125] c2) extracting a first edge gradient map of the current authenticity mark region and a second edge gradient map of the label authenticity region, determining a second norm difference value of the first edge gradient map and the second edge gradient map, determining a sequence of adjacent character center distances in the current authenticity mark region, and determining a variance value of the sequence of adjacent character center distances; and determining a loss function value according to the second norm difference value and the variance value in combination with a set loss function formula.
[0126] The first edge gradient map can be understood as the edge gradient map of the current authenticity mark region, and the second edge gradient map can be considered as the edge gradient map of the label authenticity region. The sequence of adjacent character center distances can be considered as a sequence composed of the center distances between every two adjacent characters. If there are N2 characters, there are N2-1 elements in the sequence of adjacent character center distances, and each element can represent the center distance between two adjacent characters.
[0127] It can be understood that the printed characters and handwritten signatures in the bill often have ink diffusion, deformation or occlusion. The traditional CIoU loss function only optimizes the overall overlap of the bounding box, and ignores the internal structure of the characters, resulting in blurred character detection boundaries, especially in low-quality scanned images, the stroke break or ink diffusion phenomenon further aggravates the detection error.
[0128] In the embodiment, the continuity of character strokes is first constrained. The edge gradient map of the bounding box of the current authenticity mark region can be extracted as the first edge gradient map through the Sobel operator, and the edge gradient map of the bounding box of the label authenticity region can be extracted as the second edge gradient map through the Sobel operator, and the second norm difference value of the first edge gradient map and the second edge gradient map is determined.
[0129] Secondly, the consistency of character spacing is constrained. The sequence of adjacent character center distances in the current authenticity mark region can be determined, and the variance value of the sequence of adjacent character center distances is determined, so as to punish the layout deformation.
[0130] Finally, the weighted sum formula of the second norm difference value and the variance value can be determined as the loss function formula, or the weighted sum formula of the second norm difference value, the variance value and the CIoU loss function value can be determined as the loss function formula, which considers the similarity of the character structure and constrains the character structure.
[0131] An exemplary second-order norm difference value may be specifically expressed as:
[0132]
[0133] wherein, is the first edge gradient map; is the second edge gradient map; (i,j) is the position of the pixel point on the boundary frame Edge of the current true-false mark region; N1 is the number of pixel points on the boundary frame of the current true-false mark region.
[0134] An exemplary adjacent character center distance sequence the variance value of the adjacent character center distance sequence may be specifically expressed as:
[0135]
[0136] wherein, is the mean value of all adjacent character center distances in the adjacent character center distance sequence, and N2 is the number of characters of the current true-false mark region.
[0137] An exemplary loss function formula can be specifically expressed as:
[0138]
[0139] wherein, λ1 and λ2 are balance weights, λ1 = 0.5, and λ2 = 0.3. is the loss function value; is the Clou loss function value.
[0140] d2) reversely learning and adjusting the network parameters in the initial detection model according to the loss function value, obtaining an adjusted initial detection model, and returning to perform the input operation of the sample bill image again until a training end condition is reached; determining the initial detection model obtained after the training as a target detection model.
[0141] In the present embodiment, the training end condition can include a maximum iteration number, an accuracy threshold, an accuracy variable threshold, and / or a traversed sample bill image.
[0142] The technical scheme of the embodiment is characterized in that, according to the second-order norm difference value of the first edge gradient graph and the second edge gradient graph and the variance value of the adjacent character center distance sequence in the current true-false mark region, and in combination with the set loss function formula, the loss function value is determined, the constraint on the character structure is introduced in the loss function, the detection boundary of the character region is optimized, the limitation of the traditional loss function in the dense character scene for bill anti-counterfeiting detection is overcome, and the detection accuracy of printed characters and handwritten signatures and other character type anti-counterfeiting marks in the true-false mark region is improved.
[0143] Figure 2 A structural schematic diagram of a bill anti-counterfeiting detection device provided by the embodiment of the application is shown in FIG. 1. Figure 2 As shown in the figure, the device comprises an image acquisition module 21, a region identification module 22 and a true-false detection module 23, wherein,
[0144] The image acquisition module 21 is configured to acquire a bill image of a target bill.
[0145] The region identification module 22 is configured to identify a true-false mark region in the bill image by using a trained target detection model based on resolution feature enhancement, multi-modal feature fusion and adjustment of a feature pooling kernel on the bill image.
[0146] The true-false detection module 23 is configured to compare the region content of the true-false mark region with set reference mark content, and determine a true-false detection result of the target bill according to the comparison result.
[0147] The technical scheme of the embodiment is characterized in that, the bill image of the target bill is acquired, the true-false mark region in the bill image is identified by using the trained target detection model based on resolution feature enhancement, multi-modal feature fusion and adjustment of a feature pooling kernel on the bill image, the region content of the true-false mark region is compared with the set reference mark content, and the true-false detection result of the target bill is determined according to the comparison result. By using this method, the resolution feature enhancement processing is performed on the bill image based on the target detection model, the feature capturing capability for small-scale true-false mark regions is improved, the multi-modal feature fusion processing is performed on the bill image, the features related to optics are fused, the sensitivity of the model to dynamic anti-counterfeiting features is enhanced, the false detection rate of the true-false mark region is reduced, in addition, based on the adjustment of the feature pooling kernel, the appropriate pooling kernel can be dynamically selected based on the scale difference of the true-false mark region in the target bill, so as to reduce the calculation redundancy while maintaining the multi-scale detection capability of the bill image, improve the efficiency of identifying the true-false mark region, and ensure the real-time detection, thereby meeting the high-precision and high-efficiency requirements of bill anti-counterfeiting detection in complex scenes.
[0148] Further, the target detection model comprises: a feature extraction sub-model, an attention perception sub-model, a feature pooling sub-model, and a region prediction sub-model.
[0149] Correspondingly, the region identification module 22 can specifically comprise:
[0150] A first feature map acquisition unit is configured to acquire a first feature map by performing feature extraction on the bill image according to a semantic feature extraction module and a resolution feature enhancement module in the feature extraction sub-model at a set resolution granularity.
[0151] A second feature map acquisition unit is configured to determine an optical reflection feature in the bill image according to a multi-modal attention mechanism by using the attention perception sub-model, and to generate a second feature map by fusing the optical reflection feature and the first feature map.
[0152] A third feature map acquisition unit is configured to determine a target pooling kernel according to a feature map size of the second feature map by using the feature pooling sub-model, and to perform feature splicing on the second feature map according to the target pooling kernel to obtain a third feature map.
[0153] A true-false mark region determination unit is configured to divide the bill image into regions according to the third feature map by using the region prediction sub-model, to obtain a true-false mark confidence corresponding to each region formed by the division, and to determine a region corresponding to the highest true-false mark confidence as a true-false mark region.
[0154] Further, the feature extraction sub-model comprises at least a semantic feature extraction module and a resolution feature enhancement module; a cross-layer skip connection is used between the resolution feature enhancement module and the semantic feature extraction module, the cross-layer skip connection comprises channel dimension splicing of the resolution feature enhancement module and the semantic feature extraction module and channel adjustment by one convolution; the resolution feature enhancement module comprises at least one hollow convolution layer, each hollow convolution layer has a set hollow rate and a convolution kernel, and is configured to expand a resolution receptive field to a set range.
[0155] Further, the first feature map acquisition unit can be specifically configured to:
[0156] input the bill image into the feature extraction sub-model, expand a resolution receptive field of the bill image to a set range according to each hollow convolution layer in the resolution feature enhancement module, and form a resolution enhanced image with a resolution range expanded to the set resolution granularity;
[0157] perform semantic feature extraction on the resolution enhanced image according to the semantic feature extraction module to obtain a first feature map.
[0158] Further, the second feature map acquisition unit can be specifically configured to:
[0159] input the bill image into the attention perception sub-model, extract the optical reflection feature from the bill image to generate a gray gradient map;
[0160] determine the channel weight of the first feature map through the channel attention mechanism possessed;
[0161] determine the spatial weight of the gray gradient map through the spatial attention mechanism possessed;
[0162] perform feature splicing on the first feature map and the gray gradient map according to the channel weight and the spatial weight, and obtain a second feature map.
[0163] Further, each pooling layer in the feature pooling sub-model adopts serial cascade;
[0164] Correspondingly, the third feature map acquisition unit can be specifically configured to:
[0165] input the second feature map into the feature pooling sub-model, and determine the feature map size of the second feature map;
[0166] find the preset pooling kernel determination information to determine a target pooling kernel size matched with the feature map size;
[0167] perform splicing on the feature vectors in the second feature map according to the target pooling kernel size through the serially cascaded pooling layers, and obtain a spliced third feature map.
[0168] Further, the device further includes a model training module, which can be specifically configured to:
[0169] obtain an initial detection model and a sample training set, the sample training set including at least one sample bill image and a corresponding label true-false region;
[0170] input the sample bill image into the initial detection model to obtain a current true-false mark region;
[0171] extract a first edge gradient map of the current true-false mark region, and extract a second edge gradient map of the label true-false region, and determine a second-order norm difference value of the first edge gradient map and the second edge gradient map;
[0172] determine a sequence of adjacent character center distances in the current true-false mark region, and determine a variance value of the sequence of adjacent character center distances;
[0173] determine a loss function value according to the second-order norm difference value and the variance value combined with a preset loss function formula;
[0174] According to the loss function value, the network parameters in the initial detection model are learned and adjusted in reverse, an adjusted initial detection model is obtained, and the input operation of the sample bill image is re-executed until a training end condition is reached.
[0175] The initial detection model obtained after training is determined as a target detection model.
[0176] The bill anti-counterfeiting detection device provided in the embodiments of the present application can execute the bill anti-counterfeiting detection method provided in any of the embodiments of the present application, and has the corresponding function modules and beneficial effects of the execution method.
[0177] Figure 3 A structural schematic diagram of an electronic device 30 that can be used to implement embodiments of the present application is shown. The electronic device is intended to represent various forms of digital computers, such as laptops, desktops, tablets, personal digital assistants, servers, blade servers, mainframes, and other appropriate computers. The electronic device can also represent various forms of mobile devices, such as personal digital processors, cellular telephones, smart phones, wearable devices (e.g., headsets, glasses, watches, etc.), and other similar computing devices. The components shown here, their connections and relationships, and their functions, are meant to be examples only, and are not intended to limit the implementations of the present application described and / or claimed in this document.
[0178] As shown in Figure 3 The electronic device 30 includes at least one processor 31, and a memory, such as a read-only memory (ROM) 32, a random access memory (RAM) 33, etc., which is communicatively connected to the at least one processor 31, wherein the memory stores a computer program that can be executed by the at least one processor. The processor 31 can perform various appropriate actions and processes according to the computer program stored in the read-only memory (ROM) 32 or loaded into the random access memory (RAM) 33 from the storage unit 38. In the RAM 33, various programs and data required for the operation of the electronic device 30 can also be stored. The processor 31, the ROM 32, and the RAM 33 are connected to each other through a bus 34. An input / output (I / O) interface 35 is also connected to the bus 34.
[0179] A plurality of components in the electronic device 30 are connected to the I / O interface 35, including: an input unit 36, such as a keyboard, a mouse, etc.; an output unit 37, such as various types of displays, speakers, etc.; a storage unit 38, such as a magnetic disk, an optical disk, etc.; and a communication unit 39, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 39 allows the electronic device 30 to exchange information / data with other devices through a computer network, such as the Internet, and / or various telecommunications networks.
[0180] The processor 31 can be various general and / or special purpose processing components with processing and computing capabilities. Some examples of the processor 31 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various specialized artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The processor 31 performs various methods and processes described above, such as the bill anti-forgery detection method.
[0181] In some embodiments, the bill anti-forgery detection method can be implemented as a computer program tangibly embodied in a computer readable storage medium, such as the storage unit 38. In some embodiments, part or all of the computer program can be loaded and / or installed onto the electronic device 30 via the ROM 32 and / or the communication unit 39. When the computer program is loaded onto the RAM 33 and executed by the processor 31, one or more steps of the bill anti-forgery detection method described above can be performed. Alternatively, in other embodiments, the processor 31 can be configured to perform the bill anti-forgery detection method by any other suitable means, such as by means of firmware.
[0182] Various implementations of the systems and techniques described above can be realized in digital electronic circuitry, integrated circuitry, a field programmable gate array (FPGA), an application specific integrated circuit (ASIC), a system on a chip (SOC), a programmable logic device (PLD), a computer hardware, firmware, software, and / or combinations thereof. These various implementations can include implementation in one or more computer programs that are executable and / or interpretable on a programmable system including at least one programmable processor, which can be special or general purpose, coupled to receive data and instructions from, and to transmit data and instructions to, a storage system, at least one input device, and at least one output device.
[0183] Computer programs used to implement the methods of the application can be written in any combination of one or more programming languages. These computer programs can be provided to a processor of a general purpose computer, special purpose computer, or other programmable data processing apparatus to produce a machine, such that the computer program, when executed, implements the functions / acts specified in the flowcharts and / or block diagrams. The computer program can be executed entirely on a machine, partially on a machine, partially on a machine as a stand-alone software package, partially on a machine and partially on a remote machine or entirely on a remote machine or server.
[0184] In the context of this invention, a computer-readable storage medium can be a tangible medium that may contain or store a computer program for use by or in conjunction with an instruction execution system, apparatus, or device. A computer-readable storage medium may include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination thereof. Alternatively, a computer-readable storage medium may be a machine-readable signal medium. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof.
[0185] To provide interaction with a user, the systems and techniques described herein can be implemented on an electronic device having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the electronic device. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).
[0186] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as data servers), or computing systems that include middleware components (e.g., application servers), or computing systems that include frontend components (e.g., user computers with graphical user interfaces or web browsers through which users can interact with implementations of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., communication networks). Examples of communication networks include local area networks (LANs), wide area networks (WANs), blockchain networks, and the Internet.
[0187] A computing system can include clients and servers. Clients and servers are generally located far apart and typically interact through communication networks. The client-server relationship is created by computer programs running on the respective computers and having a client-server relationship with each other. The server can be a cloud server, also known as a cloud computing server or cloud host, which is a hosting product within the cloud computing service system to address the shortcomings of traditional physical hosts and VPS services, such as high management difficulty and weak business scalability.
[0188] It should be understood that the various forms of processes shown above can be used, with steps reordered, added, or deleted. For example, the steps described in this invention can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution of this invention can be achieved, and this is not limited herein.
[0189] The specific embodiments described above do not constitute a limitation on the scope of protection of this invention. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this invention should be included within the scope of protection of this invention.
Claims
1. A method of security detection of a document, characterized in that, The method comprises the following steps: acquiring a bill image of a target bill; recognizing a true-false mark region in the bill image based on resolution feature enhancement, multi-modal feature fusion and adjustment of a feature pooling kernel through a trained target detection model; comparing the region content of the true-false mark region with a set of reference mark content, and determining the true-false detection result of the target bill according to the comparison result.
2. The method of claim 1, wherein, The target detection model comprises a feature extraction sub-model, an attention perception sub-model, a feature pooling sub-model and a region prediction sub-model. The method of recognizing a true-false mark region in the bill image based on resolution feature enhancement, multi-modal feature fusion and adjustment of a feature pooling kernel through a trained target detection model comprises the following steps: extracting features of the bill image according to a semantic feature extraction module and a resolution feature enhancement module included in the feature extraction sub-model at a set resolution granularity to obtain a first feature map; determining optical reflection features in the bill image according to a multi-modal attention mechanism through the attention perception sub-model, and fusing the optical reflection features and the first feature map to generate a second feature map; determining a target pooling kernel according to the feature map size of the second feature map through the feature pooling sub-model, and performing feature splicing on the second feature map according to the target pooling kernel to obtain a third feature map; performing region division on the bill image according to the third feature map through the region prediction sub-model to obtain a true-false mark confidence corresponding to each region formed by the division, and determining the region corresponding to the highest true-false mark confidence as the true-false mark region.
3. The method of claim 2, wherein, The feature extraction sub-model at least includes a semantic feature extraction module and a resolution feature enhancement module; The resolution feature enhancement module and the semantic feature extraction module are connected through a cross-layer skip connection, which comprises channel dimension splicing of the resolution feature enhancement module and the semantic feature extraction module and channel adjustment through one convolution; The resolution feature enhancement module comprises at least one hollow convolution layer, each hollow convolution layer has a set of hole rate and convolution kernel, and is used for expanding the resolution receptive field to a set range.
4. The method of claim 3, wherein, The method of extracting features of the bill image according to a semantic feature extraction module and a resolution feature enhancement module included in the feature extraction sub-model at a set resolution granularity to obtain a first feature map comprises the following steps: inputting the bill image into the feature extraction sub-model, expanding the resolution receptive field of the bill image to a set range according to each hollow convolution layer in the resolution feature enhancement module to form a resolution enhanced image with the resolution range expanded to the set resolution granularity; extracting semantic features of the resolution enhanced image according to the semantic feature extraction module to obtain a first feature map.
5. The method of claim 2, wherein, The method of determining optical reflection features in the bill image according to a multi-modal attention mechanism through the attention perception sub-model, and fusing the optical reflection features and the first feature map to generate a second feature map comprises the following steps: input the bill image into the attention perception sub-model, extract optical reflection features from the bill image to generate a gray scale gradient map; determine the channel weight of the first feature map through the channel attention mechanism possessed; determine the spatial weight of the gray scale gradient map through the spatial attention mechanism possessed; According to the channel weight and the spatial weight, the first feature map and the gray scale gradient map are feature spliced to obtain a second feature map.
6. The method of claim 2, wherein, Each pooling layer in the feature pooling sub-model adopts serial cascade; The feature pooling sub-model includes: input the second feature map into the feature pooling sub-model to determine the feature map size of the second feature map; determine the target pooling kernel size matched with the feature map size by searching the preset pooling kernel determination information; splicing the feature vectors in the second feature map according to the target pooling kernel size through each pooling layer in serial cascade to obtain a spliced third feature map.
7. The method according to any one of claims 1 to 6, characterized in that, The training steps of the target detection model include: obtain an initial detection model and a sample training set, the sample training set including at least one sample bill image and a corresponding label true-false region; input the sample bill image into the initial detection model to obtain a current true-false mark region; extract a first edge gradient map of the current true-false mark region and a second edge gradient map of the label true-false region, and determine the second norm difference value of the first edge gradient map and the second edge gradient map; determine the adjacent character center distance sequence in the current true-false mark region, and determine the variance value of the adjacent character center distance sequence; determine the loss function value according to the second norm difference value and the variance value combined with the preset loss function formula; According to the loss function value, the network parameters in the initial detection model are reversely learned and adjusted to obtain an adjusted initial detection model, and the input operation of the sample bill image is returned to be re-executed until the training end condition is reached; determine the initial detection model obtained after training as the target detection model.
8. A document security detection apparatus, characterized by including: an image acquisition module for acquiring a bill image of a target bill; a region identification module for identifying a true-false mark region in the bill image through a trained target detection model based on resolution feature enhancement, multi-modal feature fusion and adjustment of the feature pooling kernel on the bill image; a true-false detection module for comparing the region content of the true-false mark region with the preset reference mark content, and determining the true-false detection result of the target bill according to the comparison result.
9. An electronic device, comprising: The electronic device includes: at least one processor; and a memory connected in communication with the at least one processor; wherein the memory stores a computer program executable by the at least one processor, and the computer program is executed by the at least one processor to enable the at least one processor to execute the bill anti-fake detection method of any one of claims 1-7.
10. A computer-readable storage medium, characterized in that, The computer readable storage medium stores computer instructions for causing the processor to implement the bill anti-counterfeiting detection method in any one of claims 1-7 when executed.