Printer source identification method, device and equipment based on channel attention mechanism
Through a deep learning method based on the channel attention mechanism, using an image acquisition device and a printer source recognition model, the problems of low efficiency of traditional printer source recognition and difficulty in recognizing Chinese documents are solved, achieving more efficient and accurate printer source recognition.
Patent Information
- Application Number
- CN202510785912.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-12
- Publication Date
- 2025-09-12
AI Technical Summary
Traditional printer source recognition technology relies on scanners, which is inefficient and difficult to effectively recognize Chinese documents, especially mixed characters.
A printer source identification method based on the channel attention mechanism is adopted. The printed document image is obtained through an image acquisition device. The position information and features of mixed characters are extracted by combining the deep learning model. The channel weight value is calculated using the channel attention mechanism to determine the printer source.
It improves the efficiency and accuracy of printer source recognition and expands the scope of application, especially the recognition capability of Chinese documents.
Smart Images

Figure CN120635924A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of artificial intelligence or the Internet of Things, and more specifically to a method, apparatus, device, medium, and program product for identifying a printer source based on a channel attention mechanism. Background Art
[0002] With the continuous advancement of research in printed document authentication, significant progress has been made in methods such as texture analysis and geometric distortion detection. Traditional printer source identification techniques typically rely on scanners to acquire digital images of printed documents. Distinctive features are then manually designed and extracted from specific characters within the printed document. These features are then used to identify and classify the extracted characters, ultimately determining the printer source based on a majority vote of statistical results. However, this approach has certain limitations. First, the reliance on scanners restricts the widespread application of the technology. Second, the manual feature extraction process requires researchers to invest significant time and effort, resulting in low efficiency. Furthermore, research strategies based on specific characters are less effective for Chinese documents, as it is often difficult to find enough specific characters for effective recognition in Chinese texts. Summary of the Invention
[0003] In view of the above problems, embodiments of the present disclosure provide a method, apparatus, device, medium, and program product for identifying a printer source based on a channel attention mechanism.
[0004] According to the first aspect of the present disclosure, a method for identifying a printer source based on a channel attention mechanism is provided, the method comprising: obtaining image information of a printed document to be identified by using an image acquisition device; extracting multiple mixed characters of the printed document to be identified based on the image information, and determining the position information of each of the mixed characters in the image information; extracting target printer features corresponding to the printed document to be identified based on the attention mechanism using a pre-trained printer source recognition model according to the position information of each of the mixed characters; and determining the printer source of the printed document to be identified based on the target printer features.
[0005] According to an embodiment of the present disclosure, it also includes: performing edge detection on the image information to generate a binary edge map; detecting straight line segments in the binary edge map through Hough transform, and calculating the angle between each straight line segment and the horizontal direction; and determining the inclination angle of the printed document to be identified based on the angle with the highest frequency of occurrence, and performing rotation correction on the image information according to the inclination angle.
[0006] According to an embodiment of the present disclosure, the extracting of multiple mixed characters of the printed document to be identified and determining the position information of each of the mixed characters in the image information includes: performing morphological operations on the mixed characters in the image information, connecting the disconnected parts of each of the mixed characters, and generating multiple connected character areas; performing horizontal projection and vertical projection on each of the connected character areas, respectively, to determine the upper and lower boundaries and left and right boundaries of each of the mixed characters; and determining the position information of each of the mixed characters in the image information based on the upper and lower boundaries and the left and right boundaries.
[0007] According to an embodiment of the present disclosure, according to the position information of each of the mixed characters, a pre-trained printer source recognition model is used to extract the target printer features corresponding to the printed document to be identified based on the attention mechanism, including: inputting the image information into the printer source recognition model, and extracting a multi-scale feature map according to the position information of each of the mixed characters; calculating the channel weight value for each feature channel of the multi-scale feature map based on the channel attention mechanism; and determining the target printer features based on the channel weight value of each of the feature channels.
[0008] According to an embodiment of the present disclosure, the channel weight value of each feature channel of the multi-scale feature map is calculated based on the channel attention mechanism, including: performing global average pooling on the spatial information of each feature channel to generate a channel descriptor vector; inputting the channel descriptor vector into a one-dimensional convolution layer to obtain the output of the one-dimensional convolution layer; based on a fully connected layer, performing linear transformation and nonlinear activation on the output of the one-dimensional convolution layer to obtain the output of the fully connected layer; and calculating the output of the fully connected layer through an activation function to obtain the channel weight value corresponding to each feature channel.
[0009] According to an embodiment of the present disclosure, determining the target printer characteristics based on the channel weight value of each feature channel includes: multiplying the channel weight value with the feature channel to obtain a weighted feature map; and determining the target printer characteristics based on the weighted feature map.
[0010] According to an embodiment of the present disclosure, based on the target printer features, determining the printer source of the printed document to be identified includes: inputting the target printer features into a classifier and obtaining the printer source probability distribution output by the classifier; and determining the printer source according to the category corresponding to the maximum value of the printer source probability distribution.
[0011] The second aspect of the present disclosure provides a device for identifying a printer source based on a channel attention mechanism, the device comprising: an image acquisition module for acquiring image information of a printed document to be identified using an image acquisition device; a character extraction module for extracting multiple mixed characters of the printed document to be identified based on the image information, and determining the position information of each of the mixed characters in the image information; a feature extraction module for extracting target printer features corresponding to the printed document to be identified based on the attention mechanism using a pre-trained printer source recognition model according to the position information of each of the mixed characters; and a document recognition module for determining the printer source of the printed document to be identified based on the target printer features.
[0012] A third aspect of the present disclosure provides an electronic device comprising: one or more processors; a storage device for storing one or more programs, wherein the one or more processors execute the one or more computer programs to implement the steps of the above-mentioned method for identifying a printer source based on a channel attention mechanism.
[0013] A fourth aspect of the present disclosure further provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of the above-mentioned method for identifying a printer source based on a channel attention mechanism.
[0014] A fifth aspect of the present disclosure further provides a computer program product, comprising a computer program, which, when executed by a processor, implements the steps of the above-mentioned method for identifying a printer source based on a channel attention mechanism.
[0015] In the disclosed embodiments, an image acquisition device (such as a mobile phone camera) replaces a traditional scanner. Combined with deep learning methods, a printer source recognition model is used to extract internal printer features from printed documents. This avoids the complexity of manual feature design and extraction, making the printer source recognition method more applicable. Furthermore, the research scope is expanded from specific characters to mixed characters, addressing the difficulty of recognizing Chinese printed documents and further expanding its scope of application. BRIEF DESCRIPTION OF THE DRAWINGS
[0016] The above contents and other objects, features and advantages of the present disclosure will become more apparent through the following description of the embodiments of the present disclosure with reference to the accompanying drawings, in which:
[0017] Figure 1 A diagram schematically illustrates an application scenario of a method, apparatus, device, medium, and program product for identifying a printer source based on a channel attention mechanism according to an embodiment of the present disclosure;
[0018] Figure 2The flowchart of the method for identifying a printer source based on a channel attention mechanism according to an embodiment of the present disclosure is schematically shown;
[0019] Figure 3 Schematically illustrates a working scenario diagram of a printer source identification method, apparatus, device, medium, and program product based on a channel attention mechanism according to an embodiment of the present disclosure;
[0020] Figure 4 Schematically shows a flowchart of character extraction according to an embodiment of the present disclosure;
[0021] Figure 5 Schematically shows a framework diagram of printer source identification according to an embodiment of the present disclosure;
[0022] Figure 6 Schematically shows a structural diagram of a channel attention mechanism according to an embodiment of the present disclosure;
[0023] Figure 7 Schematically shows a structural block diagram of a device for identifying a printer source based on a channel attention mechanism according to an embodiment of the present disclosure; and
[0024] Figure 8 A block diagram of an electronic device for a method for identifying a printer source based on a channel attention mechanism according to an embodiment of the present disclosure is schematically shown. DETAILED DESCRIPTION
[0025] Hereinafter, embodiments of the present disclosure will be described with reference to the accompanying drawings. However, it should be understood that these descriptions are merely exemplary and are not intended to limit the scope of the present disclosure. In the detailed description below, for ease of explanation, many specific details are set forth to provide a comprehensive understanding of the embodiments of the present disclosure. However, it is apparent that one or more embodiments may also be implemented without these specific details. In addition, in the following description, descriptions of well-known structures and technologies are omitted to avoid unnecessary confusion of the concepts of the present disclosure.
[0026] The terms used herein are only for describing specific embodiments and are not intended to limit the present disclosure. The terms "comprise," "include," etc. used herein indicate the presence of the features, steps, operations, and / or components, but do not exclude the presence or addition of one or more other features, steps, operations, or components.
[0027] All terms used herein (including technical and scientific terms) have the meanings commonly understood by those skilled in the art unless otherwise defined. It should be noted that the terms used herein should be interpreted as having a meaning consistent with the context of this specification and should not be interpreted in an idealized or overly rigid manner.
[0028] When expressions such as "at least one of A, B, and C, etc." are used, they should generally be interpreted in accordance with the meaning commonly understood by those skilled in the art (for example, "a system having at least one of A, B, and C" should include but is not limited to a system having A alone, B alone, C alone, A and B, A and C, B and C, and / or A, B, C, etc.).
[0029] In the technical solution of the present invention, the user information involved (including but not limited to user personal information, user image information, user device information, such as location information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) are all information and data authorized by the user or fully authorized by all parties, and the collection, storage, use, processing, transmission, provision, disclosure and application of the relevant data comply with relevant laws, regulations and standards, take necessary confidentiality measures, do not violate public order and good morals, and provide corresponding operation entrances for users to choose to authorize or refuse.
[0030] In scenarios where personal information is used for automated decision-making, the methods, devices, and systems provided by the embodiments of the present disclosure all provide users with corresponding operation portals for them to choose to agree or reject the automated decision-making results; if the user chooses to reject, the expert decision-making process will be entered. The expression "automated decision-making" here refers to the activity of automatically analyzing and evaluating an individual's behavioral habits, interests and hobbies, or economic, health, credit status, etc. through computer programs and making decisions. The expression "expert decision-making" here refers to the activity of making decisions by people who specialize in a certain field, have specialized experience, knowledge, and skills, and have reached a certain level of professionalism.
[0031] Figure 1 The application scenario diagram of the printer source identification method, apparatus, device, medium and program product based on the channel attention mechanism according to an embodiment of the present disclosure is schematically shown.
[0032] like Figure 1 As shown, the application scenario 100 according to this embodiment may include a first terminal device 101, a second terminal device 102, a third terminal device 103, a network 104, and a server 105. The network 104 is used as a medium for providing a communication link between the first terminal device 101, the second terminal device 102, the third terminal device 103, and the server 105. The network 104 may include various connection types, such as wired or wireless communication links or optical fiber cables.
[0033] A user may use a first terminal device 101, a second terminal device 102, or a third terminal device 103 to interact with a server 105 via a network 104 to receive or send messages, etc. Various communication client applications may be installed on the first terminal device 101, the second terminal device 102, or the third terminal device 103, such as shopping applications, web browser applications, search applications, instant messaging tools, email clients, social platform software, etc. (for example only).
[0034] The first terminal device 101 , the second terminal device 102 , and the third terminal device 103 may be various electronic devices having display screens and supporting web browsing, including but not limited to smart phones, tablet computers, laptop computers, desktop computers, and the like.
[0035] The server 105 may be a server that provides various services, such as a background management server (for example only) that supports websites browsed by users using the first terminal device 101, the second terminal device 102, and the third terminal device 103. The background management server may analyze and process received data such as user requests, and feed back processing results (e.g., web pages, information, or data obtained or generated based on user requests) to the terminal devices.
[0036] It should be noted that the printer source identification method based on the channel attention mechanism provided in the embodiment of the present disclosure can generally be executed by the server 105. Accordingly, the printer source identification device based on the channel attention mechanism provided in the embodiment of the present disclosure can generally be set in the server 105. The printer source identification method based on the channel attention mechanism provided in the embodiment of the present disclosure can also be executed by a server or server cluster that is different from the server 105 and can communicate with the first terminal device 101, the second terminal device 102, the third terminal device 103 and / or the server 105. Accordingly, the printer source identification device based on the channel attention mechanism provided in the embodiment of the present disclosure can also be set in a server or server cluster that is different from the server 105 and can communicate with the first terminal device 101, the second terminal device 102, the third terminal device 103 and / or the server 105.
[0037] It should be understood that Figure 1 The number of terminal devices, networks and servers in the embodiment is merely illustrative. Any number of terminal devices, networks and servers may be provided as required.
[0038] The inventors have discovered through research that, in related technologies, with the popularization of printers and the reduction of printing costs, printed documents have become ubiquitous in people's lives and have become one of the important carriers of information. Therefore, the importance of identifying and tracing the source of printed documents is increasing. The purpose of printer source identification is to trace the printer source of printed documents and provide important clues for document inspection. This is achieved by extracting features between different printers. Existing printer source identification can be divided into two different methods: one is to add additional external signals to the printed document. This external signal is difficult to detect with the naked eye. For example, a mark such as a watermark that is invisible to the naked eye is embedded in the printed document. This method is equivalent to pre-embedding anti-counterfeiting information and then viewing it in a specific way; the other is to identify the printer by using the intrinsic property differences of the printer. This intrinsic feature difference of the printer is determined by the electromechanical characteristics inside the printer. It is difficult to forge and remove and can be obtained through scientific methods.
[0039] Figure 2 The flowchart of the method for identifying a printer source based on a channel attention mechanism according to an embodiment of the present disclosure is schematically shown.
[0040] like Figure 2 As shown, the printer source identification method based on the channel attention mechanism of this embodiment includes operations S210 to S240.
[0041] In operation S210, image information of a printed document to be identified is acquired using an image acquisition device.
[0042] In an embodiment of the present disclosure, before acquiring the image information of the printed document to be identified, the user's consent or authorization is obtained. For example, before operation S210, a request to acquire the image information of the printed document to be identified is issued to the user. If the user agrees or authorizes the acquisition of the image information of the printed document to be identified, operation S210 is executed.
[0043] Exemplarily, the image acquisition device may be a digital camera equipped with a lens, or a mobile phone with a camera, etc., to replace a traditional scanner and acquire image information of a printed document.
[0044] When using an image acquisition device to capture image information of printed documents, choose a location with even, bright lighting, avoiding direct strong light or shadows. If the ambient light is insufficient, use a fill light for auxiliary illumination. Ensure that the fill light is even and soft, and does not create noticeable reflections or shadows on the document surface. When shooting, keep the camera lens perpendicular to the document surface to ensure that the captured image does not exhibit perspective distortion.
[0045] In operation S220, a plurality of mixed characters of the to-be-recognized printed document are extracted based on the image information, and position information of each of the mixed characters in the image information is determined.
[0046] Currently, methods for printer source identification based on text documents can be categorized into three main categories: texture-based, noise-based, and geometric distortion-based. Many of these methods rely on character analysis, making character extraction from printed documents a critical step. For example, many texture-based methods begin by locating the letter "e," the most common letter in English printed documents. After determining the location of the "e," they accurately extract it from the printed document, then extract the desired feature information from the "e." Therefore, character extraction from printed documents is an essential step in most printer source identification methods.
[0047] In some exemplary embodiments, image preprocessing techniques can be used to optimize the captured image. For example, histogram equalization can be used to enhance image contrast, making the difference between characters and background more distinct, facilitating subsequent character segmentation and recognition. Furthermore, the image can be binarized, converting it into a black and white binary image to further simplify image information and highlight character features.
[0048] For mixed characters containing different fonts, font sizes and colors, character segmentation algorithms can be used to extract characters. At the same time, the structural features and contextual information of the characters can be combined to optimize and correct the segmentation results, so as to accurately segment the mixed characters from the image.
[0049] Furthermore, a coordinate system can be established to describe the position of characters in the image. Specifically, the upper left corner of the image is used as the coordinate origin, the horizontal axis is the x-axis, and the vertical axis is the y-axis. For each extracted mixed character, the coordinates (x, y) of its upper left corner in the coordinate system, as well as the width (width) and height (height) of the character, are recorded to determine the character's position.
[0050] In operation S230, according to the position information of each of the mixed characters, a pre-trained printer source recognition model is used to extract target printer features corresponding to the printed document to be recognized based on an attention mechanism.
[0051] For example, a printer source recognition model can be built based on a convolutional neural network, which features low power consumption and high accuracy. Furthermore, given the resource constraints of practical applications, a lightweight convolutional neural network (such as the MoblieNet v1 network) can be selected as the backbone network. This model introduces depthwise separable convolution, which splits standard convolution into depthwise convolution and pointwise convolution. The former extracts single-channel spatial features, while the latter integrates channel information. This design significantly reduces the number of parameters and computations required, allowing the model to run efficiently on resource-constrained devices while maintaining accuracy. Furthermore, to address the lower accuracy of mixed Chinese characters compared to single Chinese characters, a channel attention mechanism (UCA) can be introduced to automatically learn the importance of different channel features and weight them, allowing the network to focus more on the internal features of the printer, filtering out irrelevant information interference, and improving the recognition accuracy of mixed Chinese character documents.
[0052] In operation S240 , the printer source of the to-be-identified printed document is determined based on the target printer characteristics.
[0053] In the embodiments of the present disclosure, a corresponding operation portal is provided for the user to choose to accept or reject the automated decision result. Specifically, before proceeding with the process of determining the printer source of the print document to be identified, an instruction to accept or reject the process / decision is obtained from the user through the corresponding operation portal. If the user agrees to proceed with the process / decision, the process / decision of determining the printer source of the print document to be identified is carried out, i.e., step S240 is executed. If the user rejects the process / decision, the expert decision process is entered.
[0054] After extracting the target printer features, they can be fed into the trained classification layer. The classification layer calculates the features based on the learned parameters and outputs a probability vector, where each element corresponds to the probability of a printer source category. The larger the probability value, the more likely the document belongs to the corresponding printer source category. Ultimately, the printer source category with the highest probability value is selected as the printer source for the document to be identified.
[0055] Figure 3 A working scenario diagram of a method, apparatus, device, medium, and program product for identifying a printer source based on a channel attention mechanism according to an embodiment of the present disclosure is schematically shown.
[0056] For example, Figure 3As shown, first, a Chinese digital document can be obtained and then printed on at least two printers, resulting in at least two printed documents with unknown printer sources. Furthermore, these printed documents can be photographed using a mobile phone camera to obtain image information. For these images of documents captured by the mobile phone, mixed characters can be extracted and their corresponding location information determined. This location information is then used to extract the target printer features corresponding to the printed document to be identified using a printer source recognition model based on an attention mechanism. Finally, the printer source of each printed document is determined based on the target printer features.
[0057] It is understood that replacing traditional scanners with image acquisition devices makes printer source identification methods more valuable. Furthermore, combined with deep learning methods, the printer source identification model can better extract internal features from printed documents, thus avoiding the complexity of manual feature design and extraction. Furthermore, to address the traditional printer source identification method's reliance on specific single characters, the disclosed embodiments expand the research scope from single characters to mixed characters, resolving the difficulty of recognizing Chinese printed documents and further expanding its scope of application.
[0058] In an embodiment of the present disclosure, edge detection can be performed on the image information to generate a binary edge map; straight line segments in the binary edge map can be detected by Hough transform, and the angle between each straight line segment and the horizontal direction can be calculated; and based on the angle with the highest frequency of occurrence, the inclination angle of the printed document to be identified can be determined, and the image information can be rotationally corrected according to the inclination angle.
[0059] In some exemplary embodiments, to better align with actual image capture devices, such as mobile phone cameras, printed documents are captured using hand-photographing. Because it's difficult to maintain a fixed and accurate position for the phone during hand-photographing, the captured image often appears tilted, significantly impacting character extraction. To address this issue, a Hough transform can be used to correct the document image.
[0060] Specifically, edge detection can first be performed on the acquired image information, identifying the outlines of objects or boundaries between different regions in the image. Using specific edge detection algorithms (such as the Sobel operator or the Canny operator), the grayscale value changes of each pixel in the image and its surrounding pixels can be analyzed. When the grayscale value change exceeds a certain threshold, the pixel is identified as an edge point. This process converts the image, which originally contains rich color and detailed information, into a binary edge map. In other words, pixels in the image are represented by only two colors: black and white, with white representing edge points and black representing non-edge points. Then, the Hough transform can be used to detect straight line segments in the binary edge map. Points in the image space are mapped to a parameter space, and intersections are found in the parameter space to determine the parameters of the geometric shape. For line detection, the Hough transform converts edge points in the image space into a polar coordinate parameter space and counts the number of points corresponding to different parameter combinations. The parameter combination with the highest number of points corresponds to a straight line segment in the image. After detecting the straight line segments, the angle between each line segment and the horizontal direction can be calculated. This angle reflects the tilt of the line segment in the image. Finally, a statistical analysis of the angles between all straight line segments can be performed to identify the most frequently occurring angle. This most frequently occurring angle can be used to determine the tilt angle of the document being recognized. Once the tilt angle is determined, an image rotation algorithm can be used to perform rotation correction on the original image based on this angle, restoring the document image to a horizontal position.
[0061] It can be understood that by correcting the document image, the accuracy of character extraction is improved, thereby providing high-quality image data for subsequent printer source identification.
[0062] Based on the above embodiments, in this embodiment, the extraction of multiple mixed characters of the printed document to be identified and the determination of the position information of each of the mixed characters in the image information include: performing morphological operations on the mixed characters in the image information, connecting the disconnected parts of each of the mixed characters, and generating multiple connected character areas; performing horizontal projection and vertical projection on each of the connected character areas, respectively, to determine the upper and lower boundaries and left and right boundaries of each of the mixed characters; and determining the position information of each of the mixed characters in the image information based on the upper and lower boundaries and the left and right boundaries.
[0063] In the actual character extraction process, since Chinese characters have a more complex structure compared to English characters, there are various ways to form Chinese characters. Some Chinese characters are composed of radicals and components, making the radicals and components not connected to each other. For example, the character "休" is composed of two parts, "亻" and "木". In a printed document image, due to the influence of various factors, there may be an obvious break between "亻" and "木", resulting in a complete Chinese character being cut into two parts during the character extraction process, and thus a complete Chinese character cannot be extracted.
[0064] In view of this situation, embodiments of the present disclosure provide a method for character extraction.
[0065] As Figure 4 shown, after performing Canny edge detection on the corrected picture, morphological operations can be performed on the mixed characters in the image information to connect the unconnected parts of the left and right parts of Chinese characters. Specifically, the dilation operation in morphology can be used to select an appropriate structuring element to process the Chinese character region in the image, expand the region of character strokes, and reconnect the originally disconnected parts, thereby generating multiple connected character regions. Then, horizontal projection and vertical projection can be respectively performed on each connected character region by the projection segmentation method. Among them, horizontal projection can integrate the image in the horizontal direction to obtain the sum of pixel values of each row, forming a one-dimensional projection vector; vertical projection can integrate the image in the vertical direction to obtain the sum of pixel values of each column, also forming a one-dimensional projection vector. By analyzing the horizontal projection vector, the pixel distribution of characters in the vertical direction can be observed. When there are obvious peaks in the projection vector, it means that there are character strokes at that position, and the starting and ending positions of the peaks correspond to the upper and lower boundaries of the character. Similarly, the vertical projection vector can reflect the pixel distribution of characters in the horizontal direction, and the starting and ending positions of its peaks correspond to the left and right boundaries of the character. Therefore, by performing horizontal projection and vertical projection on each connected character region, the upper and lower boundaries and left and right boundaries of each mixed character can be determined more accurately. After determining the upper and lower boundaries and left and right boundaries of the characters, these boundary information can be converted into coordinate information in the image, so as to clarify the specific position of each character in the image. For example, a rectangular coordinate system can be established with the upper left corner of the image as the origin, and the row coordinates corresponding to the upper and lower boundaries of the character and the column coordinates corresponding to the left and right boundaries are recorded as the position information of the character. Finally, based on this position information, Chinese characters can be extracted.
[0066] It can be understood that by performing morphological operations and projection analysis on the image, Chinese characters can be extracted completely, avoiding complex feature extraction and classification processes, simplifying the processing flow, and improving the processing efficiency.
[0067] On the basis of the above embodiments, in this embodiment, according to the position information of each of the mixed characters, a pre-trained printer source recognition model is used to extract the target printer features corresponding to the printed document to be identified based on the attention mechanism, including: inputting the image information into the printer source recognition model, and extracting a multi-scale feature map according to the position information of each of the mixed characters; calculating the channel weight value for each feature channel of the multi-scale feature map based on the channel attention mechanism; and determining the target printer features based on the channel weight value of each of the feature channels.
[0068] In actual printer source recognition scenarios, especially when processing printed documents containing mixed Chinese characters, due to the large differences between mixed Chinese characters, these differences often bring a large amount of interference information, causing the printer source recognition model to focus on the differences between Chinese characters and ignore the key features that can truly distinguish different printer sources, that is, the differences between printers, which in turn hinders the improvement of recognition accuracy and leads to recognition errors.
[0069] To address this issue and improve recognition accuracy for mixed Chinese characters, the disclosed embodiment introduces a channel attention mechanism to adjust feature channels. Instead of averaging the feature information of all channels in a feature map, the algorithm assigns different weights to each channel of the feature map. Based on these weights, the model selectively enhances useful information features and suppresses less useful ones. In other words, the model's attention is focused on useful features while suppressing useless ones, thereby reducing the impact of interfering information and focusing on mining discernible printer feature information.
[0070] Figure 5 A framework diagram of printer source identification according to an embodiment of the present disclosure is schematically shown.
[0071] like Figure 5As shown, first, the document is input, and the image information containing mixed characters is input into a pre-trained printer source recognition model. Then, based on the position information of each mixed character, the model can perform multi-scale feature extraction operations on the image, extract various features from the document, such as keywords, word frequencies, and text structures, and obtain characters such as "类", "来"... "景". After obtaining the multi-scale feature map, for each feature channel of the multi-scale feature map, the model can calculate its corresponding channel weight value based on the channel attention mechanism to quantify the importance degree of each feature channel. Furthermore, based on the channel weight values of each feature channel, the model can perform "screening" and "focusing" among numerous feature information to determine the target printer features that can accurately represent the printer characteristics. Then, based on the extracted target printer features, the document is classified to output preliminary recognition results (i.e., Identification Results). For multiple recognition results, majority voting (i.e., Majority Vote) can be performed to improve the accuracy and reliability of classification. Finally, according to the results of majority voting, the final category of the document (i.e., Output Class) can be determined and output to achieve automatic classification of the input document.
[0072] It can be understood that by introducing the channel attention mechanism, the target printer features that can accurately represent the printer characteristics are obtained. These target printer features can be used as an important basis for subsequent classification or recognition tasks, which helps to improve the accuracy and reliability of printer source recognition.
[0073] Based on the above embodiments, in this embodiment, calculating the channel weight value for each feature channel of the multi-scale feature map based on the channel attention mechanism includes: performing global average pooling on the spatial information of each feature channel to generate a channel descriptor vector; inputting the channel descriptor vector into a one-dimensional convolutional layer to obtain the output of the one-dimensional convolutional layer; performing linear transformation and non-linear activation on the output of the one-dimensional convolutional layer based on a fully connected layer to obtain the output of the fully connected layer; and calculating the output of the fully connected layer through an activation function to obtain the channel weight values corresponding to each feature channel.
[0074] In order to fully learn the weight value of each channel, the embodiments of the present disclosure have made some adjustments to the classic squeeze-and-excitation network (i.e., SENet) and proposed a new attention module UCA (Useful ChannelAttention).
[0075] Figure 6 Schematically shows the structural diagram of the channel attention mechanism according to the embodiments of the present disclosure.
[0076] As Figure 6As shown, its main idea is similar to SENet. First, global average pooling can be used to aggregate the spatial information mapped by the input feature (i.e., Input Feature) to generate a feature descriptor vector, which represents the mean pooled feature. However, unlike the method of using only the fully connected layer in SENet, the embodiment of the present disclosure introduces a one-dimensional convolution with a convolution kernel size of k while using the fully connected layer to combine with the fully connected layer. In this process, the feature descriptor vector enters the one-dimensional convolution layer, and local information interaction is carried out between feature channels. The output of the one-dimensional convolution then enters the fully connected layer. The fully connected layer performs linear transformation and nonlinear activation on the input features to achieve further abstraction and transformation of the features. After processing by the one-dimensional convolution and the fully connected layer, the obtained UCA refined feature (UCA Refined Feature) is output.
[0077] After obtaining the aggregated channel features, we can calculate them through activation functions (such as the Sigmoid function) to obtain the weight of each channel. The Sigmoid function can map the input value to between 0 and 1, and the output value can be regarded as the weight of each channel.
[0078] It can be understood that by combining one-dimensional convolutional layers and fully connected layers, the information interaction between channels is enhanced. Compared with using only fully connected layers, it can better aggregate channel features and capture the dependencies between channels, ultimately increasing the model's attention to important features, thereby improving the overall recognition performance.
[0079] Based on the above embodiment, in this embodiment, the target printer characteristics are determined based on the channel weight value of each feature channel, including: multiplying the channel weight value with the feature channel to obtain a weighted feature map; and determining the target printer characteristics based on the weighted feature map.
[0080] For example, if the weight of a feature channel is 0.8 and the weight of another feature channel is 0.2, after multiplication with the feature channel, the features of the feature channel with a weight of 0.8 will be retained and enhanced to a greater extent, while the features of the feature channel with a weight of 0.2 will be weakened accordingly.
[0081] After obtaining the weighted feature map, the weighted feature map can be comprehensively analyzed and processed to extract more representative and distinctive target printer features to better reflect the unique properties of different printers, such as the printer's toner characteristics, the print head's printing mode, and the paper feeding method.
[0082] It can be understood that by weighting the feature channels, interference information can be effectively filtered out, so that the model can focus more on the features that contribute to the recognition task, thereby improving the accuracy and reliability of recognition.
[0083] In some exemplary embodiments, determining the printer source of the printed document to be identified based on the target printer features includes: inputting the target printer features into a classifier and obtaining a printer source probability distribution output by the classifier; and determining the printer source according to the category corresponding to the maximum value of the printer source probability distribution.
[0084] For example, the target printer features can be input into a pre-trained classifier, which can perform in-depth analysis and processing on these features and output a printer source probability distribution. This probability distribution can be a vector, where each dimension of the vector corresponds to a possible printer source category, and the value on each dimension represents the probability that the target printer features belong to that category. For example, if there are five different printer source categories, then the output probability distribution is a five-dimensional vector, where each element in the vector ranges from 0 to 1, and the sum of all elements is 1. Furthermore, the maximum value in the probability distribution can be found, and the category corresponding to the maximum value is the determined printer source. For example, if the probability value of a category in the probability distribution is 0.8, while the probability values of other categories are all smaller, then it can be determined that the printer source of the printed document to be identified is the category corresponding to that probability value.
[0085] It can be understood that by analyzing the target printer features using a classifier, the printer source can be accurately determined, thereby significantly improving the accuracy and stability of printer source identification.
[0086] Figure 7 The structure block diagram of the device for identifying the printer source based on the channel attention mechanism according to an embodiment of the present disclosure is schematically shown.
[0087] like Figure 7 As shown, the printer source identification device 700 based on the channel attention mechanism according to this embodiment includes an image acquisition module 710, a character extraction module 720, a feature extraction module 730 and a document recognition module 740.
[0088] The image acquisition module 710 is used to obtain image information of the printed document to be identified using an image acquisition device. In one embodiment, the image acquisition module 710 can be used to perform the operation S210 described above, which will not be described in detail here.
[0089] The character extraction module 720 is used to extract multiple mixed characters of the printed document to be recognized based on the image information and determine the position information of each mixed character in the image information. In one embodiment, the character extraction module 720 can be used to perform the operation S220 described above, which will not be repeated here.
[0090] Feature extraction module 730 is configured to extract target printer features corresponding to the printed document to be identified based on the position information of each mixed character and a pre-trained printer source identification model using an attention mechanism. In one embodiment, feature extraction module 730 can be configured to perform operation S230 described above, which will not be further described here.
[0091] The document identification module 740 is used to determine the printer source of the print document to be identified based on the target printer characteristics. In one embodiment, the document identification module 740 can be used to perform the operation S240 described above, which will not be repeated here.
[0092] In an embodiment of the present disclosure, the image acquisition module 710 can also be used to: perform edge detection on the image information to generate a binary edge map; detect straight line segments in the binary edge map through Hough transform, and calculate the angle between each straight line segment and the horizontal direction; and determine the inclination angle of the printed document to be identified based on the angle with the highest frequency of occurrence, and perform rotation correction on the image information according to the inclination angle.
[0093] In an embodiment of the present disclosure, the character extraction module 720 is specifically used to: perform morphological operations on the mixed characters in the image information, connect the disconnected parts of each of the mixed characters, and generate multiple connected character areas; perform horizontal projection and vertical projection on each of the connected character areas, respectively, to determine the upper and lower boundaries and left and right boundaries of each of the mixed characters; and determine the position information of each of the mixed characters in the image information based on the upper and lower boundaries and the left and right boundaries.
[0094] In an embodiment of the present disclosure, the feature extraction module 730 is specifically used to: input the image information into the printer source recognition model, extract a multi-scale feature map based on the position information of each mixed character; calculate the channel weight value for each feature channel of the multi-scale feature map based on the channel attention mechanism; and determine the target printer feature based on the channel weight value of each feature channel.
[0095] In an embodiment of the present disclosure, the feature extraction module 730 can also be used to: perform global average pooling on the spatial information of each of the feature channels to generate a channel descriptor vector; input the channel descriptor vector into a one-dimensional convolutional layer to obtain the output of the one-dimensional convolutional layer; based on a fully connected layer, perform linear transformation and nonlinear activation on the output of the one-dimensional convolutional layer to obtain the output of the fully connected layer; and calculate the output of the fully connected layer through an activation function to obtain the channel weight value corresponding to each of the feature channels.
[0096] In an embodiment of the present disclosure, the feature extraction module 730 may also be configured to: multiply the channel weight value by the feature channel to obtain a weighted feature map; and determine the target printer feature based on the weighted feature map.
[0097] In an embodiment of the present disclosure, the document identification module 740 is specifically used to: input the target printer features into a classifier and obtain the printer source probability distribution output by the classifier; and determine the printer source according to the category corresponding to the maximum value of the printer source probability distribution.
[0098] According to embodiments of the present disclosure, any multiple modules among the image acquisition module 710, character extraction module 720, feature extraction module 730, and document recognition module 740 may be combined into a single module, or any one of these modules may be split into multiple modules. Alternatively, at least part of the functionality of one or more of these modules may be combined with at least part of the functionality of other modules and implemented in a single module. According to embodiments of the present disclosure, at least one of the image acquisition module 710, character extraction module 720, feature extraction module 730, and document recognition module 740 may be at least partially implemented as a hardware circuit, such as a field programmable gate array (FPGA), a programmable logic array (PLA), a system on a chip, a system on a substrate, a system on a package, an application-specific integrated circuit (ASIC), or may be implemented in hardware or firmware through any other reasonable means of circuit integration or packaging, or may be implemented in any one of the three implementation methods of software, hardware, and firmware, or any appropriate combination of any of these. Alternatively, at least one of the image acquisition module 710, the character extraction module 720, the feature extraction module 730 and the document recognition module 740 may be at least partially implemented as a computer program module, which may perform corresponding functions when executed.
[0099] Figure 8 A block diagram of an electronic device for a method for identifying a printer source based on a channel attention mechanism according to an embodiment of the present disclosure is schematically shown.
[0100] like Figure 8As shown, an electronic device 800 according to an embodiment of the present invention includes a processor 801, which can perform various appropriate actions and processes based on programs stored in a read-only memory (ROM) 802 or programs loaded from a storage unit 808 into a random access memory (RAM) 803. The processor 801 may include, for example, a general-purpose microprocessor (e.g., a CPU), an instruction set processor and / or related chipsets and / or a special-purpose microprocessor (e.g., an application-specific integrated circuit (ASIC)), etc. The processor 801 may also include onboard memory for caching purposes. The processor 801 may include a single processing unit or multiple processing units for performing different actions of the method flow according to an embodiment of the present invention.
[0101] Various programs and data required for the operation of the electronic device 800 are stored in the RAM 803. The processor 801, ROM 802, and RAM 803 are connected to each other via a bus 804. The processor 801 executes the programs in the ROM 802 and / or RAM 803 to perform various operations according to the method flow of the embodiment of the present invention. It should be noted that the programs may also be stored in one or more memories other than the ROM 802 and RAM 803. The processor 801 may also execute the programs stored in the one or more memories to perform various operations according to the method flow of the embodiment of the present invention.
[0102] According to an embodiment of the present invention, electronic device 800 may further include an input / output (I / O) interface 805, which is also connected to bus 804. Electronic device 800 may also include one or more of the following components connected to I / O interface 805: an input section 806 including a keyboard, mouse, etc.; an output section 807 including devices such as a cathode ray tube (CRT), liquid crystal display (LCD), and speakers; a storage section 808 including a hard disk; and a communication section 809 including a network interface card such as a LAN card or modem. Communication section 809 performs communication processing via a network such as the Internet. A drive 810 is also connected to I / O interface 805 as needed. Removable media 811, such as a magnetic disk, optical disk, magneto-optical disk, semiconductor memory, etc., is installed in drive 810 as needed, so that computer programs read from the removable media can be installed into storage section 808 as needed.
[0103] The present invention also provides a computer-readable storage medium, which may be included in the device / apparatus / system described in the above embodiments, or may exist independently and not incorporated into the device / apparatus / system. The computer-readable storage medium carries one or more programs, which, when executed, implement the method according to the embodiments of the present invention.
[0104] According to an embodiment of the present invention, a computer-readable storage medium may be a non-volatile computer-readable storage medium, and may include, for example, but not limited to: a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination thereof. In the present invention, a computer-readable storage medium may be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution system, apparatus, or device. For example, according to an embodiment of the present invention, a computer-readable storage medium may include the ROM 802 and / or RAM 803 described above, and / or one or more memories other than ROM 802 and RAM 803.
[0105] Embodiments of the present disclosure also include a computer program product, comprising a computer program containing program code for executing the method shown in the flowchart. When the computer program product is executed in a computer system, the program code is used to cause the computer system to implement the printer source identification method based on the channel attention mechanism provided in the embodiments of the present disclosure.
[0106] The computer program executes the above functions defined in the system / device of the embodiment of the present disclosure when the computer program is executed by the processor 801. According to the embodiment of the present disclosure, the system, device, module, unit, etc. described above can be implemented by a computer program module.
[0107] In one embodiment, the computer program may be stored on a tangible storage medium such as an optical storage device or a magnetic storage device. In another embodiment, the computer program may be transmitted and distributed in the form of a signal on a network medium, downloaded and installed via the communication portion 809, and / or installed from a removable medium 811. The program code contained in the computer program may be transmitted using any appropriate network medium, including but not limited to wireless, wired, or any suitable combination thereof.
[0108] In such an embodiment, the computer program can be downloaded and installed from a network via the communication section 809, and / or installed from a removable medium 811. When the computer program is executed by the processor 801, the above-described functions defined in the system of the embodiment of the present disclosure are performed. According to the embodiment of the present disclosure, the systems, devices, means, modules, units, etc. described above can be implemented by computer program modules.
[0109] According to an embodiment of the present disclosure, the program code for executing the computer program provided by the embodiment of the present disclosure can be written in any combination of one or more programming languages. Specifically, these computer programs can be implemented using high-level procedural and / or object-oriented programming languages, and / or assembly / machine languages. Programming languages include, but are not limited to, languages such as Java, C++, Python, "C" or similar programming languages. The program code can be executed entirely on the user computing device, partially on the user device, partially on a remote computing device, or entirely on a remote computing device or server. In cases involving a remote computing device, the remote computing device can be connected to the user computing device through any type of network, including a local area network (LAN) or a wide area network (WAN), or can be connected to an external computing device (for example, using an Internet service provider to connect via the Internet).
[0110] The flowcharts and block diagrams in the accompanying drawings illustrate the possible implementation architecture, functions and operations of the systems, methods and computer program products according to various embodiments of the present disclosure. In this regard, each box in the flowchart or block diagram can represent a module, program segment, or a part of code, and the above-mentioned module, program segment, or a part of code contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in an order different from that marked in the accompanying drawings. For example, two boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram or flowchart, and the combination of boxes in the block diagram or flowchart, can be implemented with a dedicated hardware-based system that performs the specified function or operation, or can be implemented with a combination of dedicated hardware and computer instructions.
[0111] Those skilled in the art will appreciate that the features described in the various embodiments of the present disclosure may be combined and / or coupled in various ways, even if such combinations or couplings are not explicitly described in the present disclosure. In particular, the features described in the various embodiments of the present disclosure may be combined and / or coupled in various ways without departing from the spirit and teachings of the present disclosure. All such combinations and / or couplings fall within the scope of the present disclosure.
[0112] The above describes the embodiments of the present disclosure. However, these embodiments are for illustrative purposes only and are not intended to limit the scope of the present disclosure. Although each embodiment has been described separately above, this does not mean that the measures in each embodiment cannot be advantageously used in combination. Without departing from the scope of the present disclosure, those skilled in the art may make various substitutions and modifications, which should all fall within the scope of the present disclosure.
Claims
1. A printer source identification method based on channel attention mechanism, characterized in that: The method comprises: Using an image acquisition device, image information of a printed document to be identified is acquired; extracting a plurality of mixed characters of the printed document to be recognized based on the image information, and determining position information of each of the mixed characters in the image information; According to the position information of each of the mixed characters, using a pre-trained printer source recognition model, extracting target printer features corresponding to the printed document to be recognized based on an attention mechanism; and Based on the target printer characteristics, a printer source of the to-be-identified printed document is determined.
2. The method according to claim 1, characterized in that Also includes: Performing edge detection on the image information to generate a binary edge map; Detecting straight line segments in the binary edge image by Hough transform, and calculating the angle between each straight line segment and the horizontal direction; and Based on the angle with the highest frequency of occurrence, the tilt angle of the printed document to be identified is determined, and the image information is rotationally corrected according to the tilt angle.
3. The method according to claim 1 or 2, characterized in that The step of extracting a plurality of mixed characters from the printed document to be recognized and determining position information of each of the mixed characters in the image information includes: performing morphological operations on the mixed characters in the image information, connecting disconnected parts of each of the mixed characters to generate a plurality of connected character regions; Performing horizontal projection and vertical projection on each of the connected character regions to determine the upper and lower boundaries and the left and right boundaries of each of the mixed characters; and Based on the upper and lower boundaries and the left and right boundaries, position information of each of the mixed characters in the image information is determined.
4. The method according to claim 1, wherein The method of extracting target printer features corresponding to the printed document to be identified based on the attention mechanism using a pre-trained printer source identification model according to the position information of each mixed character includes: Inputting the image information into the printer source recognition model, and extracting a multi-scale feature map according to the position information of each mixed character; For each feature channel of the multi-scale feature map, calculating a channel weight value based on a channel attention mechanism; and The target printer feature is determined based on the channel weight value of each feature channel.
5. The method according to claim 4, characterized in that Calculating a channel weight value for each feature channel of the multi-scale feature map based on a channel attention mechanism includes: Performing global average pooling on the spatial information of each feature channel to generate a channel descriptor vector; Inputting the channel descriptor vector into a one-dimensional convolutional layer to obtain an output of the one-dimensional convolutional layer; Based on the fully connected layer, performing linear transformation and nonlinear activation on the output of the one-dimensional convolutional layer to obtain the output of the fully connected layer; and The output of the fully connected layer is calculated through an activation function to obtain a channel weight value corresponding to each feature channel.
6. The method according to claim 4, characterized in that The determining of the target printer characteristics based on the channel weight value of each characteristic channel includes: Multiplying the channel weight value by the feature channel to obtain a weighted feature map; and The target printer features are determined according to the weighted feature map.
7. The method according to claim 1, characterized in that Determining the printer source of the to-be-identified printed document based on the target printer characteristics includes: Inputting the target printer feature into a classifier and obtaining the printer source probability distribution output by the classifier; and The printer source is determined according to the category corresponding to the maximum value of the printer source probability distribution.
8. A printer source identification device based on a channel attention mechanism, characterized in that: The device comprises: An image acquisition module is used to acquire image information of a printed document to be identified using an image acquisition device; a character extraction module, configured to extract a plurality of mixed characters of the printed document to be recognized based on the image information, and determine position information of each of the mixed characters in the image information; a feature extraction module, configured to extract target printer features corresponding to the printed document to be identified based on the position information of each mixed character and using a pre-trained printer source recognition model based on an attention mechanism; and The document identification module is used to determine the printer source of the printed document to be identified based on the target printer characteristics.
9. An electronic device comprising: one or more processors; a storage device for storing one or more computer programs, It is characterized in that the one or more processors execute the one or more computer programs to implement the steps of the method according to any one of claims 1 to 7.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 7 are implemented.
11. A computer program product comprising a computer program, characterized in that When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 7 are implemented.