Automatic printing equipment traceability method based on deep learning
Through the automated printing equipment traceability method based on deep learning, the problems of dark mark image rotation and scale changes, and the absence of noise points and dark mark points in the prior art are solved, and efficient and accurate traceability of printing equipment is achieved, improving traceability accuracy and recognition accuracy.
Patent Information
- Application Number
- CN202411883238.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-19
- Publication Date
- 2025-05-06
AI Technical Summary
In the prior art, it is difficult to effectively trace the printing equipment due to rotation and scale changes of dark mark images caused by changes in image acquisition operations and acquisition equipment, as well as the lack of noise points and dark mark points caused by printing quality.
The traceability method of automatic printing equipment based on deep learning is adopted, and the traceability classification of printing equipment is obtained and classified by obtaining and classifying printing equipment data samples, image preprocessing is performed to extract dark mark features, and the intelligent learning method and KNN nearest neighbor classification method are used to realize traceability classification of printing equipment.
The accuracy and efficiency of the traceability of printing equipment has been improved, the efficiency, comprehensiveness and accuracy of identification have been improved, and the efficient traceability of printing equipment has been achieved. The traceability accuracy has been improved to 98.53%, and the overall recognition accuracy has been increased to 95.62%.
Smart Images

Figure CN119942569A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer vision processing technology, and in particular to a deep learning-based automated printing equipment tracing method. Background Art
[0002] Printed document tracing technology is a key application in the field of forensic science and criminal investigation. It is an important research direction in the field of court document inspection and an important technology in combating various types of criminal activities.
[0003] At present, document traceability methods based on visual technology can be roughly divided into two categories, namely passive traceability technology and active traceability technology. The former uses the technical characteristics of the printing equipment and the differences in production processes during the printing process to form certain specific features of printed documents for traceability. The latter uses specific technologies or methods to embed certain "hidden" subtle graphic features in the printed graphic content during the printing process for traceability, and the machine identification code is one of the most widely used active traceability technologies by printing equipment manufacturers.
[0004] There are several problems in the prior art regarding the machine identification code formed during the printing process of a color laser printing device. Among them, the rotation and scale changes of the hidden image caused by changes in the image acquisition operation and the acquisition device, and the noise points and missing hidden points caused by the printing quality make it difficult to trace the source printing device. Summary of the invention
[0005] The present invention provides an automated printing equipment tracing method based on deep learning, which is used to solve the defects in the prior art of rotation and scale change of hidden image caused by image acquisition operation and change of acquisition equipment, as well as noise points and missing hidden points caused by printing quality.
[0006] The present invention provides a method for tracing the source of an automated printing device based on deep learning, comprising the following steps:
[0007] Step 1, obtaining N types of printing device data samples, classifying and cutting the N types of printing device data samples to obtain a benchmark data set;
[0008] Step 2, performing image preprocessing on the reference data set to obtain a first sample data set, where the first sample data set is a black-and-white image data set of the N types of printing device data samples contained in the reference data set;
[0009] Step 3, based on one of the N types of printing device data samples contained in the first sample data set, using an intelligent learning method to extract a complete secret set, obtain neighborhood point pattern features, and then obtain training set features to obtain a traceability classification benchmark data set;
[0010] Step 4: compare the elements in the test set with the traceability classification benchmark data set, and use the KNN nearest neighbor classification method to obtain the source printing equipment traceability classification result.
[0011] Wherein, the step 1 further comprises:
[0012] Step 11, obtaining the N types of printing device data samples, automatically scanning the N types of printing device data samples based on the TWAIN protocol, and obtaining N types of scanned sample images corresponding to the N types of printing device data samples;
[0013] Step 12: randomly crop the N types of scanned sample images into k images to obtain N×k samples to form the benchmark data set.
[0014] Wherein, the step 2 further comprises:
[0015] Step 21, obtaining the reference data set, decomposing the reference data set by color channel, and obtaining a blue color channel image with the most obvious hidden mark features;
[0016] Step 22, performing an inverting and binarization process on the blue color channel image to obtain a black and white image of the imaging secret mark;
[0017] Step 24, performing morphological operations on the black and white image to obtain a clear dark image;
[0018] Step 25: Based on the clear dark mark image, denoise each pixel in the image by traversing the image to obtain the first sample data set.
[0019] Wherein, the step 3 further comprises:
[0020] Step 31, obtaining one type of data from the N types of printing device data samples contained in the first sample data set, calculating the spatial coordinates of all hidden points in the one type of data, and storing the coordinate data to obtain a hidden data file;
[0021] Step 32, based on the hidden data file, using the neighborhood point pattern NPP to perform scale normalization and rotation consistency on the hidden local features to obtain the neighborhood point pattern features of the hidden data file;
[0022] Step 33, based on the hidden data file and the corresponding neighborhood point pattern features, obtaining repeated hidden points by comparing the neighborhood point pattern features of the hidden points in the hidden data file;
[0023] Step 34, based on the repeated hidden points, using the hidden point association learning method of the greedy strategy, calculate the association relationship of all the hidden points, and determine the maximum repeated hidden point set as the hidden pattern set, wherein the maximum repeated hidden point set is the point set containing all the repeated hidden points;
[0024] Step 35, traverse the data in the first sample data set that are not detected because the neighborhood points are beyond the image range, calculate the correlation between the undetected data and other dark mark points, and supplement the dark mark pattern set to obtain a complete dark mark pattern set.
[0025] Wherein, the step 35 further comprises:
[0026] Step 351, obtaining a complete set of hidden patterns corresponding to the N types of printing device data samples;
[0027] Step 352, based on the complete dark pattern set corresponding to the N types of printing device data samples, randomly select the dark pattern set of images for the N categories as a training set, and the dark pattern set of the remaining sample images as a test set;
[0028] Step 353, obtaining the neighborhood point pattern features in the training set, storing the neighborhood point pattern features in a database, and obtaining the traceability classification benchmark data set.
[0029] Wherein, the step 4 further comprises:
[0030] Step 41, obtaining the test set and the traceability classification benchmark data set;
[0031] Step 42, for each element in the test set, traverse the traceability classification benchmark data set, use the KNN nearest neighbor method to obtain the distance relationship between the test set and the traceability classification benchmark data set, and determine the similarity result between each element in the test set and the traceability classification benchmark data set;
[0032] Step 43: classify the traceability classification benchmark data set to obtain the source printing device traceability classification result.
[0033] Wherein, the step 42 further includes:
[0034] Step 421, based on the test set, traverse the traceability classification benchmark data set;
[0035] Step 422, taking the test set as a benchmark, calculating the distance between each element in the test set and each element in the traceability classification benchmark data set to obtain a first distance;
[0036] Step 423, taking the traceability classification benchmark data set as a benchmark, calculating the distance between each element in the test set and each element in the traceability classification benchmark data set to obtain a second distance;
[0037] Step 424, traverse the first distance and the second distance corresponding to each element in the test set, calculate the difference between the first distance and the second distance, and obtain the similarity result;
[0038] Step 425, based on the similarity result, obtain multiple elements in the traceability classification benchmark data set that are closest to the test set, use the majority voting method to perform source printing device traceability classification on the test sample image, and obtain the source printing device traceability classification result.
[0039] The present invention also provides a deep learning-based automated printing equipment source tracing detection device, characterized in that the device comprises:
[0040] A data acquisition module, used to obtain N types of printing device data samples, classify and trim the N types of printing device data samples, and obtain a benchmark data set;
[0041] A data preprocessing module, configured to perform image preprocessing on the reference data set to obtain a first sample data set, wherein the first sample data set is a black-and-white image data set of the N types of printing device data samples contained in the reference data set;
[0042] A model training module, based on one of the N types of printing device data samples contained in the first sample data set, uses an intelligent learning method to extract a complete secret set, obtain neighborhood point pattern features, and then obtain training set features to obtain a traceability classification benchmark data set;
[0043] The data tracing module uses the data not randomly selected in the data enhancement module as a test set, compares the elements in the test set with the traceability classification benchmark data set, and uses the KNN nearest neighbor classification method to obtain the source printing equipment traceability classification result.
[0044] The present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, wherein when the computer program is executed by a processor, the deep learning-based automated printing equipment tracing method as described in any one of claims 1 to 7 is implemented.
[0045] The present invention also provides a computer program product, including a computer program, characterized in that when the computer program is executed by a processor, it implements the deep learning-based automated printing equipment tracing method as described in any one of claims 1 to 7.
[0046] Specifically, the method and apparatus of the present invention combine a variety of image recognition strategies with sample characteristics of printing devices to form a complete model training and sample traceability strategy, and utilize computer pattern recognition technology and automatic comparison technology to identify and compare the tracking secret features of color laser printing devices, and automatically obtain relevant information of color laser printing devices, thereby greatly improving the efficiency, comprehensiveness and accuracy of recognition.
[0047] The method and device of the present invention improve the traceability accuracy to 98.53% and the overall recognition accuracy to 95.62% by combining deep learning and pattern recognition, which greatly improves the traceability accuracy and overall recognition accuracy of the printing device and realizes efficient traceability of the printing device. BRIEF DESCRIPTION OF THE DRAWINGS
[0048] In order to more clearly illustrate the technical solutions in the present invention or the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.
[0049] Figure 1 It is a flowchart of the method for active source tracing of automated source printing equipment based on deep learning provided by the present invention.
[0050] Figure 2 The present invention provides a flowchart of image preprocessing of a method for active source tracing of an automated source printing device.
[0051] Figure 3 It is a diagram of hidden feature points and hidden image formed by printing devices of different brands and models provided by the present invention.
[0052] Figure 4 It is a schematic diagram of scale normalization of a method for active source tracing of an automated source printing device provided by the present invention.
[0053] Figure 5 It is a rotational uniformity diagram of a method for active source tracing of an automated source printing device provided by the present invention.
[0054] Figure 6 It is a schematic diagram of extracting different types of hidden point sets in the method for active source tracing of automated source printing equipment provided by the present invention. DETAILED DESCRIPTION
[0055] In order to make the purpose, technical solution and advantages of the present invention clearer, the technical solution of the present invention will be clearly and completely described below in conjunction with the drawings of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0056] Printed document traceability technology plays a vital role in the field of court document inspection and in combating criminal activities. It determines the brand and model of the source printing device by analyzing the characteristics of printed documents. At present, the document traceability methods based on vision technology are mainly divided into two categories, namely passive traceability technology and active traceability technology.
[0057] Active traceability technology achieves traceability by identifying subtle graphic features embedded in the printing process of documents, or by embedding specific identification information in printed documents. Using specific technologies or methods, some "hidden" subtle graphic features embedded in the printed graphic content are obtained to track the source of the printing device.
[0058] With respect to the machine identification code formed during the printing process of the color laser printing device, the method of the present invention effectively solves the problems of rotation and scale change of the hidden image caused by the image acquisition operation and the change of the acquisition device, as well as the problems of noise points and missing hidden points caused by the printing quality. Through the introduction of deep learning and pattern recognition technology, the automated and intelligent comparison of hidden features is realized, which is not only time-saving and low-cost, but also greatly improves the accuracy and efficiency of traceability.
[0059] Deep learning achieves high-precision recognition and analysis of document images by building a multi-layer neural network model, improving the accuracy and efficiency of document inspection. The application of this technology not only reduces the impact of human intervention and subjective judgment, but also provides a new technical means for printed document traceability. The development of printed document traceability technology, especially the progress in automated and intelligent secret feature comparison, provides strong technical support for court document inspection and combating crime, and has broad promotion value and application prospects.
[0060] Figure 1 is a flow chart of a method for tracing the source of an automated printing device based on deep learning provided by the present invention, such as Figure 1As shown, the method includes the following: step 1, obtaining N types of printing device data samples, classifying and trimming the N types of printing device data samples to obtain a benchmark data set; step 2, performing image preprocessing on the benchmark data set to obtain a first sample data set, wherein the first sample data set is a black-and-white image data set of hidden mark imaging of the N types of printing device data samples contained in the benchmark data set; step 3, based on one type of the N types of printing device data samples contained in the first sample data set, using an intelligent learning method to extract a complete hidden mark set, obtain neighborhood point pattern features, and then obtain training set features to obtain a traceability classification benchmark data set; step 4, comparing the elements in the test set with the traceability classification benchmark data set, and using the KNN nearest neighbor classification method to obtain the source printing device traceability classification result.
[0061] Wherein, the step 1 further comprises:
[0062] Step 11, obtaining the N types of printing device data samples, automatically scanning the N types of printing device data samples based on the TWAIN protocol, and obtaining N types of scanned sample images corresponding to the N types of printing device data samples;
[0063] Step 12: randomly crop the N types of scanned sample images into k images to obtain N×k samples to form the benchmark data set.
[0064] The N types of printing device data samples are paper samples printed by different types of color laser printing devices. Scanning is performed using a scanner device at a fixed DPI to remove blurry, incomplete, or color-distorted printed paper due to poor printing quality, and finally M printed paper scan images are collected. For each of the N types of printing device data samples, the scanned image is randomly cropped into k images.
[0065] Preferably, a scanner device is used to automatically scan several data samples of the N types of printing devices based on the TWAIN protocol under the scanning condition of 400 DPI, obtain several scanned sample images, and randomly crop the N types of scanned sample images into k images to obtain the benchmark data set, k≥30.
[0066] Wherein, the step 2 further comprises:
[0067] Step 21, obtaining the reference data set, decomposing the reference data set by color channel, and obtaining a blue color channel image with the most obvious hidden mark features;
[0068] Step 22, performing an inverting and binarization process on the blue color channel image to obtain a black and white image of the imaging secret mark;
[0069] Step 24, performing morphological operations on the black and white image to obtain a clear dark image;
[0070] Step 25: Based on the clear dark mark image, denoise each pixel in the image by traversing the image to obtain the first sample data set.
[0071] Figure 2 is the image preprocessing flowchart, consisting of Figure 2 It can be seen that the technical solution obtains the final dark mark image through channel decomposition, inverting processing, binarization, morphological operation, and denoising processing, thereby forming the first sample data set.
[0072] The decomposing the reference data set by color channel includes reading the scanned sample image containing the secret content in the reference data set, decomposing the scanned sample image by color channel, and obtaining separate red channel image, green channel image and blue channel image.
[0073] Furthermore, the blue color channel image with the most obvious dark mark features is selected, each pixel of the image is traversed, the brightness value of each pixel is subtracted by 255, and the dark mark in the image is highlighted by inverting processing to obtain an image of the displayed dark mark.
[0074] Furthermore, the image of the development secret mark is binarized, and after reading the image, each pixel is traversed to determine the relationship between the brightness value of each pixel and a threshold value. The threshold value is preferably set to 128, and the brightness value of the pixel with a brightness value greater than the set threshold is determined to be 255, and the brightness value of the pixel with a brightness value less than the set threshold is determined to be 0, and then the image is converted into a binary image with only black and white colors to obtain a black and white image of the development secret mark.
[0075] Furthermore, the black-and-white image of the dark mark is subjected to morphological operations, and the morphological operations include: corrosion and expansion, so as to obtain a dark mark image with clearer dark marks.
[0076] The corrosion operation refers to imaging the black and white image of the dark mark, traversing each pixel of the image, determining a 3×3 rectangle as a structural element with the current pixel as the center, that is, determining the area composed of the current pixel and its surrounding 8 neighboring pixels as the structural element, comparing all pixel values within the area covered by the structural element, assigning the minimum value within the area covered by the structural element to the pixel at the center of the structural element, thereby shrinking the white area in the image to reduce the size of the dark mark and make it more compact, thereby eliminating the burrs on the boundary and making the shape of the dark mark clearer.
[0077] The dilation operation refers to traversing each pixel of the black and white image of the hidden mark imaging, taking the current pixel as the center, determining a 3×3 rectangle as the structural element, that is, determining the area composed of the current pixel and its surrounding 8 neighboring pixels as the structural element, comparing all pixel values in the area covered by the structural element, and assigning the maximum value in the area covered by the structural element to the pixel at the center of the structural element, thereby expanding the white area in the image, making it more connected and highlighting the boundaries, thereby eliminating holes and small breaks, and facilitating the subsequent detection of hidden marks.
[0078] Furthermore, the black and white image of the dark mark after the erosion and dilation operations is subjected to denoising. After reading the image, each pixel is traversed, and the neighborhood around each pixel is checked to calculate the number of pixels occupied by the dark mark point. The dark mark point is determined to be within the set threshold range, otherwise it is determined to be a noise point, and the following is obtained: Figure 3 As shown, the black-and-white images of hidden marks after noise removal of printing devices of different brands and models; the benchmark data set is subjected to the image preprocessing operation of step 2 to obtain N×k black-and-white images of hidden mark imaging, namely the first sample data set.
[0079] Wherein, the step 3 further comprises:
[0080] Step 31, obtaining one type of data from the N types of printing device data samples contained in the first sample data set, calculating the spatial coordinates of all hidden points in the one type of data, and storing the coordinate data to obtain a hidden data file.
[0081] Step 32: Based on the hidden data file, the neighborhood point pattern NPP is used to perform scale normalization and rotation consistency on the hidden local features to obtain the neighborhood point pattern features of the hidden data file.
[0082] Step 33, based on the hidden data file and the corresponding neighborhood point pattern features, duplicate hidden points are obtained by comparing the neighborhood point pattern features of the hidden points in the hidden data file.
[0083] Step 34, based on the repeated hidden points, a hidden point association learning method with a greedy strategy is used to calculate the association relationship of all hidden points, and determine the maximum repeated hidden point set as the hidden pattern set, wherein the maximum repeated hidden point set is a point set containing all repeated hidden points.
[0084] Step 35, traverse the data in the first sample data set that are not detected because the neighborhood points are beyond the image range, calculate the correlation between the undetected data and other dark mark points, and supplement the dark mark pattern set to obtain a complete dark mark pattern set.
[0085] Among them, the calculation of the spatial coordinates of all dark mark points in the type of data includes: determining the spatial coordinates of each dark mark point in the black and white image, calculating the spatial coordinates of all dark mark points, and then storing the coordinates of each dark mark point in the format of an (x, y) coordinate pair, thereby obtaining a dark mark data file.
[0086] The step of determining the spatial coordinates of each dark mark point in the black-and-white image includes: obtaining the spatial coordinates p of a dark mark point according to a dark mark data file of a certain image; c =(x c ,y c ), the spatial distribution relationship between the 7 nearest points in the neighborhood centered on the hidden point and the corresponding hidden point is used to obtain the hidden point p c and its neighborhood point set That is, the neighborhood point pattern of the dark mark point, and then each dark mark point in the image is modeled to obtain the neighborhood point pattern of each dark mark point.
[0087] Among them, the neighborhood point pattern NPP refers to an image processing pattern recognition algorithm used for feature extraction and data dimension reduction. Based on the secret data file, the local features of the secret are scaled normalized and rotated uniformly based on the neighborhood point pattern, thereby achieving scale invariance.
[0088] Figure 4 is a schematic diagram of scale normalization, such as Figure 4 As shown, the scale normalization includes: for a certain hidden point p c and its neighborhood point set according to The calculation method is to normalize the scale of the neighborhood point pattern of the hidden point, where max(d c ) are all the nearest neighbor hidden points To the center point p c The maximum value of the distance; the scale normalization operation is performed to increase the robustness of the neighborhood point pattern to scale changes, so that the NPP operator has consistency for the hidden image features obtained by different image acquisition devices.
[0089] Figure 5 is a rotational conformity diagram, such as Figure 5 As shown, the rotation uniformity includes: for a certain hidden point p c and its neighborhood point set according to The calculation method performs a rotational uniformity operation on the neighborhood point pattern of the hidden point, where The nearest neighbor hidden point To the center point p c The distance is the nearest neighbor hidden point and the center point p c The angle between them, Δθ is the distance p c The farthest nearest neighbor hidden point and p c The angle between them; the set of dark points in the neighborhood space of each dark point is calibrated in the angle space through the rotation consistency operation, so that the NPP operator has better robustness to the rotation changes of the input image.
[0090] Among them, comparing the neighborhood point pattern features of the hidden point means: based on the characteristic that the hidden point set of a specific pattern repeats in a regular pattern, the points in the hidden point set have the same neighborhood point pattern features, and by determining whether the neighborhood point pattern features are the same, the repeated hidden points are located.
[0091] The determining of the neighborhood point pattern features specifically includes: traversing the normalized dark mark point pairs p in the dark mark image A and p B , and the neighborhood point set of the hidden point pair and
[0092] according to Method to calculate the hidden point p A and the secret point p B The distance difference of the spatial distribution of other hidden points in the neighborhood space, where Relative points in the neighborhood and The distance, ε NPP is a predefined distance threshold, is the indicator function; if D NPP (p A ,p B )=0, then the hidden point p A and the secret point p B The relative spatial positions of the seven neighboring points around it are the same, p a and p b They are two corresponding dark points with the same position in two repeated dark patterns, and then the repeated dark points with the same neighborhood point pattern characteristics and their neighborhood point sets in the dark point image are obtained.
[0093] Among them, the determining of the neighborhood point pattern characteristics also includes: associating related hidden points together through a hidden point association learning method based on a greedy strategy, and then finding the largest set of repeated hidden points, that is, a set containing as many hidden points as possible.
[0094] Figure 6Schematic diagram of extracting different types of hidden point sets. The hidden point association learning method is as follows: for the extracted repeated hidden point sets Where i = 1, 2, ..., K, is a repeated dark point with the same dark point pattern characteristics, M i is the number of repeated hidden points, K is the number of hidden point sets; calculate each repeated hidden point set The relative position pattern of any two points in 8 directions (ρ i ,θ i ), ρ i is the relative distance, θ i is the relative angle, through (ρ i ,θ i ) and the frequency of occurrence to predict the set of repeated hidden points The dark points that should appear but do not appear, or are not successfully detected, are removed, and the pseudo dark points that are misidentified due to noise are removed to obtain the final set of repeated dark points.
[0095] Furthermore, eight directions between any two points are determined based on the position vector between the two points, including eight basic directions on a two-dimensional plane, specifically, 0° direction: due right (horizontally to the right); 45° direction: upper right (diagonal direction); 90° direction: due top (vertically upward); 135° direction: upper left (diagonal direction); 180° direction: due left (horizontally to the left); 225° direction: lower left (diagonal direction); 270° direction: due bottom (vertically downward); 315° direction: lower right (diagonal direction).
[0096] Further, the secret pattern set is initialized is an empty set, select any set of repeated hidden points and will Each hidden point in is associated with the corresponding And according to Calculate the center point of all the hidden points to get the position of the hidden set, where M is the number of hidden points; for each hidden point in the unassociated repeated hidden point set, calculate the distance of each point in the set one by one The nearest set and associate its hidden points with the corresponding Update and the corresponding center point
[0097] Furthermore, for the situation in which the neighborhood points of the dark mark point image are beyond the image range and are not detected because the dark mark points are distributed at the edge of the picture, or some dark mark points are not recognized as repeated dark mark points due to noise, the dark mark pattern set Complete the code to obtain the complete secret code set.
[0098] Among them, the set of secret patterns To complete, specifically including: traversing the dark mark point image that has not yet been associated to The set of hidden points P unrelated Calculate the distance of all hidden points from the center point one by one distance, and then obtain the nearest hidden point And predict the set of repeated hidden points Among them, these hidden points and Relative to the center point have the same positional relationship; if in P unrelated If there are more than a certain threshold number of unassociated dark points in , and these dark points have the same positional relationship with the nearest point relative to the center point, then Each dark point in the point set and the dark pattern set Associate and update and the corresponding center point If in P unrelated If there are less than a certain threshold number of unassociated dark points in The points in are considered as noise points and updated Finally, a complete set of dark marks extracted based on the image of a specific dark mark point is obtained.
[0099] Wherein, the step 35 further comprises:
[0100] Step 351, obtaining a complete set of hidden patterns corresponding to the N types of printing device data samples;
[0101] Step 352, based on the complete dark pattern set corresponding to the N types of printing device data samples, randomly select the dark pattern set of images for the N categories as a training set, and the dark pattern set of the remaining sample images as a test set;
[0102] Step 353, obtaining the neighborhood point pattern features in the training set, storing the neighborhood point pattern features in a database, and obtaining the traceability classification benchmark data set.
[0103] Among them, for each type of sample image in the N types of printing device data samples, preferably, 10 sample images are randomly selected to obtain the corresponding complete hidden mark set and the neighborhood point pattern features, and the hidden mark pattern sets of the remaining sample images are used as the test set.
[0104] Among them, the traceability classification benchmark data set includes: the complete secret set and the neighborhood point pattern features.
[0105] Wherein, the step 4 further comprises:
[0106] Step 41, obtaining the test set and the traceability classification benchmark data set;
[0107] Step 42, for each element in the test set, traverse the traceability classification benchmark data set, use the KNN nearest neighbor method to obtain the distance relationship between the test set and the traceability classification benchmark data set, and determine the similarity result between each element in the test set and the traceability classification benchmark data set;
[0108] Step 43: classify the traceability classification benchmark data set to obtain the source printing device traceability classification result.
[0109] Among them, the test set is the set of hidden marks extracted from the hidden mark point image on the printed document, the traceability classification benchmark dataset is the complete secret set and the secret set formed by printing of a specific type of printing device composed of the neighborhood point pattern features; the distance relationship is, based on the test set The benchmark test set and benchmark datasets The distance and calculation are based on the benchmark dataset The benchmark test set and benchmark datasets The similarity result is that the test set With benchmark datasets The correspondence between the elements in the source printing device tracing classification result is the printing device tracing result determined according to the similarity result.
[0110] Wherein, the step 42 further includes:
[0111] Step 421, based on the test set, traverse the traceability classification benchmark data set;
[0112] Step 422, taking the test set as a benchmark, calculating the distance between each element in the test set and each element in the traceability classification benchmark data set to obtain a first distance;
[0113] Step 423, taking the traceability classification benchmark data set as a benchmark, calculating the distance between each element in the test set and each element in the traceability classification benchmark data set to obtain a second distance;
[0114] Step 424, traverse the first distance and the second distance corresponding to each element in the test set, calculate the difference between the first distance and the second distance, and obtain the similarity result;
[0115] Step 425, based on the similarity result, obtain multiple elements in the traceability classification benchmark data set that are closest to the test set, use the majority voting method to perform source printing device traceability classification on the test sample image, and obtain the source printing device traceability classification result.
[0116] Among them, the first distance is the test set The benchmark test set and benchmark datasets Distance, specifically:
[0117]
[0118] The method is calculated based on The base secret set and Secret Records The distance for and The distance of other hidden points in the neighborhood space, ε MIC is the pre-set distance threshold, is the indicator function, N A Secret code collection The number of hidden points in .
[0119] Among them, the second distance is the benchmark data set The distance and calculation are based on the benchmark dataset The benchmark test set and benchmark datasets Distance, specifically:
[0120]
[0121] The method is calculated based on The base secret set and Secret Records distance.
[0122] The specific calculation method of the similarity result is as follows:
[0123]
[0124] Calculate the secret set of the test sample A collection of secret codes printed with a specific type of printing equipment If the similarity but and The secret points in two secret sets are judged to be similar;
[0125] The multiple nearest elements are the elements in the traceability classification benchmark data set that are closest to the elements in the test set. Preferably, the five nearest elements are selected.
[0126] The majority vote is specifically to make a judgment based on the multiple nearest elements and select the largest number of the multiple nearest element categories as the source printing device traceability classification result.
[0127] It can be seen that the present invention combines deep learning technology and pattern recognition technology to perform multi-dimensional data processing on the sample image data collected by the printing device, and realizes an accurate and universal classification strategy through detailed characteristic analysis and comparison. The application of the KNN algorithm further improves the efficiency and accuracy of the traceability of printing equipment, and provides an efficient and reliable technical means for the identification of printing equipment traces.
[0128] On the other hand, the present invention also provides a deep learning-based automated printing equipment traceability detection device, characterized in that the device comprises:
[0129] A data acquisition module, used to obtain N types of printing device data samples, classify and trim the N types of printing device data samples, and obtain a benchmark data set;
[0130] A data preprocessing module, configured to perform image preprocessing on the reference data set to obtain a first sample data set, wherein the first sample data set is a black-and-white image data set of the N types of printing device data samples contained in the reference data set;
[0131] A model training module, based on one of the N types of printing device data samples contained in the first sample data set, uses an intelligent learning method to extract a complete secret set, obtain neighborhood point pattern features, and then obtain training set features to obtain a traceability classification benchmark data set;
[0132] The data tracing module uses the data not randomly selected in the data enhancement module as a test set, compares the elements in the test set with the traceability classification benchmark data set, and uses the KNN nearest neighbor classification method to obtain the source printing equipment traceability classification result.
[0133] On the other hand, the present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, wherein the computer program, when executed by a processor, implements an automated printing device tracing method based on deep learning.
[0134] The device embodiments described above are merely illustrative, wherein the units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they may be located in one place, or they may be distributed on multiple network units. Some or all of the modules may be selected according to actual needs to achieve the purpose of the scheme of this embodiment. Ordinary technicians in this field can understand and implement it without paying creative labor.
[0135] Through the description of the above implementation methods, those skilled in the art can clearly understand that each implementation method can be implemented by means of software plus a necessary general hardware platform, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solution is essentially or the part that contributes to the prior art can be embodied in the form of a software product, and the computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, a disk, an optical disk, etc., including a number of instructions for a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.
[0136] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A deep learning-based automated printing equipment traceability method, characterized in that: include: Step 1, obtaining N types of printing device data samples, classifying and cutting the N types of printing device data samples to obtain a benchmark data set; Step 2, performing image preprocessing on the reference data set to obtain a first sample data set, where the first sample data set is a black-and-white image data set of the N types of printing device data samples contained in the reference data set; Step 3, based on one of the N types of printing device data samples contained in the first sample data set, using an intelligent learning method to extract a complete secret set, obtain neighborhood point pattern features, obtain training set features, and obtain a traceability classification benchmark data set; Step 4: compare the elements in the test set with the traceability classification benchmark data set, and use the KNN nearest neighbor classification method to obtain the source printing equipment traceability classification result.
2. The deep learning-based automated printing equipment traceability method according to claim 1, characterized in that: The step 1 also includes: Step 11, obtaining the N types of printing device data samples, automatically scanning the N types of printing device data samples based on the TWAIN protocol, and obtaining N types of scanned sample images corresponding to the N types of printing device data samples; Step 12: randomly crop the N types of scanned sample images into k images to obtain samples, forming the benchmark dataset.
3. The deep learning-based automated printing equipment traceability method according to claim 1, characterized in that: The step 2 also includes: Step 21, obtaining the reference data set, decomposing the reference data set according to color channels, and obtaining a blue color channel image with obvious hidden features; Step 22, performing inverting and binarization processing on the blue color channel image to obtain a black and white image of the imaging secret mark; Step 24, performing morphological operations on the black and white image to obtain a clear dark image; Step 25: Based on the clear dark mark image, the first sample data set is obtained by traversing and denoising each pixel in the image.
4. The deep learning-based automated printing equipment traceability method according to claim 1, characterized in that: The step 3 also includes: Step 31, obtaining one type of data from the N types of printing device data samples contained in the first sample data set, calculating the spatial coordinates of all hidden points in the one type of data, and storing the coordinate data to obtain a hidden data file; Step 32, based on the hidden data file, using the neighborhood point pattern NPP to perform scale normalization and rotation consistency on the hidden local features to obtain the neighborhood point pattern features of the hidden data file; Step 33, based on the hidden data file and the corresponding neighborhood point pattern features, obtaining repeated hidden points by comparing the neighborhood point pattern features of the hidden points in the hidden data file; Step 34, based on the repeated hidden points, using the hidden point association learning method of the greedy strategy, calculate the association relationship of all the hidden points, and determine the maximum repeated hidden point set as the hidden pattern set, wherein the maximum repeated hidden point set is the point set containing all the repeated hidden points; Step 35, traverse the data in the first sample data set that are not detected because the neighborhood points are beyond the image range, calculate the correlation between the undetected data and other dark mark points, and supplement the dark mark pattern set to obtain a complete dark mark pattern set.
5. The deep learning-based automated printing equipment traceability method according to claim 4 is characterized in that: The step 35 further comprises: Step 351, obtaining a complete set of hidden patterns corresponding to the N types of printing device data samples; Step 352, based on the complete dark pattern set corresponding to the N types of printing device data samples, randomly select the dark pattern set of images for the N categories as a training set, and the dark pattern set of the remaining sample images as a test set; Step 353, obtaining the neighborhood point pattern features in the training set, storing the neighborhood point pattern features in a database, and obtaining the traceability classification benchmark data set.
6. The deep learning-based automated printing equipment traceability method according to claim 1, characterized in that: The step 4 also includes: Step 41, obtaining the test set and the traceability classification benchmark data set; Step 42, for each element in the test set, traverse the traceability classification benchmark data set, use the KNN nearest neighbor method to obtain the distance relationship between the test set and the traceability classification benchmark data set, and determine the similarity result between each element in the test set and the traceability classification benchmark data set; Step 43, classifying the reference data set to be traced and classified to obtain the source printing device traceability classification result.
7. The deep learning-based automated printing equipment traceability method according to claim 6, characterized in that: The step 42 further comprises: Step 421, based on the test set, traverse the traceability classification benchmark data set; Step 422, taking the test set as a benchmark, calculating the distance between each element in the test set and each element in the traceability classification benchmark data set to obtain a first distance; Step 423, taking the traceability classification benchmark data set as a benchmark, calculating the distance between each element in the test set and each element in the traceability classification benchmark data set to obtain a second distance; Step 424, traverse the first distance and the second distance corresponding to each element in the test set, calculate the difference between the first distance and the second distance, and obtain the similarity result; Step 425, based on the similarity result, obtain multiple elements in the traceability classification benchmark data set that are closest to the test set, use the majority voting method to perform source printing device traceability classification on the test sample image, and obtain the source printing device traceability classification result.
8. A deep learning-based automated printing equipment traceability detection device, characterized in that: The device comprises: A data acquisition module, used to obtain N types of printing device data samples, classify and trim the N types of printing device data samples, and obtain a benchmark data set; A data preprocessing module, configured to perform image preprocessing on the reference data set to obtain a first sample data set, wherein the first sample data set is a black-and-white image data set of the N types of printing device data samples contained in the reference data set; A model training module, based on one of the N types of printing device data samples contained in the first sample data set, uses an intelligent learning method to extract a complete secret set, obtain neighborhood point pattern features, obtain training set features, and obtain a traceability classification benchmark data set; The data tracing module uses the data not randomly selected in the data enhancement module as a test set, compares the elements in the test set with the traceability classification benchmark data set, and uses the KNN nearest neighbor classification method to obtain the source printing equipment traceability classification result.
9. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the deep learning-based automated printing device tracing method as described in any one of claims 1 to 7 is implemented.
10. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the deep learning-based automated printing device tracing method as described in any one of claims 1 to 7 is implemented.