Snapshot method, system and device for stamping image of physical seal

Through the improved 3D convolutional neural network and long-term memory network combined with the semantic segmentation and feature fusion algorithm of the full convolutional network, the accurate stamped image capture of the seal intelligent workstation is realized, solving the problem of inefficient identification and capture in the existing technology, and improving the accuracy of query and compliance inspection.

CN120282013APending Publication Date: 2025-07-08SHANGHAI XUANWUJI ARTIFICIAL INTELLIGENCE TECHNOLOGY CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510448757.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-10
Publication Date
2025-07-08

AI Technical Summary

Technical Problem

The existing seal intelligent workstation lacks a linkage mechanism with the camera, making it difficult to accurately identify and capture stamped images, resulting in inefficient review and compliance inspections, and the risk of human negligence.

Method used

The improved 3D convolutional neural network and long and short-term memory network detection gestures and stamping actions are adopted, combined with the full convolutional network for semantic segmentation, and combined with the fusion detection algorithm of color, shape and texture features to achieve accurate capture of stamped images.

Benefits of technology

It improves the accuracy of stamped image capture, reduces labor and time costs, enhances the accuracy of compliance inspection, and improves the safety and query efficiency of seal use.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120282013A_ABST
    Figure CN120282013A_ABST
Patent Text Reader

Abstract

The invention discloses a snapshot method for a seal image of a physical seal, and the method comprises the following steps: S1, carrying out the snapshot of a picture during sealing: collecting the image in real time through a sealing camera when a seal is sealed on an intelligent work station of the seal; when the system detects that a hand holds a seal to enter a stamping area and identifies a stamping operation action, a snapshot mechanism is automatically triggered, and a picture is snapshot to serve as a picture in the stamping to be stored; s2, picture snapshot after stamping: after stamping is completed, checking whether a hand leaves a stamped file area or not, and if it is detected that the hand leaves; detecting whether a red seal is left on the sealed file in the image in real time; and if the red stamp is detected, automatically capturing one picture again. Pictures in stamping and pictures after stamping can be automatically and accurately captured, and a stamp manager can directly check the clear pictures; the auditor can check the compliance of the seal more accurately through the captured picture; the captured picture can be used as an important mark-leaving evidence.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of seal control, and particularly relates to a method, a system and a device for capturing images of physical seal stamping. Background Art

[0002] In the existing application scenarios of intelligent seal workstations, due to the lack of a linkage mechanism between ordinary physical seals and intelligent workstations, it is difficult for the stamping cameras on the intelligent workstations to accurately identify and capture specific images of each stamping, and it is even more impossible to accurately capture the images of the documents after stamping. Currently, intelligent workstations generally use the method of recording the entire stamping process video by cameras. However, this traditional method has many drawbacks.

[0003] On the one hand, the recorded video lacks an effective association with the stamped documents. When it is necessary to consult the specific content of the stamping, the staff needs to spend a lot of time and effort in screening and searching in the long video. On the other hand, when the auditors conduct compliance inspections on the stamping, they need to view the entire long stamping video frame by frame, which not only consumes a large amount of human and time costs, but also, due to the large amount of video content, it is very easy to miss the key content of non-compliant stamping, thus triggering potential risks. Summary of the Invention

[0004] Therefore, the present invention provides a method, a system and a device for capturing images of physical seal stamping to solve the problems in the prior art.

[0005] To achieve the above object, the present invention provides the following technical solutions:

[0006] A method for capturing images of physical seal stamping includes the following steps:

[0007] S1. Capturing pictures during stamping: When stamping on an intelligent seal workstation, the stamping camera collects images in real time; Joint detection of gesture and stamping action: A gesture and stamping action detection model is constructed by combining an improved 3D convolutional neural network and a long short-term memory network;

[0008] When the system detects that a person holds the seal and enters the stamping area and recognizes the stamping operation action, the capture mechanism is automatically triggered to capture a picture as the picture during stamping for storage;

[0009] S2. Capturing pictures after stamping: After the stamping is completed, the back of the hand is detected to check whether the hand has left the stamping document area. If it is detected that the hand has left; A detection algorithm that fuses color features, shape features and texture features is used to detect in real time whether there is a red seal left on the stamped document in the image;

[0010] If a red seal is detected and it is further confirmed that the hand and the seal are not covering the stamped document on the document being stamped, then a picture is automatically captured again at this time and saved as the picture after stamping.

[0011] Furthermore: During the process of the stamping camera capturing images in real time, it is necessary to use a semantic segmentation algorithm based on a fully convolutional network to process the captured images. The semantic segmentation algorithm based on the fully convolutional network can accurately identify the area of the stamped document in the image through training on a large number of stamped document images, distinguish the document area from the background, and determine the specific position and scope of the stamped document.

[0012] Furthermore: The specific process of processing the captured images is as follows:

[0013] Let the input image be I(x, y), where (x, y) represents the pixel coordinates in the image; the fully convolutional network learns the mapping relationship from the image to the semantic class label through training on a large number of stamped document image samples and obtains the semantic segmentation model f seg ; where I i represents the i-th input image, and L i represents the semantic label corresponding to the i-th image;

[0014] For the input image I, after being processed by the fully convolutional network, the predicted probability map P(x, y) of each pixel belonging to the stamped document area is obtained, and its calculation formula is: P(x, y) = f seg (I(x, y));

[0015] The cross-entropy loss function is used to optimize the model parameters to make the prediction result close to the true label; the cross-entropy loss function formula is:

[0016]

[0017] where L i (x, y) represents the value of the true label at the pixel (x, y), which is 1 if it belongs to the stamped document area and 0 otherwise.

[0018] Furthermore: The method of constructing a gesture and stamping action detection model using a 3D convolutional neural network and a long short-term memory network is as follows:

[0019] The 3D convolutional neural network extracts spatio-temporal features in the video sequence to capture the dynamic changes of gestures and stamping actions; the long short-term memory network is used to process sequence data and further analyze and learn the features extracted by the 3D convolutional neural network to identify complex gesture and stamping action patterns;

[0020] Let the video sequence input be where v tDenote the t-th frame image, where T is the total number of frames in the video sequence; the output feature tensor of the 3D convolutional neural network is denoted as F 3d (V), representing the feature information extracted from the video;

[0021] The long short-term memory network further analyzes and learns the features extracted by the 3D convolutional neural network,

[0022] Let the hidden state of the long short-term memory network at time t be h t , and the states of the input gate, forget gate, and output gate be i t , f t , and o t , respectively, and the cell state be c t , then there are the following update formulas:

[0023] i t = σ(W i ·[h t-1 , v t + b i )

[0024] f t = σ(W f ·[h t-1 , v t + b f )

[0025] o t = σ(W o ·[h t-1 , v t + b o

[0026] c t = f t ☉c t-1 + i t ☉tanh(W c ·[h t-1 , v t + b c )

[0027] h t = o t ☉tanh(c t );

[0028] Among them, σ represents the activation function (such as the sigmoid function), W and b represent the weight matrix and bias vector respectively, ⊙ represents the element-wise multiplication operation; h t-1 is the hidden state at the previous moment; c t-1 is the cell state at the previous moment; b i , b f , b o , b cBias vectors representing the input gate, forget gate, output gate, and cell state update respectively; W i and W f and W o and W c Weight matrices representing the input gate, forget gate, output gate, and cell state update respectively;

[0029] When a certain eigenvalue output by the long short-term memory network exceeds the set threshold, it can be determined that a specific stamping action has been detected, thus triggering a capture.

[0030] Furthermore: The color feature is analyzed in detail using a color histogram and color moments based on the specific red range of the red seal; the shape feature is detected using the Hough transform to detect common seal shapes; the texture feature is extracted through a gray-level co-occurrence matrix;

[0031] The color feature, shape feature, and texture feature are fused and classified and judged through a support vector machine classifier.

[0032] Furthermore: The calculations for analysis using a color histogram and color moments are as follows: Let the color histogram of the image be H(r, g, b), where r, g, b represent the value ranges of the red, green, and blue channels respectively, and the color moments include the first-order moment, second-order moment, and third-order moment; then the formula for the first-order moment of the red channel is:

[0033]

[0034] where N represents the total number of pixels in the image, and r i represents the red channel value of the i-th pixel.

[0035] Furthermore: The calculation method for classification and judgment through a support vector machine classifier is as follows:

[0036] Let the fused feature vector be F = [f1, f2,..., f m ;

[0037] where m represents the dimension of the feature vector,

[0038] The decision function of the SVM is: f(x) = ω T φ(x) + b;

[0039] where ω is the weight vector, and ω T is the transpose of the weight vector ω; φ(x) is the kernel function, and b is the bias term.

[0040] To achieve the above object, the present invention also provides a capture system for the stamped image of a physical seal, including:

[0041] The image acquisition module consists of a stamping camera on the intelligent seal workstation and is responsible for real-time acquisition of image data during the stamping process. The camera can clearly and accurately capture each stage of stamping. At the same time, the camera supports autofocus and light compensation functions to adapt to different ambient light conditions.

[0042] The detection and analysis module is used to detect and analyze the acquired images. The detection and analysis module includes a stamping document area detection sub-module, a gesture and stamping action detection sub-module, a back of the hand detection sub-module, and a red seal detection sub-module.

[0043] The control and capture module, based on the results of the detection and analysis module, uses a control strategy based on a rule engine to control the stamping camera to perform capture operations. When the capture conditions during or after stamping are met, a capture instruction is sent to the camera to ensure accurate capture of the required pictures.

[0044] The image storage module stores the pictures during stamping and after stamping, and stores the pictures according to documents in combination with the seal management system, providing a basis for subsequent query and management.

[0045] Furthermore: The stamping document area detection sub-module analyzes the acquired images using a semantic segmentation algorithm based on FCN to determine the position and scope of the stamping document.

[0046] The gesture and stamping action detection sub-module uses an improved 3D convolutional neural network and long short-term memory network model to perform real-time recognition and judgment on the gestures and stamping actions of the operator.

[0047] The back of the hand detection sub-module uses a back of the hand detection model based on the YOLOv5 algorithm to accurately detect whether the hand has left the stamping document area.

[0048] The red seal detection sub-module uses a detection algorithm that fuses color features, shape features, and texture features, combined with an SVM classifier, to detect whether there is a red seal on the stamping document and the position and status of the red seal.

[0049] The present invention also provides a device for capturing images of physical seal stamping, including a main body, a stamping table, and an operation screen. An intelligent storage cabinet is provided at the lower part of the main body, the stamping table is above the intelligent storage cabinet, and the operation screen is provided above the stamping table. A camera for taking pictures is provided below the operation screen, a face recognition camera is provided at the upper front side of the operation screen, and a plurality of environmental monitoring cameras are provided at the top of the main body. A processor, a memory, an operating system, a seal management system, and an image processing algorithm are provided inside the main body.

[0050] The present invention has the following advantages:

[0051] 1. It can automatically and accurately capture the pictures during and after stamping. The seal management personnel can directly view these clear pictures without spending a lot of time watching the entire long stamping video. Through precise algorithms and models, the accuracy of capture is greatly improved, thus significantly enhancing the query efficiency and reducing the labor and time costs.

[0052] 2. Enhance the accuracy of compliance inspection: Through the captured pictures, the auditors can more accurately check the compliance of stamping. The clear pictures can provide detailed stamping information, avoiding missing non-compliant stamping content due to human negligence and effectively reducing the risks.

[0053] 3. Improve the security of seal use: The captured pictures can be used as important evidence for record-keeping, helping to trace the usage of the seal. In case of seal abuse or misuse, the accurate liability determination can be made by viewing the captured pictures, preventing the abuse and misuse of the seal and comprehensively enhancing the security of seal use.

[0054] Other features and advantages of the present invention will be described in the subsequent specification, and in part, will be obvious from the specification or understood by implementing the present invention. BRIEF DESCRIPTION OF THE DRAWINGS

[0055] To more intuitively illustrate the prior art and the present application, exemplary drawings are given below. It should be understood that the specific shapes and structures shown in the drawings generally should not be regarded as limiting conditions when implementing the present application. For example, those skilled in the art are capable of making routine adjustments or further optimizations to the addition / deletion / attribution division of certain units (components), specific shapes, positional relationships, connection methods, dimensional ratio relationships, etc. based on the technical concept disclosed in the present application and the exemplary drawings.

[0056] Figure 1 It is a flowchart of a method for capturing the stamping image of a physical seal provided in an embodiment of the present application.

[0057] Figure 2 It is a system block diagram of a system for capturing the stamping image of a physical seal provided in an embodiment of the present invention.

[0058] Figure 3 It is a schematic structural diagram of a device for capturing the stamping image of a physical seal provided in an embodiment of the present invention.

[0059] In the figure: 1. Main body; 2. Intelligent storage cabinet; 3. Stamping table; 4. Photographing camera; 5. Operation screen; 6. Environment monitoring camera; 7. Face recognition camera. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0060] The following specific embodiments illustrate the implementation manners of the present invention. Those skilled in the art can easily understand other advantages and effects of the present invention from the content disclosed in this specification. Obviously, the described embodiments are part of the embodiments of the present invention, rather than all embodiments. It should be understood that these embodiments are only for further explaining the present invention and cannot be construed as limiting the protection scope of the present invention. Technical engineers in this field can make some non-essential improvements and adjustments to the present invention according to the content of the above invention; based on the embodiments in the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the protection scope of the present invention.

[0061] Please refer to Figure 1 , a method for capturing the stamping image of a physical seal, comprising the following steps:

[0062] S1. Capturing pictures during stamping

[0063] (1) Precise detection of the stamping document area: When stamping on the seal intelligent workstation, the stamping camera captures images in real time.

[0064] Use the semantic segmentation algorithm based on the fully convolutional network (FCN) to process the captured images. This algorithm can accurately identify the area of the stamping document in the image through training on a large number of stamping document images, distinguish the document area from the background, and determine the specific position and scope of the stamping document.

[0065] Let the input image be I(x, y), where (x, y) represents the pixel coordinates in the image; FCN learns the mapping relationship from the image to the semantic class label through training on a large number of stamping document image samples to obtain the semantic segmentation model f seg ; where, I i represents the i-th input image, and L i represents the semantic label corresponding to the i-th image.

[0066] For the input image I, after being processed by FCN, the prediction probability map P(x, y) of each pixel belonging to the stamping document area is obtained, and its calculation formula is: P(x, y) = f seg (Ix, y).

[0067] Usually, the cross-entropy loss function is used to optimize the model parameters to make the prediction result as close as possible to the true label. The cross-entropy loss function formula is:

[0068]

[0069] where, L i(x, y) represents the value of the true label at the pixel (x, y) (1 if it belongs to the stamped document area, otherwise 0).

[0070] For example: Suppose there is a stamped document image. After being processed by the FCN model, a prediction probability map with the same size as the original image is obtained. In the probability map, the value at a certain pixel position is 0.9, indicating that the probability of this pixel belonging to the stamped document area is 0.9; if the value is 0.1, it means that the probability of belonging to the background area is relatively large. By setting a threshold (which can be 0.5), the probability map can be converted into a binary classification result, that is, pixels greater than or equal to the threshold are determined to be in the stamped document area, and pixels less than the threshold are determined to be in the background area, thereby achieving accurate detection of the stamped document area.

[0071] (2) Joint detection of gestures and stamping actions: Combine an improved 3D convolutional neural network (3D-CNN) and a long short-term memory network (LSTM) to construct a gesture and stamping action detection model.

[0072] 3D-CNN can extract spatio-temporal features in the video sequence and capture the dynamic changes of gestures and stamping actions; LSTM is used to process sequence data and further analyze and learn the features extracted by 3D-CNN to identify complex gesture and stamping action patterns.

[0073] Let the video sequence input be where v t represents the t-th frame image, and T is the total number of frames in the video sequence.

[0074] 3D-CNN is used to extract spatio-temporal features in the video sequence, and its output feature tensor is denoted as F 3d (V), representing the feature information extracted from the video.

[0075] LSTM further analyzes and learns the features extracted by 3D-CNN. Let the hidden state of LSTM at time t be h t , and the states of the input gate, forget gate, and output gate be i t , f t and o t , and the cell state be c t , then there are the following update formulas:

[0076] i t = σ(W i ·[h t-1 , v t +b i )

[0077] f t = σ(W f ·[h t-1 , v t+b f )

[0078] o t =σ(W o ·[h t-1 ,v t +b o

[0079] c t =f t ☉c t-1 +i t ☉tanh(W c ·[h t-1 ,v t +b c )

[0080] h t =o t ☉tanh(c t );

[0081] where σ represents the activation function (such as the sigmoid function), W and b represent the weight matrix and the bias vector respectively, ⊙ represents the element-wise multiplication operation; h t-1 is the hidden state at the previous moment; c t-1 is the cell state at the previous moment; b i , b f , b o , b c represent the bias vectors for the input gate, forget gate, output gate, and cell state update respectively; W i , W f , W o , W c represent the weight matrices for the input gate, forget gate, output gate, and cell state update respectively.

[0082] For example, for a video sequence containing a stamping action, the video is first input into a 3D-CNN to extract spatio-temporal features. Assume that the dimension of the feature tensor output by the 3D-CNN is C×H×W, where C represents the number of channels, and H and W represent the height and width of the feature map respectively; then these features are input into an LSTM network for processing; in the LSTM network, the hidden state and cell state at each moment are calculated according to the above formula.

[0083] When the system detects that the hand-held seal enters the stamping area and recognizes the stamping operation (such as the trajectory and force change of the seal being pressed down), the capture mechanism is automatically triggered to capture a picture and save it as the picture during stamping. This process ensures the accurate capture of the key moment of the stamping action through real-time monitoring and analysis. For example, when a certain feature value output by the LSTM network exceeds the set threshold, it can be determined that a specific stamping action has been detected, thus triggering the capture.

[0084] S2. Picture capture after stamping

[0085] Back of hand detection: After stamping is completed, a back-of-hand detection model based on the YOLOv5 algorithm is used to detect the back of the hand. YOLOv5 is an efficient object detection algorithm, characterized by fast speed and high accuracy. The model has been trained with a large number of back-of-hand image samples and can accurately identify the position and features of the back of the hand. Check whether the hand has left the stamping document area. If it is detected that the hand has left, proceed to the next detection.

[0086] Red seal detection: Continuously detect whether there is a red seal on the stamped document in the image, using a detection algorithm that fuses color features, shape features, and texture features.

[0087] a. Color features: Based on the specific red range of the red seal, detailed analysis is carried out using color histograms and color moments.

[0088] Let the color histogram of the image be H(r, g, b), where r, g, and b represent the value ranges of the red, green, and blue channels respectively. Color moments include the first moment (mean), second moment (variance), and third moment (skewness), etc. For example, the formula for the first moment of the red channel is:

[0089]

[0090] where N represents the total number of pixels in the image, and r i represents the red channel value of the i-th pixel.

[0091] b. Shape features: The Hough transform is used to detect common seal shapes such as circles and ellipses;

[0092] Let the curve of the point (x, y) in the image in the parameter space be ρ = xcosθ + ysinθ (for the case of a straight line). By finding local extreme points in the parameter space, the straight line or other shapes in the image can be determined. For circle detection, its standard equation is (x - a) 2 +(y - b) 2 = r 2 , and in the Hough transform, the parameters (a, b, r) of the circle can be determined by finding the corresponding extreme points in the parameter space.

[0093] c. The texture features are extracted through the gray-level co-occurrence matrix.

[0094] They are extracted through the gray-level co-occurrence matrix. Let the gray-level co-occurrence matrix be G(d, θ), where d represents the distance, θ represents the direction, and the matrix element G(i, j|d, θ) represents the frequency of occurrence of pixel pairs with gray levels i and j at a given direction and distance; some statistical quantities can be extracted from the gray-level co-occurrence matrix as texture features, such as energy, contrast, correlation, and entropy, etc.

[0095] The color features, shape features, and texture features are fused and classified and judged through classifiers such as the support vector machine (SVM).

[0096] Let the fused feature vector be F = [f1, f2,..., f m ;

[0097] where m represents the dimension of the feature vector,

[0098] The decision function of the SVM is: f(x) = ω T φ(x) + b;

[0099] where ω is the weight vector, ω T is the transpose of the weight vector ω; φ(x) is the kernel function (such as the linear kernel function, Gaussian kernel function, etc.), and b is the bias term.

[0100] If a red seal is detected and it is further confirmed that the hand and the seal do not cover the stamped document on the stamped document, at this time, another picture is automatically captured again and saved as the picture after stamping.

[0101] Refer to Figure 2 , a system for capturing images of physical seal stamping, including:

[0102] (1) An image acquisition module, which consists of a stamping camera on the seal intelligent workstation and is responsible for real-time acquisition of image data during the stamping process; the camera has characteristics such as high resolution and high frame rate to ensure that each stage of the stamping can be clearly and accurately captured. At the same time, the camera supports automatic focusing and light compensation functions to adapt to different ambient light conditions.

[0103] (2) A detection and analysis module, including a stamped document area detection sub-module, a gesture and stamping action detection sub-module, a back of the hand detection sub-module, and a red seal detection sub-module;

[0104] The stamped document area detection sub-module uses a semantic segmentation algorithm based on FCN to analyze the collected images and determine the position and scope of the stamped document; the semantic segmentation algorithm can improve the detection accuracy of different types of stamped documents through continuous learning and optimization.

[0105] The gesture and seal - stamping action detection sub - module, based on an improved 3D - CNN and LSTM model, real - time identifies and judges the gestures and seal - stamping actions of the operator; the model is continuously updated and optimized through an online learning mechanism to adapt to the habits and action characteristics of different operators.

[0106] The back - of - hand detection sub - module uses a back - of - hand detection model based on the YOLOv5 algorithm to accurately detect whether the hand leaves the seal - stamping document area; this model has the advantage of high real - time performance and can quickly give the detection result.

[0107] The red - seal detection sub - module uses a detection algorithm that fuses color features, shape features, and texture features, combined with an SVM classifier, to detect whether there is a red seal on the seal - stamping document and the position and status of the red seal. This algorithm is trained and optimized through a large amount of experimental data, improving the robustness of red - seal detection.

[0108] (3) The control and capture module, according to the results of the detection and analysis module, uses a control strategy based on a rule engine to control the seal - stamping camera for capture operations; when the capture conditions during or after seal - stamping are met, it sends a capture instruction to the camera to ensure accurate capture of the required pictures; the rule engine can be flexibly configured according to different application scenarios and requirements.

[0109] (4) The image storage module stores the pictures captured during and after seal - stamping, and stores the pictures according to documents in combination with the seal management system for convenient subsequent query and management; at the same time, relevant metadata, such as the seal - stamping time, seal - stamping person, seal - stamping position, etc., are added to each picture for better traceability and auditing.

[0110] The embodiment of the present invention also provides a Figure 3 capture device for the image of physical seal stamping as shown, including a main body 1, a seal - stamping table 3, and an operation screen 5. A smart storage cabinet 2 is provided at the lower part of the main body 1, the seal - stamping table 3 is above the smart storage cabinet 2, and the operation screen 5 is arranged above the seal - stamping table 3; and a photo - taking camera 4 is provided below the operation screen 5, a face - recognition camera 7 is provided at the upper front side of the operation screen 5, and a plurality of environmental monitoring cameras 6 are provided at the top of the main body 1.

[0111] A processor, a memory, an operating system, a seal management system, an image - processing algorithm, etc. are also provided inside the main body 1.

[0112] Among them, the processor uses a high - performance computer processor to run the algorithms and models in the detection and analysis module and quickly process a large amount of image data.

[0113] The memory is used to store the captured pictures and related data, and has a large storage capacity and high - speed read - write performance.

[0114] Operating system, select the stable and reliable Android system to provide basic support for the software operation of the device.

[0115] Seal management system, used to manage the stored pictures and related data, and provide efficient query, retrieval and storage functions.

[0116] Image processing algorithms, including various image processing algorithms and models, such as semantic segmentation algorithm based on FCN, improved 3D-CNN and LSTM models, dorsal hand detection model based on YOLOv5, detection algorithm that fuses color features, shape features and texture features, etc., to realize the analysis and processing of the collected images.

[0117] The above are only the preferred embodiments of the present invention, and are not intended to limit the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.

Claims

1. A method for capturing an image of a physical seal being stamped, characterized in that, It includes the following steps: S1. Picture capture during stamping: When stamping on the intelligent stamping workstation, the stamping camera captures images in real time; Joint detection of gesture and stamping action: A gesture and stamping action detection model is constructed by combining an improved 3D convolutional neural network and a long short-term memory network; When the system detects that the seal is in the stamping area and recognizes the stamping action, the capture mechanism is automatically triggered to capture a picture as the picture during stamping for storage; S2. Picture capture after stamping: After stamping is completed, the back of the hand is detected to check whether the hand has left the stamping document area. If it is detected that the hand has left, a detection algorithm that fuses color features, shape features, and texture features is used to detect in real time whether there is a red seal on the stamped document in the image; If a red seal is detected and it is further confirmed that the hand and the seal do not cover the stamped document on the stamped document, a picture is automatically captured again at this time as the picture after stamping for storage.

2. The method for capturing an image of an entity seal stamping according to claim 1, characterized in that, During the process of the stamping camera capturing images in real time, it is necessary to use a semantic segmentation algorithm based on a fully convolutional network to process the captured images. The semantic segmentation algorithm based on a fully convolutional network can accurately identify the area of the stamped document in the image through training on a large number of stamped document images, distinguish the document area from the background, and determine the specific position and scope of the stamped document.

3. A method for capturing an image of an entity seal being stamped according to claim 2, characterized in that, The specific process of processing the captured images is as follows: Let the input image be I(x, y), where (x, y) represents the pixel coordinates in the image; the fully convolutional network is trained on a large number of stamped document image samples to learn the mapping relationship from the image to the semantic class label and obtain the semantic segmentation model f seg ; where, I i represents the i-th input image, and L i represents the semantic label corresponding to the i-th image; For the input image I, after being processed by the fully convolutional network, a predicted probability map P(x, y) of each pixel belonging to the stamped document area is obtained, and its calculation formula is: P(x, y) = f seg (I(x, y)); The cross-entropy loss function is used to optimize the model parameters to make the prediction results close to the true labels; The formula for the cross-entropy loss function is: Among them, L i (x, y) represents the value of the true label at the pixel (x, y), which is 1 if it belongs to the stamped document area and 0 otherwise.

4. A method for capturing an image of a physical seal being stamped according to claim 1, characterized in that, The method of constructing a gesture and stamping action detection model by a 3D convolutional neural network and a long short-term memory network is as follows: The 3D convolutional neural network extracts spatio-temporal features in the video sequence to capture the dynamic changes of gestures and stamping actions; The long short-term memory network is used to process sequence data and further analyze and learn the features extracted by the 3D convolutional neural network to identify complex gesture and stamping action patterns; Let the video sequence input be where v t represents the t-th frame image, and T is the total number of frames in the video sequence; the output feature tensor of the 3D convolutional neural network is denoted as F 3d (V), representing the feature information extracted from the video; The long short-term memory network further analyzes and learns the features extracted by the 3D convolutional neural network, Let the hidden state of the long short-term memory network at time t be h t , and the states of the input gate, forget gate, and output gate be i t , f t , and o t , and the cell state be c t , then there are the following update formulas: i t = σ(W i · [h t-1 , v t + b i ) f t = σ(W f · [h t-1 , v t + b f ) o t = σ(W o · [h t-1 , v t + b o c t = f t ⊙c t-1 + i t ⊙tanh(W c · [h t-1 , v t + b c ) h t = o t ☉tanh(c t ); Among them, σ represents the activation function (such as the sigmoid function), W and b represent the weight matrix and the bias vector respectively, and ⊙ represents the element-wise multiplication operation; h t-1 is the hidden state at the previous moment; c t-1 is the cell state at the previous moment; b i 、b f 、b o 、b c represent the bias vectors of the input gate, forget gate, output gate, and cell state update respectively; W i 、W f 、W o 、W c represent the weight matrices of the input gate, forget gate, output gate, and cell state update respectively; When a certain feature value output by the long short-term memory network exceeds the set threshold, it can be judged that a specific stamping action has been detected, thus triggering the capture.

5. A method for capturing an image of a physical seal being stamped, according to claim 1, characterized in that, The color features are carefully analyzed using the color histogram and color moments according to the specific red range of the red seal; The shape features use the Hough transform to detect common seal shapes; The texture features are extracted through the gray-level co-occurrence matrix; The color features, shape features, and texture features are fused and classified and judged through a support vector machine classifier.

6. The method for capturing an image of an entity seal stamping according to claim 5, characterized in that, The calculations for analysis using the color histogram and color moments are as follows: Let the color histogram of the image be H(r, g, b), where r, g, b represent the value ranges of the red, green, and blue channels respectively, and the color moments include the first moment, the second moment, and the third moment; Then the formula for the first moment of the red channel is: where N represents the total number of pixels in the image, and r i represents the red channel value of the i-th pixel.

7. A method for capturing an image of a physical seal being stamped according to claim 5, characterized in that, The calculation method for classification and judgment through a support vector machine classifier is as follows: Let the fused feature vector be F = [f1, f2,..., f m ; where m represents the dimension of the feature vector, The decision function of SVM is: f(x) = ω T φ(x) + b; where ω is the weight vector, ω T is the transpose of the weight vector ω; φ(x) is the kernel function, and b is the bias term.

8. A capturing system for the stamped image of a physical seal, characterized in that, It includes: An image acquisition module, consisting of a stamping camera on the intelligent stamping workstation, responsible for capturing image data during stamping in real time; The camera can clearly and accurately capture each stage of stamping; Meanwhile, the camera supports autofocus and light compensation functions to adapt to different ambient light conditions; The detection and analysis module is used to detect and analyze the collected images; the detection and analysis module includes a sealed document area detection sub-module, a gesture and seal-stamping action detection sub-module, a back-of-hand detection sub-module, and a red seal detection sub-module; The control and capture module, according to the results of the detection and analysis module, adopts a control strategy based on a rule engine to control the seal-stamping camera to perform capture operations; when the capture conditions during or after seal-stamping are met, it sends a capture instruction to the camera to ensure accurate capture of the required pictures; The image storage module stores the pictures during and after seal-stamping, and stores the pictures according to documents in combination with the seal management system, providing a basis for subsequent query and management.

9. The capture system for the stamped image of a physical seal according to claim 8, characterized in that, The sealed document area detection sub-module analyzes the collected images using a semantic segmentation algorithm based on FCN to determine the position and scope of the sealed document; The gesture and seal-stamping action detection sub-module, based on an improved 3D convolutional neural network and long short-term memory network model, performs real-time recognition and judgment on the gestures and seal-stamping actions of the operator; The back-of-hand detection sub-module uses a back-of-hand detection model based on the YOLOv5 algorithm to accurately detect whether the hand has left the sealed document area; The red seal detection sub-module uses a detection algorithm that fuses color features, shape features, and texture features, combined with an SVM classifier, to detect whether there is a red seal on the sealed document and the position and status of the red seal.

10. A device for capturing the stamping image of a physical seal, characterized in that, It includes a main body, a seal-stamping table, and an operation screen. An intelligent storage cabinet is provided at the lower part of the main body. The seal-stamping table is above the intelligent storage cabinet. The operation screen is set above the seal-stamping table; and a photographing camera is provided below the operation screen, a face recognition camera is provided at the upper part of the front side of the operation screen, and multiple environmental monitoring cameras are provided at the top of the main body; a processor, a memory, an operating system, a seal management system, and an image processing algorithm are provided inside the main body.

Citation Information

Cited By

  • Video snapshot method, device, equipment and medium

    CN121865088A