Image information determination method and device, equipment, medium and product

By combining the shared encoder and the feature interaction module, the inaccuracy caused by environmental factors in seal image extraction is solved, enabling accurate recognition and quality evaluation of seal images, and improving the accuracy and efficiency of the task.

CN120976956APending Publication Date: 2025-11-18AGRICULTURAL BANK OF CHINA
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511096928.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-06
Publication Date
2025-11-18

AI Technical Summary

Technical Problem

In the process of handling corporate customer business, the extraction of seal images is affected by factors such as changes in shooting angle, poor lighting, and sensor noise, resulting in uneven color, blurred details, local distortion or deformation, which affects the verification and comparison in subsequent business processes.

Method used

By adopting a shared encoder structure, the two tasks of seal extraction and seal quality evaluation share the same encoder to extract features. Through the feature interaction module, information can be dynamically exchanged between the two task branches. The key areas of seal extraction are used to guide the quality evaluation and improve the task results.

Benefits of technology

This improved the accuracy of seal image extraction and quality assessment, reduced redundant calculations, enhanced seal extraction capabilities, and improved the overall performance of both tasks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120976956A_ABST
    Figure CN120976956A_ABST
Patent Text Reader

Abstract

The invention discloses an image information determination method and device, equipment, a medium and a product. The method comprises the following steps: acquiring a to-be-recognized image; performing feature extraction on the to-be-recognized image according to the shared encoder to obtain a general feature; according to the seal extraction decoder, the seal quality evaluation decoder and the feature interaction module, the general features are subjected to combined processing, image information is determined, and the image information comprises a seal image extraction result and a seal quality evaluation result. By sharing the encoder structure, the seal extraction task and the seal quality evaluation task share the same encoder to extract features, the correlation between the tasks is utilized, the general features are shared, and redundant calculation is reduced. Through the feature interaction module, information is dynamically exchanged between two task branches of seal extraction and seal quality evaluation, quality evaluation is guided by using a key area of seal extraction, the seal extraction capability is enhanced by using a quality score, the effect of the two tasks is improved, and thus the accuracy of the two tasks is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of image processing technology, and in particular to a method, apparatus, device, medium, and product for determining image information. Background Technology

[0002] In corporate banking transactions, customers often need to affix their corporate seal to a pre-registered seal card. The seal is then captured and scanned by a camera or scanner and stored in a database. However, the image acquisition process is subject to various factors, such as changing camera angles, poor lighting, and sensor noise. These factors inevitably affect seal extraction, resulting in uneven color, blurred details, localized distortion, or deformation, thus impacting the verification and comparison of seals in subsequent business processes.

[0003] With the rapid development of computer technology, image processing has become a hot topic. Deep learning methods, through models such as deep neural networks, automatically learn multi-level abstract features from data without human intervention. They are highly adaptable and more suitable for complex scenarios such as seal image extraction.

[0004] However, this method is difficult to consistently achieve accurate results when dealing with complex backgrounds for the seal image. Summary of the Invention

[0005] This invention provides a method, apparatus, device, medium, and product for determining image information, so as to achieve accurate identification and quality evaluation of seal images.

[0006] According to a first aspect of the present invention, an image information determination method is provided, comprising:

[0007] Acquire the image to be recognized;

[0008] Based on the shared encoder, feature extraction is performed on the image to be identified to obtain general features;

[0009] The general features are jointly processed by the seal extraction decoder, the seal quality evaluation decoder, and the feature interaction module to determine the image information, which includes the seal image extraction result and the seal quality evaluation result.

[0010] According to a second aspect of the present invention, an image information determining apparatus is provided, comprising:

[0011] The image acquisition module is used to acquire the image to be recognized;

[0012] The feature extraction module is used to extract features from the image to be identified based on the shared encoder to obtain general features;

[0013] The information determination module is used to jointly process the general features based on the seal extraction decoder, the seal quality evaluation decoder, and the feature interaction module to determine image information, which includes the seal image extraction result and the seal quality evaluation result.

[0014] According to a third aspect of the present invention, an electronic device is provided, the electronic device comprising:

[0015] At least one processor; and

[0016] A memory communicatively connected to the at least one processor; wherein,

[0017] The memory stores a computer program that can be executed by the at least one processor, the computer program being executed by the at least one processor to enable the at least one processor to perform the image information determination method according to any embodiment of the present invention.

[0018] According to a fourth aspect of the present invention, a computer-readable storage medium is provided, the computer-readable storage medium storing computer instructions for causing a processor to execute and implement the image information determination method according to any embodiment of the present invention.

[0019] According to a fifth aspect of the present invention, embodiments of the present invention also provide a computer program product, the computer program product including a computer program, which, when executed by a processor, implements the image information determination method of any embodiment of the present invention.

[0020] The technical solution of this invention involves acquiring an image to be recognized; extracting features from the image using a shared encoder to obtain general features; and jointly processing these general features using a seal extraction decoder, a seal quality evaluation decoder, and a feature interaction module to determine image information, which includes the seal image extraction result and the seal quality evaluation result. The shared encoder structure allows both seal extraction and seal quality evaluation tasks to share the same encoder for feature extraction, leveraging the correlation between tasks and sharing common features to reduce redundant computation. The feature interaction module enables dynamic information exchange between the seal extraction and quality evaluation task branches, using key areas of the extracted seal to guide quality evaluation and using quality scores to enhance seal extraction capabilities, thereby improving the effectiveness and accuracy of both tasks.

[0021] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of the present invention, nor is it intended to limit the scope of the invention. Other features of the invention will become readily apparent from the following description. Attached Figure Description

[0022] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0023] Figure 1 This is a flowchart of an image information determination method provided according to Embodiment 1 of the present invention;

[0024] Figure 2 This is a flowchart of an image information determination method provided according to Embodiment 2 of the present invention;

[0025] Figure 3 This is an example diagram of the first residual block of an image information determination method provided in Embodiment 2 of the present invention;

[0026] Figure 4 This is an example diagram of the second residual block of an image information determination method provided in Embodiment 2 of the present invention;

[0027] Figure 5 This is an example diagram of a shared encoder for an image information determination method according to Embodiment 2 of the present invention;

[0028] Figure 6 This is an example diagram of a seal extraction decoder for an image information determination method according to Embodiment 2 of the present invention;

[0029] Figure 7 This is an example diagram of a seal quality evaluation decoder for an image information determination method provided in Embodiment 2 of the present invention;

[0030] Figure 8 This is a schematic diagram of the structure of an image information determining device according to Embodiment 3 of the present invention;

[0031] Figure 9 This is a schematic diagram of the structure of an electronic device that implements an embodiment of the present invention. Detailed Implementation

[0032] To enable those skilled in the art to better understand the present invention, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of the present invention.

[0033] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this invention are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of the invention described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.

[0034] Example 1

[0035] Figure 1 This is a flowchart of an image information determination method provided in Embodiment 1 of the present invention. This embodiment is applicable to the extraction and quality evaluation of a seal portion in an image. The method can be executed by an image information determination device, which can be implemented in hardware and / or software and can be configured in an electronic device. Figure 1 As shown, the method includes:

[0036] S110. Obtain the image to be recognized.

[0037] In this embodiment, the image to be identified can be understood as an image stamped with a seal.

[0038] Specifically, the image to be recognized can be acquired through peripheral devices, such as by taking a picture or scanning to extract the seal. The processor can then acquire the image to be recognized acquired by the peripheral device.

[0039] S120. Based on the shared encoder, feature extraction is performed on the image to be recognized to obtain general features.

[0040] In this embodiment, the shared encoder can be understood as a feature encoder shared by the seal extraction task and the seal quality evaluation task. The general feature can be understood as the feature extraction result applicable to both tasks.

[0041] Specifically, the processor can input the image to be recognized into the shared encoder, extract features through the shared encoder, and the output result is the general feature.

[0042] S130. Based on the seal extraction decoder, seal quality evaluation decoder and feature interaction module, the general features are jointly processed to determine the image information, which includes the seal image extraction result and the seal quality evaluation result.

[0043] In this embodiment, the seal extraction decoder can be understood as a decoder used to segment the image to extract the seal outline. The image information can be understood as information used to characterize the seal in the image to be identified and the seal quality evaluation result. The seal image extraction result only contains the seal image and does not contain irrelevant areas.

[0044] Specifically, the processor can obtain the results of information exchange between the seal extraction task and the seal quality assessment task through the feature interaction module. Based on the features calculated by their respective decoders for the common features, the processor can then calculate the enhanced features for its own task. The processor can input the enhanced features of each task into the corresponding seal extraction decoder and seal quality assessment decoder for processing, obtaining the seal image extraction result output by the seal extraction decoder and the seal quality assessment result output by the seal quality assessment decoder.

[0045] The technical solution of this invention involves acquiring an image to be recognized; extracting features from the image using a shared encoder to obtain general features; and jointly processing these general features using a seal extraction decoder, a seal quality evaluation decoder, and a feature interaction module to determine image information, which includes the seal image extraction result and the seal quality evaluation result. The shared encoder structure allows both seal extraction and seal quality evaluation tasks to share the same encoder for feature extraction, leveraging the correlation between tasks and sharing common features to reduce redundant computation. The feature interaction module enables dynamic information exchange between the seal extraction and quality evaluation task branches, using key areas of the extracted seal to guide quality evaluation and using quality scores to enhance seal extraction capabilities, thereby improving the effectiveness and accuracy of both tasks.

[0046] Example 2

[0047] Figure 2 This is a flowchart of an image information determination method provided in Embodiment 2 of the present invention. This embodiment is a further refinement of the above embodiment. Figure 2 As shown, the method includes:

[0048] S201. Obtain the image to be recognized.

[0049] Furthermore, the shared encoder includes convolutional blocks, max pooling layers, and multiple cascaded residual block groups.

[0050] S202. The first feature is obtained by downsampling the image to be recognized through convolutional blocks.

[0051] In this embodiment, a convolutional block can be understood as a unit used to extract local features of the input image and perform graph feature transformation. The first feature can be understood as the feature processed by the convolutional block.

[0052] Specifically, the processor can input the image to be recognized into the shared encoder, downsample the image to be recognized through the convolutional blocks of the shared encoder, extract the local features of the image to be recognized and perform feature transformation, capture the spatial correlation of the image to be recognized through convolution operations, and enhance the expressive power of the decoder by combining nonlinear activation, and finally realize the mapping from the original data to high-level features to obtain the first feature.

[0053] S203. The first feature is downsampled through a max pooling layer to obtain the second feature.

[0054] In this embodiment, the max pooling layer can be understood as a layer used for feature dimensionality reduction and key information extraction. The second feature can be understood as the result after processing by the max pooling layer.

[0055] Specifically, the first feature is downsampled through the max pooling layer of the shared encoder to reduce its size while preserving important features of the image, thus obtaining the second feature.

[0056] S204. The second feature is extracted through each residual block group to obtain the general feature; the residual block group includes the first residual block and the second residual block.

[0057] In this embodiment, a residual block group can be understood as a combination of multiple residual blocks, and a residual block can be understood as a block group used to alleviate the gradient vanishing problem. The first residual block and the second residual block have different structures.

[0058] Specifically, the second feature can be extracted from each residual block group to obtain the general feature.

[0059] For example, a specific example is used to illustrate the first residual block and the second residual block. Figure 3 This is an example diagram of the first residual block of an image information determination method provided in Embodiment 2 of the present invention. Figure 4 This is an example diagram of the second residual block of an image information determination method provided in Embodiment 2 of the present invention, as shown below. Figure 3 and Figure 4 As shown, the second residual block consists of a series of 1×1 convolutional blocks and 3×3 convolutional blocks. Its output and input are fused using residual connections, and finally, a ReLU activation function layer is applied. Unlike the second residual block, the residual branch of the first residual block passes through a 1×1 convolutional block.

[0060] For example, to facilitate understanding of the structure of a shared encoder, a specific example can be shown. Figure 5 This is an example diagram of a shared encoder for an image information determination method provided in Embodiment 2 of the present invention, as shown below. Figure 5As shown, the shared encoder includes convolutional blocks, max-pooling layers, and four cascaded residual block groups. Each convolutional block consists of a convolutional layer, an instance normalization layer, and a ReLU layer. The first residual block group consists of a first residual block and two second residual blocks; the second residual block group consists of a first residual block and three second residual blocks; the third residual block group consists of a first residual block and five second residual blocks; and the fourth residual block group consists of a first residual block and two second residual blocks. The shared encoder's processing procedure is as follows: First, a convolutional block is used to downsample the image to be recognized. This convolutional block consists of a 3×3 kernel layer with a stride of 2, an instance normalization layer, and a ReLU activation function layer. Next, a max-pooling layer is used for further downsampling. The pooling window size is 3×3 with a stride of 2, used to reduce the size of the feature map while preserving important image features. Then, multiple residual blocks are cascaded, which helps extract features at different levels, enabling the encoder to simultaneously capture shallow local features such as edges and textures, as well as deep global semantics.

[0061] S205. The specified features of the seal extraction decoder and the seal quality evaluation decoder are exchanged through the feature interaction module to determine the first enhancement feature and the second enhancement feature.

[0062] In this embodiment, the specified feature can be understood as the feature used as input to the feature interaction module in different decoders. The first enhanced feature can be understood as the result of adding the features of the seal quality evaluation decoder to the features of the seal extraction decoder. The second enhanced feature can be understood as the result of adding the features of the seal extraction decoder to the features of the seal quality evaluation decoder.

[0063] Specifically, the processor can input general features into the stamp extraction decoder and extract specified features from the output of a specified layer as features to be enhanced for the stamp extraction task. The processor can also input general features into the stamp quality evaluation decoder and extract features from the output of a specified layer as features to be enhanced for the stamp quality evaluation task. The feature interaction module then enhances both the features to be enhanced for the stamp extraction task and the features to be enhanced for the stamp quality evaluation task, respectively, to obtain the first enhanced feature and the second enhanced feature.

[0064] Furthermore, based on the above embodiments, the step of performing feature exchange on specified features of the seal extraction decoder and the seal quality evaluation decoder through the feature interaction module to determine the first enhanced feature and the second enhanced feature can be refined as follows:

[0065] The general features are input into the seal extraction decoder to obtain the first intermediate features output by the specified layer of the seal extraction decoder; the general features are input into the seal quality evaluation decoder to obtain the second intermediate features output by the specified layer of the seal quality evaluation decoder; the first intermediate features are used as the query sequence of the feature interaction module, and the second intermediate features are used as the key-value sequence of the feature interaction module to obtain the first enhanced feature after enhancing the first intermediate features; the second intermediate features are used as the query sequence of the feature interaction module, and the first intermediate features are used as the key-value sequence of the feature interaction module to obtain the second enhanced feature after enhancing the second intermediate features.

[0066] In this embodiment, the designated layer of the stamp extraction decoder can be understood as the layer used for feature enhancement. It is typically selected from the layers corresponding to shallow features in the decoder, such as the layer corresponding to the first convolutional block in the stamp extraction module. The first intermediate feature can be understood as the feature to be enhanced output by the stamp extraction decoder. The designated layer of the stamp quality evaluation decoder can also be understood as the layer used for feature enhancement. It is typically selected from the layers corresponding to shallow features in the decoder, such as the layer corresponding to the first convolutional block in the stamp quality evaluation module. The second intermediate feature can be understood as the feature to be enhanced output by the stamp quality evaluation decoder. The Query sequence and Key-Value sequence are the three core elements in the cross-attention mechanism for realizing feature association and information extraction. The Query sequence consists of the Query vector of each element in the input sequence. Each Query vector represents the "information or feature to be found" of the current element, used to measure its relevance to other elements. The Key sequence consists of the Key vector of each element in the input sequence. Each Key vector represents the "feature identifier" or "index information" of the element, used to construct the basis for calculating the attention score. The Value sequence consists of the Value vector for each element in the input sequence. Each Value vector contains the "actual content information" of the element and is the "data source" for the final attention output.

[0067] Specifically, the processor can input general features into the stamp extraction decoder to obtain the first intermediate feature output by a specified layer of the stamp extraction decoder. For example, the feature output by the first convolutional block of the stamp extraction decoder can be used as the first intermediate feature. The processor can also input general features into the stamp quality assessment decoder to obtain the second intermediate feature output by a specified layer of the stamp quality assessment decoder. For example, the feature output by the first convolutional layer of the stamp quality assessment decoder can be used as the second intermediate feature. When enhancing the features of the stamp extraction decoder, the processor can use the first intermediate feature as the query sequence of the feature interaction module and the second intermediate feature as the key-value sequence of the feature interaction module to obtain the first enhanced feature. When enhancing the features of the stamp quality assessment decoder, the processor can use the second intermediate feature as the query sequence of the feature interaction module and the first intermediate feature as the key-value sequence of the feature interaction module to obtain the second enhanced feature.

[0068] For example, the feature interaction module can use a cross-attention mechanism to enable the decoder to dynamically exchange information between the seal recognition and quality assessment tasks, leveraging features from the other task to enhance the features of the current task. The query segmentation and key-value evaluation in the cross-attention layer come from different sequences. In multi-task learning, inputting the features of the two tasks as queries and key-value pairs respectively into the cross-attention layer allows for bidirectional information exchange. Given a query source sequence A∈R... N×dq Where N is the sequence length, d q It is a sequence dimension; the key-value source sequence B∈R M×dkv Where M is the sequence length, d kv It's a sequence dimension. First, perform a linear projection on A and B, projecting A as a query and B as a key-value pair. This process can be represented as:

[0069]

[0070] Among them, W Q W K W V This is the projection matrix. Weights are calculated using the scaled dot product attention.

[0071]

[0072] Among them, QK T Let be the dot product similarity, representing the similarity between each query and the key; d kThis is a scaling factor to prevent gradient vanishing due to excessively large dot product values. Combining the above steps, the calculation process for cross-attention for the Query source sequence A and Key-Value source sequence B can be summarized as follows:

[0073]

[0074] In this invention, a feature interaction module facilitates information exchange between the seal extraction decoder and the seal quality evaluation decoder. Each decoder uses its own task-specific features as a query sequence and features from related tasks as a key-value sequence, exchanging information through the feature interaction module to calculate enhanced features for its own task. The first enhanced feature for the seal extraction task and the second enhanced feature F′ for the seal quality evaluation task are described. task1 and F′ task2 They can be represented as follows:

[0075] F′ task1 =CrossAttention(F task1 ,F task2 )

[0076] F′ task2 =CrossAttention(F task2 ,F task1 )

[0077] Among them, F task1 This represents the first intermediate feature, i.e., the task-specific shallow feature of the stamp extraction decoder, which is the output feature of the first convolutional block in the stamp extraction decoder; F task2 This represents the second intermediate feature, which is the task-specific shallow feature of the stamp quality assessment decoder, and is the output feature of the first convolutional layer in the stamp quality assessment decoder.

[0078] Furthermore, the stamp extraction decoder includes a transposed convolutional layer and multiple concatenated convolutional layers.

[0079] S206. Determine the seal image extraction result based on the first enhanced feature, the general feature, and the seal extraction decoder.

[0080] Specifically, the seal extraction decoder can output the first intermediate feature after processing the general features through a specified layer, and input it into the feature interaction module for feature enhancement to obtain the first enhanced feature. The first enhanced feature is then processed to obtain the seal image extraction result output by the decoder.

[0081] Furthermore, based on the above embodiments, the steps for determining the seal image extraction result based on the first enhanced feature, the general feature, and the seal extraction decoder can be refined as follows:

[0082] The first decoded feature is obtained by upsampling the general features through a transposed convolutional layer; the second decoded feature is obtained by processing the first output feature, the first decoded feature, and the first enhanced feature of the first residual block group through a concatenated convolutional layer; the third decoded feature is obtained by processing the second output feature and the second decoded feature of the second residual block group through a concatenated convolutional layer; and the third output feature and the third decoded feature of the third residual block group through a concatenated convolutional layer are used to obtain the seal image extraction result.

[0083] In this embodiment, the transposed convolutional layer can be understood as a layer used for upsampling. The concatenated convolutional layer includes concat, convolutional blocks, and transposed convolutional layers, used for feature concatenation, convolution, and upsampling.

[0084] Specifically, the general features are upsampled by a transposed convolutional layer to obtain the first decoded features; the first output features, first decoded features, and first enhanced features of the first residual block group are processed by a concatenated convolutional layer to obtain the second decoded features; the second output features and second decoded features of the second residual block group are processed by a concatenated convolutional layer to obtain the third decoded features; and the third output features and third decoded features of the third residual block group are processed by a concatenated convolutional layer to obtain the seal image extraction result.

[0085] For example, the structure of the seal extraction decoder can be illustrated with a specific example. Figure 6 This is an example diagram of a seal extraction decoder for an image information determination method provided in Embodiment 2 of the present invention, as shown below. Figure 6 As shown, the concatenated convolutional layer includes concat, convolutional blocks, and a transposed convolutional layer (i.e., upsampling). Stamp extraction is essentially an image segmentation task, identifying key regions from image features and segmenting the stamp outline. First, the stamp extraction decoder processes the first output features (i.e.,...) through the transposed convolutional layer. Figure 5 The output of the rightmost residual block group is upsampled and then concatenated with the second output feature in the shared encoder (i.e., the output of the rightmost residual block group) through concatenation in the convolutional layer. Figure 5 The output of the residual block group next to the rightmost residual block group is concatenated, and this skip connection achieves the fusion of shallow details and deep semantics. The concatenated features are then fed into the convolutional blocks of the concatenated convolutional layer for refinement, and then upsampled through the transposed convolutional layer of the concatenated convolutional layer to increase the resolution of the output features. The upsampled features are then further combined with the third output feature in the shared encoder (i.e., Figure 5 The outputs of the third residual block from the right are concatenated, and the concatenated features are then convolved and upsampled before being combined with... Figure 5After the third output features of the residual block group immediately adjacent to the convolutional block are concatenated, they are then passed through the convolutional block and upsampled to finally output a feature segmentation map with the same size as the input features of the shared encoder, which is used as the stamp image extraction result. The convolutional blocks in the stamp extraction decoder have the same structure, consisting of two 3×3 convolutions, an instance normalization layer, and a ReLU layer.

[0086] S207. Determine the seal quality evaluation result based on the second enhancement feature and the seal quality evaluation decoder.

[0087] Specifically, the processor can input general features into the seal quality evaluation decoder, output second intermediate features through the convolutional layer of the seal quality evaluation decoder, input the second intermediate features into the feature interaction module to obtain second enhanced features, and process the second enhanced features through the seal quality evaluation decoder to determine the seal quality evaluation result.

[0088] Furthermore, based on the above embodiments, the seal quality evaluation decoder includes an attention pooling layer, an intermediate layer, and an activation function layer. Correspondingly, the step of determining the seal quality evaluation result based on the second enhanced feature and the seal quality evaluation decoder can be refined as follows:

[0089] The second enhanced feature is subjected to attention pooling through an attention pooling layer to obtain attention pooled features; the attention pooled features are processed through an intermediate layer to obtain the final features; the final features are processed through an activation function layer to obtain the seal quality evaluation result.

[0090] In this embodiment, the intermediate layer may include a fully connected layer, a ReLU layer, and a Dropout layer. The activation function layer may include a fully connected layer and a Sigmoid layer.

[0091] Specifically, the second enhanced feature is subjected to attention pooling through an attention pooling layer to obtain attention pooled features; the attention pooled features are processed through an intermediate layer to obtain the final features; and the final features are processed through an activation function layer to obtain the seal quality evaluation result.

[0092] For example, the structure of a seal quality assessment decoder can be illustrated with a specific example. Figure 7 This is an example diagram of a seal quality evaluation decoder for an image information determination method provided in Embodiment 2 of the present invention, as shown in the figure. Figure 7 As shown, the stamp quality assessment decoder includes convolutional layers, intermediate layers, and activation function layers. The intermediate layers include fully connected layers, ReLU layers, and Dropout layers, while the activation function layers include fully connected layers and Sigmoid layers. First, a 1×1 convolution operation is performed on the common features output by the shared encoder through the convolutional layers of the stamp quality assessment decoder, which serves as the second intermediate feature F. task2The second intermediate feature is processed by the feature interaction module to obtain the second enhanced feature F′. task2 Attention pooling is used on this enhanced feature to dynamically weight important regions within the feature. First, for the enhanced feature F′... task2 Generate attention weights:

[0093] A weighted =σ(f attn (F′ task2 ))

[0094] Among them, f attn σ is the weight generation function, and a 1×1 convolutional layer is used here; σ is the activation function, and the Softmax activation function is used here to normalize the weights to the range [0,1]. The attention weights are multiplied element-wise with the enhanced features to obtain the weighted feature F. weighted :

[0095] F weighted =F′ task2 ⊙A weighted

[0096] Here, ⊙ represents element-wise multiplication. Finally, the weighted features are summed along the spatial dimension to obtain the final attention-pooled feature F. AP :

[0097]

[0098] In the above formula, H and W are weighted features F weighted ∈R C×H×W The length and width dimensions.

[0099] The attention-pooled features are passed through a fully connected layer, a ReLU layer, and a Dropout layer to enhance feature representation and effectively prevent overfitting. Finally, the results are passed through a fully connected layer with one output channel and normalized to the range [0,1] using a sigmoid activation function to obtain the final stamp quality score.

[0100] S208. Use the seal image extraction results and seal quality evaluation results as image information.

[0101] Furthermore, in this invention, a multi-task joint learning method can be used to simultaneously learn two tasks: seal extraction and seal quality evaluation. The seal extraction decoder and the seal quality evaluation decoder can each use independent objective functions, namely segmentation loss and quality evaluation loss. By weighted summing of these two objective functions, the overall objective function for multi-task joint processing is obtained.

[0102] The stamp extraction decoder can employ BCE-Dice Loss, a loss function that combines the standard Binary Cross-Entropy Loss and Dice Loss. The Binary Cross-Entropy Loss is used to optimize the segmented region pixel-by-pixel, improving segmentation accuracy. This loss function can be expressed as:

[0103]

[0104] In the above formula, N is the total number of pixels in the image; p i ∈[0,1] represents the probability that the i-th pixel predicted by the segmentation decoder is a stamp region; g i ∈{0,1} represents the true label of the i-th pixel, where 0 represents the background and 1 represents the stamp region. The Dice loss is used to optimize the overlap between the segmentation result and the real region, improving the quality of the segmentation boundary. This loss function can be expressed as:

[0105]

[0106] In the above formula, ε is the smoothing term, which is usually taken as 1×10. -5 This is to prevent the denominator from being zero. By combining these two losses, we can obtain the representation of the BCE-Dice Loss:

[0107] L BCE-Dice =L BCE +λ1·L Dice

[0108] Where λ1 is the weight coefficient, used to balance the influence of the two losses on the training of the decoder, and the default value is λ1 = 1.

[0109] The seal quality evaluation decoder can use the Smooth L1 loss, which can be expressed as:

[0110]

[0111] In the above formula, |x| is the absolute value of the difference between the predicted quality score and the true score.

[0112] The proposed multi-task joint processing-based decoder for seal extraction and quality assessment uses two loss functions to train the network. The final optimization objective of the decoder is:

[0113] L total =L BCE-Dice +λSmoothL1

[0114] Where λ is the weight coefficient, used to balance the loss between the two decoders during training, so as to avoid the training process being dominated by the loss of one decoder, which would lead to poor training performance of the other decoder. Here, λ = 1 is assumed.

[0115] The technical solution of this invention extracts features from the image to be recognized using a shared encoder structure. This allows both seal extraction and seal quality assessment tasks to share common features extracted by the same encoder, leveraging the correlation between tasks and reducing redundant computation. A multi-task joint learning method is used to simultaneously learn both seal extraction and seal quality assessment tasks. The results of seal segmentation and extraction directly affect the quality score, which in turn indirectly assists seal segmentation. By jointly learning the shared information or features between these two tasks, the model can better capture this shared information, thereby improving the performance of each task. A feature interaction module performs feature interaction on the intermediate features output from specified layers of different decoders, enabling dynamic information exchange between the seal extraction and seal quality assessment task branches. Key regions of seal segmentation guide quality assessment, and quality scores enhance seal segmentation capabilities, thus improving the performance of both tasks. Finally, the enhanced features are processed by the two decoders to obtain the seal image extraction result set and the seal quality assessment result.

[0116] Example 3

[0117] Figure 8 This is a schematic diagram of an image information determining device provided in Embodiment 3 of the present invention. Figure 8 As shown, the device includes:

[0118] Image acquisition module 81 is used to acquire the image to be recognized;

[0119] Feature extraction module 82 is used to extract features from the image to be identified based on the shared encoder to obtain general features;

[0120] The information determination module 83 is used to jointly process the general features based on the seal extraction decoder, the seal quality evaluation decoder and the feature interaction module to determine the image information, which includes the seal image extraction result and the seal quality evaluation result.

[0121] The technical solution of this invention involves acquiring an image to be recognized; extracting features from the image using a shared encoder to obtain general features; and jointly processing these general features using a seal extraction decoder, a seal quality evaluation decoder, and a feature interaction module to determine image information, which includes the seal image extraction result and the seal quality evaluation result. The shared encoder structure allows both seal extraction and seal quality evaluation tasks to share the same encoder for feature extraction, leveraging the correlation between tasks and sharing common features to reduce redundant computation. The feature interaction module enables dynamic information exchange between the seal extraction and quality evaluation task branches, using key areas of the extracted seal to guide quality evaluation and using quality scores to enhance seal extraction capabilities, thereby improving the effectiveness and accuracy of both tasks.

[0122] Furthermore, the shared encoder includes convolutional blocks, max-pooling layers, and multiple cascaded residual block groups. Correspondingly, the feature extraction module 82 is specifically used for:

[0123] The first feature is obtained by downsampling the image to be identified using the convolutional block;

[0124] The first feature is downsampled using the max pooling layer to obtain the second feature;

[0125] The second feature is extracted by each of the residual block groups to obtain a general feature; the residual block group includes a first residual block and a second residual block.

[0126] Furthermore, the information determination module 83 includes:

[0127] The first determining unit is used to perform feature exchange on the specified features of the seal extraction decoder and the seal quality evaluation decoder through the feature interaction module to determine the first enhancement feature and the second enhancement feature.

[0128] The second determining unit is used to determine the seal image extraction result based on the first enhanced feature, the general feature and the seal extraction decoder;

[0129] The third determining unit is used to determine the seal quality evaluation result based on the second enhanced feature and the seal quality evaluation decoder.

[0130] The fourth determining unit is used to use the seal image extraction result and the seal quality evaluation result as image information.

[0131] Specifically, the first determining unit is used for:

[0132] The general features are input into the seal extraction decoder to obtain the first intermediate features output by the specified layer of the seal extraction decoder.

[0133] The general features are input into the seal quality evaluation decoder to obtain the second intermediate features output by the specified layer of the seal quality evaluation decoder.

[0134] The first intermediate feature is used as the Query sequence of the feature interaction module, and the second intermediate feature is used as the Key-Value sequence of the feature interaction module to obtain the first enhanced feature after enhancing the first intermediate feature.

[0135] The second intermediate feature is used as the query sequence of the feature interaction module, and the first intermediate feature is used as the key-value sequence of the feature interaction module to obtain the second enhanced feature after enhancing the second intermediate feature.

[0136] The seal extraction decoder includes a transposed convolutional layer and multiple concatenated convolutional layers, and the second determining unit is specifically used for:

[0137] The general features are upsampled by the transposed convolutional layer to obtain the first decoded features;

[0138] The first output feature, the first decoding feature, and the first enhancement feature of the first residual block group are processed by the concatenated convolutional layer to obtain the second decoding feature;

[0139] The second output feature and the second decoding feature of the second residual block group are processed by the concatenated convolutional layer to obtain the third decoding feature;

[0140] The third output feature and the third decoding feature of the third residual block group are processed by the spliced ​​convolutional layer to obtain the seal image extraction result.

[0141] The seal quality evaluation decoder includes an attention pooling layer, an intermediate layer, and an activation function layer. The third determining unit is specifically used for:

[0142] The second enhanced feature is then subjected to attention pooling through the attention pooling layer to obtain the attention pooled feature.

[0143] The attention pooling features are processed by the intermediate layer to obtain the final features;

[0144] The final features are processed by the activation function layer to obtain the seal quality evaluation result.

[0145] The image information determination device provided in the embodiments of the present invention can execute the image information determination method provided in any embodiment of the present invention, and has the corresponding functional modules and beneficial effects of executing the method.

[0146] Example 4

[0147] Figure 9 A schematic diagram of an electronic device 90 that can be used to implement embodiments of the present invention is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices (e.g., helmets, glasses, watches, etc.), and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the invention described and / or claimed herein.

[0148] like Figure 9 As shown, the electronic device 90 includes at least one processor 91 and a memory, such as a read-only memory (ROM) 92 or a random access memory (RAM) 93, communicatively connected to the at least one processor 91. The memory stores computer programs executable by the at least one processor. The processor 91 can perform various appropriate actions and processes based on the computer program stored in the ROM 92 or loaded into the RAM 93 from storage unit 98. The RAM 93 can also store various programs and data required for the operation of the electronic device 90. The processor 91, ROM 92, and RAM 93 are interconnected via a bus 94. An input / output (I / O) interface 95 is also connected to the bus 94.

[0149] Multiple components in electronic device 90 are connected to I / O interface 95, including: input unit 96, such as keyboard, mouse, etc.; output unit 97, such as various types of displays, speakers, etc.; storage unit 98, such as disk, optical disk, etc.; and communication unit 99, such as network card, modem, wireless transceiver, etc. Communication unit 99 allows electronic device 90 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.

[0150] Processor 91 can be a variety of general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of processor 91 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various processors running machine learning decoder algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. Processor 91 performs the various methods and processes described above, such as image information determination methods.

[0151] In some embodiments, the image information determination method may be implemented as a computer program tangibly contained in a computer-readable storage medium, such as storage unit 98. In some embodiments, part or all of the computer program may be loaded and / or mounted on electronic device 90 via ROM 92 and / or communication unit 99. When the computer program is loaded into RAM 93 and executed by processor 91, one or more steps of the image information determination method described above may be performed. Alternatively, in other embodiments, processor 91 may be configured to perform the image information determination method by any other suitable means (e.g., by means of firmware).

[0152] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SoCs), payload-programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.

[0153] Computer programs used to implement the methods of the present invention may be written in any combination of one or more programming languages. These computer programs may be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that when executed by the processor, the computer programs cause the functions / operations specified in the flowcharts and / or block diagrams to be performed. The computer programs may be executed entirely on a machine, partially on a machine, or as a standalone software package, partially on a machine and partially on a remote machine, or entirely on a remote machine or server.

[0154] In the context of this invention, a computer-readable storage medium can be a tangible medium that may contain or store a computer program for use by or in conjunction with an instruction execution system, apparatus, or device. A computer-readable storage medium may include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination thereof. Alternatively, a computer-readable storage medium may be a machine-readable signal medium. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof.

[0155] To provide interaction with a user, the systems and techniques described herein can be implemented on an electronic device having: a display device for displaying information to the user (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor); and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the electronic device. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).

[0156] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as data servers), or computing systems that include middleware components (e.g., application servers), or computing systems that include frontend components (e.g., user computers with graphical user interfaces or web browsers through which users can interact with implementations of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., communication networks). Examples of communication networks include local area networks (LANs), wide area networks (WANs), blockchain networks, and the Internet.

[0157] A computing system can include clients and servers. Clients and servers are generally located far apart and typically interact through communication networks. The client-server relationship is created by computer programs running on the respective computers and having a client-server relationship with each other. The server can be a cloud server, also known as a cloud computing server or cloud host, which is a hosting product within the cloud computing service system to address the shortcomings of traditional physical hosts and VPS services, such as high management difficulty and weak business scalability.

[0158] In one embodiment, the present invention further includes a computer program product, which includes a computer program that, when executed by a processor, implements the image information determination method of any embodiment of the present invention.

[0159] In implementing the computer program product, computer program code for performing the operations of this invention can be written in one or more programming languages ​​or a combination thereof. Programming languages ​​include object-oriented programming languages ​​such as Java, Smalltalk, and C++, as well as conventional procedural programming languages ​​such as C or similar languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including a local area network (LAN) or a wide area network (WAN)—or can be connected to an external computer (e.g., via the Internet using an Internet service provider).

[0160] It should be understood that the various forms of processes shown above can be used, with steps reordered, added, or deleted. For example, the steps described in this invention can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution of this invention can be achieved, and this is not limited herein.

[0161] The specific embodiments described above do not constitute a limitation on the scope of protection of this invention. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this invention should be included within the scope of protection of this invention.

Claims

1. A method for determining image information, characterized in that, include: Acquire the image to be recognized; Based on the shared encoder, feature extraction is performed on the image to be identified to obtain general features; The general features are jointly processed by the seal extraction decoder, the seal quality evaluation decoder, and the feature interaction module to determine the image information, which includes the seal image extraction result and the seal quality evaluation result.

2. The method according to claim 1, characterized in that, The shared encoder includes convolutional blocks, max-pooling layers, and multiple cascaded residual block groups. Correspondingly, the process of identifying the image to be recognized using the shared encoder to obtain general features includes: The first feature is obtained by downsampling the image to be identified using the convolutional block; The first feature is downsampled using the max pooling layer to obtain the second feature; The second feature is extracted by each of the residual block groups to obtain a general feature; the residual block group includes a first residual block and a second residual block.

3. The method according to claim 1, characterized in that, The step of jointly processing the general features based on the seal extraction decoder, the seal quality evaluation decoder, and the feature interaction module to determine image information includes: The feature interaction module performs feature exchange on the specified features of the seal extraction decoder and the seal quality evaluation decoder to determine the first enhancement feature and the second enhancement feature. The seal image extraction result is determined based on the first enhanced feature, the general feature, and the seal extraction decoder; The seal quality evaluation result is determined based on the second enhanced feature and the seal quality evaluation decoder. The results of the seal image extraction and the seal quality evaluation are used as image information.

4. The method according to claim 3, characterized in that, The feature interaction module performs feature exchange on specified features of the seal extraction decoder and the seal quality evaluation decoder. The first enhancement feature and the second enhancement feature are determined, including: The general features are input into the seal extraction decoder to obtain the first intermediate features output by the specified layer of the seal extraction decoder. The general features are input into the seal quality evaluation decoder to obtain the second intermediate features output by the specified layer of the seal quality evaluation decoder. The first intermediate feature is used as the Query sequence of the feature interaction module, and the second intermediate feature is used as the Key-Value sequence of the feature interaction module to obtain the first enhanced feature after enhancing the first intermediate feature. The second intermediate feature is used as the query sequence of the feature interaction module, and the first intermediate feature is used as the key-value sequence of the feature interaction module to obtain the second enhanced feature after enhancing the second intermediate feature.

5. The method according to claim 3, characterized in that, The seal extraction decoder includes a transposed convolutional layer and multiple concatenated convolutional layers. Correspondingly, determining the seal image extraction result based on the first enhanced feature, the general feature, and the seal extraction decoder includes: The general features are upsampled by the transposed convolutional layer to obtain the first decoded features; The first output feature, the first decoding feature, and the first enhancement feature of the first residual block group are processed by the concatenated convolutional layer to obtain the second decoding feature; The second output feature and the second decoding feature of the second residual block group are processed by the concatenated convolutional layer to obtain the third decoding feature; The third output feature and the third decoding feature of the third residual block group are processed by the spliced ​​convolutional layer to obtain the seal image extraction result.

6. The method according to claim 3, characterized in that, The seal quality evaluation decoder includes an attention pooling layer, an intermediate layer, and an activation function layer. Correspondingly, determining the seal quality evaluation result based on the second enhanced feature and the seal quality evaluation decoder includes: The second enhanced feature is then subjected to attention pooling through the attention pooling layer to obtain the attention pooled feature. The attention pooling features are processed by the intermediate layer to obtain the final features; The final features are processed by the activation function layer to obtain the seal quality evaluation result.

7. An image information determining device, characterized in that, include: The image acquisition module is used to acquire the image to be recognized; The feature extraction module is used to extract features from the image to be identified based on the shared encoder to obtain general features; The information determination module is used to jointly process the general features based on the seal extraction decoder, the seal quality evaluation decoder, and the feature interaction module to determine image information, which includes the seal image extraction result and the seal quality evaluation result.

8. An electronic device, characterized in that, The electronic device includes: At least one processor; and A memory communicatively connected to the at least one processor; wherein, The memory stores a computer program that can be executed by the at least one processor, the computer program being executed by the at least one processor to enable the at least one processor to perform the image information determination method according to any one of claims 1-6.

9. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer instructions that cause a processor to execute the image information determination method according to any one of claims 1-6.

10. A computer program product, characterized in that, The computer program product includes a computer program that, when executed by a processor, implements the image information determination method according to any one of claims 1-6.