Image recognition method and device, electronic equipment and storage medium

By introducing backbone networks and task branch networks into the image recognition model, the problem of low recognition efficiency of multiple independent models in the prior art is solved, and flexible and efficient image content recognition is achieved.

CN120198718APending Publication Date: 2025-06-24ZHUHAI KINGSOFT OFFICE SOFTWARE +2
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510235788.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-28
Publication Date
2025-06-24

AI Technical Summary

Technical Problem

In the prior art, multiple independent models are used to identify specific content in the image separately, resulting in low recognition efficiency and poor flexibility, and it is difficult to meet the needs in scenarios where real-time processing and limited hardware resources are limited.

Method used

Image features are extracted through the backbone network in the image recognition model and input these features into multiple task branch networks. Each task branch network independently recognizes the content attribute probability in the image, and finally determines the target content attribute in the image based on the recognition results.

Benefits of technology

It realizes the recognition tasks of multiple different content attributes that require only one network model, improves the flexibility and efficiency of image content recognition, and is suitable for real-time processing and resource-limited scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120198718A_ABST
    Figure CN120198718A_ABST
Patent Text Reader

Abstract

The invention relates to an image recognition method and device, electronic equipment and a storage medium. The method comprises the steps of obtaining a to-be-recognized image; image features of the to-be-recognized image are extracted through a backbone network in an image recognition model, the image features are input into task branch networks in the image recognition model, and the task branch networks recognize the probability of content attributes contained in the to-be-recognized image according to the image features; and according to the identification result of each task branch network, determining a target content attribute contained in the to-be-identified image. Therefore, identification tasks of various different content attributes can be completed only through one network model, and the flexibility and the identification efficiency of image content identification are greatly improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computers, and in particular, to an image recognition method, apparatus, electronic device, and storage medium. Background Art

[0002] In the field of computer vision, the text graph content recognition method plays a crucial role. This technology aims to identify specific content in images, including but not limited to watermark detection, shadow detection, handwritten character recognition, and moiré analysis. These recognition capabilities have broad application prospects in the intelligent processing of images. They not only greatly improve the efficiency of image classification but also can provide users with customized image processing function suggestions based on the specific content identified. For example, for images containing watermarks, automatic watermark removal can be recommended; for pictures with shadows, shadow removal operations are suggested. This content-based intelligent recommendation greatly enriches the application scenarios of image processing and improves the user experience.

[0003] However, the current mainstream approach to text graph content recognition is to use multiple independent models to identify each specific content item separately. Although this approach ensures the accuracy of recognition to a certain extent, it has great limitations in actual applications, especially in business scenarios that require real-time processing and have relatively limited hardware configurations. The parallel operation of multiple models will significantly increase memory occupancy and place high demands on computer resources. If a serial processing method is used, although the memory pressure can be reduced, the processing time will be prolonged, making it difficult to meet the requirements of real-time processing. Therefore, how to achieve efficient and real-time text graph content recognition while ensuring recognition accuracy has become an urgent technical problem to be solved. Summary of the Invention

[0004] This application provides an image recognition method, apparatus, electronic device, and storage medium to solve the technical problems of low recognition efficiency and poor flexibility in the prior art, where multiple independent models are used to identify each specific content item separately.

[0005] In a first aspect, this application provides an image recognition method, and the method includes:

[0006] Obtain an image to be recognized;

[0007] Extract image features of the image to be recognized through a backbone network in the image recognition model, and input the image features into each task branch network in the image recognition model respectively. The task branch network determines the probability of the content attributes contained in the image to be recognized according to the image features;

[0008] Determine the target content attributes included in the image to be recognized according to the recognition results of each task branch network.

[0009] In a possible implementation, determining the target content attributes included in the image to be recognized according to the output results of each of the task branch networks includes:

[0010] Comparing the recognition result of the task branch network with a set threshold;

[0011] In the case where the recognition result is greater than or equal to the set threshold, determining the content attribute corresponding to the task branch network as one of the target content attributes included in the image to be recognized.

[0012] In a possible implementation, the backbone network includes a downsampling module, a hierarchical feature extraction module, and a feature integration module. Extracting the image features of the image to be recognized through the backbone network in the image recognition model includes:

[0013] Performing downsampling processing on the image to be recognized by the downsampling module to obtain a downsampling result;

[0014] Performing multi-level feature extraction processing on the downsampling result by using the hierarchical feature extraction module to obtain a multi-level feature extraction result;

[0015] Performing feature integration processing on the multi-level feature extraction result by using the feature integration module to obtain the image features of the image to be recognized;

[0016] The hierarchical feature extraction module includes a shallow feature extraction module and a deep feature extraction module; performing multi-level feature extraction processing on the downsampling result by using the hierarchical feature extraction module to obtain a multi-level feature extraction result includes:

[0017] Performing shallow feature extraction processing on the downsampling result by using the shallow feature extraction module to obtain a shallow feature extraction result;

[0018] Performing deep feature extraction processing on the shallow feature extraction result by using the deep feature extraction module to obtain a multi-level feature extraction result.

[0019] In a possible implementation, the shallow feature extraction module includes a first feature expansion and compression module and N second feature expansion and compression modules connected in sequence; performing shallow feature extraction processing on the downsampling result by using the shallow feature extraction module to obtain a shallow feature extraction result includes:

[0020] Performing first feature expansion and compression processing on the downsampling result by using the first feature expansion and compression module to obtain a first feature expansion and compression processing result;

[0021] Use the first second feature expansion and compression module in the shallow feature extraction module to perform second feature expansion and compression processing on the first feature expansion and compression processing result, obtaining the first second feature expansion and compression processing result;

[0022] Use the i-th second feature expansion and compression module in the shallow feature extraction module to perform second feature expansion and compression processing on the (i - 1)-th second feature expansion and compression processing result, obtaining the i-th second feature expansion and compression processing result, where i takes values one by one within the integer range of [2, N];

[0023] Take the N-th second feature expansion and compression processing result obtained by the N-th second feature expansion and compression module in the shallow feature extraction module as the shallow feature extraction result.

[0024] In a possible implementation manner, the deep feature extraction module includes a first feature expansion and compression module, a second feature expansion and compression module, and a third feature expansion and compression module; using the deep feature extraction module to perform deep feature extraction processing on the shallow feature extraction result, obtaining a multi-level feature extraction result, including:

[0025] Use the first feature expansion and compression module to perform first feature expansion and compression processing on the shallow feature extraction result, obtaining a first feature expansion and compression processing result;

[0026] Use the second feature expansion and compression module to perform second feature expansion and compression processing on the first feature expansion and compression processing result, obtaining a second feature expansion and compression processing result;

[0027] Use the third feature expansion and compression module to perform third feature expansion and compression processing on the second feature expansion and compression processing result, obtaining a multi-level feature extraction result.

[0028] In a possible implementation manner, the first feature expansion and compression module includes: a feature expansion and activation module, a deep feature extraction module, and a feature compression module; using the first feature expansion and compression module to perform first feature expansion and compression processing on the downsampling result, obtaining a first feature expansion and compression processing result, including:

[0029] Use the feature expansion and activation module to perform feature expansion and activation processing on the downsampling result, obtaining a feature expansion and activation processing result;

[0030] Use the deep feature extraction module to perform deep feature extraction processing on the feature expansion and activation processing result, obtaining a deep feature extraction result;

[0031] Use the feature compression module to perform feature compression processing on the deep feature extraction result to obtain a feature expansion and compression processing result.

[0032] In a possible implementation, the second feature expansion and compression module includes: a feature expansion and activation module, a deep feature extraction and residual connection module, and a feature compression module; the use of the first second feature expansion and compression module in the shallow feature extraction module to perform second feature expansion and compression processing on the feature expansion and compression processing result to obtain a first second feature expansion and compression processing result includes:

[0033] Use the feature expansion and activation module to perform feature expansion and activation processing on the feature expansion and compression processing result to obtain a feature expansion and activation processing result;

[0034] Use the deep feature extraction and residual connection module to perform deep feature extraction and residual connection processing on the feature expansion and activation processing result to obtain a deep feature extraction and residual connection processing result;

[0035] Use the feature compression module to perform feature compression processing on the deep feature extraction and residual connection processing result to obtain a second feature expansion and compression processing result.

[0036] In a possible implementation, the task branch network includes: a feature extraction layer, a global average pooling layer, and a fully connected layer; the task branch network's identification of the probability that the to-be-identified image contains a corresponding content attribute based on the image feature includes:

[0037] Use the feature extraction layer to perform high-dimensional feature extraction processing on the image feature to obtain a high-dimensional feature extraction result;

[0038] Use the global average pooling layer to perform pooling processing on the high-dimensional feature extraction result to obtain a pooling processing result;

[0039] Use the fully connected layer to perform mapping processing on the pooling processing result to obtain the probability that the to-be-identified image contains a corresponding content attribute.

[0040] In a possible implementation, the task branch network includes: M sequentially connected feature extraction layers, a global average pooling layer, and a fully connected layer; the task branch network's identification of the probability that the to-be-identified image contains a corresponding content attribute based on the image feature includes:

[0041] Use the first feature extraction layer to perform high-dimensional feature extraction processing on the image feature to obtain a first high-dimensional feature extraction result;

[0042] Performing high-dimensional feature extraction processing on the (j-1)-th high-dimensional feature extraction result by using the j-th feature extraction layer to obtain the j-th high-dimensional feature extraction result; j takes values one by one within the integer range of [2, M].

[0043] Performing pooling processing on the M-th high-dimensional feature extraction result by using the global average pooling layer to obtain a pooling processing result.

[0044] Performing mapping processing on the pooling processing result by using the fully connected layer to obtain the probability that the image to be recognized contains the corresponding content attribute.

[0045] In a possible implementation manner, the image recognition model is trained through the following steps:

[0046] Obtaining a set of original text images.

[0047] Performing the following processing on each original text image in the set of original text images: adding the content of at least one content attribute to the original text image to obtain a training sample image.

[0048] Inputting the training sample image into an initial model to obtain a model recognition result of the training sample image.

[0049] Constructing a loss function for each task branch network in the initial model according to the model recognition result, and performing weighted summation processing on the loss functions of the respective task branch networks to obtain a comprehensive loss function.

[0050] Adjusting the model parameters of the initial model according to the comprehensive loss function until the initial model converges to obtain an image recognition model.

[0051] In a second aspect, the present application provides an image recognition device, and the device includes:

[0052] An image acquisition module, configured to acquire an image to be recognized.

[0053] A model recognition module, configured to extract image features of the image to be recognized by using a backbone network in the image recognition model through the image recognition model, and input the image features into respective task branch networks in the image recognition model, and the task branch network recognizes the probability of the content attribute contained in the image to be recognized according to the image features.

[0054] A result determination module, configured to determine a target content attribute included in the image to be recognized according to the recognition results of the respective task branch networks.

[0055] In a possible implementation manner, the result determination module is specifically configured to:

[0056] Compare the recognition result of the task branch network with a set threshold value;

[0057] In the case where the recognition result is greater than or equal to the set threshold value, determine the content attribute corresponding to the task branch network as one of the target content attributes included in the image to be recognized.

[0058] In a possible implementation manner, the backbone network includes a downsampling module, a hierarchical feature extraction module, and a feature integration module. Extracting the image features of the image to be recognized through the backbone network in the image recognition model includes:

[0059] Perform downsampling processing on the image to be recognized by the downsampling module to obtain a downsampling result;

[0060] Perform multi-level feature extraction processing on the downsampling result by using the hierarchical feature extraction module to obtain a multi-level feature extraction result;

[0061] Perform feature integration processing on the multi-level feature extraction result by using the feature integration module to obtain the image features of the image to be recognized;

[0062] The hierarchical feature extraction module includes a shallow feature extraction module and a deep feature extraction module; performing multi-level feature extraction processing on the downsampling result by using the hierarchical feature extraction module to obtain a multi-level feature extraction result includes:

[0063] Perform shallow feature extraction processing on the downsampling result by using the shallow feature extraction module to obtain a shallow feature extraction result;

[0064] Perform deep feature extraction processing on the shallow feature extraction result by using the deep feature extraction module to obtain a multi-level feature extraction result.

[0065] In a possible implementation manner, the shallow feature extraction module includes a first feature expansion and compression module and N second feature expansion and compression modules connected in sequence; performing shallow feature extraction processing on the downsampling result by using the shallow feature extraction module to obtain a shallow feature extraction result includes:

[0066] Perform first feature expansion and compression processing on the downsampling result by using the first feature expansion and compression module to obtain a first feature expansion and compression processing result;

[0067] Perform second feature expansion and compression processing on the first feature expansion and compression processing result by using the first of the second feature expansion and compression modules in the shallow feature extraction module to obtain a first second feature expansion and compression processing result;

[0068] The $i$-th second feature expansion and compression module in the shallow feature extraction module performs second feature expansion and compression processing on the $(i - 1)$-th second feature expansion and compression processing result to obtain the $i$-th second feature expansion and compression processing result, where $i$ takes values one by one within the integer range of $[2, N]$;

[0069] The $N$-th second feature expansion and compression processing result obtained by the $N$-th second feature expansion and compression module in the shallow feature extraction module is used as the shallow feature extraction result.

[0070] In a possible implementation, the deep feature extraction module includes a first feature expansion and compression module, a second feature expansion and compression module, and a third feature expansion and compression module; the deep feature extraction module is used to perform deep feature extraction processing on the shallow feature extraction result to obtain a multi-level feature extraction result, including:

[0071] The first feature expansion and compression module is used to perform first feature expansion and compression processing on the shallow feature extraction result to obtain a first feature expansion and compression processing result;

[0072] The second feature expansion and compression module is used to perform second feature expansion and compression processing on the first feature expansion and compression processing result to obtain a second feature expansion and compression processing result;

[0073] The third feature expansion and compression module is used to perform third feature expansion and compression processing on the second feature expansion and compression processing result to obtain a multi-level feature extraction result.

[0074] In a possible implementation, the first feature expansion and compression module includes: a feature expansion and activation module, a deep feature extraction module, and a feature compression module; the first feature expansion and compression module is used to perform first feature expansion and compression processing on the downsampling result to obtain a first feature expansion and compression processing result, including:

[0075] The feature expansion and activation module is used to perform feature expansion and activation processing on the downsampling result to obtain a feature expansion and activation processing result;

[0076] The deep feature extraction module is used to perform deep feature extraction processing on the feature expansion and activation processing result to obtain a deep feature extraction result;

[0077] The feature compression module is used to perform feature compression processing on the deep feature extraction result to obtain a feature expansion and compression processing result.

[0078] In a possible implementation manner, the second feature expansion and compression module includes: a feature expansion and activation module, a deep feature extraction and residual connection module, and a feature compression module; using the first second feature expansion and compression module in the shallow feature extraction module to perform second feature expansion and compression processing on the feature expansion and compression processing result to obtain a first second feature expansion and compression processing result, including:

[0079] Using the feature expansion and activation module to perform feature expansion and activation processing on the feature expansion and compression processing result to obtain a feature expansion and activation processing result;

[0080] Using the deep feature extraction and residual connection module to perform deep feature extraction and residual connection processing on the feature expansion and activation processing result to obtain a deep feature extraction and residual connection processing result;

[0081] Using the feature compression module to perform feature compression processing on the deep feature extraction and residual connection processing result to obtain a second feature expansion and compression processing result.

[0082] In a possible implementation manner, the task branch network includes: a feature extraction layer, a global average pooling layer, and a fully connected layer; the task branch network identifying the probability that the to-be-identified image contains a corresponding content attribute according to the image feature includes:

[0083] Using the feature extraction layer to perform high-dimensional feature extraction processing on the image feature to obtain a high-dimensional feature extraction result;

[0084] Using the global average pooling layer to perform pooling processing on the high-dimensional feature extraction result to obtain a pooling processing result;

[0085] Using the fully connected layer to perform mapping processing on the pooling processing result to obtain the probability that the to-be-identified image contains a corresponding content attribute.

[0086] In a possible implementation manner, the task branch network includes: M sequentially connected feature extraction layers, a global average pooling layer, and a fully connected layer; the task branch network identifying the probability that the to-be-identified image contains a corresponding content attribute according to the image feature includes:

[0087] Using the first feature extraction layer to perform high-dimensional feature extraction processing on the image feature to obtain a first high-dimensional feature extraction result;

[0088] Using the j-th feature extraction layer to perform high-dimensional feature extraction processing on the (j - 1)-th high-dimensional feature extraction result to obtain a j-th high-dimensional feature extraction result; j takes values one by one within the integer interval of [2, M];

[0089] Perform pooling processing on the M-th high-dimensional feature extraction result by using the global average pooling layer to obtain a pooling processing result;

[0090] Perform mapping processing on the pooling processing result by using the fully connected layer to obtain the probability that the image to be recognized contains the corresponding content attribute.

[0091] In a possible implementation manner, the image recognition model is trained through the following steps:

[0092] Obtain a set of original text images;

[0093] Perform the following processing on each original text image in the set of original text images: Add the content of at least one content attribute to the original text image to obtain a training sample image;

[0094] Input the training sample image into the initial model to obtain the model recognition result of the training sample image;

[0095] Construct the loss function of each task branch network in the initial model according to the model recognition result, and perform weighted summation processing on the loss functions of the task branch networks to obtain a comprehensive loss function;

[0096] Adjust the model parameters of the initial model according to the comprehensive loss function until the initial model converges to obtain an image recognition model.

[0097] In a second aspect, the present application provides an image recognition device, and the device includes:

[0098] An image acquisition module, configured to acquire an image to be recognized;

[0099] A model recognition module, configured to extract the image features of the image to be recognized by using a backbone network in the image recognition model through the image recognition model, and input the image features into each task branch network in the image recognition model respectively, and the task branch network recognizes the probability of the content attribute contained in the image to be recognized according to the image features;

[0100] A result determination module, configured to determine the target content attribute included in the image to be recognized according to the recognition results of the task branch networks.

[0101] In a possible implementation manner, the result determination module is specifically configured to:

[0102] Compare the recognition result of the task branch network with a set threshold;

[0103] When the recognition result is greater than or equal to the set threshold, determine the content attribute corresponding to the task branch network as one of the target content attributes included in the image to be recognized.

[0104] In a possible implementation manner, the backbone network includes a downsampling module, a hierarchical feature extraction module, and a feature integration module. Extracting the image features of the image to be recognized by the backbone network in the image recognition model includes:

[0105] Perform downsampling processing on the image to be recognized by the downsampling module to obtain a downsampling result;

[0106] Use the hierarchical feature extraction module to perform multi-level feature extraction on the downsampling result to obtain a multi-level feature extraction result;

[0107] Use the feature integration module to perform feature integration processing on the multi-level feature extraction result to obtain the image features of the image to be recognized.

[0108] In a possible implementation manner, the hierarchical feature extraction module includes a shallow feature extraction module and a deep feature extraction module; using the hierarchical feature extraction module to perform multi-level feature extraction on the downsampling result to obtain a multi-level feature extraction result includes:

[0109] Use the shallow feature extraction module to perform shallow feature extraction processing on the downsampling result to obtain a shallow feature extraction result;

[0110] Use the deep feature extraction module to perform deep feature extraction processing on the shallow feature extraction result to obtain a multi-level feature extraction result.

[0111] In a possible implementation manner, the shallow feature extraction module includes a feature expansion and compression module and N sequentially connected second feature expansion and compression modules; using the shallow feature extraction module to perform shallow feature extraction processing on the downsampling result to obtain a shallow feature extraction result includes:

[0112] Use the feature expansion and compression module to perform feature expansion and compression processing on the downsampling result to obtain a feature expansion and compression processing result;

[0113] Use the first second feature expansion and compression module in the shallow feature extraction module to perform second feature expansion and compression processing on the feature expansion and compression processing result to obtain a first second feature expansion and compression processing result;

[0114] The i-th second feature expansion and compression module in the shallow feature extraction module performs second feature expansion and compression processing on the (i - 1)-th second feature expansion and compression processing result to obtain the i-th second feature expansion and compression processing result, where i takes values one by one within the integer range of [2, N];

[0115] The N-th second feature expansion and compression processing result obtained by the N-th second feature expansion and compression module in the shallow feature extraction module is used as the shallow feature extraction result.

[0116] In a possible implementation, the deep feature extraction module includes a first feature expansion and compression module, a second feature expansion and compression module, and a second feature expansion and compression module; the deep feature extraction module is used to perform deep feature extraction processing on the shallow feature extraction result to obtain a multi-level feature extraction result, including:

[0117] The first feature expansion and compression module is used to perform feature expansion and compression processing on the shallow feature extraction result to obtain a first feature expansion and compression processing result;

[0118] The second feature expansion and compression module is used to perform second feature expansion and compression processing on the first feature expansion and compression processing result to obtain a second feature expansion and compression processing result;

[0119] The second feature expansion and compression module is used to perform feature expansion and compression processing on the second feature expansion and compression processing result to obtain a multi-level feature extraction result.

[0120] In a possible implementation, the feature expansion and compression module includes: a feature expansion and activation module, a deep feature extraction module, and a feature compression module; the feature expansion and compression module is used to perform feature expansion and compression processing on the downsampling result to obtain a feature expansion and compression processing result, including:

[0121] The feature expansion and activation module is used to perform feature expansion and activation processing on the downsampling result to obtain a feature expansion and activation processing result;

[0122] The deep feature extraction module is used to perform deep feature extraction processing on the feature expansion and activation processing result to obtain a deep feature extraction result;

[0123] The feature compression module is used to perform feature compression processing on the deep feature extraction result to obtain a feature expansion and compression processing result.

[0124] In a possible implementation manner, the second feature expansion and compression module includes: a feature expansion and activation module, a deep feature extraction and residual connection module, and a feature compression module; the use of the first second feature expansion and compression module in the shallow feature extraction module to perform second feature expansion and compression processing on the feature expansion and compression processing result to obtain a first second feature expansion and compression processing result includes:

[0125] Using the feature expansion and activation module to perform feature expansion and activation processing on the feature expansion and compression processing result to obtain a feature expansion and activation processing result;

[0126] Using the deep feature extraction and residual connection module to perform deep feature extraction and residual connection processing on the feature expansion and activation processing result to obtain a deep feature extraction and residual connection processing result;

[0127] Using the feature compression module to perform feature compression processing on the deep feature extraction and residual connection processing result to obtain a second feature expansion and compression processing result.

[0128] In a possible implementation manner, the task branch network includes: a feature extraction layer, a global average pooling layer, and a fully connected layer; the task branch network's recognition of the probability that the image to be recognized contains the corresponding content attribute according to the image feature includes:

[0129] Using the feature extraction layer to perform high-dimensional feature extraction processing on the image feature to obtain a high-dimensional feature extraction result;

[0130] Using the global average pooling layer to perform pooling processing on the high-dimensional feature extraction result to obtain a pooling processing result;

[0131] Using the fully connected layer to perform mapping processing on the pooling processing result to obtain the probability that the image to be recognized contains the corresponding content attribute.

[0132] In a possible implementation manner, the task branch network includes: M sequentially connected feature extraction layers, a global average pooling layer, and a fully connected layer; the task branch network's recognition of the probability that the image to be recognized contains the corresponding content attribute according to the image feature includes:

[0133] Using the first feature extraction layer to perform high-dimensional feature extraction processing on the image feature to obtain a first high-dimensional feature extraction result;

[0134] Using the j-th feature extraction layer to perform high-dimensional feature extraction processing on the (j - 1)-th high-dimensional feature extraction result to obtain the j-th high-dimensional feature extraction result; j takes values one by one within the integer range of [2, M];

[0135] Use the global average pooling layer to perform pooling processing on the M-th high-dimensional feature extraction result to obtain a pooling processing result;

[0136] Use the fully connected layer to perform mapping processing on the pooling processing result to obtain the probability that the image to be recognized contains the corresponding content attribute.

[0137] In a possible implementation manner, the device further includes:

[0138] A model training module, configured to train an image recognition model in the following manner:

[0139] Obtain an original text image set;

[0140] Perform the following processing on each original text image in the original text image set: Add the content of at least one content attribute to the original text image to obtain a training sample image;

[0141] Input the training sample image into the initial model to obtain a model recognition result of the training sample image;

[0142] Construct a loss function for each task branch network in the initial model according to the model recognition result, and perform weighted summation processing on the loss functions of the task branch networks to obtain a comprehensive loss function;

[0143] Adjust the model parameters of the initial model according to the comprehensive loss function until the initial model converges to obtain an image recognition model.

[0144] In a third aspect, the present application provides an electronic device, including: a processor and a memory, where the processor is configured to execute an image recognition program stored in the memory to implement the image recognition method according to any one of the first aspects.

[0145] In a fourth aspect, the present application provides a storage medium, where the storage medium stores one or more programs, and the one or more programs can be executed by one or more processors to implement the image recognition method according to any one of the first aspects.

[0146] The above technical solutions provided by the embodiments of the present application have the following advantages compared with the prior art: In the method provided by the embodiments of the present application, by inputting the image to be recognized into the image recognition model, the backbone network in the image recognition model extracts the image features of the image to be recognized, and inputs the image features into each task branch network in the image recognition model respectively. The task branch network independently recognizes the probability that the image to be recognized contains the corresponding content attribute according to the image features. Finally, according to the recognition results of each task branch network, the target content attribute included in the image to be recognized is determined, realizing a recognition task of multiple different content attributes that can be completed with only one network model, greatly improving the flexibility and recognition efficiency of image content recognition. BRIEF DESCRIPTION OF THE DRAWINGS

[0147] The accompanying drawings herein are incorporated into the specification and form a part of the specification, showing embodiments consistent with the present application and, together with the specification, are used to explain the principles of the present application.

[0148] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the accompanying drawings required for the description of the embodiments or the prior art. Obviously, for those of ordinary skill in the art, other drawings can also be obtained based on these drawings without creative efforts.

[0149] One or more embodiments are exemplarily illustrated by the pictures in the corresponding accompanying drawings. These exemplary illustrations do not limit the embodiments. Elements with the same reference numerals in the drawings are represented as similar elements, unless otherwise stated, and the drawings in the drawings do not constitute a proportional limitation.

[0150] Figure 1 It is a flowchart of an embodiment of an image recognition method provided by an embodiment of the present application;

[0151] Figure 2 It is a schematic structural diagram of an image recognition model provided by an embodiment of the present application;

[0152] Figure 3 It is a schematic structural diagram of the backbone network in an image recognition model provided by an embodiment of the present application;

[0153] Figure 4 It is a schematic structural diagram of a feature expansion and compression module provided by an embodiment of the present application;

[0154] Figure 5 It is a schematic structural diagram of a second feature expansion and compression module provided by an embodiment of the present application;

[0155] Figure 6 It is a schematic structural diagram of a backbone network provided by an embodiment of the present application;

[0156] Figure 7 A structural schematic diagram of a task branch network provided by an embodiment of the present application;

[0157] Figure 8 Another structural schematic diagram of a task branch network provided by an embodiment of the present application;

[0158] Figure 9 A flowchart of a method for training an image recognition model provided by an embodiment of the present application;

[0159] Figure 10 A schematic diagram of a watermark;

[0160] Figure 11 A schematic diagram of an image with a watermark;

[0161] Figure 12 A schematic diagram of an image with a shadow;

[0162] Figure 13 An example diagram of a scene for constructing a training sample image with moiré patterns;

[0163] Figure 14 An example diagram of a scene for constructing a training sample image with handwriting;

[0164] Figure 15 A block diagram of an embodiment of an image recognition device provided by an embodiment of the present application;

[0165] Figure 16 A structural schematic diagram of an electronic device provided by an embodiment of the present application. Detailed implementation manners

[0166] To make the objectives, technical solutions, and advantages of the embodiments of the present application clearer, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present application. Apparently, the described embodiments are some, but not all, of the embodiments of the present application. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present application without creative efforts shall fall within the protection scope of the present application.

[0167] The following disclosure provides many different embodiments or examples for implementing different structures of the present application. To simplify the disclosure of the present application, components and settings of specific examples are described below. Of course, they are only examples and are not intended to limit the present application. In addition, the present application may repeat reference numerals and / or letters in different examples. Such repetition is for the purpose of simplification and clarity and does not itself indicate the relationship between the various embodiments and / or settings discussed.

[0168] To solve the technical problems of low recognition efficiency and poor flexibility in the prior art, where multiple independent models are used to identify each specific content item separately, the present application provides an image recognition method, apparatus, electronic device, and storage medium, which can achieve the recognition task of multiple different content attributes with only one network model, greatly improving the flexibility and recognition efficiency of image content recognition.

[0169] Figure 1 It is a flowchart of an embodiment of an image recognition method provided by an embodiment of the present application. As Figure 1 shown, the method includes the following steps:

[0170] Step 101, obtain an image to be recognized.

[0171] The image to be recognized refers to an image for which it is to be determined whether it contains content with specific content attributes. Among them, specific content attributes are, for example, watermarks, shadows, moiré patterns, handwriting, etc. In practical applications, an image may contain one or two or more types of content with specific content attributes. For example, an image contains both a watermark and handwriting. The embodiments of the present application do not limit this.

[0172] In one embodiment, the image to be recognized is specified by the user, and it can be an image obtained from various sources, such as files, cameras, networks, etc. The embodiments of the present application do not limit the specific source of the image to be recognized. Here, the image to be recognized can be a part of an overall image or a specified image area.

[0173] Step 102, extract image features of the image to be recognized through the backbone network in the image recognition model, and input the image features into each task branch network in the image recognition model respectively. The task branch network determines the probability that the image to be recognized contains the corresponding content attribute according to the image features.

[0174] Step 103, determine the target content attributes included in the image to be recognized according to the recognition results of each task branch network.

[0175] For ease of understanding, the following provides a unified description of Step 102 and Step 103:

[0176] First, refer to Figure 2 , which is a schematic structural diagram of an image recognition model provided by an embodiment of the present application. As Figure 2 shown, the image recognition model 20 includes a backbone network 21, a task branch network 22, a task branch network 23, a task branch network 24, and a task branch network 25.

[0177] Among them, the input end of the backbone network 21 is the input end of the entire image recognition model 20. The output end of the backbone network 21 is respectively connected to the input ends of the task branch network 22, the task branch network 23, the task branch network 24, and the task branch network 25. The output ends of the task branch network 22, the task branch network 23, the task branch network 24, and the task branch network 25 constitute the output end of the entire image recognition model 20. It can be seen that the image recognition model 20 is a multi-task model, that is, a network model that can execute multiple tasks in parallel.

[0178] It should be noted that Figure 2 only takes the image recognition model including four task branch networks as an example for explanation. In practical applications, the number of task branch networks in the image recognition model can be set according to business needs, and the embodiments of the present application do not limit this.

[0179] Based on Figure 2 the model structure shown in the example, in step 102, the image to be recognized is input into the image recognition model. First, the backbone network in the image recognition model extracts the image features of the image to be recognized. Among them, the image features include but are not limited to: edge features, corner features, texture features, etc. of the content in the image. Here, the features of different dimensions extracted are collectively referred to as image features. By extracting the features of the image to be recognized, the deep information and more expressive information of the image to be recognized can be obtained, which is helpful for the subsequent recognition of content attributes.

[0180] Subsequently, the backbone network synchronously inputs the extracted image features into different task branch networks. Here, the task branch network is used to recognize the content attributes contained in the image according to the graphic features, such as content attributes such as watermarks, shadows, moiré patterns, or handwriting. And different task branch networks are used to recognize different content attributes. For example, Figure 2 the task branch network 22 shown in the example is used to recognize whether there is a watermark in the image, the task branch network 23 is used to recognize whether there is a shadow in the image, the task branch network 24 is used to recognize whether there is a moiré pattern in the image, and the task branch network 25 is used to recognize whether there is handwriting in the image.

[0181] It can be seen from this that in the embodiments of the present application, different task networks are integrated into one model to obtain an image recognition model. This image recognition model can support the recognition of multiple content attributes at the same time, such as supporting the recognition of four content attributes: watermarks, shadows, moiré patterns, and handwriting. Thus, it can quickly obtain which content attributes are contained in the image to be recognized, and further help to quickly classify the image to be recognized or make corresponding function recommendations according to the recognition results, such as removing watermarks, removing shadows, removing screen patterns, removing handwriting, etc., which can effectively improve the efficiency of image intelligent processing.

[0182] In one embodiment, the recognition result of the task branch network represents the probability that the image to be recognized contains the corresponding content attribute, indicating the possibility that the image to be recognized contains the corresponding content attribute. As an alternative implementation, a set threshold is set to map the recognition result of the task branch network to a definite recognition result, that is, whether it contains the corresponding content attribute. Specifically, if the recognition result of the task branch network is greater than or equal to the set threshold, it is considered that the possibility that the image to be recognized contains the content attribute corresponding to this task branch network is relatively high, and thus it can be considered that the image to be recognized contains the content attribute corresponding to this task branch network. On the contrary, if the recognition result of the task branch network is less than the set threshold, it is considered that the possibility that the image to be recognized contains the content attribute corresponding to this task branch network is relatively low, and thus it can be considered that the image to be recognized does not contain the content attribute corresponding to this task branch network.

[0183] Accordingly, in step 103, the specific implementation of determining the target content attribute included in the image to be recognized according to the output results of each task branch network includes: performing the following processing on the recognition result of each task branch network: comparing the recognition result of the task branch network with the set threshold; in the case where it is compared that the recognition result of this task branch network is greater than or equal to the set threshold, determining the content attribute corresponding to this task branch network as one of the target content attributes included in the image to be recognized.

[0184] Optionally, different set thresholds can be set for different task branch networks. In practical applications, the selection of the set threshold may be affected by various factors, including but not limited to the following points:

[0185] Performance of the task branch network: For example, a network with better performance can set a lower threshold, while a network with poorer performance needs to set a higher threshold.

[0186] Requirements of the application scenario: In some application scenarios, it may be more inclined to avoid missed detections, so a lower threshold can be set, while in other scenarios, it may be more focused on avoiding false detections, so a higher threshold can be set.

[0187] In addition, as an embodiment, before inputting the image to be recognized into the image recognition model, the image to be recognized can be preprocessed first, including but not limited to steps such as grayscale conversion, binarization, denoising, and normalization. The embodiments of the present application do not limit this. Preprocessing the image to be recognized first helps the subsequent feature extraction process.

[0188] The technical solution provided by the embodiments of this application inputs the image to be recognized into an image recognition model. The backbone network in the image recognition model extracts the image features of the image to be recognized, and inputs the image features into each task branch network in the image recognition model respectively. The task branch network independently recognizes the probability that the image to be recognized contains the corresponding content attribute according to the image features. Finally, according to the recognition results of each task branch network, the target content attribute included in the image to be recognized is determined, realizing a recognition task of multiple different content attributes that can be completed with only one network model, greatly improving the flexibility and recognition efficiency of image content recognition.

[0189] Figure 3 It is a schematic structural diagram of the backbone network in an image recognition model provided by the embodiments of this application. As Figure 3 shown, the backbone network includes a downsampling module, a hierarchical feature extraction module, and a feature integration module. Among them, the input end of the downsampling module is the input end of the entire backbone network, the output end of the downsampling module is connected to the input end of the hierarchical feature extraction module, the output end of the hierarchical feature extraction module is connected to the input end of the feature integration module, and the output end of the feature integration module is the output end of the entire backbone network.

[0190] Based on Figure 3 the structure of the exemplary backbone network, in the technical solution of the embodiments of this application, the process of extracting the image features of the image to be recognized by the backbone network in the image recognition model includes: the downsampling module performs downsampling processing on the image to be recognized to obtain a downsampling result; the hierarchical feature extraction module performs multi-level feature extraction on the downsampling result to obtain a multi-level feature extraction result; the feature integration module performs feature integration processing on the multi-level feature extraction result to obtain the image features of the image to be recognized.

[0191] Specifically, the main purpose of the downsampling module is to reduce the subsequent calculation amount by reducing the resolution of the input image, and at the same time capture the low-level features of the image (such as edge, texture and other features) to provide input for the hierarchical feature extraction module. Exemplarily, the downsampling module is implemented by a convolutional layer (such as conv2, with a kernel size of 3*3) and an activation function (such as the ReLU6 function).

[0192] The hierarchical feature extraction module is the core part of the backbone network, which is used to gradually extract different levels of features of the image. These features range from low-dimensional to high-dimensional, from local to global, and can cover all aspects of the image.

[0193] The feature integration module is located at the end of the backbone network and is used to integrate and enhance the features extracted previously to obtain a richer global feature representation. Exemplarily, the feature integration module is implemented by a convolutional layer (such as conv2, with a kernel size of 1*1) and an activation function (such as the ReLU6 function).

[0194] In summary, the downsampling module, the hierarchical feature extraction module, and the feature integration module together constitute the complete architecture of the backbone network. They cooperate with each other to jointly achieve the extraction, integration, and enhancement of image features, providing strong support for subsequent image recognition tasks.

[0195] For example, the size of the input image to be recognized is 224×224×3. Here, 224×224 represents that both the width and height of the image are 224 pixels, and 3 represents the number of color channels of the image (usually the three RGB color channels). The image feature image_feature is extracted through the backbone network in the image recognition model, and its feature dimension is [N, C, W, H]. Among them, N represents the batch size, that is, the number of images processed simultaneously in each training process. Exemplarily, it is set to 128 during training, which means that 128 images will be processed simultaneously in each training iteration. This value can be adjusted according to the available hardware resources (such as the size of GPU memory). C represents the number of channels, that is, the depth or number of feature maps. Exemplarily, C is 1280, which means there are 1280 feature values at each position (W×H). These feature values represent the responses of the image under the action of different filters (or called convolutional kernels), and each filter captures a specific feature of the image. W and H respectively represent the width and height of the feature map. Exemplarily, both W and H are 7, which means that after being processed by the backbone network, the spatial dimension of the feature map is reduced to 7×7. This is because in the backbone network, the spatial size of the feature map is usually gradually reduced through convolutional layers and pooling layers, while increasing the depth (number of channels) of the feature map to capture the high-level semantic information of the image.

[0196] In one embodiment, continuing to refer to Figure 3 the example of, the hierarchical feature extraction module includes a shallow feature extraction module and a deep feature extraction module. Among them, the input end of the shallow feature extraction module is the input end of the entire hierarchical feature extraction module, the output end of the shallow feature extraction module is connected to the input end of the hierarchical feature extraction module, and the output end of the deep feature extraction module is the output end of the entire hierarchical feature extraction module.

[0197] Based on Figure 3 the structure of the hierarchical feature extraction module in the example, in the technical solution of the embodiment of the present application, the process of using the hierarchical feature extraction module to perform multi-level feature extraction on the foregoing downsampling result to obtain the multi-level feature extraction result includes: using the shallow feature extraction module to perform shallow feature extraction processing on the downsampling result to obtain the shallow feature extraction result; using the deep feature extraction module to perform deep feature extraction processing on the shallow feature extraction result to obtain the multi-level feature extraction result.

[0198] It can be seen that the hierarchical feature extraction module adopts a step - by - step approach. First, the shallow - layer feature extraction module captures the basic and local features of the image, such as edges and textures, which are usually more intuitive and easier to extract. Subsequently, the deep - layer feature extraction module further abstracts and combines the results of the shallow - layer feature extraction to extract higher - level and more complex features, which often contain the global information and semantic content of the image. In this way, the hierarchical feature extraction module can make full use of the feature information at different levels, improve the model's understanding and recognition ability of the image. At the same time, because the shallow - layer features and deep - layer features are effectively combined and utilized in the model, the accuracy and robustness of the model can be further improved.

[0199] The above explains the structure of the hierarchical feature extraction module as a whole. Next, the structures and processing procedures of the shallow - layer feature extraction module and the deep - layer feature extraction module will be described in detail respectively.

[0200] (1) Shallow - layer feature extraction module

[0201] In one embodiment, the shallow - layer feature extraction module includes a first feature expansion and compression module and N second feature expansion and compression modules, where N is a natural number greater than 1. For example, N is 3. Among them, the input end of the first feature expansion and compression module is the input end of the entire shallow - layer feature extraction module. The output end of the first feature expansion and compression module is connected to the input end of the first second feature expansion and compression module. The N second feature expansion and compression modules are connected in sequence, and the output end of the last second feature expansion and compression module is the output end of the entire shallow - layer feature extraction module.

[0202] Based on the structure of the shallow - layer feature extraction module described above, the process of using the shallow - layer feature extraction module to perform shallow - layer feature extraction processing on the down - sampling result to obtain the shallow - layer feature extraction result includes: using the first feature expansion and compression module to perform the first feature expansion and compression processing on the down - sampling result to obtain the first feature expansion and compression processing result; using the first second feature expansion and compression module in the shallow - layer feature extraction module to perform the second feature expansion and compression processing on the first feature expansion and compression processing result to obtain the first second feature expansion and compression processing result; using the i - th second feature expansion and compression module in the shallow - layer feature extraction module to perform the second feature expansion and compression processing on the (i - 1) - th second feature expansion and compression processing result to obtain the i - th second feature expansion and compression processing result, where i takes values one by one in the integer range of [2, N]; taking the N - th second feature expansion and compression processing result obtained by the N - th second feature expansion and compression module in the shallow - layer feature extraction module as the shallow - layer feature extraction result.

[0203] To facilitate the understanding of the above processing procedure, the first feature expansion and compression module and the second feature expansion and compression module are respectively explained as follows:

[0204] In one embodiment, referring to Figure 4 , the first feature expansion and compression module includes: a feature expansion and activation module, a deep feature extraction module, and a feature compression module. Among them, the input end of the feature expansion and activation module is the input end of the entire first feature expansion and compression module. The output end of the feature expansion and activation module is connected to the input end of the deep feature extraction module. The output end of the deep feature extraction module is connected to the input end of the feature compression module. The output end of the feature compression module is the output end of the entire first feature expansion and compression module.

[0205] Based on Figure 4 the structure of the first feature expansion and compression module shown in the example, the process of using the first feature expansion and compression module to perform the first feature expansion and compression processing on the downsampling result to obtain the first feature expansion and compression processing result includes: using the feature expansion and activation module to perform feature expansion and activation processing on the downsampling result to obtain the feature expansion and activation processing result; using the deep feature extraction module to perform deep feature extraction processing on the feature expansion and activation processing result to obtain the deep feature extraction result; using the feature compression module to perform feature compression processing on the deep feature extraction result to obtain the first feature expansion and compression processing result.

[0206] Specifically, the core task of the feature expansion and activation module is to expand the input data in the channel dimension, thereby increasing the dimension of the feature space. Exemplarily, the feature expansion and activation module is implemented by a convolutional layer (such as conv2, with a convolution kernel size of 3*3) and an activation function (such as the ReLU6 function). This convolutional layer can introduce new feature channels, making the feature representation more rich and diverse. The activation function (such as the ReLU6 function) performs a non-linear transformation on the output of the convolutional layer, further enhancing the expression ability of the features. This combination not only increases the dimension of the features but also enhances the discriminability of the features through non-linear activation, laying a solid foundation for subsequent feature extraction and processing.

[0207] The deep feature extraction module is used to perform filtering in the extended high-dimensional space and extract local features. Exemplarily, the deep feature extraction module is implemented by using a depth convolution layer (such as a Dwise convolution layer with a convolution kernel size of 3*3 and a stride of 2) and an activation function (such as the ReLU6 function). The depth convolution layer independently extracts spatial information in the extended feature channels, thereby avoiding the mixing between channels, enabling each channel to focus on capturing specific spatial features, and greatly improving the computational efficiency while ensuring the richness of features. The ReLU6 activation function further non-linearly enhances the feature expression ability, making the extracted features more accurate and robust.

[0208] The feature compression module is used to effectively compress the features after high-dimensional expansion back into a lower-dimensional representation form, thereby significantly reducing the number of parameters and memory consumption required for subsequent calculations, and enabling the network to extract the most valuable features within a limited dimension. Exemplarily, the feature compression module is implemented by using a convolution layer (such as conv2 with a convolution kernel size of 1*1) and an activation function (such as the ReLU6 function). The convolution layer integrates feature information in a cross-channel manner and maps the high-dimensional features to a low-dimensional space. The ReLU6 activation function is used again to non-linearly enhance the feature expression ability, ensuring that the compressed features still have sufficient discriminability and robustness.

[0209] It can be seen that the feature expansion and activation module, the deep feature extraction module, and the feature compression module together constitute an efficient and orderly feature processing flow, providing the image recognition model with powerful feature extraction and representation capabilities.

[0210] In one embodiment, referring to Figure 5 , the second feature expansion and compression module includes: a feature expansion and activation module, a deep feature extraction and residual connection module, and a feature compression module. Among them, the input end of the feature expansion and activation module is the input end of the entire second feature expansion and compression module. The output end of the feature expansion and activation module is connected to the input end of the deep feature extraction and residual connection module. The output end of the deep feature extraction and residual connection module is connected to the input end of the feature compression module. The output end of the feature compression module is the output end of the entire second feature expansion and compression module.

[0211] Based on Figure 5The structure of the second feature expansion and compression module in the example uses the first second feature expansion and compression module in the shallow feature extraction module to perform second feature expansion and compression processing on the feature expansion and compression processing result, and the processing process for obtaining the first second feature expansion and compression processing result includes: using the feature expansion and activation module to perform feature expansion and activation processing on the feature expansion and compression processing result to obtain a feature expansion and activation processing result; using the deep feature extraction and residual connection module to perform deep feature extraction and residual connection processing on the feature expansion and activation processing result to obtain a deep feature extraction and residual connection processing result; using the feature compression module to perform feature compression processing on the deep feature extraction and residual connection processing result to obtain a second feature expansion and compression processing result.

[0212] Specifically, for the structure and functions of the feature expansion and activation module and the feature compression module, reference can be made to the above explanations about the feature expansion and compression module, which will not be elaborated here.

[0213] Compared with the above-described deep feature extraction module, the deep feature extraction and residual connection module has two significant differences. First, when using the depth convolution layer for feature extraction, its stride is set to 1. This design aims to ensure that the sizes of the front and rear feature maps are consistent, so that the residual connection technology can be seamlessly applied. Second, the deep feature extraction and residual connection module introduces a residual connection mechanism. In deep learning, this mechanism can not only effectively alleviate the problem of gradient disappearance, making it possible to build deeper neural networks, but also enables the network to better retain and reuse low-level information by directly adding the input features to the output features. This design not only enhances the feature representation ability of the network but also helps to improve the overall performance and stability of the model.

[0214] In summary, it can be seen that the shallow feature extraction module can gradually and deeply mine and refine the feature information of the input data through the first feature expansion and compression module and N sequentially connected second feature expansion and compression modules. It not only improves the efficiency and accuracy of feature processing but also effectively alleviates the problem of gradient disappearance in deep learning through mechanisms such as residual connection, enabling the network to learn and understand data at a deeper level. As a result, with the gradual transmission and processing of data, the feature representation becomes gradually richer and deeper, providing strong support for subsequent task execution.

[0215] (2) Deep Feature Extraction Module

[0216] In one embodiment, the deep feature extraction module includes a first feature expansion and compression module, a second feature expansion and compression module, and a third feature expansion and compression module. Among them, the input end of the first feature expansion and compression module is the input end of the entire deep feature extraction module. The output end of the first feature expansion and compression module is connected to the input end of the second feature expansion and compression module. The output end of the second feature expansion and compression module is connected to the input end of the third feature expansion and compression module. The output end of the third feature expansion and compression module is the output end of the entire deep feature extraction module.

[0217] Based on the structure of the deep feature extraction module described above, the process of using the deep feature extraction module to perform deep feature extraction on the shallow feature extraction result to obtain a multi-level feature extraction result includes: using the first feature expansion and compression module to perform feature expansion and compression on the shallow feature extraction result to obtain a first feature expansion and compression result; using the second feature expansion and compression module to perform a second feature expansion and compression on the first feature expansion and compression result to obtain a second feature expansion and compression result; using the third feature expansion and compression module to perform feature expansion and compression on the second feature expansion and compression result to obtain a multi-level feature extraction result.

[0218] Optionally, the first feature expansion and compression module and the third feature expansion and compression module are the same. The structures and functions of these two modules can be referred to the explanations in the above embodiments and will not be elaborated here. Similarly, the structure and function of the second feature expansion and compression module can also be referred to the explanations in the above embodiments and will not be elaborated here.

[0219] For better understanding, here through Figure 6 shows the network structure of the backbone network in the image recognition model. As Figure 6 shown, in the order from top to bottom, the backbone network sequentially includes: a downsampling module (implemented by a 3*3 convolutional layer and a ReLU6 activation function), 1 first feature expansion and compression module (denoted as InvertedResidual_v1), 3 second feature expansion and compression modules (InvertedResidual_v2), 1 first feature expansion and compression module, 1 second feature expansion and compression module, 1 third feature expansion and compression module (the same as the first feature expansion and compression module), and a feature integration module (implemented by a 1*1 convolutional layer and a ReLU6 activation function).

[0220] Based on Figure 6The structure of the backbone network in the illustrated example. The process of the backbone network in the image recognition model for extracting the image features of the image to be recognized includes: downsampling the image to be recognized by a downsampling module to obtain a downsampling result; performing first feature expansion and compression processing on the downsampling result by using the first first feature expansion and compression module to obtain a first first feature expansion and compression processing result; performing second feature expansion and compression processing on the first feature expansion and compression processing result by using the first second feature expansion and compression module to obtain a first second feature expansion and compression processing result; performing second feature expansion and compression processing on the first second feature expansion and compression processing result by using the second second feature expansion and compression module to obtain a second second feature expansion and compression processing result; performing second feature expansion and compression processing on the second second feature expansion and compression processing result by using the third second feature expansion and compression module to obtain a third second feature expansion and compression processing result; performing first feature expansion and compression processing on the third second feature expansion and compression processing result by using the second first feature expansion and compression module to obtain a second first feature expansion and compression processing result; performing second feature expansion and compression processing on the second first feature expansion and compression processing result by using the fourth second feature expansion and compression module to obtain a fourth second feature expansion and compression processing result; performing first feature expansion and compression processing on the fourth second feature expansion and compression processing result by using the third first feature expansion and compression module to obtain a multi-level feature extraction result. Finally, performing feature integration processing on the multi-level feature extraction result by using a feature integration module to obtain the image features of the image to be recognized.

[0221] The above has made a detailed explanation of the structure and function of the backbone network in the image recognition model. Next, an explanation of the task branch network in the image recognition model will be given:

[0222] In one embodiment, referring to Figure 7 , the task branch network includes: a feature extraction layer, a global average pooling layer, and a fully connected layer. Among them, the input end of the feature extraction layer is the input end of the entire task branch network. The output end of the feature extraction layer is connected to the input end of the global average pooling layer. The output end of the global average pooling layer is connected to the input end of the fully connected layer. The output end of the fully connected layer is the output end of the entire task branch network.

[0223] Based on Figure 7For the structure of the exemplified task branch network, in the technical solution of the embodiments of the present application, the processing process in which the task branch network identifies the probability that the image to be recognized contains the corresponding content attribute according to the image features includes: using the feature extraction layer to perform high-dimensional feature extraction processing on the image features to obtain a high-dimensional feature extraction result; using the global average pooling layer to perform pooling processing on the high-dimensional feature extraction result to obtain a pooling processing result; using the fully connected layer to perform mapping processing on the pooling processing result to obtain the probability that the image to be recognized contains the corresponding content attribute.

[0224] Specifically, the feature extraction layer is used to extract high-dimensional features closely related to the task in the image, and these features are usually edges, textures, shapes, etc. in the image. Exemplarily, the feature extraction layer is implemented by a convolutional layer (for example, conv2, with a convolutional kernel size of 3*3) and an activation function (for example, the ReLU6 function). The convolutional layer extracts local features in the image through convolutional operations. The activation function is used to introduce non-linearity so that the neural network can learn complex patterns.

[0225] The global average pooling layer compresses the spatial dimension of the feature map of each channel to 1 (that is, the output size is 1×1), which means that for each channel, it calculates the average value at all positions, thereby obtaining a fixed-length vector. Through the processing of the global average pooling layer, spatial information can be effectively aggregated, enhancing the robustness of the model to spatial deformation.

[0226] The fully connected layer maps the feature vector obtained by global average pooling to the number of classes. For a binary classification task, the number of classes is 1. The fully connected layer realizes the final image classification task by mapping the features to the probability scores of the corresponding classes.

[0227] Thus, Figure 7 For the exemplified task branch network, high-dimensional features are extracted through a convolutional layer and an activation function, spatial information is aggregated and the number of parameters is reduced through the global average pooling layer, and finally the features are mapped to the number of classes through the fully connected layer, so that the task branch network can be applied to the content attribute recognition task of the present application.

[0228] In another embodiment, referring to Figure 8 , the task branch network includes M successively connected feature extraction layers, a global average pooling layer, and a fully connected layer, where M is a natural number greater than 1, for example, 3. Among them, the input end of the first feature extraction layer is the input end of the entire task branch network, the output end of the last feature extraction layer is connected to the input end of the global average pooling layer, the output end of the global average pooling layer is connected to the input end of the fully connected layer, and the output end of the fully connected layer is the output end of the entire task branch network.

[0229] Based on Figure 8For the structure of the exemplified task branch network, in the technical solution of the embodiment of the present application, the processing process of the task branch network for recognizing the probability that the image to be recognized contains the corresponding content attribute according to the image features includes: using the first feature extraction layer to perform high-dimensional feature extraction processing on the image features to obtain the first high-dimensional feature extraction result; using the j-th feature extraction layer to perform high-dimensional feature extraction processing on the (j - 1)-th high-dimensional feature extraction result to obtain the j-th high-dimensional feature extraction result; j takes values one by one within the integer range of [2, M]; using the global average pooling layer to perform pooling processing on the M-th high-dimensional feature extraction result to obtain the pooling processing result; using the fully connected layer to perform mapping processing on the pooling processing result to obtain the probability that the image to be recognized contains the corresponding content attribute.

[0230] By comparing Figure 7 and Figure 8 it can be found that Figure 8 the exemplified task branch network increases the number of feature extraction layers, and increasing the number of feature extraction layers means increasing the network depth, which in turn means that the network can capture more hierarchical feature information. Each convolutional layer can extract different features in the image, and the deeper convolutional layers can further abstract higher-dimensional features based on the features of the shallower layers. This hierarchical feature extraction method helps the network better understand and recognize the targets in the image.

[0231] As an exemplary application scenario, in the handwriting recognition task, the handwriting usually occupies a small area in the image. This requires the network to be able to accurately capture these subtle features for accurate recognition. Therefore, for the handwriting recognition task, the Figure 8 exemplified task branch network structure can be adopted, and for the recognition of other content attributes such as watermarks, shadows, moiré patterns, etc., the Figure 7 exemplified task branch network structure can be adopted.

[0232] Figure 9 It is a schematic flowchart of a method for training an image recognition model provided by an embodiment of the present application. As Figure 9 shown, the image recognition model described in the above embodiment can be trained through the following steps:

[0233] Step 901, obtain the original text image set.

[0234] Step 902, perform the following processing on each original text image in the original text image set: add the content of at least one content attribute to the original text image to obtain the training sample image.

[0235] Step 903, input the training sample image into the initial model to obtain the model recognition result of the training sample image.

[0236] Step 904: Construct the loss functions for each task branch network in the initial model according to the model recognition results, and perform weighted summation processing on the loss functions of each task branch network to obtain a comprehensive loss function.

[0237] Step 905: Adjust the model parameters of the initial model according to the comprehensive loss function until the initial model converges to obtain an image recognition model.

[0238] For ease of understanding, the following provides a unified description of Steps 901 to 905:

[0239] The training of the deep processing model depends on a large number of data sets. For the image recognition model in the embodiments of the present application, its training depends on images with different content attributes (as positive samples) and images without any content attributes (as negative samples). Among them, the images with different content attributes can contain one or more content attributes at the same time. Accordingly, after obtaining the original text image set, the following processing is performed on each original text image in the original text image set: Add the content of at least one content attribute to the original text image to obtain a training sample image.

[0240] The original text image set includes multiple original text images, and the original text images here can be images without content attributes such as watermarks, shadows, moiré patterns, and handwriting. In one embodiment, multiple pictures can be obtained from an open-source picture library as the original text images to form the original text image set.

[0241] Among them, constructing a training sample image with a watermark:

[0242] Watermarks are divided into text watermarks and image watermarks. For text watermarks, multiple watermark words can be set, such as "Kingsoft Office", "Confidential", etc., and different text watermarks can be set according to different text fonts and colors. For the effect, please refer to Figure 10 the example in (a).

[0243] For image watermarks, a certain amount of icons or design element images with transparency can be obtained through image segmentation to form image watermarks. For the effect, please refer to Figure 10 the example in (b).

[0244] Perform image fusion on the text watermark or image watermark and the original text image. The parameter variables involved in the fusion include color, transparency, spacing, size, number, position, etc. For the synthesis schematic diagram, please refer to Figure 11 .

[0245] Constructing a training sample image with a shadow:

[0246] Construct a shadow illumination model. This model assumes a single main light source and ambient light as the illumination sources, and also assumes that the shadow area is only caused by occlusion by the main light source. The pixels in the non-shadow area and the pixels in the shadow area form a non-linear relationship as follows based on the model assumptions:

[0247] image_shadow = a * image_light - b

[0248] where image_light is the pixel in the non-shadow area, image_shadow is the pixel in the shadow area, and a and b represent the parameters of the shadow illumination model. These parameters depend on camera and scene attributes, such as surface material and lighting conditions.

[0249] Use the above shadow illumination model to simulate shadows. Here, there are differences between umbra and penumbra in the finally formed shadows. The specific formation range is determined by the occluding object. Here, a smoothed mask is used to represent the occluding object, and finally an image with shadows is formed. The related formula is as follows:

[0250] image_res = image_light * (1 - mark) + image_shadow * mark

[0251] where image_res is the final result image. For the shadow synthesis effects of the same original text image under different parameters, refer to the example in Figure 12 .

[0252] Construct a training sample image with moiré patterns:

[0253] In one embodiment, the original text image can be imported into a computer and displayed on a display screen, and then a photographing device, such as a mobile phone, is used to photograph the screen. Finally, a training sample image with moiré patterns can be obtained. In addition, in order to enable the model to better learn the characteristics of moiré patterns, the moiré patterns in the image with moiré patterns are removed to obtain a high-quality negative sample image without moiré patterns for model training. As shown in Figure 13 . It is an example of the scenario for constructing a training sample image with moiré patterns.

[0254] Construct a training sample image with handwriting:

[0255] In one embodiment, an image with handwriting can be obtained through image acquisition as the training sample image with handwriting. In addition, although a large number of images with and without handwriting can be obtained through image acquisition, in order to enable the model to better learn the characteristics of handwriting, the handwriting in the image with handwriting is removed to obtain a high-quality negative sample image without handwriting for model training. As shown in Figure 14 . It is an example of the scenario for constructing a training sample image with handwriting.

[0256] Furthermore, according to Figure 9 the description of the process shown, the initial model is iteratively trained using the training sample images. For each round of iterative training, the loss value of the model obtained in the current training is calculated. When the loss value does not meet the preset iterative termination condition, the next round of iterative training is continued until the loss value of the trained model meets the preset iterative termination condition, and the model obtained in the last round of training is determined as the trained image recognition model.

[0257] Among them, since this model belongs to a multi-task model, a comprehensive loss function is defined for this model, and this comprehensive function is obtained by weighted summation of the loss functions of each task. Among them, in order to balance the importance of each task in the training process and combine the difficulty of the tasks, corresponding loss weight coefficients are set for each task. For example, the corresponding loss weight coefficients of the four tasks of watermark recognition, shadow recognition, moiré recognition, and handwriting recognition are set in sequence as: 1, 0.7, 0.5, 0.7.

[0258] In addition, in practical applications, the loss weight coefficients of each task can also be dynamically and adaptively adjusted. For example, during the training process, the weight coefficients are dynamically adjusted according to the accuracy rate of the recognition results of each round of the model. Specifically: calculate the recognition accuracy rates of each task in sequence, then uniformly normalize the reciprocals of the accuracy rates to [0, 1], and then multiply by the initial weight coefficients to obtain the adjusted weight coefficients. The underlying logic of this method is that the higher the accuracy rate, the lower its weight coefficient.

[0259] Figure 9 According to the process shown, by adding content with at least one content attribute to the original text image to obtain the training sample image and using the training sample image for model training, the problem of being difficult to directly obtain training samples corresponding to various content attributes in practical applications can be solved. At the same time, by defining a comprehensive loss function for the model and using the comprehensive loss function to evaluate the performance of the model, the generalization ability of the model can be improved, enabling the model to more flexibly adapt to different image processing tasks, which helps to improve the overall performance of the model.

[0260] Figure 15 It is a block diagram of an embodiment of an image recognition device provided by an embodiment of the present application. As Figure 15 shown, the device includes:

[0261] An image acquisition module 1501, configured to acquire an image to be recognized;

[0262] The model recognition module 1502 is configured to extract the image features of the image to be recognized through the backbone network in the image recognition model, and input the image features into each task branch network in the image recognition model respectively. The task branch network determines the probability of the content attributes contained in the image to be recognized according to the image features;

[0263] The result determination module 1503 is configured to determine the target content attributes included in the image to be recognized according to the recognition results of each task branch network.

[0264] In a possible implementation manner, the result determination module 1503 is specifically configured to:

[0265] Compare the recognition result of the task branch network with a set threshold;

[0266] In the case where the recognition result is greater than or equal to the set threshold, determine the content attribute corresponding to the task branch network as one of the target content attributes included in the image to be recognized.

[0267] In a possible implementation manner, the backbone network includes a downsampling module, a hierarchical feature extraction module, and a feature integration module. The extraction of the image features of the image to be recognized by the backbone network in the image recognition model includes:

[0268] The downsampling module performs downsampling processing on the image to be recognized to obtain a downsampling result;

[0269] The hierarchical feature extraction module performs multi-level feature extraction on the downsampling result to obtain a multi-level feature extraction result;

[0270] The feature integration module performs feature integration processing on the multi-level feature extraction result to obtain the image features of the image to be recognized.

[0271] In a possible implementation manner, the hierarchical feature extraction module includes a shallow feature extraction module and a deep feature extraction module; the use of the hierarchical feature extraction module to perform multi-level feature extraction on the downsampling result to obtain a multi-level feature extraction result includes:

[0272] The shallow feature extraction module performs shallow feature extraction processing on the downsampling result to obtain a shallow feature extraction result;

[0273] The deep feature extraction module performs deep feature extraction processing on the shallow feature extraction result to obtain a multi-level feature extraction result.

[0274] In a possible implementation, the shallow feature extraction module includes a feature expansion and compression module and N second feature expansion and compression modules connected in sequence; using the shallow feature extraction module to perform shallow feature extraction processing on the downsampling result to obtain a shallow feature extraction result, including:

[0275] Using the feature expansion and compression module to perform feature expansion and compression processing on the downsampling result to obtain a feature expansion and compression processing result;

[0276] Using the first second feature expansion and compression module in the shallow feature extraction module to perform second feature expansion and compression processing on the feature expansion and compression processing result to obtain a first second feature expansion and compression processing result;

[0277] Using the i-th second feature expansion and compression module in the shallow feature extraction module to perform second feature expansion and compression processing on the (i - 1)-th second feature expansion and compression processing result to obtain the i-th second feature expansion and compression processing result, where i takes values one by one within the integer range of [2, N];

[0278] Taking the N-th second feature expansion and compression processing result obtained by the N-th second feature expansion and compression module in the shallow feature extraction module as the shallow feature extraction result.

[0279] In a possible implementation, the deep feature extraction module includes a first feature expansion and compression module, a second feature expansion and compression module, and a second feature expansion and compression module; using the deep feature extraction module to perform deep feature extraction processing on the shallow feature extraction result to obtain a multi-level feature extraction result, including:

[0280] Using the first feature expansion and compression module to perform feature expansion and compression processing on the shallow feature extraction result to obtain a first feature expansion and compression processing result;

[0281] Using the second feature expansion and compression module to perform second feature expansion and compression processing on the first feature expansion and compression processing result to obtain a second feature expansion and compression processing result;

[0282] Using the second feature expansion and compression module to perform feature expansion and compression processing on the second feature expansion and compression processing result to obtain a multi-level feature extraction result.

[0283] In a possible implementation manner, the feature expansion and compression module includes: a feature expansion and activation module, a deep feature extraction module, and a feature compression module; using the feature expansion and compression module to perform feature expansion and compression processing on the downsampling result to obtain a feature expansion and compression processing result, including:

[0284] Using the feature expansion and activation module to perform feature expansion and activation processing on the downsampling result to obtain a feature expansion and activation processing result;

[0285] Using the deep feature extraction module to perform deep feature extraction processing on the feature expansion and activation processing result to obtain a deep feature extraction result;

[0286] Using the feature compression module to perform feature compression processing on the deep feature extraction result to obtain a feature expansion and compression processing result.

[0287] In a possible implementation manner, the second feature expansion and compression module includes: a feature expansion and activation module, a deep feature extraction and residual connection module, and a feature compression module; using the first second feature expansion and compression module in the shallow feature extraction module to perform second feature expansion and compression processing on the feature expansion and compression processing result to obtain a first second feature expansion and compression processing result, including:

[0288] Using the feature expansion and activation module to perform feature expansion and activation processing on the feature expansion and compression processing result to obtain a feature expansion and activation processing result;

[0289] Using the deep feature extraction and residual connection module to perform deep feature extraction and residual connection processing on the feature expansion and activation processing result to obtain a deep feature extraction and residual connection processing result;

[0290] Using the feature compression module to perform feature compression processing on the deep feature extraction and residual connection processing result to obtain a second feature expansion and compression processing result.

[0291] In a possible implementation manner, the task branch network includes: a feature extraction layer, a global average pooling layer, and a fully connected layer; the task branch network identifying the probability that the to-be-identified image contains a corresponding content attribute according to the image feature, including:

[0292] Using the feature extraction layer to perform high-dimensional feature extraction processing on the image feature to obtain a high-dimensional feature extraction result;

[0293] Using the global average pooling layer to perform pooling processing on the high-dimensional feature extraction result to obtain a pooling processing result;

[0294] Use the fully connected layer to perform a mapping process on the pooling result to obtain the probability that the image to be recognized contains the corresponding content attribute.

[0295] In a possible implementation manner, the task branch network includes: M successively connected feature extraction layers, a global average pooling layer, and a fully connected layer; the task branch network recognizes the probability that the image to be recognized contains the corresponding content attribute according to the image feature, including:

[0296] Use the first feature extraction layer to perform a high-dimensional feature extraction process on the image feature to obtain the first high-dimensional feature extraction result;

[0297] Use the j-th feature extraction layer to perform a high-dimensional feature extraction process on the (j - 1)-th high-dimensional feature extraction result to obtain the j-th high-dimensional feature extraction result; j takes values one by one within the integer range of [2, M];

[0298] Use the global average pooling layer to perform a pooling process on the M-th high-dimensional feature extraction result to obtain a pooling result;

[0299] Use the fully connected layer to perform a mapping process on the pooling result to obtain the probability that the image to be recognized contains the corresponding content attribute.

[0300] In a possible implementation manner, the device further includes:

[0301] A model training module, configured to train an image recognition model through the following method:

[0302] Obtain a set of original text images;

[0303] For each original text image in the set of original text images, perform the following processing: Add the content of at least one content attribute to the original text image to obtain a training sample image;

[0304] Input the training sample image into the initial model to obtain the model recognition result of the training sample image;

[0305] Construct the loss function of each task branch network in the initial model according to the model recognition result, and perform a weighted summation process on the loss functions of the respective task branch networks to obtain a comprehensive loss function;

[0306] Adjust the model parameters of the initial model according to the comprehensive loss function until the initial model converges to obtain an image recognition model.

[0307] As Figure 16As shown in the figure, an embodiment of the present application provides an electronic device, including a processor 111, a communication interface 112, a memory 113, and a communication bus 114. Among them, the processor 111, the communication interface 112, and the memory 113 complete mutual communication through the communication bus 114.

[0308] The memory 113 is used to store computer programs.

[0309] In an embodiment of the present application, when the processor 111 is used to execute the program stored on the memory 113, it implements the image recognition method provided in any one of the foregoing method embodiments, including:

[0310] Obtain the image to be recognized.

[0311] Extract the image features of the image to be recognized through the backbone network in the image recognition model, and input the image features into each task branch network in the image recognition model respectively. The task branch network recognizes the probability of the content attributes contained in the image to be recognized according to the image features.

[0312] Determine the target content attributes included in the image to be recognized according to the recognition results of each task branch network.

[0313] An embodiment of the present application also provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it implements the steps of the image recognition method provided in any one of the foregoing method embodiments.

[0314] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place, or may be distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0315] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a general hardware platform, and of course, it can also be implemented by hardware. Based on this understanding, the essence of the above technical solution, or the part that contributes to the related technology, can be embodied in the form of a software product. The computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., including several instructions to enable a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.

[0316] It should be understood that the terms used herein are for the purpose of describing particular example embodiments only and are not intended to be limiting. Unless the context clearly dictates otherwise, the singular forms "a", "an", and "the" as used herein may also include the plural forms. The terms "comprising", "including", "containing", and "having" are inclusive and thus specify the presence of stated features, steps, operations, elements, and / or components, but do not preclude the presence or addition of one or more other features, steps, operations, elements, components, and / or combinations thereof. The method steps, processes, and operations described herein are not to be construed as necessarily requiring their performance in the particular order described or illustrated, unless an execution order is explicitly stated. It should also be understood that additional or alternative steps may be used.

[0317] The above are only specific embodiments of the present application, enabling those skilled in the art to understand or implement the present application. Various modifications to these embodiments will be obvious to those skilled in the art, and the general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of the present application. Therefore, the present application will not be limited to these embodiments shown herein, but rather to the broadest scope consistent with the principles and novel features claimed herein.

Claims

1. An image recognition method, characterized in that: The method comprises: Obtain an image to be recognized; Extracting image features of the image to be identified through the backbone network in the image recognition model, and inputting the image features into each task branch network in the image recognition model, so that the task branch network identifies the probability of content attributes contained in the image to be identified according to the image features; According to the recognition results of each of the task branch networks, the target content attributes contained in the image to be recognized are determined.

2. The method according to claim 1, characterized in that: Determining the target content attribute contained in the image to be identified according to the output results of each of the task branch networks includes: Comparing the recognition result of the task branch network with a set threshold; When the recognition result is greater than or equal to the set threshold, the content attribute corresponding to the task branch network is determined as one of the target content attributes contained in the image to be recognized.

3. The method according to claim 1, characterized in that: The backbone network includes a downsampling module, a hierarchical feature extraction module and a feature integration module. The image features of the image to be recognized are extracted through the backbone network in the image recognition model, including: The downsampling module performs downsampling processing on the image to be recognized to obtain a downsampling result; Using the hierarchical feature extraction module to perform multi-level feature extraction processing on the downsampling result to obtain a multi-level feature extraction result; Using the feature integration module to perform feature integration processing on the multi-level feature extraction results to obtain image features of the image to be identified; The hierarchical feature extraction module includes a shallow feature extraction module and a deep feature extraction module; the hierarchical feature extraction module is used to perform multi-level feature extraction processing on the downsampling result to obtain a multi-level feature extraction result, including: Using the shallow feature extraction module to perform shallow feature extraction processing on the downsampling result to obtain a shallow feature extraction result; The deep feature extraction module is used to perform deep feature extraction processing on the shallow feature extraction result to obtain a multi-level feature extraction result.

4. The method according to claim 3, characterized in that The shallow feature extraction module includes a first feature expansion and compression module and N second feature expansion and compression modules connected in sequence; the shallow feature extraction module is used to perform shallow feature extraction processing on the downsampling result to obtain a shallow feature extraction result, including: Using the first feature expansion and compression module to perform first feature expansion and compression processing on the downsampling result to obtain a first feature expansion and compression processing result; Using the first of the second feature expansion and compression modules in the shallow feature extraction module to perform second feature expansion and compression processing on the first feature expansion and compression processing result, to obtain a first second feature expansion and compression processing result; Using the i-th second feature expansion and compression module in the shallow feature extraction module to perform second feature expansion and compression processing on the i-1-th second feature expansion and compression processing result, to obtain the i-th second feature expansion and compression processing result, where i is taken one by one in the integer interval range of [2, N]; The Nth second feature expansion and compression processing result obtained by the Nth second feature expansion and compression module in the shallow feature extraction module is used as the shallow feature extraction result.

5. The method according to claim 3, characterized in that: The deep feature extraction module includes a first feature expansion and compression module, a second feature expansion and compression module, and a third feature expansion and compression module; the deep feature extraction module is used to perform deep feature extraction processing on the shallow feature extraction result to obtain a multi-level feature extraction result, including: Using the first feature expansion and compression module to perform first feature expansion and compression processing on the shallow feature extraction result to obtain a first feature expansion and compression processing result; Using the second feature expansion and compression module to perform a second feature expansion and compression process on the first feature expansion and compression processing result to obtain a second feature expansion and compression processing result; The third feature expansion and compression module is used to perform third feature expansion and compression processing on the second feature expansion and compression processing result to obtain a multi-level feature extraction result.

6. The method according to claim 4, characterized in that The first feature expansion and compression module includes: a feature expansion and activation module, a deep feature extraction module and a feature compression module; the first feature expansion and compression module is used to perform first feature expansion and compression processing on the downsampling result to obtain a first feature expansion and compression processing result, including: Using the feature expansion and activation module to perform feature expansion and activation processing on the downsampling result to obtain a feature expansion and activation processing result; Using the deep feature extraction module to perform deep feature extraction processing on the feature expansion and activation processing results to obtain deep feature extraction results; The feature compression module is used to perform feature compression processing on the deep feature extraction result to obtain a feature expansion and compression processing result.

7. The method according to claim 4, characterized in that The second feature expansion and compression module includes: a feature expansion and activation module, a deep feature extraction and residual connection module, and a feature compression module; the first second feature expansion and compression module in the shallow feature extraction module is used to perform a second feature expansion and compression process on the feature expansion and compression processing result to obtain a first second feature expansion and compression processing result, including: Using the feature expansion and activation module to perform feature expansion and activation processing on the feature expansion and compression processing result to obtain a feature expansion and activation processing result; Using the deep feature extraction and residual connection module to perform deep feature extraction and residual connection processing on the feature expansion and activation processing results, to obtain deep feature extraction and residual connection processing results; The feature compression module is used to perform feature compression processing on the deep feature extraction and residual connection processing results to obtain a second feature expansion and compression processing result.

8. The method according to claim 1, characterized in that The task branch network includes: a feature extraction layer, a global average pooling layer and a fully connected layer; the task branch network identifies the probability that the image to be identified contains the corresponding content attribute according to the image features, including: Using the feature extraction layer to perform high-dimensional feature extraction processing on the image features to obtain a high-dimensional feature extraction result; Using the global average pooling layer to perform pooling processing on the high-dimensional feature extraction result to obtain a pooling processing result; The pooling processing result is mapped using the fully connected layer to obtain the probability that the image to be identified contains the corresponding content attribute.

9. The method according to claim 1, characterized in that: The task branch network includes: M sequentially connected feature extraction layers, global average pooling layers, and fully connected layers; the task branch network identifies the probability that the image to be identified contains corresponding content attributes according to the image features, including: Using the first feature extraction layer to perform high-dimensional feature extraction processing on the image features to obtain a first high-dimensional feature extraction result; Using the j-th feature extraction layer to perform high-dimensional feature extraction processing on the j-1-th high-dimensional feature extraction result, to obtain the j-th high-dimensional feature extraction result; the j is taken one by one in the integer interval range of [2, M]; Using the global average pooling layer to perform pooling processing on the M-th high-dimensional feature extraction result to obtain a pooling processing result; The pooling processing result is mapped using the fully connected layer to obtain the probability that the image to be identified contains the corresponding content attribute.

10. The method according to any one of claims 1 to 9, characterized in that: The image recognition model is trained in the following way: Get the original text image set; The following processing is performed for each original text image in the original text image set: adding at least one content attribute to the original text image to obtain a training sample image; Inputting the training sample image into the initial model to obtain a model recognition result of the training sample image; Constructing the loss function of each task branch network in the initial model according to the model recognition result, and performing weighted summation processing on the loss functions of each task branch network to obtain a comprehensive loss function; The model parameters of the initial model are adjusted according to the comprehensive loss function until the initial model converges to obtain an image recognition model.

11. An image recognition device, characterized in that: The device comprises: An image acquisition module, used for acquiring an image to be identified; A model recognition module, used to extract image features of the image to be recognized through the backbone network in the image recognition model, and input the image features into each task branch network in the image recognition model, so that the task branch network recognizes the probability of the content attribute contained in the image to be recognized according to the image features; The result determination module is used to determine the target content attribute contained in the image to be identified according to the identification results of each of the task branch networks.

12. An electronic device, characterized in that: include: A processor and a memory, wherein the processor is used to execute an image recognition program stored in the memory to implement the image recognition method according to any one of claims 1 to 10.