Image recognition method, apparatus and device, storage medium, and product

Through the graphic and text interaction unit and coding unit of the image content recognition model, combined with the target audit rules, the problem of image recognition effect and inefficiency in the prior art is solved, and automatic detection and efficient recognition of metaphorical information in the image is realized.

WO2025175985A1PCT designated stage Publication Date: 2025-08-28BIGO TECH PTE LTD +1

Patent Information

Application Number
PCT/CN2025/073071
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-02-21
Filing Date
2025-01-17
Publication Date
2025-08-28

AI Technical Summary

Technical Problem

Existing image recognition technology cannot effectively detect metaphorical information in images, resulting in low recognition effect and efficiency, especially when subjective judgments are included, relying on manual review is inefficient.

Method used

The image content recognition model is used to analyze the target text features and image features through the graphic and text interaction unit, and the content recognition results are generated by the text encoding unit and the image encoding unit, and combined with the target audit rules, the metaphorical information in the image is accurately detected.

Benefits of technology

It improves the accuracy and efficiency of image recognition, can automatically detect metaphorical information in the image, reduce manual review needs, and adapt to actual business needs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2025073071_28082025_PF_FP_ABST
    Figure CN2025073071_28082025_PF_FP_ABST
Patent Text Reader

Abstract

Embodiments of the present application provide an image recognition method, apparatus and device, a storage medium, and a product. According to the technical solution provided in the embodiments of the present application, an image content recognition model uses an image-text interaction unit to perform analysis processing on a target text feature and a target image feature of an image to be recognized, to obtain a content recognition result of said image; moreover, the target text feature is generated on the basis of a set target review rule by means of a text coding unit in the image content recognition model, and the target image feature is generated on the basis of said image by means of an image coding unit in the image content recognition model. Image-text interaction can be effectively performed on the basis of the target text feature of the target review rule and the target image feature of said image, thereby accurately detecting metaphorical information in the image, and effectively improving the image recognition effect and efficiency.
Need to check novelty before this filing date? Find Prior Art

Description

Image recognition method, device, equipment, storage medium and product

[0001] This application claims priority to the Chinese patent application filed with the China Patent Office on February 21, 2024, with application number 202410193153.6, the entire contents of which are incorporated by reference into this application. Technical Field

[0002] The embodiments of the present application relate to the field of image processing technology, and in particular to an image recognition method, apparatus, device, storage medium, and product. Background Art

[0003] With the development of computer and internet technologies, information dissemination has become increasingly diverse, with platforms such as photo sharing and short video platforms. The review of image content is becoming increasingly important. To improve the quality and accuracy of image review, the current review of image content often requires subjective judgment.

[0004] Images containing subjective judgments refer to images that contain metaphorical information. Metaphorical information cannot generally be directly derived from the people, objects, or activities in the image. Currently, image content recognition generally identifies the category of objects in the image and cannot detect metaphorical information. Image content recognition containing subjective judgments is generally performed manually, resulting in low image recognition effectiveness and efficiency. Summary of the Invention

[0005] The embodiments of the present application provide an image recognition method, apparatus, device, storage medium and product to solve the technical problems of low image recognition effect and efficiency in related technologies, and effectively improve image recognition effect and efficiency.

[0006] In a first aspect, an embodiment of the present application provides an image recognition method, comprising:

[0007] Obtain the image to be recognized;

[0008] The image to be identified is input into a trained image content recognition model, and the image content recognition model uses a text-image interaction unit to analyze and process the target text features and the target image features of the image to be identified to obtain a content recognition result of the image to be identified. The target text features are generated by the text encoding unit in the image content recognition model according to the set target review rules, and the target image features are generated by the image encoding unit in the image content recognition model according to the image to be identified.

[0009] In a second aspect, an embodiment of the present application provides an image recognition device, including an image acquisition module and an image recognition module, wherein:

[0010] The image acquisition module is configured to acquire an image to be identified;

[0011] The image recognition module is configured to input the image to be recognized into a trained image content recognition model, and use the image content recognition model to analyze and process the target text features and the target image features of the image to be recognized using the image-text interaction unit to obtain the content recognition result of the image to be recognized. The target text features are generated by the text encoding unit in the image content recognition model according to the set target review rules, and the target image features are generated by the image encoding unit in the image content recognition model according to the image to be recognized.

[0012] In a third aspect, an embodiment of the present application provides an image recognition device, including: a memory and one or more processors;

[0013] The memory is used to store one or more programs;

[0014] When the one or more programs are executed by the one or more processors, the one or more processors implement the image recognition method as described in the first aspect.

[0015] In a fourth aspect, an embodiment of the present application provides a non-volatile storage medium storing computer-executable instructions, which, when executed by a computer processor, are used to perform the image recognition method as described in the first aspect.

[0016] In the fifth aspect, an embodiment of the present application provides a computer program product, which includes a computer program stored in a computer-readable storage medium. At least one processor of the device reads and executes the computer program from the computer-readable storage medium, so that the device performs the image recognition method described in the first aspect.

[0017] The embodiment of the present application uses an image-text interaction unit to analyze and process target text features and target image features of the image to be identified through an image content recognition model to obtain a content recognition result of the image to be identified, and the target text features are generated by a text encoding unit in the image content recognition model according to a set target review rule, and the target image features are generated by an image encoding unit in the image content recognition model according to the image to be identified. The target text interaction can be effectively performed based on the target text features of the target review rule and the target image features of the image to be identified, the metaphorical information in the image can be accurately detected, and the image recognition effect and efficiency can be effectively improved. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] FIG1 is a flow chart of an image recognition method provided by an embodiment of the present application;

[0019] FIG2 is a schematic diagram of the structure of an image content recognition model provided in an embodiment of the present application;

[0020] FIG3 is a schematic diagram of a training process of an image content recognition model provided in an embodiment of the present application;

[0021] FIG4 is a schematic diagram of a model structure of a first training stage provided in an embodiment of the present application;

[0022] FIG5 is a schematic diagram of a model structure of a second training stage provided in an embodiment of the present application;

[0023] FIG6 is a schematic structural diagram of an image recognition device provided in an embodiment of the present application;

[0024] FIG7 is a schematic structural diagram of an image recognition device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0025] In order to make the purpose, technical solutions and advantages of the present application clearer, the specific embodiments of the present application are further described in detail below in conjunction with the accompanying drawings. It is understood that the specific embodiments described herein are only used to explain the present application and are not intended to limit the present application. It should also be noted that, for ease of description, only some, but not all, of the contents related to the present application are shown in the accompanying drawings. Before discussing the exemplary embodiments in more detail, it should be mentioned that some exemplary embodiments are described as processes or methods depicted as flow charts. Although the flow charts describe each operation (or step) as a sequential process, many of the operations therein can be implemented in parallel, concurrently or simultaneously. In addition, the order of the operations can be rearranged. The above process can be terminated when its operation is completed, but it can also have additional steps not included in the accompanying drawings. The above process can correspond to a method, function, procedure, subroutine, subprogram, etc.

[0026] The image recognition method provided in this application can be applied to image content recognition, review, etc., and aims to perform image-text interaction based on the target text features of the target review rules and the target image features of the image to be recognized, accurately detect metaphorical information in the image, and effectively improve the image recognition effect and efficiency.

[0027] Existing image recognition solutions generally rely on models to identify the content contained in an image and output corresponding classification labels. However, these solutions are unable to effectively detect metaphorical information within images and struggle to effectively identify implicit information within video or text-based information streams, resulting in poor image recognition performance. Current image recognition technology frameworks generally consist of large models related to CV classification detection, CV retrieval, and multimodal recognition involving image-text matching. For example, when recognizing images based on CV classification detection, CV classification detection primarily identifies the object and cannot detect metaphorical information. While classification can be performed by extracting image features, it cannot achieve the desired generalization capabilities with a small number of images. CV retrieval extracts features from hundreds of accumulated image data and stores them in a feature library. While it can process recurring images identical to those previously identified, it has little ability to identify data not in the feature library. Even slightly modified original images are considered acceptable, and it is unable to accurately identify new images containing metaphorical information. Image-text retrieval based on CLIP and BLIP can recall images for thousands of specific object categories using predefined descriptions, but it cannot recall images containing metaphors. This is because they are trained by matching one description to one image, focusing on specific object categories and categorizing images. Recall is only possible when a sufficiently detailed description of a specific image is provided, and similarly, they lack generalizability. Large multimodal models based on LLaMa can conduct multiple rounds of image-text conversations, gradually focusing on the desired subcategories through multiple rounds of questioning and answering in a gradient hierarchy. In actual testing, after multiple rounds of questioning and answering, some images can be recalled, but false positives are also quite common. This not only tests the questioner's skills, but also requires large-scale computing resources and human reviewer interaction, which is not suitable for real-world business needs. Even with the addition of significant computing resources and fine-tuning the large model using artificial conversation data, so that the multimodal model only outputs binary yes / no judgments, it is still difficult to balance recall and false positives with real-world business data. Commercial multimodal applications based on GPT-4 directly invoke security shielding policies based on audit requirements, without returning recognition results. Similar to LLaMa, returning results requires multiple rounds of dialogue and interaction with the auditor, which does not meet actual business needs. Based on this, an image recognition method according to an embodiment of the present application is provided to address the technical issues of low image recognition effectiveness and efficiency.

[0028] Figure 1 shows a flowchart of an image recognition method provided in an embodiment of the present application. The image recognition method provided in an embodiment of the present application can be executed by an image recognition device, which can be implemented by hardware and / or software and integrated into an image recognition device.

[0029] The following description is based on an example of an image recognition method performed by an image recognition device. Referring to FIG1 , the image recognition method includes:

[0030] S110: Acquire an image to be recognized.

[0031] For example, when image recognition is required for short videos, graphic and text information streams, etc., the corresponding image can be obtained as the image to be recognized. For example, when image recognition is required based on set target review rules to identify subjective or metaphorical content contained in the image, the corresponding image can be obtained as the image to be recognized.

[0032] Images or videos containing subjective or metaphorical information refer to those containing metaphorical information. Metaphorical information cannot generally be directly derived from the people, objects, or events in the image. For example, an image of two people fighting, with flags of different countries attached to their heads, suggests a political message, with one person winning and the other losing. These images require image recognition based on the subjective and metaphorical content contained in the image to effectively review them.

[0033] S120: Input the image to be recognized into the trained image content recognition model, and use the image content recognition model to analyze and process the target text features and the target image features of the image to be recognized using the image-text interaction unit to obtain a content recognition result of the image to be recognized.

[0034] Among them, the target text features are generated by the text encoding unit in the image content recognition model according to the set target review rules, and the target image features are generated by the image encoding unit in the image content recognition model according to the image to be recognized.

[0035] As shown in the structural diagram of an image content recognition model provided in FIG2 , the image content recognition model provided by this solution includes a text-image interaction unit, a text encoding unit, and an image encoding unit. Among them, the text encoding unit is configured to identify the text features of the input text and submit it to the text-image interaction unit. This solution can input the audit rule text used to guide the recognition of subjective and implicit content in the image into the text encoding unit to obtain a text feature vector that records the features corresponding to the audit rule, that is, encode the audit rule text into a text feature vector (Embedding), and use it as the query vector (Query) of the text-image interaction unit. The image encoding unit is configured to extract the image features of the input image and input it into the text-image interaction unit, that is, encode the image into an image feature vector (Embedding), and use it as the input of the attention mechanism in the text-image interaction unit. The text-image interaction unit based on the attention mechanism is configured to perform text-image interaction analysis based on text features and image features, match and align the input text feature vector and image feature vector, and generate corresponding recognition results. For example, the text corresponding to the review rules could be "when there are marches, demonstrations, illegal gatherings, or signs that are anti-party, anti-government, or attacking local governments or political parties; when there is content related to the dissemination and incitement of violence, including speech mentioning violence, violent information appearing in the picture, etc.; when there are politically sensitive, reactionary social events, as well as their pronouns, symbols, obscure words, etc., the image is considered to contain political content involving reactionary behavior."

[0036] Exemplarily, the image to be identified obtained as above is input into the trained image content recognition model. After receiving the image to be identified, the image content classification model uses the image encoding unit to obtain the target graphic features corresponding to the image to be identified, and uses the image content recognition model to analyze and process the target text features and the target image features of the image to be identified using the image-text interaction unit to obtain the content recognition result of the image to be identified. Among them, the target text features are generated by the text encoding unit in the image content recognition model according to the set target review rules. The target text features can be generated in real time according to the target review rules, or they can be text features that were previously generated and reused according to the target review rules.

[0037] Among them, when different images are recognized and processed according to the same review rule, the text feature vector of the feature corresponding to the review rule only needs to be obtained once at the beginning, that is, when the image content is recognized for the first time based on the set target review rule, or when a new target review rule needs to be replaced, the text corresponding to the target review rule is input. After the text encoding unit extracts the text feature vector corresponding to the target review rule, before receiving the new target review rule, the text feature vector is submitted to the image-text interaction unit as the query vector of the image-text interaction unit.

[0038] In one embodiment, the content recognition result provided by this solution can be used to indicate the probability or degree that the image to be recognized conforms to the content described by the target audit rules. Optionally, the content recognition result can be represented by a probability, or it can be a mapping result mapped to a set range, for example, the probability or degree that the image to be recognized conforms to the content described by the target audit rules is mapped to the range of [0, 1], and as the content recognition result increases, the probability or degree that the image to be recognized conforms to the content described by the target audit rules increases.

[0039] In one possible embodiment, as shown in FIG3 , a training process diagram of an image content recognition model is provided. The training process of the image content recognition model provided by this solution includes:

[0040] S101: Fix the network parameters of the image encoding unit in the image content recognition model, and train the text encoding unit in the image content recognition model based on a first training sample set, where the first training sample set includes a first review rule, a first image sample, and a first annotation of the first image sample under the first review rule.

[0041] In one embodiment, this solution can train the image content recognition model through training samples that have been pre-reviewed and labeled in the business. For example, the total amount of audit rules and tens of millions of image data accumulated by various short video and graphic information flow related businesses can reach hundreds of copies, and the training samples of the image content recognition model can be obtained based on the audit and labeling data of these businesses. Among them, the training samples include image samples, the audit rules corresponding to the image samples, and the annotations corresponding to the audit rules (for example, a description indicating whether the image sample complies with the audit rules, when the annotation is 0, it indicates that the image sample does not comply with the description of the audit rules (that is, the training sample is a negative sample), and when the annotation is 1, it indicates that the image sample complies with the description of the audit rules (that is, the training sample is a positive sample)).

[0042] Optionally, during the training phase of the image content recognition model, the image encoding unit and text encoding unit provided by this solution can be trained encoding units with feature extraction and induction capabilities. Optionally, the image encoding unit and text encoding unit can be from trained image recognition and text processing models, or from open source models. This can fully leverage the machine-generated rules and image data of pure CV models, reducing the need to screen and clean large amounts of crawled or artificially generated, low-quality data. For example, the ViT-L / 14 module from the CLIP model can be used as the image encoding unit of this solution, and the text encoder module from the LLaMa model can be used as the text encoding unit of this solution. Using portions of these two open source projects, the CLIP and LLaMa models, as initial parameters for training optimization, can ensure that the image and text encoding units have strong feature extraction and induction capabilities, improving the training efficiency and quality of the image content recognition model.

[0043] This solution can perform multi-stage (Multi Stage) training on the image content recognition model so that the image content recognition model can effectively converge under limited data and computing resources, and effectively improve the model training efficiency while ensuring the quality of model training. For example, the training of the image content recognition model can be divided into two training stages: training the text encoding unit and training the image-text interaction unit. Optionally, a part of the training samples (for example, 20%) can be used as the first training sample set for training the text encoding unit (including multiple first training samples, the first training sample includes the first review rule, the first image sample and the first annotation of the first image sample under the first review rule), and the remaining part (for example, 80%) can be used as the second training sample set for training the image-text interaction unit (including multiple second training samples, the second training sample includes the second review rule, the second image sample and the second annotation of the second image sample under the second review rule).

[0044] Exemplarily, the network parameters of the image encoding unit in the image content recognition model are fixed, and the text encoding unit in the image content recognition model is trained using the first training sample set until the set training rounds are reached or the network parameters of the image encoding unit converge. For example, the network parameters of the text encoding unit are fine-tuned based on the similarity between the feature vector extracted by the text encoding unit for the first audit rule in the training sample and the feature vector extracted by the image encoding unit for the first image sample, gradually increasing the similarity between the feature vector extracted by the text encoding unit for the first audit rule in the training sample and the feature vector extracted by the image encoding unit for the first image sample, thereby enhancing the text encoding unit's understanding of the audit rules of the business end, and making the generated language embedding have better inductive and generalization performance.

[0045] In one possible embodiment, the image recognition method provided by this solution includes, when training the text encoding unit in the image content recognition model based on the first training sample set:

[0046] S1011: Determine the first text sample feature of the first audit rule in the first training sample set through the text encoding unit in the image content recognition model.

[0047] S1012: Determine a first image sample feature of a first image sample in a first training sample set by an image encoding unit in an image content recognition model.

[0048] S1013: Determine a text-image sample contrast loss according to the first text sample feature and the first image sample feature, and perform optimization processing on the text encoding unit based on the text-image sample contrast loss.

[0049] Exemplarily, in the first training stage of the image content recognition model, the network parameters of the image coding unit are fixed, and the image content recognition model is trained using the first training sample set. As shown in a schematic diagram of the model structure of the first training stage provided in FIG4 , in the first training stage, training is mainly performed based on the image coding unit and the text coding unit in the image content recognition model, and the network parameters of the text coding unit are optimized. For each first training sample, the first text sample feature (text embedding) of the first audit rule in the first training sample set is extracted by the text coding unit that has completed the above training, and the first image sample feature (image embedding) of the first image sample in the first training sample set is extracted by the image coding unit.

[0050] Based on the set loss function calculation method, the image-text sample contrast loss (image-text contrast loss) between the first text sample feature and the first image sample feature is determined, and the text encoding unit is tuned based on the image-text sample contrast loss to update the network parameters of the text encoding unit. Optionally, the image-text sample contrast loss for learning the sample feature embedding of the image and text can use the image-text contrast loss (ITC, Image-Text Contrastive Loss) to make the matching image-text feature pairs have a high similarity and maximize the interaction of image-text information. In one embodiment, the learning rate of the first training stage can be set to be smaller than the learning rate of the second training stage to retain the original generalization ability of the text encoding unit. This solution can set the learning rate of the first training stage to 1e-6 and attenuate the learning rate based on the cosine annealing algorithm.

[0051] This solution determines the first text sample feature of the first audit rule in the first training sample set through a text encoding unit, and determines the first image sample feature of the first image sample in the first training sample set through an image encoding unit, and optimizes the text encoding unit according to the image-text sample contrast loss of the first text sample feature and the first image sample feature, thereby enhancing the text encoding unit's understanding of the audit rules of the business end, making the generated text features have better inductive and generalization performance, and improving the image content recognition effect.

[0052] In one embodiment, when using the first training sample set for the first training stage of training, and / or using the second training sample set for the second training stage of training, each first image sample and / or second image sample can be randomly enhanced based on the set image enhancement method (for example, enhancement processing of chroma, brightness, contrast, Gaussian noise, etc.), and the enhanced image sample after the enhancement processing is used as a new image sample, and the review rules and annotations of the original image sample are corresponding, thereby effectively increasing the richness of the sample.

[0053] S102: Fix the network parameters of the image encoding unit and the text encoding unit, and train the image-text interaction unit in the image content recognition model based on the second training sample set, where the second training sample set includes a second review rule, a second image sample, and a second annotation of the second image sample under the second review rule.

[0054] Exemplarily, after completing the training of the text encoding unit, the network parameters of the image encoding unit and the text encoding unit are fixed, and the image-text interaction unit in the image content recognition model is trained based on the second training sample set until the set training rounds are reached or the network parameters of the image-text interaction unit converge. For example, the text encoding unit extracts the text sample features of the second audit rule and sends them to the image-text interaction unit, the image encoding unit extracts the image sample features of the second image sample and sends them to the image-text interaction unit, the image-text interaction unit outputs a prediction result based on the text sample features and the image sample features (the prediction result may indicate the probability or degree that the second image sample meets the description of the second audit rule), and adjusts the network parameters of the image-text interaction unit based on the comparison between the prediction result and the corresponding second label, so that the image-text interaction unit has the ability to match and align image features and text features.

[0055] This solution trains the image content recognition model in two stages. In the first training stage, the network parameters of the image coding unit in the image content recognition model are fixed, and the text coding unit in the image content recognition model is trained based on the first training sample set, so as to enhance the text coding unit's understanding of the business-side review rules and make the generated text feature vector have better inductive and generalization performance. In the second training stage, the network parameters of the image coding unit and the text coding unit are fixed, and the image-text interaction unit in the image content recognition model is trained based on the second training sample set. The second training sample set includes the second review rule, the second image sample and the second annotation of the second image sample under the second review rule, so that the image-text interaction unit has a strong ability to match and align image features and text features, and the image content recognition model can accurately detect metaphorical information in the image.

[0056] In one possible embodiment, the image recognition method provided by this solution includes, when training the image-text interaction unit in the image content recognition model based on the second training sample set:

[0057] S1021: Determine the second text sample feature of the second review rule in the second training sample set through the text encoding unit in the image content recognition model.

[0058] S1022: Determine a second image sample feature of a second image sample in a second training sample set by an image encoding unit in the image content recognition model.

[0059] S1023: Input the second text sample features and the second image sample features into the text-image interaction unit in the image content recognition model, analyze and process the second text sample features and the second image sample features through the text-image interaction unit to obtain a sample prediction result, and tune the text-image interaction unit according to the sample prediction result and the second annotation in the second training sample set.

[0060] Exemplarily, in the second training phase of the image content recognition model, the network parameters of the image encoding unit and the text encoding unit are fixed, and the image content recognition model is trained using the second training sample set. As shown in Figure 5, a schematic diagram of the model structure of the second training phase is provided. In the second training phase, training is performed based on the image encoding unit, text encoding unit, and image-text interaction unit in the image content recognition model, and the network parameters of the image-text interaction unit are optimized.

[0061] In one embodiment, for each first training sample, a text encoding unit is used to determine a second text sample feature of a second audit rule in a second training sample set, and an image encoding unit is used to determine a second image sample feature of a second image sample in a second training sample set. The second text sample feature and the second image sample feature determined above are input into a text-to-image interaction unit, and the second text sample feature and the second image sample feature are analyzed and processed by the text-to-image interaction unit to obtain a sample prediction result. The text-to-image interaction unit is tuned based on the sample prediction result and the second annotation in the second training sample set (e.g., a loss function of the sample prediction result and the second annotation) to update the network parameters of the text-to-image interaction unit until the text-to-image interaction unit reaches a set number of training rounds and / or converges. This solution determines the second text sample features of the second review rule through a text encoding unit, and determines the second image sample features of the second image sample through an image encoding unit, and analyzes and processes the second text sample features and the second image sample features through a text-image interaction unit to obtain a sample prediction result, and optimizes the text-image interaction unit according to the sample prediction result and the second annotation in the second training sample set, so that the text-image interaction unit can effectively perform text-image interaction based on the target text features of the target review rule and the target image features of the image to be identified, accurately detect metaphorical information in the image, and effectively improve image recognition effect and efficiency.

[0062] In one possible embodiment, when the image recognition method provided by the present solution analyzes and processes the second text sample features and the second image sample features through the image-text interaction unit to obtain a sample prediction result, it can be: inputting the second text sample features into the cross-attention subunit of the image-text interaction unit, inputting the second image sample features into the self-attention subunit of the image-text interaction unit, determining the sample key vector and the sample value vector according to the second image sample features through the self-attention subunit, and determining the image-text sample interaction result according to the second text sample features, the sample key vector and the sample value vector through the cross-attention layer; and analyzing and processing the image-text sample interaction result through the feedforward subunit, the fully connected subunit and the logistic regression subunit in the image content recognition model to obtain a sample prediction result.

[0063] The image-text interaction unit provided by this solution includes multiple layers of serially connected attention subunits (corresponding to N layers of attention subunits in Figure 5, for example, N can be 3), a post-processing unit, and a logistic regression subunit. Each attention subunit includes a self-attention subunit and a cross-attention subunit. In adjacent attention subunits, the cross-attention subunit of the previous attention subunit is linked to the self-attention subunit of the next attention subunit. The self-attention subunit of the first attention subunit can receive the output of the image encoding unit, and the self-attention subunit of each subsequent attention subunit can receive the output of the cross-attention subunit of the previous attention subunit. Each cross-attention subunit can receive the output of the text encoding unit. The post-processing unit can be configured as a serial combination of a feedforward subunit and a fully connected subunit, with the feedforward subunit connected to the cross-attention subunit of the last attention subunit, and the fully connected subunit connected to the logistic regression subunit. The post-processing unit can also be configured as a large language model, which can be used to realize image content recognition based on multi-round dialogue interaction.

[0064] Exemplarily, the second text sample feature is input into each cross-attention sub-unit of the image-text interaction unit, and the second image sample feature is input into the first self-attention sub-unit of the image-text interaction unit. The sample key vector (key) and sample value vector (value) are determined according to the second image sample feature through the self-attention sub-unit, and the image-text sample interaction result is determined according to the second text sample feature, the sample key vector and the sample value vector through the cross-attention layer.

[0065] In one embodiment, after obtaining the image-text sample interaction results, the image-text sample interaction results are analyzed and processed in sequence by the feedforward subunit, fully connected subunit, and logistic regression subunit in the image content recognition model to obtain sample prediction results. The feedforward subunit and fully connected subunit serve as neural network layers, which can further fuse the image-text sample interaction results and increase parameters to learn a better final result. The logistic regression subunit can map the sample prediction results to the interval [0, 1]. Optionally, the feedforward subunit can use a multi-layer perceptron to further enhance the learning of the image-text matching loss. In this solution, the image-text matching loss can use the Image-Text Matching Loss to learn the binary classification of fine-grained matching of images and texts. This solution uses the self-attention subunit to determine the sample key vector and sample value vector based on the second image sample features, and uses the cross-attention layer to determine the image-text sample interaction results based on the second text sample features, sample key vector, and sample value vector. The feedforward subunit, fully connected subunit, and logistic regression subunit in the image content recognition model analyze and process the image-text sample interaction results to accurately obtain sample prediction results, thereby improving the training quality of the image content recognition model.

[0066] In a possible embodiment, the image recognition method provided by this solution can also train the image-text interaction unit and the image encoding unit in the image content recognition model based on the second training sample set after fixing the network parameters of the image encoding unit and the text encoding unit and training the image-text interaction unit in the image content recognition model based on the second training sample set.

[0067] Exemplarily, after the training of the image-text interaction unit converges, the fixed network parameters of the image encoding unit are canceled, and the image-text interaction unit and the image encoding unit in the image content recognition model are trained using the second training sample set. This solution trains the image-text interaction unit and the image encoding unit based on the second training sample set after the training of the image-text interaction unit converges, making the processing capabilities of the image encoding unit more suitable for business scenarios and improving image recognition accuracy.

[0068] In one possible embodiment, a first learning rate for training the image-text interaction unit in the image content recognition model based on the second training sample set is greater than a second learning rate for training the image-text interaction unit and the image coding unit in the image content recognition model based on the second training sample set. The parameters corresponding to the training of the image-text interaction unit in the image content recognition model based on the second training sample set are all randomly initialized, and the first learning rate can be set to a larger value, for example, the first learning rate can be 1e-4. The training of the image-text interaction unit and the image coding unit in the image content recognition model based on the second training sample set needs to reduce the corresponding second learning rate to reduce overfitting of the image coding unit, for example, the second learning rate can be set to 1e-6.

[0069] In the above, the target text features and the target image features of the image to be identified are analyzed and processed by the image content recognition model using the text-image interaction unit to obtain the content recognition result of the image to be identified, and the target text features are generated by the text encoding unit in the image content recognition model according to the set target review rules, and the target image features are generated by the image encoding unit in the image content recognition model according to the image to be identified. The text-image interaction can be effectively performed based on the target text features of the target review rules and the target image features of the image to be identified, accurately detecting the metaphorical information in the image, and effectively improving the image recognition effect and efficiency. By directly modifying the machine review rules that can be understood by humans as the target review rules, the control update of the model capabilities can be achieved, reducing the need to retrain the model every time the machine review rules change, and improving the efficiency of image content recognition. In addition, for images containing subjective and metaphorical content, by adjusting the text of the target review rules corresponding to the abstract machine review rules to cover them, the image content recognition model can identify metaphorical image content that complies with the rules but has not appeared before, thereby improving the accuracy and generalization ability of image recognition, and realizing the recognition of subjective and metaphorical image content under the end-to-end reasoning structure, balancing the relationship between recall rate, recommendation accuracy, human review or computing resources, and helping to meet actual business implementation needs.

[0070] FIG6 is a schematic diagram of the structure of an image recognition device provided in an embodiment of the present application. Referring to FIG6 , the image recognition device includes an image acquisition module 61 and an image recognition module 62 .

[0071] Among them, the image acquisition module 61 is configured to acquire the image to be identified; the image recognition module 62 is configured to input the image to be identified into the trained image content recognition model, and use the image content recognition model to analyze and process the target text features and the target image features of the image to be identified using the image-text interaction unit to obtain the content recognition result of the image to be identified. The target text features are generated by the text encoding unit in the image content recognition model according to the set target review rules, and the target image features are generated by the image encoding unit in the image content recognition model according to the image to be identified.

[0072] In the above, the target text features and the target image features of the image to be identified are analyzed and processed by the image content recognition model using the image-text interaction unit to obtain the content recognition result of the image to be identified, and the target text features are generated by the text encoding unit in the image content recognition model according to the set target review rules, and the target image features are generated by the image encoding unit in the image content recognition model according to the image to be identified. The target text interaction can be effectively performed based on the target text features of the target review rules and the target image features of the image to be identified, and the metaphorical information in the image can be accurately detected, thereby effectively improving the image recognition effect and efficiency.

[0073] In a possible embodiment, the image recognition device further includes a model training module, which is configured to train the image content recognition model. The training process of the image content recognition model by the model training module is configured as follows:

[0074] Fixing network parameters of an image encoding unit in the image content recognition model, and training a text encoding unit in the image content recognition model based on a first training sample set, wherein the first training sample set includes a first review rule, a first image sample, and a first annotation of the first image sample under the first review rule;

[0075] The network parameters of the image encoding unit and the text encoding unit are fixed, and the image-text interaction unit in the image content recognition model is trained based on the second training sample set. The second training sample set includes a second review rule, a second image sample, and a second annotation of the second image sample under the second review rule.

[0076] In one possible embodiment, the model training module trains the text encoding unit in the image content recognition model based on the first training sample set, and is configured as follows:

[0077] Determine a first text sample feature of a first review rule in a first training sample set by a text encoding unit in an image content recognition model;

[0078] Determining, by an image encoding unit in an image content recognition model, a first image sample feature of a first image sample in a first training sample set;

[0079] The image-text sample contrast loss is determined according to the first text sample feature and the first image sample feature, and the text encoding unit is tuned based on the image-text sample contrast loss.

[0080] In a possible embodiment, the model training module trains the image-text interaction unit in the image content recognition model based on the second training sample set, and is configured as follows:

[0081] Determine, by a text encoding unit in the image content recognition model, a second text sample feature of a second review rule in a second training sample set;

[0082] determining, by an image encoding unit in the image content recognition model, a second image sample feature of a second image sample in a second training sample set;

[0083] The second text sample features and the second image sample features are input into the text-image interaction unit in the image content recognition model, the second text sample features and the second image sample features are analyzed and processed by the text-image interaction unit to obtain a sample prediction result, and the text-image interaction unit is tuned according to the sample prediction result and the second annotation in the second training sample set.

[0084] In one possible embodiment, the model training module analyzes and processes the second text sample features and the second image sample features through the image-text interaction unit to obtain a sample prediction result, and is configured as follows:

[0085] Inputting the second text sample feature into the cross-attention subunit of the image-text interaction unit, inputting the second image sample feature into the self-attention subunit of the image-text interaction unit, determining a sample key vector and a sample value vector according to the second image sample feature through the self-attention subunit, and determining a picture-text sample interaction result according to the second text sample feature, the sample key vector, and the sample value vector through the cross-attention layer;

[0086] The sample prediction results are obtained by analyzing and processing the image-text sample interaction results through the feedforward subunit, fully connected subunit and logistic regression subunit in the image content recognition model.

[0087] In one possible embodiment, after fixing the network parameters of the image encoding unit and the text encoding unit and training the image-text interaction unit in the image content recognition model based on the second training sample set, the model training module is further configured to:

[0088] The image-text interaction unit and the image encoding unit in the image content recognition model are trained based on the second training sample set.

[0089] In a possible embodiment, a first learning rate for training the image-text interaction unit in the image content recognition model based on the second training sample set is greater than a second learning rate for training the image-text interaction unit and the image encoding unit in the image content recognition model based on the second training sample set.

[0090] It is worth noting that in the embodiment of the above-mentioned image recognition device, the various units and modules included are only divided according to functional logic, but are not limited to the above-mentioned division, as long as the corresponding functions can be achieved; in addition, the specific names of the functional units are only for the convenience of distinguishing each other, and are not used to limit the scope of protection of the embodiments of this application.

[0091] The embodiment of the present application also provides an image recognition device, which can integrate the image recognition device provided by the embodiment of the present application. Figure 7 is a structural diagram of an image recognition device provided by the embodiment of the present application. Referring to Figure 7, the image recognition device includes: an input device 73, an output device 74, a memory 72 and one or more processors 71; the memory 72 is used to store one or more programs; when the one or more programs are executed by one or more processors 71, the one or more processors 71 implement the image recognition method provided by the above embodiment. The above-mentioned image recognition device, equipment and computer can be used to execute the image recognition method provided by any of the above embodiments, and have corresponding functions and beneficial effects.

[0092] The embodiments of the present application also provide a non-volatile storage medium that stores computer-executable instructions, which are used to execute the image recognition method provided in the above embodiments when executed by a computer processor. Of course, the non-volatile storage medium that stores computer-executable instructions provided in the embodiments of the present application, whose computer-executable instructions are not limited to the image recognition method provided above, can also execute the related operations in the image recognition method provided in any embodiment of the present application. The image recognition apparatus, device and storage medium provided in the above embodiments can execute the image recognition method provided in any embodiment of the present application. For technical details not described in detail in the above embodiments, please refer to the image recognition method provided in any embodiment of the present application.

[0093] Based on the above embodiments, the embodiments of the present application also provide a computer program product. The technical solution of the present application, or the part that contributes to the prior art, or all or part of the technical solution can be embodied in the form of a software product. The computer program product is stored in a storage medium and includes a number of instructions for enabling a computer device, a mobile terminal or the processor therein to execute all or part of the steps of the image recognition method provided in each embodiment of the present application.

Claims

1. An image recognition method, wherein: include: Obtain the image to be recognized; The image to be identified is input into a trained image content recognition model, and the image content recognition model uses a text-image interaction unit to analyze and process the target text features and the target image features of the image to be identified to obtain a content recognition result of the image to be identified. The target text features are generated by the text encoding unit in the image content recognition model according to the set target review rules, and the target image features are generated by the image encoding unit in the image content recognition model according to the image to be identified.

2. The image recognition method according to claim 1, wherein: The training process of the image content recognition model includes: Fixing network parameters of an image encoding unit in the image content recognition model, and training a text encoding unit in the image content recognition model based on a first training sample set, wherein the first training sample set includes a first review rule, a first image sample, and a first annotation of the first image sample under the first review rule; The network parameters of the image encoding unit and the text encoding unit are fixed, and the image-text interaction unit in the image content recognition model is trained based on a second training sample set, wherein the second training sample set includes a second review rule, a second image sample, and a second annotation of the second image sample under the second review rule.

3. The image recognition method according to claim 2, wherein: The training of the text encoding unit in the image content recognition model based on the first training sample set includes: Determining a first text sample feature of a first audit rule in a first training sample set by a text encoding unit in the image content recognition model; determining, by an image encoding unit in the image content recognition model, a first image sample feature of a first image sample in the first training sample set; The image-text sample contrast loss is determined according to the first text sample feature and the first image sample feature, and the text encoding unit is optimized based on the image-text sample contrast loss.

4. The image recognition method according to claim 2, wherein: The training of the image-text interaction unit in the image content recognition model based on the second training sample set includes: Determining a second text sample feature of a second review rule in a second training sample set by a text encoding unit in the image content recognition model; determining, by an image encoding unit in the image content recognition model, a second image sample feature of a second image sample in the second training sample set; The second text sample features and the second image sample features are input into the text-image interaction unit in the image content recognition model, the second text sample features and the second image sample features are analyzed and processed by the text-image interaction unit to obtain a sample prediction result, and the text-image interaction unit is tuned according to the sample prediction result and the second annotation in the second training sample set.

5. The image recognition method according to claim 4, wherein: The analyzing and processing the second text sample features and the second image sample features by the image-text interaction unit to obtain a sample prediction result includes: Inputting the second text sample feature into the cross-attention subunit of the image-text interaction unit, inputting the second image sample feature into the self-attention subunit of the image-text interaction unit, determining a sample key vector and a sample value vector according to the second image sample feature by the self-attention subunit, and determining an image-text sample interaction result according to the second text sample feature, the sample key vector, and the sample value vector by the cross-attention layer; The sample prediction result is obtained by analyzing and processing the image-text sample interaction result through the feedforward subunit, the fully connected subunit and the logistic regression subunit in the image content recognition model.

6. The image recognition method according to claim 2, wherein: After fixing the network parameters of the image encoding unit and the text encoding unit, and training the image-text interaction unit in the image content recognition model based on the second training sample set, the method further includes: The image-text interaction unit and the image encoding unit in the image content recognition model are trained based on the second training sample set.

7. The image recognition method according to claim 6, wherein: The first learning rate for training the image-text interaction unit in the image content recognition model based on the second training sample set is greater than the second learning rate for training the image-text interaction unit and the image encoding unit in the image content recognition model based on the second training sample set.

8. An image recognition device, wherein: It includes an image acquisition module and an image recognition module, wherein: The image acquisition module is configured to acquire an image to be identified; The image recognition module is configured to input the image to be recognized into a trained image content recognition model, and use the image content recognition model to analyze and process the target text features and the target image features of the image to be recognized using the image-text interaction unit to obtain the content recognition result of the image to be recognized. The target text features are generated by the text encoding unit in the image content recognition model according to the set target review rules, and the target image features are generated by the image encoding unit in the image content recognition model according to the image to be recognized.

9. An image recognition device, wherein: include: memory and one or more processors; The memory is used to store one or more programs; When the one or more programs are executed by the one or more processors, the one or more processors implement the image recognition method according to any one of claims 1 to 7.

10. A non-volatile storage medium storing computer-executable instructions, wherein: When the computer executable instructions are executed by a computer processor, they are used to perform the image recognition method according to any one of claims 1 to 7.

11. A computer program product comprising a computer program, wherein When the computer program is executed by a processor, the image recognition method according to any one of claims 1 to 7 is implemented.

Citation Information

Patent Citations

  • Content auditing method, training method of content auditing model and related device

    CN115565038A

  • Picture processing method and related device

    CN117011859A

  • Audit label configuration strategy generation method and device, equipment and storage medium

    CN117556100A

  • Image recognition method and device, equipment, storage medium and product

    CN117893855A

  • Multimodal Image Classifier using Textual and Visual Embeddings

    US20210264203A1

Cited By

  • Method, device and equipment for auditing claim declaration form of production breaking, and medium

    CN121118879A

  • Image sensitive content auditing method and system

    CN121353653A