Image recognition method and related device
Visual and text features are extracted through the image recognition model, and the self-attention and cross-attention mechanism are fusion, the problem of low image recognition accuracy in the prior art is solved, achieving a more efficient image recognition effect.
Patent Information
- Application Number
- CN202410021385.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-01-05
- Publication Date
- 2025-07-08
Smart Images

Figure CN120279564A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of computer technology, and particularly to an image recognition method and related device. Background Art
[0002] With the rapid development of artificial intelligence technology, in image processing scenarios such as image review scenarios, in order to improve the quality and efficiency of image review, it is very important to implement image recognition through artificial intelligence technology.
[0003] In related technologies, an image recognition method generally refers to: extracting visual features of an image through an image recognition model that has been pre-trained until convergence to obtain an image recognition result.
[0004] However, an image includes multi-modal information. The above method only extracts visual features of the image to obtain an image recognition result, and is likely to lose text information in the image, resulting in a low recognition accuracy of the image, and thus a poor recognition effect of the image. Summary of the Invention
[0005] To solve the above technical problems, this application provides an image recognition method and related device. On the basis of extracting visual features of an image, text features of text information in the image are further extracted; different attention points in multi-modal features are considered through a self-attention mechanism, and the multi-modal semantics of multi-modal features are aligned through a cross-attention mechanism to effectively deeply fuse multi-modal features and effectively independently retain multi-modal features. The multi-modal features complement and promote each other, making the feature expression ability of the image stronger; thereby improving the recognition accuracy and recognition effect of the image.
[0006] Embodiments of this application disclose the following technical solutions:
[0007] On the one hand, an embodiment of this application provides an image recognition method, and the method includes:
[0008] Extracting visual features of an image to be recognized through a first extraction network in an image recognition model to obtain visual features to be recognized; extracting text features of text information in the image to be recognized through a second extraction network in the image recognition model to obtain text features to be recognized;
[0009] Performing feature fusion of a self-attention mechanism on a first concatenated feature of the visual features to be recognized and the text features to be recognized through a multi-modal fusion network in the image recognition model to obtain a first fused feature; performing feature fusion of a cross-attention mechanism on the visual features to be recognized and the text features to be recognized through the multi-modal fusion network to obtain a second fused feature;
[0010] Through the multi-modal fusion network, perform feature concatenation on the first fusion feature, the second fusion feature, the visual feature to be recognized, and the text feature to be recognized, to obtain a fusion feature to be recognized;
[0011] Through the recognition network in the image recognition model, perform image recognition on the fusion feature to be recognized, to obtain an image recognition result of the image to be recognized.
[0012] On the other hand, an embodiment of the present application provides an image recognition device, the device includes: an extraction unit, a fusion unit, a concatenation unit, and a recognition unit;
[0013] The extraction unit is configured to, through a first extraction network in an image recognition model, perform visual feature extraction on an image to be recognized, to obtain a visual feature to be recognized; through a second extraction network in the image recognition model, perform text feature extraction on text information in the image to be recognized, to obtain a text feature to be recognized;
[0014] The fusion unit is configured to, through a multi-modal fusion network in the image recognition model, perform feature fusion of a self-attention mechanism on a first concatenation feature of the visual feature to be recognized and the text feature to be recognized, to obtain a first fusion feature; through the multi-modal fusion network, perform feature fusion of a cross-attention mechanism on the visual feature to be recognized and the text feature to be recognized, to obtain a second fusion feature;
[0015] The concatenation unit is configured to, through the multi-modal fusion network, perform feature concatenation on the first fusion feature, the second fusion feature, the visual feature to be recognized, and the text feature to be recognized, to obtain a fusion feature to be recognized;
[0016] The recognition unit is configured to, through the recognition network in the image recognition model, perform image recognition on the fusion feature to be recognized, to obtain an image recognition result of the image to be recognized.
[0017] On the other hand, an embodiment of the present application provides a computer device, the computer device includes a processor and a memory:
[0018] The memory is used to store a computer program and transmit the computer program to the processor;
[0019] The processor is configured to execute the method described in any of the foregoing aspects according to the instructions in the computer program.
[0020] On the other hand, an embodiment of the present application provides a computer-readable storage medium, the computer-readable storage medium is used to store a computer program, and when the computer program runs on a computer device, it causes the computer device to execute the method described in any of the foregoing aspects.
[0021] On the other hand, an embodiment of the present application provides a computer program product, including a computer program, which, when running on a computer device, causes the computer device to execute the method described in any of the foregoing aspects.
[0022] As can be seen from the above technical solutions, first, the image to be recognized is input into the first extraction network in the image recognition model, and the visual feature to be recognized is extracted by visual feature extraction; the text information in the image to be recognized is input into the second extraction network in the image recognition model, and the text feature to be recognized is extracted by text feature extraction; this method not only extracts the visual feature of the image to be recognized, but also extracts the text feature of the text information in the image to be recognized. Then, the first concatenated feature of the visual feature to be recognized and the text feature to be recognized is input into the multi-modal fusion network in the image recognition model, and the first fusion feature is obtained by feature fusion based on the self-attention mechanism; the visual feature to be recognized and the text feature to be recognized are input into the multi-modal fusion network, and the second fusion feature is obtained by feature fusion based on the cross-attention mechanism; the first fusion feature, the second fusion feature, the visual feature to be recognized, and the text feature to be recognized are concatenated into the fusion feature to be recognized; this method not only effectively and deeply fuses the visual feature to be recognized and the text feature to be recognized, that is, considering different attention points in the visual feature to be recognized and the text feature to be recognized through the self-attention mechanism, and aligning the visual semantics of the visual feature to be recognized and the text semantics of the text feature to be recognized through the cross-attention mechanism, realizing the mutual complementation and mutual promotion of the visual feature to be recognized and the text feature to be recognized, but also effectively and independently retains the visual feature to be recognized and the text feature to be recognized; so as to effectively comprehensively enhance the feature expression ability of the image to be recognized. Finally, the fusion feature to be recognized is input into the recognition network in the image recognition model, and the image recognition result of the image to be recognized is obtained by image recognition; this method recognizes the fusion feature to be recognized with stronger feature expression ability, making the recognition accuracy and recognition effect of the image to be recognized higher and better.
[0023] Based on this, the method further extracts the text feature of the text information in the image on the basis of extracting the visual feature of the image; considering different attention points in the multi-modal features through the self-attention mechanism, aligning the multi-modal semantics of the multi-modal features through the cross-attention mechanism, the multi-modal features complement and promote each other, so as to effectively and deeply fuse the multi-modal features and effectively and independently retain the multi-modal features, enhancing the feature expression ability of the image in multiple aspects; thereby improving the recognition accuracy and recognition effect of the image. Description of the Drawings
[0024] To more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the accompanying drawings required for the description of the embodiments or the prior art. Obviously, the accompanying drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0025] Figure 1 Schematic diagram of a system for an image recognition method provided by an embodiment of the present application;
[0026] Figure 2 Flowchart of an image recognition method provided by an embodiment of the present application;
[0027] Figure 3 Schematic diagram of obtaining visual features to be recognized based on an image to be recognized provided by an embodiment of the present application;
[0028] Figure 4 Schematic diagram of obtaining text features to be recognized based on text information in an image to be recognized provided by an embodiment of the present application;
[0029] Figure 5 Schematic diagram of obtaining fused features to be recognized based on visual features to be recognized and text features to be recognized provided by an embodiment of the present application;
[0030] Figure 6 Flowchart of an advertisement image recognition method provided by an embodiment of the present application;
[0031] Figure 7 Structural diagram of an image recognition device provided by an embodiment of the present application;
[0032] Figure 8 Structural diagram of a server provided by an embodiment of the present application;
[0033] Figure 9 Structural diagram of a terminal provided by an embodiment of the present application. Detailed implementation manners
[0034] The following describes the embodiments of the present application with reference to the accompanying drawings.
[0035] At present, in image processing scenarios such as image review scenarios, in order to improve the quality and efficiency of image review, image recognition is required. Generally, an image recognition model pre-trained until the model converges is used to extract the visual features of the image to obtain the image recognition result. For example, in the advertisement review scenario, in order to improve the quality and efficiency of advertisement review, image recognition of advertisement images is required. Generally, an image recognition model pre-trained until the model converges is used to extract the visual features of the advertisement image to obtain the advertisement recognition result.
[0036] However, through research, it is found that the image includes multimodal information. The above method only extracts the visual features of the image to obtain the image recognition result, and it is easy to lose the text information in the image, resulting in a low recognition accuracy of the image, and thus a poor recognition effect of the image. For example, only extracting the visual features of the advertisement image to obtain the advertisement recognition result is easy to lose the text information in the advertisement image, resulting in a low recognition accuracy of the advertisement image, and thus a poor recognition effect of the advertisement image.
[0037] The embodiment of the present application provides an image recognition method. On the basis of extracting the visual features of the image, the text features of the text information in the image are further extracted; the self-attention mechanism is used to consider different attention points in the multimodal features, and the cross-attention mechanism is used to align the multimodal semantics of the multimodal features. The multimodal features complement and promote each other to effectively deeply fuse the multimodal features and effectively independently retain the multimodal features, enhancing the feature expression ability of the image in multiple aspects; thereby improving the recognition accuracy and recognition effect of the image.
[0038] Next, the system architecture of the image recognition method will be introduced. Refer to Figure 1 , Figure 1 which is a schematic diagram of a system for an image recognition method provided by an embodiment of the present application. The system includes a computer device 100, and the computer device 100 is used to execute the image recognition method.
[0039] The computer device 100 extracts visual features from the image to be recognized through the first extraction network in the image recognition model to obtain the visual features to be recognized; and extracts text features from the text information in the image to be recognized through the second extraction network in the image recognition model to obtain the text features to be recognized.
[0040] As an example, the image to be recognized is an advertisement image to be recognized, the first extraction network is a deep learning model based on Transformer, and the second extraction network is a pre-trained language representation model (Bidirectional Encoder Representation from Transformers, Bert) model; the computer device 100 inputs the advertisement image to be recognized into the deep learning model based on Transformer, and the visual feature extraction obtains the visual features to be recognized as the advertisement visual features to be recognized; inputs the text information in the advertisement image to be recognized into the Bert model, and the text feature extraction obtains the text features to be recognized as the advertisement text features to be recognized.
[0041] The computer device 100 performs feature fusion of the first concatenated feature of the visual feature to be recognized and the text feature to be recognized through the multi-modal fusion network in the image recognition model by means of the self-attention mechanism, and obtains the first fusion feature; through the multi-modal fusion network, it performs feature fusion of the visual feature to be recognized and the text feature to be recognized by means of the cross-attention mechanism, and obtains the second fusion feature.
[0042] As an example, the multi-modal fusion network is a fusion model based on self-attention (Self-Attention) and cross-attention (Cross-Attention). On the basis of the above example, the computer device 100 inputs the first concatenated feature of the visual feature of the advertisement to be recognized and the text feature of the advertisement to be recognized into the fusion model based on Self-Attention and Cross-Attention in the image recognition model, and obtains the first fusion feature through the feature fusion based on the Self-Attention mechanism; the computer device 100 inputs the visual feature of the advertisement to be recognized and the text feature of the advertisement to be recognized into the fusion model based on Self-Attention and Cross-Attention in the image recognition model, and obtains the second fusion feature through the feature fusion based on the Cross-Attention mechanism.
[0043] The computer device 100 performs feature concatenation on the first fusion feature, the second fusion feature, the visual feature to be recognized, and the text feature to be recognized through the multi-modal fusion network, and obtains the fusion feature to be recognized.
[0044] As an example, on the basis of the above example, the computer device 100 concatenates the first fusion feature, the second fusion feature, the visual feature of the advertisement to be recognized, and the text feature of the advertisement to be recognized into the fusion feature to be recognized, that is, the fusion feature of the advertisement to be recognized.
[0045] The computer device 100 performs image recognition on the fusion feature to be recognized through the recognition network in the image recognition model, and obtains the image recognition result of the image to be recognized.
[0046] As an example, the recognition network is a classifier SoftMax. On the basis of the above example, the computer device 100 inputs the fusion feature of the advertisement to be recognized into SoftMax in the image recognition model, and the image recognition result of the advertisement image to be recognized obtained by image recognition is the advertisement recognition result.
[0047] That is to say, not only extract the visual features of the image to be recognized, but also extract the text features of the text information in the image to be recognized; not only effectively deeply fuse the visual features to be recognized and the text features to be recognized, that is, consider different attention points in the visual features to be recognized and the text features to be recognized through the self-attention mechanism, and align the visual semantics of the visual features to be recognized and the text semantics of the text features to be recognized through the cross-attention mechanism, so as to achieve the mutual complementarity and mutual promotion of the visual features to be recognized and the text features to be recognized, but also effectively independently retain the visual features to be recognized and the text features to be recognized; so as to effectively comprehensively enhance the feature expression ability of the image to be recognized; recognize the fused features to be recognized with stronger feature expression ability, so that the recognition accuracy and recognition effect of the image to be recognized are higher and better.
[0048] That is, on the basis of extracting the visual features of the advertisement image, the method further extracts the text features of the text information in the advertisement image; considers different attention points in the multimodal features through the self-attention mechanism, aligns the multimodal semantics of the multimodal features through the cross-attention mechanism, and the multimodal features complement and promote each other, so as to effectively deeply fuse the multimodal features and effectively independently retain the multimodal features, enhancing the feature expression ability of the advertisement image in multiple aspects; thereby improving the recognition accuracy and recognition effect of the advertisement image.
[0049] It should be noted that the image recognition method in the embodiments of this application involves artificial intelligence. Artificial intelligence uses digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, a theory, method, technology, and application system that perceives the environment, acquires knowledge, and uses knowledge to obtain the best results. In other words, artificial intelligence is a comprehensive technology in computer science that attempts to understand the essence of intelligence and produce a new intelligent machine that can react in a way similar to human intelligence. Artificial intelligence is also the study of the design principles and implementation methods of various intelligent machines, enabling machines to have the functions of perception, reasoning, and decision-making.
[0050] Artificial intelligence technology is an interdisciplinary subject, involving a wide range of fields, including both hardware-level technologies and software-level technologies. The basic technologies of artificial intelligence generally include sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, big data processing technology, pre-trained model technology, operation / interaction systems, mechatronics, etc. Among them, the pre-trained model, also known as the large model or the foundation model, can be widely applied to downstream tasks in various directions of artificial intelligence after fine-tuning. The software technologies of artificial intelligence mainly include several major directions such as computer vision technology, speech processing technology, natural language processing technology, and machine learning / deep learning. In the embodiments of this application, artificial intelligence technology mainly involves computer vision technology, natural language processing technology, and machine learning / deep learning and other technologies.
[0051] Computer vision is a science that studies how to enable machines to "see". More specifically, it refers to machine vision that uses cameras and computers to replace human eyes for tasks such as object recognition, tracking, and measurement, and further performs image processing to make the computer-processed images more suitable for human eyes to observe or be transmitted to instruments for detection. As a scientific discipline, computer vision studies related theories and technologies and attempts to build artificial intelligence systems that can obtain information from images or multi-dimensional data. The large model technology has brought important changes to the development of computer vision technology. Pretrained models in the field of vision such as swin-transformer, Vision Transformer, Vision MoE, and MAE can be quickly and widely applied to downstream specific tasks after fine-tuning. In the embodiments of this application, computer vision technology mainly involves technologies such as image processing, image recognition, image semantic understanding, image retrieval, optical character recognition, video processing, video semantic understanding, and video content / behavior recognition.
[0052] Natural language processing is an important direction in the field of computer science and artificial intelligence. It studies various theories and methods that can achieve effective communication between humans and computers in natural language. Natural language processing involves natural language, that is, the language people use in daily life, and is closely related to linguistics research; at the same time, it involves computer science and mathematics. The important technology for model training in the field of artificial intelligence, the pretrained model, is developed from large language models in the NLP field. After fine-tuning, the large language model can be widely applied to downstream tasks. In the embodiments of this application, natural language processing technology mainly involves technologies such as text processing, semantic understanding, and robot question answering.
[0053] Machine learning / deep learning is a multi-disciplinary cross-discipline that involves multiple disciplines such as probability theory, statistics, approximation theory, convex analysis, and algorithm complexity theory. It specifically studies how computers simulate or implement human learning behaviors to acquire new knowledge or skills and reorganize the existing knowledge structure to continuously improve their own performance. Machine learning is the core of artificial intelligence and the fundamental way to make computers intelligent, and its applications cover all fields of artificial intelligence. Machine learning and deep learning usually include technologies such as artificial neural networks, belief networks, reinforcement learning, transfer learning, inductive learning, and learning from demonstration. The pretrained model is the latest development result of deep learning, integrating the above technologies.
[0054] It should be noted that in the embodiments of this application, the computer device can be a server or a terminal. The method provided in the embodiments of this application can be executed independently by the terminal or the server, or can be executed in cooperation by the terminal and the server. Among them, when the method provided in the embodiments of this application is executed independently by the terminal or the server, its execution method is the same as Figure 1The corresponding embodiments are similar, mainly by replacing the computer device with a terminal or a server. In addition, when the method provided in the embodiments of the present application is executed in cooperation with a terminal and a server, the steps that need to be reflected on the front-end interface can be executed by the terminal, while some steps that require background calculation and do not need to be reflected on the front-end interface can be executed by the server.
[0055] Among them, the terminal can be a smart phone, a tablet computer, a laptop computer, a desktop computer, a smart voice interaction device, a vehicle-mounted terminal, an extended reality device, an aircraft, etc., but is not limited thereto. The server can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server providing cloud computing services, but is not limited thereto. The terminal and the server can be directly or indirectly connected through wired or wireless communication methods, and the present application does not make any restrictions here. For example, the terminal and the server can be connected through a network, and the network can be a wired or wireless network.
[0056] In addition, the embodiments of the present application can be applied to various scenarios, including but not limited to cloud technology, artificial intelligence, intelligent transportation, assisted driving, autonomous driving, digital humans, virtual humans, virtual reality, augmented reality, mixed reality, audio and video, etc.
[0057] Next, taking the computer device executing the method provided in the embodiments of the present application as an example, the image recognition method provided in the embodiments of the present application will be introduced in detail with reference to the accompanying drawings. See Figure 2 , Figure 2 which is a flowchart of an image recognition method provided in the embodiments of the present application. The method includes:
[0058] S201: Extract visual features of the image to be recognized through the first extraction network in the image recognition model to obtain the visual features to be recognized; extract text features of the text information in the image to be recognized through the second extraction network in the image recognition model to obtain the text features to be recognized.
[0059] In the related art, the image recognition method refers to: extracting visual features of an image through an image recognition model that has been pre-trained until convergence to obtain an image recognition result. However, an image includes multi-modal information. Only extracting the visual features of the image to obtain an image recognition result is likely to lose the text information in the image, resulting in a low recognition accuracy of the image and thus a poor recognition effect of the image.
[0060] Therefore, in the embodiments of the present application, to solve the above problems, considering that an image includes multi-modal information such as image information and text information, and to avoid losing the text information in the image, for the image recognition model, it is necessary to not only extract the visual features of the image, but also extract the text features of the text information in the image.
[0061] That is, the image recognition model needs to include a first extraction network for extracting visual features of the image and a second extraction network for extracting text features of the text information in the image. Based on this, for the image to be recognized, the image to be recognized is input into the first extraction network in the image recognition model, and the visual features are extracted to obtain the visual features to be recognized. The text information in the image to be recognized is input into the second extraction network in the image recognition model, and the text features are extracted to obtain the text features to be recognized.
[0062] Among them, the first extraction network includes a convolutional neural network or a deep learning model based on Transformer. The visual features to be recognized refer to the features representing the visual information of the image to be recognized, which are used to indicate the visual semantics of the image to be recognized. The visual features to be recognized are specifically represented as visual vectors to be recognized. The second extraction network can be a Bert model. The text features to be recognized refer to the features representing the text information in the image to be recognized, which are used to indicate the text semantics of the image to be recognized. The text features to be recognized are specifically represented as text vectors to be recognized.
[0063] Based on the fact that the image to be recognized includes multimodal information such as image information and text information, this S201 not only extracts the visual features of the image to be recognized, but also extracts the text features of the text information in the image to be recognized, thus avoiding the loss of text information in the image to be recognized during the feature extraction process. It provides sufficient and rich multimodal features representing multimodal information for improving the recognition accuracy and recognition effect of the image to be recognized subsequently.
[0064] See Figure 3 , Figure 3 is a schematic diagram of obtaining visual features to be recognized based on the image to be recognized provided by an embodiment of this application. One implementation means that the first extraction network is a convolutional neural network. The image to be recognized is input into the convolutional neural network in the image recognition model, and the visual features are extracted to obtain the visual features to be recognized. Another implementation means that the first extraction network is a deep learning model based on Transformer. The image to be recognized is input into the deep learning model based on Transformer in the image recognition model, and the visual features are extracted to obtain the visual features to be recognized. The deep learning model based on Transformer has the following advantages compared with the convolutional neural network: In the deep learning model based on Transformer, the self-attention mechanism allows each pixel in the image to be recognized to interact with other pixels in the image to be recognized, which helps to capture the global information in the image to be recognized. The deep learning model based on Transformer helps to process the long-range dependence relationship in the image to be recognized. The deep learning model based on Transformer does not require a fixed-size input and is more flexible when processing images to be recognized with variable sizes.
[0065] Among them, the convolutional neural network can be Resnet, Inceptionnet, etc., and the deep learning model based on Transformer can be Swin Transformer, Vision Transformer, etc.
[0066] See Figure 4 , Figure 4 This is a schematic diagram of obtaining the text features to be recognized based on the text information in the image to be recognized provided by the embodiment of the present application. The implementation method means that the second extraction network is a Bert model, and the text information in the image to be recognized is input into the Bert model in the image recognition model, and the text features are extracted to obtain the text features to be recognized. Among them, before the text information in the image to be recognized is input into the Bert model in the image recognition model, it is necessary to add a flag bit [CLS] at the front end of each word segment Token (that is, CP1, CP2,..., CP n ) and then multiply it by the masking Mask.
[0067] As an example of S201, the image to be recognized is an advertising image to be recognized, the first extraction network is a Swin Transformer model, and the second extraction network is a Bert model; the computer device inputs the advertising image to be recognized into the Swin Transformer model in the image recognition model, and the visual features are extracted to obtain the visual features to be recognized as the advertising visual features E to be recognized vision ; the text information in the advertising image to be recognized is input into the Bert model in the image recognition model, and the text features are extracted to obtain the text features to be recognized as the advertising text features E to be recognized ocr .
[0068] S202: Through the multi-modal fusion network in the image recognition model, perform feature fusion of the self-attention mechanism on the first concatenated feature of the visual features to be recognized and the text features to be recognized to obtain the first fusion feature; through the multi-modal fusion network, perform feature fusion of the cross-attention mechanism on the visual features to be recognized and the text features to be recognized to obtain the second fusion feature.
[0069] In the embodiment of the present application, in order to solve the above problems, after extracting the visual features of the image and the text features of the text information in the image, since the visual features are used to indicate the visual semantics of the image and the text features are used to indicate the text semantics of the image, in order to understand the image semantics more deeply and avoid the independence between the visual features and the text features, it is necessary to deeply fuse the visual features and the text features, not only considering different attention points in the multi-modal features, but also aligning the multi-modal semantics of the multi-modal features.
[0070] That is, the image recognition model needs to include a multi-modal fusion network for deeply fusing visual features and text features. After extracting the visual features to be recognized and the text features to be recognized, the first concatenated feature of the visual features to be recognized and the text features to be recognized is input into the multi-modal fusion network in the image recognition model. Based on the feature fusion of the self-attention mechanism, the first fusion feature is obtained, realizing the consideration of different attention points in the visual features to be recognized and the text features to be recognized; the visual features to be recognized and the text features to be recognized are input into the multi-modal fusion network, and the second fusion feature is obtained based on the feature fusion of the cross-attention mechanism, realizing the alignment of the visual semantics of the visual features to be recognized and the text semantics of the text features to be recognized.
[0071] Among them, the first fusion feature refers to the fusion feature that considers different attention points in the visual features to be recognized and the text features to be recognized. The first fusion feature is specifically represented as the first fusion vector; the second fusion feature represents the fusion feature that aligns the visual semantics of the visual features to be recognized and the text semantics of the text features to be recognized. The second fusion feature is specifically represented as the second fusion vector.
[0072] In S202, different attention points in the visual features to be recognized and the text features to be recognized are considered through the self-attention mechanism, and the visual semantics of the visual features to be recognized and the text semantics of the text features to be recognized are aligned through the cross-attention mechanism, effectively and deeply fusing the visual features to be recognized and the text features to be recognized, realizing the mutual complementarity and mutual promotion of the visual features to be recognized and the text features to be recognized; for subsequent improvement of the recognition accuracy and recognition effect of the image to be recognized, a deep fusion feature of accurate and more comprehensive multi-modal features is provided.
[0073] As an example of S202, the multi-modal fusion network includes a Self-Attention module and a Cross-Attention module. On the basis of the above S201 example, the first concatenated feature of the visual features E vision of the advertisement to be recognized and the text features E ocr of the advertisement to be recognized is E concat , that is, E concat = concat(E visi on , E ocr ). The computer device inputs E concat into the Self-Attention module, and the first fusion feature obtained based on the feature fusion of the Self-Attention mechanism is E self ; the computer device inputs E vision and E ocr into the Cross-Attention module, and the second fusion feature obtained based on the feature fusion of the Cross-Attention mechanism is E cross .
[0074] S203: Through the multi-modal fusion network, perform feature concatenation on the first fusion feature, the second fusion feature, the visual feature to be recognized, and the text feature to be recognized to obtain the fusion feature to be recognized.
[0075] In the embodiments of the present application, to solve the above problems, after deeply fusing the visual feature and the text feature, since the visual feature refers to the feature representing the visual information of the image and the text feature refers to the feature representing the text information in the image, in order to focus on the original content information of the image and avoid the weakening of the visual feature and the text feature being masked, it is also necessary to further independently retain the visual feature and the text feature to obtain the fusion feature.
[0076] That is, after fusing to obtain the first fusion feature and the second fusion feature, through the above multi-modal fusion network, the first fusion feature, the second fusion feature, the visual feature to be recognized, and the text feature to be recognized are concatenated into the fusion feature to be recognized.
[0077] Among them, the fusion feature to be recognized refers to the fusion feature that considers different focuses in the visual feature to be recognized and the text feature to be recognized, aligns the visual semantics of the visual feature to be recognized and the text semantics of the text feature to be recognized, represents the visual information of the image to be recognized, and represents the text information in the image to be recognized; the fusion feature to be recognized is specifically manifested as the fusion vector to be recognized.
[0078] Based on the fact that S202 effectively deeply fuses the visual feature to be recognized and the text feature to be recognized, realizes the mutual complement and mutual promotion of the visual feature to be recognized and the text feature to be recognized, and through further effectively independently retaining the visual feature to be recognized and the text feature to be recognized, effectively comprehensively enhances the feature expression ability of the image to be recognized; provides a more accurate and more comprehensive comprehensive fusion feature of multi-modal features for subsequent improving the recognition accuracy and recognition effect of the image to be recognized.
[0079] As an example of S203, on the basis of the example of S202 above, the computer device concatenates the first fusion feature E self , the second fusion feature E cross , the visual feature of the advertisement to be recognized E vision and the text feature of the advertisement to be recognized E ocr into the fusion feature to be recognized, that is, the fusion feature of the advertisement to be recognized E.
[0080] S204: Through the recognition network in the image recognition model, perform image recognition on the fusion feature to be recognized to obtain the image recognition result of the image to be recognized.
[0081] In the embodiments of the present application, to solve the above problems, after deeply fusing visual features and text features and further independently retaining the visual features and text features to obtain fused features, in order to improve the recognition accuracy and recognition effect of images, it is necessary to recognize the fused features to obtain an image recognition result.
[0082] That is, the image recognition model needs to include a recognition network for recognizing the fused features. The fused features to be recognized are input into the recognition network in the image recognition model, and the image recognition result of the image to be recognized is obtained through image recognition.
[0083] This S204 enables the recognition accuracy and recognition effect of the image to be recognized to be higher and better by recognizing the fused features to be recognized with stronger feature expression ability, thereby improving the recognition accuracy and recognition effect of the image to be recognized; and providing a more accurate and comprehensive recognition result for subsequent processing of the image to be recognized.
[0084] As an example of this S204, based on the above example of S203, the recognition network is the classifier SoftMax. Based on the above example, the computer device inputs the fused features E of the advertisement to be recognized into SoftMax in the image recognition model, and the image recognition result of the advertisement image to be recognized obtained through image recognition is the advertisement recognition result.
[0085] As can be seen from the above technical solution, first, the image to be recognized is input into the first extraction network in the image recognition model, and the visual feature to be recognized is extracted through visual feature extraction; the text information in the image to be recognized is input into the second extraction network in the image recognition model, and the text feature to be recognized is extracted through text feature extraction; this method not only extracts the visual feature of the image to be recognized, but also extracts the text feature of the text information in the image to be recognized. Then, the first concatenated feature of the visual feature to be recognized and the text feature to be recognized is input into the multi-modal fusion network in the image recognition model, and the first fused feature is obtained through feature fusion based on the self-attention mechanism; the visual feature to be recognized and the text feature to be recognized are input into the multi-modal fusion network, and the second fused feature is obtained through feature fusion based on the cross-attention mechanism; the first fused feature, the second fused feature, the visual feature to be recognized, and the text feature to be recognized are concatenated into the fused feature to be recognized; this method not only effectively and deeply fuses the visual feature to be recognized and the text feature to be recognized, that is, considering different focus points in the visual feature to be recognized and the text feature to be recognized through the self-attention mechanism, and aligning the visual semantics of the visual feature to be recognized and the text semantics of the text feature to be recognized through the cross-attention mechanism, realizing the mutual complement and mutual promotion of the visual feature to be recognized and the text feature to be recognized, but also effectively and independently retains the visual feature to be recognized and the text feature to be recognized; so as to effectively comprehensively enhance the feature expression ability of the image to be recognized. Finally, the fused feature to be recognized is input into the recognition network in the image recognition model, and the image recognition result of the image to be recognized is obtained through image recognition; this method recognizes the fused feature to be recognized with stronger feature expression ability, making the recognition accuracy and recognition effect of the image to be recognized higher and better.
[0086] Based on this, this method further extracts the text feature of the text information in the image on the basis of extracting the visual feature of the image; considering different focus points in the multi-modal features through the self-attention mechanism, aligning the multi-modal semantics of the multi-modal features through the cross-attention mechanism, the multi-modal features complement and promote each other, so as to effectively and deeply fuse the multi-modal features, and effectively and independently retain the multi-modal features, enhancing the feature expression ability of the image in multiple aspects; thereby improving the recognition accuracy and recognition effect of the image.
[0087] In the embodiments of the present application, considering that the self-attention mechanism is mainly used to focus on the correlation between different parts of the overall input, multiple first parameter matrices of the self-attention mechanism can map the overall input into multiple features. Therefore, in step S202 above, through the multi-modal fusion network in the image recognition model, feature fusion of the self-attention mechanism is performed on the first concatenated feature of the visual feature to be recognized and the text feature to be recognized to obtain the first fusion feature. When specifically implemented, first, the visual feature to be recognized and the text feature to be recognized are concatenated into the first concatenated feature. Then, multiple first parameter matrices of the self-attention mechanism are used to map the first concatenated feature into multiple first mapped features. Finally, the self-attention mechanism is used to fuse the multiple first mapped features into the first fusion feature. Based on this, the present application provides a possible implementation manner. The steps of obtaining the first fusion feature in step S202 above include the following S2021-S2023 (not shown in the figure):
[0088] S2021: Perform feature concatenation on the visual feature to be recognized and the text feature to be recognized to obtain the first concatenated feature.
[0089] S2022: Through the multi-modal fusion network, according to the first concatenated feature and multiple first parameter matrices of the self-attention mechanism, determine multiple first mapped features of the first concatenated feature.
[0090] S2023: Through the multi-modal fusion network, perform feature fusion on the multiple first mapped features according to the self-attention mechanism to obtain the first fusion feature.
[0091] S2021-S2023 can pay attention to the correlation between different parts of the first concatenated feature of the visual feature to be recognized and the text feature to be recognized, model the dependencies between different positions within the first concatenated feature, allow each element in the first concatenated feature to interact with other elements in the first concatenated feature, capture the long-range dependencies within the first concatenated feature, and are beneficial to global modeling of the first concatenated feature. Thus, the first fusion feature takes into account different focus points in the visual feature to be recognized and the text feature to be recognized.
[0092] As an example of S2021-S2023, on the basis of the above S202 example, multiple first parameter matrices of the Self-Attention mechanism include W1 q 、W1 k and W1 v ; the computer device concatenates the visual feature E vision of the advertisement to be recognized and the text feature E ocr of the advertisement to be recognized into the first concatenated feature E concat , inputs E concat into the Self-Attention module, and passes through W1 q 、W1k and W1 v Map E concat to multiple first mapping features including Q1, K1, and V1; fuse Q1, K1, and V1 into the first fusion feature E through the Self-Attention mechanism self . Among them, the calculation formulas for Q1, K1, and V1 are as follows:
[0093] Q1 = W1 q ×E concat
[0094] K1 = W1 k ×E concat
[0095] V1 = W1 v ×E concat
[0096] In the embodiments of the present application, considering that the cross-attention mechanism is mainly used to focus on the correlation between different inputs to process the semantic relationship between different inputs, different inputs are mapped to different features through multiple second parameter matrices of the cross-attention mechanism; therefore, in the above S202, through the multi-modal fusion network, the cross-attention mechanism is used to perform feature fusion on the visual feature to be recognized and the text feature to be recognized to obtain the second fusion feature. When specifically implemented, first, through multiple second parameter matrices of the cross-attention mechanism, the visual feature to be recognized is mapped to the second mapping feature, and the text feature to be recognized is mapped to the third mapping feature; then, the second mapping feature and the third mapping feature are fused into the second fusion feature through the cross-attention mechanism. Based on this, the present application provides a possible implementation manner. The steps of obtaining the second fusion feature in the above S202 include the following S2024-S2025 (not shown in the figure):
[0097] S2024: Through the multi-modal fusion network, determine the second mapping feature of the visual feature to be recognized and the third mapping feature of the text feature to be recognized according to the visual feature to be recognized, the text feature to be recognized, and multiple second parameter matrices of the cross-attention mechanism.
[0098] S2025: Through the multi-modal fusion network, perform feature fusion on the second mapping feature and the third mapping feature according to the cross-attention mechanism to obtain the second fusion feature.
[0099] The S2024-S2025 can handle the dependency relationship between the visual feature to be recognized and the text feature to be recognized, capture the correlation information between the visual feature to be recognized and the text feature to be recognized, and is beneficial to processing the complex dependency between the visual feature to be recognized and the text feature to be recognized; thereby enabling the second fusion feature to align the visual semantics of the visual feature to be recognized and the text semantics of the text feature to be recognized.
[0100] As an example of S2024 - S2025, based on the above S202 example, multiple second parameter matrices of the Cross - Attention mechanism include W2 q , W2 k and W2 v ; the computer device maps the visual features E q of the advertisement to be recognized to the second mapped feature Q2 through W2 k , W2 v and maps the text features E vision of the advertisement to be recognized to the third mapped features including K2 and V2; the Q2, K2 and V2 are fused into the second fused feature E ocr through the Cross - Attention mechanism. Among them, the calculation formulas of Q2, K2 and V2 are as follows: cross
[0101] Q2 = W2 q ×E vision
[0102] K2 = W2 k ×E ocr
[0103] V2 = W2 v ×E ocr
[0104] In the embodiments of the present application, since both the first fused feature and the second fused feature are fused features of the visual features to be recognized and the text features to be recognized, splicing the first fused feature and the second fused feature can effectively and deeply fuse the visual features to be recognized and the text features to be recognized in a real sense, and obtain the deep - fused feature of the visual features to be recognized and the text features to be recognized; therefore, when specifically implementing S203 above, first, the first fused feature and the second fused feature are spliced into a third fused feature through a multi - modal fusion network; then, the third fused feature, the visual features to be recognized and the text features to be recognized are spliced into the fused features to be recognized through the multi - modal fusion network. Based on this, the present application provides a possible implementation manner, and the above S203 includes S2031 - S2032 (not shown in the figure):
[0105] S2031: Through a multi - modal fusion network, perform feature splicing on the first fused feature and the second fused feature to obtain a third fused feature.
[0106] Among them, the third fused feature refers to a fused feature that considers different attention points in the visual features to be recognized and the text features to be recognized, and aligns the visual semantics of the visual features to be recognized and the text semantics of the text features to be recognized.
[0107] Among them, when the S2031 is specifically implemented, since the feature scale of the first fusion feature is the same as that of the first concatenated feature, and the feature scale of the second fusion feature is the same as that of the visual feature or text feature to be recognized; therefore, first, it is necessary to downsample the feature scales of the first fusion feature and the second fusion feature to a preset scale respectively, that is, convert the first fusion feature into a first fusion feature of the preset scale, and convert the second fusion feature into a second fusion feature of the preset scale; then, concatenate the first fusion feature of the preset scale and the second fusion feature of the preset scale into a third fusion feature. Based on this, the present application provides a possible implementation manner, and the S2031 includes the following S1-S2 (not shown in the figure):
[0108] S1: Through a multimodal fusion network, convert the first fusion feature into a first fusion feature of a preset scale; convert the second fusion feature into a second fusion feature of a preset scale.
[0109] S2: Through a multimodal fusion network, perform feature concatenation on the first fusion feature of the preset scale and the second fusion feature of the preset scale to obtain a third fusion feature.
[0110] S2032: Through a multimodal fusion network, perform feature concatenation on the third fusion feature, the visual feature to be recognized, and the text feature to be recognized to obtain a fusion feature to be recognized.
[0111] The S2031-S2032 first effectively and deeply fuses the visual feature to be recognized and the text feature to be recognized in a real sense, and then further effectively and independently retains the visual feature to be recognized and the text feature to be recognized, which can more effectively comprehensively enhance the feature expression ability of the image to be recognized.
[0112] See Figure 5 , Figure 5 is a schematic diagram of obtaining a fusion feature to be recognized based on the visual feature to be recognized and the text feature to be recognized provided by an embodiment of the present application. Among them, the multimodal fusion network includes a Self-Attention module and a Cross-Attention module; input the first concatenated feature of the visual feature to be recognized and the text feature to be recognized into the Self-Attention module in the image recognition model, and obtain the first fusion feature based on the feature fusion of the Self-Attention mechanism; input the visual feature to be recognized and the text feature to be recognized into the Cross-Attention module in the image recognition model, and obtain the second fusion feature based on the feature fusion of the Cross-Attention mechanism; concatenate the first fusion feature and the second fusion feature into a third fusion feature; concatenate the third fusion feature, the visual feature to be recognized, and the text feature to be recognized into a fusion feature to be recognized.
[0113] As an example of S2031 - S2032, based on the above S203 example, first, the first fusion feature E is obtained through a multi - modal fusion network self and the second fusion feature E cross are concatenated into the third fusion feature E self+cross ; then, through the multi - modal fusion network, E self+cross , the visual feature E of the advertisement to be recognized vision and the text feature E of the advertisement to be recognized ocr are concatenated into the fusion feature E of the advertisement to be recognized. Among them, as an example of S1 - S2, the feature scale of E vision is the x scale, the feature scale of E ocr is the y scale, the preset scale is the z scale, the E with the x scale self is converted into E with the z scale self , and the E with the y scale cross is converted into E with the z scale cross ; the E with the z scale self and the E with the z scale cross are concatenated into E self+cross .
[0114] In the embodiment of the present application, for the text information in the image to be recognized in S201 above, considering that the text in the image to be recognized can be recognized through text recognition technology, so as to convert the text information in the image to be recognized into editable text information; therefore, by performing text recognition on the image to be recognized, the text information in the image to be recognized can be obtained. Based on this, the present application provides a possible implementation manner. The steps for obtaining the text information in the image to be recognized in S201 above include S3 (not shown in the figure): performing text recognition on the image to be recognized to obtain the text information in the image to be recognized.
[0115] In practical applications, performing text recognition on the image to be recognized to obtain the text information in the image to be recognized may mean: performing optical character recognition on the image to be recognized to obtain the text information in the image to be recognized.
[0116] This S3 obtains the text information in the image to be recognized through text recognition, laying a foundation for subsequent extraction of the text features of the text information in the image to be recognized, thereby avoiding loss of the text information in the image to be recognized in the feature extraction link.
[0117] As an example of S3, based on the above S201 example, performing optical character recognition on the advertisement image to be recognized to obtain the text information in the advertisement image to be recognized.
[0118] In the embodiments of the present application, for the image recognition model in S201-S204 above, the initial recognition model includes a first extraction network, a second extraction network, a multi-modal fusion network, and a recognition network. The initial recognition model is pre-trained with sample images in the sample image set, text information in the sample images, and preset recognition labels of the sample images to obtain an image recognition model. The image recognition model is used to recognize an image based on the visual information of the image and the text information in the image.
[0119] And training the initial recognition model with sample images in the sample image set, text information in the sample images, and preset recognition labels of the sample images to obtain an image recognition model actually means:
[0120] First, it is necessary to extract not only the visual features of the sample image but also the text features of the text information in the sample image; that is, input the sample image into the first extraction network in the initial recognition model, and the visual features are extracted to obtain the sample visual features; input the text information in the sample image into the second extraction network in the initial recognition model, and the text features are extracted to obtain the sample text features.
[0121] Second, it is necessary to deeply fuse the visual features and text features, not only considering different attention points in the multi-modal features but also aligning the multi-modal semantics of the multi-modal features to be; that is, input the second concatenated features of the sample visual features and the sample text features into the multi-modal fusion network in the initial recognition model, and the feature fusion based on the self-attention mechanism obtains the fourth fusion feature, realizing the consideration of different attention points in the sample visual features and the sample text features; input the sample visual features and the sample text features into the multi-modal fusion network, and the feature fusion based on the cross-attention mechanism obtains the fifth fusion feature, realizing the alignment of the visual semantics of the sample visual features and the text semantics of the sample text features.
[0122] Then, it is also necessary to further independently retain the visual features and text features to obtain the fusion features; that is, concatenate the fourth fusion feature, the fifth fusion feature, the sample visual features, and the sample text features into the sample fusion feature.
[0123] Then, it is necessary to recognize the image; that is, input the sample fusion feature into the recognition network in the initial recognition model, and the image recognition obtains the predicted recognition result of the sample image. The predicted recognition result of the sample image is the recognition result predicted based on the visual information of the sample image and the text information in the sample image.
[0124] Finally, based on the predicted recognition results of the sample images, combined with the pre-annotated recognition labels of the sample images, that is, the preset recognition labels of the sample images, the model parameters of the initial recognition model are adjusted to achieve training, so that the predicted recognition results of the sample images are close to the preset recognition labels of the sample images to complete the model training, and the trained initial recognition model is used as the image recognition model. Therefore, the present application provides a possible implementation manner. The steps for obtaining the image recognition model in S201-S204 above include the following S4-S8 (not shown in the figure):
[0125] S4: For each sample image in the sample image set, through the first extraction network in the initial recognition model, visual features of the sample image are extracted to obtain sample visual features; through the second extraction network in the initial recognition model, text features of the text information in the sample image are extracted to obtain sample text features.
[0126] Among them, the sample visual features refer to the features representing the visual information of the sample image, which are used to indicate the visual semantics of the sample image, and the sample visual features are specifically represented as sample visual vectors; the sample text features refer to the features representing the text information in the sample image, which are used to indicate the text semantics of the sample image, and the sample text features are specifically represented as sample text vectors.
[0127] S5: Through the multi-modal fusion network in the initial recognition model, self-attention mechanism feature fusion is performed on the concatenated features of the sample visual features and the sample text features to obtain a fourth fusion feature; through the multi-modal fusion network in the initial recognition model, cross-attention mechanism feature fusion is performed on the sample visual features and the sample text features to obtain a fifth fusion feature.
[0128] Among them, the fourth fusion feature refers to the fusion feature considering different attention points in the sample visual features and the sample text features, and the fourth fusion feature is specifically represented as a fourth fusion vector; the fifth fusion feature represents the fusion feature that aligns the visual semantics of the sample visual features and the text semantics of the sample text features, and the fifth fusion feature is specifically represented as a fifth fusion vector.
[0129] In the specific implementation of S5, the present application provides a possible implementation manner. The steps for obtaining the fourth fusion feature in S5 include the following S51-S53 (not shown in the figure):
[0130] S51: Feature concatenation is performed on the sample visual features and the sample text features to obtain a second concatenated feature.
[0131] S52: According to the second concatenated feature and multiple third parameter matrices of the self-attention mechanism, multiple fourth mapped features of the second concatenated feature are determined.
[0132] S53: Feature fusion is performed on multiple fourth mapping features according to the self-attention mechanism to obtain a fourth fusion feature.
[0133] The present application provides a possible implementation manner. The steps of obtaining the fifth fusion feature in S5 include the following S54 - S55 (not shown in the figure):
[0134] S54: Determine the fifth mapping feature of the sample visual feature and the sixth mapping feature of the sample text feature according to the sample visual feature, the sample text feature, and multiple fourth parameter matrices of the cross-attention mechanism.
[0135] S55: Feature fusion is performed on the fifth mapping feature and the sixth mapping feature through the cross-attention mechanism to obtain a fifth fusion feature.
[0136] S6: Feature splicing is performed on the fourth fusion feature, the fifth fusion feature, the sample visual feature, and the sample text feature through the multi-modal fusion network in the initial recognition model to obtain a sample fusion feature.
[0137] Among them, the sample fusion feature refers to a fusion feature that considers different attention points in the sample visual feature and the sample text feature, aligns the visual semantics of the sample visual feature and the text semantics of the sample text feature, represents the visual information of the sample image, and represents the text information in the sample image; the sample fusion feature is specifically manifested as a sample fusion vector.
[0138] When specifically implementing S6, the present application provides a possible implementation manner. S6 includes S61 - S62 (not shown in the figure):
[0139] S61: Feature splicing is performed on the fourth fusion feature and the fifth fusion feature through the multi-modal fusion network to obtain a sixth fusion feature.
[0140] Among them, the sixth fusion feature refers to a fusion feature that considers different attention points in the sample visual feature and the sample text feature and aligns the visual semantics of the sample visual feature and the text semantics of the sample text feature.
[0141] When specifically implementing S61, the present application provides a possible implementation manner. S61 includes the following S61a - S61b (not shown in the figure):
[0142] S61a: Convert the fourth fusion feature to a fourth fusion feature of a preset scale; convert the fifth fusion feature to a fifth fusion feature of a preset scale.
[0143] S61b: Perform feature splicing on the fourth fusion feature of the preset scale and the fifth fusion feature of the preset scale to obtain a sixth fusion feature.
[0144] S62: Feature splicing is performed on the sixth fusion feature, the sample visual feature, and the sample text feature through a multi-modal fusion network to obtain a sample fusion feature.
[0145] S7: Image recognition is performed on the sample fusion feature through the recognition network in the initial recognition model to obtain a predicted recognition result of the sample image.
[0146] S8: According to the predicted recognition result and the preset recognition label of the sample image, the model parameters of the initial recognition model are adjusted to obtain an image recognition model.
[0147] Through S4 - S8, the initial recognition model predicts the recognition result of the sample image based on the visual information of the sample image and the text information in the sample image, obtains the predicted recognition result of the sample image, and adjusts the model parameters of the initial recognition model in the training direction that makes the predicted recognition result of the sample image close to the preset recognition label of the sample image, quickly, effectively, and accurately learns the correlation between the visual information of the sample image and the text information in the sample image and the preset recognition label of the sample image, completes the model training to obtain an image recognition model; provides a recognition model for quickly, effectively, and accurately recognizing images according to the visual information of the image and the text information in the image subsequently.
[0148] As an example of S4 - S8, based on the above example of S201 - S204, the sample image is a sample advertisement image; First, the computer device inputs the sample advertisement image into the first extraction network, the SwinTransformer model, in the initial recognition model, and visually extracts the sample visual feature as the sample advertisement visual feature; The text information in the sample advertisement image is input into the second extraction network, the Bert model, in the initial recognition model, and the text feature is extracted to obtain the sample text feature as the sample advertisement text feature. Secondly, the computer device inputs the second concatenated feature of the sample advertisement visual feature and the sample advertisement text feature into the Self - Attention module in the multi - modal fusion network, and the feature fusion based on the Self - Attention mechanism obtains the fourth fusion feature; The sample advertisement visual feature and the sample advertisement text feature are input into the Cross - Attention module in the multi - modal fusion network, and the feature fusion based on the Cross - Attention mechanism obtains the fifth fusion feature. Then, the computer device concatenates the fourth fusion feature, the fifth fusion feature, the sample advertisement visual feature, and the sample advertisement text feature into the sample fusion feature, that is, the sample advertisement fusion feature. Then, the computer device inputs the sample advertisement fusion feature into the recognition network SoftMax in the initial recognition model, and the image recognition obtains the predicted recognition result of the sample advertisement image. Finally, based on the predicted recognition result of the sample advertisement image, the computer device combines the preset recognition label of the sample advertisement image, adjusts the model parameters of the initial recognition model to achieve training, so that the predicted recognition result of the sample advertisement image is close to the preset recognition label of the sample advertisement image to complete the model training, and the trained initial recognition model is used as the image recognition model.
[0149] Among them, when specifically implementing S8, considering that there are sufficient sample images in the sample image set, in order to improve the recognition performance of the image recognition model obtained by training, the first extraction network, the second extraction network, the multi-modal fusion network, and the recognition network in the initial recognition model can all be trained from scratch; that is, the first extraction network, the second extraction network, the multi-modal fusion network, and the recognition network in the initial recognition model adopt the same training hyperparameters; therefore, when the number of images of multiple sample images in the sample image set is greater than or equal to a preset number, it indicates that there are sufficient sample images in the sample image set, and the network parameters of the first extraction network, the second extraction network, the multi-modal fusion network, and the recognition network in the initial recognition model are adjusted according to the first training hyperparameters to complete the model training, and the trained initial recognition model is used as the image recognition model. Based on this, the present application provides a possible implementation manner where the number of images of multiple sample images in the sample image set is greater than or equal to a preset number, and this S8 includes S81 (not shown in the figure): adjusting the network parameters of the first extraction network, the second extraction network, the multi-modal fusion network, and the recognition network in the initial recognition model according to the first training hyperparameters to obtain the image recognition model.
[0150] When there are sufficient sample images in the sample image set, S81 starts training the first extraction network, the second extraction network, the multi-modal fusion network, and the recognition network in the initial recognition model from scratch using the same training hyperparameters, which can improve the recognition performance of the image recognition model obtained by training; it provides a recognition model with better recognition performance for subsequent quickly, effectively, and accurately recognizing images based on the visual information of the images and the text information in the images.
[0151] Among them, when specifically implementing S8, considering the situation of insufficient sample images in the sample image set, in order to ensure the recognition performance of the image recognition model obtained by training, the first extraction network and the second extraction network in the initial recognition model can be initialized with a pre-trained model or an open-source model. However, the multi-modal fusion network and the recognition network in the initial recognition model need to be trained from scratch. That is, the first extraction network and the second extraction network in the initial recognition model adopt training hyperparameters with relatively small training performance, while the multi-modal fusion network and the recognition network in the initial recognition model adopt training hyperparameters with relatively large training performance. Therefore, when the number of images of multiple sample images in the sample image set is less than the preset number, it indicates that the sample images in the sample image set are insufficient. Determine the second training hyperparameter whose training performance is greater than that of the first training hyperparameter, adjust the network parameters of the first extraction network and the second extraction network in the initial recognition model according to the first training hyperparameter, and adjust the network parameters of the multi-modal fusion network and the recognition network in the initial recognition model according to the second training hyperparameter to complete the model training, and use the trained initial recognition model as the image recognition model. Based on this, the present application provides a possible implementation manner. The number of images of multiple sample images in the sample image set is less than the preset number. This S8 includes S82 (not shown in the figure): Adjust the parameters of the first extraction network and the second extraction network in the initial recognition model according to the first training hyperparameter, and adjust the network parameters of the multi-modal fusion network and the recognition network in the initial recognition model according to the second training hyperparameter to obtain the image recognition model; the training performance of the second training hyperparameter is greater than that of the first training hyperparameter.
[0152] When the sample images in the sample image set are insufficient, S82 fine-tunes and trains the first extraction network and the second extraction network in the initial recognition model with training hyperparameters with relatively small training performance, and trains the multi-modal fusion network and the recognition network in the initial recognition model from scratch with training hyperparameters with relatively large training performance, which can ensure the recognition performance of the obtained image recognition model; it provides a recognition model with recognition performance for subsequent quickly, effectively, and accurately recognizing images based on the visual information of the images and the text information in the images.
[0153] In addition, in the embodiments of the present application, in order to enrich the number of sample images in the sample image set, the sample image data can also be enhanced to an enhanced image, and the enhanced image is added to the sample image set as a new sample image, thereby increasing the number of sample images in the sample image set. Therefore, the present application provides a possible implementation manner. The method further includes the following S9-S10 (not shown in the figure):
[0154] S9: Perform data enhancement on the sample image to obtain an enhanced image.
[0155] Among them, the implementation methods of data augmentation include data augmentation based on geometric transformation and data augmentation based on color space transformation, etc. Data augmentation based on geometric transformation includes flipping, rotating, cropping, translating, and adding noise, etc. Data augmentation based on color space transformation includes adjusting brightness, adjusting contrast, adjusting saturation, channel separation, and grayscale conversion, etc.
[0156] S10: Determine the enhanced image as a newly added sample image and add it to the sample image set.
[0157] S9 - S10 enrich the number of sample images in the sample image set through data - augmented sample images, providing richer training data for training the initial recognition model to obtain an image recognition model.
[0158] As an example of S9 - S10, based on the above - mentioned example of S4 - S8, the sample advertisement image data is augmented into an enhanced advertisement image, and the enhanced advertisement image is added to the sample image set as a newly added sample advertisement image.
[0159] In addition, in the embodiments of the present application, since the image recognition effect of the image to be recognized is improved after performing the above - mentioned S201 - S204, that is, the accuracy rate of the image recognition result of the image to be recognized is higher. In the image review scenario, by reviewing the image to be recognized according to the image recognition result of the image to be recognized, an image review result with a higher accuracy rate of the image to be recognized can be obtained. Therefore, the present application provides a possible implementation manner. After the above - mentioned S201 - S204, the method further includes S11 (not shown in the figure): Image - review the image to be recognized according to the image recognition result to obtain the image review result of the image to be recognized.
[0160] Based on the higher accuracy rate of the image recognition result of the image to be recognized, S11 realizes image review for the image recognition result of the image to be recognized, which can improve the image review quality and image review effect of the image to be recognized.
[0161] As an example of S11, based on the above - mentioned example of S201 - S204, review the image to be recognized advertisement according to the advertisement recognition result of the image to be recognized advertisement to obtain the advertisement review result of the image to be recognized advertisement.
[0162] In summary, referring to Figure 6 , Figure 6 is a flowchart of an advertisement image recognition method provided by the embodiments of the present application. The method includes:
[0163] S601: Through the first extraction network in the image recognition model, extract visual features from the image to be recognized advertisement to obtain the visual features of the image to be recognized advertisement.
[0164] S602: Extract text features from the text information in the advertisement image to be recognized through the second extraction network in the image recognition model, and obtain the advertisement text features to be recognized.
[0165] S603: Perform feature splicing on the visual features of the advertisement to be recognized and the text features of the advertisement to be recognized, and obtain the first spliced feature.
[0166] S604: Through the multi-modal fusion network in the image recognition model, determine multiple first mapping features of the first spliced feature according to the first spliced feature and multiple first parameter matrices of the self-attention mechanism.
[0167] S605: Through the multi-modal fusion network, perform feature fusion on multiple first mapping features according to the self-attention mechanism, and obtain the first fusion feature.
[0168] S606: Through the multi-modal fusion network, determine the second mapping feature of the visual features of the advertisement to be recognized and the third mapping feature of the text features of the advertisement to be recognized according to the visual features of the advertisement to be recognized, the text features of the advertisement to be recognized, and multiple second parameter matrices of the cross-attention mechanism.
[0169] S607: Through the multi-modal fusion network, perform feature fusion on the second mapping feature and the third mapping feature according to the cross-attention mechanism, and obtain the second fusion feature.
[0170] S608: Through the multi-modal fusion network, perform feature splicing on the first fusion feature and the second fusion feature, and obtain the third fusion feature.
[0171] S609: Through the multi-modal fusion network, perform feature splicing on the third fusion feature, the visual features of the advertisement to be recognized, and the text features of the advertisement to be recognized, and obtain the fused features of the advertisement to be recognized.
[0172] S610: Through the recognition network in the image recognition model, perform image recognition on the fused features of the advertisement to be recognized, and obtain the advertisement recognition result of the advertisement image to be recognized.
[0173] It can be seen from the above technical solutions that on the basis of extracting the visual features of the advertisement image, the text features of the text information in the advertisement image are further extracted; different attention points in the multi-modal features are considered through the self-attention mechanism, and the multi-modal semantics of the multi-modal features are aligned through the cross-attention mechanism. The multi-modal features complement and promote each other to effectively deeply fuse the multi-modal features and effectively independently retain the multi-modal features, enhancing the feature expression ability of the advertisement image in multiple aspects; thereby improving the recognition accuracy and recognition effect of the advertisement image.
[0174] The advertising image recognition method provided by the above embodiment realizes advertising image recognition, which can be used for the review of advertising images, specifically for the cleaning and filtering of advertising images, effectively supporting the cleaning and filtering projects of advertising images in existing news articles, and greatly reducing the review cost of advertising images.
[0175] It should be noted that, based on the implementation manners provided in the above aspects of the present application, further combinations can be made to provide more implementation manners.
[0176] Based on Figure 2 the image recognition method provided in the corresponding embodiment, the embodiment of the present application further provides an image recognition device. Refer to Figure 7 , Figure 7 which is a structural diagram of an image recognition device provided by the embodiment of the present application. The image recognition device 700 includes: an extraction unit 701, a fusion unit 702, a splicing unit 703, and an identification unit 704;
[0177] The extraction unit 701 is configured to extract visual features of the image to be recognized through the first extraction network in the image recognition model to obtain the visual features to be recognized; and extract text features of the text information in the image to be recognized through the second extraction network in the image recognition model to obtain the text features to be recognized.
[0178] The fusion unit 702 is configured to perform feature fusion of the self-attention mechanism on the first splicing feature of the visual features to be recognized and the text features to be recognized through the multi-modal fusion network in the image recognition model to obtain the first fusion feature; and perform feature fusion of the cross-attention mechanism on the visual features to be recognized and the text features to be recognized through the multi-modal fusion network to obtain the second fusion feature.
[0179] The splicing unit 703 is configured to perform feature splicing on the first fusion feature, the second fusion feature, the visual features to be recognized, and the text features to be recognized through the multi-modal fusion network to obtain the fusion features to be recognized.
[0180] The identification unit 704 is configured to perform image recognition on the fusion features to be recognized through the recognition network in the image recognition model to obtain the image recognition result of the image to be recognized.
[0181] In a possible implementation manner, the fusion unit 702 is specifically configured to:
[0182] Perform feature splicing on the visual features to be recognized and the text features to be recognized to obtain the first splicing feature;
[0183] Determine multiple first mapping features of the first splicing feature through the multi-modal fusion network according to the first splicing feature and multiple first parameter matrices of the self-attention mechanism;
[0184] Through a multi-modal fusion network, feature fusion is performed on multiple first mapping features according to the self-attention mechanism to obtain a first fusion feature.
[0185] In a possible implementation, the fusion unit 702 is specifically configured to:
[0186] Through a multi-modal fusion network, according to the visual feature to be recognized, the text feature to be recognized, and multiple second parameter matrices of the cross-attention mechanism, determine the second mapping feature of the visual feature to be recognized and the third mapping feature of the text feature to be recognized;
[0187] Through a multi-modal fusion network, perform feature fusion on the second mapping feature and the third mapping feature according to the cross-attention mechanism to obtain a second fusion feature.
[0188] In a possible implementation, the fusion unit 702 is specifically configured to:
[0189] Through a multi-modal fusion network, perform feature splicing on the first fusion feature and the second fusion feature to obtain a third fusion feature;
[0190] Through a multi-modal fusion network, perform feature splicing on the third fusion feature, the visual feature to be recognized, and the text feature to be recognized to obtain the fusion feature to be recognized.
[0191] In a possible implementation, the fusion unit 702 is specifically configured to:
[0192] Through a multi-modal fusion network, convert the first fusion feature into a first fusion feature of a preset scale; convert the second fusion feature into a second fusion feature of a preset scale;
[0193] Through a multi-modal fusion network, perform feature splicing on the first fusion feature of the preset scale and the second fusion feature of the preset scale to obtain a third fusion feature.
[0194] In a possible implementation, the recognition unit 704 is further configured to:
[0195] Perform text recognition on the image to be recognized to obtain the text information in the image to be recognized.
[0196] In a possible implementation, the device further includes: an audit unit;
[0197] The audit unit is configured to perform image audit on the image to be recognized according to the image recognition result to obtain the image audit result of the image to be recognized.
[0198] In a possible implementation, the device further includes: an adjustment unit;
[0199] The extraction unit 701 is further configured to, for each sample image in the sample image set, extract visual features of the sample image through the first extraction network in the initial recognition model to obtain sample visual features; extract text features of the text information in the sample image through the second extraction network in the initial recognition model to obtain sample text features;
[0200] The fusion unit 702 is further configured to perform feature fusion of the self-attention mechanism on the concatenated features of the sample visual features and the sample text features through the multi-modal fusion network in the initial recognition model to obtain a fourth fusion feature; perform feature fusion of the cross-attention mechanism on the sample visual features and the sample text features through the multi-modal fusion network in the initial recognition model to obtain a fifth fusion feature;
[0201] The concatenation unit 703 is further configured to perform feature concatenation on the fourth fusion feature, the fifth fusion feature, the sample visual features, and the sample text features through the multi-modal fusion network in the initial recognition model to obtain sample fusion features;
[0202] The recognition unit 704 is further configured to perform image recognition on the sample fusion features through the recognition network in the initial recognition model to obtain a predicted recognition result of the sample image;
[0203] The adjustment unit is configured to adjust the model parameters of the initial recognition model according to the predicted recognition result and the preset recognition label of the sample image to obtain an image recognition model.
[0204] In a possible implementation manner, the number of images of multiple sample images in the sample image set is greater than or equal to a preset number, and the adjustment unit is configured to:
[0205] Adjust the network parameters of the first extraction network, the second extraction network, and the multi-modal fusion network in the initial recognition model according to the first training hyperparameter to obtain an image recognition model.
[0206] In a possible implementation manner, the number of images of multiple sample images in the sample image set is less than the preset number, and the adjustment unit is configured to:
[0207] Adjust the first extraction network and the second extraction network in the initial recognition model according to the first training hyperparameter, and adjust the network parameters of the multi-modal fusion network in the initial recognition model according to the second training hyperparameter to obtain an image recognition model; the training performance of the second training hyperparameter is greater than the training performance of the first training hyperparameter.
[0208] In a possible implementation manner, the device further includes: an enhancement unit and an addition unit;
[0209] The enhancement unit is configured to perform data enhancement on the sample image to obtain an enhanced image;
[0210] An adding unit is configured to determine the enhanced image as a newly added sample image and add it to the sample image set.
[0211] As can be seen from the above technical solutions, the image recognition device includes an extraction unit, a fusion unit, a splicing unit, and an identification unit. The extraction unit inputs the image to be recognized into the first extraction network in the image recognition model to extract the visual features to be recognized; inputs the text information in the image to be recognized into the second extraction network in the image recognition model to extract the text features to be recognized; this unit not only extracts the visual features of the image to be recognized, but also extracts the text features of the text information in the image to be recognized. The fusion unit inputs the first splicing feature of the visual features to be recognized and the text features to be recognized into the multi-modal fusion network in the image recognition model to obtain the first fusion feature based on the feature fusion of the self-attention mechanism; inputs the visual features to be recognized and the text features to be recognized into the multi-modal fusion network to obtain the second fusion feature based on the feature fusion of the cross-attention mechanism; the splicing unit splices the first fusion feature, the second fusion feature, the visual features to be recognized, and the text features to be recognized into the fusion feature to be recognized; this unit not only effectively fuses the visual features to be recognized and the text features to be recognized deeply, that is, considering different attention points in the visual features to be recognized and the text features to be recognized through the self-attention mechanism, and aligning the visual semantics of the visual features to be recognized and the text semantics of the text features to be recognized through the cross-attention mechanism, realizing the mutual complementarity and mutual promotion of the visual features to be recognized and the text features to be recognized, but also effectively retains the visual features to be recognized and the text features to be recognized independently; so as to effectively comprehensively enhance the feature expression ability of the image to be recognized. The identification unit inputs the fusion feature to be recognized into the identification network in the image recognition model to obtain the image recognition result of the image to be recognized through image recognition; this unit identifies the fusion feature to be recognized with stronger feature expression ability, making the recognition accuracy and recognition effect of the image to be recognized higher and better. Based on this, the device further extracts the text features of the text information in the image on the basis of extracting the visual features of the image; considering different attention points in the multi-modal features through the self-attention mechanism, aligning the multi-modal semantics of the multi-modal features through the cross-attention mechanism, the multi-modal features complement and promote each other, so as to effectively fuse the multi-modal features deeply and effectively retain the multi-modal features independently, enhancing the feature expression ability of the image in many aspects; thereby improving the recognition accuracy and recognition effect of the image.
[0212] An embodiment of this application further provides a computer device, which may be a server, see Figure 8 , Figure 8The structural diagram of a server provided by an embodiment of this application. Server 800 may vary significantly due to configuration or performance differences, and may include one or more processors, such as CPU 822, as well as a memory 832, and one or more storage media 830 (such as one or more mass storage devices) for storing application programs 842 or data 844. Among them, the memory 832 and the storage media 830 may be transient storage or persistent storage. The programs stored in the storage media 830 may include one or more modules (not shown in the figure), and each module may include a series of instruction operations for the server. Further, the central processing unit 822 may be configured to communicate with the storage media 830 and execute a series of instruction operations in the storage media 830 on the server 800.
[0213] Server 800 may also include one or more power supplies 826, one or more wired or wireless network interfaces 850, one or more input / output interfaces 858, and / or one or more operating systems 841, such as Windows Server TM , Mac OS X TM , Unix TM , Linux TM , FreeBSD TM and so on.
[0214] In this embodiment, the method provided in various alternative implementation manners of the above embodiment can be executed by the central processing unit 822 in the server 800.
[0215] The computer device provided by the embodiment of this application may also be a terminal. Refer to Figure 9 , Figure 9 The structural diagram of a terminal provided by an embodiment of this application. Taking a smart phone as an example of the terminal, the smart phone includes components such as a Radio Frequency (RF) circuit 910, a memory 920, an input unit 930, a display unit 940, a sensor 950, an audio circuit 960, a Wireless Fidelity (WiFi) module 970, a processor 980, and a power supply 990. The input unit 930 may include a touch panel 931 and other input devices 932. The display unit 940 may include a display panel 941. The audio circuit 960 may include a speaker 961 and a microphone 962. Those skilled in the art can understand that Figure 9 the smart phone structure shown in
[0216] The memory 920 can be used to store software programs and modules. The processor 980 executes various functional applications and data processing of the smart phone by running the software programs and modules stored in the memory 920. The memory 920 mainly includes a program storage area and a data storage area. Among them, the program storage area can store the operating system, application programs required for at least one function (such as the sound playback function, the image playback function, etc.); the data storage area can store the data created according to the use of the smart phone (such as audio data, phone book, etc.). In addition, the memory 920 can include high-speed random access memory, and can also include non-volatile memory, such as at least one magnetic disk storage device, flash memory device, or other volatile solid-state storage devices.
[0217] The processor 980 is the control center of the smart phone, connects various parts of the entire smart phone using various interfaces and lines, and executes various functions of the smart phone and processes data by running or executing the software programs and / or modules stored in the memory 920, and calling the data stored in the memory 920. Optionally, the processor 980 can include one or more processing units; preferably, the processor 980 can integrate an application processor and a modem processor. Among them, the application processor mainly processes the operating system, user interface, application programs, etc., and the modem processor mainly processes wireless communication. It can be understood that the above-mentioned modem processor may not be integrated into the processor 980.
[0218] In this embodiment, the processor 980 in the smart phone can execute the methods provided in various optional implementation manners of the above embodiment.
[0219] According to one aspect of the present application, there is provided a computer-readable storage medium for storing a computer program. When the computer program runs on a computer device, the computer device is caused to execute the methods provided in various optional implementation manners of the above embodiment.
[0220] According to one aspect of the present application, there is provided a computer program product. The computer program product includes a computer program, and the computer program is stored in a computer-readable storage medium. The processor of the computer device reads the computer program from the computer-readable storage medium, and the processor executes the computer program, so that the computer device executes the methods provided in various optional implementation manners of the above embodiment.
[0221] The descriptions of the processes or structures corresponding to the above respective drawings have their own emphases. For parts not detailed in a certain process or structure, reference can be made to the relevant descriptions of other processes or structures.
[0222] In the description of the present application and the above-mentioned drawings, terms such as "first" and "second" are used to distinguish similar objects and do not necessarily describe a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances, so that the embodiments of the present application described herein can be implemented in an order different from those illustrated or described herein. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device that includes a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.
[0223] In several embodiments provided by the present application, it should be understood that the disclosed systems, devices and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the units is only a logical function division, and there may be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection between each other can be through some interfaces, and the indirect coupling or communication connection of the device or unit can be in electrical, mechanical or other forms.
[0224] The units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0225] In addition, in each embodiment of the present application, the functional units can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above-mentioned integrated units can be implemented in the form of hardware or in the form of software functional units.
[0226] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to enable a computer device to execute all or part of the steps of the methods described in various embodiments of this application. The foregoing storage medium includes: various media that can store computer programs, such as USB flash drives, mobile hard disks, read-only memory (ROM), RAM, magnetic disks, or optical discs.
[0227] As described above, the above embodiments are only used to illustrate the technical solutions of this application and are not intended to limit them; although this application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments or perform equivalent replacements for some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of various embodiments of this application.
Claims
1. An image recognition method, characterized in that, The method includes: Performing visual feature extraction on the image to be recognized through a first extraction network in the image recognition model to obtain the visual feature to be recognized; performing text feature extraction on the text information in the image to be recognized through a second extraction network in the image recognition model to obtain the text feature to be recognized; Performing feature fusion of the self-attention mechanism on the first concatenated feature of the visual feature to be recognized and the text feature to be recognized through a multi-modal fusion network in the image recognition model to obtain a first fusion feature; performing feature fusion of the cross-attention mechanism on the visual feature to be recognized and the text feature to be recognized through the multi-modal fusion network to obtain a second fusion feature; Performing feature concatenation on the first fusion feature, the second fusion feature, the visual feature to be recognized, and the text feature to be recognized through the multi-modal fusion network to obtain a fusion feature to be recognized; Performing image recognition on the fusion feature to be recognized through a recognition network in the image recognition model to obtain an image recognition result of the image to be recognized.
2. The method according to claim 1, wherein The step of performing feature fusion of the self-attention mechanism on the first concatenated feature of the visual feature to be recognized and the text feature to be recognized through a multi-modal fusion network in the image recognition model to obtain a first fusion feature includes: Performing feature concatenation on the visual feature to be recognized and the text feature to be recognized to obtain the first concatenated feature; Determining multiple first mapped features of the first concatenated feature through the multi-modal fusion network according to the first concatenated feature and multiple first parameter matrices of the self-attention mechanism; Performing feature fusion on the multiple first mapped features according to the self-attention mechanism through the multi-modal fusion network to obtain the first fusion feature.
3. The method according to claim 1, wherein The step of performing feature fusion of the cross-attention mechanism on the visual feature to be recognized and the text feature to be recognized through the multi-modal fusion network to obtain a second fusion feature includes: Determining a second mapped feature of the visual feature to be recognized and a third mapped feature of the text feature to be recognized through the multi-modal fusion network according to the visual feature to be recognized, the text feature to be recognized, and multiple second parameter matrices of the cross-attention mechanism; Performing feature fusion on the second mapped feature and the third mapped feature according to the cross-attention mechanism through the multi-modal fusion network to obtain the second fusion feature.
4. The method according to claim 1, wherein The step of performing feature concatenation on the first fusion feature, the second fusion feature, the visual feature to be recognized, and the text feature to be recognized through the multi-modal fusion network to obtain a fusion feature to be recognized includes: Performing feature concatenation on the first fusion feature and the second fusion feature through the multi-modal fusion network to obtain a third fusion feature; Performing feature concatenation on the third fusion feature, the visual feature to be recognized, and the text feature to be recognized through the multi-modal fusion network to obtain the fusion feature to be recognized.
5. The method according to claim 4, characterized in that Performing feature concatenation on the first fusion feature and the second fusion feature through the multimodal fusion network to obtain a third fusion feature, including: Through the multimodal fusion network, converting the first fusion feature into a first fusion feature of a preset scale; converting the second fusion feature into a second fusion feature of the preset scale; Through the multimodal fusion network, performing feature concatenation on the first fusion feature of the preset scale and the second fusion feature of the preset scale to obtain the third fusion feature.
6. The method according to claim 1, wherein The steps for obtaining the text information in the image to be recognized include: Performing text recognition on the image to be recognized to obtain the text information in the image to be recognized.
7. The method according to claim 1, wherein The method further includes: Performing image review on the image to be recognized according to the image recognition result to obtain the image review result of the image to be recognized.
8. The method according to any one of claims 1-7, characterized in that, The steps for obtaining the image recognition model include: For each sample image in the sample image set, through the first extraction network in the initial recognition model, performing visual feature extraction on the sample image to obtain sample visual features; through the second extraction network in the initial recognition model, performing text feature extraction on the text information in the sample image to obtain sample text features; Through the multimodal fusion network in the initial recognition model, performing feature fusion of the self-attention mechanism on the concatenated features of the sample visual features and the sample text features to obtain a fourth fusion feature; through the multimodal fusion network in the initial recognition model, performing feature fusion of the cross-attention mechanism on the sample visual features and the sample text features to obtain a fifth fusion feature; Through the multimodal fusion network in the initial recognition model, performing feature concatenation on the fourth fusion feature, the fifth fusion feature, the sample visual features and the sample text features to obtain a sample fusion feature; Through the recognition network in the initial recognition model, performing image recognition on the sample fusion feature to obtain the predicted recognition result of the sample image; According to the predicted recognition result and the preset recognition label of the sample image, adjusting the model parameters of the initial recognition model to obtain the image recognition model.
9. The method according to claim 8, wherein The number of images of multiple sample images in the sample image set is greater than or equal to a preset number. The adjusting the model parameters of the initial recognition model according to the predicted recognition result and the preset recognition label of the sample image to obtain the image recognition model includes: Adjusting the network parameters of the first extraction network, the second extraction network, the multimodal fusion network and the recognition network in the initial recognition model according to the first training hyperparameter to obtain the image recognition model.
10. The method according to claim 8, characterized in that, The number of images of multiple sample images in the sample image set is less than the preset number. The adjusting the model parameters of the initial recognition model according to the predicted recognition result and the preset recognition label of the sample image to obtain the image recognition model includes: Adjust the parameters of the first extraction network and the second extraction network in the initial recognition model according to the first training hyperparameters, and adjust the network parameters of the multi-modal fusion network and the recognition network in the initial recognition model according to the second training hyperparameters to obtain the image recognition model; the training performance of the second training hyperparameters is greater than that of the first training hyperparameters.
11. The method according to claim 8, wherein The method further includes: Perform data augmentation on the sample image to obtain an augmented image; Determine the augmented image as a newly added sample image and add it to the sample image set.
12. An image recognition device, characterized in that, The device includes: an extraction unit, a fusion unit, a splicing unit, and a recognition unit; The extraction unit is configured to extract visual features of the image to be recognized through the first extraction network in the image recognition model to obtain visual features to be recognized; extract text features of the text information in the image to be recognized through the second extraction network in the image recognition model to obtain text features to be recognized; The fusion unit is configured to perform feature fusion of the self-attention mechanism on the first splicing feature of the visual features to be recognized and the text features to be recognized through the multi-modal fusion network in the image recognition model to obtain a first fusion feature; perform feature fusion of the cross-attention mechanism on the visual features to be recognized and the text features to be recognized through the multi-modal fusion network to obtain a second fusion feature; The splicing unit is configured to perform feature splicing on the first fusion feature, the second fusion feature, the visual features to be recognized, and the text features to be recognized through the multi-modal fusion network to obtain a fused feature to be recognized; The recognition unit is configured to perform image recognition on the fused feature to be recognized through the recognition network in the image recognition model to obtain an image recognition result of the image to be recognized.
13. A computer device, characterized in that, The computer device includes a processor and a memory: The memory is used to store a computer program and transmit the computer program to the processor; The processor is configured to execute the method according to any one of claims 1-11 according to the instructions in the computer program.
14. A computer-readable storage medium, characterized in that, The computer-readable storage medium is used to store a computer program, and when the computer program runs on a computer device, the computer device is caused to execute the method according to any one of claims 1-11.
15. A computer program product, comprising a computer program, characterized in that, When the computer program runs on a computer device, the computer device is caused to execute the method according to any one of claims 1-11.
Citation Information
Cited By
Alzheimer's disease preclinical risk quantitative evaluation method and system
CN121460166A
Intelligent invoice identification method and system based on AI technology
CN121545161A