Image recognition method and device, equipment, medium and product

By extracting and fusing features from caption images using an emotion recognition model, the problem of inaccurate parsing of multimodal information in images when no manual text input by the user is required is solved, achieving efficient and accurate image recognition and emotion recognition.

CN121838112APending Publication Date: 2026-04-10INDUSTRIAL AND COMMERCIAL BANK OF CHINA
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-12-15
Publication Date
2026-04-10

AI Technical Summary

Technical Problem

Existing technologies cannot efficiently analyze multimodal information in images without requiring users to manually input text, resulting in inaccurate image recognition results and limiting the flexibility and intelligence of AI dialogue assistants.

Method used

The emotion recognition model extracts features from the caption image, obtaining frequency and spatial domain features. Then, a multi-domain attention fusion unit is used to fuse the features, and the pixel image is output to obtain the recognition result.

Benefits of technology

It improves the accuracy of image recognition, especially the accuracy of emotion recognition in images, and enhances the flexibility and intelligence of the AI ​​dialogue assistant.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121838112A_ABST
    Figure CN121838112A_ABST
Patent Text Reader

Abstract

The invention provides an image recognition method, device and equipment, a medium and a product, and relates to the field of financial science and technology or other related fields. Comprising the following steps: acquiring a to-be-identified subtitle image; inputting the subtitle image into a coding unit of an emotion recognition model, and performing feature extraction to obtain frequency domain features and spatial domain features; the frequency domain features and the spatial domain features represent text description features of the subtitle images; according to the frequency domain features and the spatial domain features, feature fusion is carried out based on a multi-domain attention fusion unit of an emotion recognition model, and fusion enhanced features are obtained; according to the fusion enhancement feature, a decoding unit based on an emotion recognition model outputs a pixel image of the subtitle image; and obtaining an emotion recognition result of the subtitle image according to the pixel image. According to the scheme of the invention, the accuracy of image recognition is improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the field of financial technology or other related fields, and in particular, relates to an image recognition method and device, equipment, medium and product. BACKGROUND

[0002] In professional fields such as finance and medicine, users may upload pictures with professional terms for consultation, and Artificial Intelligence (AI) needs to accurately extract key information and provide professional advice combined with domain knowledge.

[0003] The current solution cannot efficiently analyze multi-modal information (such as text, emotion, and scene) in the picture without the user manually inputting text, and the result of picture recognition is not accurate. SUMMARY

[0004] The present application provides an image recognition method, device, equipment, medium and product to improve the accuracy of image recognition.

[0005] In a first aspect, the present application provides an image recognition method, comprising:

[0006] obtaining a subtitle image to be recognized;

[0007] inputting the subtitle image into an encoding unit of an emotion recognition model to perform feature extraction, obtaining frequency domain features and spatial domain features; the frequency domain features and spatial domain features represent text description features of the subtitle image;

[0008] performing feature fusion based on a multi-domain attention fusion unit of the emotion recognition model according to the frequency domain features and spatial domain features, obtaining fusion enhanced features;

[0009] outputting a pixel image of the subtitle image based on a decoding unit of the emotion recognition model according to the fusion enhanced features; and obtaining a recognition result of the subtitle image according to the pixel image.

[0010] In a second aspect, the present application provides an image recognition device, comprising:

[0011] an obtaining module configured to obtain a subtitle image to be recognized;

[0012] an extracting module configured to input the subtitle image into an encoding unit of an emotion recognition model to perform feature extraction, obtaining frequency domain features and spatial domain features; the frequency domain features and spatial domain features represent text description features of the subtitle image;

[0013] a fusion module configured to perform feature fusion based on a multi-domain attention fusion unit of the emotion recognition model according to the frequency domain feature and the spatial domain feature, to obtain a fusion enhanced feature;

[0014] a processing module configured to output a pixel image of the subtitle image based on a decoding unit of the emotion recognition model according to the fusion enhanced feature, and obtain a recognition result of the subtitle image according to the pixel image.

[0015] In a third aspect, an electronic device is provided, including a memory and a processor.

[0016] The memory stores computer execution instructions.

[0017] The processor executes the computer execution instructions stored in the memory, so that the processor performs the first aspect and / or various possible implementation manners of the first aspect.

[0018] In a fourth aspect, a computer readable storage medium is provided, which stores computer execution instructions. When the processor executes the computer execution instructions, the computer execution instructions are used to implement the first aspect and / or various possible implementation manners of the first aspect.

[0019] In a fifth aspect, a computer program product is provided, which includes a computer program. When the processor executes the computer program, the computer program implements the first aspect and / or various possible implementation manners of the first aspect.

[0020] The image recognition method, device, equipment, medium and product provided by the present application obtain a subtitle image to be recognized; input the subtitle image into an encoding unit of an emotion recognition model to perform feature extraction, to obtain a frequency domain feature and a spatial domain feature; the frequency domain feature and the spatial domain feature represent text description features of the subtitle image; perform feature fusion based on a multi-domain attention fusion unit of the emotion recognition model according to the frequency domain feature and the spatial domain feature, to obtain a fusion enhanced feature; output a pixel image of the subtitle image based on a decoding unit of the emotion recognition model according to the fusion enhanced feature; obtain a recognition result of the subtitle image according to the pixel image; the scheme of the present application performs feature extraction on the input subtitle image to be recognized based on the emotion recognition model, to obtain the frequency domain and spatial domain features of the text of the subtitle image, then inputs the frequency domain and spatial domain features of the text of the subtitle image into the attention fusion unit to perform feature fusion, to obtain a fusion feature; the fusion feature is a feature representing the emotion of the subtitle image to be recognized, and finally a recognition result is obtained by outputting a pixel feature through the decoding unit, to improve the accuracy of image recognition. BRIEF DESCRIPTION OF DRAWINGS

[0021] The accompanying drawings, which are incorporated herein and constitute part of this specification, illustrate embodiments consistent with the application and serve to explain the principles

[0022] Figure 1 An exemplary flowchart of an image recognition method is shown;

[0023] Figure 2 An exemplary flowchart of a training process of an emotion recognition model is shown;

[0024] Figure 3 An exemplary structural diagram of an image recognition device is shown;

[0025] Figure 4 An exemplary structural diagram of an electronic device is shown.

[0026] The specific embodiments of the application have been shown by way of example in the above-described figures, and will be described in more detail hereafter. These figures and written description are not meant to limit the scope of the inventive concept in any way, but merely to illustrate the inventive concept to those skilled in the art by reference to a particular embodiment. DETAILED DESCRIPTION

[0027] The exemplary embodiments will be described in detail herein below with reference to the drawings. In the following description, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The embodiments described in the following exemplary embodiments are not meant to represent all embodiments consistent with the application. Rather, they are merely examples of apparatus and methods consistent with some aspects of the application as detailed in the appended claims.

[0028] It should be noted that the brief descriptions of terms in this application are only for the convenience of understanding the embodiments described below, and are not intended to limit the embodiments of this application. Unless otherwise stated, these terms should be understood in their ordinary and common meaning. The terms "first," "second," etc., in the specification, claims, and accompanying drawings of this application are used to distinguish similar or related objects or entities, and do not necessarily mean to limit a specific order or sequence, unless otherwise indicated. It should be understood that such terms can be used interchangeably where appropriate, for example, to be implemented in an order other than those given in the illustrations or descriptions of the embodiments of this application. The terms "comprising" and "having," and any variations thereof, in the specification, claims, and accompanying drawings of this application, are intended to be omnipresent but not exclusive. For example, a product or device that comprises a series of components is not necessarily limited to those components that are explicitly listed, but may include other components that are not explicitly listed or are inherent to such products or devices. The term "module" as used in this application refers to any known or subsequently developed hardware, software, firmware, artificial intelligence, fuzzy logic, or combination of hardware and / or software code capable of performing the functions associated with that element.

[0029] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, data stored, data displayed, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties. Furthermore, the collection, storage, use, processing, transmission, provision, disclosure, and application of the relevant data all comply with the relevant laws, regulations, and standards of the relevant countries and regions, have taken necessary confidentiality measures, do not violate public order and good morals, and provide corresponding operation access points for users to choose to authorize or refuse.

[0030] Furthermore, the technical solution involved in this application, which involves big data analysis of user information (including but not limited to personal biometrics, identity data, consumption data, asset data, electronic terminal operation data, etc.) and the use of artificial intelligence technology for automated decision-making, and makes decisions that have a significant impact on personal rights based on the results of automated decision-making, provides users with corresponding operation entry points for users to choose to agree to or reject the results of automated decision-making; if the user chooses to reject, the process will proceed to the expert decision-making process.

[0031] It should be noted that the image recognition methods, devices, equipment, media and products provided in this application can be used in the field of fintech, or in any field other than fintech. The application fields of the image recognition methods, devices, equipment, media and products in this application are not limited.

[0032] In online customer service systems, users may express their needs or emotions by sending images containing text, emojis, or complex backgrounds (such as screenshots of product usage issues or images with emotional expressions). Traditional artificial intelligence (AI) dialogue assistants can only answer questions based on text input and cannot directly analyze the key information in images, resulting in low response efficiency and difficulty in meeting users' personalized needs. In the education field, students may seek help by uploading handwritten notes or images containing formulas or charts. AI needs to quickly extract the text and combine it with semantic understanding to provide answers. In social media interactions, users convey emotions through emojis or images with text. AI needs to identify emotional tendencies to generate empathetic feedback. Furthermore, in professional fields such as finance and healthcare, users may consult by uploading images with professional terminology (such as screenshots of contract terms or medical imaging reports). AI needs to accurately extract key information and combine it with domain knowledge to provide professional advice. Current technical solutions cannot efficiently analyze multimodal information (such as text, emotions, and scenes) in images without requiring users to manually input text. The results of image emotion recognition are inaccurate, thus limiting the flexibility and intelligence level of AI dialogue assistants.

[0033] The image recognition method, apparatus, device, medium, and product provided in this application extract features from the input subtitle image to be recognized based on an emotion recognition model to obtain the frequency domain and spatial domain features of the text in the subtitle image. Then, the frequency domain and spatial domain features of the text in the subtitle image are input into an attention fusion unit for feature fusion to obtain fused features. These fused features represent the emotion features of the subtitle image to be recognized. After being decoded by a decoding unit, pixel features are output to finally obtain the emotion recognition result. This improves the accuracy of image emotion recognition.

[0034] The technical solution of this application and how the technical solution of this application solves the above-mentioned technical problems are described in detail below with specific embodiments. These specific embodiments can be combined with each other, and the same or similar concepts or processes may not be described again in some embodiments. The embodiments of this application will be described below with reference to the accompanying drawings.

[0035] Example 1

[0036] Figure 1 An exemplary flowchart of an image recognition method is shown; for example... Figure 1 As shown, the method includes:

[0037] Step 101: Obtain the subtitle image to be recognized.

[0038] Optionally, the image to be identified can be a screenshot of contract terms in the financial field; a screenshot of an image in the medical field; or an image containing formulas and charts uploaded by students in the education field. This application does not impose any restrictions on this.

[0039] Optionally, the method also includes:

[0040] The subtitle image to be recognized is preprocessed, including at least one of the following: histogram equalization, contrast adjustment, color enhancement, smoothing filtering, and noise reduction.

[0041] Optionally, after acquiring the subtitle image to be recognized, preprocessing can be performed to improve its clarity. For example, histogram equalization, contrast adjustment, sharpening, and color enhancement can be used to transform the original subtitle image, making the enhanced image more distinct from the original, thereby improving image quality. The main goal of image enhancement is to improve image contrast, clarity, and readability. Image denoising involves processing the original image using methods such as average filtering, median filtering, high-pass filtering, and low-pass filtering, making the denoised image more accurately resemble the original image in terms of features, thus eliminating noise in the image. The main goal of image denoising is to improve image clarity and readability.

[0042] In this example, the clarity and quality of the subtitle image to be identified are improved by preprocessing, thereby further improving the accuracy of the recognition results.

[0043] Step 102: Input the subtitle image into the encoding unit of the emotion recognition model to extract features and obtain frequency domain features and spatial domain features.

[0044] Among them, frequency domain features and spatial domain features characterize the textual description features of the caption image. For example, the encoding unit can be an encoder, which can be any encoder capable of extracting features from an image; this application does not impose any restrictions on this. The encoder consists of convolutional layers and pooling layers. The convolutional layers generate feature maps, while the pooling layers continuously reduce the dimensionality of these feature maps to obtain more key features with more significant spatial invariance; reducing the resolution of the feature maps reduces the computational cost of the model.

[0045] Step 103: Based on the frequency domain features and spatial domain features, perform feature fusion using the multi-domain attention fusion unit of the emotion recognition model to obtain fused enhanced features.

[0046] For example, the frequency domain features and spatial domain features are further processed by feature fusion to obtain the enhanced fused features. This step fuses and enhances the frequency domain features and spatial domain features, further improving the expressiveness of the features.

[0047] Optionally, based on frequency domain features and spatial domain features, feature fusion is performed using a multi-domain attention fusion unit of the emotion recognition model to obtain enhanced fused features, including:

[0048] Based on the frequency domain features and spatial domain features, feature extraction and feature fusion are performed to obtain the first fused feature.

[0049] The first fusion feature is subjected to cross-attention calculation to obtain the fusion enhancement feature.

[0050] For example, the sentiment recognition model can be an improved U-NET model. First, a U-NET-based network segmentation model is constructed. The original U-NET model consists of two basic parts: an encoder and a decoder. To further improve the segmentation accuracy of the U-NET model, a multi-scale dual-representation alignment filter is used in the encoder-decoder connection stage. This mainly includes the following two points: Multi-scale mapping: Vertical bar convolutions of different scales are used to process frequency domain features and spatial domain features. The processed features are concatenated and subjected to 1x1 convolutions to obtain a matrix Q, K, V of uniform scale; this is the first fused feature. The first fused feature is used as input to a dual-representation alignment filter (DAF), which uses a cross-attention mechanism to achieve semantic alignment and feature selection. Attention is calculated by querying key-value pairs of the other party and itself, and feature weighting is performed to ultimately achieve feature selection, ensuring that important features are not lost and guaranteeing the accuracy of the results.

[0051] Vertical stripe convolutions of different scales are used to process frequency domain features and spatial domain features.

[0052] Step 104: Based on the fusion enhancement features, the decoding unit of the subtitle emotion recognition model outputs the pixel image of the subtitle image; based on the pixel image, the recognition result of the subtitle image is obtained.

[0053] Optionally, after obtaining the fused enhanced features from the input subtitle image through feature extraction, feature fusion, and cross-attention calculation, the fused enhanced features are then decoded based on the decoding unit of the subtitle emotion recognition model to obtain the pixel image of the subtitle image. The decoding unit in this example can be any decoder, and this application does not impose any restrictions on it. The pixel image of the subtitle image is a pixel image with the same size as the input subtitle image. Based on this pixel image, the emotion recognition result of the pixel image can be obtained. For example, the recognition result can include the text description obtained from parsing the input image, or it can include the emotion recognition result output by the model; the emotion recognition result can be happiness, sadness, surprise, etc., and subsequently, the answer to the input question can be generated based on the emotion recognized from the image, making the generated answer more accurate.

[0054] Optionally, the method also includes:

[0055] Based on the subtitle image, a text description of the subtitle image is obtained using optical character recognition technology.

[0056] For example, based on the subtitle image, optical character recognition (OCR) technology can also be used to identify the text description content contained in the input image. OCR technology uses image processing and statistical machine learning methods to convert text in an image into plain text. In practical applications, OCR technology encompasses multiple processing stages, such as text region localization, text correction, text segmentation, text recognition, and post-processing. It employs techniques such as connected component analysis for text region localization, rotation and affine transformations for text correction, binarization and noise filtering for text segmentation, logistic regression and other classifiers for text recognition, and finally, rule-based and language model-based post-processing to improve the accuracy and effectiveness of recognition.

[0057] In this example, based on an emotion recognition model, features are extracted from the input subtitle image to be recognized to obtain the frequency and spatial domain features of the text in the subtitle image. These features are then input into an attention fusion unit for feature fusion to obtain fused features. These fused features represent the emotion features of the subtitle image to be recognized. After passing through a decoding unit, pixel features are output to finally obtain the emotion recognition result, thus improving the accuracy of image emotion recognition.

[0058] Optionally, the method also includes:

[0059] The recognition results of the subtitle image, the text description of the subtitle image, and the question description are input into the large language model to obtain the answer to the question.

[0060] For example, based on the aforementioned emotion recognition model, the recognition result of the subtitle image is obtained; and after obtaining the text description of the subtitle image using optical character recognition technology, the recognition result of the subtitle image, the text description of the subtitle image, and the corresponding question can be input into a large language model to obtain the answer to the question. In one example, a user can input their question, a screenshot related to the question, and a description of the screenshot into the large language model, which will then output the answer to the question. In another example, the user can also input their question and a related screenshot into the large language model to obtain the answer to the question. This example does not impose any limitations on this.

[0061] In this example, based on the recognition results of the caption image and the text description, the large language model can output the corresponding answer based on the sentiment of the image recognition, the text description, and the question. This answer is generated based on sentiment recognition, which improves the accuracy of the generated answer.

[0062] Optionally, the method also includes:

[0063] Obtain the training caption image set.

[0064] The emotion recognition model is trained based on the training caption image set to obtain the trained caption emotion recognition model.

[0065] For example, the training caption image set can include 22,424 caption images extracted from four videos from an online platform, with sentiment added to the caption images. It can contain a variety of characters, including Thai consonants, vowels, tone marks, punctuation marks, numbers, Roman characters, and Arabic numerals. This training caption image set can also contain 157 unique characters, providing a resource for addressing the challenges of text recognition in complex contexts. It meets the growing demand for high-quality, multilingual text recognition data; variations in text length, font, and position in these images increase complexity, providing a valuable resource for developing and evaluating deep learning models. This training caption image set helps to accurately transcribe text from video content while laying the foundation for improving the computational efficiency of text recognition systems. The training caption image set can also be preprocessed using methods such as histogram equalization, contrast adjustment, sharpening, and color enhancement to improve its quality. The sentiment recognition model can then be trained based on this preprocessed training caption image set.

[0066] Optionally, the emotion recognition model is trained based on the training caption image set, including:

[0067] Based on the training caption image set, the text description corresponding to the training caption image set is identified using optical character recognition technology.

[0068] The text description and sentiment tags are used as labels for the training caption image set; the sentiment recognition model is trained based on the training caption image set and the labels.

[0069] Optionally, based on the obtained training caption image set, and using the aforementioned optical character recognition technology, the text description corresponding to each training caption image in the training caption image set is obtained. The text description and the added sentiment tags are used as labels for the training caption image set; the sentiment recognition model is trained based on the training caption image set and the corresponding labels.

[0070] In this example, by adding text descriptions and corresponding sentiment tags to the training caption image set, the effect of the sentiment recognition model is improved, thus enhancing the performance of the sentiment recognition model.

[0071] Optionally, the method also includes:

[0072] Obtain the test caption image set.

[0073] The trained emotion recognition model was validated based on the test caption image set, and the validation results were obtained.

[0074] For example, a portion of the subtitle image set obtained using the aforementioned method for obtaining the training subtitle image set can be used as the test subtitle image set. The trained emotion recognition model is then validated based on the test subtitle image set.

[0075] For example, the performance of emotion recognition models can be optimized and recognition speed and accuracy improved through model pruning techniques. Pruning makes the model sparser and lighter by identifying and removing "unimportant" connections or structures in the neural network. Common methods include amplitude pruning, gradient pruning, and iterative pruning.

[0076] For example, an evaluation metric can be used to assess a network model, as shown in the following formula:

[0077]

[0078] Wherein, TP, FP, FN, and TN represent the number of true positives, false positives, false negatives, and true negatives respectively after comparing the model's predicted values ​​with the actual values; F represents the evaluation index value.

[0079] For example, when F is greater than a preset threshold, the model is considered to be trained successfully; when F is not greater than the preset threshold, the model is considered to be trained unsuccessfully, and training continues.

[0080] In this example, the performance of the trained emotion recognition model is verified by using performance metrics, which ensures the performance of the emotion recognition model and improves the accuracy of image recognition when it is used for subsequent image recognition based on the trained emotion recognition model.

[0081] Figure 2 This application provides a schematic diagram illustrating the training process of an example emotion recognition model; as shown below. Figure 2 As shown, firstly, subtitle images from multiple videos are acquired to construct a subtitle image set, which serves as the training set for the model. Then, the acquired images undergo preprocessing including image enhancement and denoising; image enhancement includes histogram equalization, contrast adjustment, sharpening, and color enhancement. Image denoising includes average filtering, median filtering, high-pass filtering, and low-pass filtering. Next, the preprocessed images are used to identify the corresponding text descriptions using optical character recognition (OCR) technology. Finally, the model is optimized using the U-Net network, model pruning methods, and a multi-scale dual-representation alignment filter. The accuracy of the trained model is then evaluated.

[0082] The image recognition method provided in this embodiment extracts features from the input subtitle image to be recognized based on an emotion recognition model, obtaining the frequency domain and spatial domain features of the text in the subtitle image. Then, the frequency domain and spatial domain features of the text in the subtitle image are input into an attention fusion unit for feature fusion to obtain fused features. These fused features represent the emotion features of the subtitle image to be recognized. After being decoded by a decoding unit, pixel features are output to finally obtain the emotion recognition result, thus improving the accuracy of image emotion recognition.

[0083] Example 2

[0084] Figure 3 An exemplary schematic diagram of an image recognition device is shown; such as Figure 3 As shown, the device includes:

[0085] The acquisition module 21 is used to acquire the subtitle image to be recognized.

[0086] The extraction module 22 is used to input the subtitle image into the encoding unit of the emotion recognition model to extract features and obtain frequency domain features and spatial domain features; the frequency domain features and spatial domain features represent the text description features of the subtitle image.

[0087] The fusion module 23 is used to perform feature fusion based on the frequency domain features and spatial domain features, and the multi-domain attention fusion unit of the emotion recognition model to obtain fused enhanced features.

[0088] The processing module 24 is used to output the pixel image of the subtitle image based on the fusion enhancement features and the decoding unit of the emotion recognition model; and to obtain the recognition result of the subtitle image based on the pixel image.

[0089] The image recognition device provided in this embodiment can execute the image recognition method provided in the above method embodiment. Its implementation principle and technical effect are similar, and will not be described in detail here.

[0090] Example 3

[0091] Figure 4 The diagram above illustrates the structure of an electronic device, which includes:

[0092] The device includes a processor 291 and a memory 292; it may also include a communication interface 293 and a bus 294. The processor 291, memory 292, and communication interface 293 can communicate with each other via the bus 294. The communication interface 293 can be used for information transmission. The processor 291 can invoke logical instructions stored in the memory 292 to execute the methods described in the example above.

[0093] Furthermore, the logic instructions in the aforementioned memory 292 can be implemented as software functional units and, when sold or used as independent products, can be stored in a computer-readable storage medium.

[0094] The memory 292, as a computer-readable storage medium, can be used to store software programs and computer-executable programs, such as program instructions / modules corresponding to the methods in the embodiments of this application. The processor 291 executes functional applications and data processing by running the software programs, instructions, and modules stored in the memory 292, that is, it implements the methods in the above method examples.

[0095] The memory 292 may include a program storage area and a data storage area. The program storage area may store the operating system and application programs required for at least one function; the data storage area may store data created based on the use of the terminal device. Furthermore, the memory 292 may include high-speed random access memory and may also include non-volatile memory.

[0096] This application also provides a computer-readable storage medium storing computer-executable instructions, which, when executed by a processor, are used to implement the method in any of the embodiments.

[0097] This application also provides a computer program product, including a computer program that, when executed by a processor, is used to implement the method in any of the embodiments.

[0098] It should be noted that, for the sake of simplicity, the foregoing method embodiments are all described as a series of actions. However, those skilled in the art should understand that this application is not limited to the described order of actions, as some steps may be performed in other orders or simultaneously according to this application. Furthermore, those skilled in the art should also understand that the embodiments described in the specification are all optional embodiments, and the actions and modules involved are not necessarily essential to this application.

[0099] It should be further noted that although the steps in the flowchart are shown sequentially according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless explicitly stated herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least some steps in the flowchart may include multiple sub-steps or multiple stages. These sub-steps or stages are not necessarily completed at the same time, but can be executed at different times. The execution order of these sub-steps or stages is not necessarily sequential, but can be performed alternately or in turn with other steps or at least some of the sub-steps or stages of other steps.

[0100] It should be understood that the above-described device embodiments are merely illustrative, and the device of this application can also be implemented in other ways. For example, the division of units / modules in the above embodiments is only a logical functional division, and there may be other division methods in actual implementation. For example, multiple units, modules, or components may be combined, or integrated into another system, or some features may be ignored or not executed.

[0101] Furthermore, unless otherwise specified, the functional units / modules in the various embodiments of this application can be integrated into one unit / module, or each unit / module can exist physically separately, or two or more units / modules can be integrated together. The integrated units / modules described above can be implemented in hardware or as software program modules.

[0102] When integrated units / modules are implemented in hardware, the hardware can be digital circuits, analog circuits, etc. The physical implementation of the hardware structure includes, but is not limited to, transistors, memristors, etc. Unless otherwise specified, the processor can be any suitable hardware processor, such as a CPU, GPU, FPGA, DSP, and ASIC, etc. Unless otherwise specified, the storage unit can be any suitable magnetic or magneto-optical storage medium, such as Resistive Random Access Memory (RRAM), Dynamic Random Access Memory (DRAM), Static Random Access Memory (SRAM), Enhanced Dynamic Random Access Memory (EDRAM), High-Bandwidth Memory (HBM), Hybrid Memory Cube (HMC), etc.

[0103] If the integrated unit / module is implemented as a software program module and sold or used as an independent product, it can be stored in a computer-readable storage device (CMD). Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a memory and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods of the various embodiments of this application. The aforementioned memory includes various media capable of storing program code, such as a USB flash drive, read-only memory (ROM), random access memory (RAM), portable hard drive, magnetic disk, or optical disk.

[0104] In the above embodiments, the descriptions of each embodiment have their own emphasis. For parts not described in detail in a certain embodiment, please refer to the relevant descriptions of other embodiments. The technical features of the above embodiments can be combined arbitrarily. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as the combination of these technical features does not contradict each other, it should be considered within the scope of this specification.

[0105] Other embodiments of this application will readily occur to those skilled in the art upon consideration of the specification and practice of the invention disclosed herein. This application is intended to cover any variations, uses, or adaptations of this application that follow the general principles of this application and include common knowledge or customary techniques in the art not disclosed herein. The specification and examples are to be considered exemplary only, and the true scope and spirit of this application are indicated by the following claims.

[0106] It should be understood that this application is not limited to the precise structure described above and shown in the accompanying drawings, and various modifications and changes can be made without departing from its scope. The scope of this application is limited only by the appended claims.

Claims

1. An image recognition method characterized by, The method comprises: obtaining a subtitle image to be recognized; inputting the subtitle image into an encoding unit of an emotion recognition model to perform feature extraction, and obtaining frequency domain features and spatial domain features; the frequency domain features and the spatial domain features represent text description features of the subtitle image; performing feature fusion based on a multi-domain attention fusion unit of the emotion recognition model according to the frequency domain features and the spatial domain features, and obtaining fused enhanced features; outputting a pixel image of the subtitle image based on a decoding unit of the emotion recognition model according to the fused enhanced features; obtaining a recognition result of the subtitle image according to the pixel image.

2. The method of claim 1, wherein, The method further comprises: obtaining a text description of the subtitle image based on an optical character recognition technology according to the subtitle image. The method further comprises:

3. The method of claim 1, wherein, obtaining a set of training subtitle images; training the emotion recognition model according to the set of training subtitle images to obtain a trained subtitle emotion recognition model.

4. The method of claim 1, wherein, The method further comprises: identifying text descriptions corresponding to the set of training subtitle images based on an optical character recognition technology according to the set of training subtitle images; using the text descriptions and emotion labels as labels of the set of training subtitle images; and training the emotion recognition model according to the set of training subtitle images and the labels.

5. The method of claim 4, wherein, The method further comprises: obtaining a set of test subtitle images; verifying the trained emotion recognition model according to the set of test subtitle images, and obtaining a verification result.

6. The method of claim 4, wherein, The method further comprises: preprocessing the subtitle image to be recognized, wherein the preprocessing comprises at least one of histogram equalization, contrast adjustment, color enhancement, smoothing filtering, and noise removal. The method further comprises:

7. The method according to any one of claims 1 to 6, characterized in that, inputting the recognition result of the subtitle image, the text description of the subtitle image, and a question description into a large language model to obtain an answer corresponding to the question. The method comprises:

8. The method according to any one of claims 1 to 6, characterized in that, an obtaining module configured to obtain a subtitle image to be recognized; an extracting module configured to input the subtitle image into an encoding unit of an emotion recognition model to perform feature extraction, and obtain frequency domain features and spatial domain features; the frequency domain features and the spatial domain features represent text description features of the subtitle image; 9. An image recognition apparatus characterized by comprising: a fusion module configured to perform feature fusion based on a multi-domain attention fusion unit of the emotion recognition model according to the frequency domain features and the spatial domain features, and obtain fused enhanced features; a processing module configured to output a pixel image of the subtitle image based on a decoding unit of the emotion recognition model according to the fused enhanced features; obtain a recognition result of the subtitle image according to the pixel image. The method comprises: a processor and a memory connected to the processor in communication; ​ 10. An electronic device, comprising: ​ ​ The memory stores computer-executable instructions; The processor executes the computer-executable instructions stored in the memory to implement the method according to any one of claims 1 to 8.

11. A computer readable storage medium, characterized in that, The computer-readable storage medium stores computer-executable instructions, which, when executed by a processor, are used to implement the method according to any one of claims 1 to 8.

12. A computer program product, characterised in that, A computer program is included, which, when executed by a processor, implements the method according to any one of claims 1 to 8.