Image recognition method and device, training method and device, intelligent agent, equipment, medium and product
By dynamically selecting and fusion of visual features of multimodal large models, the problem of information loss in image recognition in the prior art is solved, and higher accuracy and flexibility are achieved, and various complex image tasks are adapted to.
Patent Information
- Application Number
- CN202510804877.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-16
- Publication Date
- 2025-09-02
AI Technical Summary
Existing multimodal large models have the problem of using only a single-level feature in image recognition, resulting in information loss and poor results, especially in tasks that require both global and detailed information, and the model is insufficient in flexibility and adaptability.
By dynamically selecting and fusing multiple candidate visual features based on the image recognition strategy corresponding to the input image and the problem, and adopting feature selection strategies and feature fusion strategies to ensure that the fused visual features are adapted to input problems, and improve recognition accuracy and pertinence.
It enhances the accuracy and adaptability of image recognition, can focus on key information more accurately, solves the problem of information loss, and improves the recognition effect of the model in complex tasks.
Smart Images

Figure CN120580552A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of artificial intelligence technology, in particular to the fields of large models, natural language processing and computer vision, and more specifically to an image recognition method, training method, apparatus, intelligent agent, equipment, medium and product. Background Art
[0002] With the development of artificial intelligence, multimodal technology has emerged. Multimodal technology refers to the integration of data from different modalities (such as images, text, audio, video, etc.). Because each modality can provide unique information for specific tasks, it can enhance the model's understanding and reasoning capabilities. Summary of the Invention
[0003] The present disclosure provides an image recognition method, training method, apparatus, intelligent agent, equipment, medium and product.
[0004] According to one aspect of the present disclosure, an image recognition method is provided, comprising: fusing at least two target visual features among multiple candidate visual features of the input image according to an image recognition strategy corresponding to an input image and an input problem to obtain a fused visual feature, wherein the image recognition strategy indicates a selection method of the target visual features and a fusion method for the target visual features so that the fused visual features are adapted to the input problem; and determining an image recognition result for the input problem according to the fused visual feature.
[0005] According to another aspect of the present disclosure, a training method for a multimodal large model is provided, comprising: according to an image recognition strategy corresponding to a sample input image and a sample input problem, fusing at least one sample target visual feature from the multiple layers of sample candidate visual features of the sample input image to obtain a sample fused visual feature, wherein the image recognition strategy indicates a selection method of the sample target visual feature and a fusion method for the sample target visual feature, so that the sample fused visual feature is adapted to the sample input problem; determining a sample image recognition result for the sample input problem according to the sample fused visual feature; and adjusting the model parameters of the multimodal large model to be trained according to the sample image recognition result and a reference image recognition result corresponding to the sample input image and the sample input problem to obtain a trained multimodal large model.
[0006] According to another aspect of the present disclosure, an image recognition device is provided, including: a first acquisition module, used to fuse at least two target visual features among the multiple layers of candidate visual features of the input image according to an image recognition strategy corresponding to the input image and the input problem, to obtain a fused visual feature, wherein the image recognition strategy indicates a selection method of the target visual features and a fusion method for the target visual features, so that the fused visual features are adapted to the input problem; and a first determination module, used to determine the image recognition result for the input problem according to the fused visual feature.
[0007] According to another aspect of the present disclosure, a training device for a multimodal large model is provided, including: a second acquisition module, used to fuse at least two sample target visual features of the multi-layer sample candidate visual features of the above-mentioned sample input image according to an image recognition strategy corresponding to the sample input image and the sample input problem, to obtain a sample fused visual feature, wherein the above-mentioned image recognition strategy indicates the selection method of the above-mentioned sample target visual features and the fusion method for the above-mentioned sample target visual features, so that the above-mentioned sample fused visual features are adapted to the above-mentioned sample input problem; a second determination module, used to determine the sample image recognition result for the above-mentioned sample input problem according to the above-mentioned sample fused visual feature; and a training module, used to adjust the model parameters of the multimodal large model to be trained according to the above-mentioned sample image recognition result and the reference image recognition result corresponding to the above-mentioned sample input image and the above-mentioned sample input problem, to obtain a trained multimodal large model.
[0008] According to another aspect of the present disclosure, an artificial intelligence agent is provided, which is configured to execute the steps of the above method.
[0009] According to another aspect of the present disclosure, an electronic device is provided, comprising: one or more processors; and a memory for storing one or more computer programs, wherein the one or more processors execute the one or more computer programs to implement the steps of the above method.
[0010] According to another aspect of the present disclosure, a computer-readable storage medium is provided, on which a computer program or instructions are stored. When the computer program or instructions are executed by a processor, the steps of the above method are implemented.
[0011] According to another aspect of the present disclosure, a computer program product is provided, comprising a computer program or instructions, which implement the steps of the above method when executed by a processor.
[0012] It should be understood that the contents described in this section are not intended to identify the key or important features of the embodiments of the present disclosure, nor are they intended to limit the scope of the present disclosure. Other features of the present disclosure will become readily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0013] The above and other objects, features and advantages of the present disclosure will become more apparent through the following description of the embodiments of the present disclosure with reference to the accompanying drawings, in which:
[0014] Figure 1 The system architecture to which the image recognition method and the training method of the multimodal large model can be applied according to the embodiments of the present disclosure is schematically shown;
[0015] Figure 2 The following schematically shows a flow chart of an image recognition method according to an embodiment of the present disclosure;
[0016] Figure 3 Schematically illustrates an example of a process of fusing at least two target visual features from a plurality of candidate visual features of an input image to obtain a fused visual feature according to an image recognition strategy corresponding to the input image and the input question according to an embodiment of the present disclosure;
[0017] Figure 4 The following schematically illustrates an example of a process for determining a feature selection strategy and a feature fusion strategy according to an embodiment of the present disclosure;
[0018] Figure 5 An example diagram schematically illustrates a process for determining a feature selection strategy and a feature fusion strategy according to another embodiment of the present disclosure;
[0019] Figure 6 An example diagram schematically illustrates an image recognition process according to an embodiment of the present disclosure;
[0020] Figure 7 The flowchart of the training method of the multimodal large model according to the embodiment of the present disclosure is schematically shown;
[0021] Figure 8 An example diagram schematically illustrates a training process of a multimodal large model according to an embodiment of the present disclosure;
[0022] Figure 9 Schematically shows a block diagram of an image recognition device according to an embodiment of the present disclosure;
[0023] Figure 10 A block diagram of a multimodal large model training apparatus according to an embodiment of the present disclosure is schematically shown;
[0024] Figure 11Schematically shows a structural block diagram of an intelligent agent of a large model according to an embodiment of the present disclosure; and
[0025] Figure 12 The block diagram of an electronic device suitable for implementing an image recognition method and a training method for a multimodal large model according to an embodiment of the present disclosure is schematically shown. DETAILED DESCRIPTION
[0026] Hereinafter, embodiments of the present disclosure will be described with reference to the accompanying drawings. However, it should be understood that these descriptions are merely exemplary and are not intended to limit the scope of the present disclosure. In the detailed description below, for ease of explanation, many specific details are set forth to provide a comprehensive understanding of the embodiments of the present disclosure. However, it is apparent that one or more embodiments may also be implemented without these specific details. In addition, in the following description, descriptions of well-known structures and technologies are omitted to avoid unnecessary confusion of the concepts of the present disclosure.
[0027] The terms used herein are only for describing specific embodiments and are not intended to limit the present disclosure. The terms "comprise," "include," etc. used herein indicate the presence of the features, steps, operations, and / or components, but do not exclude the presence or addition of one or more other features, steps, operations, or components.
[0028] All terms used herein (including technical and scientific terms) have the meanings commonly understood by those skilled in the art unless otherwise defined. It should be noted that the terms used herein should be interpreted as having a meaning consistent with the context of this specification and should not be interpreted in an idealized or overly rigid manner.
[0029] When expressions such as "at least one of A, B, and C, etc." are used, they should generally be interpreted in accordance with the meaning commonly understood by those skilled in the art (for example, "a system having at least one of A, B, and C" should include but is not limited to a system having A alone, B alone, C alone, A and B, A and C, B and C, and / or A, B, C, etc.).
[0030] For example, large multimodal models have been widely used in image recognition and understanding tasks. When processing visual information, these models employ structures such as the Vision Transformer (ViT) to extract visual features. However, these models have limitations in feature extraction.
[0031] First, existing models often use only the highest-level features of ViT for recognition. These features primarily encode semantic information and lack the ability to capture image details. In tasks that require attention to image details, this feature extraction method can lead to the loss of detailed information, affecting the model's recognition accuracy and effectiveness.
[0032] Secondly, although there are methods that attempt to fuse multi-scale information in pure vision tasks, these methods are usually limited to the direct splicing of features and lack research on the contribution of different feature layers in specific tasks. Therefore, it is impossible to dynamically select and fuse visual features of different layers according to actual task requirements, resulting in insufficient flexibility and adaptability of the model when processing complex image tasks.
[0033] To this end, embodiments of the present disclosure propose an image recognition solution. For example, based on an image recognition strategy corresponding to an input image and an input question, at least two target visual features from a plurality of candidate visual features of the input image are fused to obtain a fused visual feature, wherein the image recognition strategy indicates a method for selecting the target visual features and a method for fusing the target visual features so that the fused visual feature is adapted to the input question; and based on the fused visual feature, an image recognition result for the input question is determined.
[0034] According to the embodiments of the present disclosure, for different input images and input problems, appropriate target visual features are dynamically selected and effectively fused based on the corresponding image recognition strategy, so that the obtained fused visual features are more accurately focused on the key information related to the input problem. This solves the problem of information loss and poor effect caused by using only a single level of features when processing image recognition tasks that need to take into account both global and detailed information, thereby improving the accuracy and pertinence of image recognition.
[0035] In the technical solution of the present invention, the collection, storage, use, processing, transmission, provision and disclosure of user personal information involved comply with the provisions of relevant laws and regulations and do not violate public order and good morals.
[0036] In the technical solution of the present invention, the user's authorization or consent is obtained before obtaining or collecting the user's personal information.
[0037] Figure 1 The system architecture of the image recognition method and the multimodal large model training method according to the embodiment of the present disclosure is schematically shown. It should be noted that Figure 1 The examples shown are merely examples of system architectures to which the embodiments of the present disclosure may be applied, to help those skilled in the art understand the technical content of the present disclosure, but do not mean that the embodiments of the present disclosure may not be used in other devices, systems, environments or scenarios.
[0038] like Figure 1As shown, the system architecture 100 according to this embodiment may include a first terminal device 101, a second terminal device 102, a third terminal device 103, a network 104, and a server 105. The network 104 is used as a medium for providing a communication link between the first terminal device 101, the second terminal device 102, the third terminal device 103, and the server 105. The network 104 may include various connection types, such as wired or wireless communication links or fiber optic cables.
[0039] The user may use at least one of the first terminal device 101, the second terminal device 102, and the third terminal device 103 to interact with the server 105 via the network 104 to receive or send messages, etc. Various communication client applications may be installed on the first terminal device 101, the second terminal device 102, and the third terminal device 103, such as shopping applications, web browser applications, search applications, instant messaging tools, email clients, social platform software, etc. (only as examples).
[0040] The first terminal device 101 , the second terminal device 102 , and the third terminal device 103 may be various electronic devices having display screens and supporting web browsing, including but not limited to smart phones, tablet computers, laptop computers, desktop computers, and the like.
[0041] The server 105 may be a server that provides various services, such as a background management server (for example only) that supports websites browsed by users using the first terminal device 101, the second terminal device 102, and the third terminal device 103. The background management server may analyze and process received data such as user requests, and feed back processing results (e.g., web pages, information, or data obtained or generated based on user requests) to the terminal devices.
[0042] It should be noted that the image recognition method and the training method of the multimodal large model provided in the embodiments of the present disclosure can generally be executed by the server 105. Accordingly, the image recognition device and the training device of the multimodal large model provided in the embodiments of the present disclosure can generally be set in the server 105. The image recognition method and the training method of the multimodal large model provided in the embodiments of the present disclosure can also be executed by a server or server cluster that is different from the server 105 and can communicate with the first terminal device 101, the second terminal device 102, the third terminal device 103 and / or the server 105. Accordingly, the image recognition device and the training device of the multimodal large model provided in the embodiments of the present disclosure can also be set in a server or server cluster that is different from the server 105 and can communicate with the first terminal device 101, the second terminal device 102, the third terminal device 103 and / or the server 105.
[0043] Alternatively, the image recognition method and the multimodal large model training method provided in the embodiments of the present disclosure may also be executed by the first terminal device 101, the second terminal device 102, or the third terminal device 103, or may also be executed by other terminal devices different from the first terminal device 101, the second terminal device 102, or the third terminal device 103. Accordingly, the image recognition device and the multimodal large model training device provided in the embodiments of the present disclosure may also be provided in the first terminal device 101, the second terminal device 102, or the third terminal device 103, or in other terminal devices different from the first terminal device 101, the second terminal device 102, or the third terminal device 103.
[0044] It should be understood that Figure 1 The number of terminal devices, networks and servers in the embodiment is merely illustrative. Any number of terminal devices, networks and servers may be provided as required.
[0045] It should be noted that the sequence numbers of the operations in the following method are only used to indicate the operation for the purpose of description, and should not be regarded as indicating the order in which the operations should be performed. Unless explicitly stated, the method does not need to be performed in the order shown.
[0046] The above describes the system architecture that can be used to apply the image recognition method provided by the present disclosure. Figure 2 For example, the image recognition process of the present disclosure is further explained.
[0047] Figure 2 The flowchart of the image recognition method according to the embodiment of the present disclosure is schematically shown.
[0048] like Figure 2 As shown, the image recognition method 200 includes operations S210 to S220.
[0049] In operation S210, at least two target visual features among multiple candidate visual features of the input image are fused according to an image recognition strategy corresponding to the input image and the input problem to obtain a fused visual feature, wherein the image recognition strategy indicates a selection method of the target visual features and a fusion method for the target visual features so that the fused visual feature adapts to the input problem.
[0050] In operation S220 , an image recognition result for the input question is determined based on the fused visual features.
[0051] The input image refers to the image that needs to be recognized. The input question refers to the question posed for the input image. For the input image, feature extraction processing can be performed on the input image to obtain multiple candidate visual features. Candidate visual features refer to a variety of visual features extracted from the input image that may be used for recognition, which can cover information of different levels and types. For example, multiple candidate visual features can be features output by different layers of ViT, and the multiple candidate visual features can include the global visual features of the highest layer and the local visual features of each intermediate layer. The local visual features include, for example, the edges, textures, color distribution and other features of the input image.
[0052] Image recognition strategies are determined based on the input image and the input problem, and are used to determine how target visual features are selected and how they are integrated. Target visual features are selected from candidate visual features to solve the input problem.
[0053] In one embodiment, image recognition can be applied to document understanding scenarios. In this scenario, the input image may include a document image, and the input question may include a question related to document understanding. Document understanding scenarios are characterized by the fact that the information relevant to the input question may be a very small portion of the entire input image. Therefore, a thorough understanding of both the global and detailed information of the input image is required to solve the input question.
[0054] For example, the input image could be a document image containing text, graphics, and other elements, and the input question could be about the information contained in a specific paragraph within that document. Correspondingly, the image recognition strategy could be a rule for selecting and fusing visual features of an area related to that specific paragraph. Thus, by dynamically fusing multiple semantic visual features, key information within a document can be more accurately located and parsed for document understanding problems, improving the accuracy and pertinence of document image recognition.
[0055] After selecting the target visual features, the selected target visual features can be fused according to the fusion method indicated by the image recognition strategy to obtain a fused visual feature that better matches the requirements of the input problem. Furthermore, the selected target visual features can be combined in a feature concatenation method, arranging the target visual features in a certain order to form a larger feature vector. Alternatively, a neural network structure such as a multi-layer perceptron can be used to perform a nonlinear transformation on the target visual features to obtain a fused visual feature.
[0056] For example, if the input image is an image of an ancient book and the input question is to identify the author of a certain text, the image recognition strategy can indicate that both local visual features representing the stroke characteristics of the text and global visual features representing the text layout characteristics are required. Therefore, a weighted summation method can be used to assign a larger first weight to the local visual features representing the stroke characteristics of the text and a smaller second weight to the global visual features representing the text layout characteristics. The two are weightedly fused to obtain a fused visual feature. The fused feature not only reflects the text details but also contains the layout information, thereby facilitating the determination of the author. The sum of the first weight and the second weight is 1. For example, the first weight can be 0.6 and the second weight can be 0.4.
[0057] After obtaining the fused visual features, the image recognition result for the input question can be determined based on the fused visual features. For example, using the aforementioned task of identifying the author of ancient texts, the fused visual features can be input into a pre-trained classification model. The classification model will then determine that the author of a particular passage is a specific ancient scholar based on the similarity between the fused visual features and the writing characteristics of different authors. Alternatively, a similarity matching method can be used to compare the fused visual features with features in a pre-existing standard feature library, and the answer corresponding to the most similar features can be found as the image recognition result.
[0058] It should be noted that the image recognition method provided by this disclosure can be applied not only to the aforementioned document understanding scenarios, but also to scenarios requiring attention to comprehensive visual multi-scale information or detailed information. For example, the image recognition method provided by this disclosure can be applied to optical character recognition, medical image analysis, autonomous driving, or remote sensing image interpretation.
[0059] According to the embodiments of the present disclosure, for different input images and input problems, appropriate target visual features are dynamically selected and effectively fused based on the corresponding image recognition strategy, so that the obtained fused visual features are more accurately focused on the key information related to the input problem. This solves the problem of information loss and poor effect caused by using only a single level of features when processing image recognition tasks that need to take into account both global and detailed information, thereby improving the accuracy and pertinence of image recognition.
[0060] The above is a preliminary description of the image recognition method. Figure 3 For example, the process of obtaining the fused visual features disclosed herein is further described.
[0061] Figure 3The diagram schematically shows an example of a process of fusing at least two target visual features from a plurality of candidate visual features of an input image according to an image recognition strategy corresponding to the input image and the input problem, to obtain fused visual features, according to an embodiment of the present disclosure.
[0062] like Figure 3 As shown in 300 , the image recognition strategy may include a feature selection strategy 301 and a feature fusion strategy 302 .
[0063] In one example, feature selection strategy 301 and feature fusion strategy 302 can be implemented based on a mixture of experts (MOE). For example, multiple expert models for different recognition types can be trained, and the specific expert model used to implement feature selection strategy 301 and feature fusion strategy 302 can be selected based on the recognition type corresponding to the input image and input question.
[0064] In another example, the feature selection strategy 301 and the feature fusion strategy 302 may also be implemented based on a tree model such as a Random Forest (RF) / Gradient Boosting Tree (GBT), or based on a Fully Connected Neural Network (FCNN), or based on a Recurrent Neural Network (RNN) and its variants, without limitation herein.
[0065] For a plurality of candidate visual features 303 of an input image, at least two target visual features 304 may be determined from the plurality of candidate visual features according to a feature selection strategy 301. The feature selection strategy 301 may be used to determine which target visual features to select from the plurality of candidate visual features 303, i.e., the feature selection strategy 301 indicates at least one of the identity and number of the target visual features 304, and the degree of difference between each two target visual features 304.
[0066] Since the multiple candidate visual features 303 are obtained by extracting features from the input image using a neural network structure, the difference can be determined based on the layer interval in the neural network structure used to obtain each candidate visual feature 303. For example, taking a 5-layer neural network structure as an example, if the number of layers between the candidate visual features obtained in the first layer and the candidate visual features obtained in the third layer is 2, the difference can be 2; if the number of layers between the candidate visual features obtained in the first layer and the candidate visual features obtained in the fifth layer is 4, the difference can be 4.
[0067] For the at least two selected target visual features 304, multiple target visual features 304 may be fused according to a feature fusion strategy 304 to obtain a fused visual feature 305. Feature fusion strategy 305 may be used to determine how to fuse multiple target visual features 304. That is, feature fusion strategy 305 indicates a weight to be assigned to each target visual feature 304. For example, when fusing text features and layout features in a document image, the feature fusion strategy may indicate a weighted summation of these two features according to a certain weight.
[0068] According to the embodiments of the present disclosure, since the image recognition strategy includes a feature selection strategy and a feature fusion strategy, it can flexibly select and fuse visual features according to the task requirements of the specific input image and input problem, and can better capture and utilize key visual information, thereby improving the accuracy and adaptability of image recognition.
[0069] In the embodiment of the present disclosure, the feature selection strategy and the feature fusion strategy can be dynamically determined according to the input image and the input question. Figure 4 The process of determining the feature selection strategy and feature fusion strategy in a dynamic manner is exemplified.
[0070] Figure 4 An example diagram schematically illustrates a process for determining a feature selection strategy and a feature fusion strategy according to an embodiment of the present disclosure.
[0071] like Figure 4 As shown in 400, taking the input question 410 and multiple candidate visual features of the input image including candidate visual feature 421, candidate visual feature 422, ..., candidate visual feature 42P, where P is a positive integer as an example, the determination process of the feature selection strategy 430 and the feature fusion strategy 450 is explained.
[0072] In one embodiment, feature selection strategy 430 can be determined based on the correlation between candidate visual features and text features of input question 410. Text features of an input question refer to feature representations obtained by textually processing the input question. For example, the question "Who is the author of this picture?" is converted into word vectors using the Word2Vec model. Correlation refers to the degree of association between candidate visual features and text features of input question 410, and can be used to assess the value of candidate visual features in answering the question.
[0073] For each of the candidate visual features described above, the correlation between the candidate visual feature and the text features of input question 410 can be determined, resulting in correlation 4211 corresponding to candidate visual feature 421, correlation 4222 corresponding to candidate visual feature 422, ..., and correlation 42P1 corresponding to candidate visual feature 42P. Each correlation can be calculated using a model based on an attention mechanism or using cosine similarity, which is not limited here.
[0074] For each of the aforementioned correlations, a feature selection strategy 430 can be determined based on candidate visual features corresponding to correlations that satisfy a preset correlation condition. The preset correlation condition is a pre-set correlation criterion for screening target visual features. For example, the preset correlation condition may be that the correlation score must be greater than 0.7. Thus, based on feature selection strategy 430, target visual features 441, 442, ..., and 44Q that are relevant to input question 410 can be selected from the multiple candidate visual features, where Q is a positive integer less than or equal to P.
[0075] In another embodiment, feature fusion strategy 450 can be determined based on the contribution of the target visual feature to input problem 410. Contribution refers to the importance of the target visual feature in solving input problem 410 and can be used to determine the weight distribution during feature fusion. For example, the weight of each target visual feature is positively correlated with its contribution, i.e., the higher the contribution, the greater the weight.
[0076] For each of the aforementioned target visual features, its contribution to the input question 410 can be determined, yielding contribution 4411 corresponding to target visual feature 441, contribution 4412 corresponding to target visual feature 442, and so on, and contribution 44Q1 corresponding to target visual feature 44Q. Each contribution can be calculated using a gradient boosting decision tree model or Shapley Additive Explanations (SHAP), without limitation. For each of the aforementioned contribution degrees, a feature fusion strategy 450 can be determined based on the contribution of each target visual feature.
[0077] According to the embodiments of the present disclosure, by determining the feature selection strategy based on correlation and the feature fusion strategy based on contribution, it is ensured that the model can accurately screen out the visual features that are most valuable for answering questions and reasonably allocate weights. This not only improves the accuracy of image recognition, but also enhances the adaptability and flexibility of the model, enabling it to more effectively handle various complex visual tasks.
[0078] In one embodiment, multiple candidate visual features are obtained by performing visual feature extraction on multiple region images in the input image. Region images refer to a number of sub-regions that are divided or selected from the input image. For example, a document image can be divided into multiple small regions according to a grid.
[0079] The multiple candidate visual features include local visual features related to information of multiple regional images in the input image, and global visual features obtained based on the local visual features. Local visual features refer to information features related to specific regional images in the input image, which can reflect the local details of the image. Global visual features refer to features obtained based on the comprehensive analysis of local visual features, which can reflect the overall information of the entire image. For example, local visual features may include the font features, font size features, line features of a table, etc. of a paragraph in a document image, and global visual features may include the overall layout features, text typesetting, theme style features, chart position, etc. of the document image.
[0080] In one example, multiple candidate visual features can be extracted from an input image using a ViT structure, with different layers outputting features at different levels. The shallower layers output local visual features, while the deeper layers output global visual features. Alternatively, multiple candidate visual features can be generated using a convolutional neural network (CNN), which first extracts local visual features through convolutional layers and then gradually generates global visual features through pooling layers and fully connected layers.
[0081] For any input question, global visual features are essential, while the need for additional local visual features can be determined by whether the input question involves local information. Local information refers to the area of the image corresponding to the region involved in the input question. For example, if a question is asked about the name of a person at a specific location in an image, that location would be the area corresponding to the local information.
[0082] If the input question does not involve local information, that is, the input question focuses on the overall content of the image rather than specific local details, global visual features can be directly selected as the target visual features. For example, if the input question is "What is the theme of this slide?", since it does not involve specific local elements, global visual features can be directly used as the target visual features to answer the input question.
[0083] When the input question involves local information, that is, the input question involves local details of the image, it is necessary to select both global visual features and local visual features corresponding to the local information as target visual features. For example, if the input question is "What is the data in the second row and third column of the table?", it is necessary to select both global visual features and local visual features of the corresponding cell in the table as target visual features.
[0084] In one embodiment, the local visual feature corresponding to the local information can be determined by the following operations: determining at least one area image related to the input question in the input image; and determining a candidate visual feature corresponding to the at least one area image among multiple candidate visual features as a target visual feature.
[0085] When the input question involves local information, we can analyze the input question to identify the region of the input image that is relevant to the question. For example, if the input question is "Who is the author of this image?", analysis can determine that the region of the image related to the author information may be the title or footnote of the document.
[0086] The region image can be an image corresponding to a region in the input image that is relevant to the input question and identified using a pre-trained model or keyword matching. Alternatively, the region image can be an image of a region in the input image that is relevant to the input question and located using an object detection algorithm.
[0087] According to the embodiments of the present disclosure, by accurately locating the area image related to the input problem and screening the target visual features, the processing of irrelevant areas is avoided, and more focus can be placed on the image part related to the input problem, which reduces the waste of computing resources, improves processing efficiency, and effectively improves the response speed and accuracy to input problems.
[0088] According to the embodiments of the present disclosure, by flexibly selecting target visual features based on whether the input question involves local information, the focus can be dynamically adjusted according to the specific needs of the input question, taking into account both the integrity of the global information and the importance of local details, thereby effectively improving the pertinence and accuracy of image recognition.
[0089] In the embodiment of the present disclosure, the feature selection strategy and feature fusion strategy can be determined according to the recognition type corresponding to the input image and the input question. The recognition type refers to the recognition task category determined according to the input question and the image content. Figure 5 The process of determining feature selection strategy and feature fusion strategy based on recognition type is illustrated by way of example.
[0090] Figure 5 An example schematic diagram schematically illustrates a process for determining a feature selection strategy and a feature fusion strategy according to another embodiment of the present disclosure.
[0091] like Figure 5 As shown, in 500, taking the input image 510 and the input question 520 as an example, the determination process of the feature selection strategy 550 and the feature fusion strategy 570 is explained.
[0092] In one embodiment, a feature selection strategy 550 can be determined from a plurality of candidate feature selection strategies based on a recognition type 530 corresponding to an input image 510 and an input question 520. A candidate feature selection strategy refers to a predefined strategy for selecting a target feature from candidate visual features. For example, the plurality of candidate feature selection strategies may include a candidate feature selection strategy 541, a candidate feature selection strategy 542, ..., a candidate feature selection strategy 54S, where S is a positive integer.
[0093] The candidate feature selection strategies each have a corresponding candidate recognition type. According to the recognition type 530 , multiple candidate recognition types can be matched respectively, and the candidate feature selection strategy corresponding to the candidate recognition type matching the recognition type 530 is determined as the feature selection strategy 550 .
[0094] In another embodiment, a feature fusion strategy 570 can be determined from a plurality of candidate feature fusion strategies based on a recognition type 530 corresponding to an input image 510 and an input question 520. A candidate feature fusion strategy refers to a predefined strategy for fusing selected target visual features. For example, the plurality of candidate feature fusion strategies may include a candidate feature fusion strategy 561, a candidate feature fusion strategy 562, ..., a candidate feature fusion strategy 56T, where T is a positive integer.
[0095] The candidate feature fusion strategies each have a corresponding candidate recognition type. According to the recognition type 530 , multiple candidate feature fusion strategies can be matched separately, and the candidate feature fusion strategy corresponding to the candidate recognition type matching the recognition type 530 is determined as the feature fusion strategy 570 .
[0096] According to the embodiments of the present disclosure, by selecting corresponding feature selection strategies and feature fusion strategies according to the recognition type, this flexible strategy selection mechanism can adapt to different image recognition tasks, ensuring that key visual features can be effectively extracted and utilized when dealing with various types of image recognition problems, thereby improving the adaptability and accuracy of image recognition.
[0097] The above describes how to determine the feature selection strategy and feature fusion strategy, as well as how to obtain the fused visual features. Figure 6 For example, the overall process of image recognition disclosed herein is further described.
[0098] Figure 6 An example diagram of an image recognition process according to an embodiment of the present disclosure is schematically shown.
[0099] like Figure 6As shown, in 600 , after receiving an input image 601 and an input question 602 , visual feature extraction may be performed on the input image 601 to obtain a plurality of candidate visual features 603 .
[0100] Based on the input image 601 and the input question 602, a feature selection strategy 604 and a feature fusion strategy 606 are determined. Based on the feature selection strategy 604, at least two target visual features 605 are determined from multiple candidate visual features 603. Based on the feature fusion strategy 606, the multiple target visual features 605 are fused to obtain a fused visual feature 607.
[0101] In one embodiment, fusing multiple target visual features 605 to obtain a fused visual feature 607 may include the following operations: weighting the target visual features 605 according to their respective corresponding weights to obtain weighted visual features of the target visual features 605; and fusing the weighted visual features of the target visual features 605 to obtain a fused visual feature 607. A weight refers to the relative importance coefficient assigned to each target visual feature during the feature fusion process, used to measure its contribution to the final fusion result. For example, the higher the weight, the greater the impact of the target visual feature on the fusion result.
[0102] For example, target visual feature 6051, target visual feature 6052 and target visual feature 6053 are selected by feature selection strategy feature 604, and feature fusion strategy 606 indicates weight 6054 corresponding to target visual feature 6051, weight 6055 corresponding to target visual feature 6052 and weight 6056 corresponding to target visual feature 6053.
[0103] Thus, target visual feature 6051 can be weighted according to weight 6054 to obtain weighted visual feature 1 of target visual feature 6051; target visual feature 6052 can be weighted according to weight 6055 to obtain weighted visual feature 2 of target visual feature 6052; and target visual feature 6053 can be weighted according to weight 6056 to obtain weighted visual feature 3 of target visual feature 6053. On this basis, weighted visual feature 1, weighted visual feature 2, and weighted visual feature 3 can be fused to obtain fused visual feature 607. After obtaining fused visual feature 607, fused visual feature 607 can be input into the trained multimodal large model M610 to obtain image recognition result 608.
[0104] The fusion of weighted visual features can be achieved through weighted summation or feature concatenation, or by sampling such as a multi-layer perceptron (MLP) or a long short-term memory (LSTM) network, and performing nonlinear transformation on the weighted visual features. This is not limited here.
[0105] According to the embodiments of the present disclosure, through weighted fusion, the importance of different target visual features in the fusion process can be flexibly adjusted, so that the fused visual features can more accurately reflect the needs of the input problem, which not only improves the accuracy and pertinence of image recognition, but also can adapt to a variety of complex image recognition tasks, thereby enhancing the flexibility and practicality of image recognition.
[0106] The above are merely exemplary embodiments, but are not limited thereto. Other image recognition methods known in the art may also be included as long as they can improve the accuracy and pertinence of image recognition.
[0107] The above describes the image recognition method. Figure 7 For example, the training process of the multimodal large model disclosed in the present invention is further explained.
[0108] Figure 7 The flowchart of the training method of the multimodal large model according to the embodiment of the present disclosure is schematically shown.
[0109] like Figure 7 As shown, the training method 700 of the multimodal large model includes operations S710 to S730.
[0110] In operation S710, at least one sample target visual feature among the multiple layers of sample candidate visual features of the sample input image is fused according to an image recognition strategy corresponding to the sample input image and the sample input problem to obtain a sample fused visual feature, wherein the image recognition strategy indicates a selection method of the sample target visual feature and a fusion method for the sample target visual feature so that the sample fused visual feature adapts to the sample input problem.
[0111] In operation S720 , a sample image recognition result for the sample input question is determined based on the sample fused visual features.
[0112] In operation S730 , model parameters of the multimodal large model to be trained are adjusted according to the sample image recognition result and the reference image recognition result corresponding to the sample input image and the sample input question to obtain a trained multimodal large model.
[0113] For explanations of sample input images, sample input problems, image recognition strategies, sample candidate visual features, sample target visual features, sample fused visual features and sample image recognition results, please refer to the above content regarding input images, input problems, image recognition strategies, candidate visual features, target visual features, fused visual features and image recognition results, which will not be repeated here.
[0114] Reference image recognition results refer to predefined or annotated correct image recognition results, used to evaluate the accuracy of the model output. After obtaining the sample image recognition results output by the model, the difference between the reference and sample image recognition results can be compared based on the loss function. This difference can be used to adjust the model parameters of the multimodal large model to be trained until the predetermined termination conditions are met.
[0115] According to the embodiments of the present disclosure, through feature fusion based on image recognition strategy and parameter adjustment based on result feedback, the multimodal large model can dynamically adapt to different image recognition tasks, thereby improving the accuracy and adaptability of the model.
[0116] The above is a preliminary explanation of the training method of the multimodal large model. Figure 8 As an example, the overall training process of the multimodal large model disclosed in the present invention is further explained.
[0117] Figure 8 An example diagram of the training process of a multimodal large model according to an embodiment of the present disclosure is schematically shown.
[0118] like Figure 8 As shown, in 800 , after receiving a sample input image 801 and a sample input question 802 , visual features may be extracted from the sample input image 801 to obtain a plurality of sample candidate visual features 803 .
[0119] Based on the sample input image 801 and the sample input question 802, a feature selection strategy 804 and a feature fusion strategy 806 are determined. Based on the feature selection strategy 804, at least two sample target visual features 805 are determined from the multiple sample candidate visual features 803. Based on the feature fusion strategy 806, the multiple sample target visual features 805 are fused to obtain a sample fused visual feature 807.
[0120] In one embodiment, fusing multiple sample target visual features 805 to obtain sample fused visual features 807 may include the following operations: weighting the sample target visual features 805 according to the weights corresponding to the sample target visual features 805 to obtain sample weighted visual features of each sample target visual feature 805; and fusing the sample weighted visual features of each sample target visual feature 805 to obtain sample fused visual features 807. A weight refers to the relative importance coefficient assigned to each sample target visual feature during the feature fusion process, used to measure its contribution to the final fusion result. For example, the higher the weight, the greater the impact of the sample target visual feature on the fusion result.
[0121] For example, sample target visual feature 8051, sample target visual feature 8052 and sample target visual feature 8053 are selected through feature selection strategy feature 804, and feature fusion strategy 808 indicates weight 8054 corresponding to sample target visual feature 8051, weight 8055 corresponding to sample target visual feature 8052 and weight 8056 corresponding to sample target visual feature 8053.
[0122] Thus, sample target visual feature 8051 can be weighted according to weight 8054 to obtain weighted visual feature 1 of sample target visual feature 8051; sample target visual feature 8052 can be weighted according to weight 8055 to obtain weighted visual feature 2 of sample target visual feature 8052; and sample target visual feature 6053 can be weighted according to weight 8056 to obtain weighted visual feature 3 of sample target visual feature 8053. On this basis, sample weighted visual feature 1, sample weighted visual feature 2, and sample weighted visual feature 3 can be fused to obtain sample fused visual feature 807. After obtaining sample fused visual feature 807, sample fused visual feature 807 can be input into the trained multimodal large model M810 to obtain sample image recognition result 808.
[0123] After obtaining the sample image recognition result 808, the model parameters of the multimodal large model M810 to be trained can be adjusted according to the sample image recognition result 808 and the reference image recognition result 809 corresponding to the sample input image 801 and the sample input question 802 to obtain a trained multimodal large model.
[0124] The above are merely exemplary embodiments, but are not limited thereto. Other multimodal large model training methods known in the art may also be included, as long as they can improve the accuracy and adaptability of the model.
[0125] Based on the above image recognition method 200, the present invention also provides an image recognition device. Figure 9The device is described in detail.
[0126] Figure 9 The block diagram schematically shows an image recognition device according to an embodiment of the present disclosure.
[0127] like Figure 9 As shown, the image recognition device 900 may include a first obtaining module 910 and a first determining module 920 .
[0128] The first acquisition module 910 is used to fuse at least two target visual features from the multiple layers of candidate visual features of the input image according to an image recognition strategy corresponding to the input image and the input problem to obtain a fused visual feature, wherein the image recognition strategy indicates a method for selecting the target visual features and a method for fusing the target visual features so that the fused visual feature adapts to the input problem.
[0129] The first determination module 920 is configured to determine an image recognition result for the input question based on the fused visual features.
[0130] According to an embodiment of the present disclosure, the image recognition strategy includes a feature selection strategy and a feature fusion strategy.
[0131] According to an embodiment of the present disclosure, the first obtaining module 910 may include a determining submodule and a fusing submodule.
[0132] The first determination submodule is configured to determine at least two target visual features from a plurality of candidate visual features according to a feature selection strategy, wherein the feature selection strategy indicates at least one of the identity and number of the target visual features, and the difference between each two target visual features.
[0133] The fusion submodule is used to fuse multiple target visual features according to a feature fusion strategy to obtain a fused visual feature, wherein the feature fusion strategy indicates a weight for each target visual feature.
[0134] According to an embodiment of the present disclosure, the feature selection strategy and the feature fusion strategy are determined in the following manner: based on the correlation between the candidate visual features and the text features of the input question, a feature selection strategy is determined to determine at least two target visual features whose correlations meet preset correlation conditions among multiple candidate visual features; and based on the contribution of the target visual features to the input question, a feature fusion strategy is determined to make the weight and contribution of each target visual feature positively correlated.
[0135] According to an embodiment of the present disclosure, a plurality of candidate visual features are obtained by performing visual feature extraction processing on a plurality of regional images in an input image, respectively. The plurality of candidate visual features include local visual features related to information of the plurality of regional images in the input image, and global visual features obtained based on the local visual features.
[0136] According to an embodiment of the present disclosure, the image recognition apparatus 900 may include a third determination module and a fourth determination module.
[0137] The third determination module is used to determine the global visual feature as the target visual feature when the input question does not involve local information.
[0138] The fourth determination module is used to determine the global visual features and the local visual features of the regional image corresponding to the local information as target visual features when the input question involves local information.
[0139] According to an embodiment of the present disclosure, the fourth determining module may include a second determining submodule and a third determining submodule.
[0140] The second determining submodule is configured to determine at least one region image in the input image that is related to the input question.
[0141] The third determining submodule is configured to determine a candidate visual feature corresponding to at least one regional image among the plurality of candidate visual features as a target visual feature.
[0142] According to an embodiment of the present disclosure, a feature selection strategy and a feature fusion strategy are determined in the following manner: according to the recognition type corresponding to the input image and the input question, a feature selection strategy is determined from a plurality of candidate feature selection strategies, wherein the candidate feature selection strategies each have a corresponding candidate recognition type; and according to the recognition type corresponding to the input image and the input question, a feature fusion strategy is determined from a plurality of candidate feature fusion strategies, wherein the candidate feature fusion strategies each have a corresponding candidate recognition type.
[0143] According to an embodiment of the present disclosure, the fusion submodule may include a weighting unit and a fusion unit.
[0144] The weighting unit is used to weight the target visual features according to the weights corresponding to the target visual features to obtain weighted visual features of the target visual features.
[0145] The fusion unit is used to fuse the weighted visual features of the target visual features to obtain a fused visual feature.
[0146] According to an embodiment of the present disclosure, the input image includes a document image, and the input question includes a question related to document understanding.
[0147] Based on the above-mentioned multimodal large model training method 700, the present invention also provides a multimodal large model training device. Figure 10 The device is described in detail.
[0148] Figure 10 A block diagram of a training device for a multimodal large model according to an embodiment of the present disclosure is schematically shown.
[0149] like Figure 10 As shown, the training device 1000 for a multimodal large model may include a second obtaining module 1010 , a second determining module 1020 and a training module 1030 .
[0150] The second acquisition module 1010 is used to fuse at least two sample target visual features in the multi-layer sample candidate visual features of the sample input image according to an image recognition strategy corresponding to the sample input image and the sample input problem to obtain a sample fused visual feature, wherein the image recognition strategy indicates a selection method for the sample target visual features and a fusion method for the sample target visual features so that the sample fused visual feature adapts to the sample input problem.
[0151] The second determination module 1020 is configured to determine a sample image recognition result for the sample input problem based on the sample fusion visual features.
[0152] The training module 1030 is used to adjust the model parameters of the multimodal large model to be trained based on the sample image recognition results and the reference image recognition results corresponding to the sample input image and the sample input problem to obtain a trained multimodal large model.
[0153] Figure 11 The structural block diagram of the intelligent agent of the large model according to the embodiment of the present disclosure is schematically shown.
[0154] In the embodiments of the present disclosure, inspired by the von Neumann structure in modern computer theory, such as Figure 11 As shown, the AI agent 1100 may include five core modules: an input module 1110 , a control module 1120 , a storage module 1130 , a calculation module 1140 and an output module 1150 .
[0155] Input module 1110 is responsible for receiving or perceiving information such as queries, requests, instructions, signals, or data from the outside world (e.g., users or the external environment) and converting it into a format that AI agent 1100 can understand and process. Input module 1110 is the primary link for AI agent 1100 to interact with the outside world. It enables AI agent 1100 to efficiently and accurately obtain necessary "sensory" information from the outside world and respond to this information.
[0156] In an example, the input module 1110 may input the input image and input question described above.
[0157] Control module 1120 is the core support for AI agent 1100's ability to handle complex tasks. During the model training phase, control module 1120 can execute the multimodal large model training method described above; during the model inference phase, control module 1120 can execute the image recognition method described above.
[0158] In the example, the control module 1120 will continuously interact with the storage module 1130, the computing module 1140, and / or the output module 1150 during operation. However, it should be noted that in the embodiment of the present disclosure, the control module 1120 acts as a single initiator to initiate communication with the storage module 1130, the computing module 1140, and / or the output module 1150, and there is no communication coupling between the storage module 1130, the computing module 1140, and the output module 1150.
[0159] In this example, the performance of control module 1120 may be closely related to the large model underlying AI agent 1100. To fully leverage the capabilities of the large language model, the internal structure of control module 1120 may be designed to be highly configurable and extensible to handle a variety of different tasks and requirements in real-world scenarios.
[0160] The storage module 1130 may be responsible for memorizing the trained multimodal large model. The candidate feature selection strategy and candidate feature fusion strategy as described above may also be included in the storage module 1130.
[0161] In this example, after receiving the input image and input question, the AI agent 1100 can trigger the image recognition process, obtain the trained multimodal large model from the storage module 1130, and feed it back to the control module 1120. The control module 1120 can then pass the fed-back trained multimodal large model to the output module 1150.
[0162] The operation module 1140 can be regarded as a predefined tool library. The tools used for performing visual feature extraction as described above can be included in the operation module 1140.
[0163] In the example, when the AI agent 1100 needs to process data, it can call the relevant tools from the operation module 1140 and feed it back to the control module 1120. The control module 1120 can then use the fed-back tools to process the input image to obtain multiple candidate visual features. It is understandable that although the large language model has excellent language understanding and generation capabilities, it is just like a human, and the tasks it can solve without any tools are very limited. When the AI agent 1100 is given the ability to call tools, it can achieve tasks such as visual feature extraction with the help of tools for visual feature extraction.
[0164] During the model training phase, the output module 1150 may output the image recognition results described above.
[0165] The AI agent 1100 according to the embodiment of the present disclosure can simply and effectively improve the level of intelligence, and enhance flexibility and versatility.
[0166] Figure 12 A block diagram of an electronic device suitable for implementing an image recognition method and a training method for a multimodal large model according to an embodiment of the present disclosure is schematically shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital processing, cellular phones, smart phones, wearable devices and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present disclosure described and / or required herein.
[0167] like Figure 12 As shown, device 1200 includes a computing unit 1201, which can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 1202 or a computer program loaded from a storage unit 1208 into a random access memory (RAM) 1203. RAM 1203 may also store various programs and data required for the operation of device 1200. Computing unit 1201, ROM 1202, and RAM 1203 are connected to each other via a bus 1204. An input / output (I / O) interface 1205 is also connected to bus 1204.
[0168] Various components in device 1200 are connected to I / O interface 1205, including an input unit 1206, such as a keyboard and mouse; an output unit 1207, such as various types of displays and speakers; a storage unit 1208, such as a magnetic disk and optical disk; and a communication unit 1209, such as a network card, a modem, a wireless communication transceiver, etc. Communication unit 1209 allows device 1200 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.
[0169] Computing unit 1201 can be any general-purpose and / or specialized processing component with processing and computing capabilities. Some examples of computing unit 1201 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various specialized artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. Computing unit 1201 performs the various methods and processes described above, such as the image recognition method and the multimodal large model training method. For example, in some embodiments, the image recognition method and the multimodal large model training method can be implemented as a computer software program tangibly embodied in a machine-readable medium, such as storage unit 1208. In some embodiments, part or all of the computer program can be loaded and / or installed onto device 1200 via ROM 1202 and / or communication unit 1209. When the computer program is loaded into RAM 1203 and executed by computing unit 1201, one or more steps of the image recognition method and the multimodal large model training method described above can be performed. Alternatively, in other embodiments, the computing unit 1201 may be configured to execute an image recognition method or a multimodal large model training method in any other appropriate manner (eg, by means of firmware).
[0170] Various embodiments of the systems and techniques described above can be implemented in digital electronic circuit systems, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), system-on-a-chip systems (SOCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include being implemented in one or more computer programs that are executable and / or interpreted on a programmable system that includes at least one programmable processor, which can be a special purpose or general purpose programmable processor that can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit data and instructions to the storage system, the at least one input device, and the at least one output device.
[0171] The program code for implementing the method of the present disclosure can be written in any combination of one or more programming languages. These program codes can be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing device so that when the program code is executed by the processor or controller, the functions / operations specified in the flow chart and / or block diagram are implemented. The program code can be executed entirely on the machine, partially on the machine, as a stand-alone software package, partially on the machine and partially on a remote machine, or entirely on a remote machine or server.
[0172] In the context of the present disclosure, a machine-readable medium may be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium may be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium may include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination of the foregoing. More specific examples of machine-readable storage media may include an electrical connection based on one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), optical fibers, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0173] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user can provide input to the computer. Other types of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).
[0174] The systems and techniques described herein can be implemented in a computing system that includes back-end components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes front-end components (e.g., a user computer with a graphical user interface or a web browser through which a user can interact with implementations of the systems and techniques described herein), or a computing system that includes any combination of such back-end components, middleware components, or front-end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include a local area network (LAN), a wide area network (WAN), and the Internet.
[0175] A computer system may include a client and a server. The client and server are generally remote from each other and typically interact through a communication network. The client-server relationship arises by virtue of computer programs running on the respective computers and having a client-server relationship to each other. The server may be a cloud server, a server in a distributed system, or a server integrated with a blockchain.
[0176] It should be understood that the various forms of the processes shown above can be used to reorder, add, or delete steps. For example, the steps described in this disclosure can be performed in parallel, sequentially, or in a different order, as long as the desired results of the technical solutions disclosed in this disclosure can be achieved. This is not a limitation herein.
[0177] The above specific embodiments do not constitute a limitation on the scope of protection of this disclosure. Those skilled in the art will appreciate that various modifications, combinations, sub-combinations, and substitutions may be made based on design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this disclosure shall be included within the scope of protection of this disclosure.
Claims
1. An image recognition method, comprising: fusing at least two target visual features from a plurality of candidate visual features of the input image according to an image recognition strategy corresponding to an input image and an input question to obtain a fused visual feature, wherein the image recognition strategy indicates a method for selecting the target visual features and a method for fusing the target visual features, so that the fused visual feature is adapted to the input question; and An image recognition result for the input question is determined based on the fused visual features.
2. The method according to claim 1, wherein The image recognition strategy includes a feature selection strategy and a feature fusion strategy; The step of fusing at least two target visual features from a plurality of candidate visual features of the input image according to an image recognition strategy corresponding to the input image and the input question to obtain a fused visual feature comprises: determining the at least two target visual features from the plurality of candidate visual features according to the feature selection strategy, wherein the feature selection strategy indicates at least one of an identity, a number, and a degree of difference between each two of the target visual features; and According to the feature fusion strategy, multiple target visual features are fused to obtain the fused visual feature, wherein the feature fusion strategy indicates a weight for each of the target visual features.
3. The method according to claim 2, wherein: The feature selection strategy and the feature fusion strategy are determined in the following manner: determining, based on the correlation between the candidate visual features and the text features of the input question, the feature selection strategy to determine, from the plurality of candidate visual features, the at least two target visual features whose correlations satisfy a preset correlation condition; as well as The feature fusion strategy is determined according to the contribution of the target visual features to the input question, so that the weight and contribution of each target visual feature are positively correlated.
4. The method according to claim 3, wherein: The plurality of candidate visual features are obtained by performing visual feature extraction processing on a plurality of regional images in the input image, respectively, and the plurality of candidate visual features include local visual features related to information of the plurality of regional images in the input image, and global visual features obtained based on the local visual features; The method further comprises: In a case where the input question does not involve local information, determining the global visual feature as the target visual feature; and In a case where the input question involves the local information, the global visual feature and the local visual feature of the regional image corresponding to the local information are determined as the target visual feature.
5. The method according to claim 4, wherein The determining the global visual feature and the local visual feature corresponding to the local information as the target visual feature includes: determining at least one region image in the input image that is relevant to the input question; and A candidate visual feature corresponding to the at least one region image among the plurality of candidate visual features is determined as the target visual feature.
6. The method according to claim 2, wherein: The feature selection strategy and the feature fusion strategy are determined in the following manner: determining the feature selection strategy from a plurality of candidate feature selection strategies according to the recognition type corresponding to the input image and the input question, wherein each of the candidate feature selection strategies has a corresponding candidate recognition type; and The feature fusion strategy is determined from a plurality of candidate feature fusion strategies according to the recognition type corresponding to the input image and the input question, wherein each of the candidate feature fusion strategies has a corresponding candidate recognition type.
7. The method according to any one of claims 2 to 6, wherein The step of fusing the at least two target visual features according to the feature fusion strategy to obtain the fused visual feature includes: weighting the target visual features according to the weights corresponding to the respective target visual features to obtain weighted visual features of the respective target visual features; and The weighted visual features of the target visual features are fused to obtain the fused visual feature.
8. The method according to claim 1, wherein The input image includes a document image, and the input question includes a question related to document understanding.
9. A method for training a large multimodal model, comprising: According to an image recognition strategy corresponding to a sample input image and a sample input problem, at least one sample target visual feature among the multiple layers of sample candidate visual features of the sample input image is fused to obtain a sample fused visual feature, wherein the image recognition strategy indicates a selection method for the sample target visual feature and a fusion method for the sample target visual feature, so that the sample fused visual feature is adapted to the sample input problem; determining a sample image recognition result for the sample input problem based on the sample fused visual features; and According to the sample image recognition result and the reference image recognition result corresponding to the sample input image and the sample input problem, the model parameters of the multimodal large model to be trained are adjusted to obtain a trained multimodal large model.
10. An image recognition device, comprising: a first obtaining module, configured to fuse at least two target visual features from the multiple layers of candidate visual features of the input image according to an image recognition strategy corresponding to the input image and the input question to obtain a fused visual feature, wherein the image recognition strategy indicates a method for selecting the target visual features and a method for fusing the target visual features, so that the fused visual feature is adapted to the input question; and The first determination module is used to determine an image recognition result for the input question based on the fused visual features.
11. A multimodal large model training device, comprising: a second obtaining module, configured to fuse at least two sample target visual features from the multiple layers of sample candidate visual features of the sample input image according to an image recognition strategy corresponding to the sample input image and the sample input problem, to obtain a sample fused visual feature, wherein the image recognition strategy indicates a selection method for the sample target visual features and a fusion method for the sample target visual features, so that the sample fused visual feature is adapted to the sample input problem; A second determination module is configured to determine a sample image recognition result for the sample input problem based on the sample fusion visual features; and A training module is used to adjust the model parameters of the multimodal large model to be trained based on the sample image recognition results and the reference image recognition results corresponding to the sample input image and the sample input problem to obtain a trained multimodal large model.
12. An artificial intelligence agent configured to execute the method according to any one of claims 1 to 9.
13. An electronic device comprising: one or more processors; a memory for storing one or more computer programs, It is characterized in that the one or more processors execute the one or more computer programs to implement the steps of the method according to any one of claims 1 to 9.
14. A computer-readable storage medium having a computer program or instruction stored thereon, characterized in that: When the computer program or instruction is executed by a processor, the steps of the method according to any one of claims 1 to 9 are implemented.
15. A computer program product comprising a computer program or instructions, characterized in that When the computer program or instruction is executed by a processor, the steps of the method according to any one of claims 1 to 9 are implemented.
Citation Information
Cited By
Illegal content auditing method and device based on multi-modal data, equipment and medium
CN121278126A