Image recognition model training method, question and answer method based on large model and intelligent agent

By combining image recognition model and multimodal model, the sample images are comprehensively processed and trained, and the existing multimodal large language model is solved, and the existing multimodal large language model lacks visual recognition capabilities in specific visual fields is achieved, achieving stronger robustness and generalization.

CN120197705APending Publication Date: 2025-06-24BEIJING BAIDU NETCOM SCI & TECH CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510337872.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-20
Publication Date
2025-06-24

AI Technical Summary

Technical Problem

Existing multimodal large language models lack visual recognition and perception capabilities in specific visual fields, especially in fine-grained or professional fields of tasks, with poor generalization performance.

Method used

By combining image recognition model and multimodal model, the image recognition model is used to initially recognize sample images, and the multimodal model is used to process preset prompt words and sample images, and the image recognition model is comprehensively trained to enhance its ability to perceive and understand the subject's depth.

Benefits of technology

Improve the robustness and generalization of image recognition models in specific visual fields, and can better solve visual recognition tasks in fine-grained or professional fields.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120197705A_ABST
    Figure CN120197705A_ABST
Patent Text Reader

Abstract

The invention discloses an image recognition model training method, a large-model-based question and answer method and an intelligent agent. The invention provides a training method of an image recognition model, relates to the technical field of artificial intelligence, in particular to the technical fields of computer vision, deep learning, large models and the like, and can be applied to intelligent question and answer scenes. The training method comprises the steps that image recognition processing is conducted on a sample image through an image recognition model, a first sample image category for the sample image is obtained, and the sample image category represents the category to which a target object in the sample image belongs; performing image recognition processing on a preset cue word and the sample image by using the multi-modal model to obtain a second sample image category for the sample image; and training an image recognition model according to the first sample image category, the second sample image category and the label of the sample image to obtain a trained image recognition model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of artificial intelligence technology, particularly to technical fields such as computer vision, deep learning, large models, etc., and can be applied to intelligent question and answer scenarios. Specifically, it relates to a method for training a model, a question and answer method based on a large model, and an intelligent agent. Background Art

[0002] With the rapid development of artificial intelligence, multimodal large language models (MLLMs, Multimodal Large Language Models) have shown great potential in multimodal content understanding and human conversations. With its massive data training and large parameter scale, it has broad application space in various fields such as natural language processing, computer vision, and artificial intelligence generation. Therefore, conducting in-depth research on large models, exploring their greater potential, promoting their implementation in more scenarios, and realizing greater value are of great significance for the intelligent transformation of various industries. Summary of the Invention

[0003] The present disclosure provides a method for training a model, a question and answer method based on a large model, and an intelligent agent.

[0004] According to one aspect of the present disclosure, there is provided a method for training an image recognition model, including: performing image recognition processing on a sample image using the image recognition model to obtain a first sample image category for the sample image, where the sample image category represents the category to which the target object in the sample image belongs; using a multimodal model to perform image recognition processing on a preset prompt and the sample image to obtain a second sample image category for the sample image; and training the image recognition model based on the first sample image category, the second sample image category, and the label of the sample image to obtain a trained image recognition model.

[0005] According to another aspect of the present disclosure, there is provided a question and answer method based on a large model, including: obtaining question information, where the question information includes at least one target image; using the trained image recognition model to perform image recognition processing on the at least one target image to obtain a target image category for each of the at least one target image, where the target image category represents the category to which the target object in the target image belongs, and the trained image recognition model is trained based on the first sample image category obtained by performing image recognition processing on the sample image using the image recognition model, the second sample image category obtained by performing image recognition processing on the preset prompt and the sample image using the multimodal model, and the label of the sample image; using the large model to process the question information and the enhanced information of the target image to obtain target answer information for the question information, where the enhanced information of the target image is determined based on the target image category.

[0006] According to another aspect of the present disclosure, an intelligent agent of artificial intelligence is provided, which is configured to execute the method as above.

[0007] According to another aspect of the present disclosure, a training device for an image recognition model is provided, including: a first processing module, configured to perform image recognition processing on a sample image by using the image recognition model to obtain a first sample image category for the sample image, where the sample image category represents the category to which the target object in the sample image belongs; a second processing module, configured to perform image recognition processing on a preset prompt word and the sample image by using a multimodal model to obtain a second sample image category for the sample image; and a training module, configured to train the image recognition model according to the first sample image category, the second sample image category, and the label of the sample image to obtain a trained image recognition model.

[0008] According to another aspect of the present disclosure, a question-answering device for a large model is provided, including: a first acquisition module, configured to acquire question information, where the question information includes at least one target image; a third processing module, configured to perform image recognition processing on the at least one target image by using the trained image recognition model to obtain a target image category for each of the at least one target image, where the target image category represents the category to which the target object in the target image belongs, and the trained image recognition model is trained by using the first sample image category obtained by performing image recognition processing on the sample image by using the image recognition model, the second sample image category obtained by performing image recognition processing on the preset prompt word and the sample image by using the multimodal model, and the label of the sample image; and a fourth processing module, configured to process the question information and the enhanced information of the target image by using the large model to obtain target answer information for the question information, where the enhanced information of the target image is determined according to the target image category.

[0009] According to another aspect of the present disclosure, an electronic device is provided, including: at least one processor; and a memory communicatively connected to the at least one processor; where the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the method as above.

[0010] According to another aspect of the present disclosure, a non-transitory computer-readable storage medium storing computer instructions is provided, where the computer instructions are used to cause a computer to execute the method as above.

[0011] According to another aspect of the present disclosure, a computer program product is provided, including a computer program that implements the method as above when executed by a processor.

[0012] It should be understood that the content described in this part is not intended to identify the key or important features of the embodiments of the present disclosure, nor is it used to limit the scope of the present disclosure. Other features of the present disclosure will become readily understood through the following description. Description of the Drawings

[0013] The drawings are used to better understand the present solution and do not constitute a limitation to the present disclosure. Among them:

[0014] Figure 1 Schematically shows an exemplary system architecture of a training method of an applicable model, a question-and-answer method based on a large model, and an intelligent agent according to an embodiment of the present disclosure;

[0015] Figure 2 Schematically shows a flowchart of a training method of an image recognition model according to an embodiment of the present disclosure;

[0016] Figure 3 Schematically shows a schematic diagram of determining a first sample image category according to an embodiment of the present disclosure;

[0017] Figure 4 Schematically shows a schematic diagram of extracting sample image features according to an embodiment of the present disclosure;

[0018] Figure 5 Schematically shows a schematic diagram of determining a second sample image category according to an embodiment of the present disclosure;

[0019] Figure 6 Schematically shows a schematic diagram of a training method of an image recognition model according to an embodiment of the present disclosure;

[0020] Figure 7 Schematically shows a schematic diagram of a training method of an image recognition model according to another embodiment of the present disclosure;

[0021] Figure 8 Schematically shows a flowchart of a question-and-answer method based on a large model according to an embodiment of the present disclosure;

[0022] Figure 9 Schematically shows a schematic diagram of a question-and-answer method based on a large model according to an embodiment of the present disclosure;

[0023] Figure 10 Schematically shows a schematic diagram of a method for generating a target prompt word according to an embodiment of the present disclosure;

[0024] Figure 11 Schematically shows a block diagram of a training device of an image recognition model according to an embodiment of the present disclosure;

[0025] Figure 12 Schematically shows a block diagram of a question-and-answer device based on a large model according to an embodiment of the present disclosure;

[0026] Figure 13 Schematically shows a structural block diagram of an agent of artificial intelligence according to an embodiment of the present disclosure; and

[0027] Figure 14 Schematically shows a schematic block diagram of an example electronic device that can be used to implement the embodiments of the present disclosure. Detailed implementation manners

[0028] The following makes an explanation of exemplary embodiments of the present disclosure in conjunction with the accompanying drawings, including various details of the embodiments of the present disclosure to facilitate understanding, which should be considered merely exemplary. Therefore, those of ordinary skill in the art should recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of the present disclosure. Similarly, descriptions of well-known functions and structures are omitted in the following description for clarity and conciseness.

[0029] In the technical solution of the present disclosure, the processing of collection, storage, use, processing, transmission, provision, disclosure, and application of the user's personal information involved all comply with the provisions of relevant laws and regulations, adopt necessary confidentiality measures, and do not violate public order and good customs.

[0030] In the technical solution of the present disclosure, before obtaining or collecting the user's personal information, the authorization or consent of the user has been obtained.

[0031] To clearly describe the technical solution of the present disclosure, first, the nouns involved in the present disclosure are defined:

[0032] Multimodal large model: It is a large artificial intelligence system that can understand and integrate various types of data (such as text, images, videos, voices, etc.). It realizes stronger information processing capabilities through joint learning of cross-modal representations.

[0033] Prompt: Keywords or sentences used to guide the large model to accurately understand the task requirements and make corresponding responses.

[0034] Token: For text, it is the basic unit of data segmentation, which can be words, characters, or sub-words. For visual information, it is the feature vector after encoding of one piece after image cutting.

[0035] In the current context of the rapid development of information technology, large models, as the core technology of artificial intelligence, demonstrate powerful capabilities. Multimodal large language models (MLLMs) show great potential in multimodal (graphic-text) content understanding and human conversations.

[0036] In the process of implementing the present disclosure, it is found that since the basic technology of the multimodal large language model is based on the large language model (LLM), it has great limitations in visual perception ability. In some specific visual fields, the visual recognition and perception ability is still not as good as that of pure visual models. Here, the specific visual fields may include fields of fine-grained or specialized visual recognition and perception tasks, such as the field of medical image analysis, the field of microscopic image analysis, the field of transportation and autonomous driving, etc.

[0037] In related examples, Q-Former is used to encode visual features and text features into a fused feature as the input of the retrieval model, and the output of the retrieval model is given to the MLLM for content enhancement. However, this method supplements general knowledge and is not specialized enough.

[0038] In another example, the image-text alignment pre-trained CLIP (Contrastive Language-Image Pre-training) model is used and used as a basic retriever to retrieve content related to the question information from the information source, and then the MLLM model is used to re-rank the retrieval results to improve the retrieval accuracy, thereby improving the information reliability of the multimodal model. However, the retrieval model in this method has poor generalization performance. Only using the retrieval model trained on general tasks as a tool for information mining, it is difficult to generalize well in the information source in specific scenarios.

[0039] The present disclosure provides a method for training an image recognition model, including: using the image recognition model to perform image recognition processing on a sample image to obtain a first sample image category for the sample image, where the sample image category represents the category to which the target object in the sample image belongs; using a multimodal model to perform image recognition processing on a preset prompt word and the sample image to obtain a second sample image category for the sample image; and training the image recognition model according to the first sample image category, the second sample image category and the label of the sample image to obtain a trained image recognition model.

[0040] According to embodiments of the present disclosure, large models generally possess knowledge in multiple domains and the ability to handle tasks in multiple domains. However, the large model's ability to handle complex tasks in a specific domain is relatively poor compared to visual models in that specific domain. Visual models in a specific domain are restricted by the amount of labeled training data, and the model's generalization ability is weak, making it prone to overfitting. On the other hand, multi-modal large models, due to not being restricted by the amount of labeled training data and with a vast amount of unlabeled data participating in pre-training, possess strong generalization and transfer capabilities. Therefore, the present disclosure trains a picture recognition model for a specific domain based on sample pictures by utilizing a multi-modal model, which helps enhance the image recognition model's ability to deeply perceive and understand the subject, and also helps enhance the robustness and generalization of the image recognition model, thereby enabling better resolution of fine-grained or professional domain visual recognition tasks.

[0041] Figure 1 Schematically shows an exemplary system architecture to which the training method of the model, the large model-based question-and-answer method, and the agent according to embodiments of the present disclosure can be applied.

[0042] It should be noted that Figure 1 What is shown is only an example of the system architecture to which embodiments of the present disclosure can be applied, to help those skilled in the art understand the technical content of the present disclosure. However, it does not mean that embodiments of the present disclosure cannot be used in other devices, systems, environments, or scenarios. For example, in another embodiment, the exemplary system architecture to which the training method of the model, the large model-based question-and-answer method, and the agent can be applied may include a terminal device, but the terminal device can implement the model training method, the large model-based question-and-answer method, and the agent provided by embodiments of the present disclosure without interacting with the server.

[0043] As Figure 1 As shown, the system architecture 100 according to this embodiment may include a first terminal device 101, a second terminal device 102, a third terminal device 103, a network 104, and a server 105. The network 104 is used to provide a medium for communication links between the first terminal device 101, the second terminal device 102, the third terminal device 103, and the server 105. The network 104 may include various connection types, such as wired and / or wireless communication links, etc.

[0044] Users can use the first terminal device 101, the second terminal device 102, and the third terminal device 103 to interact with the server 105 through the network 104 to receive or send messages, etc. Various communication client applications may be installed on the first terminal device 101, the second terminal device 102, and the third terminal device 103, such as knowledge reading applications, web browser applications, search applications, instant messaging tools, email clients, and / or social platform software, etc. (only as examples).

[0045] The first terminal device 101, the second terminal device 102, and the third terminal device 103 may be various electronic devices having display screens and supporting web browsing, including but not limited to smart phones, tablet computers, laptop computers, desktop computers, and the like.

[0046] The server 105 may be a server that provides various services, such as a background management server (only as an example) that provides support for the content browsed by the user using the first terminal device 101, the second terminal device 102, and the third terminal device 103. The background management server may analyze and process the received data such as user requests, and feed back the processing results (such as web pages, information, or data obtained or generated according to user requests) to the terminal device.

[0047] It should be noted that the model training method and the large model-based question-answering method provided in the embodiments of the present disclosure can generally be executed by the first terminal device 101, the second terminal device 102, and the third terminal device 103. Correspondingly, the model training device and the large model-based question-answering device provided in the embodiments of the present disclosure can also be set in the first terminal device 101, the second terminal device 102, and the third terminal device 103.

[0048] Alternatively, the model training method and the large model-based question-answering method provided in the embodiments of the present disclosure may also be generally performed by the server 105. Accordingly, the model training device and the large model-based question-answering device provided in the embodiments of the present disclosure may generally be arranged in the server 105. The model training method and the large model-based question-answering method provided in the embodiments of the present disclosure may also be performed by a server or server cluster that is different from the server 105 and can communicate with the first terminal device 101, the second terminal device 102, the third terminal device 103 and / or the server 105. Accordingly, the model training device and the large model-based question-answering device provided in the embodiments of the present disclosure may also be arranged in a server or server cluster that is different from the server 105 and can communicate with the first terminal device 101, the second terminal device 102, the third terminal device 103 and / or the server 105.

[0049] It should be understood that Figure 1 The number of terminal devices, networks and servers in the embodiment is only for illustration. Any number of terminal devices, networks and servers may be provided according to the implementation requirements.

[0050] The model training method, large model-based question-answering method and intelligent agent provided by the present disclosure can be applied to any multimodal intelligent question system based on images and texts, helping it to improve the authority and timeliness of the content generated by the multimodal model. For example, it can be applied to many application scenarios such as intelligent visual question-answering based on multimodal images and texts, intelligent medical health, cross-modal information search, cross-modal recommendation system, etc.

[0051] Figure 2 Schematically shows a flowchart of a method for training an image recognition model according to an embodiment of the present disclosure.

[0052] As Figure 2 shown, the training method 200 of this embodiment includes operations S210 to S230.

[0053] In operation S210, an image recognition model is used to perform image recognition processing on a sample image to obtain a first sample image category for the sample image.

[0054] The sample image can be a sample image in any field. For example, the sample image can be an image in the field of medical image analysis such as an ultrasound image, a pathological section image, etc. Again, the sample image can also be an image in the agricultural field such as images of different kinds of plants. Again, the sample image can also be a product image such as the sample image can be images of different kinds of cars, mobile phones, etc.

[0055] It should be noted that the training set composed of sample images can include only sample images in a single field, or can include sample images in multiple fields.

[0056] Exemplarily, the training set can include only images of different kinds of cars, or can include different car images and images of different kinds of plants. In the case where the training set includes only images of different kinds of cars, the obtained image recognition model is a visual model for the automotive field. In the case where the training set includes different car images and images of different kinds of plants, the obtained image recognition model is a visual model for the automotive field and the plant field.

[0057] The sample image category represents the category to which the target object in the sample image belongs. Among them, the target object can be the main object in the target image. For example, if the target image contains a cat, the target object can be the cat; if the target image contains a person, the target object can be the person; if the target object contains a plant, the target object can be the plant.

[0058] Exemplarily, if the sample image is an image containing a car, the target object can be the car, and the sample image category can be the model, brand, etc. of the car. Again, if the sample image is an image containing a cat, the target object can be the cat, and the sample image category can be the breed of the cat. Again, if the sample image is a pathological section image, the target object can be the pathological section, and the sample image category can be the category to which the pathological section belongs, such as a lung nodule.

[0059] In operation S220, a multimodal model is used to perform image recognition processing on a preset prompt and a sample image to obtain a second sample image category for the sample image.

[0060] The purpose of the multi-modal model to perform image recognition processing on the sample image is to determine the category of the target object in the sample image. Therefore, the preset prompt can be used to determine the target object in the sample image. For example, the prompt can be "What is the main body in the image?".

[0061] In operation S230, according to the first sample image category, the second sample image category, and the label of the sample image, train the image recognition model to obtain the trained image recognition model.

[0062] According to the embodiments of the present disclosure, by respectively using the image recognition model and the multi-modal model to perform image recognition on the sample image to obtain the first sample image category and the second sample image category of the sample image, and then fixing the parameters of the multi-modal model, using the first sample image category and the second sample image category to train the image recognition model to obtain the trained image recognition model. By using the multi-modal model, a picture recognition model in a specific field is trained according to the sample pictures, which helps to enhance the ability of the image recognition model to deeply perceive and understand the main body, and helps to enhance the robustness and generalization of the image recognition model, so as to better solve the fine-grained or professional field visual recognition tasks.

[0063] The following refers to Figures 3 - 7 , and further illustrates the Figure 2 shown method in combination with specific embodiments.

[0064] According to the embodiments of the present disclosure, the image recognition model includes a feature extraction module and an image recognition module; using the image recognition model to perform image recognition processing on the sample image to obtain the first sample image category of the sample image, including: using the feature extraction module to extract the sample image features of the sample image, where the sample image features include the global features and the local features of the sample image; using the image recognition module to perform image recognition processing on the sample image features to obtain the first sample image category.

[0065] According to the embodiments of the present disclosure, the image recognition model may further include an image encoding module. The image encoding module is used to transform the sample image into a token sequence, that is, divide the sample image into multiple pixel blocks. Based on this, extracting the sample image features of the sample image may include: performing feature extraction on the multiple sample pixel blocks obtained by dividing the sample image to obtain the respective local features of the sample pixel blocks; splicing the respective local features of the sample pixel blocks according to the relative position relationship between the sample pixel blocks to obtain the global features.

[0066] The pixel block can be a small square in the sample image. These small squares have clear positions and assigned color values, and the color and position of the small squares determine the appearance of the sample image.

[0067] The multiple sample pixel blocks obtained by dividing the sample image may include: dividing the sample image into multiple pixel blocks with overlapping regions. For example, after dividing the sample image, it includes pixel block 1, pixel block 2, pixel block 3, and pixel block 4. Among them, the same region 1 may be included in pixel block 1 and pixel block 2, the same region 2 may be included in pixel block 2 and pixel block 3, and the same region 3 may be included in pixel block 3 and pixel block 4.

[0068] Feature extraction is performed on the multiple sample pixel blocks to obtain the respective local features of the sample pixel blocks, which may include: based on the self-attention mechanism, each pixel block is processed to obtain the local feature corresponding to each pixel block.

[0069] For example, after the sample image is divided into pixel block 1, pixel block 2, pixel block 3, and pixel block 4, then based on the self-attention mechanism, processing each pixel block to obtain the local feature corresponding to each pixel block may include: based on the self-attention mechanism, processing pixel block 1 to obtain the local feature 1 corresponding to pixel block 1; based on the self-attention mechanism, processing pixel block 2 to obtain the local feature 2 corresponding to pixel block 2; based on the self-attention mechanism, processing pixel block 3 to obtain the local feature 3 corresponding to pixel block 3; based on the self-attention mechanism, processing pixel block 4 to obtain the local feature 4 corresponding to pixel block 4.

[0070] According to an embodiment of the present disclosure, since each pixel block has image coordinate information. After obtaining the local feature corresponding to each pixel block, the relative position relationship between them can be determined according to the image coordinate information of the pixel block and stitched together to obtain the global feature.

[0071] For example, as described above, the sample image includes pixel block 1, pixel block 2, pixel block 3, and pixel block 4. Among them, the same region 1 is included in pixel block 1 and pixel block 2, the same region 2 is included in pixel block 2 and pixel block 3, and the same region 3 is included in pixel block 3 and pixel block 4. Therefore, the regional feature a of region 1 is included in the local feature 1 corresponding to pixel block 1, and the regional feature a of region 1 is also included in the local feature 2 corresponding to pixel block 2. At this time, according to the regional feature a, the relative position of the local feature 1 and the local feature 2 can be determined, so as to determine the relative position of pixel block 1 and pixel block 2. According to this relative position relationship, pixel block 1 and pixel block 2 can be stitched together to obtain the image feature at the positions of pixel block 1 and pixel block 2; according to this method, after stitching pixel block 1, pixel block 2, pixel block 3, and pixel block 4, the global feature can be obtained.

[0072] According to an embodiment of the present disclosure, since the global features are used to determine the general category of an object and the local features are refined to specific sub-categories, by using the feature extraction module to separately obtain the global features and local features of the sample image, it helps to improve the robustness, recognition accuracy, and generalization ability of the image recognition model.

[0073] Figure 3 Schematically shows a diagram for determining the category of the first sample image according to an embodiment of the present disclosure.

[0074] As Figure 3 shown, in this embodiment 300, the sample image 310 is input into the image recognition model 320. First, the image encoding module 321 divides the sample image 310 to obtain a plurality of pixel blocks; then the plurality of pixel degrees are input into the feature extraction module 322 for feature extraction, and the sample image features are output, and the sample image features may include the local features of each pixel block and the global features of the sample image; thereafter, the local features and the global features are input into the image recognition module 323 for image recognition, and the first sample image category 330 for the sample image is output.

[0075] The feature extraction module may include a plurality of Vision Transformer Blocks, and the global features may further include the global features output by each Transformer Block. Specifically, the token sequence of the sample image processed by the image encoding module is input into a plurality of Vision Transformer Blocks to obtain local features, and at the same time, after each Vision Transformer Block, the cls token of the intermediate result is extracted to form the global features.

[0076] Figure 4 Schematically shows a diagram for extracting the sample image features according to an embodiment of the present disclosure.

[0077] As Figure 4 shown, the feature extraction module of this embodiment 400 includes N + 1 Transformer Blocks, namely Block0 to BlockN.

[0078] The specific process may include: inputting the token sequence of the sample image into the feature extraction module 410, first extracting the global feature 422 in the token sequence, where the global feature 422 is obtained by splicing local features 421; then the token sequence is processed by Block0 and outputs feature 430, and feature 430 includes local feature 431 and global feature 432, where the global feature 432 is obtained by splicing local features 431, and then the global feature 432 is extracted. Repeat this process. Each time passing through a Block, a global feature can be extracted until the token sequence is processed by BlockN and outputs local feature 422, and the local feature 422 is spliced to obtain the global feature for BlockN. Then the extracted global features are spliced to obtain feature 441, and feature 411 is spliced with local feature 44 again to obtain the sample image feature 440 of the sample image.

[0079] According to an embodiment of the present disclosure, the above method further includes: performing average fusion processing on the local features and global features of each sample pixel block to obtain a sample image fusion feature; using an image recognition module to perform image recognition processing on the sample image feature, including: using the image recognition module to perform image recognition processing on the sample image fusion feature.

[0080] By performing average fusion processing on the local features and fusion features, the sample image feature can take into account both local features and global features, be able to spatially align different features, weaken the influence of random errors and outliers, reduce noise, and help improve the recognition accuracy.

[0081] According to an embodiment of the present disclosure, the multimodal model includes a modality conversion model and a large language model; using the multimodal model to perform image recognition processing on a preset prompt and a sample image, including: using the modality conversion model to perform modality conversion processing on the sample image feature to obtain a sample image language modality feature; using the large language model to perform image recognition processing on the preset prompt and the sample image language modality feature.

[0082] Using the large language model to perform image recognition processing on the preset prompt and the sample image language modality feature may include: splicing the preset prompt such as "What is the main body in the image?" with the sample image recognition model feature and then inputting it to the large language model to make the large language model answer the category of the sample image.

[0083] Figure 5 Schematically shows a schematic diagram for determining the second sample image category according to an embodiment of the present disclosure.

[0084] As Figure 5As shown, in Embodiment 500, the sample image feature 510 is input into the modality conversion model 520. After converting the sample image feature 510 from the image modality feature to the language modality feature, the sample image language modality feature 530 is output; then the sample image language modality feature 530 is concatenated with the preset prompt 540 to obtain the concatenated information 550; the concatenated information 550 is input into the LLM model 560, and the second sample image category 570 for the sample image is output.

[0085] By converting the sample image feature from the image modality feature to the language modality feature and combining it with a fixed prompt for VQA (Visual Question Answering) task learning, the strong generalization ability of the large language model can be utilized to enhance the depth perception and understanding of the subject by the image recognition model, and the robustness and generalization of the image recognition model can be enhanced.

[0086] According to an embodiment of the present disclosure, an image recognition model is trained based on the first sample image category, the second sample image category, and the label of the sample image to obtain a trained image recognition model, including: processing the first sample image category and the label of the sample image using a first loss function to obtain a first loss value; processing the second sample image category and the label of the sample image using a second loss function to obtain a second loss value; and based on the first loss value and the second loss value, adjusting the model parameters of the image recognition model to obtain a trained image recognition model.

[0087] The first loss function and the second loss function can be any loss function, and the first loss function and the second loss function can be the same or different. In some embodiments, the first loss function can be a cross-entropy loss function, and the second loss function can be a least squares loss function. In other embodiments, both the first loss function and the second loss function can be cross-entropy loss functions.

[0088] According to an embodiment of the present disclosure, by comprehensively calculating the loss by combining the first sample image category output by the image recognition module and the second sample image category output by the multi-modal model to train the image recognition model, the robustness and generalization of the image recognition model can be improved.

[0089] Figure 6 Schematically shows a schematic diagram of a method for training an image recognition model according to an embodiment of the present disclosure.

[0090] As Figure 6As shown, in Embodiment 600, the sample image 610 is respectively input into the image recognition model 620 and the multimodal model 660; then, the image recognition model 620 performs image recognition processing on the sample image 610 and outputs the first sample image category 630; the first loss function is used to determine the first loss value 650 between the first sample image category 630 and the label 640 of the sample image; thereafter, the multimodal model 660 performs image recognition processing on the sample image 610 and outputs the second sample image category 670; the second loss function is used to determine the second loss value 680 between the second sample image category 670 and the label 640 of the sample image; thereafter, the total loss value 690 is determined according to the first loss value 650 and the second loss value 680, and the parameters of the image recognition model 620 are adjusted according to the total loss value 690. The above operations are cycled until the image recognition model 620 meets the preset iteration condition, and then the trained image recognition model is obtained.

[0091] Figure 7 FIG. schematically shows a schematic diagram of a method for training an image recognition model according to another embodiment of the present disclosure.

[0092] As Figure 7 shown, in Embodiment 700, the sample image 701 is input into the image encoding module 702 to convert the sample image 701 into a token sequence; then the token sequence is input into the feature extraction module 703 for feature extraction, and the sample image features are output; thereafter, the sample image features are respectively input into the fusion module 704 and the modality conversion model 707. After the fusion model 704 performs fusion processing on the local features and global features in the sample image features, the fused features are input into the image recognition module 705 for category classification to obtain the first sample image category 706; at the same time, the modality conversion model 707 performs modality conversion on the sample image features, outputs the sample image language modality features 708, and splices the sample image language modality features 708 with the preset prompt word 709 to form the splicing information 710; the splicing information 710 is input into the LLM model 711 for category classification to obtain the second sample image 712; thereafter, the total loss value 712 is determined according to the label of the sample image 701, the first sample image category 706, and the second sample image category 712, and the parameters of the image recognition model 710 are adjusted according to the total loss value 713. The above operations are cycled until the image recognition model 710 meets the preset iteration condition, and then the trained image recognition model is obtained.

[0093] According to an embodiment of the present disclosure, by first passing a sample image through an image patch encoding module to transform the image into a token sequence, and then inputting the token sequence into multiple Vision Transformer Blocks to obtain a local feature sequence V_l. At the same time, after each Vision Transformer Block, the cls token of the intermediate result is extracted to form a global feature sequence V_g. Finally, the local feature sequence V_l and the global feature sequence V_g are fused by feature averaging to obtain a fused feature, which contains the perception of the global perception of the image and the local fine-grained information; then the fused feature is input into a classifier, that is, an image recognition module for feature classification. At the same time, the local feature sequence V_l and the global feature sequence V_g are input into a modality converter Projector to transform the image modality features into language modality features, and combined with the fixed query "What is the main body in the image?" to perform VQA task learning, the image recognition model can be trained, and the strong generalization ability of the large language model can be utilized to enhance the depth perception and understanding of the main body of the image recognition model, and enhance the robustness and generalization of the image recognition model.

[0094] Figure 8 Schematically shows a flowchart of a large model-based question and answer method according to an embodiment of the present disclosure.

[0095] As Figure 8 shown, the question and answer method 800 of this embodiment includes operations S810 to S830.

[0096] In operation S810, question information is obtained, where the question information includes at least one target image.

[0097] The question information can be obtained by capturing the input data submitted by the user to the large model. For example, the user inputs the question information through the interaction interface, and after the user submits the input question information, the question information is obtained.

[0098] The question information may include one or more target pictures. In one embodiment, a question information may simultaneously include multiple target pictures. For example, the question information includes Picture A, Picture B, and a piece of text, where the text is used to prompt the large model to describe the content of Picture A and Picture B.

[0099] In some embodiments, the question information may further include a video, and at this time, the video can be regarded as multiple target pictures to execute the question and answer method of the present disclosure.

[0100] In operation S820, the trained image recognition model is used to perform image recognition processing on at least one target image to obtain the target image category of each of the at least one target image.

[0101] The target image category represents the category to which the target object in the target image belongs. The target object can be the main object in the target image. For example, if the target image contains a cat, the target object can be the cat; if the target image contains a person, the target object can be the person; if the target object contains a plant, the target object can be the plant.

[0102] It should be noted that the meaning of the target image category is the same as that of the above-mentioned sample image category, and will not be elaborated here.

[0103] The trained image recognition model is trained by the following method: based on the first sample image category obtained by performing image recognition processing on the sample image using the image recognition model, the second sample image category obtained by performing image recognition processing on the preset prompt word and the sample image using the multi-modal model, and the label of the sample image, the image recognition model is trained.

[0104] The training method of the image recognition model is the same as that of the image recognition model described above, and will not be elaborated here.

[0105] In operation S830, using the large model, the problem information and the enhanced information of the target image are processed to obtain the target answer information for the problem information, where the enhanced information of the target image is determined according to the target image category.

[0106] The enhanced information of the target image can be a detailed description information for the target image category. For example, if the target image category is a cat of breed A, the enhanced information can be the detailed information about the cat of breed A, such as the body size information, color information, characteristic information, personality information, etc. of the cat of breed A, which are any information related to the cat of breed A.

[0107] The enhanced information of the target image can be obtained by calling an external interface, relying on professional model calculation, database retrieval, etc. For example, by calling an external interface, an external software can be used to process the target icon category required, and the processing result can be determined as the enhanced information.

[0108] According to the embodiments of the present disclosure, after determining the enhanced information and using it as supplementary information for the problem information, the intention of the user submitting the problem information to the large model can be more clearly represented, so that the large model can process the problem information and the enhanced information, and obtain the target answer information for the problem information. Since the enhanced information is determined according to the target image category in the problem information, the relevance between the enhanced information and the problem information can be ensured, and the accuracy and usability of the target answer information can be improved.

[0109] In addition, by using the large model to process the question information and enhanced information again, it is also possible to utilize the understanding and analysis capabilities of the large model to discover additional information not included in the enhanced information during the process of generating the target answer information.

[0110] According to an embodiment of the present disclosure, the above method further includes: calling a preset knowledge base to obtain detailed information for the target image category to obtain enhanced information.

[0111] The preset knowledge base can be constructed according to the field specificity corresponding to the image recognition model. The preset knowledge base can include more comprehensive and detailed knowledge related to the field corresponding to the image recognition model. For example, if the image recognition model is used to recognize images in the medical field, the preset knowledge base can include relatively comprehensive and detailed knowledge related to the medical field. Again, if the image recognition model is used to recognize animal images, the preset knowledge base can include relatively comprehensive and detailed knowledge related to animals. Again, if the image recognition model is used to recognize animal images and plant images, the preset knowledge base can include relatively comprehensive and detailed knowledge related to animals, as well as relatively comprehensive and detailed knowledge related to plants.

[0112] By constructing a knowledge base based on the image recognition model, more comprehensive, authoritative, and detailed world knowledge can be retrieved, thereby enhancing the professional knowledge ability of the large language model and effectively improving the authority and professionalism of the content generated by the large model.

[0113] According to an embodiment of the present disclosure, the above method further includes: generating a target prompt word according to the question information, enhanced information, and prompt word rules; wherein, using the large model to process the question information and the enhanced information of the target image includes: using the large model to process the target prompt word.

[0114] The prompt word rules can be pre-set prompt word generation rules. For example, the prompt word rules can be "Please refer to the encyclopedic information [A] and answer the question [a]", where [A] represents the enhanced information and [a] represents the text information in the question information.

[0115] By using the preset prompt word rules to generate the target prompt word, it can help the large model better understand the task requirements, avoid the large model's misunderstanding of vague instructions, reduce irrelevant or unrelated outputs, and improve the accuracy and usability of the target answer information.

[0116] Figure 9 A schematic diagram of a large model-based question-answering method according to an embodiment of the present disclosure is schematically shown.

[0117] As Figure 9As shown, in Embodiment 900, problem information 910 is input into image recognition model 920. Image recognition model 920 performs image recognition on target picture 911 in problem information 910 and outputs target image type 930. Then, knowledge base 940 is called to retrieve encyclopedic information for target image category 930, thereby obtaining enhanced information 950. Then, the text information "What is this? Please introduce it." in problem information 910 and enhanced information 950 are used to generate target prompt word 960 according to the prompt word rules. And target prompt word 960 is input into large model 970 for processing, and target answer information 980 is output.

[0118] According to an embodiment of the present disclosure, the problem information includes multiple target images, and the multiple target images each have their own enhanced information; generating a target prompt word according to the problem information, the enhanced information, and the prompt word rules includes: sequentially splicing the enhanced information of each of the multiple target images to obtain splicing information; generating a target prompt word according to the problem information, the splicing information, and the prompt word rules.

[0119] For example, the multiple target images include Image 1 to Image 4. Sequentially splicing the enhanced information of each of the multiple target images may include splicing the enhanced information of Image 2 after the enhanced information of Image 1, splicing the enhanced information of Image 3 after the enhanced information of Image 2, and splicing the enhanced information of Image 4 after the enhanced information of Image 3, thereby obtaining splicing information.

[0120] According to an embodiment of the present disclosure, the above method further includes: in response to the problem information including a target video, performing frame splitting on the target video to obtain at least one video frame; wherein, using a trained image recognition model to perform image recognition processing on at least one target image includes: using a trained image recognition model to perform image recognition processing on at least one video frame.

[0121] For the case where the problem information contains a video, the video can be frame-split to obtain multiple video frames, and then processed according to the above situation of multiple target images.

[0122] Figure 10 Schematically shows a schematic diagram of a method for generating a target prompt word according to an embodiment of the present disclosure.

[0123] As Figure 10As shown, in Embodiment 1000, the problem information 1010 includes a video 1011 and text 1012. The video 1010 is frame-processed to obtain N video frames 1020. Then, the N video frames 1020 are input into an image recognition model 1030 for image recognition, and the video frame categories 1030 of each of the N video frames 1020 are output, obtaining N video frame categories 1030. Then, the detailed information of each of the N video frame categories 1030 is retrieved from the knowledge base, obtaining the enhanced information 1050 of each of the N video frame categories 1030, obtaining N enhanced information 1050. After that, according to the text information 1012 and the N enhanced information 1050, a target prompt word 1060 is generated according to the prompt word rule. For example, the target prompt word 1060 can be "Please refer to the encyclopedia information [N enhanced information 1050] and answer the question [text 1012]".

[0124] By splicing the enhanced information of multiple images, more complete information can be obtained, which is convenient for helping the large model better understand the task requirements and improve the accuracy and availability of the target answer information.

[0125] Figure 11 A block diagram of a training device for an image recognition model according to an embodiment of the present disclosure is schematically shown.

[0126] As Figure 11 shown, the training device 1100 of the image recognition model in this embodiment includes a first processing module 1110, a second processing module 1120, and a training module 1130.

[0127] The first processing module 1110 is configured to perform image recognition processing on a sample image using an image recognition model to obtain a first sample image category for the sample image, where the sample image category represents the category to which the target object in the sample image belongs.

[0128] The second processing module 1120 is configured to perform image recognition processing on a preset prompt word and a sample image using a multimodal model to obtain a second sample image category for the sample image.

[0129] The training module 1130 is configured to train the image recognition model according to the first sample image category, the second sample image category, and the label of the sample image to obtain a trained image recognition model.

[0130] According to an embodiment of the present disclosure, the image recognition model includes a feature extraction module and an image recognition module.

[0131] The feature extraction module is configured to extract sample image features of the sample image, where the sample image features include global features and local features of the sample image.

[0132] An image recognition module for performing image recognition processing on the features of a sample image to obtain a first sample image category.

[0133] According to an embodiment of the present disclosure, the feature extraction module includes: a feature extraction sub-module and a first splicing sub-module.

[0134] The feature extraction sub-module is used to perform feature extraction on multiple sample pixel blocks obtained by dividing the sample image to obtain the local features of each sample pixel block.

[0135] The first splicing sub-module is used to splice the local features of each sample pixel block according to the relative position relationship between the sample pixel blocks to obtain a global feature.

[0136] According to an embodiment of the present disclosure, the training device further includes: a fusion module.

[0137] The fusion module is used to perform average fusion processing on the local features and the global features of each sample pixel block to obtain a sample image fusion feature.

[0138] According to an embodiment of the present disclosure, the first processing module is further used to utilize the image recognition module to perform image recognition processing on the sample image fusion feature.

[0139] According to an embodiment of the present disclosure, the multi-modal model includes a modality conversion model and a large language model.

[0140] The modality conversion model is used to perform modality conversion processing on the sample image features to obtain sample image language modality features.

[0141] The large language model is used to perform image recognition processing on a preset prompt word and the sample image language modality features.

[0142] According to an embodiment of the present disclosure, the training module includes: a first processing sub-module, a second processing sub-module, and an adjustment sub-module.

[0143] The first processing sub-module is used to process the first sample image category and the label of the sample image using a first loss function to obtain a first loss value.

[0144] The second processing sub-module is used to process the second sample image category and the label of the sample image using a second loss function to obtain a second loss value.

[0145] The adjustment sub-module is used to obtain a trained image recognition model by adjusting the model parameters of the image recognition model based on the first loss value and the second loss value.

[0146] It should be noted that the training device part of the image recognition model in the embodiments of the present disclosure corresponds to the training method part of the image recognition model in the embodiments of the present disclosure. For the description of the training device part of the image recognition model, please specifically refer to the training method part of the image recognition model, and details will not be repeated here.

[0147] Figure 12 A block diagram of a question - answering device based on a large model according to an embodiment of the present disclosure is schematically shown.

[0148] As Figure 12 shown, the question - answering device 1200 based on a large model in this embodiment includes a first acquisition module 1210, a third processing module 1220, and a fourth processing module 1230.

[0149] The first acquisition module 1210 is configured to acquire question information, where the question information includes at least one target image.

[0150] The third processing module 1220 is configured to perform image recognition processing on at least one target image by using a trained image recognition model to obtain the target image category of each of the at least one target image, where the target image category represents the category to which the target object in the target image belongs. The trained image recognition model is trained based on the first sample image category obtained by performing image recognition processing on a sample image by using the image recognition model, the second sample image category obtained by performing image recognition processing on a preset prompt word and the sample image by using a multimodal model, and the label of the sample image.

[0151] The fourth processing module 1230 is configured to use the large model to process the question information and the enhanced information of the target image to obtain target answer information for the question information, where the enhanced information of the target image is determined according to the target image category.

[0152] According to an embodiment of the present disclosure, the above - mentioned question - answering device further includes: a second acquisition module.

[0153] The second acquisition module is configured to call a preset knowledge base to acquire detailed information for the target image category to obtain enhanced information.

[0154] According to an embodiment of the present disclosure, the above - mentioned question - answering device further includes: a generation module

[0155] The generation module is configured to generate a target prompt word according to the question information, the enhanced information, and a prompt word rule.

[0156] According to an embodiment of the present disclosure, the first processing module is further configured to process the target prompt word by using the large model.

[0157] According to an embodiment of the present disclosure, the question information includes a plurality of target images, and the plurality of target images have respective enhancement information.

[0158] According to an embodiment of the present disclosure, the generation module includes: a second splicing sub-module and a generation sub-module.

[0159] The second splicing sub-module is configured to sequentially splice the respective enhancement information of the plurality of target images to obtain splicing information.

[0160] The generation sub-module is configured to generate a target prompt word according to the question information, the splicing information, and the prompt word rule.

[0161] According to an embodiment of the present disclosure, the above question-and-answer device further includes: a frame division module.

[0162] The frame division module is configured to perform frame division processing on the target video in response to the question information including the target video to obtain at least one video frame.

[0163] According to an embodiment of the present disclosure, the third processing module is further configured to perform image recognition processing on at least one video frame by using a trained image recognition model.

[0164] It should be noted that in the embodiments of the present disclosure, the part of the question-and-answer device based on the large model corresponds to the part of the question-and-answer method based on the large model in the embodiments of the present disclosure. For the description of the part of the question-and-answer device based on the large model, please refer to the part of the question-and-answer method based on the large model, and details are not described herein again.

[0165] An embodiment of the present disclosure further provides an AI (Artificial Intelligence) intelligent agent, which is configured to execute the above-mentioned model training method and the question-and-answer method based on the large model.

[0166] Figure 13 A structural block diagram of the intelligent agent of artificial intelligence according to an embodiment of the present disclosure is schematically shown.

[0167] In the embodiments of the present disclosure, inspired by the von Neumann architecture in modern computer theory, as Figure 13 shown, the AI intelligent agent 1300 may include five core modules: an input module 1310, a control module 1320, a storage module 1330, an operation module 1340, and an output module 1350.

[0168] The input module 1310 is responsible for receiving or perceiving information such as queries, requests, instructions, signals, or data from the external world (e.g., users or the external environment), and converting it into a format that the AI agent 1300 can understand and process. The input module 1310 is the primary link for the AI agent 1300 to interact with the external world. It enables the AI agent 1300 to efficiently and accurately obtain necessary "sensory" information from the external world and respond to this information.

[0169] In the example, the input module 1310 can input the problem information described above.

[0170] In the example, the control module 1320 is the core support for the AI agent 1300's ability to handle complex tasks. The control module 1320 can execute the training method of the model and the question-and-answer method based on the large model described above.

[0171] In the example, during operation, the control module 1320 will continuously interact with the storage module 1330, the operation module 1340, and / or the output module 1350. However, it should be noted that in the embodiments of the present disclosure, the control module 1320 acts as a single initiator to initiate communication with the storage module 1330, the operation module 1340, and / or the output module 1350, and there is no communication coupling between the storage module 1330, the operation module 1340, and the output module 1350.

[0172] In the example, the performance of the control module 1320 can be closely related to the large model on which the AI agent 1300 is based. To fully utilize the capabilities of the large language model, the internal structure of the control module 1320 can be designed to be highly configurable and extensible to handle various different types of tasks and requirements in real-world scenarios.

[0173] The storage module 1330 can be responsible for memorizing information such as historical conversations and event streams. The operation area and the preset prompt word configuration file described above can be included in the storage module 1330.

[0174] In the example, after the AI agent 1300 obtains the problem information, the AI agent 1300 can determine the information enhancement method according to the information type in the problem information and store the information enhancement method in the storage module 1330. The AI agent 1300 can determine the information enhancement method from the storage module 1330 and feedback it to the control module 1320. Then, the control module 1320 can use the feedback target image category to determine the enhancement information for the problem information, control the large model to process the problem information and the enhancement information, obtain the target answer information for the problem information, and transmit the target answer information to the output module 1350.

[0175] The operation module 1340 can be regarded as a predefined tool library. An arithmetic unit for determining the jump parameters and / or search parameters as described above can be included in the operation module 1340.

[0176] In the example, when the AI agent 1300 needs to determine the jump parameters and / or search parameters for the information enhancement method, the relevant arithmetic unit can be called from the operation module 1340 and fed back to the control module 1320. Then, the control module 1320 can utilize the fed-back arithmetic unit to generate the jump parameters and / or search parameters for the information enhancement method, obtain the enhanced information, and based on the enhanced information and the question information, control the large model to perform corresponding operations to obtain the target answer information, and transmit the target answer information to the output module 1350. It can be understood that although the large language model has excellent language understanding and generation capabilities, like humans, the tasks it can solve without any tools are very limited. When the AI agent 1300 is given the ability to call tools, tasks such as performing mathematical operations with the help of a calculator, performing data analysis with the help of python, and performing human recognition in images with the help of an expert model can be achieved.

[0177] In the example, the output module 1350 can output the target answer information described above.

[0178] The AI agent 1300 according to the embodiments of the present disclosure can simply and effectively improve the degree of intelligence, and improve flexibility and versatility.

[0179] According to the embodiments of the present disclosure, the present disclosure also provides an electronic device, a readable storage medium, and a computer program product.

[0180] According to the embodiments of the present disclosure, an electronic device includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein, the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the method as above.

[0181] According to the embodiments of the present disclosure, a non-transitory computer-readable storage medium storing computer instructions, wherein the computer instructions are used to cause a computer to execute the method as above.

[0182] According to the embodiments of the present disclosure, a computer program product includes a computer program, and the computer program implements the method as above when executed by a processor.

[0183] Figure 14FIG. schematically shows a schematic block diagram of an exemplary electronic device 1400 that may be used to implement embodiments of the present disclosure. The electronic device is intended to represent various forms of digital computers, such as, for example, laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device may also represent various forms of mobile devices, such as, for example, personal digital processors, cellular telephones, smart phones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely exemplary and are not intended to limit the implementation of the present disclosure described and / or claimed herein.

[0184] As Figure 14 shown, the device 1400 includes a computing unit 1401 that can perform various appropriate actions and processes in accordance with a computer program stored in a read-only memory (ROM) 1402 or a computer program loaded from a storage unit 1408 into a random access memory (RAM) 1403. In the RAM 1403, various programs and data required for the operation of the device 1400 can also be stored. The computing unit 1401, the ROM 1402, and the RAM 1403 are connected to each other via a bus 1404. An input / output (I / O) interface 1405 is also connected to the bus 1404.

[0185] A plurality of components in the device 1400 are connected to the input / output (I / O) interface 1405, including: an input unit 1406, such as, for example, a keyboard, a mouse, etc.; an output unit 1407, such as, for example, various types of displays, speakers, etc.; a storage unit 1408, such as, for example, a magnetic disk, an optical disk, etc.; and a communication unit 1409, such as, for example, a network card, a modem, a wireless communication transceiver, etc. The communication unit 1409 allows the device 1400 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.

[0186] The computing unit 1401 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 1401 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 1401 executes the various methods and processes described above, such as the large model-based question answering method and the large model training method. For example, in some embodiments, the large model-based question answering method and the large model training method can be implemented as a computer software program tangibly embodied in a machine-readable medium, such as the storage unit 1408. In some embodiments, part or all of the computer program can be loaded and / or installed onto the device 1400 via the ROM 1402 and / or the communication unit 1409. When the computer program is loaded into the RAM 1403 and executed by the computing unit 1401, one or more steps of the large model-based question answering method and the large model training method described above can be executed. Alternatively, in other embodiments, the computing unit 1401 can be configured to execute the large model-based question answering method and the large model training method by any other suitable means (e.g., by means of firmware).

[0187] The various embodiments of the systems and techniques described above in this document can be implemented in digital electronic circuit systems, integrated circuit systems, field programmable gate arrays (FPGA), application specific integrated circuits (ASIC), application specific standard products (ASSP), system on a chip systems (SOC), complex programmable logic devices (CPLD), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include: being implemented in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which can be a special or general programmable processor, receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting the data and instructions to the storage system, the at least one input device, and the at least one output device.

[0188] The program code for implementing the methods of the present disclosure may be written in any combination of one or more programming languages. These program codes may be provided to a processor or controller of a general purpose computer, special purpose computer, or other programmable data processing device, such that the program codes, when executed by the processor or controller, cause the functions / operations specified in the flowchart and / or block diagram to be implemented. The program code may execute entirely on the machine, partly on the machine, as a stand-alone software package partly on the machine and partly on a remote machine, or entirely on the remote machine or server.

[0189] In the context of the present disclosure, a machine-readable medium may be a tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. A machine-readable medium may be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium may include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of a machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0190] In order to provide interaction with a user, the systems and techniques described herein may be implemented on a computer having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and a pointing device (e.g., a mouse or a trackball) by which the user can provide input to the computer. Other kinds of devices may also be used to provide interaction with the user; for example, the feedback provided to the user may be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user may be received in any form (including acoustic input, speech input, or tactile input).

[0191] The systems and techniques described herein can be implemented in a computing system including backend components (e.g., as a data server), or a computing system including middleware components (e.g., an application server), or a computing system including frontend components (e.g., a user computer having a graphical user interface or a web browser through which a user can interact with an implementation of the systems and techniques described herein), or a computing system including any combination of such backend components, middleware components, or frontend components. The components of the system can be interconnected with each other by digital data communication in any form or medium (e.g., a communication network). Examples of communication networks include: local area network (LAN), wide area network (WAN), and the Internet.

[0192] A computer system can include a client and a server. The client and the server are generally far from each other and typically interact through a communication network. The client-server relationship is created by computer programs running on the respective computers and having a client-server relationship with each other. The server can be a cloud server, a server of a distributed system, or a server incorporating a blockchain.

[0193] It should be understood that various forms of the processes shown above can be used, steps can be reordered, added, or deleted. For example, the steps recited in this disclosure can be executed in parallel, sequentially, or in a different order, as long as the desired results of the technical solutions disclosed in this disclosure can be achieved, and this is not limited herein.

[0194] The above specific embodiments do not constitute a limitation on the protection scope of this disclosure. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this disclosure shall be included within the protection scope of this disclosure.

Claims

1. A method for training an image recognition model, comprising: Performing image recognition processing on the sample image using an image recognition model to obtain a first sample image category for the sample image, wherein the sample image category represents the category to which the target object in the sample image belongs; Using a multimodal model, performing image recognition processing on a preset prompt word and the sample image to obtain a second sample image category for the sample image; and The image recognition model is trained according to the first sample image category, the second sample image category and the label of the sample image to obtain a trained image recognition model.

2. The method according to claim 1, wherein: The image recognition model includes a feature extraction module and an image recognition module; The using the image recognition model to perform image recognition processing on the sample image to obtain a first sample image category for the sample image includes: Utilizing the feature extraction module to extract sample image features of the sample image, wherein the sample image features include global features of the sample image and local features of the sample image; The image recognition module is used to perform image recognition processing on the sample image features to obtain the first sample image category.

3. The method according to claim 2, wherein: The step of extracting the sample image features of the sample image comprises: Extracting features from a plurality of sample pixel blocks obtained by dividing the sample image to obtain local features of each of the sample pixel blocks; The local features of the sample pixel blocks are spliced ​​according to the relative position relationship between the sample pixel blocks to obtain the global feature.

4. The method according to claim 3, further comprising: Performing average fusion processing on the local features of each of the sample pixel blocks and the global features to obtain sample image fusion features; The using the image recognition module to perform image recognition processing on the sample image features includes: The image recognition module is used to perform image recognition processing on the sample image fusion features.

5. The method according to claim 2, wherein: The multimodal model includes a modal conversion model and a large language model; The method of using the multimodal model to perform image recognition processing on the preset prompt words and the sample image includes: Using the modality conversion model, performing modality conversion processing on the sample image features to obtain sample image language modality features; The large language model is used to perform image recognition processing on the preset prompt words and the language modality features of the sample images.

6. The method according to claim 1, wherein: The step of training the image recognition model according to the first sample image category, the second sample image category and the label of the sample image to obtain a trained image recognition model comprises: Processing the first sample image category and the label of the sample image using a first loss function to obtain a first loss value; Processing the second sample image category and the label of the sample image using a second loss function to obtain a second loss value; and Based on the first loss value and the second loss value, a trained image recognition model is obtained by adjusting the model parameters of the image recognition model.

7. A question answering method based on a large model, comprising: Acquiring question information, wherein the question information includes at least one target image; Using the trained image recognition model, performing image recognition processing on at least one of the target images to obtain a target image category of at least one of the target images, wherein the target image category represents the category to which the target object in the target image belongs, and the trained image recognition model is trained based on a first sample image category obtained by performing image recognition processing on a sample image using the image recognition model, a second sample image category obtained by performing image recognition processing on a preset prompt word and the sample image using a multimodal model, and a label of the sample image; The question information and the enhanced information of the target image are processed by using a large model to obtain target answer information for the question information, wherein the enhanced information of the target image is determined according to the category of the target image.

8. The method according to claim 7, further comprising: A preset knowledge base is called to obtain detailed information for the target image category to obtain the enhanced information.

9. The method according to claim 7, further comprising: Generate a target prompt word according to the question information, the enhanced information and the prompt word rule; Wherein, the step of using the large model to process the problem information and the enhanced information of the target image includes: The target cue word is processed using the large model.

10. The method according to claim 9, wherein: The problem information includes a plurality of target images, and the plurality of target images have respective enhancement information; The step of generating a target prompt word according to the question information, the enhanced information and the prompt word rule includes: Sequentially stitching the enhanced information of the plurality of target images to obtain stitching information; A target prompt word is generated according to the question information, the splicing information and the prompt word rule.

11. The method according to claim 7, further comprising: In response to the problem information including a target video, performing frame processing on the target video to obtain at least one video frame; The step of performing image recognition processing on at least one of the target images using the trained image recognition model includes: Using the trained image recognition model, image recognition processing is performed on at least one of the video frames.

12. An artificial intelligence agent configured to execute the method according to any one of claims 1 to 11.

13. A training device for an image recognition model, comprising: A first processing module is used to perform image recognition processing on the sample image using an image recognition model to obtain a first sample image category for the sample image, wherein the sample image category represents the category to which the target object in the sample image belongs; a second processing module, configured to perform image recognition processing on a preset prompt word and the sample image using a multimodal model to obtain a second sample image category for the sample image; and A training module is used to train the image recognition model according to the first sample image category, the second sample image category and the label of the sample image to obtain a trained image recognition model.

14. A question-answering device based on a large model, comprising: A first acquisition module, used to acquire question information, wherein the question information includes at least one target image; A third processing module is used to perform image recognition processing on at least one of the target images using a trained image recognition model to obtain a target image category of at least one of the target images, wherein the target image category represents the category to which the target object in the target image belongs, and the trained image recognition model is trained based on a first sample image category obtained by performing image recognition processing on a sample image using an image recognition model, a second sample image category obtained by performing image recognition processing on a preset prompt word and the sample image using a multimodal model, and a label of the sample image; The fourth processing module is used to use the large model to process the question information and the enhanced information of the target image to obtain the target answer information for the question information, wherein the enhanced information of the target image is determined according to the category of the target image.

15. An electronic device, comprising: at least one processor; as well as a memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method according to any one of claims 1 to 11.

16. A non-transitory computer-readable storage medium storing computer instructions, wherein: The computer instructions are used to cause the computer to execute the method according to any one of claims 1-11.

17. A computer program product, comprising a computer program, wherein the computer program is stored on at least one of a readable storage medium and an electronic device, and when the computer program is executed by a processor, the method according to any one of claims 1 to 11 is implemented.

Citation Information

Cited By

  • Agricultural environment intelligent control system based on artificial intelligence and control method thereof

    CN121300097A