Intelligent question answering method based on multi-modal data, electronic equipment and storage medium
Through the intelligent question-and-answer method of multimodal data, combined with coarse-grained and fine-grained feature extraction, the information overload and details loss of visual language models when processing large-scale image data is solved, efficient and flexible image analysis and interaction are achieved, and the accuracy of question-and-answer results is improved.
Patent Information
- Application Number
- CN202510116952.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-23
- Publication Date
- 2025-05-30
AI Technical Summary
When existing visual language models deal with large-scale or high-dimensional image data, there are problems such as image information overload, image details loss and insufficient interactivity, which limits their application potential in complex scenarios.
Using an intelligent question-and-answer method based on multimodal data, by acquiring image data and text instructions, using coarse-grained and fine-grained feature extraction modules, combined with visual language processing modules, dynamically adjusting feature extraction granularity, accurately locate and analyze key areas in the image, and generating question-and-answer results.
It significantly reduces the computational complexity and memory consumption of the model, improves the operation efficiency and flexibility, enhances the processing ability of local details, improves the accuracy of Q&A results, and meets the dynamic and personalized needs of users.
Smart Images

Figure CN120069067A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of artificial intelligence technology, and particularly relates to an intelligent question-answering method, an electronic device, and a storage medium based on multi-modal data. Background Art
[0002] Currently, in the field of artificial intelligence, especially in the application of Visual Large Language Models (vLLMs), processing and analyzing image data is a fundamental and crucial task. With the development of technology, vLLMs have made remarkable progress in image recognition, classification, and interaction with users. Through deep learning technology, the model can understand and process complex information in images, providing users with rich visual content and interactive experiences.
[0003] However, existing vLLMs still face some technical challenges when dealing with large-scale or high-dimensional image data, such as 3D medical images, including problems like image information overload, loss of image details, and insufficient interactivity. These problems limit the application potential of vLLMs in more complex scenarios.
[0004] To address the information overload problem, existing methods may adopt image downsampling (pooling) techniques, but this will result in the loss of image detail information, affecting the model's in-depth understanding and analysis of images. In fields such as medical image analysis, the loss of detail information may cause key diagnostic information to be overlooked, thus affecting the final medical decision-making.
[0005] In addition, existing vLLMs also have deficiencies in terms of interactivity. When processing images, vLLMs often only perform one image analysis and lack the ability to dynamically adjust according to user instructions, which limits the depth and flexibility of interaction between the model and users. In practical applications, users may need to conduct in-depth analysis of different parts of the image according to different needs, and existing vLLMs cannot meet such dynamic and personalized requirements.
[0006] Correspondingly, a new technical solution is needed in this field to solve the above problems. Summary of the Invention
[0007] In order to overcome the above defects, the present application is proposed to provide a solution to solve or at least partially solve the technical problems of image information overload, loss of image details, and insufficient interactivity existing in existing visual large language models when dealing with large-scale or high-dimensional image data.
[0008] In a first aspect, an intelligent question-answering method based on multi-modal data is provided, characterized in that the method includes:
[0009] Obtain multimodal data; the multimodal data includes image data and text instructions, and the image data includes three-dimensional images and two-dimensional images;
[0010] Input the multimodal data into an intelligent question-answering model; the intelligent question-answering model includes a coarse-grained feature extraction module, a fine-grained feature extraction module, and a vision-language processing module;
[0011] Obtain the coarse-grained features of the image data based on the coarse-grained feature extraction module;
[0012] Obtain the fine-grained features of the key regions in the image data based on the coarse-grained features of the image data and the text instructions;
[0013] Obtain the question-answering result corresponding to the text instructions based on the coarse-grained features of the image data, the text instructions, and the fine-grained features of the key regions; the question-answering result includes answer text.
[0014] In a technical solution of the above intelligent question-answering method based on multimodal data, the obtaining of the fine-grained features of the key regions in the image data based on the coarse-grained features of the image data and the text instructions includes:
[0015] Obtain the key-region images of the image data based on the coarse-grained features of the image data, the text instructions, and the vision-language processing module;
[0016] Obtain the fine-grained features of the key-region images based on the fine-grained feature extraction module.
[0017] In a technical solution of the above intelligent question-answering method based on multimodal data, the vision-language processing module includes a decoding sub-module; the obtaining of the key-region images of the image data based on the coarse-grained features of the image data, the text instructions, and the vision-language processing module includes:
[0018] Obtain a first multimodal sequence based on the coarse-grained features of the image data and the text instructions;
[0019] Input the first multimodal sequence into the vision-language processing module, and process the first multimodal sequence based on the decoding sub-module to obtain the position information of the key regions of the image data;
[0020] Crop the image data based on the position information of the key regions to obtain the key-region images.
[0021] In a technical solution of the above intelligent question-answering method based on multimodal data, the obtaining of the first multimodal sequence based on the coarse-grained features of the image data and the text instructions includes:
[0022] Encode the text instruction to obtain a sequence of word vectors;
[0023] Concatenate the coarse-grained feature and the sequence of word vectors to obtain a first multi-modal sequence.
[0024] In a technical solution of the above intelligent question-answering method based on multi-modal data, obtaining the question-answering result corresponding to the text instruction based on the coarse-grained feature of the image data, the text instruction, and the fine-grained feature of the key region includes:
[0025] Concatenate the coarse-grained feature, the sequence of word vectors, and the fine-grained feature to obtain a second multi-modal sequence;
[0026] Input the second multi-modal sequence into the vision-language processing module, and process the second multi-modal sequence based on the decoding sub-module to obtain the question-answering result corresponding to the text instruction.
[0027] In a technical solution of the above intelligent question-answering method based on multi-modal data, the decoding sub-module includes multiple layers of decoders, and the question-answering result further includes the position information of the key region; processing the second multi-modal sequence based on the decoding sub-module to obtain the question-answering result corresponding to the text instruction includes:
[0028] Based on the multi-head self-attention mechanism of each layer of decoder, obtain the interaction information between the text instruction in the second multi-modal sequence, the coarse-grained feature of the image data, and the fine-grained feature of the key region image;
[0029] Based on the interaction information, obtain the answer text corresponding to the text instruction, or obtain the answer text corresponding to the text instruction and the position information of the key region.
[0030] In a technical solution of the above intelligent question-answering method based on multi-modal data, encoding the text instruction to obtain a sequence of word vectors includes:
[0031] Tokenize the text instruction and convert it into a sequence of token identifiers;
[0032] Convert the sequence of token identifiers into the sequence of word vectors.
[0033] In a technical solution of the above intelligent question-answering method based on multi-modal data, the image data is a standard three-dimensional image or a two-dimensional image of a preset size; obtaining multi-modal data includes:
[0034] Obtain initial image data;
[0035] Preprocess the initial image data to obtain a standard three-dimensional image or two-dimensional image of the preset size;
[0036] Obtain a text instruction input by the user; the text instruction includes a question text or a requirement text;
[0037] Obtain the multi-modal data based on the standard three-dimensional image or two-dimensional image of the preset size and the text instruction input by the user.
[0038] In a second aspect, there is provided an electronic device, which includes a processor and a memory. The memory is adapted to store multiple program codes, and the program codes are adapted to be loaded and run by the processor to execute the intelligent question-answering method based on multi-modal data according to any one of the technical solutions in the technical solution of the intelligent question-answering method based on multi-modal data described above.
[0039] In a third aspect, there is provided a computer-readable storage medium, in which multiple program codes are stored, and the program codes are adapted to be loaded and run by a processor to execute the intelligent question-answering method based on multi-modal data according to any one of the technical solutions in the technical solution of the intelligent question-answering method based on multi-modal data described above.
[0040] One or more of the above technical solutions of the present application have at least one or more of the following Beneficial effects:
[0041] In implementing the technical solution of this application, multi-modal data is first obtained. The multi-modal data includes image data and text instructions. The image data includes three-dimensional images and two-dimensional images. Then, the multi-modal data is input into an intelligent question-answering model. The intelligent question-answering model includes a coarse-grained feature extraction module, a fine-grained feature extraction module, and a vision-language processing module. Based on the coarse-grained feature extraction module, the coarse-grained features of the image data are obtained. Based on the coarse-grained features of the image data and the text instructions, the fine-grained features of the key regions in the image data are obtained. Finally, based on the coarse-grained features of the image data, the text instructions, and the fine-grained features of the key regions, the question-answering result corresponding to the text instructions is obtained. The question-answering result includes the answer text. Through the above implementation method, after the multi-modal data is input into the intelligent question-answering model, the model can initially extract the coarse-grained features of the image data. When processing large-scale or high-dimensional image data, it can significantly reduce the computational complexity and memory consumption of the model and improve the running efficiency. By obtaining the fine-grained features of the key regions in the image through the coarse-grained features and the user instructions, it can dynamically adjust the granularity of feature extraction according to the user's needs, adapt to different task requirements, improve flexibility and the user's interaction experience, and at the same time can accurately locate and analyze the key regions in the image and enhance the processing ability of local details. By obtaining the question-answering result corresponding to the text instructions through the coarse-grained features, the text instructions, and the fine-grained features, it can combine the coarse-grained feature extraction and the fine-grained feature extraction, retain the key details in the image, and significantly improve the accuracy of the question-answering result. BRIEF DESCRIPTION OF THE DRAWINGS
[0042] With reference to the accompanying drawings, the disclosure of this application will become more understandable. It is easily understood by those skilled in the art that these drawings are only for illustrative purposes and are not intended to limit the protection scope of this application. Among them:
[0043] Figure 1 is a schematic diagram of the main step flow of an intelligent question-answering method based on multi-modal data according to an embodiment of this application;
[0044] Figure 2 is a schematic diagram of the main flow of inputting multi-modal data into an intelligent question-answering model to obtain the question-answering result corresponding to the text instructions according to an embodiment of this application;
[0045] Figure 3 is a schematic diagram of the main step flow of obtaining the fine-grained features of the key regions in the image data based on the coarse-grained features of the image data and the text instructions according to an embodiment of this application;
[0046] Figure 4 is a schematic diagram of the main step flow of obtaining the question-answering result corresponding to the text instructions based on the coarse-grained features of the image data, the text instructions, and the fine-grained features of the key regions according to an embodiment of this application;
[0047] Figure 5 It is a schematic diagram of the main structure of an electronic device according to an embodiment of the present application.
[0048] List of reference numerals:
[0049] 21: Coarse-grained feature extraction module; 22: Fine-grained feature extraction module; 23: Vision-language processing module; 51: Processor; 52: Memory. Detailed implementation manners
[0050] Some implementation manners of the present application will be described below with reference to the accompanying drawings. Those skilled in the art should understand that these implementation manners are only used to explain the technical principle of the present application and are not intended to limit the protection scope of the present application.
[0051] In the description of the present application, "module" and "processor" may include hardware, software, or a combination of both. A module may include a hardware circuit, various suitable sensors, communication ports, memory, and may also include a software part, such as program code, or a combination of software and hardware. The processor may be a central processing unit, a microprocessor, an image processor, a digital signal processor, or any other suitable processor. The processor has data and / or signal processing functions. The processor may be implemented in software, in hardware, or in a combination of both. The non-transitory computer-readable storage medium includes any suitable medium for storing program code, such as magnetic disks, hard disks, optical discs, flash memories, read-only memories, random access memories, and so on.
[0052] The term "A and / or B" represents all possible combinations of A and B, such as only A, only B, or A and B. The term "at least one A or B" or "at least one of A and B" has a similar meaning to "A and / or B" and may include only A, only B, or A and B. The singular terms "a" and "this" may also include the plural form.
[0053] Some terms related to the present application will be explained here first.
[0054] vLLM: vLLM is a vision-language large model based on the Transformer architecture, capable of processing joint inputs of images and texts. The model extracts image features and combines them with user instructions to generate corresponding outputs. The vLLM model usually uses pre-trained vision encoders (such as ResNet, ViT) to extract image features and then generates text outputs through a Transformer decoder.
[0055] Image Pooling: Image pooling is an image processing technique used to reduce the size and computational complexity of an image. Common pooling operations include Max Pooling and Average Pooling. Pooling operations can effectively reduce the length of the input sequence but may result in the loss of image detail information.
[0056] token: A basic unit of feature representation, which is a representative information segment obtained after processing the data.
[0057] <focus>Token: Used to indicate the location of the area in the image data that requires key attention. Each " <focus>Token" will contain 6 characters after it, in the form of "xxyyzz" to indicate the percentage position of the position in the three dimensions of the image data.
[0058] Existing vLLMs have the following problems when processing large-scale or high-dimensional image data, such as 3D medical images, which limits the application potential of vLLMs in more complex scenarios:
[0059] 1. The input information is too long, resulting in high computational complexity: When processing large-scale images, such as 3D medical images, existing vLLMs usually need to use all image information as input. Due to the large amount of 3D image data, the input sequence length increases significantly, resulting in a sharp increase in the computational complexity and memory consumption of the model. Due to the high demand for computing resources, it is difficult to run efficiently on ordinary hardware, which limits the actual application scenarios of the model.
[0060] 2. Pooling operation leads to loss of image details: In order to reduce the length of the input sequence, existing methods usually perform pooling operations on the image, such as maximum pooling or average pooling. Although the pooling operation reduces the input length, it also loses a lot of detail information, especially for scenes that require high-precision analysis such as medical images. Due to the limited performance of the model, it is difficult to capture the key details in the image, affecting the accuracy of the task.
[0061] 3. Static feature extraction cannot be adjusted dynamically: Existing methods usually extract image features once and cannot dynamically adjust the granularity of feature extraction according to user instructions or task requirements. This static feature extraction method lacks flexibility when dealing with complex tasks. Since the model cannot dynamically optimize the feature extraction process according to task requirements, it leads to resource waste or insufficient information.
[0062] 4. Lack of fine-grained attention to local areas: Existing methods usually perform global feature extraction on the entire image, lacking fine-grained attention to specific local areas. For tasks that require high-precision analysis, such as lesion detection in medical images, this global feature extraction method may not meet the needs. Since the model performs poorly when processing tasks that require local details, it is difficult to accurately locate and analyze key areas.
[0063] In order to solve the above problems, the present application provides an intelligent question-answering method, electronic device and storage medium based on multimodal data.
[0064] See attached Figure 1 , Figure 1 FIG. 1 is a flow chart of the main steps of an intelligent question-answering method based on multimodal data according to an embodiment of the present application. Figure 1 As shown, the intelligent question-answering method based on multimodal data in the embodiment of the present application mainly includes the following steps S101 to S105.
[0065] Step S101: Obtain multimodal data;
[0066] Among them, the multimodal data includes image data and text instructions. The image data includes three-dimensional images and two-dimensional images, specifically, it can be a standard three-dimensional image or two-dimensional image of a preset size. The text instructions are text instructions input by the user, including task text or question text.
[0067] Step S102: Input the multimodal data into the intelligent question-answering model;
[0068] Among them, the intelligent question-answering model includes a coarse-grained feature extraction module, a fine-grained feature extraction module, and a vision-language processing module.
[0069] Step S103: Obtain the coarse-grained features of the image data based on the coarse-grained feature extraction module;
[0070] Step S104: Obtain the fine-grained features of the key regions in the image data based on the coarse-grained features of the image data and the text instructions;
[0071] Step S105: Obtain the question-answering result corresponding to the text instructions based on the coarse-grained features of the image data, the text instructions, and the fine-grained features of the key regions.
[0072] Among them, the question-answering result can include answer text, or it can include the answer text corresponding to the text instructions and the location information of the key regions.
[0073] Based on the method described in the above steps S101 to S105, after inputting the multimodal data into the intelligent question-answering model, the model can initially extract the coarse-grained features of the image data. When processing large-scale or high-dimensional image data, it can significantly reduce the computational complexity and memory consumption of the model, and improve the operation efficiency; by obtaining the fine-grained features of the key regions in the image through the coarse-grained features and user instructions, it can dynamically adjust the granularity of feature extraction according to user needs, adapt to different task requirements, improve flexibility and the user's interaction experience, and at the same time can accurately locate and analyze the key regions in the image, enhancing the processing ability of local details; by obtaining the question-answering result corresponding to the text instructions through the coarse-grained features, text instructions, and fine-grained features, it can combine the coarse-grained feature extraction and fine-grained feature extraction, retain the key details in the image, and significantly improve the accuracy of the question-answering result.
[0074] The following further explains the above steps S101 to S105.
[0075] In some embodiments of the above step S101, the obtained multimodal data includes image data and text instructions.
[0076] Among them, the image data is a standard three-dimensional image or two-dimensional image of a preset size. Specifically, the initial image data can be obtained first, and then the initial image data can be preprocessed to obtain a standard three-dimensional image or two-dimensional image of a preset size.
[0077] In some embodiments, the initial image data can be a high-resolution three-dimensional image (such as 512*512*512 voxels) or a large-size two-dimensional image (such as 512*512 pixels).
[0078] For example, in the medical field, the initial image data can be multi-channel volume data such as Computed Tomography (CT) images, Magnetic Resonance Imaging (MRI) images, etc., and its format can be common medical image formats such as DICOM or NIfTI, etc.
[0079] In addition, the initial image data can also be three-dimensional images or two-dimensional images in the fields of computer, industrial manufacturing, and intelligent driving, which are not limited here.
[0080] Furthermore, the preprocessing of the initial image data can include normalization processing and size adjustment.
[0081] Performing normalization processing on the initial image data can make its pixel values between 0 and 1, so that different initial image data have a unified scale.
[0082] In some embodiments, each initial pixel value in the initial image data can be subtracted by the minimum pixel value, and then divided by the difference between the maximum pixel value and the minimum pixel value.
[0083] Specifically, the initial image data can be normalized through the following formula (1):
[0084]
[0085] Among them, x is any initial pixel value in the initial image data, and x norm is the pixel value after normalization processing, min is the minimum value of all pixel values in the initial image, and max is the maximum value of all pixel values in the initial image.
[0086] In addition, if the initial pixel values include negative values, such as the Hounsfield units (HU) in CT images, which reflect the degree of tissue absorption of X-rays, the HU values of different tissues are different. The HU value of air is about -1000, the HU value of water is about 0, and the HU value of bone can reach more than 1000, which makes negative values exist in CT images. At this time, truncation can be performed first to adjust the initial pixel values exceeding the specified range in the initial image to the preset range, such as within [-1000, 1000], and then the above-mentioned normalization process is performed to obtain a standard three-dimensional image or two-dimensional image.
[0087] Furthermore, the size of the standard three-dimensional image or two-dimensional image can be adjusted to a unified size, such as a three-dimensional image of 224*224*224 voxels or a two-dimensional image of 224*224 pixels, to adapt to subsequent feature extraction. Specifically, bilinear interpolation or trilinear interpolation can be used to scale the standard three-dimensional image or two-dimensional image, retaining the main structural information of the image, to obtain a standard three-dimensional image or two-dimensional image of the preset size.
[0088] The multi-modal data obtained in step S101 also includes text instructions input by the user, mainly including problem text or requirement text.
[0089] Among them, the text instructions are used to describe the user's problem or requirement. For example, "What is the diagnosis result of this patient?", "Is there a tumor in the brain?", "Please analyze the area around the liver.", etc.
[0090] Through the standard three-dimensional image or two-dimensional image of the preset size, and the text instructions input by the user, multi-modal data can be obtained.
[0091] The above is a further description of step S101. Next, step S102 will be further described.
[0092] In some embodiments of the above step S102, an intelligent question-answering model can be constructed first. The intelligent question-answering model includes a coarse-grained feature extraction module, a fine-grained feature extraction module, and a vision-language processing module.
[0093] Among them, the coarse-grained feature extraction module is used to extract rough feature representations from the input image data. Specifically, the coarse-grained feature extraction module can be a 3D convolutional neural network for processing three-dimensional images, such as 3D ResNet-50, 3D Vision Transformer (ViT). 3D ResNet-50 is suitable for processing three-dimensional images and can capture features in the spatial dimension; 3D ViT can extract global features through the self-attention mechanism, and the input image size can be relatively large, such as 32, that is, each 32*32*32 voxels corresponds to a token. The coarse-grained feature extraction module can also be a 2D convolutional neural network for processing two-dimensional images, such as 2D ResNet-50, 2D Vision Transformer (ViT), etc. In addition, the coarse-grained feature extraction module can also include both a 3D convolutional neural network for processing three-dimensional images and a 2D convolutional neural network for processing two-dimensional images to extract rough features of three-dimensional images or two-dimensional images through the coarse-grained feature extraction module, which is not limited here.
[0094] The fine-grained feature extraction module is used to extract fine-grained feature representations from the key region images. The fine-grained feature extraction module can be a high-resolution visual encoder, such as the 3D HRNet network for processing three-dimensional images. The 3D HRNet network is a neural network architecture specifically designed for processing three-dimensional data and is suitable for processing three-dimensional images. The fine-grained feature extraction module can also be a 2D HRNet network for processing two-dimensional images, or include both a 3D network for processing three-dimensional images and a 2D network for processing two-dimensional images to extract fine-grained features of three-dimensional images or two-dimensional images through the fine-grained feature extraction module, which is not limited here.
[0095] The vision-language processing module is used to process text instructions and image features to generate text or <focus>token. The visual language processing module can be vLLM, which is a large visual language model based on the Transformer architecture and can process the joint input of images and text. The model extracts image features and combines them with the user's instructions to generate corresponding outputs.
[0096] It should be noted that the above examples of the coarse-grained feature extraction module, the fine-grained feature extraction module, and the visual language processing module are only illustrative. In actual applications, those skilled in the art can make selections according to specific scenarios, which are not limited here.
[0097] Furthermore, after the intelligent question-answering module is constructed, the intelligent question-answering module can be trained through the following steps S1021 to step S1023.
[0098] Step S1021: Obtain a training data set;
[0099] In some embodiments, the training data set includes training text instructions and image data key region labels.
[0100] Among them, the training text instructions can include text instructions written by relevant field experts according to the content of 3D images or 2D images, text instructions related to 3D images or 2D images extracted from existing question-answering data sets (such as the medical visual question-answering data set VQA-Med, the radiology field data set Radiology Objects in Context, etc.), and text instructions generated according to the content of 3D images or 2D images using a natural language generation model (such as GPT). The training text instructions have diversity to ensure that they cover various task types (such as diagnosis, localization, description, etc.) and diverse language styles (such as declarative sentences, interrogative sentences, etc.). In addition, the training text instructions should concisely and clearly describe the user's questions or needs.
[0101] The image data key region labels can include those obtained by relevant field experts annotating the regions that need attention according to the content of 3D images or 2D images <focus>Token, and automatically detecting key regions in three-dimensional or two-dimensional images using an object detection model (such as Faster R-CNN) or a segmentation model (such as U-Net), generating <focus>Tokens and the like. And, <focus>Tokens are diverse to ensure coverage of multiple key regions in the image data and to ensure multiple <focus>The positions of the tokens are evenly distributed, covering all parts of the image data.
[0102] Step S1022: Train the coarse-grained feature extraction module, the fine-grained feature extraction module, and the vision-language processing module based on the training dataset;
[0103] Step S1023: Optimize the intelligent question-answering model until the intelligent question-answering model converges to a preset error, and then complete the training of the intelligent question-answering model.
[0104] In some embodiments, the parameters of the model can be updated using the Adaptive Moment Estimation (Adam) optimizer. The Adam optimizer is a widely used deep learning optimization algorithm that can efficiently update the parameters of the model, enabling the model to converge to the optimal solution faster.
[0105] Specifically, the parameters of the Adam optimizer can be set according to the specific scenario and actual requirements. For example, set the initial learning rate to 1e-4, the weight decay to 1e-5, the momentum parameter to β1 = 0.9, and the gradient clipping to the maximum value = 1.0.
[0106] Among them, the initial learning rate determines the step size of the model parameter update during each training. Weight decay is a regularization technique that limits the magnitude of the parameters in the model by adding a penalty term to the loss function to prevent overfitting. The momentum parameter helps to make the parameter update smoother, speeds up the training, and avoids getting stuck in local optimal solutions. The maximum value of gradient clipping can limit the norm of the gradient within a maximum value, preventing the problem of gradient explosion during training and ensuring the stability and predictability of the training process.
[0107] When the intelligent question-answering model converges to the preset error, the training of the intelligent question-answering model is completed.
[0108] Furthermore, multi-modal data can be input into the intelligent question-answering model, and the question-answering result corresponding to the text instruction can be obtained through the intelligent question-answering model.
[0001] The above is a further description of step S102.
[0002] After inputting the multi-modal data into the intelligent question-answering model, steps S103 to S105 can be executed through the coarse-grained feature extraction module, the fine-grained feature extraction module, and the vision-language processing module of the intelligent question-answering model to obtain the question-answering result corresponding to the text instruction.
[0109] Refer to the appendix Figure 2 , Figure 2 which is the main flowchart of inputting multi-modal data into the intelligent question-answering model to obtain the question-answering result corresponding to the text instruction according to an embodiment of the present application.
[0110] As shown Figure 2 in the figure, the intelligent question-answering model includes a coarse-grained feature extraction module 21, a fine-grained feature extraction module 22, and a vision-language processing module 23.
[0111] Further, in some embodiments of the above step S103, the preprocessed standard three-dimensional image or two-dimensional image of a preset size can be input into the coarse-grained feature extraction module 21.
[0112] Specifically, the shape of the three-dimensional image input into the coarse-grained feature extraction module 21 can be (Batch_size, Channels, Depth, Height, Width), for example, (1, 1, 224, 224, 224); the shape of the two-dimensional image input into the coarse-grained feature extraction module 21 can be (Batch_size, Channels, Height, Width), for example, (1, 1, 224, 224).
[0113] Among them, the "shape" of the input and output of the model or module is used to clearly define the dimensional structure and scale of the tensor. It intuitively shows the number of elements of the data in each dimension. Taking a common three-dimensional image data tensor as an example, its shape may be (Batch_size, Channels, Depth, Height, Width). Among them, Batch_size is the batch size, representing the number of data samples processed at one time. Channels is the number of channels, which is used to illustrate the number of channels of the data in a specific dimension, and different channels store different types of information. Height (H) is the height, Width (W) is the width, and Depth (D) is the depth. In addition, the "shape" of the input to the model or module can also include the sequence length Sequence_length, the feature dimension Feature_dim, the embedding dimension Embedding_dim, etc., which are not limited here.
[0114] Further, the coarse-grained features of the image data can be obtained based on the coarse-grained feature extraction module 21.
[0115] Taking the example of obtaining the coarse-grained features of a three-dimensional image in the medical field based on the coarse-grained feature extraction module 21, the low-level to high-level features of the three-dimensional image can be gradually extracted through the convolutional layers of 3D ResNet-50. Specifically, as the 3D ResNet-50 network is calculated layer by layer, through convolutional operations with different sizes and numbers of convolutional kernels, it is possible to start from simple features (such as edges and textures) and gradually extract more advanced features, such as the local and global shapes of objects in the three-dimensional image, the contours of organs, and the complex structures of tissues.
[0116] The output of 3D ResNet-50 is a feature map, with the shape of (Batch_size, Channels, D, H, W), such as (1, 768, 7, 7, 7). After obtaining the feature map, the spatial dimensions of the feature map can be merged into one dimension, such as (1, 768, 7*7*7), and then the three-dimensional feature map is transposed to (1, 7*7*7, 768). The elements at each position can be regarded as a token, that is to say, there are 7*7*7 tokens here, and the dimension of each token is 768, finally obtaining a coarse-grained feature token sequence.
[0117] The global features can also be extracted through the multi-head self-attention mechanism of 3D ViT to capture the long-range dependencies in 3D images. Specifically, the multi-head self-attention mechanism allows 3D ViT to simultaneously focus on the information of multiple parts in 3D images. Through the self-attention operations of multiple "heads", it can capture different aspects of dependencies in different representation subspaces, thereby extracting global features and capturing the information associations at long distances, such as the potential connections between lesions at different positions.
[0118] The output of 3D ViT is a serialized feature vector, with the shape of (Batch_size, Sequence_length, Feature_dim), such as (1, 343, 768). The elements at each position can be regarded as a token, that is to say, there are 343 tokens here, and the dimension of each token is 768, finally obtaining a coarse-grained feature token sequence.
[0003] The above is a further description of step S103. Next, step S104 will be further described.
[0119] In some embodiments of the above step S104, refer to the appendix Figure 3 , Figure 3 is a schematic diagram of the main step flow for obtaining the fine-grained features of the key regions in the image data based on the coarse-grained features and text instructions of the image data according to an embodiment of the present application. As Figure 3 shown, it mainly includes the following steps S301 to step S302.
[0120] Step S301: Obtain the key region image of the image data based on the coarse-grained features of the image data, text instructions, and the vision-language processing module;
[0121] Among them, the vision-language processing module (vLLM) 23 may include a decoding sub-module.
[0122] In some embodiments, step S301 mainly includes the following steps S3011 to S3013.
[0123] Step S3011: Obtain a first multi-modal sequence based on the coarse-grained features of the image data and the text instruction;
[0124] Among them, step S3011 may include the following steps S3011-1 to S3011-2.
[0125] Step S3011-1: Encode the text instruction to obtain a sequence of word vectors;
[0126] Specifically, it may include the following two steps:
[0127] (1) Segment the text instruction and convert it into a sequence of token identifiers (token ID);
[0128] In some embodiments, a pre-trained language model (such as BERT, GPT, etc.) can be used to encode the text instruction, segment the text instruction, split it into individual words or sub-words as tokens, and then assign a unique identifier, i.e., token ID, to each token, thus forming a sequence of token IDs.
[0129] (2) Convert the sequence of token identifiers into a sequence of word vectors.
[0130] The embedding layer of the language model can map the token ID to the word vector space, convert each token ID into a vector with a certain dimension, and form a sequence of word vectors.
[0131] The shape of the sequence of word vectors is (Batch_size, Sequence_length, Embedding_dim), such as (1, 32, 768).
[0132] Step S3011-2: Concatenate the coarse-grained features and the sequence of word vectors to obtain a first multi-modal sequence.
[0133] Specifically, the sequence of word vectors of the text instruction (with a length of Ltext) and the coarse-grained feature token sequence (with a length of Lc) can be concatenated to obtain a first multi-modal sequence, and the shape of the first multi-modal sequence is (Batch_size, Lc + Ltext, D).
[0134] The above is a further description of step S3011.
[0135] Step S3012: Input the first multi-modal sequence into the vision-language processing module 23, and process the first multi-modal sequence based on the decoding sub-module to obtain the position information of the key regions of the image data.
[0136] Among them, the decoding sub-module includes multiple layers of decoders, specifically, it can be a Transformer decoder. Each layer of the Transformer decoder includes a multi-head self-attention mechanism and a feed-forward neural network.
[0137] Specifically, the Transformer decoder can capture the interaction information between the text instructions and the coarse-grained features through the self-attention mechanism.
[0138] Taking the intelligent question-answering model applied in the medical field as an example, when the text instruction is "Analyze the lung region", the Transformer decoder can, through the self-attention mechanism, associate the semantic information related to "lung" in the text with the coarse-grained features of the lung region in the three-dimensional image, understand the association between the text and the image, and obtain the position information of multiple key regions in the three-dimensional image.
[0139] Among them, the position information of the key regions can be in the form of <focus>Output in the form of tokens. For three-dimensional images, each <focus>The token will contain 6 characters (xxyyzz) afterwards, representing the percentage coordinates xx%, yy%, zz% of this position in the three directions of x, y, and z in the three-dimensional image; for a two-dimensional image, each <focus>The token will contain 4 characters (xxyy) later, indicating the percentage coordinates xx%, yy% in the xy directions of the two-dimensional image at this position.
[0140] The above is a further description of step S3012.
[0141] Step S3013: Crop the image data based on the position information of the key area to obtain the key area image.
[0142] Taking a three-dimensional image as an example, it is possible to <focus>The percentage coordinates of the token are converted into specific three-dimensional image coordinates to obtain the center point coordinates of the key area. Among them, the center point coordinates (x coord , y coord , z coord ) of the key area can be calculated through the following formulas (2)-(4):
[0143] x coord = round(xx * width / 100) (2)
[0144] y coord = round(yy * height / 100) (3)
[0145] z coord = round(zz * depth / 100) (4)
[0146] Among them, round is a mathematical function used to round the value in the parentheses and approximate a value to the nearest integer; the overall meaning of (xx * width / 100) is to convert the position information xx given in percentage form into the pixel position of the three-dimensional image. For example, if the size of the three-dimensional image is 512 * 512 * 512 and the position information of the key area is 255075, then the center point coordinates of the key area are (128, 256, 384), where 128 = 512 * 0.25; (yy * height / 100) and (zz * depth / 100) are the same.
[0147] Furthermore, the image data can be cropped based on the center point coordinates of the key area, and the size of the cropped area can be adjusted according to the task requirements, such as a three-dimensional image of 64 * 64 * 64 voxels or a two-dimensional image of 64 * 64 pixels, so as to obtain the key area image.
[0148] The above is a further description of step S301.
[0149] Step S302: Obtain the fine-grained features of the key area image based on the fine-grained feature extraction module.
[0150] Specifically, the key area image obtained after cropping can be input into the fine-grained feature extraction module 22.
[0151] Among them, the shape of the three-dimensional image input to the fine-grained feature extraction module 22 can be (Batch_size, Channels, Depth, Height, Width), such as (1, 1, 64, 64, 64); the shape of the two-dimensional image input to the fine-grained feature extraction module 22 can be (Batch_size, Channels, Height, Width), such as (1, 1, 64, 64).
[0152] Furthermore, the fine-grained features of the key region image can be obtained based on the fine-grained feature extraction module 22.
[0153] Taking the acquisition of the fine-grained features of the three-dimensional image in the medical field based on the fine-grained feature extraction module 22 as an example, the key region image can be convolved through the convolutional layer of 3D HRNet to extract local features in the key region image, such as the boundaries and textures of different tissues. As the network layer deepens, the lower-level features will gradually combine to form more complex and representative high-level features. For example, the early layers may identify simple line or plane structures, while the subsequent layers can identify the shape of specific organs or the characteristic patterns of lesions.
[0154] The output of 3D HRNet is a high-resolution feature map with a shape of (Batch_size, Channels, D, H, W), such as (1, 768, 8, 8, 8). After obtaining the feature map, the spatial dimensions of the feature map can be merged into one dimension, such as (1, 768, 8*8*8), and then the three-dimensional feature map is transposed to (1, 8*8*8, 768). The elements at each position can be regarded as a token, that is, there are 8*8*8 tokens here, and the dimension of each token is 768, and finally a fine-grained feature token sequence is obtained.
[0155] The above is a further description of step S104. Next, step S105 will be further described.
[0156] In some embodiments of the above step S105, refer to the attached Figure 4 , Figure 4 is a schematic diagram of the main step process for obtaining the Q&A result corresponding to the text instruction based on the coarse-grained feature, text instruction, and fine-grained feature of the key region of the image data according to an embodiment of the present application. As Figure 4 shown, it mainly includes the following steps S401 to step S402.
[0157] Step S401: Concatenate the coarse-grained feature, word vector sequence, and fine-grained feature to obtain a second multi-modal sequence;
[0158] Specifically, the word vector sequence of the text instruction (with a length of Ltext), the coarse-grained feature token sequence (with a length of Lc), and the fine-grained feature token sequence (with a length of Lf) can be concatenated to obtain a second multi-modal sequence, and the shape of the second multi-modal sequence is (Batch_size, Lc + Ltext + Lf, D).
[0159] Step S402: Input the second multi-modal sequence into the vision-language processing module, and process the second multi-modal sequence based on the decoding sub-module to obtain the Q&A result corresponding to the text instruction.
[0160] In some embodiments, step S402 may include the following steps S4021 to S4022.
[0161] Step S4021: Based on the multi-head self-attention mechanism of each layer of the decoder, obtain the interaction information between the text instruction in the second multi-modal sequence and the coarse-grained features of the image data and the fine-grained features of the key-region image.
[0162] Specifically, through the multi-head self-attention mechanism of each layer of the Transformer decoder, the connection between the semantic information in the text instruction and the coarse-grained features of the image data and the fine-grained features of the key-region image can be found.
[0163] Taking the intelligent Q&A model applied in the medical field as an example, when the text instruction is "Identify the tumor in the brain", the multi-head self-attention mechanism can associate words such as "brain" and "tumor" with the coarse-grained features such as the shape and texture of the brain region in the three-dimensional image, and associate "tumor" with the fine-grained features of the tumor in the key-region image, and find the correlation between the text instruction and the image features.
[0164] Step S4022: Based on the interaction information, obtain the answer text corresponding to the text instruction, or obtain the answer text corresponding to the text instruction and the position information of the key region.
[0165] Taking the intelligent Q&A model applied in the medical field as an example, when the input image data of the intelligent Q&A model is a CT image of 512*512*512, and the text instruction input by the user is "What is the diagnosis result of this patient?", the answer text output by the intelligent Q&A model is "There is a tumor in the brain of this patient", and the position information of the key region is <focus>Token: 505050.
[0166] When the image data input into the intelligent Q&A model is a 512*512*512 MRI image and the text instruction input by the user is "Is there a lesion in the liver?", the answer text output by the intelligent Q&A model is "Yes, there is a lesion in the liver.", and the location information of the key area is <focus>Token: 203040。
[0167] In some embodiments of step S105, the key region of the image data may be obtained again based on the coarse-grained features of the image data, the text instruction, and the fine-grained features of the key region, and the fine-grained features of the key region may be obtained again based on the fine-grained feature extraction module 22, so as to repeat step S105 based on the fine-grained features of the key region obtained again until the key region of the image data cannot be obtained, and then output the question-and-answer result corresponding to the text instruction.
[0168] The above is a further description of step S105.
[0169] It should be noted that the above-listed technical solutions are described by taking three-dimensional images in the medical field as an example, and the methods described in the above steps S101 to S105 are also applicable to other fields and two-dimensional images.
[0170] Through the intelligent question-and-answer method based on multi-modal data provided by the present application, it is possible to dynamically adjust the granularity of the image data input into the intelligent question-and-answer model through the combination of coarse-grained feature extraction and fine-grained feature extraction, and generate <focus>The Token indicates the image region that the model needs to further focus on, and can fuse text instructions, coarse-grained features, and fine-grained features into a multi-modal input sequence. It uses the Transformer architecture to process the multi-modal input, captures the interaction information between text and image, realizes efficient and accurate image understanding and generation, improves the depth and flexibility of the interaction between the model and the user, and meets the dynamic and personalized needs of the user.
[0171] It should be noted that although the above embodiments describe the various steps in a specific order, those skilled in the art can understand that in order to achieve the effects of this application, it is not necessary for different steps to be executed in such an order. They can be executed simultaneously (in parallel) or in other orders, and these variations are within the protection scope of this application.
[0172] Those skilled in the art can understand that all or part of the processes in the method of the above-mentioned embodiment of this application can also be completed by instructing relevant hardware through a computer program. The computer program can be stored in a computer-readable storage medium. When the computer program is executed by a processor, the steps of the above-mentioned method embodiments can be realized. Among them, the computer program includes computer program code, and the computer program code can be in the form of source code, object code, executable file, or some intermediate form, etc. The computer-readable storage medium can include: any entity or device, medium, USB flash drive, mobile hard disk, magnetic disk, optical disk, computer memory, read-only memory, random access memory, electrical carrier signal, telecommunication signal, and software distribution medium, etc., that can carry the computer program code.
[0173] Furthermore, this application also provides an electronic device. Refer to the appendix Figure 5 , Figure 5 is a schematic diagram of the main structure of an electronic device according to an embodiment of this application. As Figure 5 shown, the electronic device in the embodiment of this application mainly includes a processor 51 and a memory 52. The memory 52 can be configured to store a program for executing the intelligent question-answering method based on multi-modal data in the above-mentioned method embodiment. The processor 51 can be configured to execute the program in the memory 52, and the program includes but is not limited to the program for executing the intelligent question-answering method based on multi-modal data in the above-mentioned method embodiment. For the sake of convenience of description, only the parts related to the embodiment of this application are shown. For the specific technical details not disclosed, please refer to the method part of the embodiment of this application.
[0174] In some possible embodiments of the present application, the electronic device may include a plurality of processors 51 and a plurality of memories 52. The program for executing the intelligent question-answering method based on multimodal data in the above method embodiments may be divided into multiple sub-programs, and each sub-program may be loaded and run by the processor 51 respectively to execute different steps of the intelligent question-answering method based on multimodal data in the above method embodiments. Specifically, each sub-program may be stored in a different memory 52 respectively, and each processor 51 may be configured to execute the programs in one or more memories 52 to jointly implement the intelligent question-answering method based on multimodal data in the above method embodiments, that is, each processor 51 executes different steps of the intelligent question-answering method based on multimodal data in the above method embodiments respectively to jointly implement the intelligent question-answering method based on multimodal data in the above method embodiments.
[0175] The above-mentioned plurality of processors 51 may be processors deployed on the same device. For example, the above-mentioned electronic device may be a high-performance device composed of a plurality of processors, and the above-mentioned plurality of processors 51 may be the processors configured on the high-performance device. In addition, the above-mentioned plurality of processors 51 may also be processors deployed on different devices. For example, the above-mentioned electronic device may be a server cluster, and the above-mentioned plurality of processors 51 may be the processors on different servers in the server cluster.
[0176] Furthermore, the present application also provides a computer-readable storage medium. In an embodiment of the computer-readable storage medium according to the present application, the computer-readable storage medium may be configured to store a program for executing the intelligent question-answering method based on multimodal data in the above method embodiments, and this program may be loaded and run by the processor to implement the above intelligent question-answering method based on multimodal data. For the sake of convenience of description, only the parts related to the embodiments of the present application are shown. For the specific technical details not disclosed, please refer to the method part of the embodiments of the present application. The computer-readable storage medium may be a memory device formed by various electronic devices. Optionally, the computer-readable storage medium in the embodiments of the present application is a non-transitory computer-readable storage medium.
[0177] It should be noted that the relevant user personal information that may be involved in the embodiments of the present application is strictly in accordance with the requirements of laws and regulations, follows the principles of legality, legitimacy, and necessity, and is based on reasonable purposes of business scenarios to process the personal information actively provided by the user during the use of the product / service or generated due to the use of the product / service, as well as the personal information obtained with the user's authorization.
[0178] The user personal information processed by this application may vary depending on the specific product / service scenario. It shall be subject to the specific scenario of the user using the product / service, and may involve the user's account information, device information, input image data, text instructions or other relevant information. This application will treat the user's personal information and its processing with a high degree of diligence.
[0179] This application attaches great importance to the security of user personal information and has taken security protection measures that meet industry standards and are reasonable and feasible to protect the user's information and prevent personal information from being accessed, publicly disclosed, used, modified, damaged or lost without authorization.
[0180] So far, the technical solution of this application has been described in conjunction with one embodiment shown in the drawings. However, it is easy for those skilled in the art to understand that the protection scope of this application is obviously not limited to these specific embodiments. Without departing from the principle of this application, those skilled in the art can make equivalent changes or substitutions to the relevant technical features, and the technical solutions after these changes or substitutions will fall within the protection scope of this application.< / focus> < / focus> < / focus> < / focus> < / focus> < / focus> < / focus> < / focus> < / focus> < / focus> < / focus> < / focus> < / focus> < / focus>
Claims
1. An intelligent question-answering method based on multimodal data, characterized in that: The method comprises: Acquiring multimodal data; the multimodal data includes image data and text instructions, and the image data includes a three-dimensional image and a two-dimensional image; Inputting the multimodal data into an intelligent question-answering model; the intelligent question-answering model includes a coarse-grained feature extraction module, a fine-grained feature extraction module and a visual language processing module; Acquire the coarse-grained features of the image data based on the coarse-grained feature extraction module; Acquire fine-grained features of key areas in the image data based on the coarse-grained features of the image data and the text instructions; Based on the coarse-grained features of the image data, the text instructions and the fine-grained features of the key area, a question-and-answer result corresponding to the text instruction is obtained; the question-and-answer result includes an answer text.
2. The intelligent question-answering method based on multimodal data according to claim 1, characterized in that: The acquiring of fine-grained features of key areas in the image data based on the coarse-grained features of the image data and the text instruction comprises: Acquire a key area image of the image data based on the coarse-grained features of the image data, the text instructions, and the visual language processing module; The fine-grained features of the key area image are acquired based on the fine-grained feature extraction module.
3. The intelligent question-answering method based on multimodal data according to claim 2, characterized in that: The visual language processing module includes a decoding submodule; the acquiring of the key area image of the image data based on the coarse-grained features of the image data, the text instructions and the visual language processing module includes: Acquire a first multimodal sequence based on the coarse-grained features of the image data and the text instructions; Inputting the first multimodal sequence into the visual language processing module, and processing the first multimodal sequence based on the decoding submodule to obtain position information of the key area of the image data; The image data is cropped based on the position information of the key area to obtain the key area image.
4. The intelligent question-answering method based on multimodal data according to claim 3, characterized in that: The acquiring a first multimodal sequence based on the coarse-grained features of the image data and the text instruction comprises: Encoding the text instruction to obtain a word vector sequence; The coarse-grained features and the word vector sequence are concatenated to obtain a first multimodal sequence.
5. The intelligent question-answering method based on multimodal data according to claim 3, characterized in that: The step of obtaining the question-answer result corresponding to the text instruction based on the coarse-grained features of the image data, the text instruction and the fine-grained features of the key area includes: Concatenating the coarse-grained features, the word vector sequence, and the fine-grained features to obtain a second multimodal sequence; The second multimodal sequence is input into the visual language processing module, and the second multimodal sequence is processed based on the decoding submodule to obtain a question and answer result corresponding to the text instruction.
6. The intelligent question-answering method based on multimodal data according to claim 5, characterized in that: The decoding submodule includes a multi-layer decoder, and the question-answering result also includes the location information of the key area; The step of processing the second multimodal sequence based on the decoding submodule to obtain a question-answer result corresponding to the text instruction includes: Based on a multi-head self-attention mechanism of each layer of decoders, obtaining interactive information between the text instructions in the second multimodal sequence and the coarse-grained features of the image data and the fine-grained features of the key area image; Based on the interaction information, the answer text corresponding to the text instruction is obtained, or the answer text corresponding to the text instruction and the position information of the key area are obtained.
7. The intelligent question-answering method based on multimodal data according to claim 4, characterized in that: The encoding of the text instruction to obtain a word vector sequence includes: Segmenting the text instruction and converting it into a sequence of word unit identifiers; The word element identifier sequence is converted into the word vector sequence.
8. The intelligent question-answering method based on multimodal data according to claim 1, characterized in that: The image data is a standard three-dimensional image or a two-dimensional image of a preset size; and the acquiring of multimodal data includes: Obtaining initial image data; Preprocessing the initial image data to obtain a standard three-dimensional image or a two-dimensional image of the preset size; Acquire a text instruction input by a user; the text instruction includes a question text or a requirement text; and obtain the multimodal data based on the standard three-dimensional image or two-dimensional image of the preset size and the text instruction input by the user.
9. An electronic device comprising a processor and a memory, wherein the memory is suitable for storing a plurality of program codes, wherein: The program code is suitable for being loaded and run by the processor to execute the intelligent question-answering method based on multimodal data according to any one of claims 1 to 8.
10. A computer-readable storage medium storing a plurality of program codes, characterized in that: The program code is suitable for being loaded and run by a processor to execute the intelligent question-answering method based on multimodal data according to any one of claims 1 to 8.
Citation Information
Cited By
Video understanding processing method and device, equipment and storage medium
CN120564105A