Intelligent question answering method and device based on multi-modal information processing, electronic equipment and storage medium

By adopting multimodal information processing methods and search enhancement technology in the intelligent question and answer system, the shortcomings of the existing system in multimodal data processing are solved, and more efficient and accurate information retrieval and answer generation are achieved.

CN119988563APending Publication Date: 2025-05-13中国邮政储蓄银行股份有限公司
View PDF 0 Cites 3 Cited by

Patent Information

Application Number
CN202510139808.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-08
Publication Date
2025-05-13

AI Technical Summary

Technical Problem

The existing intelligent question-and-answer system has problems such as insufficient comprehensiveness in information retrieval, inaccurate intention identification and low retrieval efficiency when processing multimodal data.

Method used

Using an intelligent question-and-answer method based on multimodal information processing, user problems are processed through the LLM model, vectorized processing is used and search enhanced with multimodal information. The joint vectorization model is used to convert data from different modalities into vector representations in a unified space, and a hybrid vector index library is built to achieve efficient information retrieval and answer generation.

Benefits of technology

The intelligent question-and-answer system's processing capability on multimodal data is improved, the accuracy and retrieval efficiency of intention recognition are enhanced, and the high correlation results are quickly returned.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119988563A_ABST
    Figure CN119988563A_ABST
Patent Text Reader

Abstract

The invention discloses an intelligent question answering method and device based on multi-modal information processing, electronic equipment and a storage medium, and the method comprises the steps: responding to a question input by a user, employing an LLM model to process the question, and obtaining a processed question; the processed problem is vectorized, a vectorized problem is obtained, knowledge base vectorization processing is carried out, the knowledge base vectorization processing comprises two processing ways of carrying out separate parallel processing on each mode and carrying out mixed common processing on multiple modes, and a joint vectorization model is adopted when the multiple modes are subjected to mixed common processing; after the vectorization problem and multi-modal information are fused, the vectorization problem and the multi-modal information are input into an LMM model to generate response content, and the multi-modal information is obtained through retrieval according to the vectorization problem. According to the method, the intelligent question-answering system is optimized, the retrieval efficiency is improved, and high-correlation results are returned.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of intelligent question answering systems, and in particular to an intelligent question answering method, device, electronic device, and storage medium based on multimodal information processing. Background Art

[0002] Intelligent Question Answering System, Intelligent Question Answering System (QA System) is an artificial intelligence application based on natural language processing (NLP) and machine learning (ML) technologies, which aims to simulate the ability of human interlocutors to answer user questions. Its core function is to understand the questions raised by users and filter out the most relevant answers from massive amounts of information, and even generate appropriate answers in some cases. The goal of the intelligent question answering system is to provide concise, accurate, and real-time answers to enhance the user experience.

[0003] Retrieval-enhanced RAG is an artificial intelligence technology that combines information retrieval technology with language generation models. RAG retrieves relevant information from external knowledge bases and inputs it as prompts to large language models (LLMs) to enhance the model's ability to handle knowledge-intensive tasks such as question answering, text summarization, and content.

[0004] In the field of intelligent question answering system and retrieval enhancement, the existing technologies mainly have the following shortcomings:

[0005] (1) Single modality limitation: Traditional question-answering systems usually rely only on text information and ignore data from other modalities such as images and audio, resulting in insufficient comprehensiveness of information retrieval.

[0006] (2) Intent recognition is inaccurate. Existing systems have limitations in understanding and parsing user intent and are unable to fully extract the user’s real needs, resulting in inaccurate responses.

[0007] (3) The retrieval efficiency is low. Existing retrieval algorithms are inefficient when processing massive amounts of data and are unable to quickly return highly relevant results. Summary of the invention

[0008] The embodiments of the present application provide an intelligent question and answer method, device, electronic device, and storage medium based on multimodal information processing to process multiple data modalities and generate intelligent questions and answers based on retrieval enhancement.

[0009] The present application embodiment adopts the following technical solutions:

[0010] In a first aspect, an embodiment of the present application provides an intelligent question-answering method based on multimodal information processing, wherein the method comprises:

[0011] In response to a question input by a user, the problem after the question is processed using the LLM model;

[0012] Vectorizing the processed problem to obtain a vectorized problem and performing knowledge base vectorization processing, wherein the knowledge base vectorization processing includes two processing paths: processing each mode separately in parallel and processing multiple modes together, wherein a joint vectorization model is used when processing multiple modes together;

[0013] After the vectorized question is fused with the multimodal information, it is input into the LMM model to generate response content, and the multimodal information is retrieved according to the vectorized question.

[0014] In some embodiments, the joint vectorization model is used when the multiple modes are mixed and processed together, including:

[0015] According to the joint vectorization model, data of different modalities are converted into vector representations in a unified space, and a hybrid vector index library is constructed as a second data storage. The vector database stores vectors generated by the joint vectorization model, and the joint vectorization model includes a CLIP model and a Whisper model.

[0016] In some embodiments, each modality is processed separately and in parallel, including:

[0017] Based on multimodal data collection, text data, image data and audio data from different sources are collected;

[0018] The text data is divided into blocks, a large language model is used to generate a summary of the document file where each data block is located and corresponding context information is generated for each data block, and the summary of the document where each block is located and the context of each block are combined together to form a complete text block description;

[0019] Generating text data of an audio summary for the audio data, translating the audio into a text description, and combining the summary, the translated text, and the storage address of each audio to form a complete audio block description;

[0020] Generate image tags for the image data, provide relevant image understanding and image background text information for the image, and put the tag, text description, and storage address of each image together to form a complete image block description;

[0021] The text blocks, audio description blocks and picture description blocks are processed through a text vectorization model, and the text data is converted into a vector representation to form a vector data index library and store it as the first data in a vector database.

[0022] In some embodiments, the text vectorization model includes a bge-m3 vectorization model,

[0023] The text block, the audio description block and the picture description block are processed by TF-IDF encoding in the BM25 algorithm to form a TF-IDF index library and stored in the ES library.

[0024] In some embodiments, after fusing the vectorized question with the multimodal information, the multimodal information is input into an LMM model to generate response content, and the multimodal information is retrieved according to the vectorized question, including:

[0025] When the user's query question is text only, the text vectorization model converts the query question into a vector and performs a similarity search with the vector data in the vector database stored in the first data store. At the same time, the query question is segmented and TF-IDF is calculated. Combined with the ES library in the first data store, the relevance score is obtained to implement Top K search;

[0026] When the user query question is multimodal, the joint vectorization model directly converts the query question into a vector representation in a hybrid space, performs a similarity search with the vector data in the vector database stored in the second data store, and finally removes duplicates and aggregates the retrieved results;

[0027] The multimodal information is a fusion of the results of context and question enhancement sorting and multi-method search.

[0028] The context information is integrated by combining the context information of the text and the question. The enhanced sorting adopts content-based sorting, which weights the information of different modalities and prioritizes the most relevant information. The result fusion of multi-method search refers to integrating the results from different search methods through simple splicing or weighted averaging.

[0029] In some embodiments, the problem after processing the problem using the LLM model in response to the question input by the user includes:

[0030] Processing the user input question using a processing module including multiple components for generating tasks based on language understanding;

[0031] The LLM model is used to complete each component, wherein the component includes at least one of the following: context acquisition, question rewriting, question expansion, question decomposition, and question enhancement;

[0032] Among them, the context acquisition refers to extracting the description summary of each file from the data set, the question rewriting refers to rewriting the user's query question based on the context information, the question expansion refers to expanding the question to add relevant details or background information on the basis of the rewriting, the question decomposition refers to decomposing the query question into multiple more specific questions, and the question enhancement refers to performing enhancement operations based on the context and analysis results.

[0033] In some embodiments, the joint vectorization model further includes: an encoder and contrastive learning,

[0034] The encoder includes a text encoder, an image encoder and an audio encoder. The basic architecture of the encoder uses Transformer to map the input data into the same vector space, so that similar images, audio and text are close in space;

[0035] The contrastive learning is performed by maximizing the cosine similarity between images, audio, and text from the same sample while minimizing the similarity between different samples.

[0036] In a second aspect, an embodiment of the present application further provides an intelligent question-answering device based on multimodal information processing, wherein the device comprises:

[0037] A response module, used for responding to a question input by a user and processing the question using the LLM model;

[0038] A processing module, used for vectorizing the processed problem to obtain a vectorized problem and perform knowledge base vectorization processing, wherein the knowledge base vectorization processing includes two processing paths: processing each mode separately in parallel and processing multiple modes together, wherein a joint vectorization model is used when processing multiple modes together;

[0039] A generation module is used to fuse the vectorized question with multimodal information and input the result into an LMM model to generate response content, wherein the multimodal information is retrieved according to the vectorized question.

[0040] In a third aspect, an embodiment of the present application further provides an electronic device, comprising: a processor; and a memory arranged to store computer executable instructions, wherein the executable instructions, when executed, cause the processor to perform the above method.

[0041] In a fourth aspect, an embodiment of the present application further provides a computer-readable storage medium, which stores one or more programs. When the one or more programs are executed by an electronic device including multiple application programs, the electronic device executes the above method.

[0042] At least one of the above-mentioned technical solutions adopted in the embodiments of the present application can achieve the following beneficial effects: First, in response to the question input by the user, the LLM model is used to process the question to obtain the processed question. Then, the processed question is vectorized to obtain the vectorized question and the knowledge base is vectorized. Finally, after the vectorized question is fused with the multimodal information, it is input into the LMM model to generate the response content. Through the above method, a retrieval-enhanced generation intelligent question-answering system capable of processing multiple data modalities such as text, images and audio is constructed. Specifically, the large language model LLM is used to deeply analyze and rewrite the user's query to improve the accuracy of intent recognition, and the advanced vectorized indexing technology and the multimodal large model LMM understanding generation capability are used to improve the efficiency of the retrieval process and ensure the rapid return of highly relevant results. BRIEF DESCRIPTION OF THE DRAWINGS

[0043] The drawings described herein are used to provide a further understanding of the present application and constitute a part of the present application. The illustrative embodiments of the present application and their descriptions are used to explain the present application and do not constitute an improper limitation on the present application. In the drawings:

[0044] Figure 1 A flowchart of an intelligent question-answering method based on multimodal information processing in an embodiment of the present application;

[0045] Figure 2 A schematic diagram of an intelligent question-answering process based on multimodal information processing in an embodiment of the present application;

[0046] Figure 3 A schematic diagram of a multimodal data preparation module in an embodiment of the present application;

[0047] FIG4( a) is a schematic diagram of a multimodal data preparation process according to an embodiment of the present application;

[0048] FIG4( b ) is a second schematic diagram of a multimodal data preparation process in an embodiment of the present application;

[0049] Figure 5 A schematic diagram of multimodal joint vectorization model training in an embodiment of the present application;

[0050] Figure 6 Schematic diagram of a multimodal data retrieval module in an embodiment of the present application;

[0051] Figure 7 Schematic diagram of the structure of an intelligent question-answering device based on multimodal information processing in an embodiment of the present application;

[0052] Figure 8 It is a schematic diagram of the structure of an electronic device in an embodiment of the present application. DETAILED DESCRIPTION

[0053] In order to make the purpose, technical solution and advantages of the present application clearer, the technical solution of the present application will be clearly and completely described below in combination with the specific embodiments of the present application and the corresponding drawings. Obviously, the described embodiments are only part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of the present application.

[0054] The technical terms involved in this application are as follows:

[0055] Large Language Model (LLM): is a deep learning algorithm that usually contains tens to hundreds of billions of parameters. It is pre-trained on large-scale text datasets and can perform various natural language processing tasks such as text generation, translation, question answering, etc. LLM is based on the Transformer architecture and processes input sequences in parallel through the self-attention mechanism, thereby efficiently understanding and generating language.

[0056] Large Multimodal Model (LMM): It is an emerging concept in the field of artificial intelligence, which aims to simultaneously process and understand multiple types of data modalities, such as images, text, audio, etc. By integrating the data processing capabilities of different modalities, the model can more comprehensively simulate human perception and understanding capabilities.

[0057] Retrieval-Augmented Generation (RAG): is an architecture designed to enhance the generative capabilities of a large language model (LLM) by retrieving relevant information from external data sources. This approach enables the model to not only rely on its training data, but also access the latest information in real time, thereby providing more accurate and contextually relevant answers.

[0058] Intelligent Question-Answering System: It is a technology in the field of natural language processing (NLP) and artificial intelligence (AI) that aims to understand and answer questions asked by users in natural language. These systems analyze user queries, understand their intent, and provide relevant and accurate answers.

[0059] Traditional retrieval-enhanced question-answering systems mainly rely on structured data and single-modal information processing, and cannot meet users' query needs for multimodal information. Especially in the Internet era with the surge in information volume, users hope to obtain richer and more intuitive answers through multimodal input (such as text, images, and audio). Therefore, it is particularly important to develop a system that can process multimodal information and provide intelligent answers.

[0060] For the multimodal intelligent question-answering system in the related technology, unified processing of multimodal information (including text, images and videos) is realized, and its main steps are as follows: First, the multimodal information receiving module receives multimodal information questions input by users, including voice, text, pictures and videos, etc. Second, the classification module classifies the input information according to its type, such as converting voice into text, extracting features from pictures, etc. Third, mapping processing maps information of different modes to a unified vector space for subsequent analysis. Fourth, the answer selection module calculates similarity based on the mapped vector and selects the most relevant answer from the knowledge base.

[0061] Although the solutions of related technologies have certain multimodal processing capabilities, they still have some limitations in practical applications, such as insufficient support for complex query scenarios and unresponsiveness to real-time data updates.

[0062] In view of the above shortcomings, an intelligent question-answering method is provided in the embodiments of the present application. By introducing a retrieval enhancement mechanism, the intelligent question-answering system can not only process multimodal input, but also obtain the latest information from an external knowledge base in real time, thereby improving the accuracy and relevance of the answer. In addition, since the present application is implemented based on a large language model and a multimodal model, the intelligent question-answering system is more intelligent and humane.

[0063] The technical solutions provided by various embodiments of the present application are described in detail below in conjunction with the accompanying drawings.

[0064] The present application embodiment provides an intelligent question-answering method based on multimodal information processing, such as Figure 1 As shown, a flow chart of an intelligent question-answering method based on multimodal information processing in an embodiment of the present application is provided, and the method at least includes the following steps S110 to S140:

[0065] Step S110 , in response to a question input by a user, using the LLM model to process the question to obtain a processed question.

[0066] The intelligent question-answering system responds to the questions entered by users and rewrites and analyzes them to refine the intent and optimize the expression. Specifically, it uses the powerful natural language processing capabilities of LLM (Large Language Model) to rewrite, expand, analyze, and enhance the query questions with the description summary of each file in the data set as the context, with the purpose of generating more accurate answers, such as Figure 2 The “problem rewriting, expansion, decomposition, and enhancement modules” are shown.

[0067] Step S120, vectorize the processed problem to obtain a vectorized problem and perform knowledge base vectorization processing, the knowledge base vectorization processing includes two processing methods: separate and parallel processing of each modality and mixed processing of multiple modalities, and a joint vectorization model is used when mixed processing of multiple modalities.

[0068] The processed questions are vectorized and converted into a computer-understandable form. At the same time, the multimodal data module is executed. The intelligent question-answering system retrieves relevant information based on the vectorized questions and integrates information from different modalities to form a comprehensive understanding, such as Figure 2 The “Multimodal Data Preparation Module” is shown.

[0069] Step S130: After fusing the vectorized question with the multimodal information, the multimodal information is input into an LMM model to generate response content, wherein the multimodal information is retrieved according to the vectorized question.

[0070] The vectorized question is fused with the multimodal information and then input into the LMM model. The intelligent question answering system generates the final answer using the LMM (multimodal large model) and generates the user-visible response content based on the output. Figure 2 As shown, Large Multimodal Model (LMM).

[0071] Through the above method, the LLM model is used to process the processed question, thereby improving the query efficiency and accuracy. The language model (LLM) is used to rewrite, expand, decompose and enhance the original question to better capture the user's intention.

[0072] Through the above method, in terms of data processing, the processed problem is vectorized to obtain a vectorized problem and perform knowledge base vectorization processing. The knowledge base vectorization processing includes two processing paths: separate parallel processing of each modality and mixed and processed multiple modalities together. When the multiple modalities are mixed and processed together, a joint vectorization model is used. A dual-path multimodal data processing scheme is adopted, which are separate parallel processing and mixed and processed together. The separate parallel processing scheme processes data of different modalities separately, specifically generating summaries from text and generating contexts in blocks, converting images into labels and descriptions, and transcribing audio into text content and summaries, and then vectorizing them separately. The mixed and processed scheme integrates multimodal data and processes them together, and maps the data to the same vector space through a unified model to support subsequent information retrieval and fusion operations.

[0073] Different from the related technologies, there are problems such as insufficient support for complex query scenarios and slow response to real-time data updates. Through this application, by introducing the retrieval enhancement mechanism RAG, the intelligent question-answering system can not only process multimodal input, but also obtain the latest information from the external knowledge base in real time, thereby improving the accuracy and relevance of the answer. And the above is achieved based on a large language model and a multimodal model, making the intelligent question-answering system more intelligent and humane.

[0074] In one embodiment of the present application, a joint vectorization model is used when mixing and processing multiple modalities, including: converting data of different modalities into vector representations in a unified space according to the joint vectorization model, and constructing a mixed vector index library as a second data storage, wherein the vector database stores vectors generated by the joint vectorization model, and the joint vectorization model includes a CLIP model and a Whisper model.

[0075] As shown in Figure 4(b), multiple modalities are mixed and processed together to obtain data storage 2. Specifically, a multimodal joint vectorization model (Joint Embedding Model) is used to convert data sets of different modalities (text, image, audio) into vector representations in a unified space, a hybrid vector index library is constructed, and vectors generated by the joint vectorization model are stored in the vector database to support efficient retrieval.

[0076] like Figure 5 As shown in the figure, the design of the multimodal joint vectorization model is based on the CLIP and Whisper models, and performs modal fusion and expansion, mainly including two main parts: encoder and contrastive learning. In the encoder part, there are three independent encoders: text encoder, image encoder and audio encoder. The encoder infrastructure uses Transformer, which maps the input data to the same vector space, so that similar images, audio and text are close in space.

[0077] For example, images, text, and audio of the same topic will form clusters in the vector space. The contrastive learning part refers to maximizing the cosine similarity between images, audio, and text from the same sample, while minimizing the similarity between different samples. This method enables the model to learn the common features between images, audio, and text.

[0078] The design of the multimodal joint vectorization model is based on the CLIP and Whisper models, and includes three independent encoders for processing text, images, and audio. Through contrastive learning training, the model can map data from different modalities into the same vector space, thereby achieving cross-modal information retrieval and fusion.

[0079] In hybrid joint processing, data of different modalities are directly converted into vector representations in a unified space by a multimodal joint vectorization model, and a hybrid vector index library is established. This multimodal joint vectorization model is based on the CLIP and Whisper models and is designed to extend and include three independent encoders for text, image, and audio. It uses the Transformer architecture and contrastive learning training method to maximize the similarity between different modalities of the same sample, while minimizing the similarity between different samples.

[0080] In one embodiment of the present application, each modality is processed separately and in parallel, including: based on multimodal data collection, collecting text data, image data and audio data from different sources; segmenting the text data, using a large language model to generate a summary of the document file where each data block is located and generating corresponding context information for each data block, and combining the summary of the document where each block is located and the context of each block together to form a complete text block description; generating text data of audio summaries for the audio data, translating the audio into text descriptions, and combining the summary of each audio, the translated text, and the storage address where the audio is located to form a complete audio block description; generating image labels for the image data, providing text information related to image understanding and image background for the image, and putting together the label, text description, and storage address of each image to form a complete image block description; processing the text blocks, audio description blocks, and image description blocks through a text vectorization model, converting the text data into vector representations, forming a vector data index library and storing it as the first data in the vector database.

[0081] like Figure 3 As shown in the figure, the multimodal data preparation module consists of four parts: multimodal data collection, data cleaning and processing, index library construction, and vector index library construction. Please continue to refer to Figure 3 ,Multimodal data collection is responsible for collecting data from different sources, including but not limited to text, images, audio, etc. Data cleaning and processing are responsible for denoising, removing noise and irrelevant information in the data to improve data quality, and format standardization, converting the collected data into a unified format for subsequent processing.

[0082] The index library construction part is responsible for further processing of multimodal documents or files. The first is document segmentation, which divides the document into multiple manageable blocks for easy retrieval and processing. The second is summary, label and context generation, which generates a summary and label for each block to provide context information. The vector index library creation part is responsible for establishing vector indexes based on the split blocks, generated features and summaries to support efficient retrieval. After the preparation of the data and index library is completed, the subsequent retrieval module can be carried out.

[0083] As shown in FIG4(a) and FIG4(b), the multimodal data preparation module mainly includes two solutions for processing multimodal data. One is to process each modality separately and in parallel, and the other is to process multiple modalities together. The following is a specific description of processing each modality separately and in parallel:

[0084] The data set collects data in multiple formats, including text (such as PDF, Word), images (such as JPG, PNG), and audio (such as MP3, WAV). Each modality is processed separately and in parallel to obtain data storage1.

[0085] Specifically, for text data, firstly, data segmentation is performed, and the collected data is divided into multiple chunks (Chunk 1, Chunk 2, ..., Chunk N) for easy processing and analysis. Secondly, context and summary generation is performed, and the summary (Abstract) of the document file where each data chunk is located is generated using a large language model. Corresponding context information (Context) is generated for each data chunk to better understand its position and meaning in the overall document. In particular, the context information is 50 to 100 tokens, and the prompt word for generating the context is:

[0086] “ <document>

[0087] {{Data Collection Summary}}

[0088] < / document>

[0089] This is the snippet we want to place throughout the document.

[0090] <chunk>

[0091] {{Snippet content}}

[0092] < / chunk>

[0093] Please provide a brief context to locate this snippet within the document, in order to improve search retrieval of the snippet. Answers should only provide brief context and should not include anything else. "

[0094] Combine the abstract of the document where each chunk is located and the context of each chunk together to form a complete text chunk description (e.g. Abstract 1+Context 1+Chunk 1). Perform the same processing on all chunks to ensure the consistency and integrity of the information.

[0095] For pictures and audio data, audio and image data from different sources are input into the multimodal generator for processing.

[0096] The audio data is processed as follows:

[0097] 1. Generate an audio summary, generate a summary for the input audio, the summary is text data (such as "Abstract1").

[0098] 2. Generate a transcribed text, and describe the audio transcribed text to better understand the processing content (such as "Context1").

[0099] Put together the summary, translated text, and storage address of each audio to form a complete audio block description. The output format is required to be integrated into the format of "Abstract 1+Context 1+Audio 1URL" and so on.

[0100] The image data is processed as follows:

[0101] 1. Generate image labels and generate description labels (such as "Label 1") for the input images, such as "car", "animal", etc.

[0102] 2. Generate a text description of the image understanding, providing relevant text information such as image understanding and image background (such as "Context 1") for the image.

[0103] Put together the label, text description, and storage address of each image to form a complete image block description. The format is required to be integrated into the format of "Label 1+Context 1+Image1 URL". Process the text block, audio description block, and image description block through the text vectorization model, convert the text data into vector representation, form a vector data index library and store it in the vector database to facilitate subsequent support for efficient similarity retrieval.

[0104] Based on the summary of the data set, contextual information is provided to further optimize the query, and the LLM is guided to complete the tasks of each optimization stage through carefully designed prompt words. The above dual-path multimodal data processing solution includes two modes: separate parallel processing and mixed joint processing. In separate parallel processing, text data is described by block segmentation and summary generation; image data generates labels and descriptions and combines them with image URLs; audio data also generates summaries and combines them with audio URLs. All data are vectorized using the bge-m3 model, and vector index libraries and TF-IDF index libraries are built.

[0105] In one embodiment of the present application, the text vectorization model includes a bge-m3 vectorization model, and the text block, the audio description block, and the image description block are processed through the TF-IDF encoding in the BM25 algorithm to form a TF-IDF index library and store it in the ES library.

[0106] In the embodiment of the present application, the bge-m3 vectorization model is adopted. At the same time, the text block, audio description block and picture description block are processed by TF-IDF encoding in the BM25 algorithm to form a TF-IDF (term frequency-inverse document frequency) index library and stored in the ElasticSearch library, which is suitable for traditional keyword retrieval.

[0107] It should be noted that the BM25 used is a ranking algorithm for information retrieval and text mining, which is used to calculate the relevance score between the query and the document. The basic idea is to segment the query to obtain a series of feature words, and then calculate the relevance score between each feature word and the document. Finally, the scores of all feature words are weighted and summed to obtain the overall relevance score between the query and the document.

[0108] In one embodiment of the present application, after the vectorized question is fused with the multimodal information, it is input into the LMM model to generate response content. The multimodal information is retrieved according to the vectorized question, including: when the user query question is only text, the text vectorization model converts the query question into a vector, and performs a similarity search with the vector data in the vector database stored in the first data store. At the same time, the query question is segmented and TF-IDF is calculated, and the relevance score is obtained in combination with the ES library in the first data store to realize the Top K search; when the user query question is multimodal, the joint vectorization model directly converts the query question into a vector representation of the mixed space, and performs a similarity search with the vector data in the vector database stored in the second data store, and finally removes duplicates and summarizes the retrieved results; wherein, the result fusion of context and question enhanced sorting and multi-method search is used as the multimodal information, the context information is integrated by combining the context information of the text and the question, the enhanced sorting adopts content-based sorting, and the most relevant information is displayed first by weighting the information of different modalities, and the result fusion of the multi-method search refers to the integration of results from different search methods through simple splicing or weighted averaging.

[0109] like Figure 6 As shown in the figure, the retrieval module processes separately according to the question modality input by the user. When the user query question is text only, the text vectorization model bge-m3 converts the query question into a vector and performs a similarity search with the vector data in the vector database in data storage 1. At the same time, the BM25 algorithm segments the query question and calculates TF-IDF. Combined with the ES library in data storage 1, the relevance score is obtained to realize the Top K search.

[0110] like Figure 6As shown in FIG. 1 , when the user query question is multimodal, the joint vectorization model directly converts the query question into a vector representation in a hybrid space and performs a similarity search with the vector data in the vector database in data storage 2. Finally, the retrieved results are deduplicated and summarized and input into the next module.

[0111] In the multimodal information fusion module, context and question enhancement ranking and multi-method search result fusion are adopted. Context information integration enhances the ranking of search results by combining the context information of text and question.

[0112] Context information may include the user's historical queries, the relevance of the current task, etc.

[0113] The enhanced sorting mechanism adopts content-based sorting, which ensures that the most relevant information is displayed first by weighting information of different modalities.

[0114] The fusion of multi-method search results refers to the integration of results from different search methods (keyword matching, semantic search) through simple splicing or weighted averaging to improve the accuracy and relevance of the overall search.

[0115] In one embodiment of the present application, in response to a question input by a user, the problem after the question is processed by using an LLM model includes: using a processing module including multiple components based on language understanding to generate tasks to process the question input by the user; using the LLM model to complete each component, and the component includes at least one of the following: context acquisition, question rewriting, question expansion, question decomposition, and question enhancement; wherein, the context acquisition refers to extracting a description summary of each file from a data set, the question rewriting refers to rewriting the user's query question based on context information, the question expansion refers to expanding the question to add relevant details or background information based on the rewriting, the question decomposition refers to decomposing the query question into multiple more specific questions, and the question enhancement refers to performing enhancement operations in combination with the context and analysis results.

[0116] The question rewriting, expansion, decomposition, and enhancement module specifically includes five components: context acquisition, question rewriting, question expansion, question decomposition, and question enhancement. Each component is essentially a language understanding generation task, which is completed based on a large language model. Context acquisition refers to extracting a description summary of each file from a data set. The summary can be generated through a large model, and these summaries serve as context information for subsequent processing. Question rewriting refers to rewriting the user's query question based on context information to make it more accurate, clear, and easy to understand. The prompt word design example for question rewriting is:

[0117] <document>

[0118] {{Data Collection Summary}}

[0119] < / document>

[0120] This is a query from a user

[0121] <query>

[0122] {{Original query question}}

[0123] < / query>

[0124] Please rephrase the query question based on the contextual information provided in the above data set summary to make it more accurate, clear, and easy to understand. Only answer the rephrased query question, no other content.

[0125] Question expansion refers to expanding the question on the basis of rewriting, adding relevant details or background information to improve the comprehensiveness of the search. Question decomposition refers to breaking down the query question into multiple more specific questions in order to gradually solve the user's needs. Question enhancement refers to combining the context and analysis results to perform enhancement operations, including possible keyword extraction and semantic enhancement, to improve the quality of the query. The prompt word design method of these components is similar to question rewriting. For example, the prompt word for question decomposition is:

[0126] <document>

[0127] {{Data Collection Summary}}

[0128] < / document>

[0129] This is a query from a user

[0130] <query>

[0131] {{Original query question}}

[0132] < / query>

[0133] Please decompose the query into multiple more specific questions based on the contextual information provided by the data collection summary, so as to gradually solve the user's needs. Only answer the decomposed questions, and no other content.

[0134] In terms of system architecture, the above method integrates multimodal data preparation, query optimization, retrieval, information fusion and answer generation modules, and uses a multimodal large model (LMM) to generate the final answer. The retrieval module selects different retrieval paths according to the query question modality, and performs deduplication and aggregation processing on the results to ensure the relevance and accuracy of the query results. These designs enable the present invention to provide efficient and accurate services when processing complex multimodal data.

[0135] In one embodiment of the present application, the joint vectorization model also includes: an encoder and contrastive learning, the encoder includes a text encoder, an image encoder and an audio encoder, the basic architecture of the encoder uses Transformer to map the input data to the same vector space, so that similar images, audio and text are close in space; the contrastive learning maximizes the cosine similarity between images, audio and text from the same sample while minimizing the similarity between different samples.

[0136] Training of multimodal joint vectorization model:

[0137] The first step is feature extraction. The text encoder extracts text features T through a multi-head self-attention mechanism. The image encoder divides the image into multiple patches and inputs these patches as sequences into the Transformer's multi-head attention mechanism to extract image features I. For the audio encoder, the audio is converted into a log-Mel spectrum map so that it can be input into the multi-head self-attention mechanism to map it into a series of hidden states, namely, audio features A.

[0138] The second step is modality fusion and contrastive learning, which fuses the feature vectors of different modalities to form a joint representation: V = [T; I; A]. The objective function of contrastive learning is defined to maximize the similarity between different modalities of the same instance and minimize the similarity between different instances.

[0139] The loss function L is expressed as:

[0140]

[0141] N is the sample size;

[0142] V i + Is with V i The associated eigenvectors of other modes;

[0143] sim(x, y) calculates the cosine similarity between two vectors;

[0144] τ is a temperature parameter that controls the sharpness of the similarity distribution.

[0145] The third step is model training. The model is trained using the optimization algorithm Adam to update the parameters of all encoders and fusion layers to minimize the loss function L.

[0146] The above method uses a multi-head self-attention mechanism to extract features from text, divides the image into patches and extracts features through a multi-head attention mechanism, and converts the audio into a log-Mel spectrogram and maps it into a hidden state, and finally fuses these features into a unified representation V.

[0147] The present application embodiment also provides an intelligent question-answering device 700 based on multimodal information processing, such as Figure 7 As shown, a schematic diagram of the structure of an intelligent question-answering device based on multimodal information processing in an embodiment of the present application is provided. The intelligent question-answering device based on multimodal information processing 700 at least includes: a response module 710, a processing module 720, and a generation module 730, wherein:

[0148] In one embodiment of the present application, the response module 710 is specifically used to: respond to a question input by a user and use the LLM model to process the question to obtain a processed question.

[0149] The intelligent question-answering system responds to the questions entered by users and rewrites and analyzes them to refine the intent and optimize the expression. Specifically, it uses the powerful natural language processing capabilities of LLM (Large Language Model) to rewrite, expand, analyze, and enhance the query questions with the description summary of each file in the data set as the context, with the purpose of generating more accurate answers, such as Figure 2 The “problem rewriting, expansion, decomposition, and enhancement modules” are shown.

[0150] In one embodiment of the present application, the processing module 720 is specifically used to: vectorize the processed problem to obtain a vectorized problem and perform knowledge base vectorization processing, the knowledge base vectorization processing includes: two processing paths: separate and parallel processing of each modality and mixed and processed multiple modalities together, and a joint vectorization model is used when mixing and processing multiple modalities together.

[0151] The processed questions are vectorized and converted into a computer-understandable form. At the same time, the multimodal data module is executed. The intelligent question-answering system retrieves relevant information based on the vectorized questions and integrates information from different modalities to form a comprehensive understanding, such as Figure 2 The “Multimodal Data Preparation Module” is shown.

[0152] In one embodiment of the present application, the generation module 730 is specifically used to: fuse the vectorized question with the multimodal information, and input the result into an LMM model to generate response content, wherein the multimodal information is retrieved based on the vectorized question.

[0153] The vectorized question is fused with the multimodal information and then input into the LMM model. The intelligent question answering system generates the final answer using the LMM (multimodal large model) and generates the user-visible response content based on the output. Figure 2 As shown, Large Multimodal Model (LMM).

[0154] In one embodiment of the present application, the processing module 720 is also used to

[0155] According to the joint vectorization model, data of different modalities are converted into vector representations in a unified space, and a hybrid vector index library is constructed as a second data storage. The vector database stores vectors generated by the joint vectorization model, and the joint vectorization model includes a CLIP model and a Whisper model.

[0156] In one embodiment of the present application, the processing module 720 is also used to

[0157] Based on multimodal data collection, text data, image data and audio data from different sources are collected;

[0158] The text data is divided into data blocks, a large language model is used to generate a summary of the document file where each data block is located and corresponding context information is generated for each data block, and the summary of the document where each block is located and the context of each block are combined together to form a complete text block description;

[0159] Generating text data of an audio summary for the audio data, translating the audio into a text description, and combining the summary, the translated text, and the storage address of each audio to form a complete audio block description;

[0160] Generate image tags for the image data, provide relevant image understanding and image background text information for the image, and put the tag, text description, and storage address of each image together to form a complete image block description;

[0161] The text blocks, audio description blocks and picture description blocks are processed through a text vectorization model, and the text data is converted into a vector representation to form a vector data index library and store it as the first data in a vector database.

[0162] In one embodiment of the present application, the text vectorization model includes a bge-m3 vectorization model.

[0163] The text block, the audio description block and the picture description block are processed by TF-IDF encoding in the BM25 algorithm to form a TF-IDF index library and stored in the ES library.

[0164] In one embodiment of the present application, the generating module 730 is also used to

[0165] When the user's query question is text only, the text vectorization model converts the query question into a vector and performs a similarity search with the vector data in the vector database stored in the first data store. At the same time, the query question is segmented and TF-IDF is calculated. Combined with the ES library in the first data store, the relevance score is obtained to implement Top K search;

[0166] When the user query question is multimodal, the joint vectorization model directly converts the query question into a vector representation in a hybrid space, performs a similarity search with the vector data in the vector database stored in the second data store, and finally removes duplicates and aggregates the retrieved results;

[0167] The multimodal information is a fusion of the results of context and question enhancement sorting and multi-method search.

[0168] The context information is integrated by combining the context information of the text and the question. The enhanced sorting adopts content-based sorting, which weights the information of different modalities and prioritizes the most relevant information. The result fusion of multi-method search refers to integrating the results from different search methods through simple splicing or weighted averaging.

[0169] In one embodiment of the present application, the response module 710 is also used to

[0170] Processing the user input question using a processing module including multiple components for generating tasks based on language understanding;

[0171] The LLM model is used to complete each component, wherein the component includes at least one of the following: context acquisition, question rewriting, question expansion, question decomposition, and question enhancement;

[0172] Among them, the context acquisition refers to extracting the description summary of each file from the data set, the question rewriting refers to rewriting the user's query question based on the context information, the question expansion refers to expanding the question to add relevant details or background information on the basis of the rewriting, the question decomposition refers to decomposing the query question into multiple more specific questions, and the question enhancement refers to performing enhancement operations based on the context and analysis results.

[0173] In one embodiment of the present application, the joint vectorization model further includes: an encoder and contrastive learning,

[0174] The encoder includes a text encoder, an image encoder and an audio encoder. The basic architecture of the encoder uses Transformer to map the input data into the same vector space, so that similar images, audio and text are close in space;

[0175] The contrastive learning is performed by maximizing the cosine similarity between images, audio, and text from the same sample while minimizing the similarity between different samples.

[0176] It can be understood that the above-mentioned intelligent question and answer device based on multimodal information processing can implement the various steps of the intelligent question and answer method based on multimodal information processing provided in the aforementioned embodiments, and the relevant explanations about the intelligent question and answer method based on multimodal information processing are applicable to the intelligent question and answer device based on multimodal information processing, which will not be repeated here.

[0177] Figure 8 This is a schematic diagram of the structure of an electronic device according to an embodiment of the present application. Figure 8At the hardware level, the electronic device includes a processor, and optionally also includes an internal bus, a network interface, and a memory. The memory may include a memory, such as a high-speed random access memory (RAM), and may also include a non-volatile memory (non-volatile memory), such as at least one disk storage. Of course, the electronic device may also include hardware required for other services.

[0178] The processor, network interface and memory can be interconnected through an internal bus, which can be an ISA (Industry Standard Architecture) bus, a PCI (Peripheral Component Interconnect) bus or an EISA (Extended Industry Standard Architecture) bus. The bus can be divided into an address bus, a data bus, a control bus, etc. For ease of representation, Figure 8 Only one bidirectional arrow is used in the diagram, but this does not mean that there is only one bus or only one type of bus.

[0179] The memory is used to store the program. Specifically, the program may include a program code, and the program code includes a computer operation instruction. The memory may include a memory and a non-volatile memory, and provides instructions and data to the processor.

[0180] The processor reads the corresponding computer program from the non-volatile memory into the memory and then runs it, forming an intelligent question-answering device at the logical level. The processor executes the program stored in the memory and is specifically used to perform the following operations:

[0181] In response to a question input by a user, the problem after the question is processed using the LLM model;

[0182] Vectorizing the processed problem to obtain a vectorized problem and performing knowledge base vectorization processing, wherein the knowledge base vectorization processing includes two processing paths: processing each mode separately in parallel and processing multiple modes together, wherein a joint vectorization model is used when processing multiple modes together;

[0183] After the vectorized question is fused with the multimodal information, it is input into the LMM model to generate response content, and the multimodal information is retrieved according to the vectorized question.

[0184] The above application Figure 1The method performed by the intelligent question-answering device based on multimodal information processing disclosed in the illustrated embodiment can be applied to a processor or implemented by a processor. The processor may be an integrated circuit chip with signal processing capabilities. In the implementation process, each step of the above method can be completed by an integrated logic circuit of hardware in the processor or an instruction in the form of software. The above processor may be a general-purpose processor, including a central processing unit (CPU), a network processor (NP), etc.; it may also be a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA) or other programmable logic devices, discrete gates or transistor logic devices, discrete hardware components. The methods, steps and logic block diagrams disclosed in the embodiments of the present application can be implemented or executed. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc. The steps of the method disclosed in conjunction with the embodiments of the present application can be directly embodied as being executed by a hardware decoding processor, or executed by a combination of hardware and software modules in a decoding processor. The software module can be located in a storage medium mature in the art such as a random access memory, a flash memory, a read-only memory, a programmable read-only memory, or an electrically erasable programmable memory, a register, etc. The storage medium is located in the memory, and the processor reads the information in the memory and completes the steps of the above method in combination with its hardware.

[0185] The electronic device may also perform Figure 1 A method for executing an intelligent question-answering device based on multimodal information processing in Figure 1 The functions of the illustrated embodiment will not be described in detail in the embodiments of the present application.

[0186] The present application also provides a computer-readable storage medium, which stores one or more programs, wherein the one or more programs include instructions, which, when executed by an electronic device including multiple application programs, enable the electronic device to execute Figure 1 The method performed by the intelligent question-answering device based on multimodal information processing in the illustrated embodiment is specifically used to perform:

[0187] In response to a question input by a user, the problem after the question is processed using the LLM model;

[0188] Vectorizing the processed problem to obtain a vectorized problem and performing knowledge base vectorization processing, wherein the knowledge base vectorization processing includes two processing paths: processing each mode separately in parallel and processing multiple modes together, wherein a joint vectorization model is used when processing multiple modes together;

[0189] After the vectorized question is fused with the multimodal information, it is input into the LMM model to generate response content, and the multimodal information is retrieved according to the vectorized question.

[0190] Those skilled in the art will appreciate that embodiments of the present invention may be provided as methods, systems, or computer program products. Therefore, the present invention may take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware. Moreover, the present invention may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0191] The present invention is described with reference to flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to embodiments of the present invention. It should be understood that each process and / or block in the flowchart and / or block diagram, as well as the combination of processes and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 A process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.

[0192] These computer program instructions may also be stored in a computer-readable memory capable of directing a computer or other programmable data processing device to operate in a specific manner, so that the instructions stored in the computer-readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 A process or multiple processes and / or boxes Figure 1 A function specified in one or more boxes.

[0193] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operating steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing instructions for implementing the process. Figure 1 A process or multiple processes and / or boxes Figure 1The steps for the functions specified in one or more boxes.

[0194] In a typical configuration, a computing device includes one or more processors (CPU), input / output interfaces, network interfaces, and memory.

[0195] The memory may include non-permanent storage in a computer-readable medium, random access memory (RAM) and / or non-volatile memory in the form of read-only memory (ROM) or flash RAM. The memory is an example of a computer-readable medium.

[0196] Computer readable media include permanent and non-permanent, removable and non-removable media that can be implemented by any method or technology to store information. Information can be computer readable instructions, data structures, program modules or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, compact disk read-only memory (CD-ROM), digital versatile disk (DVD) or other optical storage, magnetic cassettes, magnetic tape magnetic disk storage or other magnetic storage devices or any other non-transmission media that can be used to store information that can be accessed by a computing device. As defined herein, computer readable media does not include temporary computer readable media (transitory media), such as modulated data signals and carrier waves.

[0197] It should also be noted that the terms "include", "comprises" or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, method, commodity or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, commodity or device. In the absence of more restrictions, the elements defined by the sentence "comprises a ..." do not exclude the existence of other identical elements in the process, method, commodity or device including the elements.

[0198] Those skilled in the art will appreciate that the embodiments of the present application may be provided as methods, systems or computer program products. Therefore, the present application may adopt the form of a complete hardware embodiment, a complete software embodiment or an embodiment in combination with software and hardware. Moreover, the present application may adopt the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) that contain computer-usable program code.

[0199] The above is only an embodiment of the present application and is not intended to limit the present application. For those skilled in the art, the present application may have various changes and variations. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application should be included in the scope of the claims of the present application.

Claims

1. An intelligent question-answering method based on multimodal information processing, wherein: The method comprises: In response to a question input by a user, the problem after the question is processed using the LLM model; Vectorizing the processed problem to obtain a vectorized problem and performing knowledge base vectorization processing, wherein the knowledge base vectorization processing includes two processing paths: processing each mode separately in parallel and processing multiple modes together, wherein a joint vectorization model is used when processing multiple modes together; After the vectorized question is fused with the multimodal information, it is input into the LMM model to generate response content, and the multimodal information is retrieved according to the vectorized question.

2. The method of claim 1, wherein: The joint vectorization model is used when the multiple modes are mixed and processed together, including: According to the joint vectorization model, data of different modalities are converted into vector representations in a unified space, and a hybrid vector index library is constructed as a second data storage. The vector database stores vectors generated by the joint vectorization model, and the joint vectorization model includes a CLIP model and a Whisper model.

3. The method of claim 2, wherein: Each modality is processed separately and in parallel, including: Based on multimodal data collection, text data, image data and audio data from different sources are collected; The text data is divided into blocks, a large language model is used to generate a summary of the document file where each data block is located and corresponding context information is generated for each data block, and the summary of the document where each block is located and the context of each block are combined together to form a complete text block description; Generating text data of an audio summary for the audio data, translating the audio into a text description, and combining the summary, the translated text, and the storage address of each audio to form a complete audio block description; Generate image tags for the image data, provide relevant image understanding and image background text information for the image, and put the tag, text description, and storage address of each image together to form a complete image block description; The text blocks, audio description blocks and picture description blocks are processed through a text vectorization model, and the text data is converted into a vector representation to form a vector data index library and store it as the first data in a vector database.

4. The method of claim 3, wherein: The text vectorization model includes a bge-m3 vectorization model, The text block, the audio description block and the picture description block are processed by TF-IDF encoding in the BM25 algorithm to form a TF-IDF index library and stored in the ES library.

5. The method of claim 3, wherein: After fusing the vectorized question with the multimodal information, the multimodal information is input into the LMM model to generate response content, wherein the multimodal information is retrieved according to the vectorized question, including: When the user's query question is text only, the text vectorization model converts the query question into a vector and performs a similarity search with the vector data in the vector database stored in the first data store. At the same time, the query question is segmented and TF-IDF is calculated. Combined with the ES library in the first data store, the relevance score is obtained to implement Top K search; When the user query question is multimodal, the joint vectorization model directly converts the query question into a vector representation in a hybrid space, performs a similarity search with the vector data in the vector database stored in the second data store, and finally removes duplicates and aggregates the retrieved results; The multimodal information is a fusion of the results of context and question enhancement sorting and multi-method search. The context information is integrated by combining the context information of the text and the question. The enhanced sorting adopts content-based sorting, which weights the information of different modalities and prioritizes the most relevant information. The result fusion of multi-method search refers to integrating the results from different search methods through simple splicing or weighted averaging.

6. The method of claim 1, wherein: The problem after the problem is processed by using the LLM model in response to the problem input by the user includes: Processing the user input question using a processing module including multiple components for generating tasks based on language understanding; The LLM model is used to complete each component, wherein the component includes at least one of the following: context acquisition, question rewriting, question expansion, question decomposition, and question enhancement; Among them, the context acquisition refers to extracting the description summary of each file from the data set, the question rewriting refers to rewriting the user's query question based on the context information, the question expansion refers to expanding the question to add relevant details or background information on the basis of the rewriting, the question decomposition refers to decomposing the query question into multiple more specific questions, and the question enhancement refers to performing enhancement operations based on the context and analysis results.

7. The method of claim 2, wherein: The joint vectorization model also includes: an encoder and contrastive learning, The encoder includes a text encoder, an image encoder and an audio encoder. The basic architecture of the encoder uses Transformer to map the input data into the same vector space, so that similar images, audio and text are close in space; The contrastive learning is performed by maximizing the cosine similarity between images, audio, and text from the same sample while minimizing the similarity between different samples.

8. An intelligent question-answering device based on multimodal information processing, wherein: The device comprises: A response module, used for responding to a question input by a user and processing the question using the LLM model; A processing module, used for vectorizing the processed problem to obtain a vectorized problem and perform knowledge base vectorization processing, wherein the knowledge base vectorization processing includes two processing paths: processing each mode separately in parallel and processing multiple modes together, wherein a joint vectorization model is used when processing multiple modes together; A generation module is used to fuse the vectorized question with multimodal information and input the resultant information into an LMM model to generate response content, wherein the multimodal information is retrieved based on the vectorized question.

9. An electronic device, comprising: processor; as well as A memory arranged to store computer executable instructions, which when executed cause the processor to perform the method of any one of claims 1 to 7.

10. A computer-readable storage medium storing one or more programs, which, when executed by an electronic device including a plurality of application programs, causes the electronic device to execute any one of the methods of claims 1 to 7.

Citation Information

Cited By

  • MLLM-based energy storage battery data automatic retrieval analysis method and system

    CN120407753A

  • Question and answer method and device, equipment, medium and program product

    CN120821815A

  • Question and answer method, apparatus, device, medium, and program product

    CN120821815B