Intelligent query method and device, computer equipment and storage medium
By building a multimodal corpus and utilizing a pre-trained visual language model, the problem of inaccurate query results in the healthcare and fintech fields of multimodal query technology is solved, achieving more efficient information integration and answer generation.
Patent Information
- Application Number
- CN202510844782.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-23
- Publication Date
- 2025-09-26
AI Technical Summary
Existing multimodal query technologies in the fields of healthcare and financial technology lack accuracy and effectiveness in query results and cannot meet the diverse needs of users.
Build a multimodal corpus consisting of text, image, and video corpora, identify the target modality corpus through content analysis, retrieve relevant content, and generate query answers using a pre-trained large-scale visual language model.
The accuracy and effectiveness of query results are improved, and can more comprehensively and accurately meet users' query needs in complex multimodal scenarios.
Smart Images

Figure CN120705367A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of artificial intelligence technology, and in particular to an intelligent query method, apparatus, computer equipment, and computer-readable storage medium. Background Art
[0002] With the rapid development of information technology, users' needs for information retrieval are becoming increasingly diverse and complex. Traditional single-modal information query systems (such as text-only search engines) are no longer able to meet users' needs for multimodal information (including text, images, and videos). In recent years, multimodal information query technology has gradually emerged, aiming to provide richer and more accurate query results by integrating information from multiple modalities. However, existing multimodal query technologies still have many shortcomings in terms of the accuracy and effectiveness of query results.
[0003] Multimodal information querying has significant application value in healthcare. For example, during diagnosis, doctors need to comprehensively consider a patient's medical records, medical images (such as X-rays, CT scans, and MRIs), and video consultation records. However, existing multimodal query systems still need to improve the accuracy and effectiveness of query results when processing medical data.
[0004] In the fintech sector, multimodal information querying faces similar challenges. For example, financial institutions need to comprehensively analyze multimodal data such as customer financial records, transaction data, video interviews, and image recognition information (such as ID images) during risk assessment, customer identity verification, and investment decision-making. However, existing multimodal query systems still need to improve the accuracy and effectiveness of query results when processing fintech data.
[0005] Based on this, how to provide an intelligent query method, device, computer equipment and computer-readable storage medium that can effectively improve the accuracy and effectiveness of query results when users perform question queries is an urgent problem to be solved by technical personnel in this field. Summary of the Invention
[0006] In view of the above-mentioned deficiencies in the prior art, the purpose of the present invention is to provide an intelligent query method, apparatus, computer equipment and computer-readable storage medium, aiming to solve the problem of how to effectively improve the accuracy and effectiveness of query results when users perform question queries.
[0007] In order to achieve the above object, the present invention adopts the following technical solutions:
[0008] In a first aspect, the present invention provides an intelligent query method, which includes:
[0009] Construct a multimodal corpus consisting of text corpus, image corpus and video corpus;
[0010] Acquiring a query question from a target user, performing content analysis on the query question, and identifying a target modal corpus adapted to the query question based on the multimodal corpus;
[0011] Retrieving content related to the query question based on the target modal corpus and generating candidate content;
[0012] The query question and the candidate content are integrated to obtain query content, and the query content is used as input to generate a target answer to the query question through a pre-trained large-scale visual language model.
[0013] In a second aspect, the present invention provides an intelligent query device, wherein the device comprises:
[0014] A construction module for constructing a multimodal corpus consisting of a text corpus, an image corpus, and a video corpus;
[0015] an acquisition module, configured to acquire a query question of a target user, perform content analysis on the query question, and identify a target modal corpus adapted to the query question based on the multimodal corpus;
[0016] A retrieval module, configured to retrieve content related to the query question based on the target modal corpus and generate candidate content;
[0017] An answer generation module is used to integrate the query question with the candidate content to obtain query content, use the query content as input, and generate a target answer to the query question through a pre-trained large-scale visual language model.
[0018] In a third aspect, the present invention provides a computer device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the intelligent query method as described above when executing the computer program.
[0019] In a fourth aspect, the present invention provides a computer-readable storage medium storing a computer program, wherein the computer program implements the intelligent query method as described above when executed by a processor.
[0020] Compared with the prior art, the present invention provides an intelligent query method, apparatus, computer device and computer-readable storage medium, wherein a multimodal corpus comprising a text corpus, an image corpus and a video corpus is constructed; a query question of a target user is obtained, content analysis is performed on the query question, and a target modal corpus adapted to the query question is identified based on the multimodal corpus; content related to the query question is retrieved based on the target modal corpus to generate candidate content; the query question and the candidate content are content-integrated to obtain query content, the query content is used as input, and a target answer to the query question is generated through a pre-trained large-scale visual language model; thereby, the present invention can effectively improve the accuracy and effectiveness of query results (i.e., target answers) when users perform question queries. BRIEF DESCRIPTION OF THE DRAWINGS
[0021] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are only some embodiments recorded in the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0022] Figure 1 A schematic diagram of an application environment of an intelligent query method provided by an embodiment of the present invention.
[0023] Figure 2 A flowchart of an intelligent query method provided by one embodiment of the present invention.
[0024] Figure 3 A schematic diagram of program modules of an intelligent query device provided by one embodiment of the present invention.
[0025] Figure 4 A schematic diagram of the structure of a computer device provided in one embodiment of the present invention.
[0026] Figure 5 Another structural diagram of a computer device provided by an embodiment of the present invention. DETAILED DESCRIPTION
[0027] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of them. All other embodiments obtained by ordinary technicians in this field based on the embodiments of the present invention without making any creative efforts shall fall within the scope of protection of the present invention.
[0028] It should be understood that when used in the present specification and the appended claims, the term "comprising" indicates the presence of described features, integers, steps, operations, elements and / or components, but does not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components and / or groups thereof.
[0029] It will also be understood that the term "and / or" used in the present description and appended claims refers to and includes any and all possible combinations of one or more of the associated listed items.
[0030] As used in the present specification and the appended claims, the term "if" may be interpreted as "when" or "upon" or "in response to determining" or "in response to detecting," depending on the context. Similarly, the phrase "if it is determined" or "if [described condition or event] is detected" may be interpreted as meaning "upon determination" or "in response to determining" or "upon detection of [described condition or event]" or "in response to detecting [described condition or event]," depending on the context.
[0031] In addition, in the description of the present specification and the appended claims, the terms "first", "second", "third", etc. are only used to distinguish the descriptions and cannot be understood as indicating or implying relative importance.
[0032] References to "one embodiment" or "some embodiments" in the present specification mean that a particular feature, structure, or characteristic described in conjunction with that embodiment is included in one or more embodiments of the present invention. Thus, phrases such as "in one embodiment," "in some embodiments," "in other embodiments," and "in yet other embodiments" appearing in various places in this specification do not necessarily refer to the same embodiment, but rather mean "one or more, but not all, embodiments," unless otherwise specifically emphasized. The terms "including," "comprising," "having," and variations thereof mean "including but not limited to," unless otherwise specifically emphasized.
[0033] It should be understood that the order of execution of the steps in the following embodiments does not necessarily mean the order in which they are executed. The order in which each process is executed should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present invention.
[0034] In order to illustrate the technical solution of the present invention, specific embodiments are provided below.
[0035] An intelligent query method provided by an embodiment of the present invention can be applied in Figure 1In the application environment shown, the client and server communicate via a network. The client includes, but is not limited to, PDAs, desktop computers, laptops, ultra-mobile personal computers (UMPCs), netbooks, cloud computing devices, personal digital assistants (PDAs), and other computer devices. The server can be a standalone server or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, content delivery networks (CDNs), and big data and artificial intelligence platforms.
[0036] See also Figure 2 An embodiment of the present invention provides an intelligent query method, wherein the method comprises the following steps:
[0037] S100, constructing a multimodal corpus including a text corpus, an image corpus, and a video corpus;
[0038] S200: Obtain a query question of a target user, perform content analysis on the query question, and identify a target modal corpus adapted to the query question based on the multimodal corpus;
[0039] S300, based on the target modal corpus, searching for content related to the query question and generating candidate content;
[0040] S400: Integrate the query question with the candidate content to obtain query content, use the query content as input, and generate a target answer to the query question through a pre-trained large-scale visual language model.
[0041] In specific implementation, the intelligent query method of this embodiment effectively improves the accuracy and effectiveness of the query results (i.e., the target answer) when the user queries the question through a series of innovative steps, which is specifically reflected in the following aspects:
[0042] 1. Construction of a multimodal corpus
[0043] By building a multimodal corpus encompassing multiple modalities, including text, image, and video corpora, this approach is able to cover a wider range of information types. This multimodal corpus design enables the system to understand user questions from multiple perspectives and retrieve relevant information from this rich data, providing a solid data foundation for generating more comprehensive and accurate answers.
[0044] 2. Accurate modality adaptation and retrieval
[0045] After receiving a user's query, the system uses content analysis to identify a target modal corpus that matches the query. This modal adaptation mechanism accurately selects the most appropriate modal corpus (text, image, or video) for retrieval based on the semantic characteristics and information requirements of the question. For example, for questions involving visual content, the system prioritizes image or video corpora, thus avoiding the modal bias common in traditional single-modal retrieval systems and improving the targetedness and accuracy of retrieval.
[0046] 3. Multimodal content integration
[0047] After generating candidate content, the system integrates the query with the candidate content. Through semantic alignment and multimodal fusion strategies, the system can organically combine information from different modalities to generate unified query content. This fusion strategy not only considers textual information but also fully utilizes visual information such as images and videos, thereby generating richer and more accurate answers. For example, in a medical diagnosis scenario, the system can combine medical records with medical images to generate more accurate diagnostic recommendations.
[0048] 4. Application of Pre-trained Large Visual Language Models
[0049] This approach leverages advanced deep learning techniques to process complex multimodal information by generating target answers using a pre-trained large-scale visual-language model. Pre-trained models typically possess strong semantic understanding and generation capabilities, enabling them to generate high-quality, user-specific answers based on the aggregated query content. The application of this model further enhances the accuracy and logicality of answers, ensuring that the results effectively address the user's query.
[0050] In this embodiment, the intelligent query method achieves comprehensive optimization from the data level to the generation level through the construction of a multimodal corpus, precise modality adaptation, multimodal fusion content integration, and the application of pre-trained models. These innovative designs not only improve the accuracy and effectiveness of query results, but also enhance the adaptability and flexibility of the system, effectively meeting user query needs in complex multimodal scenarios.
[0051] It is understandable that the intelligent query method provided by the embodiment of the present invention can be applied to intelligent query scenarios related to the medical and health field. The following is a specific example:
[0052] Application examples in the medical and health field (medical knowledge retrieval)
[0053] Scenario description: Medical staff need to quickly retrieve relevant medical literature, case studies, and imaging data during medical research or clinical practice.
[0054] Application:
[0055] Multimodal corpus: Construct a multimodal corpus that includes medical literature (text corpus), typical case images (image corpus) and teaching videos (video corpus).
[0056] Query question: Medical staff enter a query question, such as "What are the typical imaging features of a certain disease?"
[0057] Modality adaptation: The system identifies the target modality corpus adapted to the query question and selects the image corpus for retrieval.
[0058] Content integration: The retrieved typical case images are integrated with the relevant descriptions in the query question to generate query content.
[0059] Target answer generation: Generate answers through a pre-trained visual language model, such as "Typical imaging features of this disease include..., and related literature and case images are as follows..."
[0060] Technical effect: The system can quickly provide multimodal medical knowledge, helping medical staff to obtain information more efficiently and improve work efficiency.
[0061] It is understandable that the intelligent query method provided by the embodiment of the present invention can also be applied to intelligent query scenarios related to the financial technology field. The following is a specific example:
[0062] Application examples in the field of FinTech (intelligent customer service and customer consultation)
[0063] Scenario description: Financial institutions' intelligent customer service needs to handle diverse customer inquiries, including text inquiries, video interviews, and image recognition.
[0064] Application:
[0065] Multimodal corpus: Build a multimodal corpus containing customer consultation texts (text corpus), video interview records (video corpus), and identity document images (image corpus).
[0066] Inquiry questions: Customers enter inquiry questions such as "What is my account balance?" and "How do I verify my identity?"
[0067] Modal adaptation: The system identifies the target modal corpus adapted to the query question and selects the text corpus and image corpus for retrieval.
[0068] Content integration: The customer's query question is integrated with the key information in the account information and ID card image to generate query content.
[0069] Target answer generation: Generate answers using a pre-trained visual language model, such as “Your account balance is…, and authentication can be performed in the following ways.”
[0070] Technical effect: The system can provide personalized customer service, improve customer satisfaction and the service efficiency of financial institutions.
[0071] It can be seen from the above specific application examples that the intelligent query method provided by the embodiments of the present invention can effectively integrate multimodal information, accurately adapt to user needs, and generate high-quality query results (i.e., target answers), significantly improving the accuracy and effectiveness of the query results.
[0072] Furthermore, in one embodiment, the intelligent query method, wherein the step of obtaining a query question of a target user, performing content analysis on the query question, and identifying a target modal corpus adapted to the query question based on the multimodal corpus, specifically comprises the following steps:
[0073] receiving a query question input by a target user through a user interface;
[0074] Performing content analysis on the query question to extract content features of the query question;
[0075] evaluating, by a modality adaptation algorithm, the degree of matching between the content features and different types of modal corpora in the multimodal corpus;
[0076] According to the evaluation result, a target modal corpus adapted to the query question is determined.
[0077] During specific implementation, the specific implementation process of the steps in this embodiment is roughly as follows:
[0078] 1. Receive user inquiries
[0079] Receive query questions input by the target user through the user interface.
[0080] Users can make query requests through text input, voice input or uploading images / videos.
[0081] The system performs preliminary processing on the received query questions, such as word segmentation and standardization of text, and preprocessing of images / videos (such as adjusting resolution, format conversion, etc.).
[0082] 2. Content Analysis and Feature Extraction
[0083] Conduct content analysis on query questions and extract their content features.
[0084] For text queries, natural language processing technology is used to extract keywords, phrases, and semantic vectors; for image or video queries, computer vision technology is used to extract visual features, such as CNN features of images or key frame features of videos.
[0085] The extracted features are expressed as vectors for subsequent matching evaluation.
[0086] 3. Modal Adaptation and Matching Evaluation
[0087] Through the modal adaptation algorithm, the matching degree between the extracted content features and different types of modal corpora in the multimodal corpus is evaluated.
[0088] Modality adaptation algorithms can dynamically select the most suitable modality corpus based on similarity calculations between feature vectors (such as cosine similarity) or semantic relevance assessments. For example, if the query question mainly involves visual content, the system will give priority to image or video corpora.
[0089] 4. Determine the target modal corpus
[0090] According to the evaluation results of the modality adaptation algorithm, the target modality corpus for query question adaptation is determined.
[0091] The system will select the most relevant modal corpus based on the matching score to ensure the accuracy and efficiency of the subsequent retrieval process.
[0092] If the query question involves multiple modal information, the system can dynamically combine multiple related modal corpora to provide more comprehensive retrieval results.
[0093] Through the above process, this embodiment can accurately identify the target modal corpus adapted to the user query question, providing a solid foundation for subsequent retrieval and answer generation.
[0094] Furthermore, in one embodiment, the intelligent query method, wherein the step of retrieving content related to the query question based on the target modal corpus and generating candidate content, specifically comprises the steps of:
[0095] Retrieving content related to the query question in the target modal corpus using a content retrieval algorithm;
[0096] Sort the retrieved content by relevance, and select the top N content results most relevant to the query question based on a preset relevance threshold; where N is a positive integer;
[0097] All the content results are integrated to generate candidate content.
[0098] During specific implementation, the specific implementation process of the steps in this embodiment is roughly as follows:
[0099] 1. Content retrieval
[0100] In the target modality corpus, content retrieval algorithms are used to retrieve content related to the query question.
[0101] Retrieval algorithms can match query features with the features of the corpus based on similarity. For example, for text corpora, retrieval algorithms based on pre-trained language models such as the Vector Space Model (VSM) or BERT can be used; for image and video corpora, deep learning-based feature matching algorithms (such as CNN feature matching) can be used.
[0102] 2. Relevance Sorting
[0103] Sort the retrieved content by relevance.
[0104] Relevance can be determined by calculating the similarity between the query feature vector and the retrieved content feature vector. For example, similarity can be calculated using methods such as cosine similarity or Euclidean distance. Based on a preset relevance threshold, the top N content results most relevant to the query are selected, where N is a positive integer.
[0105] The preset correlation threshold can be adjusted according to specific application scenarios and user needs.
[0106] 3. Content screening
[0107] According to the preset relevance threshold, the top N most relevant content results are screened out from the sorted search results.
[0108] The filtering process can be based on similarity scores, ensuring that only results that are highly relevant to the query are retained. For example, if the similarity score is below a preset threshold, the result will be excluded.
[0109] 4. Content Integration
[0110] The first N filtered content results are integrated to generate candidate content.
[0111] The integration process can be optimized based on the content modality. For example, for text content, key sentences or paragraphs can be extracted; for image and video content, key frames can be extracted or descriptive text can be generated.
[0112] The integrated candidate content should be able to fully reflect the multimodal information related to the query question.
[0113] 5. Candidate content generation
[0114] The resulting candidate content will serve as input to subsequent steps for further content fusion and answer generation. The candidate content should contain enough information so that the pre-trained large-scale visual language model can generate accurate and comprehensive target answers.
[0115] Through the above process, this embodiment can retrieve content related to the query question from the target modal corpus, and generate high-quality candidate content through relevance sorting and screening, providing support for subsequent multimodal fusion and answer generation.
[0116] Furthermore, in one embodiment, the intelligent query method, wherein the query question and the candidate content are integrated to obtain query content, the query content is used as input, and a target answer to the query question is generated by a pre-trained large-scale visual language model, specifically comprises the following steps:
[0117] Performing semantic alignment processing on the query question and the candidate content, and fusing the semantically aligned query question and the candidate content to generate query content;
[0118] Formatting the query content, and inputting the formatted query content into a pre-trained large-scale visual language model to generate a target answer to the query question;
[0119] The target answer is displayed.
[0120] Furthermore, the intelligent query method, wherein the semantic alignment of the query question and the candidate content is performed, and the semantically aligned query question and the candidate content are fused to generate query content, specifically comprises the steps of:
[0121] Extracting semantic features from the query question and the candidate content to obtain semantic vector representations of the query question and the candidate content;
[0122] Calculating the similarity between the query question and the semantic vector representation of the candidate content, and adjusting the semantic vector representation of the candidate content according to the calculation result so that the candidate content is semantically aligned with the query question;
[0123] The query question and the candidate content after semantic alignment are fused using a multimodal fusion strategy to generate query content.
[0124] Furthermore, the intelligent query method, wherein the multimodal fusion strategy is used to fuse the semantically aligned query question and the candidate content to generate query content, specifically includes the following steps:
[0125] Performing multimodal feature extraction on the semantically aligned query question and the candidate content to obtain multimodal features of the query question and the candidate content;
[0126] Using the multimodal fusion strategy, the query question and the multimodal features of the candidate content are fused to obtain fused features;
[0127] Based on the fused features, the query content is generated through a pre-trained multimodal generation model.
[0128] Furthermore, the intelligent query method, wherein the displaying of the target answer specifically comprises the steps of:
[0129] Before generating the target answer, pre-acquire the answer display strategy input by the target user through the client;
[0130] The target answer is sent to the client for display according to the answer display strategy.
[0131] During specific implementation, the specific implementation process of the steps in this embodiment is roughly as follows:
[0132] 1. Semantic alignment processing
[0133] Semantic feature extraction: Semantic features are extracted from the query and candidate content to obtain their semantic vector representations. For text content, pre-trained language models (such as BERT) are used to extract semantic vectors; for image or video content, pre-trained visual models (such as ResNet) are used to extract visual feature vectors.
[0134] Similarity calculation and adjustment: Calculate the similarity (e.g., cosine similarity) between the semantic vector representation of the query and the semantic vector representation of the candidate content. Based on the similarity calculation results, adjust the semantic vector representation of the candidate content to align it with the semantics of the query. For example, through weighted adjustment or optimization algorithms, the semantic vector of the candidate content can be brought closer to the semantic vector of the query.
[0135] 2. Content Integration
[0136] Multimodal feature extraction: Multimodal feature extraction is performed on the semantically aligned query and candidate content to obtain their multimodal features, including text features, image features, and video features.
[0137] Feature fusion: Utilize multimodal fusion strategies to fuse the multimodal features of the query and candidate content. Fusion strategies can include feature concatenation, weighted summation, or attention mechanisms to generate unified fused features.
[0138] Generate query content: Based on the fused features, query content is generated using a pre-trained multimodal generative model. The generative model can be a multimodal model based on the Transformer architecture, which can process multimodal input and generate high-quality text output.
[0139] 3. Formatting and target answer generation
[0140] Formatting: The generated query content is formatted to ensure that it meets the input requirements of the pre-trained large-scale visual language model. Formatting may include text encoding, sequence length adjustment, etc.
[0141] Target answer generation: The formatted query content is input into a pre-trained large-scale visual language model to generate the target answer to the query question.
[0142] 4. Target answer display
[0143] Obtaining the answer display strategy: Before generating the target answer, obtain the answer display strategy entered by the target user through the client. The answer display strategy can include the answer presentation format (such as text, chart, image, etc.), sorting method (such as relevance, chronological order, etc.), and interaction method (such as expandable details, links, etc.).
[0144] Answer display: Based on the user-specified answer display strategy, the target answer is sent to the client for display. The system can flexibly adjust the target answer display method based on user preferences and needs to enhance the user experience.
[0145] Through the above process, this embodiment effectively semantically aligns the query and candidate content, integrating their content to generate high-quality query content. Furthermore, a pre-trained large-scale visual language model is used to generate accurate target answers. Ultimately, based on the user's answer display strategy, the system displays the answers in the desired format, ensuring the accuracy and effectiveness of the query results.
[0146] As can be seen from the above method embodiments, the intelligent query method provided by the present invention includes: constructing a multimodal corpus including a text corpus, an image corpus, and a video corpus; obtaining the query question of the target user, performing content analysis on the query question, and identifying a target modal corpus adapted to the query question based on the multimodal corpus; based on the target modal corpus, retrieving content related to the query question and generating candidate content; integrating the query question with the candidate content to obtain query content, using the query content as input, and generating a target answer to the query question through a pre-trained large-scale visual language model. In this way, the method of the present invention can effectively improve the accuracy and effectiveness of the query results (i.e., the target answer) when the user performs a question query.
[0147] It should be understood that although the present application provides method operation steps as described in the embodiments or flowcharts, more or fewer operation steps may be included based on conventional or non-creative work, and these operation steps are not necessarily performed in the order of the embodiments or flowcharts. The order of steps listed in the embodiments or flowcharts is only one way of executing the steps among many steps and does not represent the only execution order. It should be noted that there is not necessarily a certain order between the above steps. Those of ordinary skill in the art can understand from the description of the embodiments of the present invention that in different embodiments, the above steps may have different execution orders, that is, they may be executed in parallel, or they may be executed in an interchangeable manner, etc. Moreover, at least a portion of the steps in the embodiments or flowcharts may include multiple sub-steps or multiple stages, and these sub-steps or stages are not necessarily executed at the same time, but may be executed at different times. The execution order of these sub-steps or stages is not necessarily to be performed in sequence, but may be executed in turn, alternately or synchronously with other steps or at least a portion of the sub-steps or stages of other steps.
[0148] Based on the above method embodiment, please refer to Figure 3 Another embodiment of the present invention further provides an intelligent query device, wherein the device includes:
[0149] A construction module 11 is used to construct a multimodal corpus including a text corpus, an image corpus and a video corpus;
[0150] An acquisition module 12 is configured to acquire a query question of a target user, perform content analysis on the query question, and identify a target modal corpus adapted to the query question based on the multimodal corpus;
[0151] A retrieval module 13 is configured to retrieve content related to the query question based on the target modal corpus and generate candidate content;
[0152] The answer generation module 14 is used to integrate the query question with the candidate content to obtain query content, use the query content as input, and generate a target answer to the query question through a pre-trained large-scale visual language model.
[0153] Furthermore, in one embodiment, the intelligent query device, wherein the acquiring of the query question of the target user, performing content analysis on the query question, and identifying a target modal corpus adapted to the query question based on the multimodal corpus, specifically includes:
[0154] receiving a query question input by a target user through a user interface;
[0155] Performing content analysis on the query question to extract content features of the query question;
[0156] evaluating, by a modality adaptation algorithm, the degree of matching between the content features and different types of modal corpora in the multimodal corpus;
[0157] According to the evaluation result, a target modal corpus adapted to the query question is determined.
[0158] Furthermore, in one embodiment, the intelligent query device, wherein the step of retrieving content related to the query question based on the target modal corpus and generating candidate content specifically includes:
[0159] Retrieving content related to the query question in the target modal corpus using a content retrieval algorithm;
[0160] Sort the retrieved content by relevance, and select the top N content results most relevant to the query question based on a preset relevance threshold; where N is a positive integer;
[0161] All the content results are integrated to generate candidate content.
[0162] Furthermore, in one embodiment, the intelligent query device, wherein the step of integrating the query question with the candidate content to obtain query content, taking the query content as input, and generating a target answer to the query question through a pre-trained large-scale visual language model, specifically includes:
[0163] Performing semantic alignment processing on the query question and the candidate content, and fusing the semantically aligned query question and the candidate content to generate query content;
[0164] Formatting the query content, and inputting the formatted query content into a pre-trained large-scale visual language model to generate a target answer to the query question;
[0165] The target answer is displayed.
[0166] Furthermore, the intelligent query device, wherein the semantic alignment of the query question and the candidate content is performed, and the semantically aligned query question and the candidate content are fused to generate query content, specifically includes:
[0167] Extracting semantic features from the query question and the candidate content to obtain semantic vector representations of the query question and the candidate content;
[0168] Calculating the similarity between the query question and the semantic vector representation of the candidate content, and adjusting the semantic vector representation of the candidate content according to the calculation result so that the candidate content is semantically aligned with the query question;
[0169] The query question and the candidate content after semantic alignment are fused using a multimodal fusion strategy to generate query content.
[0170] Furthermore, the intelligent query device, wherein the multimodal fusion strategy is used to fuse the semantically aligned query question and the candidate content to generate query content, specifically including:
[0171] Performing multimodal feature extraction on the semantically aligned query question and the candidate content to obtain multimodal features of the query question and the candidate content;
[0172] Using the multimodal fusion strategy, the query question and the multimodal features of the candidate content are fused to obtain fused features;
[0173] Based on the fused features, the query content is generated through a pre-trained multimodal generation model.
[0174] Furthermore, in the intelligent query device, the step of displaying the target answer specifically includes:
[0175] Before generating the target answer, pre-acquire the answer display strategy input by the target user through the client;
[0176] The target answer is sent to the client for display according to the answer display strategy.
[0177] It should be noted that, in the embodiment of the device of the present invention, the information interaction, execution process and other contents between the above modules are based on the same concept as the embodiment of the method of the present invention. Their specific functions and technical effects can be found in the aforementioned method embodiment part and will not be repeated here.
[0178] Based on the above method embodiment, another embodiment of the present invention further provides a computer device, which can be a server, and its internal structure diagram can be as follows: Figure 4As shown. The computer device includes a processor, a memory, a network interface and a database connected via a system bus. The processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program and a database. The internal memory provides an environment for the operation of the operating system and computer program in the non-volatile storage medium. The network interface of the computer device is used to communicate with an external terminal via a network connection. When the computer program is executed by the processor, the functions or steps on the server side of the intelligent query method in any of the above method embodiments are implemented.
[0179] Based on the above method embodiment, another embodiment of the present invention further provides a computer device, which can be a client, and its internal structure diagram can be as follows: Figure 5 As shown. The computer device includes a processor, memory, network interface, display screen and input device connected via a system bus. The processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and computer program in the non-volatile storage medium. The network interface of the computer device is used to communicate with an external terminal via a network connection. When the computer program is executed by the processor, the functions or steps of the client side of the intelligent query method in any of the above method embodiments are implemented.
[0180] Those skilled in the art will understand that Figure 4 and Figure 5 The structural diagram shown in the figure is only a schematic diagram of a part of the structure related to the solution of the present invention, and does not constitute a limitation on the computer device to which the solution of the present invention is applied. The specific computer device may include more components than shown in the figure, or combine certain components, or have a different component arrangement.
[0181] The processor referred to herein may be a CPU, other general-purpose processors, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. A general-purpose processor may be a microprocessor, or any conventional processor, etc.
[0182] The memory includes a readable storage medium, an internal memory, etc., wherein the internal memory can be the memory of a computer device, and the internal memory provides an environment for the operation of the operating system and computer-readable instructions in the readable storage medium. The readable storage medium can be a hard disk of the computer device, and in other embodiments, it can also be an external storage device of the computer device, for example, a plug-in hard disk, a smart memory card (Smart Media Card, SMC), a secure digital (Secure Digital, SD) card, a flash card (Flash Card), etc. equipped on the computer device. Furthermore, the memory can also include both an internal storage unit of the computer device and an external storage device. The memory is used to store an operating system, an application program, a boot loader (BootLoader), data, and other programs, such as the program code of a computer program. The memory can also be used to temporarily store data that has been output or is about to be output.
[0183] Based on the above method embodiments, another embodiment of the present invention further provides a computer-readable storage medium storing a computer program. When executed by a processor, the computer program implements the intelligent query method described in any of the above method embodiments. The computer-readable storage medium may be non-volatile or volatile.
[0184] It should be noted that the above-mentioned functions or steps that can be implemented by computer-readable storage media or computer devices, and the technical effects brought about by the functions / steps, can be found in the relevant descriptions in the aforementioned method embodiments. To avoid repetition, they will not be described one by one here.
[0185] Those skilled in the art will appreciate that all or part of the processes in the above-mentioned embodiments can be implemented by instructing the relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to memory, storage, database or other media used in the embodiments provided in this application can include non-volatile and / or volatile memory. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM) or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link (Synchlink) DRAM (SLDRAM), memory bus (Rambus) direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM). The disclosed memory components or memories of the operating environments described herein are intended to comprise one or more of these and / or any other suitable types of memory.
[0186] Those skilled in the art can clearly understand that, for the convenience and simplicity of description, in the embodiment of the device of the present invention, only the division of the above-mentioned functional units and modules is used as an example for illustration. In actual application, the above-mentioned functions can be distributed and completed by different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above. The functional units and modules in the embodiment can be integrated into a processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit. The above-mentioned integrated unit can be implemented in the form of hardware or in the form of software functional units. In addition, the specific names of the functional units and modules are only for the convenience of distinguishing each other and are not used to limit the scope of protection of the present invention. The specific working process of the units and modules in the above-mentioned device can refer to the corresponding process in the above-mentioned method embodiment, which will not be repeated here. If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium.
[0187] Those skilled in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professionals and technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of the present invention.
[0188] In the embodiments provided by the present invention, it should be understood that the disclosed apparatus / computer equipment and methods can be implemented in other ways. For example, the apparatus / computer equipment embodiments described above are merely illustrative. For example, the division of modules or units is merely a logical function division. In actual implementation, there may be other division methods, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or units, which can be electrical, mechanical or other forms.
[0189] Units described as separate components may or may not be physically separate, and components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of these units may be selected to achieve the purpose of this embodiment according to actual needs.
[0190] It should be noted that if software tools or components other than those of the Company appear in the embodiments of this application, they are merely for illustration and do not represent actual use. The above embodiments are intended only to illustrate the technical solutions of the present invention, not to limit them. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or replace some of the technical features therein with equivalents. These modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the various embodiments of the present invention, and should all be included in the scope of protection of the present invention.
Claims
1. An intelligent query method, characterized in that: include: Construct a multimodal corpus consisting of text corpus, image corpus and video corpus; Acquiring a query question from a target user, performing content analysis on the query question, and identifying a target modal corpus adapted to the query question based on the multimodal corpus; Retrieving content related to the query question based on the target modal corpus and generating candidate content; The query question and the candidate content are integrated to obtain query content, and the query content is used as input to generate a target answer to the query question through a pre-trained large-scale visual language model.
2. The intelligent query method according to claim 1, characterized in that: The step of obtaining a query question of a target user, performing content analysis on the query question, and identifying a target modal corpus adapted to the query question based on the multimodal corpus includes: receiving a query question input by a target user through a user interface; Performing content analysis on the query question to extract content features of the query question; evaluating, by a modality adaptation algorithm, the degree of matching between the content features and different types of modal corpora in the multimodal corpus; According to the evaluation result, a target modal corpus adapted to the query question is determined.
3. The intelligent query method according to claim 1, characterized in that: The step of retrieving content related to the query question based on the target modal corpus and generating candidate content includes: Retrieving content related to the query question in the target modal corpus using a content retrieval algorithm; Sort the retrieved content by relevance, and select the top N content results most relevant to the query question based on a preset relevance threshold; where N is a positive integer; All the content results are integrated to generate candidate content.
4. The intelligent query method according to claim 1, characterized in that: The step of integrating the query question with the candidate content to obtain query content, taking the query content as input, and generating a target answer to the query question through a pre-trained large-scale visual language model includes: Performing semantic alignment processing on the query question and the candidate content, and fusing the semantically aligned query question and the candidate content to generate query content; Formatting the query content, and inputting the formatted query content into a pre-trained large-scale visual language model to generate a target answer to the query question; The target answer is displayed.
5. The intelligent query method according to claim 4, characterized in that: The semantically aligning the query question with the candidate content, and fusing the semantically aligned query question and the candidate content to generate query content, includes: Extracting semantic features from the query question and the candidate content to obtain semantic vector representations of the query question and the candidate content; Calculating the similarity between the query question and the semantic vector representation of the candidate content, and adjusting the semantic vector representation of the candidate content according to the calculation result so that the candidate content is semantically aligned with the query question; The query question and the candidate content after semantic alignment are fused using a multimodal fusion strategy to generate query content.
6. The intelligent query method according to claim 5, characterized in that: The method of utilizing a multimodal fusion strategy to fuse the semantically aligned query question and the candidate content to generate query content includes: Performing multimodal feature extraction on the semantically aligned query question and the candidate content to obtain multimodal features of the query question and the candidate content; Using the multimodal fusion strategy, the query question and the multimodal features of the candidate content are fused to obtain fused features; Based on the fused features, the query content is generated through a pre-trained multimodal generation model.
7. The intelligent query method according to claim 4, characterized in that: The displaying of the target answer includes: Before generating the target answer, pre-acquire the answer display strategy input by the target user through the client; The target answer is sent to the client for display according to the answer display strategy.
8. An intelligent query device, characterized in that: The device comprises: A construction module for constructing a multimodal corpus consisting of a text corpus, an image corpus, and a video corpus; an acquisition module, configured to acquire a query question of a target user, perform content analysis on the query question, and identify a target modal corpus adapted to the query question based on the multimodal corpus; A retrieval module, configured to retrieve content related to the query question based on the target modal corpus and generate candidate content; An answer generation module is used to integrate the query question with the candidate content to obtain query content, use the query content as input, and generate a target answer to the query question through a pre-trained large-scale visual language model.
9. A computer device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the computer program, the intelligent query method according to any one of claims 1 to 7 is implemented.
10. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the intelligent query method according to any one of claims 1 to 7 is implemented.