Question and answer method, apparatus, device, medium, and program product
By constructing a multimodal index and selecting an appropriate indexing strategy, the problems of structural destruction and resource consumption in enterprise heterogeneous data question and answering were solved, achieving efficient and accurate question and answer responses.
Patent Information
- Application Number
- CN202511328795.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-17
- Publication Date
- 2025-12-12
- Estimated Expiration
- 2045-09-17
AI Technical Summary
Existing technologies struggle to meet the needs for accurate, comprehensive, and context-sensitive question answering when processing heterogeneous enterprise data. Furthermore, the retrieval process can easily disrupt the structural relationships of the data, increase computational resource consumption, and degrade the user experience.
By constructing multimodal indexes, the structural integrity and semantic coherence of multimodal data are preserved. Appropriate indexes are selected based on intent and available computing resources, and accurate responses are generated using large models.
It improves the accuracy of question answering and user experience, while balancing computational efficiency and resource consumption, and generates more accurate responses.
Smart Images

Figure CN120821815B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of artificial intelligence, and more particularly to a question and answer method, device, equipment, medium and program product. BACKGROUND
[0002] With the development of artificial intelligence technology, the natural language understanding and generation capabilities of artificial intelligence models can be used in combination with retrieval enhancement generation technology to process question and answer tasks such as dialogue, abstract generation and knowledge retrieval. Due to the widespread application of digital technology, enterprises generate and accumulate massive amounts of heterogeneous data in the process of operation. When processing heterogeneous data in enterprises based on artificial intelligence models and retrieval enhancement generation technology, it is difficult to meet the needs of enterprises for accurate, comprehensive and context-related question and answer. SUMMARY
[0003] In view of the above problems, the present application provides a question and answer method, device, equipment, medium and program product capable of improving the retrieval accuracy and question and answer quality for heterogeneous data.
[0004] According to a first aspect of the present application, a question and answer method is provided, comprising: in response to receiving a question from a target object, determining at least one target index matching at least one of an intent represented by the question and available computing resources for performing retrieval from a plurality of indexes according to the intent and the available computing resources, wherein the plurality of indexes are obtained by respectively processing multi-modal information with a plurality of index accuracies, and the multi-modal information includes structured representations of a plurality of modal data obtained by processing target files; retrieving target multi-modal information based on the intent through the at least one target index, wherein the target multi-modal information includes a plurality of candidate modal data and respective metadata, and the metadata indicates description information indicating that the corresponding modal data matches the intent; processing the target multi-modal information based on the question using a preset model, and outputting a reply to the intent.
[0005] A second aspect of the present application provides a question and answer device, comprising: an index module configured to determine at least one target index matching at least one of an intent represented by a question and available computing resources for performing retrieval from a plurality of indexes in response to receiving the question from a target object, wherein the plurality of indexes are obtained by respectively processing multi-modal information with a plurality of index accuracies, and the multi-modal information includes structured representations of a plurality of modal data obtained by processing target files; a retrieval module configured to retrieve target multi-modal information based on the intent through the at least one target index, wherein the target multi-modal information includes a plurality of candidate modal data and respective metadata, and the metadata indicates description information indicating that the corresponding modal data matches the intent; and a generation module configured to process the target multi-modal information based on the question using a preset model, and output a reply to the intent.
[0006] The third aspect of the present application provides an electronic device, comprising: one or more processors; a memory for storing one or more computer programs, wherein the one or more processors execute the one or more computer programs to implement the steps of the method.
[0007] The fourth aspect of the present application further provides a computer-readable storage medium having stored thereon a computer program or instructions, which, when executed by a processor, implement the steps of the method.
[0008] The fifth aspect of the present application further provides a computer program product comprising a computer program or instructions, which, when executed by a processor, implement the steps of the method.
[0009] Through the embodiments of the present application, the target file can be processed to obtain the structured representation of the plurality of modal data to form the multi-modal information, the structural integrity and semantic coherence of the multi-modal data are effectively preserved, the suitable target index is selected according to at least one of the intention and the available computing resource, and the flexible multi-index matching is considered, so that the accurate data meeting the intention can be obtained in the retrieval process considering the constraint of the available computing resource, the calculation efficiency, the calculation resource consumption and the retrieval accuracy are taken into account, the reply generated by the large model is more accurate, and the question and answer quality and user experience are improved. BRIEF DESCRIPTION OF DRAWINGS
[0010] The above content and other purposes, features and advantages of the present application will be more apparent through the following description of the embodiments of the present application with reference to the accompanying drawings, in which:
[0011] Figure 1 An application scenario diagram suitable for implementing a question and answer method according to an embodiment of the present application is schematically shown;
[0012] Figure 2 A flowchart of a question and answer method according to an embodiment of the present application is schematically shown;
[0013] Figure 3 An acquisition of multi-modal information according to an embodiment of the present application is schematically shown;
[0014] Figure 4 An acquisition of a plurality of indexes according to an embodiment of the present application is schematically shown;
[0015] Figure 5 A flowchart of a determination of a target index according to an embodiment of the present application is schematically shown;
[0016] Figure 6 A question and answer method according to another embodiment of the present application is schematically shown;
[0017] Figure 7A structural block diagram of a question-answering device according to an embodiment of the present application is schematically shown.
[0018] Figure 8 A block diagram of an electronic device suitable for implementing a question-answering method according to an embodiment of the present application is schematically shown. DETAILED DESCRIPTION
[0019] Hereinafter, embodiments of the present application will be described with reference to the accompanying drawings. It should be understood, however, that the description which follows is merely exemplary and is not intended to limit the scope of the application. In the following detailed description of the embodiments of the present application, numerous specific details are set forth in order to provide a thorough understanding of the present application. However, it will be apparent to one skilled in the art that one or more embodiments of the present application can be practiced without these specific details. In other instances, well-known structures and functions have not been described in detail in order to avoid obscuring aspects of the present application.
[0020] The terms used herein are merely used to describe specific embodiments and are not intended to limit the present application. The terms "include", "comprise" and the like as used herein specify the presence of features, steps, operations and / or components but do not preclude the presence or addition of one or more other features, steps, operations or components.
[0021] All terms used herein, including technical and scientific terms, have the same meanings as those generally understood by those skilled in the art unless otherwise defined. It should be noted that the terms used herein should be interpreted as having meanings consistent with the context of the present description, and should not be interpreted in an idealized or excessively formal manner.
[0022] In the case of using expressions similar to "at least one of A, B and C, etc.", it should be generally interpreted as including one or more of the items enumerated in the list (e.g., "a system having at least one of A, B and C" should include, but not be limited to, a system having A alone, a system having B alone, a system having C alone, a system having both A and B, a system having both A and C, a system having both B and C, and / or a system having A, B, and C, etc.).
[0023] In the technical solutions of the present application, the user information (including but not limited to user personal information, user image information, user device information, such as location information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved are information and data authorized by the user or authorized by all parties, and the collection, storage, use, processing, transmission, provision, disclosure and application of related data comply with relevant laws, regulations and standards, necessary security measures are taken, do not violate public order and good customs, and provide corresponding operation portal for user to choose authorization or refusal.
[0024] In the scenario of making automated decisions by using personal information, the method, device and system provided by the embodiments of the present application all provide corresponding operation entrances for the user to select to agree or reject the automated decision result; if the user selects to reject, the expert decision process is entered. The expression "automated decision" here refers to the activity of making decisions by automatically analyzing and evaluating the personal behavior habits, interests and hobbies, or economic, health and credit conditions, etc. by a computer program. The expression "expert decision" here refers to the activity of making decisions by personnel who are engaged in a certain field of work, have special experience, knowledge and skills and have reached a certain professional level.
[0025] The retrieval-augmented generation (RAG) framework enhances the relevance and factuality of responses by enabling the model to access external knowledge during reasoning when combining retrieval mechanisms with artificial intelligence models. However, the general RAG framework destroys the structural relationship of data when processing structured and semi-structured data in enterprises, and it is difficult to meet the needs of enterprises for accurate, comprehensive and context-related question and answer. For example, heterogeneous data generated in the operation of an enterprise covers various types such as human resource records, structured reports and table documents. When processing multi-modal data, the RAG framework usually loses structured information and semantic information, such as blocking text documents according to fixed token lengths without considering the semantic coherence of the text, converting tables into continuous text streams, completely losing the row-column correspondence and position information, and turning data into unstructured text sequences. Moreover, a single index is constructed based on multi-modal data, and the retrieval accuracy and retrieval method based on a single index are fixed, which is difficult to meet the diversified question and answer needs, may increase the retrieval time and computing resource consumption, and reduce the user experience.
[0026] Therefore, in view of the problems of destroying semantic coherence and structured information, and possibly increasing retrieval time, computing resource consumption and reducing user experience, the embodiments of the present application provide a question and answer method, which can process target files to obtain structured representations of multiple modal data to form multi-modal information, effectively preserve the structural integrity and semantic coherence of multi-modal data, and flexibly match multiple indexes according to at least one of the intent and available computing resources to select a suitable target index, so that accurate data meeting the intent can be obtained in the retrieval process considering the constraints of available computing resources, taking into account computing efficiency, computing resource consumption and retrieval accuracy, making the generated reply of the large model more accurate and improving the user experience.
[0027] Figure 1 An application scenario suitable for implementing the question and answer method according to the embodiments of the present application is schematically shown.
[0028] As Figure 1As shown, the application scenario 100 according to this embodiment can include a first terminal device 101, a second terminal device 102, a third terminal device 103, a network 104, and a server 105. The network 104 is a medium for providing a communication link between the first terminal device 101, the second terminal device 102, the third terminal device 103, and the server 105. The network 104 can include various connection types, such as wired, wireless communication links, or fiber optic cables, and the like.
[0029] A user can use the first terminal device 101, the second terminal device 102, the third terminal device 103 to interact with the server 105 through the network 104 to send a question or receive a reply, and the like. Various communication client applications can be installed on the first terminal device 101, the second terminal device 102, the third terminal device 103, such as a question and answer application, a shopping application, a web browser application, a search application, an instant messaging tool, an email client, a social platform software, and the like (only as examples).
[0030] The first terminal device 101, the second terminal device 102, the third terminal device 103 can be various electronic devices with a display screen and supporting web browsing, including but not limited to a smartphone, a tablet computer, a laptop computer, a desktop computer, and the like.
[0031] The server 105 can be a server providing various services, such as a background management server providing support for a website browsed by a user using the first terminal device 101, the second terminal device 102, the third terminal device 103 (only as an example). The background management server can analyze and process received user requests, questions, and the like, such as identifying a question intent, performing a search, calling a preset model to generate a reply, and feeding back a processing result (such as a webpage, information, or data generated according to a user request) to a terminal device.
[0032] It should be noted that the question and answer method provided by the embodiments of the present application can generally be executed by the server 105. Correspondingly, the question and answer device provided by the embodiments of the present application can generally be arranged in the server 105. The question and answer method provided by the embodiments of the present application can also be executed by a server or a server cluster different from the server 105 and capable of communicating with the first terminal device 101, the second terminal device 102, the third terminal device 103, and / or the server 105. Correspondingly, the question and answer device provided by the embodiments of the present application can also be arranged in a server or a server cluster different from the server 105 and capable of communicating with the first terminal device 101, the second terminal device 102, the third terminal device 103, and / or the server 105.
[0033] It should be understood that, Figure 1The number of terminal devices, networks and servers in the above scenario is merely illustrative. Any number of terminal devices, networks and servers can be provided according to implementation needs.
[0034] The following will be based on Figure 1 the described scenario, by Figures 2-6 the question and answer method according to the embodiments of the application is described in detail.
[0035] Figure 2 The flowchart of the question and answer method according to the embodiments of the application is schematically shown.
[0036] As Figure 2 shown, the question and answer method of this embodiment includes operations S210-S230.
[0037] At operation S210, in response to receiving a question from a target object, at least one target index matching at least one of an intent of the question representation and available computing resources for performing retrieval is determined from a plurality of indexes according to at least one of the intent and the available computing resources, wherein the plurality of indexes are obtained by respectively processing multi-modal information with a plurality of index precisions, and the multi-modal information includes structured representations of a plurality of modal data obtained by processing a target file.
[0038] For example, the target object can include a user (such as an enterprise user, a system administrator or a terminal consumer, etc.) or a computer (such as an interface invoker), and the question can include user input content. For example, the question can be identified and the intent can be determined using a model based on a machine learning algorithm, such as an information query intent (such as "query sales data"), an operation instruction intent (such as "generate a monthly report"), an analysis intent (such as "compare the performance difference between product A and product B"), etc. The available computing resources include system resources that can be allocated when performing retrieval, such as the number of central processing cores, memory capacity, graphics processor video memory and network bandwidth, etc. The index is an ordered structure constructed based on the multi-modal information, and using the index can quickly access specific information in the multi-modal information.
[0039] The plurality of modal data includes a plurality of data format information, such as text, pictures, tables, audio and video in the target file, optical character recognition text extraction, table structure recognition, image feature extraction, speech transcription, etc. The structured representation refers to a representation form that maintains the original data structure relationship, such as the semantic relationship between texts, the relationship and pixel distribution in each region of the picture, the row and column structure in the table, the time sequence in the audio, the relationship between texts, the frame-to-frame picture in the video, the time sequence, and the relationship between texts, etc.
[0040] Exemplarily, the plurality of indexes can include data index structures of different index precision levels, such as a high-precision vector index, a medium-precision inverted index, and a low-precision hash index, and the like. Due to different index precisions, the intent requirements that can be met and the computing resources consumed are also different. For example, when system resources are sufficient, a high-precision vector index is selected for fine-grained retrieval; when system resources are tight, a medium-precision inverted index or a low-precision hash index is used. In addition, the low-precision hash index can be used for simple fact queries, and the high-precision vector index can be used for complex analysis queries.
[0041] In operation S220, target multi-modal information is retrieved based on the intent through at least one target index, the target multi-modal information including a plurality of candidate modal data and respective metadata, the metadata indicating description information of the corresponding modal data matching the intent.
[0042] Exemplarily, the intent refers to the purpose, target, or expected result of the target object, reflecting a specific requirement, which can be expressed according to the original content of the question, or can be rewritten and extended based on the original content. The candidate modal data includes multi-modal data segments retrieved in relation to the question and possibly displayed in the reply, such as relevant paragraphs in a document, specific rows and columns in a table, key regions in a picture, and key frames in a video, and the like.
[0043] In operation S230, the target multi-modal information is processed based on the question using a preset model, and a reply to the intent is output.
[0044] Exemplarily, the preset model can include a machine learning model, such as a large model, which can specifically refer to a deep learning model with a large number of model parameters. A large model generally contains hundreds of millions, billions, tens of billions, hundreds of billions, or even tens of trillions of model parameters. The large model can include a large-scale language model (LLM), a GPT (Generative Pre-trained Transformer), a visual large model, a multi-modal large model, and the like. The large model involved in the embodiments of the present application can be a general-purpose large model, or can also be a specialized large model obtained by fine tuning based on requirements, and the embodiments of the present application do not limit this.
[0045] According to the embodiments of the present application, the target file can be processed to obtain structured representations of multiple modal data to form multi-modal information, effectively preserving the structural integrity and semantic coherence of the multi-modal data, and flexibly selecting a suitable target index according to at least one of an intention and available computing resources, so that the accurate data meeting the intention can be obtained in the retrieval process considering the constraints of available computing resources, and the calculation efficiency, computing resource consumption and retrieval accuracy are taken into account, so that the reply generated by the large model is more accurate, and the user experience is improved.
[0046] Next, combined with Figure 3 Embodiments for obtaining multi-modal information are introduced.
[0047] Figure 3 An example of obtaining multi-modal information according to the embodiments of the present application is shown.
[0048] In some embodiments, the multi-modal information is obtained in the following way:
[0049] The text data in the target file is divided into multiple text blocks, and the structured representation of the multiple text blocks includes that a predetermined number of characters are overlapped between adjacent text blocks, and the predetermined number of characters are used to represent the semantic coherence between adjacent text blocks; the table data in the target file is split into multiple table blocks, and the structured representation of the multiple table blocks includes column structure information of table rows corresponding to each table block and table content based on column structure distribution. Metadata is added to the multiple text blocks and the multiple table blocks respectively to obtain multi-modal information.
[0050] For example, additional metadata such as document type, department and security level can be simulated for testing environment, and real enterprise metadata can be replaced in application of production environment. Rich metadata helps to improve the relevance of retrieval.
[0051] Exemplarily, different parsing methods can be determined according to the file type of the document, and the corresponding parsing method is used to parse the document, and the block length corresponding to the target file is set according to the actual application scenario, for example, in some scenarios that require long text context, a larger block size can be set; while in scenarios that require fast retrieval and processing, a smaller block size can be set.
[0052] In some embodiments, the predetermined number of characters overlapping between adjacent text blocks includes that the proportion of the predetermined number to the total number of characters of each text block is greater than or equal to one fourth. For example, the extracted text is segmented by a recursive character text segmenter, with a block size of 2000 characters and an overlap of 500 characters, which meets the input constraints of the large model while ensuring semantic coherence. In the case where the predetermined number is greater than or equal to one fourth, the relevant semantic content can be covered to some extent, so that the large model can identify the semantic connection between adjacent text blocks and retain meaningful context. For example, when processing complex documents such as policy manuals or technical manuals, complete and coherent information can be retrieved.
[0053] In some embodiments, the predetermined number of characters overlapping between adjacent text blocks can refer to the characters at the head and tail of any text block and the characters at the tail and head of the adjacent text block having the same predetermined number of characters. Referring to Figure 3 , the page includes text data 301 and table data 302, and the tail characters of text block 1 and the head characters of text block 2 are the same (filled with the same pattern).
[0054] In other embodiments, it is not limited to the head and tail overlapping, such as Figure 3 The text block 2 and the text block 3 are filled with the same pattern for a plurality of characters. The adjacent text blocks include a first text block and a second text block, and the predetermined number of characters overlapping between the adjacent text blocks is obtained by: obtaining semantic content represented by the first text block and the second text block; extracting a set of keywords representing the semantic content from the first text block and the second text block; and adding missing keywords to the first text block and the second text block based on the set of keywords.
[0055] For example, the first semantic content of the first text block and the second semantic content of the second text block can be extracted according to a machine learning model as the semantic content of the embodiment. Then, a first set of keywords representing the semantics of the first text block is extracted according to the first semantic content, and a second set of keywords representing the semantics of the second text block is extracted according to the second semantic content. Then, missing keywords are identified and selected from the other set of keywords to supplement the missing keywords. By cross-supplementing keywords, the semantic coherence between adjacent text blocks is enhanced, the semantic fragmentation problem caused by text block division is solved, and the semantic consistency of the overall text is improved.
[0056] For example, the first text block, the second text block, and the third text block are sequentially distributed, the cross-supplementing of keywords between the third text block and the second text block is based on semantics, and the cross-supplementing of keywords between the first text block and the third text block can not be performed, which can reduce the storage space occupation and improve the data processing speed.
[0057] In some embodiments, splitting the table data in the target file into the plurality of table blocks comprises: evaluating a structure complexity of the table data according to at least one of a table image, a cell style, a table content format and a table embedding manner of the table data; selecting a target splitting component matching the structure complexity from a plurality of preset splitting components, and splitting the table data into the plurality of table blocks, wherein the plurality of splitting components are used to provide matching splitting services for table data with different structure complexities.
[0058] Exemplarily, a data table with uniform format contains rectangular cells, and the cells record regular content, which can be considered as structured data; a complex table containing merged cells, irregular format and mixed data types can be considered as semi-structured data; a table embedded in a document can be regarded as part of unstructured data when processing. The structure complexity of structured data, semi-structured data and unstructured data gradually increases.
[0059] For example, the table data includes a set of structured information organized in the form of rows and columns, and can include a plurality of cells. The table image includes an image representation of the table on the visual presentation, which can be obtained by optical character recognition technology. The cell style includes the visual attributes of the cells in the table, including font, color, border, background and alignment, etc. The table content format includes the organization and representation form of the data in the table, such as data type, separator and nested structure, etc. The structure complexity can be determined using regular expressions or pre-trained machine learning models, which can represent the complexity of the table structure, and can be evaluated based on factors such as nesting level, number of merged cells, irregular rows / columns, etc. The splitting component includes a functional module for performing table splitting operations, for example, a basic splitting component can process a simple two-dimensional table, and an advanced splitting component can process nested and irregular tables.
[0060] Referring to Figure 3 Table 1 can include multiple cells in the same row, such as H11, which has information of column 1 and coordinate information belonging to the first column of the second row; H12, which has information of column 2 and coordinate information belonging to the second column of the second row. Table 2 can include multiple cells in the same row, such as H21, which has information of column 1 and coordinate information belonging to the first column of the third row; H22, which has information of column 2 and coordinate information belonging to the second column of the third row.
[0061] For example, the boundaries, table headers, data areas of the table data can be analyzed, the split points such as obvious blank lines, specific header lines are identified, and then the splitting operation is performed to generate multiple independent table blocks. Multiple splitting components can provide the ability of matching analysis, identifying split points and performing splitting operations for table data of different complexities. For example, a first splitting component can process a simple two-dimensional table, a second splitting component can identify the table structure by analyzing the coordinate positions of table texts, find the blank areas between texts as an alternative indicator of table lines, can determine the cell boundaries by detecting lines to process the bordered table and infer the structure by text alignment and spacing to process the unbordered table, and the split structure can be converted into CSV, JSON and other formats and the table structure information is retained; a third splitting component not only analyzes the text position, but also considers the visual elements (such as lines, rectangles) in the target file, can identify the table border drawn by graphics and determine the cell content in combination with the text position, can process cross-page and complex layout tables, and can also retain the font size, color and other style information of the text, and can extract other information (such as pictures, text blocks) of the page where the table is located to facilitate the retention of semantic coherence.
[0062] According to the embodiments of the present application, through multi-dimensional complexity evaluation and matching of multiple splitting components, various formats and structures of table data can be processed, and resource waste caused by using high complexity processing strategy for all tables can be avoided.
[0063] In some embodiments, metadata is added to the plurality of text blocks and the plurality of table blocks respectively to obtain the multi-modal information, which includes:
[0064] According to at least one of the entity information of the target file, the attribute information of the target file and the text block content, named entity recognition is performed to obtain the plurality of metadata of the plurality of text blocks, wherein the attribute information of the target file includes at least one of file source, file type, security level and file creation date.
[0065] According to at least one of the entity information of the target file, the attribute information of the target file, the position information of the table data in the target file, the structured feature of the table data, the column structure information in the table block and the table content, named entity recognition is performed to obtain the plurality of metadata of the plurality of table blocks, wherein the structured feature of the table data is obtained according to the position of the cell in the table.
[0066] According to the plurality of text blocks and the plurality of metadata thereof, the plurality of table blocks and the plurality of metadata thereof, and the cross-modal information, the multi-modal information is obtained, and the cross-modal information indicates the connection between the text blocks and the table blocks.
[0067] Exemplarily, the entity information includes the creator of the target file and specific objects, persons, organizations, places, and other identifiable entities mentioned in the file. The column structure information includes the attributes and characteristic descriptions of each column in the table, such as column name, column width, and data type. For example, for a research report containing multiple chapters, the key terms, research objects, time range, and other information of each chapter text block can be automatically identified as metadata, while the document source, security level, and other file attributes are associated as metadata. The title, data type, unit of measurement, and other information of each table can be identified as metadata, and the page number position in the document, whether there are merged cells, and other structural characteristics, as well as the name of each column, data format, and other column structure information are recorded, and further combined with the cell position and content of the table block as the metadata of the table block.
[0068] In some embodiments, the text blocks and table blocks with metadata can also be integrated, for example, to establish an association between different types of data blocks (for example, to use the same metadata as a link to analyze the semantic association between the text blocks and the table blocks, and to establish an association relationship across modalities), to realize cross-modality information fusion, and to create a unified index structure for the integrated multi-modal information.
[0069] For example, the text blocks and table blocks can be segmented, tagged with parts of speech, and recognized with entities, that is, all entities are recognized, each entity includes text content, entity type, and starting / ending character position in the original text. Each text block and table block can be matched with the position information of the entity to determine which entities are contained in the block, and then the corresponding block is labeled and annotated with entity annotations.
[0070] According to embodiments of the present application, by combining entity information, attribute information, text block content, and table block content of the target file for named entity recognition, adding semantic-rich metadata tags, the association between the text blocks and the overall properties of the file can be established, the structural characteristics and semantic information of the table can be captured, the context understanding can be enhanced, and the relevance of the search in the search stage can be improved.
[0071] The following will be described in conjunction with Figure 4 Embodiments of obtaining multiple indexes are introduced.
[0072] Figure 4 A schematic diagram of obtaining multiple indexes according to embodiments of the present application is shown.
[0073] The multiple indexes include a first index and a second index, as shown in Figure 4 The embodiment of obtaining multiple indexes can include operations S410-S420.
[0074] In operation S410, the multiple text blocks and the multiple table blocks are represented by dense vectors based on a semantic embedding method, and a first index is obtained.
[0075] At operation S420, the plurality of text blocks and the plurality of table blocks are represented by sparse vectors based on a word embedding manner, to obtain a second index; wherein the first index and the second index are used to provide a block-level index of the text data and a row-level index of the table data, and the index accuracy of the first index and the second index is different.
[0076] Exemplarily, the semantic embedding manner includes a method of converting text or table data into a vector representation capable of representing its semantic meaning. For example, a pre-trained language model is used to convert text into a vector, capturing deep semantic information of the text. The word embedding manner includes a method of converting a single word or phrase into a vector representation. The block-level index refers to an index structure constructed in units of text blocks. The row-level index refers to an index structure constructed for each row of table data. The index accuracy of the first index is greater than the index accuracy of the second index.
[0077] For example, referring to Figure 4 , the first index is obtained by capturing more subtle semantic associations (such as synonym differences, context implicit meanings) between the contents of the same modality data and between different modalities of data through a pre-trained language model, extracting basic semantic features of the text, and fusing the underlying extracted scattered features (such as the attributes of individual words, simple collocation relationships) into global semantic information. The second index is obtained by another pre-trained language model based on words or phrases to filter key semantics, rather than focusing on the semantics and implicit semantics of all elements globally.
[0078] According to embodiments of the present application, by combining the semantic embedding and word embedding two different vector representation methods, the first index and the second index are constructed, which can meet the needs of semantic understanding and accurate matching. Through the cooperative work of multiple indexes, the entire text block can be retrieved as needed, and specific row data in the table can also be accurately located. It can adapt to different types of query requirements and data types, has stronger adaptability, and takes into account the calculation efficiency, calculation resource consumption and retrieval accuracy.
[0079] Next, combined with Figure 5 further illustrate the embodiment of determining the target index.
[0080] Figure 5 The flowchart for determining the target index according to the embodiments of the present application is schematically shown.
[0081] As Figure 5 indicated, determining at least one target index matching at least one of the intent and the available computing resources from the plurality of indexes includes operation S510 to operation S520.
[0082] At operation S510, the target index accuracy satisfying the intent is determined according to the problem complexity indicated by the intent.
[0083] At operation S520, at least one target index is determined based on the target index precision and the available computing resources under the constraint.
[0084] For example, the problem complexity refers to the difficulty of the problem involved in the intent, which is usually related to the semantic depth of the problem, the breadth and complexity of the required information, etc. The problem complexity of the user intent can be evaluated using a pre-trained machine learning model or a predefined complexity evaluation rule. For example, the evaluation rule can determine the complexity level according to the length of the problem, the number of professional terms used, the number of information sources to be associated, the depth of reasoning to be performed, etc. The complexity level may, for example, include low complexity (simple keyword query), medium complexity (query requiring certain semantic understanding), and high complexity (query requiring deep semantic understanding and cross-document analysis), so that the corresponding target index precision retrieval can be determined.
[0085] As shown in Figure 5 There is a mapping relationship between the index, the index precision, and the consumed resource value, such as index 1, index precision 1, and consumed resource value 1, index 2, index precision 2, and consumed resource value 2, index 3, index precision 3, and consumed resource value 3. One or more target indexes can be selected according to the target index precision and the available computing resources in combination with the mapping relationship.
[0086] For another example, the central processor utilization rate, memory usage, storage remaining space, network bandwidth, etc. are collected to generate the resource constraint condition of the current system. For example, when the memory usage exceeds a certain threshold, the index with large memory usage is limited, and at least one target index that meets the target index precision is selected. For example, the target index precision and the available computing resource value can also be assigned weights, and then a weighted sum is obtained to obtain a predicted consumed resource value. According to the preset mapping relationship between each index and the consumed resource value, at least one target index is determined.
[0087] According to the embodiments of the present application, by associating the problem complexity of the intent, the index precision, and the computing resources, the index precision level that matches the user intent of different complexity can be selected, and the computing resources can also be reasonably allocated and utilized under the premise of meeting the retrieval requirements.
[0088] It can be understood that the manner of determining the target index is not limited to the embodiments of Figure 5 The following further describes.
[0089] In some embodiments, the multi-modal data to be obtained can be determined according to at least one of the intent and the available computing resources, and then at least one target index of each modal data in the multi-modal data is determined.
[0090] For example, when the intent is "find product reviews" and computing resources are sufficient, it is determined to obtain text reviews, rating data, and related pictures; when the intent is "quickly understand news highlights" and computing resources are limited, it is determined to only obtain text titles and summaries, and to give up obtaining pictures and videos; when it is detected that computing resources are tight, the intent contains image analysis requirements, and it is also possible to prioritize low-resolution images or only obtain text descriptions. Then, for example, product reviews mainly focus on text, and text reviews are retrieved using high-precision vector indexing, while pictures are retrieved using medium-precision inverted indexing.
[0091] By selecting a suitable target index for each type of modal data, data retrieval and processing time can be reduced, response speed can be improved, and different question and answer requirements and computing resource conditions can be adapted to.
[0092] In some embodiments, a large model can be used to simulate the human step-by-step thinking process, and complex problems can be decomposed into a logically coherent and serialized multiple sub-problems. Then, multiple sub-intents can be identified for multiple sub-problems. Then, at least one target index for each sub-intent can be determined for each sub-intent and the available computing resources at the time of retrieving the corresponding sub-intent. The target indexes between the sub-intents can be the same or different, for example, the sub-int Figure 1 that requires deep analysis uses high-precision vector indexing; the sub-int Figure 2 that requires original content uses medium-precision inverted indexing to quickly retrieve marketing documents; the sub-int Figure 3 When the available computing resources are less, a low-precision hash index is used. Alternatively, when the available resources are more, at least two sub-intents can also be retrieved in parallel in the corresponding target index, for example, first retrieve the data of the sub-int Figure 1 through high-precision vector indexing, and then retrieve the data of the sub-int Figure 2 through medium-precision inverted indexing in parallel, and retrieve the data of the sub-int Figure 3 through low-precision hash indexing.
[0093] The following further illustrates embodiments of retrieving target multi-modal information.
[0094] In some embodiments, retrieving target multi-modal information based on the intent through at least one target index comprises:
[0095] Perform named entity recognition based on the question and answer context of the target object to obtain screening information of the intent, the screening information including at least one of entity information, file source, file type, security level, file creation date, specific position of a cell in a table row and column, and geometric position of a cell in a table image, the question and answer context including the question and historical question and answer content of the target object.
[0096] The candidate multi-modal information includes a candidate text block and a candidate table block.
[0097] The target multi-modal information is obtained based on matching the screening information with metadata of at least one of the candidate text block and the candidate table block in the candidate multi-modal information.
[0098] Exemplarily, the historical question-and-answer content of the target object includes content generated in the current session of the target object, and can also include question-and-answer content generated by the target object based on other questions before the current session. Each session starts from, for example, a time when the user opens a question-and-answer interface to input a first question, and ends at a time when the user closes the question-and-answer interface or does not input new content within a predetermined time period.
[0099] According to the embodiments of the present application, accurate screening information is obtained through named entity recognition, and secondary screening is performed in combination with metadata matching, which effectively avoids interference of irrelevant information, improves matching degree of the retrieval result and the user intent, and reduces operation burden and waiting time of the user.
[0100] The target index can determine one or more. If multiple target indexes are determined, in some embodiments, when the target index includes a first index and a second index, the target multi-modal information is obtained based on the intent through at least one target index, including:
[0101] The dense retrieval and the sparse retrieval are performed on the question in the first index and the second index respectively, to obtain a plurality of candidate retrieval results, a plurality of dense ranking scores and a plurality of sparse ranking scores. The candidate retrieval result includes a result retrieved in the first index and the second index.
[0102] The first ranking result of the plurality of candidate retrieval results is determined according to weighted fusion results of the dense ranking score and the sparse ranking score of each candidate retrieval result.
[0103] The target multi-modal information is obtained according to the first ranking result from the plurality of candidate retrieval results to screen at least one target retrieval result.
[0104] Exemplarily, the dense retrieval mainly uses a dense vector to represent the semantics of the text, for example, the retrieval is implemented through a neighbor search of the vector. The sparse retrieval mainly uses a sparse vector of a size of a word table to represent the text, and the index is established based on an inverted index. The dense ranking score includes a score for measuring a semantic relevance degree of the retrieval result and the user question, which is calculated based on a dense vector similarity; and the sparse ranking score includes a score for measuring a semantic relevance degree of the retrieval result and the user question, which is calculated based on a sparse vector similarity.
[0105] For example, the question is converted into a dense vector and a sparse vector, an approximate nearest neighbor search is performed in the first index, a dense similarity score is calculated, and K1 results are obtained based on the K1 results. Extract the keywords in the question, perform an inverted index search in the second index, calculate the sparse matching score, and obtain K2 results. K1 and K2 are integers greater than or equal to 1. Take the same multiple candidate search results from the K1 results and the K2 results. Then, determine the weight of the dense ranking score and the sparse ranking score (which can be pre-determined according to business needs or predicted in real time by a machine learning model), and for each candidate search result, calculate the weighted sum of its dense ranking score and sparse ranking score, such as the dense ranking score multiplied by 0.6 and then added to the sparse ranking score multiplied by 0.4 (the weight is only an example). According to the weighted fusion result, all candidate search results are sorted in descending order to generate the final first ranking result list, and the top 3 (only an example) are returned.
[0106] According to the embodiments of the present application, by means of weighted fusion, the semantic relevance and keyword matching degree can be considered comprehensively, the semantic and lexical relevance can be balanced, and the retrieved results can be sorted more reasonably.
[0107] In some embodiments, determining the first ranking results of the multiple candidate search results according to the weighted fusion result of the dense ranking score and the sparse ranking score of each candidate search result comprises:
[0108] Processing the multiple dense ranking scores based on multiple first preset weights to obtain multiple first modified scores, the first preset weight indicating the contribution degree of the semantic matching between the target search result and the question.
[0109] Processing the multiple sparse ranking scores based on multiple second preset weights to obtain multiple second modified scores, the second preset weight indicating the contribution degree of the word matching between the target search result and the question.
[0110] Obtaining the first ranking result according to the first modified score and the second modified score of each candidate search result.
[0111] For example, the multiple first preset weights can include the same or different weights, and the multiple second preset weights can also include the same or different weights. The weights can be pre-set according to the frequency of the search result in the historical question and answer content of the target object, and the higher the frequency, the greater the weight value. The preset weights can be allocated according to different modal data. The first preset weight and the second preset weight can also be allocated to adjust the contribution of semantics and words.
[0112] According to the embodiments of the present application, the weight values can be preset according to different application scenarios, user demands or data characteristics, and the contribution degrees of semantic matching and word matching in the final sorting can be flexibly adjusted. The way of introducing preset weights for correction can further take into account the deep understanding at the semantic level and the accurate matching at the vocabulary level, effectively make up for the limitations of a single retrieval dimension, and thus significantly improve the relevance and accuracy of the retrieval results and the user questions.
[0113] In some embodiments, the obtaining the target multi-modal information according to the at least one target retrieval result from the plurality of candidate retrieval results includes:
[0114] The cross-encoder model is used to calculate a plurality of relevancies between the question-answer context and the plurality of candidate retrieval results, and the question-answer context includes a question and historical question-answer content of the target object.
[0115] The first sorting result is reordered based on the plurality of relevancies to obtain a second sorting result of the plurality of candidate retrieval results.
[0116] The at least one target retrieval result is selected from the plurality of candidate retrieval results according to the second sorting result to obtain the target multi-modal information.
[0117] For example, a pre-trained cross-encoder model based on a transformer can be used to calculate the interaction features between the question-answer context and the plurality of candidate retrieval results through a multi-head cross-attention mechanism, align each query (query vector) with the plurality of candidate retrieval results to capture semantic associations, and output a relevance score to achieve reordering.
[0118] According to the embodiments of the present application, the reordering mechanism can fully consider the most relevant information to improve the matching degree and relevance of the retrieval results.
[0119] In the embodiments of the present application, the target multi-modal information can be obtained by rewriting or expanding the question, which will be further described below.
[0120] In some embodiments, the obtaining the target multi-modal information based on the intent through the at least one target index includes:
[0121] The question-answer context of the target object and a feedback result of at least one historical reply of the target object to the question-answer context are obtained, the question-answer context includes a question and historical question-answer content of the target object, and the feedback result includes positive feedback and negative feedback.
[0122] The question is rewritten or expanded according to at least one of the question-answer context, the feedback result of the at least one historical reply, a keyword in the question and semantic information of a representation of the question, and the intent, and the target multi-modal information is obtained based on the rewritten or expanded question.
[0123] Exemplarily, the rewriting refers to an optimized adjustment of the original question to make it clearer and more accurate to express the user's intention. The expansion refers to supplementing relevant information or refining the angle of questioning on the basis of maintaining the core content of the original question.
[0124] For example, after the reply is displayed to the user, the user can score the reply, or like or dislike it, so as to collect the feedback result. The rewritten or expanded question output by the large model can make up for the problems such as brevity and ambiguity of the original question, and can also rewrite the original question into content more in line with the retrieval requirements, thereby improving the coverage rate and accuracy. For example, when the original question is expressed ambiguously, the information is incomplete, or there is ambiguity, the question rewriting is preferentially performed (which can be further expanded after rewriting); when the original question is relatively specific but needs to supplement relevant information, the question expansion is performed. The question understanding process can be embedded with the question and answer context in the rewriting or expansion process, thereby improving the accuracy of understanding.
[0125] According to the embodiments of the present application, by combining positive feedback and negative feedback, the user's preferences can be determined and the content not meeting the requirements can be excluded, thereby significantly improving the accuracy and relevance of the retrieval result to the actual needs of the user, and further serving as a basis for rewriting or expanding the original question. The problems such as unclear expression, incomplete information, or improper words of the user's question can be effectively made up, the retrieval can be more accurate, the irrelevant results can be reduced, and the retrieval efficiency can be improved.
[0126] In the embodiments of the present application, all questions can be rewritten or expanded, or the questions can be selectively rewritten or expanded. In some embodiments, at least one of an intention accuracy and a content completeness of the intention is evaluated, the intention accuracy indicates an accuracy degree of a target object intention expressed by the question, and the content completeness indicates a completeness degree of the target object intention expressed by the question content; in response to the question satisfying at least one of a first threshold value, the content completeness being less than a second threshold value, the question is rewritten or expanded, so that the intention accuracy of the rewritten or expanded question is greater than the first threshold value or the content completeness is greater than the second threshold value.
[0127] Exemplarily, the intention accuracy indicates an accuracy degree of a target object intention expressed by the question. The content completeness indicates a completeness degree of the target object intention expressed by the question content. For example, the user question is input into a trained intention recognition model, and a score between 0 and 1 is output, indicating the accuracy degree of the intention expression. For the content completeness evaluation, it can be determined by checking whether the necessary entity information is contained in the question. The first threshold value and the second threshold value can be determined in advance according to historical data statistics or expert experience.
[0128] According to the embodiments of the present application, by automatically identifying and optimizing unclear expressions and incomplete information, the need for multiple information supplements by the user can be reduced, the questioning habits and expression abilities of different users can be adapted, ambiguity understanding and invalid calculations in the processing process can be reduced, and the interaction efficiency and user satisfaction can be improved.
[0129] The following describes an embodiment of generating a reply by a preset model.
[0130] In some embodiments, the preset model includes a first preset model and a second preset model, and processing target multi-modal information based on the question by using the preset model includes:
[0131] Processing the target multi-modal information based on the question by using the first preset model obtains a first reply.
[0132] Processing the target multi-modal information based on the question by using the second preset model obtains a second reply, and the second preset model is different from the first preset model in at least one of inference accuracy, inference speed, and calculation resource consumption.
[0133] Based on the quality score result of the first reply and the second reply, a reply is obtained.
[0134] Illustratively, inference accuracy refers to the ability of a model to generate correct and relevant replies based on input information. Inference speed refers to the time required for a model to generate an output from receiving an input. Calculation resource consumption refers to the amount of processor, memory, video memory, and other computing resources occupied by the model when running.
[0135] For example, the first preset model is a large model with more parameters, higher inference accuracy, and more calculation resource consumption than the second preset model, and the second preset model is a large model with fewer parameters, faster inference speed, and less calculation resource consumption.
[0136] For example, the third preset model can be used to score the accuracy, relevance, completeness, fluency, etc. of the first reply and the second reply to obtain a quality score result, and the one with a higher score is taken as the reply.
[0137] In some embodiments, a prompt word can also be input to the first preset model and the second preset model, such as "answer strictly based on the retrieved sources; use bullet points to ensure clarity; provide citations of source documents; if the response exceeds three sentences, include an abstract." In addition, the prompt word can be replaced according to different questions or execution effects.
[0138] In some embodiments, the first reply and the second reply can also be respectively decomposed into multiple sub-replies, and quality evaluation results of the multiple sub-replies are obtained, so that the sub-replies output by the first preset model and the second preset model are combined to obtain a reply displayed to the user.
[0139] According to the embodiments of the present application, by using two models with different characteristics to process information and generate replies respectively, and then selecting the reply with the best quality, the quality of the output reply can be effectively improved, and the limitations of a single model can be avoided. The use of computing resources can be optimized while ensuring the quality of the reply.
[0140] In some embodiments, the reply can also be evaluated based on the question and answer context and the feedback result of the at least one historical reply; and in response to the evaluation result indicating that the reply does not meet the preset condition, the candidate multi-modal information is re-retrieved based on the rewritten or expanded question through the at least one target index.
[0141] According to the embodiments of the present application, the evaluation in combination with historical feedback and question and answer context can better understand the question and answer requirements and provide a reply that better meets expectations, thereby improving user satisfaction and use experience. Through the closed-loop mechanism of re-evaluation and re-retrieval, problems existing in the system can be continuously found and corrected, and self-learning and continuous optimization can be achieved.
[0142] In some embodiments, the feedback result of the target object on the reply can also be obtained; and in response to the feedback result of the reply being negative feedback, the question is re-executed for rewriting or expansion. For example, the query intention of the user is sometimes complex or implicit, and a single question rewriting or expansion may not be able to fully capture it. Through the multiple optimization mechanisms triggered by negative feedback, the real needs of the user can be gradually approached.
[0143] According to the embodiments of the present application, dynamic adjustment can be made according to actual interaction. When the target object provides negative feedback, it can be identified that the current reply fails to meet its needs, and by re-executing question rewriting or expansion, the query expression can be optimized so that the query accurately reflects the real intention of the user, thereby continuously learning and optimizing according to the real-time feedback of the user and improving the adaptability to different user needs.
[0144] Figure 6 A schematic diagram of a question and answer method according to another embodiment of the present application is shown.
[0145] As Figure 6As shown, first, one or more of the multi-modal data in the target file is extracted, such as text, table, formula, picture, audio, and video, etc. Then, a vector database is constructed using multiple components, for example, text data is processed using a text processing component to obtain multiple text blocks, table data is processed using a table processing component to obtain multiple table blocks, and the semantic coherence between blocks, structured information, and cross-modal connections are preserved. Then, a metadata augmentation component is used to add metadata to the text blocks and table blocks to obtain multi-modal information. Then, an indexing component is used to construct multiple indexes based on the multi-modal information and store them in the vector database.
[0146] With reference to the foregoing Figure 6 When the user asks a question, at least one target index can be determined according to at least one of the question and available computing resources, wherein the question can be rewritten or expanded before determination. Then, when multiple target indexes are determined, mixed retrieval can be performed to improve the ability of the large model to balance between semantic understanding and accurate matching, and to avoid missing relevant information or false detection of irrelevant information. For example, dense retrieval is performed in a high-precision vector index, sparse retrieval is performed in a medium-precision inverted index, and then the results of the dense retrieval and the sparse retrieval are weighted and fused to obtain a mixed retrieval result. Then, a ranking model, such as a cross-encoder model, is used to reorder the mixed retrieval result. The reordered result is input to the large model to obtain a reply output by the large model, which is displayed to the user. Then, if the user provides negative feedback, the large model is triggered to reconstruct (i.e., rewrite) the question and expand it, and the question is re-retrieved based on the rewritten and expanded question. This cycle continues until the user is satisfied.
[0147] For example, the embodiments of the present application can be used in enterprise scenarios, and can also be widely used in the fields of medical health and law. In the field of medical health, heterogeneous data such as electronic medical records and medical literature can be processed, key information semantics and structures are preserved through structure perception blocking, similar cases, treatment plans, etc. are quickly located by combining mixed retrieval and metadata filtering, doctors are assisted in diagnosis and decision-making, and feedback cycles can continuously optimize the system; in the legal field, legal provisions and case documents can be retrieved, key information of contract clauses can be accurately extracted using semantic blocking and line-level indexing, and lawyers can be assisted in efficiently handling cases through entity filtering and query optimization. In addition, there are also application values in the fields of education and finance. In the field of education, resources such as teaching materials and test item banks can be integrated, semantic coherence of text blocking ensures knowledge coherence, line-level indexing facilitates test paper compilation, metadata filtering supports personalized resource recommendation, and teaching and learning are assisted; in the field of finance, financial reports and transaction records can be processed, market trends and risk information can be quickly captured with the help of structure perception blocking and mixed retrieval, and line-level indexing can achieve accurate customer analysis to assist investment decision-making and risk assessment.
[0148] According to the embodiment of the present application, the hybrid retrieval mechanism combines the advantages of dense retrieval and sparse retrieval, balances semantic depth and lexical accuracy, ensures the semantic coherence of text and the structural integrity of table data, lays a foundation for accurate retrieval, the metadata-driven filtering further improves the relevance of retrieval through rich entity information and document attributes, the cross-encoder reordering optimizes the retrieval results, and improves the context matching degree, and the dynamic query optimization can adjust the query according to user feedback and conversation conditions, improves the adaptability and accuracy of retrieval.
[0149] Based on the above question and answer method, the present application also provides a question and answer device. The following will be combined with Figure 7 The device is described in detail.
[0150] Figure 7 The structural block diagram of the question and answer device according to the embodiment of the present application is schematically shown.
[0151] As Figure 7 shown, the question and answer device 700 of the embodiment includes an index module 710, a retrieval module 720 and a generation module 730.
[0152] The index module 710 can perform operation S210 for determining at least one target index matched with at least one of the intent and the available computing resource for performing retrieval from a plurality of indexes in response to receiving a question from a target object according to at least one of the intent of the question representation and the available computing resource, wherein the plurality of indexes are obtained by respectively processing the multi-modal information with a plurality of index precisions, and the multi-modal information includes a structured representation of a plurality of modal data obtained by processing a target file.
[0153] The retrieval module 720 can perform operation S220 for retrieving target multi-modal information including a plurality of candidate modal data and respective metadata indicating description information of the corresponding modal data matching the intent through the at least one target index based on the intent.
[0154] The generation module 730 can perform operation S230 for processing the target multi-modal information based on the question using a preset model to output a reply to the intent.
[0155] In some embodiments, the question and answer device 700 can further include a preprocessing module configured to obtain the multi-modal information by: dividing text data in the target file into a plurality of text blocks, and a structured representation of the plurality of text blocks includes that a predetermined number of characters are overlapped between adjacent text blocks, and the predetermined number of characters are used to represent semantic coherence between adjacent text blocks; splitting table data in the target file into a plurality of table blocks, and a structured representation of the plurality of table blocks includes column structure information corresponding to table rows of each table block and table content based on column structure distribution; and adding metadata to the plurality of text blocks and the plurality of table blocks respectively to obtain the multi-modal information.
[0156] In some embodiments, adding metadata to the plurality of text blocks and the plurality of table blocks respectively to obtain the multi-modal information includes: performing named entity recognition according to at least one of entity information of the target file, attribute information of the target file, and text block content to obtain a plurality of metadata of the plurality of text blocks, wherein the attribute information of the target file includes at least one of file source, file type, security level, and file creation date; performing named entity recognition according to at least one of entity information of the target file, attribute information of the target file, position information of the table data in the target file, structured features of the table data, column structure information in the table block, and table content to obtain a plurality of metadata of the plurality of table blocks, wherein the structured features of the table data are obtained according to positions of cells in the table; and obtaining the multi-modal information according to the plurality of text blocks and the plurality of metadata thereof, the plurality of table blocks and the plurality of metadata thereof, and cross-modal information, wherein the cross-modal information indicates a connection between the text blocks and the table blocks.
[0157] In some embodiments, the retrieval module 720 includes a screening information unit, a candidate unit, and a screening unit. The screening information unit is configured to perform named entity recognition according to a question and answer context of the target object to obtain screening information of the intent, the screening information including at least one of entity information, file source, file type, security level, file creation date, specific positions of cells in table rows and columns, and geometric positions of cells in table images, and the question and answer context including a question and historical question and answer content of the target object. The candidate unit is configured to retrieve candidate multi-modal information based on the intent through at least one target index, and the candidate multi-modal information including candidate text blocks and candidate table blocks. The screening unit is configured to match the screening information with metadata of at least one of the candidate text blocks and the candidate table blocks in the candidate multi-modal information to obtain the target multi-modal information.
[0158] In some embodiments, the predetermined number of characters overlapped between adjacent text blocks includes that the predetermined number accounts for more than or equal to one fourth of a total number of characters in each text block.
[0159] In some embodiments, the adjacent text blocks include a first text block and a second text block, and the predetermined number of characters overlapping between the adjacent text blocks is obtained by: obtaining semantic content represented by the first text block and the second text block; extracting a keyword set representing the semantic content from the first text block and the second text block; and adding missing keywords to the first text block and the second text block based on the keyword set.
[0160] In some embodiments, splitting the table data in the target file into a plurality of table blocks includes: evaluating a structural complexity of the table data according to at least one of a table image, a cell style, a table content format, and a table embedding manner of the table data; and selecting a target splitting component matching the structural complexity from a plurality of preset splitting components, and splitting the table data into a plurality of table blocks, wherein the plurality of splitting components are configured to provide matching splitting services for table data with different structural complexities.
[0161] In some embodiments, the question and answer device 700 can further include an index construction module configured to construct a plurality of indexes, specifically a first index and a second index, wherein the plurality of text blocks and the plurality of table blocks are represented by dense vectors based on a semantic embedding manner to obtain the first index, and the plurality of text blocks and the plurality of table blocks are represented by sparse vectors based on a word embedding manner to obtain the second index; wherein the first index and the second index are configured to provide block-level indexes of the text data and row-level indexes of the table data, and the index accuracy of the first index and the second index is different.
[0162] In some embodiments, when the target index includes the first index and the second index, the retrieval module 720 is further configured to perform dense retrieval in the first index and sparse retrieval in the second index based on the question to obtain a plurality of candidate retrieval results, a plurality of dense ranking scores, and a plurality of sparse ranking scores, wherein the candidate retrieval results include results retrieved in the first index and the second index; determine a first ranking result of the plurality of candidate retrieval results according to weighted fusion of the dense ranking score and the sparse ranking score of each candidate retrieval result; and filter at least one target retrieval result from the plurality of candidate retrieval results according to the first ranking result to obtain the target multi-modal information.
[0163] In some embodiments, determining the first ranking result of the plurality of candidate retrieval results according to the weighted fusion of the dense ranking scores and the sparse ranking scores of each candidate retrieval result comprises: processing the plurality of dense ranking scores based on a plurality of first preset weights to obtain a plurality of first modified scores, the first preset weights indicating the contribution degree of semantic matching between the target retrieval result and the question; processing the plurality of sparse ranking scores based on a plurality of second preset weights to obtain a plurality of second modified scores, the second preset weights indicating the contribution degree of word matching between the target retrieval result and the question; and obtaining the first ranking result according to the first modified score and the second modified score of each candidate retrieval result.
[0164] In some embodiments, the question and answer device 700 can further comprise a reordering module configured to calculate a plurality of relevancies of the question and answer context of the target object and the plurality of candidate retrieval results respectively using a cross-encoder model, the question and answer context comprising the question and historical question and answer content of the target object; reorder the first ranking result based on the plurality of relevancies to obtain a second ranking result of the plurality of candidate retrieval results; and the preprocessing module is configured to filter at least one target retrieval result from the plurality of candidate retrieval results according to the second ranking result to obtain the target multi-modal information.
[0165] In some embodiments, the index module 710 is further configured to determine a target index accuracy satisfying the intent according to the question complexity indicated by the intent; and determine at least one target index based on constraint limitation of the target index accuracy and available computing resources.
[0166] In some embodiments, the preset model comprises a first preset model and a second preset model, and the generation module 730 is further configured to process the target multi-modal information based on the question using the first preset model to obtain a first reply; process the target multi-modal information based on the question using the second preset model to obtain a second reply, the second preset model being different from the first preset model in at least one of inference accuracy, inference speed, and computing resource consumption; and obtain the reply based on a quality score result of the first reply and the second reply.
[0167] In some embodiments, the retrieval module 720 is further configured to obtain a question and answer context of the target object and a feedback result of at least one historical reply in the question and answer context by the target object, the question and answer context comprising the question and historical question and answer content of the target object, and the feedback result comprising positive feedback and negative feedback; perform rewriting or expansion on the question according to at least one of the question and answer context, the feedback result of the at least one historical reply, a keyword in the question, and semantic information representing the question, and the intent, to retrieve the target multi-modal information based on the rewritten or expanded question.
[0168] In some embodiments, the question-answering apparatus 700 can further comprise a first self-optimization module configured to evaluate the reply based on the question-answering context and a feedback result of the at least one historical reply; and in response to an evaluation result indicating that the reply does not meet a preset condition, re-retrieve candidate multi-modal information based on the rewritten or expanded question through the at least one target index.
[0169] In some embodiments, the question-answering apparatus 700 can further comprise a second self-optimization module configured to evaluate at least one of an intention accuracy and a content completeness of the intention, the intention accuracy indicating an accuracy degree of the target object intention expressed by the question, and the content completeness indicating a completeness degree of the target object intention expressed by the question content; and in response to the question meeting at least one of the intention accuracy being less than a first threshold value and the content completeness being less than a second threshold value, perform rewriting or expansion on the question so that the intention accuracy of the rewritten or expanded question is greater than the first threshold value or the content completeness is greater than the second threshold value.
[0170] In some embodiments, the question-answering apparatus 700 can further comprise a third self-optimization module configured to obtain a feedback result of the target object on the reply; and in response to the feedback result of the reply being negative feedback, re-perform rewriting or expansion on the question.
[0171] According to embodiments of the present application, any of the index module 710, the retrieval module 720 and the generation module 730 can be combined in one module, or any of the modules can be split into multiple modules. Alternatively, at least part of the function of one or more of the modules can be combined with at least part of the function of the other modules, and implemented in one module. According to embodiments of the present application, at least one of the index module 710, the retrieval module 720 and the generation module 730 can be at least partially implemented as a hardware circuit, such as a field programmable gate array (FPGA), a programmable logic array (PLA), a system on chip, a system on board, a system on package, an application specific integrated circuit (ASIC), or any other reasonable way of hardware or firmware that can be integrated or packaged, or implemented in any one of software, hardware and firmware or in a proper combination of any of them. Alternatively, at least one of the index module 710, the retrieval module 720 and the generation module 730 can be at least partially implemented as a computer program module that can perform corresponding functions when the computer program module is run.
[0172] Figure 8 A block diagram of an electronic device suitable for implementing the question-answering method according to embodiments of the present application is schematically shown.
[0173] As Figure 8As shown, the electronic device 800 according to the embodiments of the present application includes a processor 801, which can perform various appropriate actions and processes according to a program stored in a read only memory (ROM) 802 or a program loaded into a random access memory (RAM) 803 from a storage section 808. The processor 801 can include, for example, a general purpose microprocessor (e.g., a CPU), an instruction set processor, and / or a dedicated microprocessor (e.g., an application specific integrated circuit (ASIC)), and so on. The processor 801 can also include an on-board memory for cache use. The processor 801 can include a single processing unit or multiple processing units for performing different actions of the method processes according to the embodiments of the present application.
[0174] In the RAM 803, various programs and data required for the operation of the electronic device 800 are stored. The processor 801, the ROM 802, and the RAM 803 are connected to each other via a bus 804. The processor 801 performs various operations of the method processes according to the embodiments of the present application by executing the programs in the ROM 802 and / or the RAM 803. Note that the programs can also be stored in one or more memories other than the ROM 802 and the RAM 803. The processor 801 can also perform various operations of the method processes according to the embodiments of the present application by executing the programs stored in the one or more memories.
[0175] According to the embodiments of the present application, the electronic device 800 can further include an input / output (I / O) interface 805, which is also connected to the bus 804. The electronic device 800 can further include one or more of the following components connected to the input / output (I / O) interface 805: an input section 806 including a keyboard, a mouse, etc.; an output section 807 including a display such as a cathode ray tube (CRT), a liquid crystal display (LCD), etc., and a speaker, etc.; a storage section 808 including a hard disk, etc.; and a communication section 809 including a network interface card such as a LAN card, a modem, etc. The communication section 809 performs communication processing via a network such as the Internet. A drive 810 is also connected to the input / output (I / O) interface 805 as necessary. A removable recording medium 811 such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc. is attached to the drive 810 as necessary, so that a computer program read therefrom is installed into the storage section 808 as necessary.
[0176] The present application also provides a computer readable storage medium, which can be included in the device / apparatus / system described in the above embodiments; or can exist separately without being assembled into the device / apparatus / system. The above computer readable storage medium carries one or more programs, which when executed, implement the method according to the embodiments of the present application.
[0177] According to an embodiment of the present application, the computer readable storage medium can be a non-transitory computer readable storage medium, for example, can include but not limited to: a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing. In this application, a computer readable storage medium can be any tangible medium that contains or stores a program for use by or in connection with an instruction execution system, apparatus, or device. For example, according to an embodiment of the present application, the computer readable storage medium can include the ROM 802 and / or the RAM 803 described above and / or one or more memory other than the ROM 802 and the RAM 803.
[0178] Embodiments of the present application also include a computer program product, which includes a computer program containing program codes for executing the methods shown in the flowcharts. When the computer program product is run in a computer system, the program codes are used to make the computer system implement the question answering method provided by the embodiments of the present application.
[0179] The above functions defined in the system / device of the embodiments of the present application are performed when the computer program is executed by the processor 801. According to an embodiment of the present application, the system, device, module, unit, etc. described above can be implemented by computer program modules.
[0180] In one embodiment, the computer program can rely on tangible storage media such as optical storage media, magnetic storage media, etc. In another embodiment, the computer program can also be transmitted, distributed, and downloaded in the form of signals on a network medium. The computer program containing program codes can be transmitted by any appropriate network medium, including but not limited to wireless, wired, etc., or any suitable combination of the foregoing.
[0181] In such an embodiment, the computer program can be downloaded and installed from the network by the communication part 809, and / or installed from the detachable medium 811. When the computer program is executed by the processor 801, the above functions defined in the system of the embodiments of the present application are performed. According to an embodiment of the present application, the system, device, apparatus, module, unit, etc. described above can be implemented by computer program modules.
[0182] According to embodiments of the present application, program code for implementing the computer programs provided by embodiments of the present application can be written in any combination of one or more programming languages, and can be implemented in a high-level procedural and / or object-oriented programming language, and / or in assembly / machine language. Programming languages include, but are not limited to, Java, C++, python, "C", or the like. Program code can execute entirely on a user's computing device, partly on the user's device, as a stand-alone software package, partly on a remote computing device, or entirely on the remote computing device or server. In the latter scenario, the remote computing device can be connected to the user's computing device through any type of network, including a local area network (LAN) or a wide area network (WAN), or the connection can be made to an external computing device, such as through the Internet using an Internet Service Provider.
[0183] The computer program instructions can also be loaded onto a computer or other programmable information processing apparatus to cause a series of operations to be performed on the computer or other programmable information processing apparatus to produce a computer implemented process such that the instructions which execute on the computer or other programmable information processing apparatus implement the functions / acts specified in the flowchart and / or block diagram block or blocks.
[0184] Those skilled in the art will appreciate that features recited in the various embodiments of the present application can be combined and / or integrated in various combinations, even if such combinations have not been explicitly recited in the present application. In particular, the features recited in the various embodiments of the present application can be combined and / or integrated in various combinations, without departing from the spirit and scope of the present application. All such combinations are within the scope of the present application.
Claims
1. A question and answer method, characterized by, The method comprises the following steps: in response to receiving a question from a target object, determining at least one target index matching the intent represented by the question and the available computing resources for retrieval from a plurality of indexes, wherein the plurality of indexes comprise data index structures of different index precision levels, which are obtained by processing multi-modal information at a plurality of index precisions, and the multi-modal information comprises structured representations of a plurality of modal data obtained by processing a target file; retrieving target multi-modal information based on the intent through the at least one target index, wherein the target multi-modal information comprises a plurality of candidate modal data and respective metadata, and the metadata indicates description information indicating that the corresponding modal data matches the intent; processing the target multi-modal information based on the question using a preset model to output a reply to the intent; wherein the step of determining at least one target index matching the intent represented by the question and the available computing resources for retrieval from a plurality of indexes comprises: determining a target index precision that satisfies the intent according to the complexity of the question indicated by the intent; determining at least one target index of each modal data in the multi-modal data to be obtained based on the target index precision and the available computing resources.
2. The method of claim 1, wherein, The multi-modal information is obtained in the following manner: segmenting text data in the target file into a plurality of text blocks, wherein the structured representation of the plurality of text blocks comprises a plurality of characters overlapping between adjacent text blocks, and the plurality of characters are used to represent semantic coherence between adjacent text blocks; splitting table data in the target file into a plurality of table blocks, wherein the structured representation of the plurality of table blocks comprises column structure information of each table block corresponding to a table row and table content based on column structure distribution; adding metadata to the plurality of text blocks and the plurality of table blocks respectively to obtain the multi-modal information.
3. The method of claim 2, wherein, The step of adding metadata to the plurality of text blocks and the plurality of table blocks respectively to obtain the multi-modal information comprises: performing named entity recognition based on at least one of entity information of the target file, attribute information of the target file, and text block content to obtain a plurality of metadata of the plurality of text blocks, wherein the attribute information of the target file comprises at least one of file source, file type, security level, and file creation date; performing named entity recognition based on at least one of entity information of the target file, attribute information of the target file, position information of the table data in the target file, structured features of the table data, column structure information in the table block, and table content to obtain a plurality of metadata of the plurality of table blocks, wherein the structured features of the table data are obtained based on the position of a cell in the table; obtaining the multi-modal information based on the plurality of text blocks and their metadata, the plurality of table blocks and their metadata, and cross-modal information, wherein the cross-modal information indicates the relationship between the text blocks and the table blocks.
4. The method of claim 3, wherein, The retrieving the target multi-modal information based on the intention through the at least one target index comprises: performing named entity recognition according to a question and answer context of the target object to obtain screening information of the intention, the screening information comprising at least one of the entity information, a file source, a file type, a security level, a file creation date, a specific position of a cell in a table row and column, and a geometric position of the cell in a table image, the question and answer context comprising the question and historical question and answer content of the target object; retrieving candidate multi-modal information based on the intention through the at least one target index, the candidate multi-modal information comprising candidate text blocks and candidate table blocks; matching the screening information with metadata of at least one of the candidate text blocks and the candidate table blocks in the candidate multi-modal information to obtain the target multi-modal information.
5. The method of claim 2, wherein, The overlapping of the predetermined number of characters between the adjacent text blocks comprises: The predetermined number accounts for more than or equal to one fourth of the total number of characters in each text block.
6. The method of claim 2, wherein, The adjacent text blocks comprise a first text block and a second text block, and the predetermined number of characters overlapping between the adjacent text blocks is obtained by: obtaining semantic content represented by the first text block and the second text block; extracting a keyword set representing the semantic content from the first text block and the second text block; adding missing keywords to the first text block and the second text block based on the keyword set.
7. The method of claim 2, wherein, The splitting the table data in the target file into a plurality of table blocks comprises: evaluating structural complexity of the table data according to at least one of a table image, a cell style, a table content format, and a table embedding manner of the table data; selecting a target splitting component matching the structural complexity from a plurality of preset splitting components, and splitting the table data into the plurality of table blocks, wherein the plurality of splitting components are configured to provide matching splitting services for table data with different structural complexities.
8. The method of claim 2, wherein, The plurality of indexes comprise a first index and a second index, and the method further comprises: performing dense vector representation of the plurality of text blocks and the plurality of table blocks based on a semantic embedding manner to obtain the first index; performing sparse vector representation of the plurality of text blocks and the plurality of table blocks based on a word embedding manner to obtain the second index; wherein the first index and the second index are configured to provide block-level indexing of the text data and row-level indexing of the table data, and the index accuracy of the first index and the second index is different.
9. The method of claim 8, wherein, In a case where the target index comprises the first index and the second index, the retrieving the target multi-modal information based on the intention through the at least one target index comprises: performing dense retrieval in the first index and sparse retrieval in the second index based on the question to obtain a plurality of candidate retrieval results, a plurality of dense ranking scores, and a plurality of sparse ranking scores, the candidate retrieval results comprising results retrieved in both the first index and the second index; determine a first ranking result of the plurality of candidate retrieval results according to weighted fusion results of the dense ranking scores and the sparse ranking scores of each candidate retrieval result; screen at least one target retrieval result from the plurality of candidate retrieval results according to the first ranking result to obtain the target multi-modal information.
10. The method of claim 9, wherein, The determining a first ranking result of the plurality of candidate retrieval results according to weighted fusion results of the dense ranking scores and the sparse ranking scores of each candidate retrieval result includes: processing the plurality of dense ranking scores based on a plurality of first preset weights to obtain a plurality of first modified scores, the first preset weights indicating a contribution degree of semantic matching between a target retrieval result and the question; processing the plurality of sparse ranking scores based on a plurality of second preset weights to obtain a plurality of second modified scores, the second preset weights indicating a contribution degree of word matching between a target retrieval result and the question; obtaining the first ranking result according to the first modified score and the second modified score of each candidate retrieval result.
11. The method of claim 9, wherein, The screening at least one target retrieval result from the plurality of candidate retrieval results according to the first ranking result to obtain the target multi-modal information includes: calculating a plurality of relevancies of the target object's question and answer context and the plurality of candidate retrieval results respectively by using a cross-encoder model, the question and answer context including the question and the target object's historical question and answer content; re-ranking the first ranking result based on the plurality of relevancies to obtain a second ranking result of the plurality of candidate retrieval results; screening at least one target retrieval result from the plurality of candidate retrieval results according to the second ranking result to obtain the target multi-modal information.
12. The method of claim 1, wherein, The preset model includes a first preset model and a second preset model, and the processing the target multi-modal information based on the question by using the preset model to output a reply to the question includes: processing the target multi-modal information based on the question by using the first preset model to obtain a first reply; processing the target multi-modal information based on the question by using the second preset model to obtain a second reply, the second preset model being different from the first preset model in at least one of inference accuracy, inference speed, and calculation resource consumption; obtaining the reply based on a quality score result of the first reply and the second reply.
13. The method of claim 1, wherein, The retrieving target multi-modal information based on the intent by using the at least one target index includes: obtaining a question and answer context of the target object and a feedback result of the target object to at least one historical reply in the question and answer context, the question and answer context including the question and the target object's historical question and answer content, and the feedback result including positive feedback and negative feedback; performing rewriting or expansion on the question according to at least one of the question and answer context, the feedback result of the at least one historical reply, keywords in the question, semantic information represented by the question, and the intent, to retrieve the target multi-modal information based on the rewritten or expanded question.
14. The method of claim 13, wherein, The method further includes: evaluate the reply based on the feedback result of the question-answer context and the at least one historical reply; in response to an evaluation result indicating that the reply does not meet a preset condition, re-retrieve candidate multi-modal information based on the rewritten or expanded question through the at least one target index.
15. The method of claim 13, wherein, The method further comprises: evaluating at least one of an intent accuracy and a content completeness of the intent, the intent accuracy indicating an accuracy degree of the intent of the target object expressed by the question, and the content completeness indicating a completeness degree of the intent of the target object expressed by the question content; in response to the question satisfying at least one of the intent accuracy being less than a first threshold value and the content completeness being less than a second threshold value, performing rewriting or expansion on the question, so that an intent accuracy of the rewritten or expanded question is greater than the first threshold value or a content completeness is greater than the second threshold value.
16. The method of claim 13, wherein, The method further comprises: obtaining a feedback result of the target object on the reply; in response to the feedback result of the reply being negative feedback, re-performing rewriting or expansion on the question.
17. A question and answer apparatus, characterized by comprise: an index module, configured to, in response to receiving a question from a target object, determine at least one target index matching an intent expressed by the question and available computing resources for performing retrieval from a plurality of indexes according to the intent and the available computing resources, wherein the plurality of indexes comprise data index structures of different index precision levels, obtained by processing multi-modal information at a plurality of index precisions respectively, and the multi-modal information comprises structured representations of a plurality of modal data obtained by processing a target file; a retrieval module, configured to retrieve target multi-modal information comprising a plurality of candidate modal data and respective metadata indicating description information of corresponding modal data matching the intent through the at least one target index based on the intent; a generation module, configured to process the target multi-modal information based on the question by using a preset model to output a reply for the intent; wherein the determining at least one target index matching the intent and the available computing resources for performing retrieval from a plurality of indexes according to the intent and the available computing resources comprises: determining a target index precision satisfying the intent according to a question complexity indicated by the intent; determining at least one target index of each modal data in the multi-modal data to be obtained based on constraint restrictions of the target index precision and the available computing resources.
18. An electronic device comprising: one or more processors; a memory for storing one or more computer programs, characterized in that the one or more processors execute the one or more computer programs to implement the steps of the method according to any one of claims 1-16.
19. A computer readable storage medium having stored thereon a computer program, characterized in that, The computer program is executed by the processor to implement the steps of the method according to any one of claims 1-16. The computer program is executed by the processor to implement the steps of the method according to any one of claims 1-16.
Citation Information
Patent Citations
Semantic retrieval model fusion method and system based on adaptive weight
CN117076598A
Intelligent question answering method and device based on multi-modal information processing, electronic equipment and storage medium
CN119988563A