Data processing method, apparatus, device, and medium
By dividing documents into fragments and vectorizing them, a document fragment vector library is constructed, which solves the matching accuracy and flexibility problems of traditional question-answering systems and realizes the automated construction of intelligent question-answering systems and the generation of high-quality responses.
Patent Information
- Application Number
- CN202411898053.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-20
- Publication Date
- 2026-08-25
- Estimated Expiration
- 2044-12-20
AI Technical Summary
Traditional question-answering systems rely on keyword matching, which results in poor accuracy and flexibility when faced with diverse user questions. Furthermore, manual maintenance of the question-answering database limits the coverage of answers, and expanding and updating the corpus is a huge undertaking.
The document is segmented into document fragments, and a document fragment vector library is built using vectorized document fragments. User query requests are retrieved from the document fragment vector library to generate response information, thereby realizing the automated construction of the question-answering system corpus.
It improves the accuracy and flexibility of the question-and-answer system, enhances the matching accuracy and quality of the generated response information, and is suitable for intelligent question-and-answer systems for multimodal documents.
Smart Images

Figure CN119807396B_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to the field of artificial intelligence technology, particularly to the fields of natural language processing, deep learning, and large models, and specifically to a data processing method, a data processing device, an electronic device, a computer-readable storage medium, and a computer program product. Background Technology
[0002] Artificial intelligence (AI) is the study of enabling computers to simulate certain human thought processes and intelligent behaviors (such as learning, reasoning, thinking, and planning). It encompasses both hardware and software technologies. AI hardware technologies generally include sensors, dedicated AI chips, cloud computing, distributed storage, and big data processing. AI software technologies mainly include natural language processing, computer vision, speech recognition, machine learning / deep learning, big data processing, and knowledge graph technologies.
[0003] Intelligent question answering is an important research direction in the field of artificial intelligence, involving the deep integration of natural language processing, knowledge retrieval, and generation technologies. It is widely used in customer service, education, and healthcare, gradually overcoming the technical bottlenecks of traditional systems in semantic understanding, multimodal data processing, and personalized responses by accurately understanding user intent and providing efficient answers.
[0004] The methods described in this section are not necessarily methods that had been previously conceived or adopted. Unless otherwise specified, no method described in this section should be assumed to be prior art simply because it is included in this section. Similarly, unless otherwise specified, the issues mentioned in this section should not be considered to be accepted in any prior art. Summary of the Invention
[0005] This disclosure provides a data processing method, a data processing apparatus, an electronic device, a computer-readable storage medium, and a computer program product.
[0006] According to one aspect of this disclosure, a data processing method is provided, comprising: acquiring a user query request and a document fragment vector library, wherein the document fragment vector library is obtained using a construction operation, the construction operation comprising: slicing a document to obtain multiple document fragments; vectorizing the multiple document fragments to obtain multiple fragment vectors; storing the multiple fragment vectors into the document fragment vector library; retrieving from the document fragment vector library based on the user query request to obtain at least one target vector among the multiple fragment vectors; and generating response information corresponding to the user query request based on at least one target fragment corresponding to the at least one target vector.
[0007] According to another aspect of this disclosure, a data processing apparatus is provided, comprising: a first acquisition unit configured to acquire a user query request and a document fragment vector library, wherein the document fragment vector library is obtained using a construction operation, the construction operation including: slicing a document to obtain multiple document fragments; vectorizing the multiple document fragments to obtain multiple fragment vectors; and storing the multiple fragment vectors into the document fragment vector library; a retrieval unit configured to perform a retrieval in the document fragment vector library based on the user query request to obtain at least one target vector among the multiple fragment vectors; and a generation unit configured to generate response information corresponding to the user query request based on at least one target fragment corresponding to the at least one target vector.
[0008] According to another aspect of this disclosure, an electronic device is provided, comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor to enable the at least one processor to perform the methods described above.
[0009] According to another aspect of this disclosure, a non-transitory computer-readable storage medium is provided storing computer instructions, wherein the computer instructions are used to cause a computer to perform the above-described method.
[0010] According to another aspect of this disclosure, a computer program product is provided, including a computer program, wherein the computer program implements the above-described method when executed by a processor.
[0011] According to one or more embodiments of this disclosure, a document is segmented into document fragments, and a document fragment vector library is constructed using vectorized document fragments. Then, based on a user query request, a retrieval is performed in the document fragment vector library to obtain the target fragment related to the user query request and generate response information. This realizes the automated construction of the corpus of the question-and-answer system, improves the accuracy and flexibility of the question-and-answer system, and enhances the matching degree and quality of the generated response information.
[0012] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of this disclosure, nor is it intended to limit the scope of this disclosure. Other features of this disclosure will become readily apparent from the following description. Attached Figure Description
[0013] The accompanying drawings exemplify embodiments and form part of the specification, serving together with the textual description to explain exemplary implementations of the embodiments. The illustrated embodiments are for illustrative purposes only and do not limit the scope of the claims. Throughout the drawings, the same reference numerals refer to similar but not necessarily identical elements.
[0014] Figure 1 A schematic diagram of an exemplary system in which the various methods described herein may be implemented according to embodiments of the present disclosure is shown;
[0015] Figure 2 A flowchart of a data processing method according to an embodiment of the present disclosure is shown;
[0016] Figure 3 A flowchart illustrating vectorization of multiple document fragments according to an embodiment of the present disclosure is shown;
[0017] Figure 4 A flowchart illustrating the conversion of non-textual information in a first multimodal segment into second textual information according to an embodiment of the present disclosure is shown;
[0018] Figure 5 A flowchart illustrating a training operation according to an embodiment of the present disclosure is shown;
[0019] Figure 6 A flowchart illustrating the generation of response information corresponding to a user query request according to an embodiment of the present disclosure is shown;
[0020] Figure 7 A flowchart illustrating the storage of multiple fragment vectors into a document fragment vector library according to an embodiment of the present disclosure is shown;
[0021] Figure 8 A flowchart illustrating the generation of response information corresponding to a user query request according to an embodiment of the present disclosure is shown;
[0022] Figure 9 A flowchart illustrating a retrieval in a document fragment vector library according to an embodiment of the present disclosure is shown;
[0023] Figure 10 A schematic diagram illustrating document processing and document fragmentation vector library construction according to exemplary embodiments of the present disclosure is shown;
[0024] Figure 11 A schematic diagram of a question-and-answer process according to an exemplary embodiment of the present disclosure is shown;
[0025] Figure 12 A structural block diagram of a data processing apparatus according to embodiments of the present disclosure is shown; and
[0026] Figure 13 A structural block diagram of an exemplary electronic device that can be used to implement embodiments of the present disclosure is shown. Detailed Implementation
[0027] The exemplary embodiments of this disclosure are described below with reference to the accompanying drawings, including various details of the embodiments to aid understanding, and should be considered merely exemplary. Therefore, those skilled in the art will recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope of this disclosure. Similarly, for clarity and brevity, descriptions of well-known functions and structures are omitted in the following description.
[0028] In this disclosure, unless otherwise stated, the use of terms such as "first," "second," etc., to describe various elements is not intended to limit the positional, temporal, or importance relationships of these elements; such terms are merely used to distinguish one element from another. In some examples, the first element and the second element may refer to the same instance of that element, while in other cases, based on the context, they may refer to different instances.
[0029] The terminology used in the description of the various examples described in this disclosure is for the purpose of describing particular examples only and is not intended to be limiting. Unless the context clearly indicates otherwise, an element may be one or more unless the number of elements is specifically limited. Furthermore, the term "and / or" as used in this disclosure covers any one of the listed items and all possible combinations thereof.
[0030] In related technologies, traditional question-answering systems mainly rely on keyword matching. When faced with diverse user question formats, the matching accuracy and flexibility are poor, often failing to meet user expectations. In addition, existing systems often rely on manual maintenance of the question-answer database, resulting in limited coverage of answers, and the workload of expanding and updating the corpus is enormous.
[0031] To address the aforementioned issues, this disclosure divides documents into document fragments and constructs a document fragment vector library using vectorized document fragments. Then, based on user query requests, it retrieves target fragments related to the user query requests and generates response information. This achieves automated construction of the question-and-answer system corpus, improves the accuracy and flexibility of the question-and-answer system, and enhances the matching degree and quality of the generated response information.
[0032] The embodiments of this disclosure will now be described in detail with reference to the accompanying drawings.
[0033] Figure 1 A schematic diagram of an exemplary system 100 in which the various methods and apparatus described herein can be implemented according to embodiments of this disclosure is shown. Reference Figure 1The system 100 includes one or more client devices 101, 102, 103, 104, 105 and 106, a server 120, and one or more communication networks 110 coupling the one or more client devices to the server 120. The client devices 101, 102, 103, 104, 105 and 106 can be configured to execute one or more applications.
[0034] In embodiments of this disclosure, server 120 may run one or more services or software applications that enable the execution of the methods of this disclosure.
[0035] In some embodiments, server 120 may also provide other services or software applications, which may include non-virtual and virtual environments. In some embodiments, these services may be provided as web-based services or cloud services, such as to users of client devices 101, 102, 103, 104, 105, and / or 106 under a Software as a Service (SaaS) model.
[0036] exist Figure 1 In the configuration shown, server 120 may include one or more components that implement the functions performed by server 120. These components may include software components, hardware components, or combinations thereof that can be executed by one or more processors. Users operating client devices 101, 102, 103, 104, 105, and / or 106 can sequentially interact with server 120 using one or more client applications to utilize the services provided by these components. It should be understood that various different system configurations are possible and may differ from system 100. Therefore, Figure 1 This is an example of a system used to implement the various methods described herein, and is not intended to be limiting.
[0037] Users can use client devices 101, 102, 103, 104, 105, and / or 106 for human-computer interaction. The client devices provide interfaces that enable users to interact with them. The client devices can also output information to the user through these interfaces. Although... Figure 1 Only six client devices are described, but those skilled in the art will understand that this disclosure can support any number of client devices.
[0038] Client devices 101, 102, 103, 104, 105, and / or 106 may include various types of computer devices, such as portable handheld devices, general-purpose computers (such as personal computers and laptops), workstation computers, wearable devices, smart screen devices, self-service terminal devices, service robots, gaming systems, thin clients, various messaging devices, sensors, or other sensing devices. These computer devices can run various types and versions of software applications and operating systems, such as Microsoft Windows, Apple iOS, UNIX-like operating systems, Linux or Linux-like operating systems (such as Google Chrome OS); or include various mobile operating systems, such as Microsoft Windows Mobile OS, iOS, Windows Phone, and Android. Portable handheld devices may include cellular phones, smartphones, tablets, personal digital assistants (PDAs), etc. Wearable devices may include head-mounted displays (such as smart glasses) and other devices. Gaming systems may include various handheld gaming devices, internet-enabled gaming devices, etc. Client devices are capable of executing various applications, such as various internet-related applications, communication applications (such as email applications), short message service (SMS) applications, and can use various communication protocols.
[0039] Network 110 can be any type of network well known to those skilled in the art, and can use any of a variety of available protocols (including but not limited to TCP / IP, SNA, IPX, etc.) to support data communication. By way of example only, one or more networks 110 can be a local area network (LAN), an Ethernet-based network, a token ring network, a wide area network (WAN), the Internet, a virtual network, a virtual private network (VPN), an intranet, an extranet, a blockchain network, a public switched telephone network (PSTN), an infrared network, a wireless network (e.g., Bluetooth, WIFI), and / or any combination of these and / or other networks.
[0040] Server 120 may include one or more general-purpose computers, special-purpose server computers (e.g., PC (personal computer) servers, UNIX servers, mid-range servers), blade servers, mainframe computers, server clusters, or any other suitable arrangement and / or combination. Server 120 may include one or more virtual machines running a virtual operating system, or other computing architectures involving virtualization (e.g., one or more flexible pools of logical storage devices that can be virtualized to maintain virtual storage devices for servers). In various embodiments, server 120 may run one or more services or software applications that provide the functionality described below.
[0041] The computing unit in server 120 can run one or more operating systems, including any of the aforementioned operating systems and any commercially available server operating system. Server 120 can also run any of a variety of additional server applications and / or middleware applications, including HTTP servers, FTP servers, CGI servers, JAVA servers, database servers, etc.
[0042] In some implementations, server 120 may include one or more applications to analyze and merge data feeds and / or event updates received from users of client devices 101, 102, 103, 104, 105 and / or 106. Server 120 may also include one or more applications to display data feeds and / or real-time events via one or more display devices of client devices 101, 102, 103, 104, 105 and / or 106.
[0043] In some implementations, server 120 can be a server for a distributed system or a server integrated with blockchain. Server 120 can also be a cloud server, or an intelligent cloud computing server or intelligent cloud host with artificial intelligence technology. A cloud server is a host product in the cloud computing service system, designed to address the shortcomings of traditional physical hosts and Virtual Private Server (VPS) services, such as high management difficulty and weak business scalability.
[0044] System 100 may also include one or more databases 130. In some embodiments, these databases may be used to store data and other information. For example, one or more of the databases 130 may be used to store information such as audio files and video files. Databases 130 may reside in various locations. For example, a database used by server 120 may be local to server 120, or it may be located away from server 120 and may communicate with server 120 via a network-based or dedicated connection. Databases 130 may be of different types. In some embodiments, the database used by server 120 may be, for example, a relational database. One or more of these databases may store, update, and retrieve data from and from the databases in response to commands.
[0045] In some embodiments, one or more of the databases 130 may also be used by an application to store application data. The databases used by the application may be of different types, such as key-value stores, object stores, or regular stores supported by a file system.
[0046] Figure 1The system 100 can be configured and operated in various ways to enable the application of the various methods and apparatus described in this disclosure.
[0047] According to one aspect of this disclosure, a data processing method is provided. For example... Figure 2 As shown, the construction operation of the document fragment vector library includes: step S201, slicing the document to obtain multiple document fragments; step S202, vectorizing the multiple document fragments to obtain multiple fragment vectors; and step S203, storing the multiple fragment vectors into the document fragment vector library. The data processing method includes: step S204, obtaining the user query request and the document fragment vector library, wherein the document fragment vector library is obtained using the above construction operation; step S205, searching the document fragment vector library based on the user query request to obtain at least one target vector among the multiple fragment vectors; and step S206, generating response information corresponding to the user query request based on at least one target fragment corresponding to at least one target vector.
[0048] Therefore, by segmenting documents into document fragments and constructing a document fragment vector library using vectorized document fragments, and then retrieving the target fragments related to the user's query request from the document fragment vector library based on the user's query request, and generating response information, the corpus of the question-answering system is automatically constructed, improving the accuracy and flexibility of the question-answering system, and increasing the completeness of the generated response information and the matching degree between the response information and the user's query request.
[0049] The data processing method provided in this disclosure can be used in intelligent question-answering systems targeting specific vertical categories. In one exemplary embodiment, the specific vertical category can be the document data of a department within a company. The intelligent question-answering system can answer questions based on a large number of standardized and non-standardized documents from that department, and these document data contain a wealth of relevant information. Effectively integrating these documents into the intelligent question-answering system will greatly improve the accuracy and breadth of the questions and answers.
[0050] In some embodiments, the document fragmentation vector library may be constructed using a construction operation prior to performing the data processing method described above. The documents involved in this disclosure may be plain text documents or multimodal documents including text, images, tables, charts, or other modal content information.
[0051] In practical applications, vertical domain question-answering systems involve various document types, including structured data, policy documents, Word documents, PPTs, PDFs, and web pages. This heterogeneous data is difficult to process effectively using a single model. Furthermore, the corpus contains a large amount of multimodal content such as tables and images, requiring the system to process and convey this information during the response. In addition, some business scenarios have extremely high requirements for information accuracy and cannot tolerate "illusion" in the answers.
[0052] In some embodiments, before step S201, different types of original documents can be preprocessed and converted into image format. Then, a layout analysis model is used to process the document, identifying and labeling the positions and corner coordinates of content blocks such as body text, title text, images, tables, and charts, as the basis for subsequent slicing. In some embodiments, useless information such as headers and footers can be filtered out to prevent noise from being introduced into the subsequent slicing.
[0053] In some embodiments, in step S201, the text and figures can be read and segmented from the document in sequence. During the segmentation process, in order to prevent splitting the same continuous text or splitting images and image titles into different segments, the segmentation process can be performed by combining the above layout analysis results and coordinates.
[0054] In step S201, the document can be divided into multiple document segments according to a preset segmentation method. In some embodiments, the preset segmentation method may include semantic-based segmentation, paragraph-based segmentation, and heading-based segmentation, or other segmentation methods, which are not limited here. Heading-based segmentation may include dividing the document content between two adjacent headings into the same document segment, and optionally, segmentation may be combined with periods and line breaks. The resulting document segments may include the earlier heading from the two headings.
[0055] In step S202, appropriate vectorization tools can be used to vectorize document fragments as needed. In some embodiments, word embeddings can be used for vectorization of plain text content. For multimodal document fragments, tools that support different modalities (e.g., feature extraction models, embedding models, etc.) can be used to vectorize information of different modalities in the document fragments.
[0056] According to some embodiments, multiple document fragments may include a first multimodal fragment, which may include first text information and non-text information.
[0057] In some embodiments, the first text information may include title text, body text, captions, figure captions, figure titles, table headers, etc., in the first multimodal segment. Non-text information may include information from other modalities, such as at least one of image information, table information, and chart information.
[0058] like Figure 3 As shown, step S202, vectorizing multiple document fragments to obtain multiple fragment vectors, may include: step S301, converting non-text information in the first multimodal fragment into second text information; step S302, embedding words into the first combined text including the first text information and the second text information to obtain fragment vectors corresponding to the first multimodal fragment.
[0059] Therefore, by using the above method, the segment vector can simultaneously contain semantic information from both the original text information and non-text information, thereby effectively preserving the semantics of multimodal text and improving the subsequent retrieval effect. This ensures that the retrieved target segments not only match the user's query request in terms of text semantics, but also match the user's query request in terms of the semantics of non-text modalities such as images and tables.
[0060] In some embodiments, the first text information and the second text information can be combined according to the positional relationship between the first text information and the non-text information to obtain the first combined text.
[0061] According to some embodiments, such as Figure 4 As shown, step S301, converting non-text information in the first multimodal segment into second text information, may include: step S401, extracting intermediate text information from the non-text information; step S402, obtaining information extraction prompt text corresponding to the modality of the non-text information, the information extraction prompt text instructing the first large model to summarize the non-text information according to the modality of the non-text information, the first large model being a large language model; and step S403, inputting the information extraction prompt text and intermediate text information into the first large model to obtain the second text information.
[0062] Therefore, by first extracting intermediate text information from non-textual information, and then instructing the first major model to generate summary content for non-textual information by referring to the corresponding modalities, the semantic content in non-textual information can be fully explored, improving the semantic richness of the obtained fragment vectors. This, in turn, can improve the subsequent retrieval effect, so as to obtain more accurate target fragments with stronger relevance to the user's query request, thereby generating response information with a higher degree of matching with the user's query request.
[0063] According to some embodiments, step S401, extracting intermediate text information from non-text information, may include: calling a second large model to perform cross-modal content understanding on the non-text information to obtain intermediate text information. The second large model is a cross-modal large model that supports the modality of non-text information.
[0064] Therefore, by using the above methods, we can further leverage the capabilities of cross-modal large models to perform content understanding on non-textual information and further mine the semantic content within non-textual information.
[0065] According to some embodiments, step S401, extracting intermediate text information from non-text information, may include at least one of the following: performing optical character recognition on the image in the non-text information to obtain first intermediate text information; and processing the table in the non-text information using a table extraction model to obtain second intermediate text information. The intermediate text information may include the first intermediate text information and / or the second intermediate text information.
[0066] Therefore, the above methods can be used to conveniently and directly obtain text information from non-text information.
[0067] It is understood that the various methods mentioned above for extracting intermediate text information can be used individually or in combination. Other methods besides those described above can also be used to extract intermediate text information, and these are not limited here.
[0068] In some embodiments, information extraction prompt texts corresponding to multiple modalities (e.g., images, tables, charts, etc.) can be pre-set. These prompt texts instruct the large model to generate summary content by referencing the corresponding modal information. An example prompt text corresponding to a table modality might include:
[0069] Please generate a summary of the table content below, with the following requirements…
[0070] This disclosure does not limit the specific content of the information extraction prompt text; any prompt text that can instruct the corresponding modality to generate a summary content for the large model reference is within the scope of this disclosure.
[0071] In step S203, the fragment vectors corresponding to each document fragment can be stored in a document fragment vector library for subsequent retrieval. In some embodiments, the mapping relationship between document fragments and / or document fragments and their corresponding fragment vectors can be stored in the document fragment vector library. When a fragment vector is retrieved subsequently, the document fragment can be directly obtained from the document fragment vector library, or the mapping relationship can be obtained from it and then the document fragment can be obtained through other means.
[0072] The construction operations described above enable intelligent processing of complex documents, better accommodating diverse file types and organizational formats. This allows for effective slicing and vectorization of structured and unstructured questions and answers in real-world scenarios, forming searchable structured data. Furthermore, the slicing and vectorization process fully utilizes the semantic information of non-textual modalities (such as images and tables) to improve retrieval and response effectiveness, and achieves recall of complex content that includes both textual and non-textual information.
[0073] In step S204, the user query request and document fragment vector library are obtained.
[0074] In some embodiments, a user query request may include text input by the user obtained by the intelligent question-answering system, or other forms of query requests obtained through other means.
[0075] In step S205, a search is performed in the document fragment vector library based on the user's query request to obtain at least one target vector from multiple fragment vectors.
[0076] A user query request can be a query text targeting a specific vertical category. In one exemplary embodiment, a user query request can be the referral reward details for a specific job level.
[0077] In some embodiments, vector retrieval recall can be used to determine at least one target vector in a document fragment vector library that is most relevant to the user's query request.
[0078] According to some embodiments, in the construction operation, the vectorization in step S202 can be implemented using a word embedding model. Step S205, retrieving at least one target vector from a document fragment vector library based on the user query request, may include: vectorizing the user query request using a word embedding model to obtain a query request vector; and calculating the similarity between the query request vector and the fragment vectors in the document fragment vector library.
[0079] Therefore, by using a word embedding model to vectorize document segments and user query requests, and then calculating the similarity between the query request vector and the segment vector, the target vector most relevant to the user query request can be accurately obtained, thereby determining the document segment that matches the user query request.
[0080] In some embodiments, the similarity can be, for example, the cosine similarity between two vectors, or it can be calculated using other similarity measures, which are not limited here.
[0081] According to some embodiments, such as Figure 5As shown, the data processing method may further include: step S501, obtaining sample user query requests and multiple sample document fragments, and obtaining the annotation ranking of the multiple sample document fragments, wherein the annotation ranking indicates the degree of real closeness between the multiple sample document fragments and the sample user query requests; step S502, using the word embedding model to be trained to vectorize the sample user query requests and multiple sample document fragments to obtain sample query request vectors and multiple sample fragment vectors; step S503, calculating the similarity between the sample query request vectors and multiple sample fragment vectors; and step S504, adjusting the parameters of the word embedding model to be trained based on the similarity and annotation ranking.
[0082] Understandable Figure 5 The operation and effects of steps S505-S510 can be referred to the above text. Figure 2 The descriptions of steps S201-S206 are not repeated here.
[0083] Compared to conventional word embedding models used for general text, the word embedding model obtained through the above training operations can provide stronger targeting for question answering systems of specific vertical knowledge bases, thereby improving the effectiveness of the obtained fragment vectors or query request vectors, and thus improving the accuracy of downstream retrieval.
[0084] In some embodiments, in step S501, the annotation order of multiple sample document fragments can be manually annotated or determined by other means.
[0085] In some embodiments, in step S504, multiple sets of positive and negative examples can be determined from multiple sample segment vectors based on the annotation ranking relationship, and the parameters of the word embedding model to be trained can be adjusted using the corresponding loss function. Specifically, the sample segment vector with higher similarity to the sample query request vector can be used as a positive example, and the sample segment vector with lower similarity to the sample query request vector can be used as a negative example.
[0086] Back Figure 2 In step S206, based on at least one target fragment corresponding to at least one target vector, response information corresponding to the user's query request is generated.
[0087] In some embodiments, if a single target vector can be retrieved in step S205, then the target segment corresponding to that target vector can be used as the generated response information in step S206. In some embodiments, at least one target segment includes multiple target segments, and these target segments can be combined as the generated response information. Alternatively, these target segments can be further processed to obtain more refined and targeted response information.
[0088] According to some embodiments, at least one target vector may include multiple target vectors, and at least one target slice may include multiple target slices corresponding to the multiple target vectors. For example... Figure 6 As shown, step S206, generating response information corresponding to the user query request based on at least one target segment corresponding to at least one target vector, may include: step S601, inputting the first document segment corresponding to the first target vector with the highest similarity to the user query request among multiple target vectors into the third large model to determine whether the first document segment covers the user query request; and step S602, in response to determining that the first document segment does not cover the user query request, inputting multiple target segments into the third large model to obtain response information, wherein the response information is generated by the third large model by summarizing multiple target segments.
[0089] By using the above methods, we can fully leverage the understanding capabilities of large models to generate response information that ensures coverage of user query requests.
[0090] In some embodiments, the third large model can be a large language model or a cross-modal large model. In step S601, the user query request, the first document fragment, and the corresponding prompt text can be input into the third large model. The corresponding prompt text instructs the third large model to determine whether the first document fragment covers the user query request. In this way, the understanding ability of the large model can be fully utilized to determine whether the first document fragment covers the user query request.
[0091] In some embodiments, a document sharding vector library can use a plain text-based database. Such a database is convenient for storing vectors, offers greater flexibility, and enables faster retrieval. Furthermore, such a database facilitates the storage of permission tags for access control, as will be discussed below. In one exemplary embodiment, the document sharding vector library can be built using Elasticsearch (ES).
[0092] According to some embodiments, such as Figure 7 As shown, step S202, storing multiple fragment vectors into the document fragment vector library, may include: step S701, generating a first identifier for non-text information; step S702, using the first identifier as an index, storing the non-text information into a non-text database; step S703, combining the first text information and the first identifier to obtain a second combined text; and step S704, storing the mapping relationship between the fragment vectors of the first multimodal fragment and the second combined text into the document fragment vector library.
[0093] Therefore, the above method can convert multimodal document fragments into plain text information. By generating a first identifier for non-text information and combining it with the first text information, the original structure of the document fragments can be preserved, and the length of the second combined text can be effectively reduced, saving storage resources. Furthermore, by using the first identifier as an index for storing non-text information, efficient access to non-text information can be achieved.
[0094] In some embodiments, in step S701, a first identifier can be generated by performing a hash operation on the non-text information. The first identifiers of different non-text information need to be guaranteed to be unique.
[0095] When the large model selects a complete fragment for response, if the response contains non-textual information (e.g., an image, table, or other modal identifier), the corresponding non-textual information will be retrieved from the non-textual database based on the identifier and replaced, ultimately displayed to the user in rich text format. In some embodiments, to facilitate direct access to the original policy document, the mapping relationship between the original document address and document fragments can be stored in a document fragment vector library. This allows the original file address to be retrieved while presenting the document fragments and included in the reference section of the response information.
[0096] In some embodiments, the first text information and the first identifier can be combined according to the positional relationship between the first text information and the non-text information to obtain the second combined text. This method not only preserves the original structure of the document fragments but also retains the positional relationship between different information, thereby achieving accurate restoration in subsequent steps.
[0097] According to some embodiments, at least one target slice may include a second multimodal slice. For example... Figure 8 As shown, step S206, generating response information corresponding to the user query request based on at least one target fragment corresponding to at least one target vector, may include: step S801, obtaining the second combined text corresponding to the fragment vector of the second multimodal fragment from the document fragment vector library based on the mapping relationship; step S802, disassembling the second combined text corresponding to the second multimodal fragment to obtain the second identifier of the non-text information in the second multimodal fragment; step S803, using the second identifier to look up the non-text information in the second multimodal fragment in the non-text database; and step S804, generating the response information using the non-text information in the second multimodal fragment.
[0098] Therefore, by using the above method, the original content of the document fragments can be restored during the response information generation stage, so that the first text information and non-text information in the second multimodal fragment can be treated as a whole response, providing the user with complete document content.
[0099] The question-answering system provided in this disclosure can also support access control. According to some embodiments, the construction operation may further include: obtaining access permission tags for a document; and storing the access permission tags of the document as access permission tags for multiple document fragments into a document fragment vector library.
[0100] Therefore, by using the above method, the access permissions of each document fragment can be determined and stored, thereby realizing the permission management of document fragments.
[0101] In one exemplary embodiment, access permission labels may include: workplace, whether access is limited to a specific department, level requirements, etc.
[0102] According to some embodiments, such as Figure 9 As shown, step S205, which involves searching the document fragment vector library based on the user's query request to obtain at least one target vector from multiple fragment vectors, may include: step S901, obtaining user access permissions; step S902, filtering out document fragments without access permissions in the document fragment vector library based on the user's access permissions; and step S903, searching among the unfiltered document fragments based on the user's query request.
[0103] Therefore, by using the above method, it is possible to filter out document fragments that do not have access permissions during the retrieval stage, thereby ensuring that the generated response content is consistent with the user's access permissions.
[0104] In some embodiments, user access permissions may include workplace, whether a user belongs to a specific department, user level, etc.
[0105] The above description of steps S204-S206 and their sub-steps demonstrates a weighted question-answering system based on retrieval enhancement. In these steps, multi-step orchestration generates or recalls results directly, meeting the specific requirement in business scenarios that answers must faithfully reproduce the original text. Furthermore, considering the often strict access control over policy documents within companies, rule-based tag calculation is incorporated to support the implementation of complex rule-based personnel access control requirements.
[0106] Figure 10This diagram illustrates document processing and document fragmentation vector library construction according to exemplary embodiments of the present disclosure. The original document 1002 may include various formats such as PDF, WORD, PPT, and images. Document processing 1004 includes three modules: a preprocessing module 1006, a layout analysis model 1008, and a fragmentation module 1010, to process the original document 1002 sequentially. The processing results can then be used for vector calculation and storage 1012. Specifically, the document fragments can be decomposed to obtain first text information 1020, image information 1022, and table information 1024. Image information 1022 and table information 1024 can be stored in a database 1014. Image information 1022 and table information 1024 can be subjected to OCR recognition 1028 and table recognition 1030 respectively, and the recognition results, combined with different prompt texts 1034, are input into a large model 1032 to obtain second text information 1026. The first text information 1020 and the second text information 1026 are combined and then input into the embedding model 1016 for vectorization. Finally, the result is stored in the document fragment vector library 1018.
[0107] Figure 11 A schematic diagram of a question-and-answer process according to an exemplary embodiment of this disclosure is shown. User Query 1102 inputs into the embedded model 1112, obtaining a corresponding semantic vector 1114. User identity information 1104 is used for permission rule system tag calculation 1106, resulting in a tag group 1108. Then, based on the tag group 1108, the access permission tags of document fragments in the document fragment vector library 1110, and the semantic vector 1114, knowledge retrieval 1122 is performed. This process includes permission tag filtering 1116, vector retrieval 1118, and retrieval sorting 1120, resulting in multiple document fragments 1124. Subsequently, a large model generation or retrieval output 1126 can be performed to obtain response information 1132. The large model generation or retrieval output 1126 can be implemented using prompt word engineering 1128 and a large model 1130.
[0108] According to another aspect of this disclosure, a data processing apparatus is provided. For example... Figure 12 As shown, the apparatus 1200 includes: a first acquisition unit 1210 configured to acquire a user query request and a document fragment vector library, wherein the document fragment vector library is obtained using a construction operation, the construction operation including: slicing a document to obtain multiple document fragments; vectorizing the multiple document fragments to obtain multiple fragment vectors; and storing the multiple fragment vectors into the document fragment vector library; a retrieval unit 1220 configured to perform a retrieval in the document fragment vector library based on the user query request to obtain at least one target vector among the multiple fragment vectors; and a generation unit 1230 configured to generate response information corresponding to the user query request based on at least one target fragment corresponding to at least one target vector.
[0109] It is understandable that the operation and effects of units 1210-1230 in device 1200 can be referred to the above description. Figure 2 The descriptions of steps S204-S206 are not repeated here.
[0110] According to some embodiments, multiple document fragments may include a first multimodal fragment. The first multimodal fragment may include first text information and non-text information. Vectorizing multiple document fragments to obtain multiple fragment vectors may include: converting the non-text information in the first multimodal fragment into second text information; and embedding words into a first combined text including the first text information and the second text information to obtain a fragment vector corresponding to the first multimodal fragment.
[0111] According to some embodiments, converting non-textual information in a first multimodal segment into second textual information may include: extracting intermediate textual information from the non-textual information; obtaining information extraction prompt text corresponding to the modality of the non-textual information, wherein the information extraction prompt text instructs a first large model to generate a summary of the non-textual information based on the modality of the non-textual information, and the first large model is a large language model; and inputting the information extraction prompt text and intermediate textual information into the first large model to obtain the second textual information.
[0112] According to some embodiments, extracting intermediate text information from non-textual information may include: calling a second large model to perform cross-modal content understanding on the non-textual information to obtain intermediate text information, wherein the second large model is a cross-modal large model that supports the modality of non-textual information.
[0113] According to some embodiments, extracting intermediate text information from non-text information may include at least one of the following: performing optical character recognition on an image in non-text information to obtain first intermediate text information; and processing a table in non-text information using a table extraction model to obtain second intermediate text information.
[0114] According to some embodiments, storing multiple fragment vectors into a document fragment vector library may include: generating a first identifier for non-text information; using the first identifier as an index to store the non-text information into a non-text database; combining the first text information and the first identifier to obtain a second combined text; and storing the mapping relationship between the fragment vectors of the first multimodal fragment and the second combined text into the document fragment vector library.
[0115] According to some embodiments, at least one target fragment may include a second multimodal fragment. The generation unit may include: a first acquisition subunit configured to acquire a second combined text corresponding to the fragment vector of the second multimodal fragment from a document fragment vector library based on a mapping relationship; a disassembly subunit configured to disassemble the second combined text corresponding to the second multimodal fragment to obtain a second identifier for non-textual information in the second multimodal fragment; a lookup subunit configured to use the second identifier to look up non-textual information in the second multimodal fragment in a non-text database; and a first generation subunit configured to generate response information using the non-textual information in the second multimodal fragment.
[0116] According to some embodiments, non-textual information may include at least one of image information, table information, and chart information.
[0117] According to some embodiments, the build operation may further include: obtaining the access permission tags of the document; and storing the access permission tags of the document as access permission tags for multiple document fragments into a document fragment vector library.
[0118] According to some embodiments, the retrieval unit may include: a second acquisition subunit configured to acquire user access permissions; a filtering subunit configured to filter out document fragments that do not have access permissions in the document fragment vector library based on user access permissions; and a retrieval subunit configured to perform retrieval in the unfiltered document fragments based on user query requests.
[0119] According to some embodiments, vectorization can be implemented using a word embedding model. The retrieval unit may include: a vectorization subunit configured to vectorize a user query request using a word embedding model to obtain a query request vector; and a calculation subunit configured to calculate the similarity between the query request vector and fragment vectors in a document fragment vector library.
[0120] According to some embodiments, the data processing apparatus may further include: a second acquisition unit configured to acquire sample user query requests and multiple sample document fragments, and acquire the annotation ranking of the multiple sample document fragments, wherein the annotation ranking indicates the degree of real resemblance between the multiple sample document fragments and the sample user query requests; a vectorization unit configured to vectorize the sample user query requests and multiple sample document fragments using a word embedding model to be trained, to obtain sample query request vectors and multiple sample fragment vectors; a calculation unit configured to calculate the similarity between the sample query request vectors and the multiple sample fragment vectors; and a parameter tuning unit configured to adjust the parameters of the word embedding model to be trained based on the similarity and annotation ranking.
[0121] According to some embodiments, at least one target vector may include multiple target vectors, and at least one target slice may include multiple target slices corresponding to the multiple target vectors. The generation unit may include: a determining subunit configured to input a first document slice corresponding to a first target vector among the multiple target vectors that has the highest similarity to the user query request into a third large model to determine whether the first document slice covers the user query request; and a second generation subunit configured to, in response to determining that the first document slice does not cover the user query request, input multiple target slices into the third large model to obtain response information, wherein the response information is generated by the third large model by summarizing the multiple target slices.
[0122] The collection, storage, use, processing, transmission, provision, and disclosure of user personal information involved in the technical solution disclosed herein comply with the provisions of relevant laws and regulations and do not violate public order and good morals.
[0123] According to embodiments of this disclosure, an electronic device, a readable storage medium, and a computer program product are also provided.
[0124] refer to Figure 13 The present invention describes a structural block diagram of an electronic device 1300 that can serve as a server or client of the present disclosure, which is an example of a hardware device that can be applied to various aspects of the present disclosure. The electronic device is intended to represent various forms of digital electronic computer devices, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the present disclosure described and / or claimed herein.
[0125] like Figure 13 As shown, the electronic device 1300 includes a computing unit 1301, which can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 1302 or a computer program loaded from a storage unit 1308 into a random access memory (RAM) 1303. The RAM 1303 may also store various programs and data required for the operation of the electronic device 1300. The computing unit 1301, ROM 1302, and RAM 1303 are interconnected via a bus 1304. An input / output (I / O) interface 1305 is also connected to the bus 1304.
[0126] Multiple components in electronic device 1300 are connected to I / O interface 1305, including: input unit 1306, output unit 1307, storage unit 1308, and communication unit 1309. Input unit 1306 can be any type of device capable of inputting information to electronic device 1300. Input unit 1306 can receive input digital or character information and generate key signal inputs related to user settings and / or function control of the electronic device, and may include, but is not limited to, a mouse, keyboard, touchscreen, trackpad, trackball, joystick, microphone, and / or remote control. Output unit 1307 can be any type of device capable of presenting information, and may include, but is not limited to, a monitor, speaker, video / audio output terminal, vibrator, and / or printer. Storage unit 1308 may include, but is not limited to, a hard disk and an optical disk. The communication unit 1309 allows the electronic device 1300 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks, and may include, but is not limited to, modems, network cards, infrared communication devices, wireless communication transceivers and / or chipsets, such as Bluetooth devices, 802.11 devices, WiFi devices, WiMax devices, cellular communication devices and / or the like.
[0127] The computing unit 1301 can be a variety of general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 1301 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 1301 performs the various methods, processes, and / or processes described above. For example, in some embodiments, these methods, processes, and / or processes may be implemented as computer software programs tangibly contained in a machine-readable medium, such as storage unit 1308. In some embodiments, part or all of the computer program may be loaded and / or installed on the electronic device 1300 via ROM 1302 and / or communication unit 1309. When the computer program is loaded into RAM 1303 and executed by the computing unit 1301, one or more steps of the methods, processes, and / or processes described above may be performed. Alternatively, in other embodiments, the computing unit 1301 may be configured to perform these methods, processes, and / or handling by any other suitable means (e.g., by means of firmware).
[0128] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SoCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.
[0129] The program code used to implement the methods of this disclosure may be written in any combination of one or more programming languages. This program code may be provided to a processor or controller of a general-purpose computer, special-purpose computer, or other programmable data processing apparatus, such that when executed by the processor or controller, the program code causes the functions / operations specified in the flowcharts and / or block diagrams to be implemented. The program code may be executed entirely on a machine, partially on a machine, as a standalone software package partially on a machine and partially on a remote machine, or entirely on a remote machine or server.
[0130] In the context of this disclosure, a machine-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.
[0131] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device for displaying information to the user (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor); and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the computer. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).
[0132] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as data servers), or middleware components (e.g., application servers), or frontend components (e.g., user computers with graphical user interfaces or web browsers through which users can interact with implementations of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., communication networks). Examples of communication networks include local area networks (LANs), wide area networks (WANs), the Internet, and blockchain networks.
[0133] Computer systems can include clients and servers. Clients and servers are generally geographically separated and typically interact via communication networks. The client-server relationship is created by computer programs running on the respective computers and having a client-server relationship with each other. A server can be a cloud server, also known as a cloud computing server or cloud host, a hosting product within the cloud computing service system that addresses the management difficulties and weak business scalability inherent in traditional physical hosts and VPS (Virtual Private Server) services. Servers can also be servers for distributed systems or servers incorporating blockchain technology.
[0134] It should be understood that the various forms of processes shown above can be used to rearrange, add, or delete steps. For example, the steps described in this disclosure can be performed in parallel, sequentially, or in a different order, as long as the desired result of the technical solution disclosed in this disclosure can be achieved, and this is not limited herein.
[0135] While embodiments or examples of this disclosure have been described with reference to the accompanying drawings, it should be understood that the methods, systems, and devices described above are merely exemplary embodiments or examples, and the scope of the invention is not limited by these embodiments or examples, but only by the granted claims and their equivalents. Various elements in the embodiments or examples may be omitted or replaced by their equivalents. Furthermore, the steps may be performed in a different order than that described in this disclosure. Further, various elements in the embodiments or examples may be combined in various ways. Importantly, as the technology evolves, many elements described herein can be replaced by equivalents that appear after this disclosure.
Claims
1. A data processing method, comprising: Obtain user query requests and a document sharding vector library, wherein the document sharding vector library is obtained using a construction operation, the construction operation including: The document is processed using a layout analysis model to identify multiple content blocks and their coordinate information within the document; Based on the multiple content blocks and their coordinate information, the document is sliced to obtain multiple document slices. The multiple document slices include a first multimodal slice, which includes first text information and non-text information. The multiple document fragments are vectorized to obtain multiple fragment vectors, including: Converting non-text information in the first multimodal segment into second text information includes: Extract intermediate text information from the non-text information; From multiple preset information extraction prompt texts corresponding to multiple modalities, the information extraction prompt text corresponding to the non-text information modality is obtained. These multiple preset information extraction prompt texts are used to instruct a first large model to generate summary content based on the corresponding modality. The first large model is a large language model. The extracted prompt text and the intermediate text information are input into the first large model to obtain the second text information, wherein the second text information represents the semantic content of the non-textual modality; and Word embedding is performed on the first combined text, including the first text information and the second text information, to obtain a segmentation vector corresponding to the first multimodal segmentation; and Storing the multiple fragment vectors into the document fragment vector library includes: Perform a hash operation on the non-text information to generate a first identifier for the non-text information; Using the first identifier as an index, the non-text information is stored in a non-text database; The first text information and the first identifier are combined to obtain the second combined text; and The mapping relationship between the segmentation vector of the first multimodal segment and the second combined text is stored in the document segmentation vector library; Based on the user query request, a search is performed in the document fragmentation vector database to obtain at least one target vector from the plurality of fragmentation vectors; and Based on at least one target fragment corresponding to the at least one target vector, generate response information corresponding to the user query request.
2. The method according to claim 1, wherein, The step of extracting intermediate text information from the non-text information includes: The second major model is invoked to perform cross-modal content understanding on the non-textual information in order to obtain the intermediate textual information. The second major model is a cross-modal large model that supports the modality of the non-textual information.
3. The method according to claim 1, wherein, The extraction of intermediate text information from the non-text information includes at least one of the following: Optical character recognition is performed on the image in the non-text information to obtain the first intermediate text information; and The tables in the non-text information are processed using a table extraction model to obtain the second intermediate text information.
4. The method according to claim 1, wherein, The at least one target slice includes a second multimodal slice; The step of generating response information corresponding to the user query request based on at least one target fragment corresponding to the at least one target vector includes: Based on the mapping relationship, the second combined text corresponding to the segment vector of the second multimodal segment is obtained from the document segment vector library; Disassemble the second combined text corresponding to the second multimodal segment to obtain the second identifier of the non-text information in the second multimodal segment; Using the second identifier, retrieve the non-text information in the second multimodal fragment from the non-text database; and The response information is generated using the non-textual information in the second multimodal segment.
5. The method according to any one of claims 1-4, wherein, The non-text information includes at least one of image information, table information, and chart information.
6. The method according to any one of claims 1-4, wherein, The construction operation also includes: Obtain the access permission tags for the document; and The access permission tags of the document are stored in the document fragment vector library as access permission tags for the multiple document fragments.
7. The method according to claim 6, wherein, The step of retrieving at least one target vector from the plurality of fragment vectors based on the user query request in the document fragment vector database includes: Obtain user access permissions; Based on the user access permissions, filter out document fragments that do not have access permissions from the document fragment vector library; and Based on the user's query request, a search is performed on the unfiltered document fragments.
8. The method according to any one of claims 1-4, wherein, The vectorization is achieved using a word embedding model; The step of retrieving at least one target vector from the plurality of fragment vectors based on the user query request in the document fragment vector database includes: The user query request is vectorized using the word embedding model to obtain a query request vector; and Calculate the similarity between the query request vector and the fragment vectors in the document fragment vector library.
9. The method according to claim 8, further comprising: Obtain sample user query requests and multiple sample document fragments, and obtain the annotation ranking of the multiple sample document fragments, wherein the annotation ranking indicates the degree of resemblance between the multiple sample document fragments and the sample user query requests; The word embedding model to be trained is used to vectorize the sample user query request and the multiple sample document fragments to obtain the sample query request vector and the multiple sample fragment vectors. Calculate the similarity between the sample query request vector and the multiple sample shard vectors; as well as Based on the similarity and the label ranking, the parameters of the word embedding model to be trained are adjusted.
10. The method according to any one of claims 1-4, wherein, The at least one target vector includes multiple target vectors, and the at least one target slice includes multiple target slices corresponding to the multiple target vectors; The step of generating response information corresponding to the user query request based on at least one target fragment corresponding to the at least one target vector includes: The first document segment corresponding to the first target vector with the highest similarity to the user query request from among the multiple target vectors is input into the third large model to determine whether the first document segment covers the user query request; and In response to determining that the first document fragment does not cover the user query request, the multiple target fragments are input into the third large model to obtain the response information, wherein the response information is generated by the third large model by summarizing the multiple target fragments.
11. A data processing apparatus, comprising: The first acquisition unit is configured to acquire user query requests and a document shard vector library, wherein the document shard vector library is obtained using a construction operation, the construction operation including: The document is processed using a layout analysis model to identify multiple content blocks and their coordinate information within the document; Based on the multiple content blocks and their coordinate information, the document is sliced to obtain multiple document slices. The multiple document slices include a first multimodal slice, which includes first text information and non-text information. The multiple document fragments are vectorized to obtain multiple fragment vectors, including: Converting non-text information in the first multimodal segment into second text information includes: Extract intermediate text information from the non-text information; From multiple preset information extraction prompt texts corresponding to multiple modalities, the information extraction prompt text corresponding to the non-text information modality is obtained. These multiple preset information extraction prompt texts are used to instruct a first large model to generate summary content based on the corresponding modality. The first large model is a large language model. The extracted prompt text and the intermediate text information are input into the first large model to obtain the second text information, wherein the second text information represents the semantic content of the non-textual modality; and Word embedding is performed on the first combined text, including the first text information and the second text information, to obtain a segmentation vector corresponding to the first multimodal segmentation; and Storing the multiple fragment vectors into the document fragment vector library includes: Perform a hash operation on the non-text information to generate a first identifier for the non-text information; Using the first identifier as an index, the non-text information is stored in a non-text database; The first text information and the first identifier are combined to obtain the second combined text; and The mapping relationship between the segmentation vector of the first multimodal segment and the second combined text is stored in the document segmentation vector library; The retrieval unit is configured to perform a retrieval in the document fragment vector library based on the user query request, and obtain at least one target vector from the plurality of fragment vectors; and The generation unit is configured to generate response information corresponding to the user query request based on at least one target fragment corresponding to the at least one target vector.
12. The apparatus according to claim 11, wherein, The step of extracting intermediate text information from the non-text information includes: The second major model is invoked to perform cross-modal content understanding on the non-textual information in order to obtain the intermediate textual information. The second major model is a cross-modal large model that supports the modality of the non-textual information.
13. The apparatus according to claim 11, wherein, The extraction of intermediate text information from the non-text information includes at least one of the following: Optical character recognition is performed on the image in the non-text information to obtain the first intermediate text information; and The tables in the non-text information are processed using a table extraction model to obtain the second intermediate text information.
14. The apparatus according to claim 11, wherein, The at least one target fragment includes a second multimodal fragment, and the generation unit includes: The first acquisition subunit is configured to acquire, based on a mapping relationship, the second combined text corresponding to the fragment vector of the second multimodal fragment from the document fragment vector library; The disassembly subunit is configured to disassemble the second combined text corresponding to the second multimodal segment to obtain a second identifier of non-textual information in the second multimodal segment; The lookup subunit is configured to use the second identifier to look up non-text information in the second multimodal fragment in the non-text database; and The first generation subunit is configured to generate the response information using non-textual information in the second multimodal segment.
15. The apparatus according to any one of claims 11-14, wherein, The non-text information includes at least one of image information, table information, and chart information.
16. The apparatus according to any one of claims 11-14, wherein, The construction operation also includes: Obtain the access permission tags for the document; and The access permission tags of the document are stored in the document fragment vector library as access permission tags for the multiple document fragments.
17. The apparatus according to claim 16, wherein, The retrieval unit includes: The second acquisition subunit is configured to acquire user access permissions; A filtering subunit is configured to filter out document fragments from the document fragment vector library that the user does not have access to, based on the user's access permissions; and The retrieval subunit is configured to perform a retrieval in the unfiltered document fragments based on the user query request.
18. The apparatus according to any one of claims 11-14, wherein, The vectorization is achieved using a word embedding model, and the retrieval unit includes: The vectorization subunit is configured to vectorize the user query request using the word embedding model to obtain a query request vector; and The computation subunit is configured to calculate the similarity between the query request vector and the fragment vectors in the document fragment vector library.
19. The apparatus of claim 18, further comprising: The second acquisition unit is configured to acquire sample user query requests and multiple sample document fragments, and acquire the annotation ranking of the multiple sample document fragments, wherein the annotation ranking indicates the degree of real relevance of the multiple sample document fragments to the sample user query requests; The vectorization unit is configured to use the word embedding model to be trained to vectorize the sample user query request and the multiple sample document fragments to obtain sample query request vectors and multiple sample fragment vectors. The calculation unit is configured to calculate the similarity between the sample query request vector and the plurality of sample shard vectors; as well as The parameter tuning unit is configured to adjust the parameters of the word embedding model to be trained based on the similarity and the label ranking.
20. The apparatus according to any one of claims 11-14, wherein, The at least one target vector includes multiple target vectors, and the at least one target slice includes multiple target slices corresponding to the multiple target vectors. The generation unit includes: The determining subunit is configured to input the user query request and a first document slice corresponding to the first target vector among the plurality of target vectors that has the highest similarity to the user query request into a third large model to determine whether the first document slice covers the user query request; and The second generation subunit is configured to, in response to determining that the first document fragment does not cover the user query request, input the plurality of target fragments into the third large model to obtain the response information, wherein the response information is generated by the third large model by summarizing the plurality of target fragments.
21. An electronic device, characterized in that, The electronic device includes: At least one processor; and A memory communicatively connected to the at least one processor; wherein The memory stores instructions executable by the at least one processor, which, when executed by the at least one processor, enables the at least one processor to perform the method of any one of claims 1-10.
22. A non-transitory computer-readable storage medium storing computer instructions, characterized in that, The computer instructions are used to cause the computer to perform the method according to any one of claims 1-10.
23. A computer program product comprising a computer program, wherein, When the computer program is executed by a processor, it implements the method of any one of claims 1-10.
Citation Information
Patent Citations
Question and answer method and device, equipment and medium
CN118445395A
Image-text fusion question and answer method, device and equipment based on large language model
CN118861229A