Document information extraction method and device, computer equipment and storage medium

By using semantic embedding vector database and document knowledge graph to obtain context fusion information, the existing insurance information processing methods are solved, and fast and accurate document information extraction is achieved.

CN119938898APending Publication Date: 2025-05-06CHINA PING AN PROPERTY INSURANCE CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510038693.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-09
Publication Date
2025-05-06

AI Technical Summary

Technical Problem

Existing insurance information processing methods rely on manual reading and analysis, are inefficient and error-prone.

Method used

By obtaining document query requests, the context fusion information is obtained using semantic embedding vector database and document knowledge graph, multiple related document information are obtained based on user information and context fusion information, and filtering them to obtain target document information.

Benefits of technology

It realizes the rapid and accurate acquisition of key information in the document, which significantly improves the efficiency of document information extraction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119938898A_ABST
    Figure CN119938898A_ABST
Patent Text Reader

Abstract

The invention discloses a document information extraction method and device, computer equipment and a storage medium, and the method comprises the steps: obtaining a document query request which comprises a query field and user information; acquiring context fusion information from a semantic embedded vector database and a document knowledge graph according to the query field; obtaining a plurality of related document information according to the user information and the context fusion information; and screening the multiple pieces of related document information to obtain target document information. According to the technical scheme, the context fusion information is obtained from the semantic embedded vector database and the document knowledge graph by utilizing the query field, then the multiple pieces of related document information are obtained according to the user information and the context fusion information, and finally the target document information is screened out from the multiple pieces of related document information. Therefore, the key information in the document can be quickly and accurately obtained, and the document information extraction efficiency is remarkably improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of artificial intelligence technology, and in particular to a document information extraction method, device, computer equipment and storage medium. Background Art

[0002] In the current insurance industry, especially in the field of consumer-oriented insurance products such as auto insurance, medical insurance, health insurance and pet insurance, improving customer experience and business operation efficiency have become key factors in corporate competition. With the continuous expansion of the insurance market, insurance companies have launched a variety of insurance products to meet the different needs of consumers.

[0003] However, insurance products usually involve complex terms and conditions. Most of this information exists in the form of unstructured documents, such as insurance contracts and claims guidelines. Extracting useful information from these documents and accurately interpreting them is crucial for the smooth operation of insurance business. Existing insurance information processing methods mainly rely on manual reading and analysis, which is not only inefficient but also prone to errors. Summary of the invention

[0004] The embodiments of the present invention provide a document information extraction method, device, computer equipment and storage medium to solve the problems of low efficiency and easy error in existing document information processing.

[0005] A document information extraction method, comprising: Obtaining a document query request, wherein the document query request includes a query field and user information; According to the query field, context fusion information is obtained from a semantic embedding vector database and a document knowledge graph; Acquire multiple related document information according to the user information and the context fusion information; The multiple related document information are screened to obtain target document information.

[0006] Furthermore, before acquiring context fusion information from a semantic embedding vector database and a document knowledge graph according to the query field, the document information extraction method includes: Get the original document; Extracting a semantic embedding vector of the original document; The semantic embedding vector is stored in a preset database, and the preset database in which the semantic embedding vector is stored is determined as the semantic embedding vector database.

[0007] Furthermore, extracting the semantic embedding vector of the original document includes: Acquire document structure information of the original document; According to the document structure information, the original document is divided into blocks to obtain document text blocks; Performing multimodal data analysis on the document text block to obtain document semantic information; According to the document semantic information, a semantic embedding vector of the original document is obtained.

[0008] Furthermore, the performing multimodal data analysis on the document text block to obtain document semantic information includes: Classifying the document text block to obtain the text block type; A semantic extraction strategy corresponding to the text block type is adopted to obtain global semantic information and local semantic information corresponding to the document text block, and the global semantic information and the local semantic information are determined as the document semantic information.

[0009] Furthermore, before acquiring context fusion information from a semantic embedding vector database and a document knowledge graph according to the query field, the document information extraction method includes: Initialize the knowledge graph structure; According to the preset entity relationship, the entity relationship data of the original document is obtained; According to the entity relationship data, the initialized knowledge graph structure is updated to obtain a document knowledge graph.

[0010] Furthermore, the acquiring of context fusion information from a semantic embedding vector database and a document knowledge graph according to the query field includes: Acquire first context information from the semantic embedding vector database according to the query field; Acquire second context information from the document knowledge graph according to the query field; Performing a correlation score on the first context information and the second context information to obtain a correlation result; According to the correlation result, obtaining a target fusion weight; The context fusion information is acquired according to the first context information, the second context information and the target fusion weight.

[0011] Furthermore, after acquiring context fusion information from a semantic embedding vector database and a document knowledge graph according to the query field, the document information extraction method includes: The target fusion weight is adjusted according to the query field and the user information.

[0012] A document information extraction device, comprising: A request acquisition module, used to acquire a document query request, wherein the document query request includes a query field and user information; An information fusion module, used to obtain context fusion information from a semantic embedding vector database and a document knowledge graph according to the query field; A document information module, used to obtain a plurality of related document information according to the user information and the context fusion information; The information screening module is used to screen the multiple related document information to obtain target document information.

[0013] A computer device comprises a memory, a processor and a computer program stored in the memory and executable on the processor, wherein the processor implements the document information extraction method when executing the computer program.

[0014] A computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the document information extraction method is implemented.

[0015] The document information extraction method, device, computer equipment and storage medium obtain a document query request, which includes a query field and user information. According to the query field, context fusion information is obtained from a semantic embedding vector database and a document knowledge graph. According to the user information and the context fusion information, multiple related document information is obtained, and the multiple related document information is screened to obtain target document information. By using the query field, context fusion information is obtained from a semantic embedding vector database and a document knowledge graph, and then according to the user information and the context fusion information, multiple related document information is obtained, and finally the target document information is screened from the multiple related document information, so that the key information in the document can be quickly and accurately obtained, thereby significantly improving the efficiency of document information extraction. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings required for use in the description of the embodiments of the present invention will be briefly introduced below. Obviously, the accompanying drawings in the following description are only some embodiments of the present invention. For ordinary technicians in this field, other accompanying drawings can be obtained based on these accompanying drawings without paying creative labor.

[0017] Figure 1 is a schematic diagram of an application environment of a document information extraction method in one embodiment of the present invention; Figure 2 is a flow chart of a document information extraction method in one embodiment of the present invention; Figure 3 is another flow chart of a method for extracting document information in one embodiment of the present invention; Figure 4 is another flow chart of a method for extracting document information in one embodiment of the present invention; Figure 5 is another flow chart of a method for extracting document information in one embodiment of the present invention; Figure 6 is another flow chart of a method for extracting document information in one embodiment of the present invention; Figure 7 is another flow chart of a method for extracting document information in one embodiment of the present invention; Figure 8 is a schematic diagram of a document information extraction device in one embodiment of the present invention; Fig. 9 is a schematic diagram of a computer device in one embodiment of the present invention. DETAILED DESCRIPTION

[0018] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0019] The document information extraction method provided by the embodiment of the present invention can be applied as follows: Figure 1 Specifically, the document information extraction method is applied in a document information extraction system, and the document information extraction system includes: Figure 1 The client and server shown in the figure communicate with each other through the network to realize document information extraction. The client, also known as the user end, refers to the program corresponding to the server and providing local services to the client. The client can be installed on but not limited to various personal computers, laptops, smart phones, tablets and portable wearable devices. The server can be implemented as an independent server or a server cluster consisting of multiple servers.

[0020] This embodiment provides a document information extraction method, which is applied in the above server, such as Figure 2 As shown, including: S201: Obtain a document query request, where the document query request includes query fields and user information.

[0021] S202: According to the query field, context fusion information is obtained from the semantic embedding vector database and the document knowledge graph.

[0022] S203: Acquire multiple related document information according to the user information and the context fusion information.

[0023] S204: Screening multiple related document information to obtain target document information.

[0024] Among them, the document query request refers to a request for querying document information. Exemplarily, the document may be an insurance document or a contract, etc. In this embodiment, an insurance document is used as an example. The query field is a field entered by the user and is used to query the corresponding document information. Exemplarily, the query field is such as "Help me interpret adult accident insurance" and "Help me compare insurance A and insurance B". User information includes user portrait, user age, user family, user historical query records, and user behavior analysis data.

[0025] Among them, the semantic embedding vector database is a database that stores the semantic embedding vectors of the original document. The original document is a document pre-stored in the document database. The document knowledge graph is a data structure that organizes and represents knowledge in a specific field in a graphical form. It converts the information in the original document into a series of interconnected entities, concepts and relationships. For example, the specific field is the insurance field, and the original document is an insurance document. Context fusion information refers to the context information that fuses the semantic embedding vector database and the document knowledge graph. Related document information refers to document information associated with the query field. Target document information refers to document information that matches the query field.

[0026] As an example, in step S201, the user forms a document query request through the client and sends the document query request to the server. After the server obtains the document query request, it parses the query field and user information from the document query request.

[0027] Furthermore, the server identifies the query type of the user's document query request based on the query field and determines the document query type. Exemplarily, the document query types include manual query, content query, and special query. The manual query is a query that needs to be transferred to a manual specialist. The content query is a document information query. The special query is a custom query type. For example, a renewal query in the insurance field, that is, querying the insurance that the user has purchased, including checking the validity period, determining whether the user meets the renewal conditions, etc.

[0028] It can be understood that if the document query type is a manual query, the manual specialist will be directly connected. If the document query type is a special query, the background query interface is called to query the insurance that the user has purchased, check the validity period, and determine whether the user meets the renewal conditions, etc., and give the user a voice answer through the intelligent question and answer system. Exemplarily, the intelligent question and answer system can adopt technologies known to those skilled in the art, which are not limited here. If the document query type is a content query, the corresponding target document information is obtained through steps S202 to S204 in this embodiment.

[0029] As an example, in step S202, context fusion information is obtained from the semantic embedding vector database and the document knowledge graph according to the query field. Exemplarily, according to the query field. A pre-trained machine learning model is used to obtain context information associated with the query field from the semantic embedding vector database and the document knowledge graph, respectively, and the context information corresponding to the semantic embedding vector database and the context information corresponding to the document knowledge graph are fused to obtain context fusion information. Exemplarily, the pre-trained machine learning model can be a model based on XGBoost or neural network training, which is used to obtain context information associated with the query field from the semantic embedding vector database and the document knowledge graph according to the query field, and the context information corresponding to the semantic embedding vector database and the context information corresponding to the document knowledge graph are fused. It should be noted that the method of training the machine learning model can adopt the well-known technology of those skilled in the art, which is not limited here.

[0030] As an example, in step S203, multiple relevant document information is obtained according to user information and context fusion information. Exemplarily, based on a large language model with context awareness, through the context management module of LangChain, multiple relevant document information is obtained according to user information and context fusion information through a preset diversity controller. Exemplarily, the large language module can be a BERT model or an ELMo model, etc. Specifically, a language model with context awareness is selected and configured into the LangChain framework, and a context management module is configured in LangChain for processing and storing user information and context fusion information. Exemplarily, the context management module analyzes in combination with user information and context fusion information to improve the relevance and personalization of the generated relevant document information. In this example, the context manager can dynamically adjust the generation process of relevant document information according to the query field to ensure that the generated relevant document information meets user needs and has high accuracy. LangChain's dynamic scheduling function ensures that the relevant document information can effectively integrate user information and context fusion information during the generation process, thereby improving the accuracy and relevance of the generated content. In this embodiment, the diversity controller combines the multi-round generation and screening mechanism of LangChain, and can dynamically adjust the diversity of the generated multiple related document information based on the query content, ensuring that the generated multiple related document information is both rich and consistent.

[0031] As an example, in step S204, multiple related document information is screened to obtain target document information. Exemplarily, screening multiple related document information includes cross-validation and logical consistency check to ensure that the target document information outputted in the end is of high quality and avoids self-contradictory or logically incorrect content.

[0032] In this embodiment, a document query request is obtained, and the document query request includes a query field and user information. According to the query field, context fusion information is obtained from the semantic embedding vector database and the document knowledge graph. According to the user information and the context fusion information, multiple related document information is obtained, and the multiple related document information is screened to obtain the target document information. By using the query field, context fusion information is obtained from the semantic embedding vector database and the document knowledge graph, and then according to the user information and the context fusion information, multiple related document information is obtained, and finally the target document information is screened from the multiple related document information, so that the key information in the document can be quickly and accurately obtained, thereby significantly improving the efficiency of document information extraction.

[0033] In one embodiment, if Figure 3 As shown, before step S202, before obtaining context fusion information from the semantic embedding vector database and the document knowledge graph according to the query field, the document information extraction method includes: S301: Acquire an original document.

[0034] S302: Extracting the semantic embedding vector of the original document.

[0035] S303: storing the semantic embedding vector into a preset database, and determining the preset database storing the semantic embedding vector as a semantic embedding vector database.

[0036] As an example, in step S301, the original document can be obtained through a preset document import module. Exemplarily, the preset document import module can be a PyPDFLoader module in the LangChain framework. It can be understood that the original document can include documents of different types. The number of the original documents can be selected according to actual needs. That is, a semantic embedding vector database corresponding to original documents of different types and quantities can be constructed according to actual needs.

[0037] As an example, in step S302, the semantic embedding vector of the original document is extracted. Exemplarily, the original document is semantically analyzed through natural language processing (NLP) technology, and the words or more complex language units in the original document are mapped to vectors in a high-dimensional space, so as to retain the contextual information between the words or language units in the original document through the language embedding vector. Among them, the language unit includes chapter titles and clause numbers, etc. Chapter titles are such as "insurance liability", "exclusion of liability", etc.

[0038] As an example, in step S303, the semantic embedding vector is stored in a preset database, and the preset database in which the semantic embedding vector is stored is determined as a semantic embedding vector database. The preset database is a pre-selected database for storing semantic embedding vectors. In this example, the semantic embedding vector is stored in the semantic embedding vector database so that the server can obtain context information associated with the query field from the semantic embedding vector database according to the query field and the pre-trained machine learning model.

[0039] In this embodiment, the original document is obtained, the semantic embedding vector of the original document is extracted, the semantic embedding vector is stored in a preset database, and the preset database in which the semantic embedding vector is stored is determined as a semantic embedding vector database, thereby constructing a semantic embedding vector database through the original document to support subsequent query field retrieval operations.

[0040] In one embodiment, if Figure 4 As shown, in step S302, the semantic embedding vector of the original document is extracted, including: S401: Obtain document structure information of the original document.

[0041] S402: Divide the original document into blocks according to the document structure information to obtain document text blocks.

[0042] S403: Perform multimodal data analysis on the document text block to obtain document semantic information.

[0043] S404: Obtain a semantic embedding vector of the original document according to the document semantic information.

[0044] The document structure information is the structure information of different language units in the document. The language unit includes text unit and chart unit. The text unit includes chapter, title, paragraph and sentence, etc. The chart unit includes picture and table.

[0045] As an example, in step S401, the document structure information of the original document is obtained. For example, the original document is generally in PDF format, and the original document is converted to text or HTML, and the document structure is identified by NLP technology or document analysis tools to obtain the document structure information. For example, the spaCy library and NLTK library in the NLP library are used to identify the document structure and obtain the document structure information.

[0046] As an example, in step S402, according to the document structure information, the original document is divided into blocks to obtain document text blocks. Exemplarily, by a preset text segmentation tool, according to the document structure information, by combining regular expressions and deep learning models, the original document is divided into blocks to obtain document text blocks. Specifically, the original document is preprocessed, including cleaning the text and standardizing the format. Cleaning the text includes removing unnecessary blank characters, special symbols, headers and footers, etc. The standardized format includes unified font size, style, etc., for subsequent processing. According to the document structure information, the original document is subjected to preliminary text segmentation using regular expressions. For the parts that are uncertain by the regular expression, a deep learning model is used for further analysis. It can be understood that the deep learning module can be pre-trained, such as collecting and annotating training data, including examples of the document structure and clause boundaries of the original document, and then selecting a suitable deep learning model, such as BiLSTM, BERT or Transformer, etc., and then training the model by using the annotated data.

[0047] As an example, in step S403, multimodal data analysis is performed on the document text block to obtain document semantic information. Among them, multimodal data refers to a data set composed of two or more different types of data. In this example, multimodal data analysis is performed on the document text block, that is, the text type in the document text block, such as text, image text, and table text, is analyzed to obtain document semantic information corresponding to different types of document text blocks to support multimodal semantic retrieval. In this example, LangChain's multimodal data processing capabilities can be used to fully cover the user's query needs and provide richer retrieval results.

[0048] Furthermore, after obtaining the document text block, the document semantic information of the document text block is further enhanced through the integrated Transformer architecture model. Specifically, the document text block is processed through tools such as LSTM or BERT models to ensure that the context information is fully retained when obtaining the semantic embedding vector of the original document.

[0049] As an example, in step S404, the semantic embedding vector of the original document is obtained according to the document semantic information. Exemplarily, the document semantic information is converted into a word index through an embedding model such as Word2Vec, GloVe, BERT or a custom model, and the word index is mapped to a high-dimensional vector in the embedding layer. The embedding model extracts text features through forward propagation of a multi-layer neural network, and uses a pooling layer to aggregate these text features to form a semantic embedding vector, which is normalized and reduced in dimension to be stored in a semantic embedding vector database for fast similarity search and document matching.

[0050] In this embodiment, the document structure information of the original document is obtained, the original document is divided into blocks according to the document structure information, the document text blocks are obtained, multimodal data analysis is performed on the document text blocks, the document semantic information is obtained, and the semantic embedding vector of the original document is obtained according to the document semantic information, so as to improve the multimodal retrieval capability, so that insurance practitioners can quickly and accurately obtain key information in the document, thereby significantly improving business processing efficiency, especially in complex scenarios such as claims review and risk assessment.

[0051] In one embodiment, if Figure 5 As shown, in step S403, multimodal data analysis is performed on the document text block to obtain document semantic information, including: S501: Classify the document text blocks and obtain the text block types.

[0052] S502: adopting a semantic extraction strategy corresponding to the text block type, obtaining global semantic information and local semantic information corresponding to the document text block, and determining the global semantic information and the local semantic information as document semantic information.

[0053] The text block type is the type of the document text block. Exemplarily, the text block type includes text, image text, and table text.

[0054] As an example, in step S501, a pre-trained text classification model is used to classify the document text block to obtain the text block type. Exemplarily, the text classification model includes naive Bayes, support vector machine (SVM), random forest, convolutional neural network (CNN), recurrent neural network (RNN), long short-term memory network (LSTM) or Transformer, etc. It can be understood that the training method of the text classification model can adopt the technology known to those skilled in the art, and is not limited here.

[0055] As an example, in step S501, a semantic extraction strategy corresponding to the text block type is adopted to obtain global semantic information and local semantic information corresponding to the document text block, and the global semantic information and local semantic information are determined as document semantic information. The semantic extraction strategy is a predefined strategy for obtaining global semantic information and local semantic information corresponding to the document text block according to the text block type.

[0056] Exemplarily, if the document text block is text, the global semantic information and local semantic information are directly generated through the embedding model. If the document text block is image text or table text, the image text or table text is converted into text through ORC technology or image processing module, and then the global semantic information and local semantic information are generated through the embedding model.

[0057] In this embodiment, the document text block is classified to obtain the text block type. A semantic extraction strategy corresponding to the text block type is adopted to obtain the global semantic information and local semantic information corresponding to the document text block, and the global semantic information and local semantic information are determined as document semantic information to achieve multi-level embedded vector retrieval and dynamic expansion of the knowledge graph, which can deeply analyze the complex relationship in the original document, and the generated target document information not only comprehensively covers the key information of the document, but also deeply mines the relationship and meaning behind it.

[0058] In one embodiment, if Figure 6 As shown, before step S202, before obtaining context fusion information from the semantic embedding vector database and the document knowledge graph according to the query field, the document information extraction method includes: S601: Initialize the knowledge graph structure.

[0059] S602: Acquire entity relationship data of the original document according to the preset entity relationship.

[0060] S603: Update the initialized knowledge graph structure according to the entity relationship data to obtain the document knowledge graph.

[0061] Among them, the preset entity relationship is the relationship type between entities pre-defined in the process of building the knowledge graph.

[0062] As an example, in step S601, the knowledge graph structure is initialized. Exemplarily, the basic framework of the knowledge graph is determined, including predefined entity types and relationship types. An empty knowledge graph structure is created in a graph database (such as Neo4j), and the schema of nodes and edges is defined.

[0063] As an example, in step S602, entity relationship data of the original document is obtained according to the preset entity relationship. For example, key events in the business system are monitored through Webhooks or message queues. When events such as new insurance policy issuance or legal terms update are detected, NLP technology is used to analyze relevant documents to extract entities (such as customers, insurance companies) and their relationships (such as insurance coverage, contract terms).

[0064] As an example, in step S603, the initialized knowledge graph structure is updated according to the entity relationship data to obtain the document knowledge graph. Exemplarily, the server uses entities as nodes and preset entity relationships as edges, and dynamically adds or modifies them to a graph database (such as Neo4j). For example, if a new insurance policy is issued, the server creates a new node to represent the policy and connects it to related entities (such as customers, insurance products) with edges.

[0065] In this embodiment, the knowledge graph structure is initialized, and the entity relationship data of the original document is obtained according to the preset entity relationship. The initialized knowledge graph structure is updated according to the entity relationship data to obtain the document knowledge graph. Due to the real-time update of the document knowledge graph, the latest document changes can be reflected in a timely manner, providing users with real-time and effective target document information.

[0066] In one embodiment, if Figure 7 As shown, in step S202, context fusion information is obtained from the semantic embedding vector database and the document knowledge graph according to the query field, including: S701: Acquire first context information from a semantic embedding vector database according to a query field.

[0067] S702: Obtain second context information from the document knowledge graph according to the query field.

[0068] S703: Perform relevance scoring on the first context information and the second context information to obtain a relevance result.

[0069] S704: Obtain target fusion weight according to the correlation result.

[0070] S705: Acquire context fusion information according to the first context information, the second context information and the target fusion weight.

[0071] As an example, in step S701, first context information is obtained from a semantic embedding vector database according to a query field. Exemplarily, when a user submits a query, the server uses the query field to search in the semantic embedding vector database. Using the context management module of LangChain, the semantic embedding vectors most relevant to the query field are retrieved through the embedding model, which represent the context information of the original document. Example: If a user queries about "insurance claim process", the vector representation of the document fragment that best matches the topic will be retrieved.

[0072] As an example, in step S702, the second context information is obtained from the document knowledge graph according to the query field. Exemplarily, the server searches the document knowledge graph according to the query field. A graph database query language (such as Cypher) is used to find entities and relationships related to the query field, which constitute the second context information of the query. For example, entities related to "insurance claim process" such as "claim form", "review process" and the connections between them are retrieved.

[0073] As an example, in step S703, the first context information and the second context information are scored for relevance to obtain a relevance result. Exemplarily, the server uses a machine learning model (such as XGBoost or a neural network) to score the context information obtained from the vector retrieval and the knowledge graph retrieval. Exemplarily, the timeliness of the first context information and the second context information, the relevance to the query field, and the user's historical query records are combined to evaluate the scoring result of the first context information and the scoring result of the second context information. Example: For example, the document fragment obtained by vector retrieval is very relevant to the user's query field, while some entity information in the knowledge graph is slightly less important.

[0074] As an example, in step S704, the target fusion weight is obtained according to the correlation result. Exemplarily, based on the correlation result, the server dynamically assigns the weight of the first context information and the weight of the second context information to determine which context information is dominant in the final target document information. Exemplarily, the dynamic scheduling function of LangChain automatically adjusts the weight according to the correlation result to ensure that the generated target document information is both accurate and relevant.

[0075] As an example, in step S704, context fusion information is obtained based on the first context information, the second context information and the target fusion weight. Exemplarily, the server combines the first context information, the second context information and their corresponding weights, i.e., the target fusion weight, to generate context fusion information. The context fusion information is passed as input to a large language model based on context awareness, and through the context management module of LangChain, multiple related document information is obtained through a preset diversity controller based on user information and context fusion information to optimize the generation process of multiple related document information. Exemplarily, the server combines the specific content of the document fragment and the structured information in the knowledge graph to generate a comprehensive and personalized answer, explain the insurance claim process, and provide relevant entity links, such as "Please fill out this claim form and follow the following review process."

[0076] In this embodiment, first context information is obtained from a semantic embedding vector database according to the query field. The semantic embedding vector can capture the deep semantics of the document, ensuring that the retrieved first context information is highly relevant to the user's query intention; second context information is obtained from the document knowledge graph according to the query field. The information in the knowledge graph can help users better understand the background and context of the query subject, thereby providing deeper knowledge services; the first context information and the second context information are scored for relevance to obtain a relevance result; based on the relevance result, a target fusion weight is obtained; based on the first context information, the second context information and the target fusion weight, context fusion information is obtained. The fused context fusion information is used as the input of a large language model to generate more coherent, accurate and useful target document information.

[0077] In one embodiment, after acquiring context fusion information from a semantic embedding vector database and a document knowledge graph according to a query field, the document information extraction method includes: adjusting a target fusion weight according to the query field and user information.

[0078] As an example, according to the preset dynamic weight control mechanism, the weights of vector retrieval and knowledge graph retrieval are automatically adjusted according to factors such as the query type, context quality, and user behavior of the query field. Exemplarily, the weights of vector retrieval and knowledge graph retrieval are automatically adjusted through the policy control module of LangChain.

[0079] For example, when the user's question is "Help me compare Insurance A and Insurance B". The intent is recognized as a comparison intent, and then the corresponding entity recognition module is called to extract the names of Insurance A and Insurance B. The corresponding first context information and second context information are found from two data sources, the semantic embedding vector database and the document knowledge graph. After the relevance of the first context information and the second context information is scored using the machine learning model, the corresponding weights of the first context information and the second context information are given. The large language model adopts the corresponding data according to the corresponding weights of the first context information and the second context information (which can be written in PROMPT to remind the large language model which part is more important), and finally generates an answer, that is, the target document information.

[0080] In this embodiment, the target fusion weight is adjusted according to the query field and user information, and the retrieval strategy is dynamically optimized in combination with the real-time monitoring and feedback system to ensure that the best answer can be generated in different scenarios.

[0081] It should be understood that the order of execution of the steps in the above embodiment does not necessarily mean the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiment of the present invention.

[0082] In one embodiment, a document information extraction device is provided, which corresponds one-to-one to the document information extraction method in the above embodiment. Figure 8 As shown, the document information extraction device includes a request acquisition module, an information fusion module, a document information module and an information screening module. The detailed description of each functional module is as follows: A request acquisition module is used to acquire a document query request, which includes query fields and user information; The information fusion module is used to obtain context fusion information from the semantic embedding vector database and the document knowledge graph according to the query field; The document information module is used to obtain multiple related document information based on user information and context fusion information; The information screening module is used to screen multiple related document information and obtain target document information.

[0083] Furthermore, the document information extraction device also includes: A document acquisition module is used to acquire original documents; Vector extraction module, used to extract the semantic embedding vector of the original document; The vector storage module is used to store the semantic embedding vector into a preset database, and determine the preset database in which the semantic embedding vector is stored as the semantic embedding vector database.

[0084] Furthermore, the vector extraction module includes: A structure information unit, used to obtain document structure information of the original document; A block processing unit, used to perform block processing on the original document according to the document structure information to obtain document text blocks; A data analysis unit, used for performing multimodal data analysis on a document text block to obtain document semantic information; The vector acquisition unit is used to acquire the semantic embedding vector of the original document according to the document semantic information.

[0085] Furthermore, the data analysis unit comprises: The type acquisition subunit is used to classify the document text block and obtain the text block type; The extraction strategy subunit is used to adopt a semantic extraction strategy corresponding to the text block type, obtain global semantic information and local semantic information corresponding to the document text block, and determine the global semantic information and local semantic information as document semantic information.

[0086] Furthermore, the document information extraction device also includes: Initialization module, used to initialize the knowledge graph structure; A data acquisition module, used to acquire entity relationship data of the original document according to a preset entity relationship; The graph update module is used to update the initialized knowledge graph structure according to the entity relationship data and obtain the document knowledge graph.

[0087] Furthermore, the information fusion module includes: A first context unit, configured to obtain first context information from a semantic embedding vector database according to a query field; The second context unit, the information fusion module obtains the second context information from the document knowledge graph according to the query field; A scoring unit, used to score the relevance of the first context information and the second context information to obtain a relevance result; A weight acquisition unit, used for acquiring a target fusion weight according to a correlation result; The fusion information unit is used to obtain context fusion information according to the first context information, the second context information and the target fusion weight.

[0088] Furthermore, the document information extraction device also includes: The weight adjustment module is used to adjust the target fusion weight according to the query field and user information.

[0089] For the specific definition of the document information extraction device, please refer to the definition of the document information extraction method above, which will not be repeated here. Each module in the above-mentioned document information extraction device can be implemented in whole or in part by software, hardware and a combination thereof. The above-mentioned modules can be embedded in or independent of the processor in the computer device in the form of hardware, or can be stored in the memory of the computer device in the form of software, so that the processor can call and execute the operations corresponding to the above modules.

[0090] In one embodiment, a computer device is provided. The computer device may be a server, and its internal structure diagram may be as follows: Fig. 9 As shown. The computer device includes a processor, a memory, a network interface and a database connected through a system bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The database of the computer device is used for document information extraction. The network interface of the computer device is used to communicate with an external terminal through a network connection. When the computer program is executed by the processor, a document information extraction method is implemented.

[0091] In one embodiment, a computer device is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the document information extraction method in the above embodiment is implemented. To avoid repetition, it is not described here. Alternatively, when the processor executes the computer program, the functions of each module / unit in the document information extraction device in this embodiment are implemented. To avoid repetition, it is not described here.

[0092] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored, and when the computer program is executed by a processor, the document information extraction method in the above embodiment is implemented, which will not be described in detail here to avoid repetition. Alternatively, when the computer program is executed by a processor, the functions of each module / unit in the embodiment of the document information extraction device are implemented, which will not be described in detail here to avoid repetition.

[0093] Those of ordinary skill in the art can understand that all or part of the processes in the above-mentioned embodiment methods can be implemented by instructing the relevant hardware through a computer program, and the computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to memory, storage, database or other media used in the embodiments provided in this application may include non-volatile and / or volatile memory. Non-volatile memory may include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM) or flash memory. Volatile memory may include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in many forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link (Synchlink) DRAM (SLDRAM), memory bus (Rambus) direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM), etc.

[0094] Those skilled in the art can clearly understand that for the convenience and simplicity of description, only the division of the above-mentioned functional units and modules is used as an example. In actual applications, the above-mentioned functions can be distributed and completed by different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above.

[0095] The embodiments described above are only used to illustrate the technical solutions of the present invention, rather than to limit the same. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that the technical solutions described in the aforementioned embodiments may still be modified, or some of the technical features may be replaced by equivalents. Such modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present invention, and should all be included in the protection scope of the present invention.

Claims

1. A document information extraction method, characterized in that: include: Obtaining a document query request, wherein the document query request includes a query field and user information; According to the query field, context fusion information is obtained from a semantic embedding vector database and a document knowledge graph; Acquire multiple related document information according to the user information and the context fusion information; The multiple related document information are screened to obtain target document information.

2. The document information extraction method according to claim 1, characterized in that: Before acquiring context fusion information from a semantic embedding vector database and a document knowledge graph according to the query field, the document information extraction method includes: Get the original document; Extracting a semantic embedding vector of the original document; The semantic embedding vector is stored in a preset database, and the preset database in which the semantic embedding vector is stored is determined as the semantic embedding vector database.

3. The document information extraction method according to claim 2, characterized in that: The extracting the semantic embedding vector of the original document comprises: Acquire document structure information of the original document; According to the document structure information, the original document is divided into blocks to obtain document text blocks; Performing multimodal data analysis on the document text block to obtain document semantic information; According to the document semantic information, a semantic embedding vector of the original document is obtained.

4. The document information extraction method according to claim 3, characterized in that: The performing multimodal data analysis on the document text block to obtain document semantic information includes: Classifying the document text block to obtain the text block type; A semantic extraction strategy corresponding to the text block type is adopted to obtain global semantic information and local semantic information corresponding to the document text block, and the global semantic information and the local semantic information are determined as the document semantic information.

5. The document information extraction method according to claim 1, characterized in that: Before acquiring context fusion information from a semantic embedding vector database and a document knowledge graph according to the query field, the document information extraction method includes: Initialize the knowledge graph structure; According to the preset entity relationship, the entity relationship data of the original document is obtained; According to the entity relationship data, the initialized knowledge graph structure is updated to obtain a document knowledge graph.

6. The document information extraction method according to claim 1, characterized in that: The step of acquiring context fusion information from a semantic embedding vector database and a document knowledge graph according to the query field includes: Acquire first context information from the semantic embedding vector database according to the query field; Acquire second context information from the document knowledge graph according to the query field; Performing a correlation score on the first context information and the second context information to obtain a correlation result; According to the correlation result, obtaining a target fusion weight; The context fusion information is acquired according to the first context information, the second context information and the target fusion weight.

7. The document information extraction method according to claim 6, characterized in that: After acquiring context fusion information from a semantic embedding vector database and a document knowledge graph according to the query field, the document information extraction method includes: The target fusion weight is adjusted according to the query field and the user information.

8. A document information extraction device, characterized in that: include: A request acquisition module, used to acquire a document query request, wherein the document query request includes a query field and user information; An information fusion module, used to obtain context fusion information from a semantic embedding vector database and a document knowledge graph according to the query field; A document information module, used to obtain a plurality of related document information according to the user information and the context fusion information; The information screening module is used to screen the multiple related document information to obtain target document information.

9. A computer device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the computer program, the document information extraction method according to any one of claims 1 to 7 is implemented.

10. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the document information extraction method according to any one of claims 1 to 7 is implemented.