Intelligent question and answer method and system for digital and electronic tickets based on large model FAQ extraction
Through the FAQ extraction technology based on the big model, the knowledge base of the digital ticket intelligent question-and-answer system is automatically constructed and optimized, which solves the problem that the existing intelligent customer service system relies on manual operations, and realizes efficient and automated knowledge base management and precise matching, reducing operational costs and improving customer satisfaction.
Patent Information
- Application Number
- CN202411939512.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-26
- Publication Date
- 2025-05-23
AI Technical Summary
The existing intelligent customer service system relies heavily on manual operations in question-and-answer dialogue setting and knowledge base maintenance and update, resulting in long business start-up time, high speech maintenance costs, poor human-computer interaction experience, and difficulty in closing the business process.
The intelligent question-and-answer method based on large-model FAQ extraction is adopted. By obtaining different types of source analysis files, using the large language model interface for FAQ extraction, obtaining question-and-answer pairs and embedding processing, it is stored in the Elasticsearch vector database, integrated into the knowledge base, and realizing automated knowledge base construction and optimization.
It improves the precise matching degree of the knowledge base, reduces the work burden of manual customer service, reduces operating costs, realizes efficient automation of intelligent customer service systems, and improves customer satisfaction and service quality.
Smart Images

Figure CN120030113A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of artificial intelligence technology, and more specifically, to an intelligent question-answering method and system for digital invoices based on large-model FAQ extraction. Background Art
[0002] The fully digitalized electronic invoice (abbreviated as "digital electronic invoice") is a new form of invoice with the same legal effect as traditional paper invoices. It is completely digital and can reduce the printing, storage and transportation of paper invoices, thereby improving the efficiency of invoice circulation, contributing to the digitization and intelligence of tax collection and management, and promoting the modernization of tax governance. Precisely because of its digital characteristics, since the pilot work of digital electronic invoices was carried out in 2021, all existing digital electronic invoice platforms on the market are required to provide services to taxpayers online 24 hours a day for free.
[0003] Therefore, the original online manual customer service can no longer meet the needs of the existing platform for 24-hour continuous service, and an intelligent question-and-answer system must be used to provide services around the clock. The digital ticket intelligent question-and-answer system is an important customer service tool for the digital platform. It can handle multiple users' requests at the same time without affecting the service quality due to excessive workload, reducing dependence on manpower and reducing the company's operating costs. The knowledge base is the core component of the intelligent question-and-answer system. It provides the system with the necessary information and data to generate accurate answers. A well-designed and properly maintained knowledge base is the key to its intelligence and efficiency.
[0004] However, at present, some links of the traditional intelligent customer service system, such as question-and-answer dialogue setting, knowledge base maintenance and update, have always been highly dependent on manual operations. For example, due to the lack of vertical field data accumulation, the cold start of the knowledge base is usually done by manual sorting of business problems, manually looking through various documents and historical chat records, precipitating initial questions and answers, and then sorting them into the form of "standard questions-similar questions" and importing them into the knowledge base. These manual operations are not only time-consuming and labor-intensive, but also have various limitations, resulting in long business startup time, high maintenance costs, poor human-computer interaction experience, and difficult closed-loop business processes.
[0005] Therefore, there is a need for an intelligent question-answering method for digital invoices based on large-model FAQ extraction. Summary of the invention
[0006] The present invention proposes a digital electronic invoice intelligent question-answering method and system based on large model FAQ extraction to solve the problem of how to efficiently realize the intelligent question-answering of digital electronic invoice questions.
[0007] In order to solve the above problems, according to one aspect of the present invention, a method for intelligent question answering of electronic invoices based on large model FAQ extraction is provided, the method comprising:
[0008] Acquiring different types of source analysis files for digital electrical analysis, and determining analysis document data based on the source analysis files;
[0009] Acquire the analysis document data using a large language model interface, and perform FAQ extraction based on the analysis document data to acquire at least one question-answer pair;
[0010] Embed the question and answer pairs in the imported file and store them in the Elasticsearch vector database, and integrate the results into the knowledge base;
[0011] Obtain questions related to electronic invoice counting input by the user, match question-answer pairs in the knowledge base based on the questions, and return the matched question-answer pairs to the user.
[0012] Preferably, the step of acquiring different types of source analysis files for digital electronic analysis and determining analysis document data based on the meta-analysis files comprises:
[0013] Various operation documents of the digital electricity system used for digital electricity analysis, historical chat records of digital electricity customer service, and policy documents related to digital electricity invoices are obtained as source analysis files, and Python scripts are used to collect and organize data from the source analysis files to obtain analysis document data in plain text format.
[0014] Preferably, the method further comprises:
[0015] Before forming a file to be imported in a preset format based on the acquired question-answer pairs, each question-answer pair is verified to verify whether each question-answer pair is compliant and legal and meets the conditions for being written into the knowledge base, and the question-answer pairs that are compliant and legal and meet the conditions for being written into the knowledge base are retained.
[0016] Preferably, the method further comprises:
[0017] The document analysis data is segmented and marked based on preset characters, and the document analysis data is segmented into multiple tokens through a word segmentation generator for use in the Embedding task.
[0018] Preferably, the method uses an Elasticsearch vector database to build a knowledge base to store and save the Embedding results of question-answer pairs, and uses msmarco-MiniLM-L12-cos-v5 as an Embedding model to represent a sentence or paragraph with a 384-dimensional vector.
[0019] According to another aspect of the present invention, there is provided an intelligent question-answering system for counting electronic invoices based on large model FAQ extraction, the system comprising:
[0020] A file analysis unit, used to obtain different types of source analysis files for digital electronic analysis, and determine analysis document data based on the source analysis files;
[0021] A question-answer pair acquisition unit, configured to acquire the analysis document data using a large language model interface, and to perform FAQ extraction based on the analysis document data to acquire at least one question-answer pair;
[0022] The knowledge base determination unit is used to perform embedding processing on the question-answer pairs in the imported file and store them in the Elasticsearch vector database, and integrate the results into the knowledge base;
[0023] The question-answering unit is used to obtain questions related to electronic invoices input by users, match question-answer pairs in the knowledge base based on the questions, and return the matched question-answer pairs to the user.
[0024] Preferably, the file analysis unit acquires different types of source analysis files for digital electronic analysis, and determines analysis document data based on the meta-analysis file, including:
[0025] Various operation documents of the digital electricity system used for digital electricity analysis, historical chat records of digital electricity customer service, and policy documents related to digital electricity invoices are obtained as source analysis files, and Python scripts are used to collect and organize data from the source analysis files to obtain analysis document data in plain text format.
[0026] Preferably, the system further comprises:
[0027] A verification unit is used to verify each question and answer pair before forming a file to be imported in a preset format based on the acquired question and answer pairs, to verify whether each question and answer pair is compliant and legal and meets the conditions for writing into the knowledge base, and to retain the question and answer pairs that are compliant and legal and meet the conditions for writing into the knowledge base.
[0028] Preferably, the system further comprises:
[0029] The word unit acquisition unit is used to segment and mark the document analysis data based on preset characters, and divide the document analysis data into multiple word units through a word segmentation generator for use in the Embedding task.
[0030] Preferably, the knowledge base determination unit uses the Elasticsearch vector database to build a knowledge base to store and save the Embedding results of the question-answer pairs, and uses msmarco-MiniLM-L12-cos-v5 as the Embedding model to represent a sentence or paragraph with a 384-dimensional vector.
[0031] The present invention provides a method and system for intelligent question and answer of digital electricity bills based on FAQ extraction of a large model, including: obtaining different types of source analysis files for digital electricity analysis, and determining analysis document data based on the source analysis files; obtaining the analysis document data using a large language model interface, and performing FAQ extraction based on the analysis document data to obtain at least one question and answer pair; performing embedding processing on the question and answer pairs in the imported file and storing them in the Elasticsearch vector database, and integrating the results into the knowledge base; obtaining questions related to digital electricity bills input by users, matching the question and answer pairs in the knowledge base based on the questions, and returning the matched question and answer pairs to the user. The method of the present invention introduces a large language model to provide high-quality intelligent question and answer services, assist in the construction and optimization of the knowledge base of the customer service system, improve the accurate matching degree of the knowledge base, and effectively improve the efficiency of customer service operation and maintenance. By automating the processing of common problems and repetitive tasks, the automatic construction of the intelligent customer service knowledge base can effectively reduce the workload of manual customer service and reduce operating costs. Enterprises can invest more resources in the processing of complex problems and the maintenance of high-value customers. Multi-round dialogues of intelligent customer service can provide fast and accurate responses and reduce customer waiting time. At the same time, multiple rounds of conversations can enable a deeper understanding of customer needs, provide personalized solutions, and significantly improve customer satisfaction. BRIEF DESCRIPTION OF THE DRAWINGS
[0032] A more complete understanding of exemplary embodiments of the present invention may be obtained by referring to the following drawings:
[0033] Figure 1 It is a flowchart of a method 100 for intelligent question and answering of electronic invoices based on large model FAQ extraction according to an embodiment of the present invention;
[0034] Figure 2 A flowchart of building a knowledge base according to an embodiment of the present invention;
[0035] Figure 3 It is a structural diagram of the electronic invoice intelligent question-answering system 300 based on large model FAQ extraction according to an embodiment of the present invention. DETAILED DESCRIPTION
[0036] Now, exemplary embodiments of the present invention are described with reference to the accompanying drawings. However, the present invention can be implemented in many different forms and is not limited to the embodiments described herein. These embodiments are provided to disclose the present invention in detail and completely and to fully convey the scope of the present invention to those skilled in the art. The terms used in the exemplary embodiments shown in the accompanying drawings are not intended to limit the present invention. In the accompanying drawings, the same units / elements are marked with the same reference numerals.
[0037] Unless otherwise specified, the terms (including technical terms) used herein have the commonly understood meanings to those skilled in the art. In addition, it is understood that the terms defined in commonly used dictionaries should be understood to have the same meanings as those in the context of the relevant fields, and should not be understood as idealized or overly formal meanings.
[0038] Big model technology is rapidly penetrating into all industries and fields around the world. With the deep integration with vertical industries, the application scenarios of big models are becoming increasingly diverse. With the large-scale development of "intelligent emergence", industry innovation may be just around the corner. As one of the earliest application scenarios of big models, intelligent question and answer is driving changes in customer service.
[0039] For the digital invoice business, which requires high accuracy of knowledge responses and extensive professional knowledge, it is often difficult to balance the maintenance efficiency and editing quality of the knowledge base. When faced with massive amounts of knowledge, it is inefficient to rely solely on manual sorting. The context analysis capabilities of the large model can easily achieve intelligent extraction of FAQs and automatic expansion of similar questions. Business personnel only need to review before they can use it, which not only improves maintenance efficiency but also takes into account the quality of the knowledge base.
[0040] Therefore, the present invention provides an intelligent question-answering method for digital invoices based on large model FAQ extraction, and utilizes a large language model to organize questions and answers into the form of a knowledge base to facilitate subsequent query and use, thereby enabling the intelligent customer service system to automatically and continuously learn and provide more accurate answers.
[0041] Figure 1 FIG. 1 is a flow chart of a method 100 for intelligent question answering of electronic invoices based on large model FAQ extraction according to an embodiment of the present invention. Figure 1 As shown, the digital electricity invoice intelligent question-answering method based on large model FAQ extraction provided by the embodiment of the present invention introduces a large language model to provide high-quality intelligent question-answering services, assists in the construction and optimization of the customer service system knowledge base, improves the accurate matching degree of the knowledge base, and effectively improves the efficiency of customer service operation and maintenance. By automating the processing of common problems and repetitive tasks, the automatic construction of the intelligent customer service knowledge base can effectively reduce the workload of manual customer service and reduce operating costs. Enterprises can invest more resources in the processing of complex problems and the maintenance of high-value customers. Intelligent customer service multi-round dialogue can provide fast and accurate responses and reduce customer waiting time. At the same time, multi-round dialogue can more deeply understand customer needs, provide personalized solutions, and significantly improve customer satisfaction. The digital electricity invoice intelligent question-answering method 100 based on large model FAQ extraction provided by the embodiment of the present invention starts from step 101. In step 101, different types of source analysis files for digital electricity analysis are obtained, and analysis document data is determined based on the source analysis files.
[0042] Preferably, the step of acquiring different types of source analysis files for digital electronic analysis and determining analysis document data based on the meta-analysis files comprises:
[0043] Various operation documents of the digital electricity system used for digital electricity analysis, historical chat records of digital electricity customer service, and policy documents related to digital electricity invoices are obtained as source analysis files, and Python scripts are used to collect and organize data from the source analysis files to obtain analysis document data in plain text format.
[0044] Combination Figure 2 As shown, in the present invention, the data source comes from various operation documents of the digital electricity system, historical chat records of digital electricity customer service, and policy documents related to digital electricity invoices. The system collects various operation documents of the digital electricity system, historical chat records of digital electricity customer service, and policy documents related to digital electricity invoices, etc., and uses python scripts to collect data from various types of files and organize them into plain text (txt) format, forming the content to be identified in the form of a text file.
[0045] In step 102, the analysis document data is obtained using a large language model interface, and FAQ extraction is performed based on the analysis document data to obtain at least one question-answer pair.
[0046] In step 103, embedding processing is performed on the question-answer pairs in the imported file and stored in the Elasticsearch vector database, and the results are integrated into the knowledge base.
[0047] Preferably, the method further comprises:
[0048] Before forming a file to be imported in a preset format based on the acquired question-answer pairs, each question-answer pair is verified to verify whether each question-answer pair is compliant and legal and meets the conditions for being written into the knowledge base, and the question-answer pairs that are compliant and legal and meet the conditions for being written into the knowledge base are retained.
[0049] Preferably, the method further comprises:
[0050] The document analysis data is segmented and marked based on preset characters, and the document analysis data is segmented into multiple tokens through a word segmentation generator for use in the Embedding task.
[0051] Preferably, the method uses an Elasticsearch vector database to build a knowledge base to store and save the Embedding results of question-answer pairs, and uses msmarco-MiniLM-L12-cos-v5 as an Embedding model to represent a sentence or paragraph with a 384-dimensional vector.
[0052] Combination Figure 2As shown, in the present invention, first, FAQ extraction is performed by calling the large language model interface to return the frequently asked questions and answers; then, it is checked whether the returned data is legal and compliant, and whether it meets the conditions for writing into the knowledge base; if the question and answer pairs are compliant and meet the conditions for writing into the knowledge base, they are retained; otherwise, a data error is prompted; then, an Excel format file to be imported is formed based on multiple question and answer pairs that meet the storage conditions; finally, the question and answer pairs in the file to be imported are embedded through the Elasticsearch vector database, and the results are stored in the knowledge base.
[0053] In the present invention, the text is also divided into word units through a word segmentation generator to ensure that each word unit has relatively complete and independent semantics for use in subsequent tasks such as Embedding.
[0054] The present invention uses large model technology to summarize, format, and deduplicate the original text, and extracts a number of independent, shorter knowledge points ("question-answer" pairs) from the original text as the most basic record of the question and answer, and matches it with the question. In particular, clear segmentation marks (such as punctuation, paragraphs, and chapters) can be designed in the original file, and coding formats such as html and markdown can be tried to improve the segmentation effect. If the granularity of the segmentation is too fine, the question and answer pairs will be fragmented and affect the relationship between them; if the granularity of the segmentation is too coarse, redundant information may be carried during matching, and it will also affect the efficiency of embedding, storage, and retrieval.
[0055] The present invention uses the Elasticsearch vector database to build a knowledge base for storing and saving the Embedding results of question-answer pairs. The solution uses msmarco-MiniLM-L12-cos-v5 as the Embedding model, which represents a sentence or paragraph with a 384-dimensional vector and is optimized for semantic search, providing real-time vector indexing, real-time vector update / deletion, K-nearest neighbor (KNN) search, range filtering and other functions.
[0056] In step 104, a question related to electronic invoice counting input by a user is obtained, a question-answer pair in the knowledge base is matched based on the question, and the matched question-answer pair is returned to the user.
[0057] In the present invention, after the knowledge base is constructed, when a question related to counting electric invoices input by a user is obtained, the question-answer pairs in the knowledge base are matched based on the question, and the matched question-answer pairs are returned to the user.
[0058] The method of the present invention uses a new artificial intelligence technology to replace manual assistance in the construction and optimization of the customer service system knowledge base, improves the accurate matching of the knowledge base, and effectively improves the efficiency of customer service operation and maintenance; uses large model technology to enhance the traditional single-round dialogue of intelligent customer service to multi-round dialogue, and handles customer inquiries through multi-round interactions. It can record and understand the context information of the dialogue and better understand customer intentions.
[0059] The key points of the present invention are:
[0060] In terms of technology, the large language model trained with massive corpus can conduct in-depth analysis of the text, understand the intent of the text in context, and automatically extract and process common questions, thereby quickly providing accurate answers.
[0061] In terms of service quality, the system can use large-model FAQ knowledge extraction technology to analyze new questions, unanswered questions, questions to be optimized, new policies, etc., continuously update the knowledge base content, help the intelligent question-answering system to achieve continuous learning and improvement, and improve service quality.
[0062] In terms of business operations, compared with traditional manual FAQ maintenance, the use of big model technology only requires one-time investment and daily operation and maintenance, which can reduce the operating costs of the business and enable enterprises to invest more resources in core business and innovation.
[0063] Figure 3 FIG. 3 is a schematic diagram of the structure of the digital invoice intelligent question-answering system 300 based on large model FAQ extraction according to an embodiment of the present invention. Figure 3 As shown, the digital invoice intelligent question and answer system 300 based on large model FAQ extraction provided in an embodiment of the present invention includes: a file analysis unit 301, a question and answer pair acquisition unit 302, a knowledge base determination unit 303 and a question and answer unit 304.
[0064] Preferably, the file analysis unit 301 is used to obtain different types of source analysis files for digital electronic analysis, and determine analysis document data based on the source analysis files.
[0065] Preferably, the file analysis unit 301 obtains different types of source analysis files for digital electronic analysis, and determines analysis document data based on the meta-analysis file, including:
[0066] Various operation documents of the digital electricity system used for digital electricity analysis, historical chat records of digital electricity customer service, and policy documents related to digital electricity invoices are obtained as source analysis files, and Python scripts are used to collect and organize data from the source analysis files to obtain analysis document data in plain text format.
[0067] Preferably, the question-answer pair acquisition unit 302 is used to acquire the analysis document data using a large language model interface, and perform FAQ extraction based on the analysis document data to acquire at least one question-answer pair.
[0068] Preferably, the knowledge base determination unit 303 is used to perform embedding processing on the question-answer pairs in the file to be imported and store them in the Elasticsearch vector database, and integrate the results into the knowledge base.
[0069] Preferably, the system further comprises:
[0070] A verification unit is used to verify each question and answer pair before forming a file to be imported in a preset format based on the acquired question and answer pairs, to verify whether each question and answer pair is compliant and legal and meets the conditions for writing into the knowledge base, and to retain the question and answer pairs that are compliant and legal and meet the conditions for writing into the knowledge base.
[0071] Preferably, the system further comprises:
[0072] The word unit acquisition unit is used to segment and mark the document analysis data based on preset characters, and divide the document analysis data into multiple word units through a word segmentation generator for use in the Embedding task.
[0073] Preferably, the knowledge base determination unit uses the Elasticsearch vector database to build a knowledge base to store and save the Embedding results of the question-answer pairs, and uses msmarco-MiniLM-L12-cos-v5 as the Embedding model to represent a sentence or paragraph with a 384-dimensional vector.
[0074] Preferably, the question-and-answer unit 304 is used to obtain questions related to electronic invoice counting input by the user, match question-and-answer pairs in the knowledge base based on the questions, and return the matched question-and-answer pairs to the user.
[0075] The intelligent question-answering system 300 for counting electronic invoices based on large model FAQ extraction in an embodiment of the present invention corresponds to the intelligent question-answering method 100 for counting electronic invoices based on large model FAQ extraction in another embodiment of the present invention, which will not be repeated here.
[0076] The invention has been described above with reference to a few embodiments. However, it is readily apparent to a person skilled in the art that other embodiments than the ones disclosed above are equally within the scope of the invention, as defined by the appended patent claims.
[0077] Generally, all terms used in the claims are to be interpreted according to their ordinary meaning in the technical field, unless explicitly defined otherwise therein. All references to "a / said / the [means, components, etc.]" are to be openly interpreted as at least one instance of said means, components, etc., unless explicitly stated otherwise. The steps of any method disclosed herein do not necessarily have to be performed in the exact order disclosed, unless explicitly stated otherwise.
[0078] Those skilled in the art will appreciate that embodiments of the present invention may be provided as methods, systems, or computer program products. Therefore, the present invention may take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware. Moreover, the present invention may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0079] The present invention is described with reference to flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to embodiments of the present invention. It should be understood that each process and / or block in the flowchart and / or block diagram, as well as the combination of processes and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 A process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.
[0080] These computer program instructions may also be stored in a computer-readable memory capable of directing a computer or other programmable data processing device to operate in a specific manner, so that the instructions stored in the computer-readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 A process or multiple processes and / or boxes Figure 1 A function specified in one or more boxes.
[0081] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operating steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing instructions for implementing the process. Figure 1 A process or multiple processes and / or boxes Figure 1 The steps for the functions specified in one or more boxes.
[0082] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention rather than to limit it. Although the present invention has been described in detail with reference to the above embodiments, ordinary technicians in the relevant field should understand that the specific implementation methods of the present invention can still be modified or replaced by equivalents. Any modification or equivalent replacement that does not depart from the spirit and scope of the present invention should be covered within the scope of protection of the claims of the present invention.
Claims
1. An intelligent question-answering method for counting electronic invoices based on large model FAQ extraction, characterized in that: The method comprises: Acquiring different types of source analysis files for digital electrical analysis, and determining analysis document data based on the source analysis files; Acquire the analysis document data using a large language model interface, and perform FAQ extraction based on the analysis document data to acquire at least one question-answer pair; Based on the acquired question-answer pairs, a file to be imported in a preset format is formed, the question-answer pairs in the file to be imported are embedded and stored in the Elasticsearch vector database, and the results are integrated into the knowledge base; Obtain questions related to electronic invoice counting input by the user, match question-answer pairs in the knowledge base based on the questions, and return the matched question-answer pairs to the user.
2. The method according to claim 1, characterized in that The acquiring of different types of source analysis files for digital electronic analysis and determining analysis document data based on the meta-analysis files includes: Various operation documents of the digital electricity system used for digital electricity analysis, historical chat records of digital electricity customer service, and policy documents related to digital electricity invoices are obtained as source analysis files, and Python scripts are used to collect and organize data from the source analysis files to obtain analysis document data in plain text format.
3. The method according to claim 1, characterized in that The method further comprises: Before forming a file to be imported in a preset format based on the acquired question-answer pairs, each question-answer pair is verified to verify whether each question-answer pair is compliant and legal and meets the conditions for being written into the knowledge base, and the question-answer pairs that are compliant and legal and meet the conditions for being written into the knowledge base are retained.
4. The method according to claim 1, characterized in that: The method further comprises: The document analysis data is segmented and marked based on preset characters, and the document analysis data is segmented into multiple tokens through a word segmentation generator for use in the Embedding task.
5. The method according to claim 1, characterized in that The method uses the Elasticsearch vector database to build a knowledge base to store and save the Embedding results of question-answer pairs, and uses msmarco-MiniLM-L12-cos-v5 as an Embedding model to represent a sentence or paragraph with a 384-dimensional vector.
6. An intelligent question-answering system for electronic invoices based on large model FAQ extraction, characterized in that: The system comprises: A file analysis unit, used to obtain different types of source analysis files for digital electronic analysis, and determine analysis document data based on the source analysis files; A question-answer pair acquisition unit, configured to acquire the analysis document data using a large language model interface, and to perform FAQ extraction based on the analysis document data to acquire at least one question-answer pair; The knowledge base determination unit is used to perform embedding processing on the question-answer pairs in the imported file and store them in the Elasticsearch vector database, and integrate the results into the knowledge base; The question-answering unit is used to obtain questions related to electronic invoices input by users, match question-answer pairs in the knowledge base based on the questions, and return the matched question-answer pairs to the user.
7. The system according to claim 6, characterized in that The file analysis unit obtains different types of source analysis files for digital electronic analysis and determines analysis document data based on the meta-analysis file, including: Various operation documents of the digital electricity system used for digital electricity analysis, historical chat records of digital electricity customer service, and policy documents related to digital electricity invoices are obtained as source analysis files, and Python scripts are used to collect and organize data from the source analysis files to obtain analysis document data in plain text format.
8. The system according to claim 6, characterized in that The system further comprises: A verification unit is used to verify each question and answer pair before forming a file to be imported in a preset format based on the acquired question and answer pairs, to verify whether each question and answer pair is compliant and legal and meets the conditions for writing into the knowledge base, and to retain the question and answer pairs that are compliant and legal and meet the conditions for writing into the knowledge base.
9. The system according to claim 6, characterized in that The system further comprises: The word unit acquisition unit is used to segment and mark the document analysis data based on preset characters, and divide the document analysis data into multiple word units through a word segmentation generator for use in the Embedding task.
10. The system according to claim 6, characterized in that The knowledge base determination unit uses the Elasticsearch vector database to build a knowledge base to store and save the Embedding results of the question-answer pairs, and uses msmarco-MiniLM-L12-cos-v5 as the Embedding model to represent a sentence or paragraph with a 384-dimensional vector.
Citation Information
Patent Citations
Customer service knowledge base expansion method, system and device based on large model
CN118245613A
Vector database retrieval method and device based on large model, terminal and medium
CN118312594A
Question and answer pair generation method and system based on large language model
CN118332086A