Intelligent customer service robot based on multi-modal data fusion and knowledge graph enhancement and construction method thereof

Through the methods of multimodal data fusion and knowledge graph enhancement, the problems of modal fragmentation, knowledge staticization and insufficient reasoning depth of large models in intelligent question-answering systems were solved, and real-time and accurate query services for competition information were realized.

CN120705264APending Publication Date: 2025-09-26Chinese People's Liberation Army Cyberspace Force Information Engineering University
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510797911.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-13
Publication Date
2025-09-26

AI Technical Summary

Technical Problem

Existing large models in intelligent question-answering systems suffer from modal splitting problems, knowledge staticization defects, and insufficient reasoning depth, which makes it impossible to effectively handle consulting needs for competition-related information.

Method used

Adopting the method of multimodal data fusion and knowledge graph enhancement, the PDF files are preprocessed through the large language model and Miner U tool to construct the knowledge graph and weaviate vector database. Combining the question classification model, knowledge graph query engine and large model interaction interface, the access links of data layer, model layer and application layer are realized, and the knowledge graph is dynamically updated and managed.

Benefits of technology

It achieves minute-level synchronization and dynamic optimization of competition information, improves the system's ability to deeply understand competition professional knowledge, and provides real-time, efficient and accurate information query services.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120705264A_ABST
    Figure CN120705264A_ABST
Patent Text Reader

Abstract

The invention provides an intelligent customer service robot based on multi-modal data fusion and knowledge graph enhancement and a construction method thereof. The method comprises the following steps: constructing a data layer of the intelligent customer service robot, wherein the data layer comprises a knowledge graph and a weavate vector database; constructing a model layer of the intelligent customer service robot, wherein the model layer comprises a question classification model, a knowledge graph query engine, a vector retrieval module and a large model interaction interface; constructing an application layer of the intelligent customer service robot, wherein the application layer comprises a Web end interaction interface, an API service and a management background; and constructing access links of the data layer, the model layer and the application layer. According to the method, structured query of the knowledge graph and semantic retrieval of the vector database are communicated, a'accurate matching + semantic generalization 'mixed question and answer mode is realized, and the answer quality of open questions is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of artificial intelligence technology, and in particular to an intelligent customer service robot based on multimodal data fusion and knowledge graph enhancement and a construction method thereof. Background Art

[0002] In recent years, with the popularity of various professional knowledge competitions, large models are used to build intelligent customer service robots to provide efficient and timely services, so as to quickly respond to contestants' consultation needs for competition-related information. However, as contestants' consultation needs for competition-related information continue to increase, large models have the following shortcomings in intelligent question answering: (1) Modal fragmentation problem. For example, there is a lack of unified representation for text parsing and image / table data, and heterogeneous data such as competition rules documents, historical competition videos, and scoring tables are stored in a scattered manner, resulting in knowledge fragmentation during retrieval; (2) Knowledge static defects. For example, new competition rules need to be manually re-blocked and vectorized, and minute-level knowledge synchronization cannot be achieved during the competition; (3) Insufficient reasoning depth. For example, for open innovation competition questions, the system can only output template solution ideas and lacks adaptive strategy optimization based on scoring criteria. Summary of the Invention

[0003] In response to the problems of modal fragmentation, knowledge staticization defects, insufficient reasoning depth, etc. in current large-model intelligent question-answering applications, the present invention provides an intelligent customer service robot based on multimodal data fusion and knowledge graph enhancement and its construction method, exploring a new paradigm of combining multimodal learning with knowledge graphs, and providing real-time, efficient and accurate information query services for events.

[0004] In a first aspect, the present invention provides a method for constructing an intelligent customer service robot based on multimodal data fusion and knowledge graph enhancement, comprising:

[0005] Building the data layer for the intelligent customer service robot involves: obtaining PDF files related to the target domain knowledge, preprocessing all PDF files using a large language model and the Miner U tool to obtain structured data for each PDF file; and constructing a knowledge graph and weaviate vector database for the target domain knowledge based on the structured data of all PDF files.

[0006] Build the model layer of the intelligent customer service robot, including the question classification model, knowledge graph query engine, vector retrieval module and large model interaction interface;

[0007] Build the application layer of the intelligent customer service robot, including the web-side interactive interface, API service, and management backend;

[0008] Build access links for the data layer, model layer, and application layer, including: links between the knowledge graph and the knowledge graph query engine, links between the vector database and the vector retrieval module, links between the question classification model and the knowledge graph query engine, links between the question classification model and the vector retrieval module, links between the knowledge graph query engine and the large model interaction interface, links between the vector retrieval module and the large model interaction interface, links between the large model interaction interface and the Web-side interaction interface, and links between the large model interaction interface and the API service.

[0009] Furthermore, all PDF files are preprocessed using a large language model and the Miner U tool, specifically including:

[0010] Design prompt words based on the metadata to be obtained, and use a large language model to extract text from each PDF file based on the prompt words to obtain structured data for each PDF file;

[0011] Use the Magic-PDF library to process each PDF file and obtain the structured data of each PDF file.

[0012] Furthermore, the structured data of all PDF files is used to construct a knowledge graph of target domain knowledge, specifically including: using a large language model to construct a lightweight knowledge graph for type questions including basic query types, statistical analysis types and open types.

[0013] Furthermore, it also includes: updating the knowledge graph, specifically including:

[0014] Determine the data type of the input multi-source data. If it is structured data, use SQL / API for incremental extraction; if it is unstructured data, use Magic-PDF / NLP for parsing, and then perform knowledge extraction and standardization.

[0015] Use the difference detection engine to locate the changed entities, relationships, and structures and determine whether there are any updates;

[0016] If an update is detected, conflict resolution strategies are used to handle possible entity ambiguities, relationship contradictions, and attribute conflicts.

[0017] Differential tap technology is used to encapsulate knowledge changes into entity taps, relationship taps, and structure taps. Combined with a progressive synchronization algorithm, the central node generates and verifies the legality of differential taps, dynamically allocates taps according to node load, and the edge node returns the verification result after completing the local update, ultimately completing the incremental update of the knowledge graph.

[0018] Furthermore, the difference detection engine locates the changed entities, relationships, and structures and determines whether there are any updates, including:

[0019] Calculate the subgraph embedding vectors of the new and old graphs through graph neural networks to locate the changed entities, relationships, and structures;

[0020] For unstructured data, the Magic-PDF library is used to parse the hierarchical structure of PDF documents, and the CLIP model is combined to extract cross-modal semantic vectors, and the cosine similarity is used to calculate the difference between modalities.

[0021] Furthermore, the conflict resolution strategy includes: when the conflict type is entity ambiguity, entity linking is adopted; when the conflict type is relationship contradiction, a combination of timestamp priority and manual review is adopted; when the conflict type is attribute contradiction, if the attribute is numerical, the average is taken; if the attribute is enumeration type, voting decision is adopted.

[0022] In a second aspect, the present invention provides an intelligent customer service robot based on multimodal data fusion and knowledge graph enhancement, which is constructed using the construction method described in the first aspect.

[0023] In a third aspect, the present invention provides an electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the method described in the first aspect when executing the program.

[0024] In a fourth aspect, the present invention provides a non-transitory computer-readable storage medium having a computer program stored thereon, wherein the computer program, when executed by a processor, implements the method described in the first aspect.

[0025] The beneficial effects of the present invention are:

[0026] To address the problem of modal fragmentation, this paper proposes a PDF data preprocessing method based on multi-technology fusion. By combining a large language model with the Magic-PDF library, it achieves unified representation and joint understanding of text, images, tables and other data.

[0027] To address the problem of knowledge staticization defects, the present invention proposes a RAG knowledge base dynamic update and management method based on multimodal differential synchronization. Through the closed loop of "data perception-difference detection-update execution-quality verification", minute-level synchronization and dynamic optimization of event information are achieved.

[0028] To address the problem of insufficient reasoning depth, this paper designs a large-model reasoning enhancement mechanism based on multi-dimensional data governance and knowledge graphs. By constructing a semantic network knowledge graph in the competition field and combining it with the retrieval enhancement technology of the Weaviate vector database, the system's ability to deeply understand the competition's professional knowledge is improved. BRIEF DESCRIPTION OF THE DRAWINGS

[0029] Figure 1 A flowchart of a method for constructing a competitive intelligent customer service robot based on multimodal data fusion and knowledge graph enhancement provided by an embodiment of the present invention;

[0030] Figure 2 A flowchart of PDF data preprocessing based on multi-technology fusion provided by an embodiment of the present invention;

[0031] Figure 3 The workflow for parsing PDF documents using a large model provided by an embodiment of the present invention;

[0032] Figure 4 Visualization of the knowledge graph of the four pieces of information provided by the embodiment of the present invention;

[0033] Figure 5 The Magic-pdf library extraction process provided by the embodiment of the present invention;

[0034] Figure 6 An example diagram of PDF file recognition and an example diagram of a partially converted document provided in an embodiment of the present invention;

[0035] Figure 7 A partial schematic diagram of the knowledge graph provided by an embodiment of the present invention;

[0036] Figure 8 A global schematic diagram of the semantic network knowledge graph for the competition field provided by an embodiment of the present invention;

[0037] Figure 9 The document processing metadata information record representation provided by the embodiment of the present invention;

[0038] Figure 10 A schematic diagram of storing some event information based on a vector database provided by an embodiment of the present invention;

[0039] Figure 11 A schematic diagram of metadata storage details for a specific competition project document provided by an embodiment of the present invention;

[0040] Figure 12 A schematic diagram of the incremental update process provided by an embodiment of the present invention;

[0041] Figure 13 The system architecture of a competition intelligent customer service robot based on multimodal data fusion and knowledge graph enhancement provided by the embodiment of the present invention;

[0042] Figure 14 This is a structural block diagram of an electronic device provided by an embodiment of the present invention. DETAILED DESCRIPTION

[0043] To make the objectives, technical solutions, and advantages of the present invention more clear, the technical solutions in the embodiments of the present invention will be clearly described below in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.

[0044] Take the competition field as an example, Figure 1 As shown, an embodiment of the present invention provides a method for constructing a competition intelligent customer service robot based on multimodal data fusion and knowledge graph enhancement, comprising the following steps:

[0045] S101: Build the data layer for the intelligent customer service robot, including: obtaining PDF files related to the target domain knowledge, preprocessing all PDF files using a large language model and Miner U tools to obtain structured data for each PDF file; based on the structured data of all PDF files, constructing a knowledge graph (such as Neo4j) and a weaviate vector database for the target domain knowledge;

[0046] Specifically, the competition domain knowledge is centrally managed and integrated, and through a combination of retrieval and generation, accurate, comprehensive, and detailed information is provided to users, meeting their diverse query needs. Therefore, extracting valuable information from PDF files, similar to unstructured data, and performing effective preprocessing are key to building intelligent question-answering systems and conducting data analysis. In the task of building intelligent customer service robots, preprocessing unstructured PDF documents is a key step in achieving information extraction and knowledge modeling.

[0047] S102: Build the model layer of the intelligent customer service robot, including the question classification model (such as FastText), the knowledge graph query engine, the vector retrieval module (such as Sentence-BERT), and the large model interaction interface (such as GPT-4);

[0048] S103: Build the application layer of the intelligent customer service robot, including the web-based interactive interface, API services (such as Excel import / export support), and management backend (such as data monitoring and auditing);

[0049] S104: Construct access links for the data layer, model layer, and application layer, including: links between the knowledge graph and the knowledge graph query engine, links between the vector database and the vector retrieval module, links between the question classification model and the knowledge graph query engine, links between the question classification model and the vector retrieval module, links between the knowledge graph query engine and the large model interaction interface, links between the vector retrieval module and the large model interaction interface, links between the large model interaction interface and the Web-side interaction interface, and links between the large model interaction interface and the API service.

[0050] The embodiment of the present invention provides a method for constructing a highly available intelligent customer service system based on a three-layer architecture of "data layer-model layer-application layer". To address the problem of modal fragmentation, the intelligent robot provided by the embodiment of the present invention pre-processes PDF files and achieves unified representation and joint understanding of data such as text, images, and tables through the combination of a large language model and the Magic-PDF library. In addition, the method for constructing an intelligent customer service robot provided by the embodiment of the present invention connects the structured query of the knowledge graph with the semantic retrieval of the vector database, realizing a hybrid question-answering model of "precise matching + semantic generalization", and improving the quality of answers to open-ended questions.

[0051] In one embodiment, all PDF files are preprocessed using a large language model and the Miner U tool, specifically including: designing prompt words based on the metadata to be acquired, extracting text from each PDF file using the large language model based on the prompt words to obtain structured data of each PDF file; and processing each PDF file using the Magic-PDF library to obtain structured data of each PDF file.

[0052] Specifically, this embodiment designs a preprocessing framework that integrates text parsing, image recognition and structured storage for competition PDF files. Figure 2 shown.

[0053] PDF files are not plain text files, but are transmitted and stored in binary format. Further knowledge mining of PDF document information in a particular field requires the identification and extraction of PDF text content. This embodiment provides knowledge extraction of PDF files from two perspectives.

[0054] The first aspect is the PDF text extraction method based on a large language model. This method integrates the semantic understanding and layout perception capabilities of large model documents through a three-level architecture of "preprocessing-semantic analysis-structured output" to achieve efficient extraction of unstructured and semi-structured PDFs. This method uses the "prompt word engineering" to convert the extraction task into natural language instructions, combines multimodal input processing (text coordinates, font features, etc.) with post-processing calibration technology, breaks through the traditional method's dependence on fixed templates, and has the advantages of template-free generalization, low annotation cost adaptation and complex layout processing. It can also be used with the RAG system to more conveniently build a competition question-answering robot, such as Figure 3 shown.

[0055] Specifically, competition-related PDF files contain a wealth of information, such as the event name, track, release date, registration period, organizing unit, official website, and possibly multimodal content such as images. To extract these six pieces of basic information (i.e., six metadata), we analyzed the data files and designed the following prompt words using the "prompt word project":

[0056] (1) Competition name: Combine the file name and the title of the first page of the document to extract the full competition name (including the number of sessions and year), for example, the 7th National Youth Artificial Intelligence Innovation Challenge.

[0057] (2) Track: Extracted from the title or competition introduction. Tracks are subdivided competition items based on themes, tasks or fields in the competition. They need to be accurately extracted from the title or introduction of the competition document and directly reflect the core type or goal of the competition. They must include keywords such as "special competition" and "challenge competition", such as "smart application special competition" (only retain the core type and remove prefixes such as "future campus").

[0058] (3) Release time: The year and month information at the bottom of the first page of the document, in the format of 'XXXX year X month'. You can filter based on all the dates in the entire document. The release time must be earlier than the registration time. If there is no such information, 'None' will be returned.

[0059] (4) Registration time: Search for the keyword 'registration time' in the format of 'XXXX-XXXX'. It must be after the release time. If not, 'None' will be returned.

[0060] (5) Organizing unit: The full name of the sponsoring unit above the publication date on the first page of the document, such as "China Children and Teenagers Development Service Center".

[0061] (6) Official website: Find the specific URL corresponding to the 'organizer website' or 'challenge website', such as 'http: / / www.china61.org.cn'.

[0062] Through PDF text extraction based on a large model, we achieved the goal of extracting valuable information from competition-related PDF files and converting it into structured data. This preprocessing process has good performance in terms of processing efficiency and data quality, providing a solid data foundation for the subsequent construction of an intelligent customer service system. Figure 4 .

[0063] The second aspect is document parsing based on the Magic-PDF library. Traditional knowledge base construction methods often directly extract and store the plain text content of PDF documents. This simple processing method ignores the inherent hierarchical structure information and contextual semantic associations of the document, resulting in problems such as logical discontinuities and semantic fragmentation in the constructed knowledge base, which in turn seriously affects the accuracy and response effect of the question-answering system. In response to the above problems, this embodiment proposes a solution for PDF document parsing based on the open source tool Magic-PDF library, such as Figure 5 By parsing the document's hierarchical structure (such as chapter titles, paragraph indentations, list relationships, etc.) and capturing contextual semantic associations (such as referential relationships, logical connectives, and co-occurrence of domain terms), this method can effectively preserve the document's semantic structure information, providing key support for the construction of a high-quality knowledge base. Compared with traditional methods, the parsing technology based on the Magic-PDF library not only maintains the integrity of the text content, but also mines the deep semantic structure of the document, thereby improving the quality of knowledge base construction and laying a solid foundation for the subsequent construction of the question-answering system. Figure 6 The following diagram shows an example of document identification and conversion into a corresponding Markdown file, which better preserves the hierarchical structure of the competition questions and the file format that is easy for machines to understand.

[0064] In one embodiment, to address the problem of insufficient reasoning depth, this embodiment designs a large-model reasoning enhancement mechanism based on multi-dimensional data governance and knowledge graph. By constructing a semantic network knowledge graph in the competition field and combining the retrieval enhancement technology of the Weaviate vector database, the system's ability to deeply understand the competition's professional knowledge is improved.

[0065] Specifically, this embodiment uses a large model to enhance the reasoning function of the knowledge graph, mainly by combining the semantic understanding ability of the large language model and the structured relationship network of the knowledge graph, so that the system can not only "remember" knowledge, but also "associate" and "reason". For example, when a user asks a question, the large model can quickly understand the intention of the question, and the knowledge graph provides relevant entities and relationships as background information to help the system find answers more accurately. In addition, the large model can also complete the missing logical chains in the knowledge graph, such as deducing "A competition is related to C city" from "A competition is hosted by B organization" and "B organization is located in C city". This combination not only improves the accuracy of the answer, but also can handle more complex open-ended questions, such as analyzing the logic behind the rules of the competition or giving personalized suggestions. Below, this embodiment further explains the construction method of the present invention mainly from two aspects: the construction of the knowledge graph and the construction of the vector database.

[0066] First, knowledge graph construction can be summarized into the following steps: ontology construction, entity extraction, relationship extraction, and knowledge storage. Traditional knowledge graph construction is complex and resource-intensive. However, with the rapid development of large models, it is becoming increasingly clear that combining them with large models is both faster and more efficient, and exhibits greater accuracy when processing lightweight data. Figure 7 In order to use the large model as an auxiliary means to build a lightweight knowledge graph based on the cleaned data.

[0067] Specifically, this embodiment aims to construct a multi-layer knowledge graph containing entities, relationships, and attributes for the three types of questions in the competition (basic queries, statistical analysis, and open questions), and realize the semantic upgrade from data to knowledge. Figure 8 This is the semantic network knowledge graph of the competition field constructed based on the PDF document in this embodiment.

[0068] The second aspect is retrieval enhancement based on the Weaviate vector database. The core of Weaviate's retrieval enhancement lies in converting unstructured data into vector embeddings, measuring data relevance through indicators such as cosine similarity and Euclidean distance between vectors, and combining hybrid retrieval capabilities to support simultaneous query of vector similarity and traditional structured fields (such as keywords, time, categories, etc.), achieving dual precise matching of "semantics + syntax". Its technical framework supports the dynamic construction of knowledge graphs, allowing data objects to be associated through attributes and relationships, and is suitable for complex query scenarios that require deep semantic understanding.

[0069] Faced with the need for semantic parsing of large numbers of PDF procedural documents in intelligent customer service competition scenarios, weaviate has advantages over other methods in unstructured data processing, hybrid retrieval capabilities, and scalable architectural design, enabling it to support vectorized storage and retrieval of multimodal data such as text, images, and PDFs; it supports combining vector retrieval with traditional SQL-like filtering conditions, taking into account both semantic relevance and structured screening; it supports horizontal cluster expansion to adapt to real-time updates of event data, and can ensure stable retrieval performance when the knowledge base is dynamically expanded.

[0070] From the aspects of document processing metadata information recording, partial event information storage based on vector database, and specific competition project document metadata storage details, Figure 9 It shows the data storage details during the vector database construction process. Figure 10 This is a schematic diagram of storing some event information based on a vector database. Figure 11 A diagram showing the details of how metadata is stored for a specific competition project document.

[0071] In one embodiment, to address the problem of knowledge static defects, this embodiment proposes a RAG knowledge base dynamic update and management method based on multimodal differential synchronization. Through the "data perception-difference detection-update execution-quality verification" closed loop, minute-level synchronization and dynamic optimization of event information are achieved.

[0072] Specifically, if Figure 12 As shown, the incremental update process for the knowledge graph is rigorous and efficient, ensuring the timeliness and reliability of the knowledge graph. The update process specifically involves: First, after multi-source data is input, data type determination is performed. Structured data is incrementally extracted using SQL / API, while unstructured data is parsed using Magic-PDF / NLP. Knowledge extraction and standardization are then performed. Next, a difference detection engine employs techniques such as graph neural network-based subgraph matching and multimodal data difference benchmarking to locate changed entities, relationships, and structures and determine whether an update has occurred. If an update is detected, the conflict resolution module utilizes resolution strategies to address potential entity ambiguity, relationship inconsistencies, and attribute conflicts. Subsequently, differential tap technology is used to encapsulate knowledge changes into entity taps, relationship taps, and structure taps. Combined with a progressive synchronization algorithm, the central node generates and verifies the validity of differential taps, dynamically allocating taps based on node load. After completing local updates, edge nodes return verification results, ultimately completing the incremental update of the knowledge graph.

[0073] Entity taps include complete descriptions of added or deleted entities, such as the ID and attributes of "Added player Zhang San." Relationship taps include edge information for added or deleted relationships, such as the weight of "Zhang San - Participation - 2025 Season." Structural taps include subgraph reorganization instructions, such as "Merge the 'Group Stage' subgraph into the 'Schedule' parent node."

[0074] As an implementation method, a difference detection engine is used to locate changed entities, relationships, and structures, and determine whether there are any updates. Specifically, it includes: (1) Calculating the subgraph embedding vectors of the new and old graphs through a graph neural network to locate the changed entities, relationships, and structures; for example, in the event knowledge graph, a player transfer event triggers a change in the "belonging team" relationship, and the associated subgraph modules such as schedules and results are quickly located through subgraph matching. (2) For unstructured data (such as PDF rule documents and event images), the Magic-PDF library is used to parse the PDF document hierarchical structure, combined with the CLIP model to extract cross-modal semantic vectors, and the difference between modalities is calculated through cosine similarity.

[0075] As an implementable approach, the conflict resolution strategy includes: when the conflict type is entity ambiguity, entity linking is used; when the conflict type is relationship conflict, a combination of timestamp priority and manual review is used; when the conflict type is attribute conflict, the average is taken if the attribute is numeric, and voting is used if the attribute is enumerated. This is shown in Table 1.

[0076] Table 1 Conflict resolution examples

[0077]

[0078] With the dynamic changes in competition information (such as new events, registration time adjustments, etc.), the knowledge base needs to have real-time update capabilities and be effectively managed to ensure the accuracy and timeliness of answers. The dynamic update management of the knowledge base is affected by the update management of the knowledge graph. As long as the update management of the knowledge graph is reliable, the questions answered by the robot will meet the requirements. However, the update management of the knowledge graph is not achieved overnight. This embodiment solves the problems of incremental data identification, conflict resolution, and semantic consistency maintenance by constructing a closed-loop knowledge graph update process of "data perception-difference detection-update execution-quality verification".

[0079] On this basis, the present invention also provides optimization strategies for knowledge graph updating and management from the following aspects.

[0080] (1) On the technical level, the synergistic mechanism between hierarchical indexing and GPU parallel computing can be further optimized to increase the speed of difference detection and reduce update time. For noise processing of open domain dynamic knowledge, the training of self-supervised learning models can be strengthened to improve their ability to identify and filter complex noise. At the same time, the application of multimodal evidence chains can be strengthened to improve anti-interference performance.

[0081] (2) In terms of management mechanisms, the data quality management system can be further improved, and more stringent quality assessment and screening of multi-source data can be carried out to ensure the reliability of input data. A more intelligent update demand identification mechanism can be established to achieve self-evolving graphs, automatically detect data changes, and trigger the update process in a timely manner.

[0082] (3) We should pay more attention to the deep integration with downstream tasks. According to the actual needs of applications such as RAG systems and intelligent question answering, we should optimize the update frequency and content of the knowledge graph to provide them with higher-quality knowledge support that is more suitable for actual application scenarios. At the same time, we can also strengthen the monitoring and evaluation of the update process, establish a scientific evaluation indicator system, constantly summarize experience, and continuously improve the update and management strategies to achieve a comprehensive improvement in the knowledge quality of the knowledge graph.

[0083] like Figure 13 As shown, an embodiment of the present invention also provides an intelligent customer service robot based on multimodal data fusion and knowledge graph enhancement, which is constructed using the above-mentioned construction method.

[0084] Figure 14 An example of a physical structure diagram of an electronic device is shown below. Figure 14As shown, the electronic device may include: a processor (processor) 1401, a communication interface (Communications Interface) 1402, a memory (memory) 1403 and a communication bus 1404, wherein the processor 1401, the communication interface 1402, and the memory 1403 communicate with each other through the communication bus 1404. The processor 1401 can call the logic instructions in the memory 1403 to execute the method for constructing an intelligent customer service robot based on multimodal data fusion and knowledge graph enhancement, the method comprising: constructing a data layer of the intelligent customer service robot, including: obtaining PDF files related to target domain knowledge, using a large language model and Miner The U tool preprocesses all PDF files to obtain the structured data of each PDF file; based on the structured data of all PDF files, it constructs the knowledge graph and weaviate vector database of the target domain knowledge; builds the model layer of the intelligent customer service robot, including the question classification model, knowledge graph query engine, vector retrieval module and large model interaction interface; builds the application layer of the intelligent customer service robot, including the Web-side interaction interface, API service and management background; builds access links among the data layer, model layer and application layer, including: the link between the knowledge graph and the knowledge graph query engine, the link between the vector database and the vector retrieval module, the link between the question classification model and the knowledge graph query engine, the link between the question classification model and the vector retrieval module, the link between the knowledge graph query engine and the large model interaction interface, the link between the vector retrieval module and the large model interaction interface, the link between the large model interaction interface and the Web-side interaction interface, and the link between the large model interaction interface and the API service.

[0085] In addition, when the logic instructions in the above-mentioned memory 1403 are implemented in the form of a software functional unit and sold or used as an independent product, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art, or the part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for enabling a computer device (which can be a personal computer, a server, or a network device, etc.) to perform all or part of the steps of the method described in each embodiment of the present invention. The aforementioned storage medium includes various media that can store program codes, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk.

[0086] An embodiment of the present invention also provides a computer program product, which includes a computer program stored on a non-transitory computer-readable storage medium, and the computer program includes program instructions. When the program instructions are executed by a computer, the computer can execute the method for constructing an intelligent customer service robot based on multimodal data fusion and knowledge graph enhancement provided by the above-mentioned method embodiments.

[0087] An embodiment of the present invention also provides a non-transitory computer-readable storage medium on which a computer program is stored. When the computer program is executed by a processor, it implements the method for constructing an intelligent customer service robot based on multimodal data fusion and knowledge graph enhancement provided by the above-mentioned method embodiments.

[0088] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a necessary general hardware platform, or of course, by hardware. Based on this understanding, the essence of the above technical solution or the part that contributes to the existing technology can be embodied in the form of a software product. The computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, a magnetic disk, an optical disk, etc., and includes a number of instructions for enabling a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or certain parts of the embodiments.

[0089] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the various embodiments of the present invention.

Claims

1. A method for constructing an intelligent customer service robot based on multimodal data fusion and knowledge graph enhancement, characterized in that: include: Building the data layer for the intelligent customer service robot involves: obtaining PDF files related to the target domain knowledge, preprocessing all PDF files using a large language model and the Miner U tool to obtain structured data for each PDF file; and constructing a knowledge graph and weaviate vector database for the target domain knowledge based on the structured data of all PDF files. Build the model layer of the intelligent customer service robot, including the question classification model, knowledge graph query engine, vector retrieval module and large model interaction interface; Build the application layer of the intelligent customer service robot, including the web-side interactive interface, API service, and management backend; Build access links for the data layer, model layer, and application layer, including: links between the knowledge graph and the knowledge graph query engine, links between the vector database and the vector retrieval module, links between the question classification model and the knowledge graph query engine, links between the question classification model and the vector retrieval module, links between the knowledge graph query engine and the large model interaction interface, links between the vector retrieval module and the large model interaction interface, links between the large model interaction interface and the Web-side interaction interface, and links between the large model interaction interface and the API service.

2. The method for constructing an intelligent customer service robot based on multimodal data fusion and knowledge graph enhancement according to claim 1 is characterized in that: All PDF files are preprocessed using a large language model and Miner U tools, including: Design prompt words based on the metadata to be obtained, and use a large language model to extract text from each PDF file based on the prompt words to obtain structured data for each PDF file; Use the Magic-PDF library to process each PDF file and obtain the structured data of each PDF file.

3. The method for constructing an intelligent customer service robot based on multimodal data fusion and knowledge graph enhancement according to claim 1 is characterized in that: The structured data based on all PDF files is used to construct a knowledge graph of target domain knowledge, specifically including: using a large language model to construct a lightweight knowledge graph for type questions including basic query types, statistical analysis types and open types.

4. The method for constructing an intelligent customer service robot based on multimodal data fusion and knowledge graph enhancement according to claim 1 is characterized in that: Also includes: Update the knowledge graph, including: Determine the data type of the input multi-source data. If it is structured data, use SQL / API for incremental extraction; if it is unstructured data, use Magic-PDF / NLP for parsing, and then perform knowledge extraction and standardization. Use the difference detection engine to locate the changed entities, relationships, and structures and determine whether there are any updates; If an update is detected, conflict resolution strategies are used to handle possible entity ambiguities, relationship contradictions, and attribute conflicts. Differential tap technology is used to encapsulate knowledge changes into entity taps, relationship taps, and structure taps. Combined with a progressive synchronization algorithm, the central node generates and verifies the legality of differential taps, dynamically allocates taps according to node load, and the edge node returns the verification result after completing the local update, ultimately completing the incremental update of the knowledge graph.

5. The method for constructing an intelligent customer service robot based on multimodal data fusion and knowledge graph enhancement according to claim 4 is characterized in that: The difference detection engine locates the changed entities, relationships, and structures and determines whether there are any updates. Specifically, it includes: Calculate the subgraph embedding vectors of the new and old graphs through graph neural networks to locate the changed entities, relationships, and structures; For unstructured data, the Magic-PDF library is used to parse the hierarchical structure of PDF documents, and the CLIP model is combined to extract cross-modal semantic vectors, and the cosine similarity is used to calculate the difference between modalities.

6. The method for constructing an intelligent customer service robot based on multimodal data fusion and knowledge graph enhancement according to claim 4 is characterized in that: The conflict resolution strategy includes: when the conflict type is entity ambiguity, entity linking is adopted; when the conflict type is relationship contradiction, a combination of timestamp priority and manual review is adopted; when the conflict type is attribute contradiction, if the attribute is numerical, the average is taken; if the attribute is enumeration type, voting decision is adopted.

7. An intelligent customer service robot based on multimodal data fusion and knowledge graph enhancement, characterized by: The method is constructed according to any one of claims 1 to 6.

8. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the program, the method according to any one of claims 1 to 6 is implemented.

9. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the method according to any one of claims 1 to 6 is implemented.

Citation Information

Cited By

  • Dynamic AI knowledge graph system

    CN120892581A

  • Community service robot-oriented multi-mode perception fusion man-machine interaction system

    CN122064273A