Communication customer service question answering system and method based on deep knowledge graph

By using deep knowledge graphs and related technologies in the communication customer service question and answer system, the existing system cannot deeply understand user problems and provide answers accurately, and more efficient semantic understanding and answer generation are achieved.

CN120179778APending Publication Date: 2025-06-20NANJING LONGYUAN INFORMATION TECH CO LTD

Patent Information

Application Number
CN202510246833.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-04
Publication Date
2025-06-20

AI Technical Summary

Technical Problem

The existing communication field question and answer system based on knowledge graphs cannot deeply understand user problems and cannot provide answers accurately, and there are problems such as incomplete knowledge graph construction, insufficient accuracy of entity recognition and links, and poor text classification algorithm effects.

Method used

The communication customer service question and answer system based on deep knowledge graph is adopted, including deep knowledge graph construction module, entity recognition module, link module, question classification module, word order diagram matching module and answer generation module. Through technologies such as CNN model and naive Bayes algorithm, entity recognition, linking and problem classification are realized, forming a complete query logic diagram, and query answers in the graph database.

Benefits of technology

It improves the system's deep understanding of user problems and the accuracy of answers, enhances the overall performance of the system, and can more effectively deal with complex semantic understanding and logical reasoning problems.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120179778A_ABST
    Figure CN120179778A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of communication customer service question answering systems, in particular to a communication customer service question answering system and method based on a deep knowledge graph. The deep knowledge graph is established through a deep knowledge graph establishing module, after a user inputs a question, the user question is recognized through an entity recognition module, and an entity in the user question is obtained; entities in user questions and candidate entities in the deep knowledge graph are selected to be linked by utilizing a linking module, the user questions are classified by utilizing a question classification module, and the entities subjected to entity recognition and linking are filled into a word order graph template by utilizing a word order graph matching module to form a query logic graph; according to the technical scheme, the intelligent level of the communication field can be improved, the competitiveness of communication operators and service providers can be enhanced, and better communication service experience can be provided for users.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of communication customer service question - answering systems, and particularly to a communication customer service question - answering system and method based on a deep knowledge graph. Background Art

[0002] With the rapid development of communication technologies, the types of services provided by communication operators are becoming increasingly diverse, and the problems encountered by users during use are also becoming more complex and diverse. Traditional communication customer service methods mainly rely on human customer service, which has problems such as slow response speed, and the accuracy of answers being limited by the professional level and experience of customer service staff. At the same time, in the face of a large number of user inquiries, the work pressure on human customer service is huge, and it is difficult to meet the immediate needs of users.

[0003] In recent years, knowledge graph technology has been applied to a certain extent in the field of intelligent question - answering. However, in the communication field, existing question - answering systems based on knowledge graphs still have many deficiencies. For example, the construction of the knowledge graph is not perfect enough, the depth and breadth of coverage of professional knowledge in the communication field are limited, and it is unable to effectively handle complex semantic understanding and logical reasoning problems; the accuracy of entity recognition and linking needs to be improved, resulting in deviations in question classification and answer retrieval; the text classification algorithm has poor performance when dealing with the unique features of communication texts, affecting the overall performance of the system. Therefore, it is of great practical significance to develop a communication customer service question - answering system based on a deep knowledge graph that can deeply understand user questions and accurately provide answers. Summary of the Invention

[0004] The purpose of the present invention is to provide a communication customer service question - answering system and method based on a deep knowledge graph, so as to solve the problems that existing question - answering systems based on knowledge graphs in the communication field cannot deeply understand user questions and cannot accurately provide answers.

[0005] To achieve the above object, the present invention provides a communication customer service question and answer system and method based on a deep knowledge graph. The communication customer service question and answer system based on a deep knowledge graph includes a deep knowledge graph construction module, an entity recognition module, a linking module, a question classification module, a word order graph matching module, and an answer generation module. The deep knowledge graph construction module is used to establish a deep knowledge graph. The entity recognition module is used to identify entities in the user's question. The linking module is used to obtain vectors for the entities in the user's question and candidate entities in the deep knowledge graph based on a CNN model, and select entity links. The question classification module is used to preprocess the user's question, represent the question as a feature vector based on a CNN model, and classify it using the naive Bayes algorithm to obtain the category label to which the question belongs. The word order graph matching module, based on the category label obtained by the question classification module, fills the entities after entity recognition and linking into the word order graph template to form a complete query logic graph. The answer generation module is used to convert the constructed query logic graph into a query statement for a graph database, query in the graph database storing the deep knowledge graph, and obtain the answer to the question.

[0006] Among them, the deep knowledge graph construction module includes a data source selection unit, a data collection unit, a knowledge extraction unit, a knowledge graph schema layer design unit, a knowledge fusion unit, and a knowledge storage unit. The data source selection unit is used to integrate multi-source data. The data collection unit is used to collect multi-source data. The knowledge extraction unit performs knowledge extraction by combining natural language processing techniques and domain-specific rules. The knowledge graph schema layer design unit is used to design the schema layer of the knowledge graph. The knowledge fusion unit fuses knowledge through entity alignment strategies to ensure the consistency and accuracy of knowledge. The knowledge storage unit is used to store the deep knowledge graph using a suitable graph database.

[0007] Among them, the specific function of the entity recognition unit is as follows: after the user's question is input, it is first segmented using a Chinese word segmentation tool, then irrelevant words are removed by combining with a stop word list, and then the entities in the question are recognized through matching with a communication domain entity dictionary and a named entity recognition model based on a conditional random field.

[0008] Among them, the specific content of obtaining vectors through the CNN model is as follows: represent the entity text in the user's question or knowledge graph as a matrix X, whose dimension is n×m, where n is the length of the text and m is the dimension of the word vector. For a text containing n words, each word x i is represented as an m-dimensional vector, that is, x i ∈ R m , and the input matrix X is represented as:

[0009]

[0010] In the convolutional layer, a convolutional kernel (filter) W is used to extract features. Assume the size of the convolutional kernel is h×m, where h represents the number of words covered by the convolutional kernel, and its elements are w ij , then the convolution operation can be expressed as

[0011]

[0012] where f is the activation function and b is the bias term. By sliding the convolutional kernel over the input matrix X, if k different convolutional kernels are used, k feature maps will be obtained. Each feature map can be represented as C1 (l = 1, 2,..., k). To reduce the dimension of the feature maps, a max-pooling operation is performed on each feature map C1, which can be expressed as p1 = max(C1). The pooled feature vectors are concatenated together to form a vector p = [p1, p2,..., p k , whose dimension is k×1. Then, it is mapped to the final output vector y through the fully connected layer. If the weight matrix of the fully connected layer is W f , with dimension k×d, where d is the dimension of the final output vector, and the bias term is b f , then: y = f(W f p + b f ).

[0013] Among them, the specific content of the entity linking module for entity linking through vectors is as follows:

[0014] Calculate the cosine similarity between the vector y1 obtained by the CNN model for the entity in the user question and the vector y2 obtained by the CNN model for the candidate entity in the knowledge graph:

[0015]

[0016] Select entity linking according to the similarity magnitude and the method of semantic similarity calculation.

[0017] Among them, the specific content of the question classification module is: after preprocessing the user question by word segmentation and stop word filtering, represent the question as a feature vector through the CNN model, and use the Naive Bayes algorithm for classification to obtain the category label to which the question belongs.

[0018] Among them, the word order graph matching module includes a word order graph template library construction unit and a matching word order graph unit. The word order graph template library construction unit constructs a template library containing multiple word order graph templates based on the problem types in the communication field and the structure of the knowledge graph. The matching word order graph unit is used to select the corresponding word order graph template from the word order graph template library according to the category label obtained by the question classification module, and then fill the entities after entity recognition and linking into the word order graph template to form a complete query logic graph.

[0019] The present invention also provides a communication customer service question - answering method based on a deep knowledge graph, which is applied to the communication customer service question - answering system based on the deep knowledge graph as described above, and includes the following steps:

[0020] A user inputs a question into the communication customer question - answering system based on the deep knowledge graph;

[0021] The communication customer question - answering system based on the deep knowledge graph performs word segmentation based on the question input by the user;

[0022] The entity recognition module is used to recognize the segmented user question to obtain the entities in the user question;

[0023] The link module is used to select entity links between the entities in the user question and the candidate entities in the deep knowledge graph in the system;

[0024] The question classification module is used to classify the user question;

[0025] The word order graph matching module is used to fill the entities after entity recognition and linking into the word order graph template to form a complete query logic graph;

[0026] The answer generation module is used to convert the constructed query logic graph into a query statement of the graph database, query in the graph database storing the deep knowledge graph, and obtain the answer to the question.

[0027] For a communication customer service question - answering system and method based on a deep knowledge graph of the present invention, the deep knowledge graph construction module is used to establish a deep knowledge graph. A user inputs a question into the communication customer question - answering system based on the deep knowledge graph. The entity recognition module is used to recognize the segmented user question to obtain the entities in the user question. The link module is used to select entity links between the entities in the user question and the candidate entities in the deep knowledge graph in the system. The question classification module is used to classify the user question. The word order graph matching module is used to fill the entities after entity recognition and linking into the word order graph template to form a complete query logic graph. The answer generation module is used to convert the constructed query logic graph into a query statement of the graph database, query in the graph database storing the deep knowledge graph, and obtain the answer to the question. By adopting this technical solution, the intelligent level in the communication field can be improved, the competitiveness of communication operators and service providers can be enhanced, and a better communication service experience can be provided for users. Description of the Drawings

[0028] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the accompanying drawings required for the description of the embodiments or the prior art. Obviously, the accompanying drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other accompanying drawings can be obtained based on these drawings.

[0029] Figure 1 It is the principle block diagram of the communication customer service Q&A system based on the deep knowledge graph provided by the present invention.

[0030] Figure 2 It is the operation flowchart of the link module provided by the present invention.

[0031] Figure 3 It is the step flowchart of the communication customer service Q&A method based on the deep knowledge graph provided by the present invention.

[0032] 101 - Deep knowledge graph construction module, 102 - Entity recognition module, 103 - Link module, 104 - Question classification module, 105 - Word order diagram matching module, 106 - Answer generation module, 107 - Data source selection unit, 108 - Data acquisition unit, 109 - Knowledge extraction unit, 110 - Knowledge graph schema layer design unit, 111 - Knowledge fusion unit, 112 - Knowledge storage unit, 113 - Word order diagram template library construction unit, 114 - Matching word order diagram unit. Detailed implementation manners

[0033] The following will describe in detail the embodiments of the present invention. The examples of the embodiments are shown in the accompanying drawings, where the same or similar reference numerals represent the same or similar elements or elements with the same or similar functions from beginning to end. The embodiments described below with reference to the accompanying drawings are exemplary and are intended to explain the present invention, but should not be construed as a limitation to the present invention.

[0034] Please refer to Figure 1 and Figure 2, the present invention provides a communication customer service Q&A system based on a deep knowledge graph. The communication customer service Q&A system based on a deep knowledge graph includes a deep knowledge graph construction module 101, an entity recognition module 102, a linking module 103, a question classification module 104, a word order graph matching module 105, and an answer generation module 106. The deep knowledge graph construction module 101 is used to build a deep knowledge graph. The entity recognition module 102 is used to recognize entities in the user's question. The linking module 103 is used to obtain vectors for the entities in the user's question and candidate entities in the deep knowledge graph based on a CNN model, and select entity links. The question classification module 104 is used to preprocess the user's question, represent the question as a feature vector based on a CNN model, and classify it using the naive Bayes algorithm to obtain the category label to which the question belongs. The word order graph matching module 105 fills the entities after entity recognition and linking into the word order graph template based on the category label obtained by the question classification module 104 to form a complete query logic graph. The answer generation module 106 is used to convert the constructed query logic graph into a query statement for the graph database, query in the graph database storing the deep knowledge graph, and obtain the answer to the question.

[0035] In this embodiment, the deep knowledge graph construction module 101 is used to build a deep knowledge graph. The user inputs a question into the communication customer service Q&A system based on a deep knowledge graph. The entity recognition module 102 is used to recognize the entities in the segmented user's question. The linking module 103 is used to select entity links for the entities in the user's question and candidate entities in the deep knowledge graph in the system. The question classification module 104 is used to classify the user's question. The word order graph matching module 105 fills the entities after entity recognition and linking into the word order graph template to form a complete query logic graph. The answer generation module 106 is used to convert the constructed query logic graph into a query statement for the graph database, query in the graph database storing the deep knowledge graph, and obtain the answer to the question. By adopting this technical solution, the intelligent level in the communication field can be improved, the competitiveness of communication operators and service providers can be enhanced, and a better communication service experience can be provided for users.

[0036] Further, the deep knowledge graph construction module 101 includes a data source selection unit 107, a data collection unit 108, a knowledge extraction unit 109, a knowledge graph schema layer design unit 110, a knowledge fusion unit 111, and a knowledge storage unit 112. The data source selection unit 107 is used to integrate multi-source data. The data collection unit 108 is used to collect multi-source data. The knowledge extraction unit 109 performs knowledge extraction by combining natural language processing technology and domain-specific rules. The knowledge graph schema layer design unit 110 is used to design the schema layer of the knowledge graph. The knowledge fusion unit 111 fuses knowledge through entity alignment strategies to ensure the consistency and accuracy of knowledge. The knowledge storage unit 112 is used to select a suitable graph database to store the deep knowledge graph.

[0037] In this embodiment, the data source selection unit 107 includes the integration of communication network device data and communication service data. Communication network device data: Obtain detailed information of communication devices from the network management system (NMS) of communication operators, the device management platform of device suppliers, and network operation and maintenance logs, including device models, device configurations, software versions, geographical locations, device status (normal, faulty, under maintenance), device connection relationships, etc. At the same time, integrate the own data such as customer service records, work order information, and operation and maintenance data within communication operators to ensure the comprehensiveness and authority of the data.

[0038] Communication service data: Collect user service subscription information, service usage records (call duration, data usage, number of text messages sent, etc.), service package details (package type, package fee, included service content, etc.), service billing data, and service marketing activity data in the business operation support system (BOSS) of communication operators.

[0039] The data collection unit 108 uses ETL (Extract, Transform, Load) tools for regular extraction and transformation of structured data (such as data in databases) to ensure the accuracy and consistency of the data. For unstructured data (such as text files, web page content, etc.), web crawler technology, text parsing tools, and natural language processing (NLP) technology are used for data collection and preprocessing. A web crawler is written using the Scrapy framework based on Python to crawl product technical documents and update information on the websites of communication device manufacturers according to predefined rules. The NLP library (NLTK) is used to perform word segmentation, stop word removal, part-of-speech tagging, named entity recognition, etc. on the collected text data for subsequent knowledge extraction.

[0040] The knowledge extraction unit 109 performs knowledge extraction by combining natural language processing techniques and domain-specific rules. For structured data, entities, attributes, and relationships are extracted through parsing and transformation rules; for semi-structured data, such as tables and lists on web pages, template matching and information extraction algorithms are used to obtain knowledge; for unstructured text data, named entity recognition (NER) tools are used to identify entities such as communication device names, package types, fault phenomena, and communication protocols, and dependency syntactic analysis and semantic role labeling are used to extract relationships between entities, such as "mobile phone model A supports communication protocol B" and "package C includes traffic D".

[0041] The knowledge graph schema layer designed by the knowledge graph schema layer design unit 110 includes:

[0042] 1. Entity classification and attribute definition:

[0043] Communication device entities: including base stations (macro base stations, micro base stations, pico base stations, etc.), core network devices (such as mobile switching centers, packet core network devices, etc.), transmission devices (optical cables, optical transmission devices, microwave transmission devices, etc.), and terminal devices (mobile phones, tablets, Internet of Things terminals, etc.). Their attributes cover device names, device types, device IDs, device manufacturers, device production dates, device hardware configurations (such as CPU models, memory capacities, storage capacities, etc.), device software versions, device geographical locations (longitude, latitude, altitude, etc.), device statuses (normal operation, fault, maintenance in progress, shutdown, etc.), device connection relationships (connected to which devices, connection methods, connection ports, etc.), and device historical maintenance records (maintenance times, maintenance contents, maintenance personnel, etc.).

[0044] Communication technology entities: such as wireless communication technologies (2G, 3G, 4G, 5G, 6G, etc.), wired communication technologies (optical fiber communication, coaxial cable communication, twisted pair communication, etc.), network protocols (TCP / IP, HTTP, FTP, etc.), and communication coding technologies (such as CDMA, TDMA, etc.). Their attributes include technology names, technology standard numbers, technology development histories (start times, evolution stages, key technology breakthroughs, etc.), technology application scenarios (applicable to which communication services, widely used in which fields, etc.), technology advantages (such as high speed, low latency, large capacity, etc.), technology limitations (such as limited coverage, signal interference problems, etc.), and related patent information (patent numbers, patent owners, patent application dates, etc.).

[0045] Communication service entities: Voice call service, SMS service, data services (such as mobile data traffic packages, fixed broadband services, etc.), value-added services (such as video ringback tones, caller ID, mobile payment, Internet of Things application services, etc.). Its attributes include service name, service type (basic service, value-added service, etc.), service rules (such as billing rules, usage restrictions, service terms, etc.), service package details (package contents, package price, package validity period, etc.), service usage statistics (such as user usage duration, usage frequency, traffic consumption, etc.), service satisfaction evaluation (user ratings, complaint rates, etc.), service correlation relationships (which other services are interrelated, complementary or substitutive relationships, etc.).

[0046] User entities: User basic information (name, gender, age, ID number, contact information, home address, etc.), user account information (account ID, account balance, credit rating, package subscription history, payment records, etc.), user behavior information (service usage habits, application usage preferences, call time distribution, Internet access time distribution, location movement trajectory, etc.), user social relationships (friend list, group membership, social interaction frequency, etc.), user feedback information (complaint records, suggestion content, satisfaction evaluation, etc.).

[0047] Fault entities: Network faults (such as signal coverage problems, network congestion, disconnection, network outage, etc.), equipment faults (such as hardware faults, software faults, power supply faults, interface faults, etc.), service faults (such as services cannot be used normally, billing errors, service interruptions, etc.). Its attributes include fault name, fault code, fault occurrence time, fault discovery method (user report, system monitoring, etc.), fault impact scope (number of affected users, geographical area, service type, etc.), fault cause analysis (possible hardware reasons, software reasons, human reasons, environmental reasons, etc.), fault resolution measures (repair steps taken, repair time, repair personnel, etc.), fault history records (number of occurrences of the same or similar faults, last occurrence time, solution methods, etc.).

[0048] 2. Relationship definition and semantic description:

[0049] Device - Technology relationship: Represents the communication technologies adopted or supported by communication devices, such as "base station - adopts - 5G technology", "optical transmission device - supports - fiber optic communication technology". This relationship reflects the dependence and adaptation between devices and technologies, which helps to make reasonable decisions during technology upgrades or device selections.

[0050] Device - Device relationship: Describes the physical connection relationship, logical association relationship or cooperation relationship between communication devices, such as "base station - connects - core network device" (physical connection), "core network device - cooperates - billing system" (logical cooperation). Through this relationship, the topology of the communication network can be constructed, which is convenient for network management and fault troubleshooting.

[0051] Device - Service Relationship: Describes the bearing and support capabilities of communication devices for communication services, such as "Base Station - Supports - Voice Call Service", "Data Center Server - Provides - Cloud Storage Service". This relationship helps analyze the matching of services and device resources, and optimize service deployment and device resource allocation.

[0052] Service - User Relationship: Reflects the subscription, usage, and evaluation relationships between users and the communication services they use, such as "User - Subscribes to - Data Traffic Package", "User - Uses - Video Call Service", "User - Evaluates - SMS Service (Satisfaction: High)". Based on this relationship, personalized service recommendations and user demand analysis can be realized.

[0053] Technology - Technology Relationship: Reflects the evolution, substitution, or complementary relationships between communication technologies, such as "4G Technology - Develops into - 5G Technology", "Optical Fiber Communication Technology - Complementary to - Wireless Communication Technology (Collaborates in different scenarios)". This relationship helps track technology development trends and plan technology upgrade paths.

[0054] User - User Relationship: Describes the social, group, or similar behavior relationships between users, such as "User A - Friend - User B", "User C - Belongs to - Game Lover Group", "User D - Similar Behavior - User E (Has similar service usage habits)". Using these relationships, social network analysis can be conducted to mine user group characteristics and potential needs.

[0055] Fault - Device Relationship: Indicates that a fault occurs in a specific communication device, such as "Hardware Fault - Occurs in - Base Station Equipment (Specific Device ID)". This helps quickly locate the fault source and improve fault repair efficiency.

[0056] Fault - Service Relationship: Explains the impact of a fault on communication services, such as "Network Congestion Fault - Affects - Data Service (Leads to a decrease in data transmission rate)". Through this relationship, the service impact scope of a fault can be evaluated, and corresponding service restoration measures can be taken in a timely manner.

[0057] User - Fault Relationship: Reflects users' perception and feedback of faults, such as "User - Reports - Poor Call Quality Fault". The fault information feedback by users is an important basis for fault troubleshooting and optimizing network services.

[0058] The entity pair its strategy in the knowledge fusion unit 111 includes:

[0059] Attribute similarity calculation: Compare the attribute values of entities in different data sources to determine whether the entities are aligned. For numerical attributes (such as the hardware configuration parameters of devices, the age of users, etc.), the numerical difference calculation method "Euclidean distance" is used to evaluate the similarity; for string attributes (such as device names, user names, etc.), the string similarity algorithm "cosine similarity" is used to calculate the similarity. Different weights are set according to the importance of the attributes, and the overall similarity between entities is calculated comprehensively. For example, in the alignment of device entities, the weights of device models and device IDs are relatively high, while the weight of the device production date is relatively low.

[0060] Semantic similarity calculation: Use the word vector model to convert the entity name or description into a vector representation, and calculate the cosine similarity between the vectors. For entities with semantic ambiguity or aliases (such as "Mobile Switching Center" and "MSC"), semantic similarity calculation can more accurately determine whether they are the same entity.

[0061] Furthermore, the specific function of the entity recognition unit is as follows: After the user's question is input, first perform word segmentation processing using a Chinese word segmentation tool, then combine with a stop word list to remove irrelevant words, and then identify the entities in the question through matching with the entity dictionary in the communication field and the named entity recognition model based on conditional random fields.

[0062] In this embodiment, after the user's question is input, first perform word segmentation processing using Chinese word segmentation tools such as Jieba, then combine with a stop word list (including common modal particles, conjunctions, and general vocabulary irrelevant to the communication field) to remove irrelevant words, and then identify the entities in the question through matching with the entity dictionary in the communication field and the named entity recognition model based on conditional random fields (CRF). For example, for the question "What is the download speed of the Huawei P40 mobile phone under the 5G network", it can accurately identify "Huawei P40" (mobile phone entity), "5G network" (communication technology entity), etc.

[0063] Furthermore, the specific content of the vector obtained through the CNN model is as follows: Represent the entity text in the user's question or knowledge graph as a matrix X, whose dimension is n×m, where n is the length of the text and m is the dimension of the word vector. For a text containing n words, each word x i is represented as an m-dimensional vector, that is, x i ∈R m , and the input matrix X is represented as:

[0064]

[0065] In the convolutional layer, use the convolutional kernel (filter) W to extract features. Assume that the size of the convolutional kernel is h×m, h represents the number of words covered by the convolutional kernel, and its elements are wij , the convolution operation can be expressed as

[0066]

[0067] where f is the activation function, b is the bias term. By sliding the convolution kernel on the input matrix X, if k different convolution kernels are used, k feature maps will be obtained. Each feature map can be expressed as C1 (l = 1, 2, …, k). To reduce the dimension of the feature maps, a max-pooling operation is performed on each feature map C1, which can be expressed as p1 = max(C1). The pooled feature vectors are concatenated together to form a vector p = [p1, p2, …, p k , whose dimension is k×1. Then, it is mapped to the final output vector y through a fully connected layer. If the weight matrix of the fully connected layer is W f , with dimension k×d, where d is the dimension of the final output vector, and the bias term is b f , then: y = f(W f p + b f ).

[0068] Furthermore, the specific content of the entity linking by the linking module 103 through vectors is as follows:

[0069] Calculate the cosine similarity between the vector y1 obtained by the CNN model for the entity in the user question and the vector y2 obtained by the CNN model for the candidate entity in the knowledge graph:

[0070]

[0071] Select entity linking according to the similarity magnitude and the method of calculating semantic similarity.

[0072] In this embodiment, the link determination criterion is specifically as follows:

[0073] High similarity case (cosine similarity > 0.6):

[0074] Set two key cosine similarity discrimination thresholds, which are 0.4 and 0.6 respectively. When the calculated cosine similarity between the entity in the user question and the candidate entity is higher than 0.6, a direct preference strategy is adopted. Specifically, among the candidate entity set that meets this condition, select the entity with the largest similarity value as the final linking target. This is because a higher cosine similarity indicates that the two are very similar in the vector space, and selecting the candidate entity with the largest similarity can ensure the accuracy of the link.

[0075] Low similarity case (cosine similarity < 0.4):

[0076] If the calculated cosine similarities are all lower than 0.4, the backup plan is activated. First, the feature vector of the user question entity is calculated through a specific algorithm, and then the top 5 entities with a cosine similarity of more than 0.4 with the feature vector are screened out from the massive candidate entities. Then, the Jaro distance (edit distance) similarity between these 5 entities and the corresponding entities in the entity dictionary is further calculated. If there is an entity with a Jaro distance similarity greater than 0.6 in this round of calculation, it will be used as the final output result to complete the entity linking; if not, it is determined that the user question entity currently under investigation does not participate in this linking process. This step aims to explore potential matching entities that may exist through feature vectors and more detailed cosine similarity calculations.

[0077] Intermediate state (0.4≤cosine similarity≤0.6):

[0078] When the cosine similarity is in the range of 0.4 to 0.6, the candidate entities that meet this condition are first properly saved, and their feature vectors are calculated using the established algorithm, from which the entities corresponding to the first five values ​​with similarity exceeding 0.4 are found. After that, it is prioritized to check whether there is a case where the Jaro distance similarity with the target entity is exactly 1 among these five entities. If so, the entity is directly output; if not, the Jaro distance similarity of these five entities and the previously saved entities is calculated again. Once an entity with a similarity greater than 0.6 appears in the recalculation, it is established as the link target, otherwise the link is abandoned. This intermediate state processing method comprehensively considers the candidate entities with similarity in the medium range, and improves the accuracy and reliability of entity linking through multiple rounds of screening and calculation.

[0079] Furthermore, the specific content of the question classification module 104 is: after pre-processing the user's question by word segmentation and stop word filtering, the question is represented as a feature vector through the CNN model, and classified using the naive Bayes algorithm to obtain the category label to which the question belongs.

[0080] In this embodiment, since the communication field corpus has the characteristics of strong professionalism, many terms, and relatively concentrated question types, in terms of feature selection, in addition to using traditional text abstraction and preprocessed words as features, special attention is also paid to the extraction and weight setting of key terms in the communication field, business keywords, and core verbs in the questions (such as query, handle, troubleshooting, etc.). After preprocessing the user's questions by word segmentation and stop word filtering, the questions are represented as feature vectors through the above-mentioned CNN model, and classified using the naive Bayes algorithm to obtain the category label to which the questions belong, such as business consultation, technical failure, account problem, etc., so as to determine the intention of the user's questions.

[0081] Further, the word order graph matching module 105 includes a word order graph template library construction unit 113 and a matching word order graph unit 114. The word order graph template library construction unit 113 constructs a template library containing multiple word order graph templates based on the problem types in the communication field and the knowledge graph structure. The matching word order graph unit 114 is used to select the corresponding word order graph template from the word order graph template library through the category label obtained by the problem classification module 104, and then fill the entities after entity recognition and linking into the word order graph template to form a complete query logic graph.

[0082] In this embodiment, the word order graph is a directed graph in which the subject points to the object and is connected by the predicate. The subject and the object are usually entities, and the predicate is the relationship between the entities. For example, for the package query problem, the template is "package name - query - package attributes (traffic, call duration, fee, etc.)"; for the device failure problem, the template is "device name - appears - failure phenomenon - solution method", etc.

[0083] Further, the answer generation module 106 specifically means: converting the constructed query logic graph into a query statement of the graph database Neo4j, and querying in the graph database storing the deep knowledge graph to obtain the answer to the question.

[0084] In this embodiment, for some complex problems, it may be necessary to perform multi-step queries and inferences by combining the inference rules and association relationships in the knowledge graph to obtain accurate answers. For example, when the user asks "why is the signal of my mobile phone bad in this area", the system first searches for the knowledge related to the mobile phone model, the area where it is located, and the signal in the knowledge graph, and then analyzes the possible factors that may cause the bad signal, such as nearby base station failures, building blockages, mobile phone hardware problems, etc., according to the communication principle and the inference rules of common failure reasons, and generates a detailed answer to return to the user by integrating this information.

[0085] Please refer to Figure 3 , the present invention also provides a communication customer service question-answering method based on a deep knowledge graph, which is applied to the communication customer service question-answering system based on the deep knowledge graph as described above, and includes the following steps:

[0086] S1. The user inputs a question to the communication customer question-answering system based on the deep knowledge graph;

[0087] S2. The communication customer question-answering system based on the deep knowledge graph performs word segmentation based on the question input by the user;

[0088] S3. Use the entity recognition module 102 to recognize the segmented user question to obtain the entities in the user question;

[0089] S4. Use the link module 103 to perform entity linking on the entities in the user question and the candidate entities in the deep knowledge graph in the system;

[0090] S5. Use the question classification module 104 to classify the user question;

[0091] S6. Use the word order graph matching module 105 to fill the entities after entity recognition and linking into the word order graph template to form a complete query logic graph;

[0092] S7. Use the answer generation module 106 to convert the constructed query logic graph into a query statement for the graph database, query in the graph database storing the deep knowledge graph, and obtain the answer to the question.

[0093] In this embodiment, the deep knowledge graph is established by using the deep knowledge graph construction module 101. The user inputs a question to the communication customer Q&A system based on the deep knowledge graph. The entity recognition module 102 is used to recognize the entities in the segmented user question to obtain the entities in the user question. The link module 103 is used to perform entity linking on the entities in the user question and the candidate entities in the deep knowledge graph in the system. The question classification module 104 is used to classify the user question. The word order graph matching module 105 is used to fill the entities after entity recognition and linking into the word order graph template to form a complete query logic graph. The answer generation module 106 is used to convert the constructed query logic graph into a query statement for the graph database, query in the graph database storing the deep knowledge graph, and obtain the answer to the question. By adopting this technical solution, the intelligent level in the communication field can be improved, the competitiveness of communication operators and service providers can be enhanced, and a better communication service experience can be provided for users.

[0094] The above-disclosed is only a preferred embodiment of the present invention. Of course, it cannot be used to limit the scope of the rights of the present invention. Those of ordinary skill in the art can understand the entire or partial processes of implementing the above embodiments, and the equivalent changes made according to the claims of the present invention still fall within the scope covered by the present invention.

Claims

1. A communication customer service question-answering system based on deep knowledge graph, characterized in that: It includes a deep knowledge graph construction module, an entity recognition module, a link module, a question classification module, a syntax graph matching module and an answer generation module. The deep knowledge graph construction module is used to establish a deep knowledge graph. The entity recognition module is used to identify entities in user questions. The link module is used to obtain vectors based on the CNN model for entities in user questions and candidate entities in the deep knowledge graph, and select entity links. The question classification module is used to pre-process user questions, represent the questions as feature vectors based on the CNN model, and classify them using the naive Bayes algorithm to obtain the category label to which the questions belong. The syntax graph matching module fills the entities after entity recognition and linking into the syntax graph template based on the category labels obtained by the question classification module to form a complete query logic graph. The answer generation module is used to convert the constructed query logic graph into a query statement of a graph database, query in the graph database storing the deep knowledge graph, and obtain the answer to the question.

2. The communication customer service question and answer system based on deep knowledge graph according to claim 1, characterized in that: The deep knowledge graph construction module includes a data source selection unit, a data collection unit, a knowledge extraction unit, a knowledge graph pattern layer design unit, a knowledge fusion unit and a knowledge storage unit. The data source selection unit is used to integrate multi-source data, the data collection unit is used to collect multi-source data, the knowledge extraction unit extracts knowledge by combining natural language processing technology with domain-specific rules, the knowledge graph pattern layer design unit is used to design the pattern layer of the knowledge graph, the knowledge fusion unit fuses knowledge through an entity alignment strategy to ensure the consistency and accuracy of knowledge, and the knowledge storage unit is used to select an appropriate graph database to store the deep knowledge graph.

3. The communication customer service question and answer system based on deep knowledge graph according to claim 2, characterized in that: The specific function of the entity recognition unit is: after the user inputs the question, it first uses the Chinese word segmentation tool to perform word segmentation processing, then combines the stop word list to remove irrelevant words, and then identifies the entity in the question by matching with the entity dictionary in the communication field and the named entity recognition model based on conditional random fields.

4. The communication customer service question and answer system based on deep knowledge graph according to claim 3, characterized in that: The specific content of the vector obtained by the CNN model is as follows: the entity text in the user question or knowledge graph is represented as a matrix X with a dimension of n×m, where n is the length of the text and m is the dimension of the word vector. For a text containing n words, each word x i are all represented as an m-dimensional vector, namely x i ∈R m , the input matrix X is expressed as: In the convolution layer, a convolution kernel (filter) W is used to extract features. Assume that the size of the convolution kernel is h×m, where h represents the number of words covered by the convolution kernel and its elements are w. ij , then the convolution operation can be expressed as Where f is the activation function and b is the bias term. By sliding the convolution kernel on the input matrix X, if k different convolution kernels are used, k feature maps will be obtained. Each feature map can be represented as C1(l=1,2,…,k). In order to reduce the dimension of the feature map, the maximum pooling operation is performed on each feature map C1, which can be represented as p1=max(C1). The pooled feature vectors are concatenated together to form a vector p=[p1,p2,…,p k ], whose dimension is k×1, and then it is mapped to the final output vector y through the fully connected layer. If the weight matrix of the fully connected layer is W f , the dimension is k×d, d is the dimension of the final output vector, and the bias term is b f , then: y=f(W f p+b f ).

5. The communication customer service question-and-answer system based on deep knowledge graph according to claim 4, characterized in that: The specific content of the link module performing entity linking through vectors is: The cosine similarity is calculated by using the vector y1 obtained by the CNN model for the entity in the user's question and the vector y2 obtained by the CNN model for the candidate entity in the knowledge graph: Entity links are selected based on the similarity size and the method of calculating semantic similarity.

6. The communication customer service question and answer system based on deep knowledge graph according to claim 5, characterized in that: The specific content of the question classification module is: after preprocessing the user's questions by word segmentation and stop word filtering, the questions are represented as feature vectors through the CNN model, and classified using the Naive Bayes algorithm to obtain the category label to which the questions belong.

7. The communication customer service question and answer system based on deep knowledge graph according to claim 3, characterized in that: The syntax graph matching module includes a syntax graph template library construction unit and a syntax graph matching unit. The syntax graph template library construction unit constructs a template library containing multiple syntax graph templates based on the communication field problem type and knowledge graph structure. The syntax graph matching unit is used to select the corresponding syntax graph template from the syntax graph template library through the category label obtained by the question classification module, and then fill the entity after entity recognition and linking into the syntax graph template to form a complete query logic diagram.

8. A communication customer service question and answer method based on deep knowledge graph, applied to the communication customer service question and answer system based on deep knowledge graph as claimed in claim 1, characterized in that: The steps include: The user inputs a question into the communication customer question-answering system based on the deep knowledge graph; The communication customer question-answering system based on the deep knowledge graph performs word segmentation based on the question input by the user; Using the entity recognition module to recognize the user question after word segmentation, and obtain the entity in the user question; Utilizing the linking module to select entity links between the entity in the user's question and the candidate entity in the deep knowledge graph in the system; Using the question classification module to classify user questions; Using the syntax graph matching module to fill the entities after entity recognition and linking into the syntax graph template to form a complete query logic graph; The answer generation module is used to convert the constructed query logic graph into a query statement of the graph database, and a query is performed in the graph database storing the deep knowledge graph to obtain the answer to the question.

Citation Information

Patent Citations

  • Knowledge graph-based interactive question and answer method and system

    CN107766483A

  • Intelligent question and answer method and system based on pet knowledge graph

    CN110209787A

  • General knowledge graph construction method and system for specific field

    CN118485140A

Cited By

  • Communication field question and answer system fusing rag

    CN121456096A