Telecommunication field knowledge recall method and system

By combining knowledge graphs and deep learning methods, a knowledge retrieval system for the telecommunications field was constructed, which solved the problems of low accuracy, slow response speed and poor scalability of traditional telecommunications knowledge retrieval systems, and achieved efficient and intelligent knowledge management and retrieval.

CN120892580APending Publication Date: 2025-11-04CHINA TELECOM CORP LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511008540.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-22
Publication Date
2025-11-04

AI Technical Summary

Technical Problem

Traditional telecommunications knowledge retrieval systems suffer from insufficient accuracy, slow response speed, and poor scalability, especially under the complex semantic relationships and large-scale knowledge bases in the telecommunications field.

Method used

This approach combines knowledge graphs with deep learning. It constructs a knowledge graph through data preprocessing, entity recognition, and relation extraction. It then uses user intent recognition and semantic matching for knowledge retrieval and optimizes the system through an intelligent update module to achieve self-adaptation and dynamic expansion.

Benefits of technology

It significantly improves the accuracy and response speed of knowledge retrieval, enhances the scalability and user experience of the system, can accurately match user needs under complex and fuzzy queries, and supports dynamic updates and self-learning.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120892580A_ABST
    Figure CN120892580A_ABST
Patent Text Reader

Abstract

The invention discloses a telecom field knowledge recall method and system. The method comprises the steps that word segmentation is conducted on a data text, feature vectors are extracted, and entities and relation sets in knowledge are obtained through recognition; constructing the preprocessed data into a knowledge graph; each node represents a telecommunication knowledge entity, and each edge represents the relationship between two entities; converting a query problem input by a user into an intention vector and a key entity set queried by the user; performing knowledge recall processing through graph search and semantic matching according to the data identified by the user intention to obtain a candidate knowledge set; the semantic similarity between the user intention vector and the knowledge graph node is calculated, and recalled knowledge entries are reordered according to the user query requirement and the intention vector, so that the entry with the highest relevancy is located in the front. According to the method, the query intention of the user can be accurately identified and extracted, and intelligent multi-level knowledge recall can be realized through combination of the knowledge graph and the semantic matching model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of knowledge management and retrieval technology, and in particular to a knowledge retrieval method and system in the telecommunications field. Background Technology

[0002] In modern telecommunications companies, with the continuous expansion of technology and business, a vast amount of professional knowledge has accumulated. This knowledge includes business processes, network configuration, equipment maintenance, signal optimization, and frequently asked questions. The content of telecommunications knowledge bases has the following characteristics: Vast and complex: Telecommunications knowledge covers everything from basic network principles to advanced signal processing technologies, with a wide range of dimensions and complex content. Unstructured and fragmented: Much knowledge base consists of unstructured data, stored in natural language, charts, etc., lacking a unified format. Professional and specific: The telecommunications industry has a large number of technical terms and unique business processes, which are not easily understood by non-industry personnel. Faced with these characteristics, traditional knowledge retrieval methods (such as keyword search) cannot accurately match user needs, leading to low information retrieval accuracy and poor user experience.

[0003] In the field of knowledge retrieval, traditional methods include keyword-based TF-IDF algorithms, vector space models, and semantic matching models. In recent years, the development of deep learning has led to the increasing application of knowledge graphs and semantic matching models in knowledge retrieval, but these methods still face the following challenges: Insufficient matching accuracy: Traditional methods struggle to capture the complex semantic relationships in the telecommunications domain, resulting in low accuracy in knowledge retrieval. Limited system scalability: The ever-expanding scale of telecommunications knowledge bases leads to decreased retrieval efficiency and system response speed. Insufficient intelligence: Many systems lack deep-level intelligent analysis capabilities and cannot automatically adapt to changes in user needs. Summary of the Invention

[0004] The purpose of this invention is to solve the problems of insufficient accuracy, slow response speed and poor scalability in traditional telecommunications knowledge retrieval systems, and to provide a method and system for knowledge retrieval in the telecommunications field, so as to improve the management and retrieval efficiency of telecommunications knowledge bases.

[0005] The technical solution adopted in this invention is:

[0006] A knowledge retrieval method in the telecommunications field includes the following steps:

[0007] Step 1, Knowledge Data Preprocessing: The data text used to build the knowledge base is segmented into words and feature vectors are extracted. The entities and relation sets in the knowledge are then identified through natural language processing techniques (such as Named Entity Recognition, NER).

[0008] Step 2, Knowledge Graph Construction: The preprocessed data is used to construct a knowledge graph, which is a directed graph structure consisting of nodes and edges; each node represents a telecommunications knowledge entity, and each edge represents the relationship between two entities;

[0009] Step 3, User Intent Recognition: Transform the user's query into structured data that the knowledge retrieval system can process; the structured data includes the user's query intent vector I and the key entity set K;

[0010] Step 4, Knowledge Retrieval: Based on the user intent, structured data is identified and retrieved through graph search and semantic matching to obtain a set of candidate knowledge.

[0011] Step 5, Knowledge Reordering: Calculate the semantic similarity between the user intent vector and the knowledge graph nodes, and reorder the recalled knowledge items according to the user's query needs and intent vector, so that the most relevant items are at the top.

[0012] Furthermore, step 1 specifically includes the following steps:

[0013] Step 1-1: Use regular expressions or machine learning word segmentation tools to break down knowledge content into words, while removing useless words and symbols.

[0014] Steps 1-2 use natural language processing technology to identify core entities in the knowledge base, such as "base station", "frequency band", "signal strength", etc., as well as the relationships between entities.

[0015] Furthermore, in steps 1-2, BERT is used to segment the input text and extract feature vectors, and NER data with entity labels is used for fine-tuning; finally, the fine-tuned data is used to generate entity labels and relation sets in knowledge through the trained BERT model.

[0016] Furthermore, step 2 specifically includes the following steps:

[0017] Step 2-1: Initialize the knowledge graph G = (V, E), where G represents the directed graph structure; V represents the set of nodes, i.e., the set of telecommunications knowledge entities; and E represents the set of edges.

[0018] Step 2-2, traverse each pair of entities (v) in the knowledge base. i ,v j When entity v i With entity v j There is a relationship r ij If ∈R, then add a path from vertex v to graph G. i Pointing to vertex v j The edge represents the relation r. ij Where R is a set of relations;

[0019] Steps 2-3 involve storing the constructed knowledge graph in the database.

[0020] Furthermore, step 3 specifically includes the following steps:

[0021] Step 3-1: Segment the input query Q into words and extract the query features;

[0022] Specifically, in step 3-1, before performing intent recognition, the user query is preprocessed, including stop word removal, stemming, and phrase segmentation, in order to standardize the user's input format.

[0023] Step 3-2: Input the query features into the BERT model to generate an intent vector; the intent vector can not only capture the semantic information of the user's query, but also understand fuzzy queries based on context analysis.

[0024] Step 3-3: Identify the set of key entities K associated with the user's intent from the user query by combining contextual information;

[0025] Specifically, by using an entity extraction algorithm that combines rule-based and deep learning, the set of entities most relevant to the user's intent is found from the user query. For example, when the user enters "query base station frequency band information", the two key entities "base station" and "frequency band" are extracted.

[0026] Steps 3-4 involve matching the query vector with a pre-trained intent classifier to determine the intent category of the query. Intent categories include "finding information" and "troubleshooting," which helps to accurately match user needs with relevant content in the knowledge base.

[0027] Steps 3-5: The classified intent vector I and the key entity set K are used as structured data to be processed (input to the knowledge recall module).

[0028] Specifically, the expression for calculating the intent vector I in step 3 is as follows:

[0029] I = VERT(Q);

[0030] Where Q = (w1, w2, ..., wm) represents the input query sentence, and m represents the total number of words after the sentence is segmented, i.e. the sequence length.

[0031] Furthermore, step 4 specifically includes the following steps:

[0032] Step 4-1: Perform path search in the knowledge graph based on the key entity set K to find related knowledge nodes;

[0033] Step 4-2: For each associated knowledge node, use the BERT model to calculate the similarity between the intent vector I and the node vector, and retain the top N knowledge items with the highest similarity; the similarity calculation expression is:

[0034]

[0035] Where V = v1,…,v i ,…,v n represents a node in the knowledge graph, and n represents the total number of nodes in the knowledge graph;

[0036] Step 4-3 returns the candidate knowledge set C formed by the retained knowledge entries.

[0037] Furthermore, in step 5, the reordering is achieved by calculating the semantic similarity between the intent vector and each knowledge item. A similarity threshold is set based on the similarity score, and finally, highly relevant knowledge items with a similarity greater than the set similarity threshold are returned.

[0038] Specifically, the trained BERT model is used to rank the recall results.

[0039] A knowledge retrieval system in the telecommunications field includes the following modules:

[0040] Data preprocessing module: Normalizes massive amounts of unstructured or semi-structured data in the telecommunications field to obtain a set of entities and relationships in the knowledge;

[0041] Specifically, the data preprocessing module is used for data cleaning, word segmentation, entity recognition, and relation extraction. Data cleaning includes removing noisy data, redundant and irrelevant information, ensuring the purity of the input data and facilitating subsequent knowledge graph construction and improved recall accuracy. Word segmentation uses tools such as Jieba, NLTK, or SpaCy to break the text into word sequences for subsequent entity recognition and relation extraction. For telecommunications terms and abbreviations (such as "5G" and "base station"), a custom dictionary can be used to ensure accuracy. Entity recognition utilizes Named Entity Recognition (NER) technology to extract key entities (such as equipment type, frequency band, signal strength, etc.) from the data. BERT or BiLSTM-CRF models are well-suited for domain-specific datasets, and annotation and training improve the accuracy of entity recognition in the telecommunications field. Relation extraction uses relation extraction models to discover relationships between entities, such as "base station - coverage - area," achieved through algorithms such as dependency parsing or convolutional neural networks (CNNs), ensuring accurate capture of semantic relationships between entities.

[0042] Knowledge graph construction module: As the core foundation of the entire recall system, it creates graph nodes for each entity based on the entity information (extracted by the data preprocessing module), adds edges to each entity pair based on the relationships extracted from the relation extraction, and marks the attributes of the edges to construct the knowledge graph; at the same time, it stores the constructed knowledge graph in the graph database;

[0043] Specifically, the knowledge graph is the core foundation of the entire retrieval system. It uses a graph structure to represent entities and relationships for semantic reasoning and relationship lookup, enabling efficient semantic reasoning and relationship lookup. During node creation, a graph node is created for each entity based on the entity information extracted by the data preprocessing module. Node attributes include name, type, and association information. Based on the characteristics of the telecommunications field, node types mainly include "device," "frequency band," and "signal type." During relationship generation, edges are added to each entity pair based on the relationships identified by the relationship extraction module, and the edge attributes are labeled. Triples (entity1, relation, entity2) can be used to represent relationships between nodes, such as (base station, coverage, area). The constructed knowledge graph can be stored in a graph database (such as Neo4j or JanusGraph) for fast retrieval and querying. Graph databases improve the query efficiency of complex relationships through optimized graph indexing mechanisms.

[0044] User intent recognition module: Extracts user query intent and key entities by analyzing the text of the user query;

[0045] Specifically, the user intent recognition module uses BERT-based intent recognition and key entities. The specific process is as follows: 1) Query preprocessing: Before intent recognition, the user query is preprocessed, including stop word removal, stemming, and phrase segmentation, to standardize the user's input format. 2) BERT semantic encoding: The query is semantically encoded using the BERT model to generate an intent vector. The intent vector not only captures the semantic information of the user query but also enables understanding of fuzzy queries based on contextual analysis. 3) Key entity extraction: By using an entity extraction algorithm that combines rule-based and deep learning, the set of entities most relevant to the intent is found from the user query. For example, when the user inputs "query base station frequency band information," the two key entities "base station" and "frequency band" are extracted. 4) Intent classification: The query vector is matched with a pre-trained intent classifier to determine the intent category of the query. Intent categories include "finding information," "troubleshooting," etc., which helps to accurately match user needs with relevant content in the knowledge base.

[0046] Knowledge Retrieval Module: Based on user intent, structured data is identified and knowledge retrieval processing is performed in the knowledge graph through graph search and semantic matching to find the knowledge points that best match the user query and form a candidate knowledge set;

[0047] Specifically, path search is performed based on the set of key entities in the knowledge graph to find the set of nodes associated with the user's query.

[0048] The knowledge reordering module calculates the semantic similarity between the user's intent vector and the knowledge graph nodes. Based on the user's actual query needs and intent vector, it reorders the recalled knowledge items so that the most relevant items are placed at the top, thus providing the optimal retrieval response content.

[0049] Specifically, BERT or Sentence-BERT models are used to calculate the semantic similarity between the user's intent vector and the knowledge graph nodes. For each node, its cosine similarity to the user's query intent is calculated, and the results are filtered, retaining only nodes with similarity higher than a threshold. For the recalled knowledge items, they are reordered based on the user's actual query needs and intent vector. The reordering algorithm employs a learned ranking algorithm (such as a Pairwise or Listwise LTR model) to prioritize the most relevant items. During the reordering process, a similarity score formula is introduced:

[0050] Score=w1·Sim(I,v)+w2·Rel(I,v);

[0051] Where Sim(I,v) represents the similarity between intent and knowledge node, Rel(I,v) represents the relevance of recalled items, and w1 and w2 are weight parameters.

[0052] Furthermore, the system also includes an intelligent update module: using scheduled tasks to periodically capture and organize new telecommunications domain knowledge, adding new knowledge entities and relationships to the knowledge graph, and automatically adjusting the recall strategy based on user feedback and behavior analysis to improve the recall effect of dynamic data.

[0053] Specifically, user clicks on the recall results are used as reward signals to continuously optimize recall accuracy and response speed; the reinforcement learning model's update strategy (using Q-learning or DQN (Deep Q-Network) models) is used for online learning, with user feedback data serving as reinforcement signals to improve the model's recall performance on dynamic data. The updated Q-value function formula is as follows:

[0054] Where Q(s,a) represents the Q-value of the state-action pair, α is the learning rate, γ is the discount factor, and r is the immediate reward; s represents the state, and a represents the action to be performed in state s. ′ This indicates the next state that is transitioned to after performing action a in the current state s.

[0055] This invention provides a highly intelligent solution for telecommunications knowledge management by combining deep learning with knowledge graphs.

[0056] This invention, employing the above technical solutions, offers the following advantages compared to existing technologies: 1) By combining knowledge graphs and deep learning models for intelligent knowledge retrieval, this invention significantly improves the accuracy of retrieval results. The combination of semantic association in the knowledge graph and intent recognition in deep learning enables the system to accurately match user needs even under complex and fuzzy queries, effectively improving retrieval accuracy. 2) By optimizing the data index and knowledge graph structure, this invention improves system response speed by more than 30% compared to traditional retrieval systems. Optimization based on the index structure significantly reduces query time, maintaining high efficiency even with large-scale knowledge bases. 3) The introduction of knowledge graphs and adaptive optimization modules in this invention simplifies and facilitates knowledge base content expansion, supporting dynamic updates and model self-learning. The system can adaptively optimize based on user needs and feedback, ensuring the scalability of the knowledge base and the long-term performance of the system. 4) Through the deep learning-based semantic analysis module of this invention, user-inputted fuzzy queries can be understood and parsed more accurately, effectively improving the system's ability to understand non-standardized queries and significantly enhancing the overall user experience.

[0057] This invention can not only accurately identify and extract user query intent, but also achieve intelligent multi-level knowledge retrieval by combining knowledge graphs with semantic matching models. Attached Figure Description

[0058] The present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments;

[0059] Figure 1 This is a flowchart illustrating a knowledge retrieval method in the telecommunications field according to the present invention. Detailed Implementation

[0060] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings.

[0061] In recent years, deep learning and knowledge graphs have provided new solutions for knowledge retrieval technology: Knowledge Graph Construction: Representing telecommunications knowledge as a network structure of nodes and relationships through knowledge graph construction can significantly improve knowledge manageability and system retrieval accuracy. Deep Learning Matching Models: Using deep learning models (such as BERT, GPT, etc.) can effectively improve semantic matching accuracy, providing richer contextual information support for knowledge retrieval. Knowledge Representation and Structuring: The structured representation of knowledge graphs enables multi-level reasoning of knowledge, providing more intelligent support for retrieval systems.

[0062] like Figure 1 As shown, this invention discloses a knowledge retrieval method in the telecommunications field, which includes the following steps:

[0063] Step 1, Knowledge Data Preprocessing: The data text used to build the knowledge base is segmented into words and feature vectors are extracted. The entities and relation sets in the knowledge are then identified through natural language processing techniques (such as Named Entity Recognition, NER).

[0064] Step 2, Knowledge Graph Construction: The preprocessed data is used to construct a knowledge graph, which is a directed graph structure consisting of nodes and edges; each node represents a telecommunications knowledge entity, and each edge represents the relationship between two entities;

[0065] Step 3, User Intent Recognition: Transform the user's query into structured data that the knowledge retrieval system can process; the structured data includes the user's query intent vector I and the key entity set K;

[0066] Step 4, Knowledge Retrieval: Based on the user intent, structured data is identified and retrieved through graph search and semantic matching to obtain a set of candidate knowledge.

[0067] Step 5, Knowledge Reordering: Calculate the semantic similarity between the user intent vector and the knowledge graph nodes, and reorder the recalled knowledge items according to the user's query needs and intent vector, so that the most relevant items are at the top.

[0068] Furthermore, step 1 specifically includes the following steps:

[0069] Step 1-1: Use regular expressions or machine learning word segmentation tools to break down knowledge content into words, while removing useless words and symbols.

[0070] Steps 1-2 use natural language processing technology to identify core entities in the knowledge base, such as "base station", "frequency band", "signal strength", etc., as well as the relationships between entities.

[0071] Furthermore, in steps 1-2, BERT is used to segment the input text and extract feature vectors, and NER data with entity labels is used for fine-tuning; finally, the fine-tuned data is used to generate entity labels and relation sets in knowledge through the trained BERT model.

[0072] Specifically, let the input text be T = (t1, t2, ..., t... n The entity label is L = (l1, l2, ..., l n The training objective is:

[0073]

[0074] Furthermore, step 2 specifically includes the following steps:

[0075] Step 2-1: Initialize the knowledge graph G = (V, E), where G represents the directed graph structure; V represents the set of nodes, i.e., the set of telecommunications knowledge entities; and E represents the set of edges.

[0076] Step 2-2, traverse each pair of entities (v) in the knowledge base. i ,v j When entity v i With entity v j There is a relationship r ij If ∈R, then add a path from vertex v to graph G. i Pointing to vertex v j The edge represents the relation r. ij Where R is a set of relations;

[0077] Steps 2-3 involve storing the constructed knowledge graph in the database.

[0078] Furthermore, step 3 specifically includes the following steps:

[0079] Step 3-1: Segment the input query Q into words and extract the query features;

[0080] Specifically, in step 3-1, before performing intent recognition, the user query is preprocessed, including stop word removal, stemming, and phrase segmentation, in order to standardize the user's input format.

[0081] Step 3-2: Input the query features into the BERT model to generate an intent vector; the intent vector can not only capture the semantic information of the user's query, but also understand fuzzy queries based on context analysis.

[0082] Step 3-3: Identify the set of key entities K associated with the user's intent from the user query by combining contextual information;

[0083] Specifically, by using an entity extraction algorithm that combines rule-based and deep learning, the set of entities most relevant to the user's intent is found from the user query. For example, when the user enters "query base station frequency band information", the two key entities "base station" and "frequency band" are extracted.

[0084] Steps 3-4 involve matching the query vector with a pre-trained intent classifier to determine the intent category of the query. Intent categories include "finding information" and "troubleshooting," which helps to accurately match user needs with relevant content in the knowledge base.

[0085] Steps 3-5: The classified intent vector I and the key entity set K are used as structured data to be processed (input to the knowledge recall module).

[0086] Specifically, the expression for calculating the intent vector I in step 3 is as follows:

[0087] I = BERT(Q);

[0088] Where Q = (w1, w2, ..., wm) represents the input query sentence, and m represents the total number of words after the sentence is segmented, i.e. the sequence length.

[0089] Furthermore, step 4 specifically includes the following steps:

[0090] Step 4-1: Perform path search in the knowledge graph based on the key entity set K to find related knowledge nodes;

[0091] Step 4-2: For each associated knowledge node, use the BERT model to calculate the similarity between the intent vector I and the node vector, and retain the top N knowledge items with the highest similarity; the similarity calculation expression is:

[0092]

[0093] Where V = v1,…,v i ,…,v n represents a node in the knowledge graph, and n represents the total number of nodes in the knowledge graph;

[0094] Step 4-3 returns the candidate knowledge set C formed by the retained knowledge entries.

[0095] Furthermore, in step 5, the reordering is achieved by calculating the semantic similarity between the intent vector and each knowledge item. A similarity threshold is set based on the similarity score, and finally, highly relevant knowledge items with a similarity greater than the set similarity threshold are returned.

[0096] Specifically, the trained BERT model is used to rank the recall results.

[0097] A knowledge retrieval system in the telecommunications field includes the following modules:

[0098] Data preprocessing module: Normalizes massive amounts of unstructured or semi-structured data in the telecommunications field to obtain a set of entities and relationships in the knowledge;

[0099] Specifically, the data preprocessing module is used for data cleaning, word segmentation, entity recognition, and relation extraction. Data cleaning includes removing noisy data, redundant and irrelevant information, ensuring the purity of the input data and facilitating subsequent knowledge graph construction and improved recall accuracy. Word segmentation uses tools such as Jieba, NLTK, or SpaCy to break the text into word sequences for subsequent entity recognition and relation extraction. For telecommunications terms and abbreviations (such as "5G" and "base station"), a custom dictionary can be used to ensure accuracy. Entity recognition utilizes Named Entity Recognition (NER) technology to extract key entities (such as equipment type, frequency band, signal strength, etc.) from the data. BERT or BiLSTM-CRF models are well-suited for domain-specific datasets, and annotation and training improve the accuracy of entity recognition in the telecommunications field. Relation extraction uses relation extraction models to discover relationships between entities, such as "base station - coverage - area," achieved through algorithms such as dependency parsing or convolutional neural networks (CNNs), ensuring accurate capture of semantic relationships between entities.

[0100] Knowledge graph construction module: As the core foundation of the entire recall system, it creates graph nodes for each entity based on the entity information (extracted by the data preprocessing module), adds edges to each entity pair based on the relationships extracted from the relation extraction, and marks the attributes of the edges to construct the knowledge graph; at the same time, it stores the constructed knowledge graph in the graph database;

[0101] Specifically, the knowledge graph is the core foundation of the entire retrieval system. It uses a graph structure to represent entities and relationships for semantic reasoning and relationship lookup, enabling efficient semantic reasoning and relationship lookup. During node creation, a graph node is created for each entity based on the entity information extracted by the data preprocessing module. Node attributes include name, type, and association information. Based on the characteristics of the telecommunications field, node types mainly include "device," "frequency band," and "signal type." During relationship generation, edges are added to each entity pair based on the relationships identified by the relationship extraction module, and the edge attributes are labeled. Triples (entity1, relation, entity2) can be used to represent relationships between nodes, such as (base station, coverage, area). The constructed knowledge graph can be stored in a graph database (such as Neo4j or JanusGraph) for fast retrieval and querying. Graph databases improve the query efficiency of complex relationships through optimized graph indexing mechanisms.

[0102] User intent recognition module: Extracts user query intent and key entities by analyzing the text of the user query;

[0103] Specifically, the user intent recognition module uses BERT-based intent recognition and key entities. The specific process is as follows: 1) Query preprocessing: Before intent recognition, the user query is preprocessed, including stop word removal, stemming, and phrase segmentation, to standardize the user's input format. 2) BERT semantic encoding: The query is semantically encoded using the BERT model to generate an intent vector. The intent vector not only captures the semantic information of the user query but also enables understanding of fuzzy queries based on contextual analysis. 3) Key entity extraction: By using an entity extraction algorithm that combines rule-based and deep learning, the set of entities most relevant to the intent is found from the user query. For example, when the user inputs "query base station frequency band information," the two key entities "base station" and "frequency band" are extracted. 4) Intent classification: The query vector is matched with a pre-trained intent classifier to determine the intent category of the query. Intent categories include "finding information," "troubleshooting," etc., which helps to accurately match user needs with relevant content in the knowledge base.

[0104] Knowledge Retrieval Module: Based on user intent, structured data is identified and knowledge retrieval processing is performed in the knowledge graph through graph search and semantic matching to find the knowledge points that best match the user query and form a candidate knowledge set;

[0105] Specifically, path search is performed based on the set of key entities in the knowledge graph to find the set of nodes associated with the user's query.

[0106] And the knowledge reordering module: calculates the semantic similarity between the user's intent vector and the knowledge graph nodes, and reorders the recalled knowledge items according to the user's actual query needs and intent vector, so that the most relevant items are placed at the top, in order to provide the best retrieval response content.

[0107] Specifically, BERT or Sentence-BERT models are used to calculate the semantic similarity between the user's intent vector and the knowledge graph nodes. For each node, its cosine similarity to the user's query intent is calculated, and the results are filtered, retaining only nodes with similarity higher than a threshold. For the recalled knowledge items, they are reordered based on the user's actual query needs and intent vector. The reordering algorithm employs a learned ranking algorithm (such as a Pairwise or Listwise LTR model) to prioritize the most relevant items. During the reordering process, a similarity score formula is introduced:

[0108] Score=w1·Sim(I,v)+w2·Rel(I,v);

[0109] Where Sim(I,v) represents the similarity between intent and knowledge node, Rel(I,v) represents the relevance of recalled items, and w1 and w2 are weight parameters.

[0110] In the technical solution of this invention, the knowledge retrieval module and the re-ranking module are among the core technologies, aiming to provide efficient and accurate knowledge retrieval and information re-ranking services. The knowledge retrieval and re-ranking module of this invention combines Natural Language Processing (NLP) and machine learning technologies to retrieve the most relevant information from massive amounts of structured and unstructured data, and dynamically ranks it according to specific business needs, policy environment, and user goals. This technical solution can accurately and quickly provide decision-makers with the information they need in complex scenarios, improving business process efficiency and decision quality.

[0111] Unlike traditional knowledge retrieval methods, keyword-based retrieval systems are often limited by the accuracy of search terms and their relevance to actual needs, leading to unsatisfactory retrieval results. However, by employing deep learning and semantic understanding technologies, the knowledge retrieval and re-ranking module of this invention can intelligently retrieve and rank knowledge based on contextual information, user historical data, and real-time feedback. This helps governments and enterprises save significant time and improve task processing efficiency during information retrieval.

[0112] The beneficial effects of the knowledge retrieval and re-ranking modules are as follows: 1) Precise knowledge retrieval: Combining advanced semantic analysis models (such as BERT, GPT, etc.), it can accurately retrieve relevant content from massive amounts of information such as government policy documents, corporate reports, and big data, ensuring that the returned information is highly relevant to the current task and greatly improving the accuracy of information retrieval. Intelligent re-ranking: By introducing an intelligent re-ranking mechanism, this module not only ranks the retrieval results based on user needs, but also adjusts the content according to specific business needs (such as policy priorities, corporate strategic directions, etc.), ensuring that decision-makers obtain the most valuable information. 2) High-efficiency processing capabilities: Under large-scale data processing and high-concurrency requests, this module can still guarantee a fast response and flexibly handle the retrieval needs of different data sources (such as text, images, tables, etc.), adapting to the diverse application scenarios of governments and enterprises. 3) Dynamic learning and optimization: Through real-time data feedback, the module can continuously learn and optimize retrieval and ranking strategies, ensuring that it maintains high efficiency and accuracy in a changing environment, helping governments and enterprises cope with rapidly changing needs and challenges.

[0113] Furthermore, the system also includes an intelligent update module: using scheduled tasks to periodically capture and organize new telecommunications domain knowledge, adding new knowledge entities and relationships to the knowledge graph, and automatically adjusting the recall strategy based on user feedback and behavior analysis to improve the recall effect of dynamic data.

[0114] Specifically, user clicks on the recall results are used as reward signals to continuously optimize recall accuracy and response speed; the reinforcement learning model's update strategy (using Q-learning or DQN (Deep Q-Network) models) is used for online learning, with user feedback data serving as reinforcement signals to improve the model's recall performance on dynamic data. The updated Q-value function formula is as follows:

[0115] Where Q(s,a) represents the Q-value of the state-action pair, α is the learning rate, γ is the discount factor, and r is the immediate reward; s represents the state, and a represents the action to be performed in state s. ′ This indicates the next state that is transitioned to after performing action a in the current state s.

[0116] The intelligent update module integrates dynamic learning and adaptive optimization algorithms, enabling the system to automatically update and optimize based on real-time data, feedback, and emerging needs without human intervention. The goal of the intelligent update module is to allow the system to continuously adapt to new challenges and tasks, providing dynamic and personalized support. Through intelligent updates, the system can quickly adjust relevant knowledge bases, models, and strategies in the context of government policy adjustments and changes in corporate strategies, ensuring the real-time nature and accuracy of services and supporting governments and enterprises in efficiently responding to various policy, market, and social changes.

[0117] The intelligent update module introduces an automated optimization mechanism that automatically adjusts system parameters and strategies based on real-time data without manual intervention, ensuring optimal system performance over long-term use. This mechanism enables AI applications in government and enterprises to continuously adapt to changing environments and needs, ensuring efficiency and stability. Through the intelligent update algorithm, the system can acquire external data feedback in real time and update the model, quickly responding to changes in the external environment such as government policy changes and adjustments in enterprise needs. The intelligent update module is highly flexible, automatically adjusting optimization strategies according to the specific needs of different government departments and enterprises. Whether it's policy updates, enterprise strategy adjustments, or changes in the market environment, the system can react quickly, maintaining robustness and adaptability.

[0118] This invention provides a highly intelligent solution for telecommunications knowledge management by combining deep learning with knowledge graphs.

[0119] This invention, employing the above technical solutions, offers the following advantages compared to existing technologies: 1) By combining knowledge graphs and deep learning models for intelligent knowledge retrieval, this invention significantly improves the accuracy of retrieval results. The combination of semantic association in the knowledge graph and intent recognition in deep learning enables the system to accurately match user needs even under complex and fuzzy queries, effectively improving retrieval accuracy. 2) By optimizing the data index and knowledge graph structure, this invention improves system response speed by more than 30% compared to traditional retrieval systems. Optimization based on the index structure significantly reduces query time, maintaining high efficiency even with large-scale knowledge bases. 3) The introduction of knowledge graphs and adaptive optimization modules in this invention simplifies and facilitates knowledge base content expansion, supporting dynamic updates and model self-learning. The system can adaptively optimize based on user needs and feedback, ensuring the scalability of the knowledge base and the long-term performance of the system. 4) Through the deep learning-based semantic analysis module of this invention, user-inputted fuzzy queries can be understood and parsed more accurately, effectively improving the system's ability to understand non-standardized queries and significantly enhancing the overall user experience.

[0120] This invention can not only accurately identify and extract user query intent, but also achieve intelligent multi-level knowledge retrieval by combining knowledge graphs with semantic matching models.

[0121] Obviously, the described embodiments are only a part of the embodiments of this application, not all of them. Without conflict, the embodiments and features in the embodiments of this application can be combined with each other. The components of the embodiments of this application described and illustrated herein can generally be arranged and designed in various different configurations. Therefore, the detailed description of the embodiments of this application is not intended to limit the scope of the claimed application, but merely to illustrate selected embodiments of this application. All other embodiments obtained by those skilled in the art based on the embodiments of this application without inventive effort are within the scope of protection of this application.

Claims

1. A knowledge retrieval method in the telecommunications field, characterized in that: It includes the following steps: Step 1, Knowledge Data Preprocessing: The data text used to build the knowledge base is segmented into words and feature vectors are extracted. Natural language processing technology is then used to identify the entity and relation set in the knowledge. Step 2, Knowledge Graph Construction: The preprocessed data is used to construct a knowledge graph, which is a directed graph structure consisting of nodes and edges; each node represents a telecommunications knowledge entity, and each edge represents the relationship between two entities; Step 3, User Intent Recognition: Transform the user's query into structured data that the knowledge retrieval system can process; the structured data includes the user's query intent vector I and the key entity set K; Step 4, Knowledge Retrieval: Based on the user intent, structured data is identified and retrieved through graph search and semantic matching to obtain a set of candidate knowledge. Step 5, Knowledge Reordering: Calculate the semantic similarity between the user intent vector and the knowledge graph nodes, and reorder the recalled knowledge items according to the user's query needs and intent vector, so that the most relevant items are at the top.

2. The telecommunications knowledge retrieval method according to claim 1, characterized in that: Step 1 specifically includes the following steps: Step 1-1: Use regular expressions or machine learning word segmentation tools to break down knowledge content into words, while removing useless words and symbols. Steps 1-2 involve using natural language processing technology to identify the core entities in the knowledge base and the relationships between them.

3. The telecommunications knowledge retrieval method according to claim 2, characterized in that: In steps 1-2, BERT is used to segment the input text and extract feature vectors, and then fine-tuned using NER data with entity labels. Finally, the fine-tuned data is used to generate entity labels and relation sets in the knowledge base through the trained BERT model.

4. The telecommunications knowledge retrieval method according to claim 1, characterized in that: Step 2 specifically includes the following steps: Step 2-1: Initialize the knowledge graph G = (V, E), where G represents the directed graph structure; V represents the set of nodes, i.e., the set of telecommunications knowledge entities; and E represents the set of edges. Step 2-2, traverse each pair of entities (v) in the knowledge base. i ,v j When entity v i With entity v j There is a relationship r ij If ∈R, then add a path from vertex v to graph G. i Pointing to vertex v j The edge represents the relation r. ij Where R is a set of relations; Steps 2-3 involve storing the constructed knowledge graph in the database.

5. A knowledge retrieval method in the telecommunications field according to claim 1, characterized in that: Step 3 specifically includes the following steps: Step 3-1: Segment the input query Q into words and extract the query features; Step 3-2: Input the query features into the BERT model to generate an intent vector; Step 3-3: Identify the set of key entities K associated with the user's intent from the user query by combining contextual information; Steps 3-4: Match the query vector with the pre-trained intent classifier to obtain the intent category of the query; Steps 3-5: Use the classified intent vector I and the key entity set K as structured data to be processed.

6. A knowledge retrieval method in the telecommunications field according to claim 1, characterized in that: Step 4 specifically includes the following steps: Step 4-1: Perform path search in the knowledge graph based on the key entity set K to find related knowledge nodes; Step 4-2: For each associated knowledge node, use the BERT model to calculate the similarity between the intent vector I and the node vector, and retain the top N knowledge items with the highest similarity. The similarity calculation expression is: Where V = v1,…,v i ,…,v n represents a node in the knowledge graph, and n represents the total number of nodes in the knowledge graph; Step 4-3 returns the candidate knowledge set C formed by the retained knowledge entries.

7. A knowledge retrieval method in the telecommunications field according to claim 1, characterized in that: In step 5, the reordering is achieved by calculating the semantic similarity between the intent vector and each knowledge item. A similarity threshold is set based on the similarity score, and finally, highly relevant knowledge items with a similarity greater than the set similarity threshold are returned.

8. A telecommunications domain knowledge retrieval system, comprising a telecommunications domain knowledge retrieval method according to any one of claims 1 to 7, characterized in that: The system includes the following modules: Data preprocessing module: Normalizes massive amounts of unstructured or semi-structured data in the telecommunications field to obtain a set of entities and relationships in the knowledge; Knowledge graph construction module: As the core foundation of the entire recall system, it creates graph nodes for each entity based on entity information, adds edges to each entity pair based on the relationships extracted from the relationships, and marks the attributes of the edges to construct the knowledge graph; at the same time, it stores the constructed knowledge graph in a graph database; User intent recognition module: Extracts user query intent and key entities by analyzing the text of the user query; Knowledge Retrieval Module: Based on user intent, structured data is identified and knowledge retrieval processing is performed in the knowledge graph through graph search and semantic matching to find the knowledge points that best match the user query and form a candidate knowledge set; The knowledge reordering module calculates the semantic similarity between the user's intent vector and the knowledge graph nodes. Based on the user's actual query needs and intent vector, it reorders the recalled knowledge items so that the most relevant items are placed at the top, thus providing the optimal retrieval response content.

9. A knowledge retrieval system in the telecommunications field according to claim 8, characterized in that: The knowledge re-ranking module incorporates similarity scores during the re-ranking process and filters the recalled knowledge entries based on these scores, retaining only nodes with similarity scores above a threshold. The formula for calculating the similarity score is as follows: Score=w1·Sim(I,v)+w2·Rel(I,v); Where Sim(I,v) represents the similarity between intent and knowledge node, Rel(I,v) represents the relevance of recalled items, and w1 and w2 are weight parameters.

10. A knowledge retrieval system in the telecommunications field according to claim 8, characterized in that: The system also includes an intelligent update module: it uses scheduled tasks to periodically capture and organize new telecommunications domain knowledge, adds new knowledge entities and relationships to the knowledge graph, and automatically adjusts the recall strategy based on user feedback and behavior analysis to improve the recall effect of dynamic data.