An Interactive Intelligent Analysis Method, Device, and Medium Based on a Knowledge Graph
Through the interactive intelligent analysis method based on knowledge graph, the problem of traditional analysis tools having high requirements for user input content and being unable to recommend correlation analysis dimensions is solved, and intelligent guidance from natural language query to visual analysis results is realized, which improves analysis efficiency and user experience.
Patent Information
- Application Number
- CN202510465832.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-15
- Publication Date
- 2025-07-01
- Estimated Expiration
- 2045-04-15
AI Technical Summary
The traditional interactive intelligent analysis process requires high requirements for the input content in the user interaction process, and cannot recommend the correlation analysis dimension based on business logic, resulting in inefficient analysis.
Using an interactive intelligent analysis method based on knowledge graph, the user's natural language query text is obtained, intent detection and fuzzy confidence determination are performed, and the target query channel is matched. In the boot query channel, the business knowledge graph constructed by pre-learning is used for path derivation, generate a personalized analysis card sequence, switch to the precise query channel, call the query mode knowledge base to generate target query statements, and determine the visual analysis results.
It lowers the threshold for non-technical personnel to use business intelligence tools, improves human-machine collaboration efficiency, realizes intelligent guidance from fuzzy needs to precise analysis, actively recommends correlation analysis dimensions for users, and improves analysis efficiency.
Smart Images

Figure CN120045686B_ABST
Abstract
Description
Technical Field
[0001] This specification relates to the technical field of data processing, and particularly to an interactive intelligent analysis method, device, and medium based on a knowledge graph. Background Art
[0002] There is a significant gap between user cognition and system capabilities in the current business intelligence field. Traditional business intelligence (BI) tools rely on users to pre-construct a complete analysis idea and require mastery of professional query languages such as SQL, resulting in too high a usage threshold for non-technical personnel. Existing interactive analysis tools show a polarized situation. Basic question-and-answer systems only support simple instructions (such as "query sales amount") and cannot handle complex conversations; while systems with high-order analysis capabilities require users to forcibly confirm parameters such as analysis dimensions and filtering conditions step by step, breaking down natural semantics into discrete input sequences, destroying the interaction fluency, resulting in low human-machine collaboration efficiency, and further increasing the usage threshold. Traditional systems lack the ability to actively derive analysis paths. When users put forward vague requests, they need to rely on manual experience to explore step by step, resulting in low analysis efficiency. In summary, the traditional interactive intelligent analysis process has high requirements for the input content in the user interaction process, and cannot recommend associated analysis dimensions based on business logic for fuzzy input scenarios, resulting in low analysis efficiency. Summary of the Invention
[0003] One or more embodiments of this specification provide an interactive intelligent analysis method, device, and medium based on a knowledge graph to solve the following technical problems: The traditional interactive intelligent analysis process has high requirements for the input content in the user interaction process, and cannot recommend associated analysis dimensions based on business logic for fuzzy input scenarios, resulting in low analysis efficiency.
[0004] One or more embodiments of this specification adopt the following technical solutions:
[0005] One or more embodiments of this specification provide an interactive intelligent analysis method based on a knowledge graph. The method includes: obtaining a natural language query text input by a user, performing intent detection on the natural language query text to determine a corresponding intent ambiguity confidence level, and based on the intent ambiguity confidence level, matching a target query channel; when the target query channel is a guided query channel, performing path derivation on the natural language query text through a pre-learned business knowledge graph to generate a personalized analysis card sequence, so as to obtain a target analysis dimension in the personalized analysis card sequence triggered by the user; under the trigger of the target analysis dimension, switching the guided query channel to an accurate query channel, and through the accurate query channel, calling a pre-learned query mode knowledge base to generate a target query statement corresponding to the target analysis dimension, so as to determine a corresponding visual analysis result through the target query statement.
[0006] One or more embodiments of this specification provide an interactive intelligent analysis device based on a knowledge graph, including:
[0007] At least one processor; and,
[0008] A memory communicatively connected to the at least one processor; wherein,
[0009] The memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the above method.
[0010] A non-volatile computer storage medium provided by one or more embodiments of this specification stores computer-executable instructions, and the computer-executable instructions are set to: execute the above method.
[0011] One or more of the above technical solutions adopted in the embodiments of this specification can achieve the following beneficial effects: Through the technical solutions in the embodiments of this specification, users do not need to master professional query languages such as SQL, and only need to input natural language query texts, which reduces the threshold for non-technical personnel to use business intelligence tools; By performing intent detection and fuzzy confidence determination on natural language query texts, the target query channel can be automatically matched. In the guided query channel, the business knowledge graph is used for path derivation to generate a personalized analysis card sequence, without the need for users to confirm parameters such as analysis dimensions and filtering conditions step by step, avoiding disassembling natural semantics into discrete input sequences, thus maintaining the fluency of the interaction and improving the human-computer collaboration efficiency; After the user triggers the target analysis dimension in the personalized analysis card sequence, switch to the precise query channel, call the query mode knowledge base to generate the target query statement, and determine the visual analysis result. From the guided to the precise query method, it not only allows users to get guidance and inspiration in the stage of fuzzy requirements, but also can quickly obtain precise visual analysis results after the requirements are clarified. The entire interaction process is natural and fluent, conforming to the user's thinking habits; Based on the dynamic path exploration of the business knowledge graph, it can actively recommend associated analysis dimensions for users, without relying on manual experience for gradual exploration, greatly improving the analysis efficiency; When the user determines the target analysis dimension, the precise query channel uses the pre-learned query mode knowledge base to generate the target query statement and quickly obtains the visual analysis result, reducing the user's waiting time and further improving the analysis efficiency. BRIEF DESCRIPTION OF THE DRAWINGS
[0012] To more clearly illustrate the technical solutions in the embodiments of this specification or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments recorded in this specification. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings. In the drawings:
[0013] Figure 1 It is a schematic flowchart of an interactive intelligent analysis method based on a knowledge graph provided by an embodiment of this specification;
[0014] Figure 2 It is a schematic structural diagram of an interactive intelligent analysis device based on a knowledge graph provided by an embodiment of this specification. Detailed implementation manners
[0015] To enable those skilled in the art to better understand the technical solutions in this specification, the following will clearly and completely describe the technical solutions in the embodiments of this specification with reference to the drawings in the embodiments of this specification. Obviously, the described embodiments are only a part of the embodiments of this specification, rather than all of them. Based on the embodiments of this specification, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of this specification.
[0016] An embodiment of this specification provides an interactive intelligent analysis method based on a knowledge graph. It should be noted that the execution subject in the embodiments of this specification can be a server or any device with data processing capabilities. Figure 1 It is a schematic flowchart of an interactive intelligent analysis method based on a knowledge graph provided by an embodiment of this specification. As Figure 1 shown, it mainly includes the following steps:
[0017] Step S101: Obtain the natural language query text input by the user, perform intent detection on the natural language query text, determine the corresponding intent ambiguity confidence, and match the target query channel according to the intent ambiguity confidence.
[0018] The intention detection is performed on the natural language query text to determine the corresponding intention fuzzy confidence, specifically including: performing word segmentation on the natural language query text to determine the word vector features, and performing syntactic parsing on the natural language query file to extract syntactic structure features; determining the word vector sequence corresponding to the word vector features to capture the bidirectional context dependency; utilizing a multi-head attention mechanism to fuse the bidirectional context dependency and the syntactic structure features, and inputting them into a pre-trained BERT model to obtain a CLS tag vector, so as to determine the intention fuzzy confidence corresponding to the natural language query text through the CLS tag vector.
[0019] In one embodiment of the present specification, the intention ambiguity detection model based on deep learning is used to evaluate the clarity of user input. The model is based on the pre-trained BERT model as the encoder, and a specific classification layer structure is added to achieve an accurate assessment of the clarity of the user's query intention. The model architecture uses BERT-base-uncased as the basic encoder, which includes a 12-layer transformer structure, a 768-dimensional hidden layer and 12 attention heads, and is adapted to the characteristics of the BI query scenario through targeted fine-tuning. On top of the BERT encoding layer, a bidirectional GRU layer is designed with a 256-dimensional hidden state to effectively capture contextual dependencies, and then a multi-head self-attention mechanism is connected, and 8 attention heads are configured to further extract and strengthen key semantic features. During the training process, more than 10,000 annotated samples are used, including two types of data: clear queries and fuzzy queries, and the data set is divided into a training set, a validation set and a test set at a ratio of 7:1:2. The data preprocessing stage includes text standardization, word segmentation, and conversion of text into a BERT vocabulary index, while generating a fuzziness label with a continuous value of 0-1. The model training adopts the mini-batch gradient descent method with a batch size of 32 and uses the Adam W optimizer. The corresponding initial learning rate here is 2e-5 and the weight decay is 0.01. The learning rate scheduling strategy combines linear warm-up (accounting for the first 10% of the training steps) and cosine decay mechanism. To avoid overfitting, an early stopping mechanism is introduced, and the training is automatically terminated when the performance of the validation set has not improved significantly within 5 consecutive epochs. Although the maximum number of training rounds is set to 30, in practice the model usually reaches convergence within 15-20 rounds. In terms of key parameter settings, the maximum length of the input sequence is limited to 128 tokens, the dropout rate of the hidden layer and the attention mechanism is 0.1, the loss function uses binary cross entropy loss, and the gradient clipping threshold is set to 1.0 to ensure training stability. The model finally calculates the confidence score of the user's query intent clarity through the sigmoid function. The specific calculation formula is: ,in The 768-dimensional vector representation corresponding to the [CLS] tag of BERT, is a weight matrix (768×1), is a bias term, represents the Sigmoid activation function, enabling the model to output continuous values between 0 and 1. The confidence score of intention ambiguity is obtained by subtracting the confidence score of clear intention from 1, which precisely quantifies the clarity of the query intention. According to the confidence score of intention ambiguity, the target query channel is matched. A channel selection threshold is preset. When the confidence score of intention ambiguity is greater than 0.3, it indicates that the clarity of intention is less than 0.7, and the user input is determined to be a fuzzy intention, and the target query channel is determined to be the guided query channel; correspondingly, when the confidence score of intention ambiguity is not greater than 0.3, the target query channel is determined to be the precise query channel.
[0020] By performing word segmentation on the natural language query text to determine the word vector features and conducting syntactic parsing to extract the syntactic structure features, the multi-dimensional analysis method can comprehensively understand the semantic information of the text, laying a foundation for accurately detecting the user's intention. Compared with the intention detection methods that rely only on a single feature, such as relying only on word vectors or only on syntactic structures, it can more accurately grasp the user's true intention and reduce the misjudgment of intention caused by information loss; determine the word vector sequence corresponding to the word vector features, thereby capturing the bidirectional context dependency relationship, enabling the model to consider the front and back association information between the words in the text; use the multi-head attention mechanism to fuse the bidirectional context dependency relationship and the syntactic structure features, which can fully integrate different types of feature information and play a synergistic role. At the same time, relying on the pre-training results of the BERT model on a large-scale corpus, it can learn rich language knowledge and semantic representations, more accurately obtain the CLS token vector, and then determine the confidence score of intention ambiguity through this vector, which can more reasonably reflect the degree of intention ambiguity of the user input text. According to different intention ambiguity situations, select the appropriate query channel, such as the guided query channel or the precise query channel, to improve the efficiency and effect of the system in processing user queries and provide more personalized and demand-compliant services for users.
[0021] Step S102, when the target query channel is the guided query channel, through the business knowledge graph constructed by pre-learning, perform path derivation on the natural language query text to generate a personalized analysis card sequence, so as to obtain the target analysis dimension in the personalized analysis card sequence triggered by the user.
[0022] Before generating a personalized analysis card sequence by performing path derivation on the natural language query text using the business knowledge graph constructed through pre-learning, the method further includes: obtaining historical query information, and automatically extracting all business data tables through a preset database metadata interface, where the all business data tables include table information, field information, and data types; constructing a multi-level knowledge base according to the all business data tables and the historical query information through a preset pre-learning framework, where the multi-level knowledge base includes a business knowledge graph, a business field semantic knowledge base, and a query pattern knowledge base. Constructing a multi-level knowledge base according to the all business data tables and the historical query information through a preset pre-learning framework specifically includes: constructing initial schema graph data according to the all business data tables, and performing edge weight assignment on the initial schema graph according to the historical query frequency in the historical query information to determine business schema graph data; determining the business knowledge graph of the table-level data structure according to a pre-trained deep graph learning model and the business schema graph data; extracting multi-modal semantic features of fields in the all business data tables through a field-level semantic modeling layer, and constructing a multi-dimensional similarity matrix based on the multi-modal semantic features, where the multi-dimensional similarity matrix includes field name similarity, data distribution similarity, and content overlap; determining field semantic vectors using an attention-enhanced autoencoder according to the field feature descriptors obtained in advance in the all business data tables and the multi-dimensional similarity matrix to construct a business field semantic knowledge base; obtaining high-frequency query pattern data in the historical query information to transform the query statements in the high-frequency query pattern data and determine statement syntax trees; performing semantic transformation on the statement syntax trees through a preset semantic encoder to determine the query semantic information of the query statements, and constructing a query pattern knowledge base according to the query semantic information and the field semantic vectors, where the query pattern knowledge base includes multiple query statement templates and the query semantic information corresponding to each query statement template; determining a multi-level knowledge base according to the business knowledge graph of the table-level data structure, the business field semantic knowledge base, and the query pattern knowledge base.
[0023] In one embodiment of this specification, a dataset pre-training framework is adopted. Through the integration of deep learning and knowledge graph technology, all-round semantic modeling of enterprise data assets is realized. The underlying foundation of the pre-training framework is the table-level knowledge acquisition layer, which uses an improved graph structure learning algorithm to model the database schema. First, complete table structure information is obtained through the database metadata API, including table names, field names, data types, primary and foreign key relationships, etc. Subsequently, an initial schema graph is constructed, where nodes represent tables or fields, and edges represent inter-table relationships (such as foreign key constraints, JOIN paths, etc.). A usage-frequency-based edge weight calculation mechanism is introduced. By analyzing historical query logs, weight values are assigned to different inter-table relationships to reflect their importance in actual business. It should be noted that the table-level knowledge acquisition layer uses a deep graph learning model based on R-GCN (Relational Graph Convolutional Network) for training. The input is the original schema graph, and the output is the low-dimensional vector representation of each table and field. The dimension here can be 256. The training process uses an edge prediction task as the self-supervised learning objective, and the model parameters are optimized by minimizing the triple loss function. Specifically, for each known relationship (table A, relationship type, table B), the model needs to predict the most likely table B given (table A, relationship type). A business knowledge graph of the table-level data structure is obtained through the table-level knowledge acquisition layer.
[0024] The middle layer is the field-level semantic modeling layer, which is used to capture the semantic associations between fields and solve the problems of term differences and concept mapping. First, the actual data content of each field is analyzed by sampling to generate field feature descriptors. 1000 records are randomly selected for each field, and for large tables, sampling is performed proportionally. Through statistical analysis, data distribution characteristics are extracted, including data type statistics, distribution parameters (the distribution parameters can be mean, variance, quantiles, etc.), value patterns, and uniqueness indicators, etc. For text fields, TF-IDF and Word2Vec technologies are also applied to extract semantic features. On this basis, a field similarity matrix is constructed, and an innovative multi-modal similarity calculation method is adopted, comprehensively considering field name similarity (based on Word2Vec embedding), data distribution similarity (based on Jensen-Shannon divergence of distribution parameters), and content overlap (based on the min-hash algorithm). The core model of the field-level semantic modeling layer adopts an attention-enhanced autoencoder architecture. The input is the field feature descriptor and the similarity matrix, and the output is a 512-dimensional field semantic vector to construct a business field semantic knowledge base. The training uses a weighted combination of reconstruction loss and contrastive learning loss. The former ensures the fidelity of the encoded information, and the latter ensures the proximity of semantically similar fields in the embedding space.
[0025] The top layer is the query pattern mining layer, which is used to extract high-frequency query patterns and business analysis paradigms from historical query data. First, a dedicated SQL parsing engine is established to convert the original SQL statements in the historical query data into an abstract syntax tree (AST). The parsing process is implemented using an improved ANTLR4 grammar parser, which supports the accurate extraction of complex SQL structures, including nested subqueries, window functions, and complex aggregations. Subsequently, the syntax trees are clustered using the tree edit distance algorithm to identify high-frequency query patterns. Then, based on the Transformer-based SQL semantic encoder, the SQL syntax tree is converted into a 768-dimensional vector representation to capture the semantic intent of the query rather than just focusing on the syntax structure. This encoder adopts a hierarchical attention mechanism to encode SQL components such as the SELECT clause, WHERE conditions, JOIN relationships, etc., and fuses the overall semantics through the self-attention mechanism. The model is trained using a contrastive learning framework, where the positive sample pairs are different SQL expressions with equivalent semantics, such as variants generated by SQL normalization tools, and the negative samples are SQLs with different semantic intents. The training data contains 500,000 SQL query records from the actual system, which are preprocessed, normalized, and enhanced before being used for training. Through the query pattern mining layer, a query pattern knowledge base is constructed. It should be noted that the query pattern knowledge base contains query statement templates and the corresponding query semantic information for each query statement template. Here, the query semantic information corresponds to the query business mode of the query statement.
[0026] It should be noted that the data flow between the three-layer architecture follows a bottom-up processing flow and information enhancement mechanism. The table-level knowledge layer first constructs a basic data structure graph, which serves as the input for the field-level semantic layer and jointly guides the learning of semantic vectors with the field content features. The field semantic vectors further serve as auxiliary features for the query pattern mining layer to enhance the SQL semantic understanding ability. In addition, there is not only a forward data flow between the three layers, but also a backward optimization mechanism is implemented. The commonly used table field combinations extracted from the query patterns are fed back to the field layer to optimize the judgment of field semantic relevance; while the newly discovered semantic associations between fields are also updated to the surface layer to enrich the graph structure. The entire framework maintains the latest state through an asynchronous update strategy. By default, the surface layer is updated once a week, the field layer is updated once every three days, and the query layer is updated daily to ensure that the system can adapt to changes in data patterns and business requirements in a timely manner.
[0027] The output of the final pre-training framework is a multi-level knowledge graph, which provides services in the form of a GraphQL interface. This interface supports three core query modes. First, given a natural language description, it returns the most relevant tables and fields (supporting Top-K queries); second, given a combination of some table fields, it recommends semantically related extended fields; third, given a description of the analysis intention, it recommends the most matching query pattern template. The interface layer implements an efficient vector retrieval mechanism, using the Hierarchical Navigable Small World (HNSW) algorithm to construct an approximate nearest neighbor index, with the average query latency controlled within 15 ms to meet the real-time interaction requirements. In addition, the framework deployment supports two modes, the online learning mode and the batch processing mode. In the online learning mode, the system can continuously learn from interaction feedback and update field semantics and query patterns in real time; the batch processing mode is applicable to scenarios of large-scale data schema changes, triggering a full-scale reconstruction. In practical applications, the two modes are usually combined, with batch updates executed regularly and incremental learning maintained daily, taking into account both system stability and adaptability.
[0028] Through the above multi-level pre-training architecture, a complete semantic portrait of the enterprise's data assets can be constructed, providing a solid knowledge foundation for subsequent natural language query understanding and SQL generation. The actual deployment results show that the introduction of the pre-training framework has increased the accuracy of the system in processing complex business queries by 35.2 percentage points, especially showing significant advantages in scenarios involving technical terms, implicit associations, and multi-table joint analysis.
[0029] Through the business knowledge graph constructed by pre-training, path derivation is performed on the natural language query text to generate a personalized analysis card sequence, specifically including: entity extraction is performed on the natural language query text to determine at least one corresponding business entity; centered on each such business entity, dynamic path exploration is performed in the business knowledge graph to determine a real-time exploration path, so as to obtain path parameters corresponding to the real-time exploration path, where the path parameters include the total number of connected edges of the path node in the business knowledge graph and the edge weight parameter corresponding to the exploration path; through the total number of connected edges and the edge weight parameter, the dynamic path score corresponding to the real-time exploration path is dynamically calculated, and based on the dynamic path score, path screening is performed to determine multiple target analysis paths; the user's historical interaction information is obtained, and based on the user's historical interaction information, user-dimensional matching is performed on the multiple target analysis paths to determine the recommendation index corresponding to each target analysis path; according to the recommendation index of each target analysis path and the pre-obtained recommended quantity corresponding to the user, the personalized analysis card sequence is generated, where the personalized analysis card sequence includes analysis dimensions corresponding to multiple specified target analysis paths that meet preset requirements.
[0030] When the target query channel is the guided query channel, the business knowledge graph constructed through pre-learning is used to perform path derivation on the natural language query text to generate a personalized analysis card sequence. In an embodiment of this specification, the natural language query text input by the user is parsed multi-dimensionally, and the domain entity recognition model is used to extract the core business entities. The entity recognition model can be built based on a bidirectional long short-term memory network and a conditional random field, and can accurately identify industry-specific terms such as "inventory turnover rate" and "user repurchase cycle". After the extracted entities are standardized, they are matched with the pre-constructed business knowledge graph to determine the node positions of each entity in the graph. For example, when the user inputs "Analyze the characteristics of best-selling products", "product" is extracted as the core entity and located at the "commodity entity" node of the knowledge graph. Centered on each business entity node, a bidirectional breadth-first search algorithm is used to explore multi-hop paths in the graph, and the maximum number of hops is set to three-level association. During the exploration process, path parameters are collected in real time, including node degree and edge weight parameters. Among them, the node degree is used to reflect the connection density of the current node in the graph, and the higher the degree, the more extensive the business relevance. For each exploration path, the path score is calculated dynamically, and the calculation formula is , where is the relationship weight,[[]]END]] i is the node degree of the i-th node. By comprehensively considering the positive contribution of the edge weight and the reverse adjustment effect of the node degree, it is ensured that paths with high weights and low interference obtain the priority recommendation qualification. Paths with scores lower than the preset threshold will be automatically filtered, and multiple high-value candidate paths will be retained. The user's historical interaction records are retrieved, and based on the user's historical interaction information, the user dimension matching is performed on the multiple target analysis paths to determine the recommendation index corresponding to each target analysis path. The multiple target analysis paths are sorted in descending order according to the recommendation index of each target analysis path, and the recommended number of target analysis paths corresponding to the user is obtained in sequence to generate the personalized analysis card sequence. Each card can contain the following elements. First is the analysis dimension description, which describes the analysis logic in business terms, such as "Regional sales comparison: Decompose sales by geographical level"; it also includes parameter controls and data previews. The parameter controls provide interactive adjustment components, such as time range selectors and index weight sliders; the data preview displays the sample data distribution chart under this dimension. The analysis dimension description describes the analysis logic in business terms, which is convenient for users to understand; the parameter controls provide interactive adjustment components, allowing users to make personalized settings according to their own needs; the data preview displays the sample data distribution chart under this dimension, enabling users to have an intuitive understanding of the data before conducting a detailed analysis. The personalized analysis card sequence can improve the user experience and help users obtain the required information more efficiently.
[0031] Through the above technical solution, the domain entity recognition model can accurately identify the core business entities in the natural language query text, especially industry-specific terms, standardize the extracted entities and match them with the pre-constructed business knowledge graph, and can determine the node positions of the entities in the graph, which helps to transform the user's natural language query into specific nodes and relationships in the knowledge graph. Centering on the business entity nodes, the bidirectional breadth-first search algorithm is used to explore multi-hop paths in the knowledge graph, which can comprehensively explore the paths related to the business entities within a reasonable range, avoid over-searching or missing important information. Through multi-hop path exploration, potential relationships and indirect associations between business entities can be discovered; during the exploration process, path parameters are collected in real time and path scores are dynamically calculated, taking into account both the positive contribution of edge weights and the reverse adjustment effect of node degrees, ensuring that paths with high weights and low interference obtain the priority recommendation qualification, effectively screening out the most valuable paths and reducing the workload of subsequent analysis; the historical interaction records of the user are retrieved, and the user dimension matching is performed on multiple target analysis paths based on the user's historical interaction information to determine the recommendation index corresponding to each target analysis path, so that the recommended analysis paths can fully consider the user's personalized needs and preferences, improving the accuracy and relevance of the recommendation; the target analysis paths are sorted in descending order of the recommendation index, and the recommended number of target analysis paths corresponding to the user are obtained in sequence to generate a personalized analysis card sequence. Through accurate entity extraction, efficient path exploration and screening, and personalized recommendation, the analysis of a large amount of irrelevant information is avoided. In the scenario of fuzzy input, based on the business logic, the associated analysis dimensions are automatically recommended, realizing the intelligent guidance from fuzzy requirements to accurate analysis.
[0032] Based on the user's historical interaction information, perform user - dimension matching on the multiple target analysis paths to determine the recommendation index corresponding to each target analysis path, specifically including: counting the historical query dimensions corresponding to the user's historical interaction information, constructing the user portrait feature vector of the user based on the historical query dimensions, and determining the corresponding interaction preference type of the user, where the interaction preference type includes exploratory users and conservative users; determining the embedding vector of the knowledge node corresponding to the analysis dimension and the user portrait feature vector according to the analysis dimension corresponding to each target analysis path, performing personalized matching on the target analysis path to determine the personalized matching index corresponding to each target analysis path; counting the number of times the analysis dimension has been recommended within a preset historical time period in the user's historical interaction information, calculating the exposure frequency of the analysis dimension relative to the user through the number of times recommended to determine the novelty matching index corresponding to each target analysis path; counting the cumulative number of queries of the analysis dimension within the preset historical time period to calculate the heat index corresponding to the analysis dimension through the cumulative number of queries; matching the corresponding weight parameter combination through the interaction preference type corresponding to the user, and based on the weight parameter combination, weighting the personalized matching index, the novelty matching index, and the heat index to determine the recommendation index corresponding to each target analysis path.
[0033] In an embodiment of the present specification, first retrieve the user's historical interaction records, and count the analysis dimensions and operation characteristics selected within a preset period (such as 30 days), including but not limited to query dimension types, parameter adjustment frequencies, result export times, etc. Based on the distribution law of historical query dimensions, use feature engineering methods to construct the user portrait feature vector, which comprehensively reflects the user's focus of attention and operation habits in the business field. For example, a user who frequently uses "regional sales comparison" and often adjusts the time range will have significant weights in the spatial dimension and time dimension of the feature vector. Further, through a clustering algorithm, users are divided into two categories: "exploratory" or "conservative". Based on the user's historical interaction behavior characteristics, including the frequency of trying new dimensions, the number of times of modifying query conditions, the diversity of result exports, etc., construct a multi - dimensional feature matrix, and use a clustering algorithm to divide users into exploratory and conservative categories. Among them, exploratory users show behavior patterns such as frequently trying new analysis dimensions and actively adjusting complex query parameters, while conservative users exhibit operation characteristics such as repeatedly using fixed analysis templates and rarely modifying default settings.
[0034] For each candidate analysis path, a triple-index evaluation is conducted. The feature vector of the user profile is compared with the embedding vector of the knowledge node corresponding to the target analysis dimension in terms of similarity. The similarity comparison method here can be cosine similarity. The knowledge node embedding vector is generated by a pre-trained graph neural network, representing the semantic position of this dimension in the business graph. The similarity calculation reflects the degree of fit between the user's historical preferences and the current analysis dimension, thereby obtaining a personalized matching index. The exposure times of this analysis dimension to the current user in the recent period (such as 7 days) are counted, and the novelty value is calculated using an inverse function. The novelty value , N shown is the number of times this dimension has been recommended within the historical time period. Based on the cumulative usage times of this analysis dimension by all platform users during the same period, its popularity is quantified through normalization. Among them, high-heat-value dimensions reflect the common needs of business scenarios. The index weights are dynamically adjusted according to the user's preferred type. Different weight templates are preset for different preferred types. Exploratory users adopt a "novelty-dominated" weight combination, such as a novelty weight of 0.5, a personalized matching weight of 0.3, and a popularity weight of 0.2. For exploratory users, a higher weight is given to the novelty index to stimulate the exploration of unknown needs; conservative users adopt a "popularity-first" weight combination, such as a popularity weight of 0.6, a personalized matching weight of 0.3, and a novelty weight of 0.1. For conservative users, the popularity index is emphasized, and mature analysis methods are recommended first. At the same time, a dynamic adjustment mechanism is established to continuously monitor the signs of behavior deviation in the user's recent interaction data. For example, if a conservative user continuously selects a new dimension 3 times, it will trigger a type reclassification, and the activation intensity of the weight template will be automatically adjusted. And a feedback learning loop is set up. By comparing the recommendation effects of different weight combinations in the same user group through A / B testing, the optimization plan with a click-through rate increase of more than 15% is selected and updated to the weight strategy library to form a continuously evolving weight configuration system. The three indexes are weighted and summed according to the preset weight combination to generate a final recommendation index, and the candidate paths are sorted based on this. Through the deep integration of user behavior characteristics and the business knowledge graph, an accurate mapping from historical behavior to personalized recommendation is achieved.
[0035] Before generating the personalized analysis card sequence according to the recommendation index of each such target analysis path and the pre-obtained recommended quantity corresponding to this user, the method further includes: obtaining the historical recommendation data corresponding to the guided query channel; statistically analyzing the historical recommendation data to determine the number of historical recommendation dimensions corresponding to the guided query channel during each interaction process, and determining the average recommended quantity through the number of historical recommendation dimensions and the number of historical interactions; obtaining the historical recommendation dimension information and historical selected dimension information in the user's historical interaction information, and generating a personalized selection recommendation ratio according to the ratio of the selection dimension sequence identifier in the historical selected dimension information and the historical recommended quantity in the historical recommendation dimension information; correcting the average recommended quantity through the personalized selection recommendation ratio to determine the recommended quantity corresponding to this user.
[0036] In one embodiment of the present specification, historical recommendation data of the guided query channel is obtained, and the number distribution characteristics of the recommendation dimensions in each interaction process are determined through statistical analysis. The arithmetic mean of the historical recommended dimension numbers is calculated as the basic recommendation quantity. Further, the recommendation and selection logs in the user's historical interaction records are extracted, and the proportional relationship between the sequential position of the actually selected dimension in the past recommended list and the total number of recommendations is statistically analyzed to generate a personalized selection recommendation ratio index. For example, if a user selects the third position when the recommended list has 6 items in a certain interaction, the corresponding selection recommendation ratio index for that interaction is 3 / 6 = 0.5. In another historical interaction process, when the recommended list has 8 items and the user selects the fifth position, the corresponding selection recommendation ratio index for that interaction is 5 / 8 = 0.625. The selection recommendation ratio corresponding to each interaction process is statistically analyzed, and the maximum selection recommendation ratio index in the historical interaction process is calculated. After multiplying this ratio index by the basic recommendation quantity and rounding up, the final recommended quantity is determined. For example, if the maximum selection recommendation ratio index for User A is 0.5 and that for User B is 0.625, and the basic recommendation quantity is 8, then for User A, 4 items need to be recommended to meet the maximum ratio index in the historical interaction process, while for User B, 5 items need to be recommended.
[0037] Through the above technical solution, by analyzing the historical recommendation data of the guided query channel, it is possible to understand the appropriate number of recommended dimensions in each interaction under normal circumstances. On this basis, by combining the selected dimension information in the user's own historical interaction data, the personalized selection recommendation ratio is calculated, taking into account the acceptance and selection preferences of different users for the recommended content; the average recommended quantity is corrected according to the user's personalized selection recommendation ratio, avoiding the one-size-fits-all approach of a unified fixed recommended quantity. The final recommended quantity is determined by multiplying the maximum selection recommendation ratio index of different users by the basic recommendation quantity and rounding up, which can provide a recommended quantity that better meets the individual needs of each user; making the recommended quantity more personalized and dynamic, taking into account both the overall historical recommendation situation and the unique interaction behaviors of each user.
[0038] Step S103, under the trigger of the target analysis dimension, the guided query channel is switched to the precise query channel. Through the precise query channel, the pre-learned query mode knowledge base is called to generate a target query statement corresponding to the target analysis dimension, so as to determine the corresponding visual analysis result through the target query statement.
[0039] In one embodiment of the present specification, after generating a personalized analysis card sequence and presenting it to the user, corresponding parameter adjustment controls are provided in each card, such as a price range slider, a geographical level selector, etc. The target analysis dimension in the personalized analysis card sequence triggered by the user is obtained, where the user trigger operation refers to the target analysis dimension selected by the user in the personalized analysis card sequence.
[0040] In one embodiment of the present specification, upon triggering of the target analysis dimension, through a dynamic routing mechanism, the guiding query channel is switched to a precise query channel. Through the precise query channel, a pre-learned query pattern knowledge base is called to generate a target query statement corresponding to the target analysis dimension, so as to determine the corresponding visual analysis result through the target query statement. It should be noted that the switching between the two channels is controlled by a dynamic routing mechanism, which makes decisions based on a confidence threshold. For example, in a financial risk assessment scenario, when the user enters a vague query intention, the system will automatically switch to the guiding channel and use visual cards to help the user gradually clarify the analysis dimension. After the user selects a specific analysis direction, it will then switch to the precise channel to execute a detailed data query.
[0041] Through this precise query channel, a pre-learned query pattern knowledge base is called to generate a target query statement corresponding to the target analysis dimension, which specifically includes: performing intention recognition on the target analysis dimension, extracting business entities and intention recognition results in the target analysis dimension, so as to map the business entities in the pre-learned and constructed business field semantic knowledge base to determine entity link results; through the query pattern knowledge base, determining a query statement template according to the intention recognition result and the entity link result; and performing syntax optimization on the query statement template to determine the target query statement corresponding to the target analysis dimension.
[0042] In one embodiment of this specification, when the target query channel is the precise query channel or is switched to the precise query channel, the target query statement is generated as follows: Based on the deep integration of pre-trained large language models and pre-learned knowledge from the dataset in the precise query channel, end-to-end conversion from natural language to SQL is achieved, significantly enhancing the system's semantic understanding depth and SQL generation accuracy. The introduction of a context-aware query understanding mechanism and dataset-specific knowledge injection solves the bottleneck problems of traditional methods in scenarios such as handling complex business logic, multi-dimensional analysis, and non-standard expressions. When the large model generates SQL, it automatically retrieves and integrates multi-level knowledge bases obtained during the pre-learning stage, mainly including: data schema knowledge, that is, the business knowledge graph of table-level data structures, such as table structures, field types, constraint conditions, etc.; business field semantic knowledge bases, which can also be called entity-relationship knowledge, such as inter-table associations, field mappings, and business concept hierarchies; and query pattern knowledge bases, containing frequent query patterns, common aggregation methods, and query templates for typical business scenarios. These structured knowledges are embedded into the model's prompt input sequence through special encoding methods, enabling the large model to accurately grasp the data characteristics and query constraints of a specific business domain while maintaining its powerful semantic understanding ability.
[0043] First is the entity linking and intent recognition stage. In the precise query channel, it is stated that the user's query requirements are clear. The model precisely maps the business entities in the user's query to the tables and fields in the data schema, and at the same time extracts the core analysis intents, such as comparison, aggregation, sorting, grouping, etc. According to the recognized intents and entity linking results, combined with the pre-learned query pattern knowledge, a preliminary SQL logical skeleton is generated, clarifying the JOIN relationship, WHERE conditions, and GROUP BY clauses. Finally is the SQL refinement and optimization stage. By iteratively refining the initial SQL, inserting appropriate subqueries, window functions, and advanced SQL structures, and at the same time tuning the performance of the generated SQL in combination with the feedback from the database optimizer, ensuring the efficient execution of the query.
[0044] In addition, in addition to the above methods, the target query statement can also be determined by constructing a question and answer knowledge base. Aiming at the problems of automatic generation and pattern abstraction of question-answer pairs in intelligent question-answering systems, through the organic combination of automatic annotation and manual review, the large-scale generation and standardized management of high-quality question-answer pairs are realized.
[0045] In the automatic annotation stage, based on a pre-established multi-layer pre-learning architecture, data features are automatically recognized and candidate questions that conform to business logic are generated. First, an automatic trigger task is created based on a large model. When a dataset is imported into the system, the system extracts key information from data fields and business rules, converts this information into natural language questions, and generates corresponding SQL query statements or structured answers. To enhance the diversity and authenticity of the generated questions, an innovative adversarial training mechanism is adopted. In this mechanism, the generator is responsible for generating question variants with different expressions, while the discriminator is responsible for evaluating the rationality and usability of these questions. Through the repeated adversarial training of the generator and the discriminator, the quality and diversity of question generation can be continuously optimized. In the manual review stage, a visual annotation platform is designed to support business experts in systematically reviewing and optimizing the automatically generated questions. The platform provides an intuitive interface that enables experts to conveniently correct questions, merge similar questions, add business tags (such as "financial analysis type", "operation monitoring type", etc.), and adjust answer templates. When the guided query channel is switched to the precise query channel, the target analysis dimension is converted into the corresponding question description, and the corresponding query statement is matched in the Q&A knowledge base through the converted question description. The visualization display of the analysis results is achieved through the query statement.
[0046] The embodiment of this specification uses the retail data analysis scenario as an example for explanation. When the user enters the fuzzy query "analyze the characteristics of best-selling products", the intention fuzziness detection model is first started for evaluation. By processing the word vector features and syntactic structure of the input text, it is recognized that the query lacks a specific analysis dimension definition, and the calculated confidence level is 0.35, which is higher than the precise query threshold of 0.3. This indicates that the user's analysis intention is relatively vague and requires active guidance by the system. Subsequently, the active guidance channel is activated, and the pre-built retail commodity knowledge graph is called. The knowledge graph contains basic commodity attributes (such as categories, brands, specifications, etc.), sales attributes (such as price, sales volume, gross profit, etc.), and associated attributes (such as promotions, seasonality, complementary commodities, etc.). Based on the current analysis topic "Best-selling product characteristics", the path is derived in the knowledge graph through a bidirectional breadth-first search algorithm. At the same time, combined with the user's historical analysis preferences, three most valuable analysis paths are generated: the price band distribution analysis dimension, which reveals the sales performance of products at different prices through price range division, which helps to optimize pricing strategies; the regional sales comparison dimension, which analyzes the distribution characteristics of best-selling products from a geographical perspective and supports regional operation decisions; the repurchase rate association analysis dimension, which combines transaction frequency and shopping cart analysis to mine association rules between products. These analysis paths are presented to users in the form of visual cards, each of which is equipped with corresponding parameter adjustment controls, such as price range sliders, geographic level selectors, etc. When the user selects the "regional sales comparison" analysis dimension, the system immediately switches to the precise query channel, retrieves the SQL template with the highest matching degree from the query mode knowledge base, and automatically generates queries containing multi-table associations, time series comparisons, and composite indicator calculations. The execution results generate a dashboard of associated charts through the visualization engine, and automatically generate core analysis information using a large model.
[0047] Through the technical solutions in the embodiments of this specification, users do not need to master professional query languages such as SQL. They only need to input natural language query texts, which reduces the threshold for non-technical personnel to use business intelligence tools. By performing intent detection and fuzzy confidence determination on natural language query texts, the target query channel can be automatically matched. In the guided query channel, the business knowledge graph is used for path derivation to generate a personalized analysis card sequence. Users do not need to confirm parameters such as analysis dimensions and filtering conditions step by step, avoiding disassembling natural semantics into discrete input sequences, thus maintaining the fluency of the interaction and improving the human-computer collaboration efficiency. After the user triggers the target analysis dimension in the personalized analysis card sequence, it switches to the precise query channel, calls the query mode knowledge base to generate the target query statement, and determines the visual analysis result. From the guided to the precise query method, it not only enables users to get guidance and inspiration in the stage of fuzzy requirements but also quickly obtains precise visual analysis results after the requirements are clarified. The entire interaction process is natural and fluent, conforming to the user's thinking habits. Based on the dynamic path exploration of the business knowledge graph, it can actively recommend associated analysis dimensions for users without relying on manual experience for gradual exploration, greatly improving the analysis efficiency. When the user determines the target analysis dimension, the precise query channel uses the pre-learned query mode knowledge base to generate the target query statement and quickly obtains the visual analysis result, reducing the user's waiting time and further improving the analysis efficiency.
[0048] The embodiments of this specification also provide an interactive intelligent analysis device based on a knowledge graph, as Figure 2 shown. The device includes: at least one processor; and a memory communicatively connected to the at least one processor. Wherein, the memory stores instructions executable by the at least one processor. The instructions are executed by the at least one processor to enable the at least one processor to execute the above method.
[0049] The embodiments of this specification also provide a non-volatile computer storage medium storing computer-executable instructions, and the computer-executable instructions are set to: execute the above method.
[0050] The embodiments in this specification are all described in a progressive manner. For the same or similar parts among the embodiments, reference can be made to each other. Each embodiment focuses on the differences from other embodiments. In particular, for the embodiments of the device, equipment, and non-volatile computer storage medium, since they are basically similar to the method embodiments, the description is relatively simple, and the relevant parts can refer to the partial description of the method embodiments.
[0051] The above description is only one or more embodiments of this specification and is not intended to limit this specification. For those skilled in the art, various changes and modifications can be made to one or more embodiments of this specification. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of one or more embodiments of this specification shall be included within the scope of the claims of this specification.
Claims
1. An interactive intelligent analysis method based on knowledge graph, characterized in that: The method comprises: Acquire a natural language query text input by a user, perform intent detection on the natural language query text, determine a corresponding intent fuzzy confidence, and match a target query channel according to the intent fuzzy confidence; When the target query channel is a guided query channel, the path of the natural language query text is derived through the business knowledge graph constructed by pre-learning, and a personalized analysis card sequence is generated to obtain the target analysis dimension in the personalized analysis card sequence triggered by the user; Under the triggering of the target analysis dimension, the guided query channel is switched to a precise query channel, and through the precise query channel, a pre-learned query pattern knowledge base is called to generate a target query statement corresponding to the target analysis dimension, so as to determine the corresponding visual analysis result through the target query statement; Through the business knowledge graph constructed by pre-learning, the path of the natural language query text is derived to generate a personalized analysis card sequence, which specifically includes: Performing entity extraction on the natural language query text to determine at least one corresponding business entity; Taking each of the business entities as the center, dynamic path exploration is performed in the business knowledge graph to determine a real-time exploration path to obtain path parameters corresponding to the real-time exploration path, wherein the path parameters include the total number of connection edges of the path node in the business knowledge graph and the edge weight parameter corresponding to the exploration path; Dynamically calculating a dynamic path score corresponding to the real-time exploration path through the total number of connected edges and the edge weight parameter, so as to perform path screening based on the dynamic path score and determine multiple target analysis paths; Acquire user history interaction information of the user, perform user dimension matching on the multiple target analysis paths based on the user history interaction information, and determine a recommendation index corresponding to each of the target analysis paths; The personalized analysis card sequence is generated according to the recommendation index of each target analysis path and the number of recommendations corresponding to the user obtained in advance, wherein the personalized analysis card sequence includes analysis dimensions corresponding to multiple designated target analysis paths that meet preset requirements.
2. According to claim 1, an interactive intelligent analysis method based on knowledge graph is characterized in that: Performing intent detection on the natural language query text to determine the corresponding intent fuzzy confidence level specifically includes: Performing word segmentation processing on the natural language query text to determine word vector features, and performing syntactic analysis on the natural language query text to extract syntactic structure features; Determining a word vector sequence corresponding to the word vector feature to capture a bidirectional context dependency; The bidirectional context dependency and the syntactic structure features are fused by using a multi-head attention mechanism and input into a pre-trained BERT model to obtain a CLS tag vector, so as to determine the intent fuzzy confidence corresponding to the natural language query text through the CLS tag vector.
3. The interactive intelligent analysis method based on knowledge graph according to claim 1 is characterized in that: Before the business knowledge graph constructed by pre-learning is used to deduce the path of the natural language query text and generate a personalized analysis card sequence, the method further includes: Acquire historical query information, and automatically extract a full business data table through a preset database metadata interface, wherein the full business data table includes table information, field information, and data type; Through a pre-set pre-learning framework, a multi-level knowledge base is constructed based on the full business data table and the historical query information, wherein the multi-level knowledge base includes a business knowledge graph, a business field semantic knowledge base and a query pattern knowledge base.
4. The interactive intelligent analysis method based on knowledge graph according to claim 1, characterized in that: Based on the user historical interaction information, user dimension matching is performed on the multiple target analysis paths to determine the recommendation index corresponding to each target analysis path, specifically including: Counting the historical query dimensions corresponding to the historical interaction information of the user, so as to construct a user portrait feature vector of the user based on the historical query dimensions, and determine the interaction preference type corresponding to the user, wherein the interaction preference type includes an exploratory user and a conservative user; According to the analysis dimension corresponding to each of the target analysis paths, the embedding vector of the knowledge node corresponding to the analysis dimension and the user portrait feature vector are determined, personalized matching is performed on the target analysis path, and the personalized matching index corresponding to each of the target analysis paths is determined; Counting the number of times the analysis dimension in the user's historical interaction information is recommended within a preset historical time period, and calculating the exposure frequency of the analysis dimension relative to the user based on the number of times recommended, so as to determine a novelty matching index corresponding to each of the target analysis paths; Counting the cumulative number of queries for the analysis dimension within the preset historical time period, so as to calculate the heat index corresponding to the analysis dimension according to the cumulative number of queries; The corresponding weight parameter combination is matched by the interaction preference type corresponding to the user, so as to weight the personalized matching index, the novelty matching index and the heat index based on the weight parameter combination to determine the recommendation index corresponding to each of the target analysis paths.
5. The interactive intelligent analysis method based on knowledge graph according to claim 1, characterized in that: Before generating the personalized analysis card sequence according to the recommendation index of each target analysis path and the number of recommendations corresponding to the user obtained in advance, the method further includes: Obtain historical recommendation data corresponding to the guided query channel; Performing statistical analysis on the historical recommendation data to determine the number of historical recommendation dimensions corresponding to the guided query channel in each interaction process, and determining the mean value of the recommendation quantity through the number of historical recommendation dimensions and the number of historical interactions; Obtaining historical recommendation dimension information and historical selection dimension information in the user's historical interaction information, and generating a personalized selection recommendation ratio according to a ratio of a selection dimension sequence identifier in the historical selection dimension information and a historical recommendation quantity in the historical recommendation dimension information; The average of the recommended quantity is corrected by the personalized selection recommendation ratio to determine the recommended quantity corresponding to the user.
6. The interactive intelligent analysis method based on knowledge graph according to claim 1, characterized in that: Through the precise query channel, the pre-learned query pattern knowledge base is called to generate a target query statement corresponding to the target analysis dimension, specifically including: Performing intent recognition on the target analysis dimension, extracting business entities and intent recognition results in the target analysis dimension, mapping the business entities in a pre-learned and constructed business field semantic knowledge base, and determining entity linking results; Determine a query statement template through the query pattern knowledge base according to the intent recognition result and the entity linking result; The query statement template is grammatically optimized to determine a target query statement corresponding to the target analysis dimension.
7. The interactive intelligent analysis method based on knowledge graph according to claim 3 is characterized in that: Through the pre-set pre-learning framework, a multi-level knowledge base is constructed according to the full business data table and the historical query information, specifically including: According to the full business data table, construct initial pattern graph data, and according to the historical query frequency in the historical query information, perform edge weight allocation on the initial pattern graph to determine the business pattern graph data; Determine a business knowledge graph of a table-level data structure based on a pre-trained deep graph learning model and the business model graph data; Through the field-level semantic modeling layer, the multimodal semantic features of the fields in the full business data table are extracted, and based on the multimodal semantic features, a multidimensional similarity matrix is constructed, wherein the multidimensional similarity matrix includes field name similarity, data distribution similarity and content overlap; According to the field feature descriptors and the multi-dimensional similarity matrix obtained in advance in the full business data table, the field semantic vector is determined by using an attention-enhanced autoencoder to construct a business field semantic knowledge base; In the historical query information, high-frequency query pattern data is obtained to convert query statements in the high-frequency query pattern data to determine a statement syntax tree; By using a preset semantic encoder, the sentence syntax tree is semantically converted to determine the query semantic information of the query statement, so as to construct a query pattern knowledge base according to the query semantic information and the field semantic vector, wherein the query pattern knowledge base includes a plurality of query statement templates and the query semantic information corresponding to each of the query statement templates; A multi-level knowledge base is determined according to the business knowledge graph of the table-level data structure, the business field semantic knowledge base and the query pattern knowledge base.
8. An interactive intelligent analysis device based on knowledge graph, characterized in that: The device comprises: at least one processor; and, a memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can perform the method according to any one of claims 1 to 7.
9. A non-volatile computer storage medium storing computer executable instructions, characterized in that: The computer executable instructions are configured to execute the method according to any one of claims 1 to 7.
Citation Information
Patent Citations
Systems and methods for adaptive question answering related applications
EP3855320A1
Method and apparatus for querying writing material, and storage medium
US20220335070A1