Enterprise intelligent consultation management method and system based on big data analysis
By improving the text concept semantic fusion algorithm and the positional focus-bidirectional collaborative entity information extraction algorithm, the accuracy and efficiency problems of text understanding and key information extraction in the enterprise intelligent consulting management system have been solved, and efficient consulting suggestion generation and decision support have been achieved.
Patent Information
- Application Number
- CN202510867229.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-26
- Publication Date
- 2025-12-09
- Estimated Expiration
- 2045-06-26
AI Technical Summary
Existing enterprise intelligent consulting management systems suffer from insufficient accuracy and efficiency in text understanding and key information extraction, making it difficult to effectively utilize big data analysis to generate efficient consulting recommendations.
An improved text concept semantic fusion algorithm and positional focus-bidirectional collaborative entity information extraction algorithm are adopted. By generating text representation vectors through Transformer embedding and multiple attention heads, combined with CRF decoding and attention focus, the system can accurately understand user inquiries and extract key information.
It improves the accuracy of text understanding and the ability to extract key information in the enterprise intelligent consulting management system, generates structured and actionable consulting suggestions, and provides more comprehensive and accurate decision support.
Smart Images

Figure CN120763292B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The application relates to the field of natural language processing and deep learning, in particular to an enterprise intelligent consultation management method and system based on big data analysis. BACKGROUND
[0002] Natural language processing technology is an important branch of artificial intelligence, aiming to study and implement effective interaction between computers and human natural language, and its core goal is to enable machines to understand, generate and manipulate human language, thereby bridging the gap between human communication methods and computer formalized language, which involves multi-dimensional analysis and processing from the basic units of language (such as vocabulary, grammar) to complex semantics, context, emotion and pragmatic level, with the rise of machine learning, especially statistical methods, NLP turns to data-driven paradigm, using large-scale corpus to learn the statistical rules of language, for example, hidden Markov model (HMM), conditional random field (CRF) in sequence labeling tasks (such as part-of-speech tagging, named entity recognition) have achieved remarkable results, in the past decade, deep learning has completely reshaped the NLP field, distributed representation technology represented by word embedding (such as Word2Vec, GloVe) enables semantic information of words to be effectively captured by dense vectors, opening up the era of wide application of neural network models. In particular, the emergence of the Transformer architecture and its derivative pre-training language models (such as BERT, GPT, T5), the mutual cooperation of these technologies enables the enterprise intelligent consultation management system to have more accurate and efficient performance.
[0003] Deep learning technology is to automatically learn high-level features and abstract representations of data from a large amount of data by constructing a neural network model with multiple hidden layers, so as to realize the recognition and prediction of complex patterns, deep learning technology includes long short-term memory network (LSTM), gated recurrent unit network, multi-LSTM network of physical information and physically guided convolutional neural network, the mutual cooperation of these technologies jointly constructs the enterprise intelligent consultation management system, so that the enterprise intelligent consultation management system has more professional and accurate performance. SUMMARY
[0004] In view of the above problems, the application aims to provide an enterprise intelligent consultation management method and system based on big data analysis.
[0005] In order to achieve the above object, the present application provides the following technical scheme: a kind of enterprise intelligent consultation management method and system based on big data analysis, including data acquisition module, data preprocessing module, consultation question input and analysis module, big data intelligent analysis engine module, intelligent consultation suggestion generation module, visual module and feedback learning and system optimization module, data acquisition module is used to collect enterprise internal data and external network data, data preprocessing module is used to clean, convert, integrate and store the data collected, consultation question input and analysis module proposes to improve text concept semantic fusion algorithm to receive the consultation question proposed by user and carries out text understanding and structuring, big data intelligent analysis engine module includes task scheduling and management unit and text analysis engine unit, task scheduling and management unit is used to receive structured consultation task to coordinate the required data resources, model resources and computing resources, text analysis engine unit proposes to improve bit sequence focus-bi-directional collaborative entity information extraction algorithm to carry out opinion mining and key information extraction, intelligent consultation suggestion generation module is used to generate structured, executable consultation suggestion, visual module is used to visually show analysis result and consultation suggestion for user, feedback learning and system optimization module is used to optimize the performance of system model according to user feedback.
[0006] Further, data acquisition module accesses enterprise internal system data by API, database connection and log file reading mode, and obtains external network data by network crawler, API interface, data subscription and the like.
[0007] Further, data preprocessing module cleans, converts and integrates the collected data, extracts key information from unstructured text data, and stores it in structured / semi-structured form, to prepare high-quality data basis for text understanding.
[0008] Further, the consultation question input and analysis module proposes to improve text concept semantic fusion algorithm to receive the consultation question proposed by user and carries out text understanding and structuring.
[0009] Further, the improved text concept semantic fusion algorithm is as follows: first, receive the text consultation question input by user, then, input the consultation question for text conceptualization, the representation of input text in concept is as follows: Wherein, c is the weight vector represented by input text in conceptual space, t is term, P (c|t) is the conditional probability that the term belongs to concept c under the condition that the term t appears, c i To traverse all concepts, C is the concept set, the data source of C comes from structured knowledge base and corpus, count (t, c) is the co-occurrence number of term t and concept c in knowledge base, The total co-occurrence number of the standardized term t and all concepts is standardized, that is The statistical time window of the co-occurrence number is a sliding window, and the conceptualized text is converted from the word vector space to the concept space, that is wherein, is the corresponding weight of the text under the concept c1, is the corresponding weight of the text under the concept c2, is the corresponding weight of the text under the concept c k In order to make the term t and the concept c have better distinguishing ability, the intra-concept inverse term frequency and the inverse concept frequency are proposed to improve the relevance between the concept and the text, so that the association strength between the concept and the text is wherein, is the corresponding weight of the text under the concept c i is the weight of the current text, ST is the current text, ST i is the ith short text, t i is the ith term in the text, T i is the term t i belongs to the concept c i , that is, T i = P(c i |t i ), is the corresponding weight of the text under the concept c i in the knowledge base, idf c (t i ) is the intra-concept inverse term frequency, icf(c i ) is the inverse concept frequency, then, a score function is defined to enrich the semantic knowledge in the text, and a dynamic weight and a domain adaptive term are proposed to improve the score function to improve the accuracy of short text understanding, that is Scoure(x|y,s) = α(s)P co-occun (x|y) + (1-α(s))P smmetaic ((x|y,s) + β·Domain(x|y), wherein, Scoure(x|y,s) is the improved score function, x and y are terms, s is a short text, α(s) is a dynamic weight coefficient, P ro-occur (x|y) is the global co-occurrence probability of the terms x and y, P semantic( (x|y,s) is the semantic association probability of the terms x and y in the text s, β is a domain knowledge weight coefficient, β is initialized as 0.5, and is adjusted according to user feedback as β←β±0.1·sign(accuracy delta ), wherein, sign(·) is a sign function, accuracy deltaThe accuracy rate change value fed back to the user, Domain(x|y) is the association strength of terms x and y in the domain knowledge base, i.e. Where, P Eit (x,y) is the co-occurrence probability of terms x and y in the enterprise knowledge base, P Ent (x) is the independent occurrence probability of term x in the enterprise knowledge base, P Ent (y) is the independent occurrence probability of term y in the enterprise knowledge base.
[0010] Then, the conceptualized information is embedded by Transformer, with embedding dimension d model to further capture the context semantic information. Assuming that the length of the text consultation question input by the user is n, the embedding matrix is M, and the correlation calculation between the query vector query and the key vector key in the self-attention mechanism of Transformer is Similarity(query,key)=Q·K, where Similarity(query,key) is a similarity function, Q is a numerical vector of query, K is a numerical vector of key, and Q·K is a dot product operation. In order to realize attention weight normalization, a softmax function is used to determine the correlation calculation result, i.e. Where, r is the weight coefficient of the correlation calculation result, and the attention calculation is Attention(Q,K,V)=∑r·V, where Attention(Q,K,V) is an attention output vector, V is a value vector, and a multi-head attention mechanism is used for multi-subspace parallel calculation, i.e. MultiHeadAttention(ST i )=Concat(head1,…,head k )W i , where MultiHeadAttention(ST i ) is the final output result after multi-head attention calculation on the i-th text ST i , Concat(·) is a concatenation operation, head1 is the first attention head result, head k is the k-th attention head result, and W i is a projection matrix that projects the i-th text ST i of the embedding matrix M to the query, key and value vectors. The head calculation is Where, head i is the i-th attention head result, is the projection matrix that projects the i-th text ST i to the query vector, is the projection matrix that projects the i-th text ST i to the key vector, To obtain the i-th text ST i The projection matrix of the projection to the value vector, thus obtaining the text representation vector rich in context information T j = Norm(Norm(T j-1 + P j )+ Transition(Norm(T j-1 + P j ))) where T j is the k-th layer text representation vector, T j-1 is the j-1-th layer text representation vector, Norm(·) is the normalization operation, P j is the two-dimensional coordinate embedding, and Transition(·) is the depth separable convolution. The improved text concept semantic fusion algorithm first proposes the intra-concept inverse term frequency and inverse concept frequency to improve the relevance between concepts and texts, then proposes dynamic weights and domain adaptive terms to improve the score function to improve the accuracy of short text understanding, and then embeds the conceptualized information into the Transformer to further capture the context semantic information, and uses multiple attention heads to generate the text representation vector, so as to accurately understand the user input text consultation question.
[0011] Further, the task scheduling and management unit judges the problem type through the concept weight distribution and semantic density information in the vector, dynamically calculates the priority by fusing user metadata, business rules and vector features, and then dynamically allocates the execution path according to the problem type, priority and resource occupation estimation.
[0012] Further, the text analysis engine unit proposes an improved bit order focusing-bidirectional collaborative entity information extraction algorithm for opinion mining and key information extraction.
[0013] Further, the improved bit order focusing-bidirectional collaborative entity information extraction algorithm is as follows: first, the consultation question input and the text representation vector processed by the analysis module are taken as input, that is, where T is a text representation vector sequence, is the first text representation vector, is the second text representation vector, is the n-th text representation vector, n is the length of the text representation vector sequence, and is the positioning word boundary set for entity extraction. When the position encoding is greater than 0.7, the boundary detection is activated. An absolute position encoding is proposed to add position information to the text representation vector. i = PositionEmbedding(pos i ), i is the position index of the word in the sequence, and p iis the position embedding vector, PositionEmbedding(·) is the position embedding function, and pos i is the absolute position index of the current word in the text, i.e., pos i is the position embedding vector, and i is the fusion of semantic and position information, i.e., wherein, is the enhanced semantic vector after splicing, and concat(·) is the splicing operation, is the i-th text representation vector, and p i is the position embedding vector, in order to better capture the joint influence of the entity boundary of the front and rear words, a bidirectional long short-term memory network is introduced to capture the context dependence, and the present definition time step is t, and t = i, the forward LSTM calculation is wherein, is the output hidden state of the forward LSTM at time step t, LSTM(·) is the long short-term memory unit, is the input vector of time step t, is the output hidden state of the forward LSTM at time step t-1, and the backward LSTM calculation is wherein, is the output hidden state of the backward LSTM at time step t, is the output hidden state of the backward LSTM at time step t+1, and the output of the bidirectional LSTM is
[0014] The CRF (conditional random field) label transition score coefficient λ (0≤λ≤1) is introduced to fuse with the features of the LSTM output, i.e., m t = λ·h t + (1-λ)s t wherein, m t is the mixed state vector, h t is the output vector of the bidirectional LSTM, s t is the hidden state vector of the CRF at time step t, and then the CRF (conditional random field) decoding is performed to meet the label sequence constraint, i.e. wherein, score(T,y) is the CRF sequence scoring function, y is the candidate label sequence, is the label confidence of being labeled as the candidate label y i at the position index i, is the label transition score from the candidate label y i to y i+1 , and n is the length of the text representation vector sequence, and then the global optimal label after decoding is calculated as wherein, y *For the optimal label sequence, argmax(·) is to solve the label sequence y when the value of score(T,y) is maximum, and then when performing relation extraction, attention focusing is proposed to avoid the dilution of key words by redundant words, that is Wherein, alpha i is a normalized attention weight, u is an attention vector, u T is the transpose of u, tanh(·) is a hyperbolic tangent activation function, W is a weight matrix of attention focusing, W is initialized by Xavier initialization, and diagonal elements are forced to be 1.0, is a relation feature vector of the current focus position index i, is a relation feature vector of the traversal comparison position index j, and the weighted aggregation feature vector is Then, the semantic core refined based on attention is used for relation determination, that is Wherein, is a relation probability distribution vector, W rel is a relation classification weight matrix, b rel is a relation classification bias vector, and finally a relation triple is generated. The improved position sequence focusing-bidirectional collaborative entity information refining algorithm first proposes absolute position encoding to add position information of a text representation vector, then proposes a bidirectional long short-term memory network to capture the joint influence of front and rear words on entity boundaries, and proposes CRF decoding to meet the label sequence constraint, then proposes attention focusing to avoid the dilution of key words by redundant words, and finally the semantic core refined based on attention is used for relation determination to generate a relation triple, so as to realize the key information refining of the text.
[0015] Further, the intelligent consultation suggestion generation module reasons according to the relation triple generated by the text engine unit to generate structured and executable consultation suggestions, and the visualization module is used to visually display the analysis results and consultation suggestions generated by the intelligent consultation suggestion generation module for the user.
[0016] Further, the feedback learning and system optimization module is used to optimize the performance of the system model according to the user feedback, the user feedback is fed back to the system according to the obtained consultation suggestion, and the system re-answers the consultation question according to the user's feedback opinion, so as to optimize the performance of the system model.
[0017] Compared with the prior art, the beneficial effects of the present application are:
[0018] 1. The application proposes an improved text concept semantic fusion algorithm to receive user's consultation questions and perform text understanding and structuring. The innovation of the application is that the improved text concept semantic fusion algorithm first proposes intra-concept inverse term frequency and inverse concept frequency to improve the relevance between concepts and texts, then proposes dynamic weight and domain adaptive term to improve the score function to improve the accuracy of short text understanding, and then embeds the conceptualized information into Transformer to further capture the context semantic information, and uses multiple attention heads to generate text representation vectors, so as to accurately understand the user's input text consultation questions.
[0019] 2. The application proposes an improved bit sequence focusing-bidirectional collaborative entity information extraction algorithm to perform opinion mining and key information extraction. The innovation of the application is that the improved bit sequence focusing-bidirectional collaborative entity information extraction algorithm first proposes absolute position encoding to add position information of the text representation vector, then proposes a bidirectional long short-term memory network to capture the joint influence of the front and rear words on the entity boundary, and proposes CRF decoding to meet the label sequence constraint, then proposes attention focusing to avoid dilution of key words by redundant words, and finally performs relationship determination based on the semantic core extracted by attention to generate relationship triples, so as to realize key information extraction of the text. BRIEF DESCRIPTION OF DRAWINGS
[0020] The application is further described with reference to the accompanying drawings, but the embodiments in the drawings do not constitute any limitation on the application. For ordinary skilled in the art, other drawings can be obtained without creative labor based on the following drawings.
[0021] Figure 1 The application is further described with reference to the accompanying drawings, but the embodiments in the drawings do not constitute any limitation on the application. For ordinary skilled in the art, other drawings can be obtained without creative labor based on the following drawings. DETAILED DESCRIPTION
[0022] The technical solutions in the embodiments of the application will be described clearly and completely with reference to the accompanying drawings in the embodiments of the application. Obviously, the described embodiments are only part of the embodiments of the application, not all. Based on the embodiments in the application, all other embodiments obtained by ordinary skilled in the art without creative labor are within the scope of protection of the application.
[0023] The application discloses an enterprise intelligent consultation management method and system based on big data analysis, which comprises a data collection module, a data preprocessing module, a consultation question input and analysis module, a big data intelligent analysis engine module, an intelligent consultation suggestion generation module, a visualization module and a feedback learning and system optimization module. The data collection module is used for collecting enterprise internal data and external network data. The data preprocessing module is used for cleaning, converting, integrating and storing the collected data. The consultation question input and analysis module proposes an improved text concept semantic fusion algorithm to receive the consultation question input by a user and performs text understanding and structuring. The big data intelligent analysis engine module comprises a task scheduling and management unit and a text analysis engine unit. The task scheduling and management unit is used for receiving the structured consultation task to coordinate the calling of required data resources, model resources and computing resources. The text analysis engine unit proposes an improved bit sequence focusing-bidirectional collaborative entity information extraction algorithm to perform opinion mining and key information extraction. The intelligent consultation suggestion generation module is used for generating structured and executable consultation suggestions. The visualization module is used for visually displaying the analysis result and the consultation suggestion for the user. The feedback learning and system optimization module is used for optimizing the performance of the system model according to the user feedback.
[0024] Preferably, the data collection module accesses enterprise internal system data (such as financial system data, inventory database) through API, database connection and log file reading mode, and obtains external network data (such as news information, policies and regulations) through network crawler, API interface, data subscription and the like.
[0025] Preferably, the data preprocessing module cleans, converts and integrates the collected data, and simultaneously extracts key information (named entity recognition: company, person name, location, time, product, numerical value; event recognition; sentiment analysis) from unstructured text data and stores the same in a structured / semi-structured form, so as to prepare high-quality data basis for text understanding.
[0026] Preferably, the consultation question input and analysis module proposes an improved text concept semantic fusion algorithm to receive the consultation question input by a user and performs text understanding and structuring.
[0027] Specifically, the improved text concept semantic fusion algorithm is as follows: first, receiving the text consultation question input by a user, then, performing text conceptualization on the input consultation question, and the representation of the input text in the concept is as follows: wherein c is a weight vector represented by the input text in the conceptual space, t is a term (i.e. a basic language unit in the text), P(c|t) is a conditional probability that the term belongs to the concept under the condition that the term t appears, and c iTo iterate through all concepts, C represents the complete set of concepts, with data for C sourced from a structured knowledge base and a corpus. count(t,c) represents the number of times term t and concept c co-occur in the knowledge base. Standardization is the total number of occurrences of the standardized term t and all concepts. The statistical time window for co-occurrence frequency is a sliding window (default window size is 1 year). After conceptualization, the text is transformed from the word vector space to the concept space, i.e. in, The corresponding weights of the text under concept c1, The corresponding weights of the text under concept c2, For concept c k To better distinguish between term t and concept c, corresponding weights are assigned to the text. In this context, inverse term frequency and inverse concept frequency are proposed to enhance the correlation between concepts and text. Therefore, the correlation strength between concepts and text is... in, For concept c i The weight of the current text, ST is the weight of the current text, ST i For the i-th short text, t i Let T be the i-th term in the text. i For the term t i Belongs to concept c i The probability, i.e., T i =P(c i |t i ), For concept c i In all relevant terms in the knowledge base, idf c (t i ) is the inverse term frequency within the concept (i.e., term t) i In concept c i scarcity within), icf(c i ) is the inverse concept frequency (i.e., concept c) i (Representativeness within the entire knowledge base) Next, a score function is defined to enrich the semantic knowledge in the text, and dynamic weights and domain-adaptive terms are proposed to improve the score function to enhance the accuracy of short text understanding, i.e., Scoure(x|y,s)=α(s)P co-occur (x|y)+(1-α(s))P semantic( (x|y,s)+β·Domain(x|y), where Scoure(x|y,s) is the improved fractional function, x and y are terms, s is the short text, α(s) is the dynamic weighting coefficient, and P co-occur (x|y) represents the global co-occurrence probability of terms x and y. Psemantic((x|y,s) is the semantic association probability of terms x and y in text s (semantic association probability = concept weight x intra-concept association probability), and β is the domain knowledge weight coefficient, β is initialized as 0.5, and is adjusted according to user feedback as β <- β ± 0.1 sign(accuracy delta ) where sign(·) is a sign function, accuracy delta is the accuracy change value of user feedback, and Domain(x|y) is the association strength of terms x and y in the domain knowledge base, i.e. where P Ent (x,y) is the co-occurrence probability of terms x and y in the enterprise knowledge base, P Ent (x) is the independent occurrence probability of term x in the enterprise knowledge base, and P Ent (y) is the independent occurrence probability of term y in the enterprise knowledge base.
[0028] Then, the conceptual representation provides concept-level information, but lacks modeling of dynamic context dependence, so the self-attention mechanism of the Transformer is used to capture long-distance dependency between terms in the text and generate context-aware vector representation, and the conceptual information is embedded into the Transformer with embedding dimension d model to further capture context semantic information, assuming that the length of the text consultation question input by the user is n, the embedding matrix is M, and the correlation calculation between the query vector query and the key vector key in the self-attention mechanism of the Transformer is Similarity(query,key) = Q K, where Similarity(query,key) is a similarity function, Q is a numerical vector of query, K is a numerical vector of key, and Q K is a dot product operation. In order to realize attention weight normalization, the correlation calculation result is determined using a softmax function, i.e. where r is the weight coefficient of the correlation calculation result, the attention calculation is Attention(Q,K,V) = ∑r V, where Attention(Q,K,V) is an attention output vector, and V is a value vector. Multi-head attention mechanism is used for multi-subspace parallel calculation, i.e. MultiHeadAttention(ST i ) = Concat(head1,…,head k ) W i , where MultiHeadAttention(ST i ) is the i-th text ST iThe final output result after multi-head attention calculation, Concat(·) is the concatenation operation, head1 is the first attention head result, head k is the kth attention head result, W i is the projection matrix, the ith text ST i is projected to the query, key and value vectors, and the head calculation is where head i is the ith attention head result, is the projection matrix for projecting the ith text ST i to the query vector, is the projection matrix for projecting the ith text ST i to the key vector, is the projection matrix for projecting the ith text ST i to the value vector, and thus the context information-rich text representation vector T j is obtained. j-1 +P j +Transition(Norm(T j-1 +P j )), where T j is the jth layer text representation vector, T j-1 is the (j-1)th layer text representation vector, Norm(·) is the normalization operation, P j is the two-dimensional coordinate embedding, and Transition(·) is the depth separable convolution. The improved text concept semantic fusion algorithm first proposes the intra-concept inverse term frequency and inverse concept frequency to improve the relevance between concepts and texts, then proposes dynamic weights and domain adaptive terms to improve the score function to improve the accuracy of short text understanding, and then embeds the conceptualized information into the Transformer to further capture the context semantic information, and uses multiple attention heads to generate text representation vectors, so as to accurately understand the user input text consultation question.
[0029] Preferably, the task scheduling and management unit determines the problem type through the concept weight distribution and semantic density information in the vector, dynamically calculates the priority by fusing user metadata, business rules and vector features, and dynamically allocates the execution path according to the problem type, priority and resource occupation estimation.
[0030] Preferably, the text analysis engine unit proposes an improved bit sequence focusing-bidirectional collaborative entity information extraction algorithm for opinion mining and key information extraction.
[0031] Specifically, the improved bit sequence focusing-bidirectional collaborative entity information extraction algorithm is as follows: first, the text representation vector processed by the consultation question input and analysis module is taken as input, that is wherein, T is a text representation vector sequence, is the first text representation vector, is the second text representation vector, is the nth text representation vector, n is the length of the text representation vector sequence, is the positioning word boundary of the set entity extraction, the boundary detection is activated when the position code is greater than 0.7, the absolute position code is proposed to add the position information of the text representation vector, p i =PositionEmbedding(pos i ), i is the position index of the word in the sequence, p i is the output position embedding vector, PositionEmbedding(·) is a position embedding function, pos i is the absolute position index of the current word in the text, that is, pos i =i, is the fusion of semantic and position information, the vector splicing is performed, that is wherein, is the spliced enhanced semantic vector, concat(·) is a splicing operation, is the ith text representation vector, p i is the position embedding vector, in order to better capture the influence of the entity boundary of the front and rear words, a bidirectional long short-term memory network is introduced to capture the context dependence, and the time step is defined as t, and t = i, the forward LSTM calculation is wherein, is the output hidden state of the forward LSTM at time step t, LSTM(·) is a long short-term memory unit, is the input vector of time step t, is the output hidden state of the forward LSTM at time step t-1, the backward LSTM calculation is wherein, is the output hidden state of the backward LSTM at time step t, is the output hidden state of the backward LSTM at time step t+1, the output of the bidirectional LSTM is
[0032] The CRF (conditional random field) label transition score coefficient λ (0≤λ≤1) is introduced to fuse the features with the LSTM output, that is, m t =λ·h t +(1-λ)s t , wherein, m t is a mixed state vector, ht is the output vector of bidirectional LSTM, s t is the hidden state vector of CRF at time step t, then CRF (Conditional Random Field) decoding is performed to meet the label sequence constraint, i.e. where score(T, y) is the CRF sequence scoring function, y is the candidate label sequence, is the label confidence of being labeled as candidate label y i at position index i, is the label transition score from candidate label y i to y i+1 , n is the length of the text representation vector sequence, then the global optimal label after decoding is calculated as where y * is the optimal label sequence, argmax(·) is the label sequence y that makes score(T, y) maximum, then attention focusing is proposed to highlight the relationship keywords to avoid dilution of the keywords by redundant words when performing relationship extraction, i.e. where a i is the normalized attention weight, u is the attention vector, u T is the transpose of u, tanh(·) is the hyperbolic tangent activation function, W is the weight matrix of attention focusing, W is initialized by Xavier initialization, and the diagonal elements are forced to be 1.0, is the relationship feature vector of the current focusing position index i obtained by bidirectional LSTM output, as in entity extraction, is the relationship feature vector of the traversal comparison position index j, and the weighted aggregation feature vector is The weighted aggregation feature vector c strengthens the features of high-weight positions and weakens the features of low-weight positions to generate a semantic core representation that removes redundant words, then relationship determination is performed based on the semantic core refined by attention, i.e. where, is the relationship probability distribution vector, W rel is the relationship classification weight matrix, b rel is the relationship classification bias vector, and finally the relationship triple is generated. The improved position sequence focusing-bidirectional collaborative entity information refinement algorithm first proposes absolute position encoding to add position information of the text representation vector, then proposes bidirectional long short-term memory network to capture the joint influence of the front and rear words on the entity boundary, and proposes CRF decoding to meet the label sequence constraint, then proposes attention focusing to avoid dilution of the keywords by redundant words, and finally performs relationship determination based on the semantic core refined by attention to generate the relationship triple, so as to realize key information refinement of the text.
[0033] Preferably, the intelligent consultation suggestion generation module reasons according to the relation triple generated by the text engine unit to generate structured and executable consultation suggestions, and the visualization module is used to visually display the analysis results and consultation suggestions generated by the intelligent consultation suggestion generation module for the user.
[0034] Preferably, the feedback learning and system optimization module is used to optimize the performance of the system model according to the user feedback, the user feedback is given to the system according to the obtained consultation suggestions, and the system re-answers the consultation questions according to the feedback of the user, so as to optimize the performance of the system model.
[0035] An enterprise intelligent consultation management method and system based on big data analysis are provided for managing user consultation questions. Through the fusion of the data acquisition module, the data preprocessing module, the consultation question input and analysis module, the big data intelligent analysis engine module, the intelligent consultation suggestion generation module, the visualization module and the feedback learning and system optimization module, an enterprise intelligent consultation management method and system based on big data analysis are provided. An improved text concept semantic fusion algorithm is proposed to receive the consultation questions proposed by the user and perform text understanding and structuring. The innovation of the present application lies in that the improved text concept semantic fusion algorithm first proposes the concept internal inverse term frequency and inverse concept frequency to improve the relevance between the concept and the text, then proposes the dynamic weight and the field self-adaptive item to improve the score function to improve the accuracy of short text understanding, and then embeds the conceptual information into the Transformer to further capture the context semantic information, and uses multiple attention heads to generate the text representation vector, so as to accurately understand the text consultation questions input by the user. An improved bit sequence focus-bidirectional collaborative entity information extraction algorithm is proposed for opinion mining and key information extraction. The innovation of the present application lies in that the improved bit sequence focus-bidirectional collaborative entity information extraction algorithm first proposes the absolute position encoding to add the position information of the text representation vector, then proposes the bidirectional long short-term memory network to capture the common influence of the entity boundary of the previous and subsequent words, and proposes the CRF decoding to meet the label sequence constraint, then proposes the attention focus to avoid the dilution of the key words by the redundant words, and finally performs relation determination based on the semantic core extracted by the attention to generate the relation triple, so as to realize the key information extraction of the text and effectively improve the working effect of the enterprise intelligent consultation management method and system based on big data analysis, and provide more comprehensive and accurate technical support for the enterprise intelligent consultation management method and system, and provide better decision support for the scientific and efficient enterprise intelligent consultation management system. Meanwhile, the present application relates to natural language processing and deep learning technology, provides an accurate and efficient enterprise intelligent consultation management system, and contributes to the important application value of enterprise intelligent consultation management.
[0036] Although the present application has been described in detail with reference to the foregoing embodiments, the technical solutions recorded in the foregoing embodiments can be modified, or some of the technical features can be replaced by equivalent features, by those skilled in the art, any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application shall be included in the protection scope of the present application.
Claims
1. A method and system for enterprise intelligent consulting management based on big data analysis, characterized in that, The system comprises a data collection module, a data preprocessing module, an advisory question input and analysis module, a big data intelligent analysis engine module, an intelligent advisory suggestion generation module, a visualization module, and a feedback learning and system optimization module. The data collection module is used to collect enterprise internal data and external network data. The data preprocessing module is used to clean, convert, integrate, and store the collected data. The advisory question input and analysis module proposes an improved text concept semantic fusion algorithm to receive user-proposed advisory questions and perform text understanding and structuring. The big data intelligent analysis engine module comprises a task scheduling and management unit and a text analysis engine unit. The task scheduling and management unit is used to receive structured advisory tasks to coordinate the calling of required data resources, model resources, and computing resources. The text analysis engine unit proposes an improved bit sequence focus-bidirectional collaborative entity information extraction algorithm to perform opinion mining and key information extraction. The intelligent advisory suggestion generation module is used to generate structured and executable advisory suggestions. The visualization module is used to visually display analysis results and advisory suggestions for users. The feedback learning and system optimization module is used to optimize the performance of system models according to user feedback. The improved text concept semantic fusion algorithm first proposes concept internal inverse term frequency and inverse concept frequency to improve the relevance between concepts and text. Then, dynamic weights and domain adaptive items are proposed to improve the precision of short text understanding. After that, the conceptualized information is embedded into a Transformer to further capture contextual semantic information, and multiple attention heads are used to generate text representation vectors. The improved bit sequence focus-bidirectional collaborative entity information extraction algorithm first proposes absolute position encoding to add position information to the text representation vector. Then, a bidirectional long short-term memory network is proposed to capture the joint influence of previous and subsequent words on entity boundaries. CRF decoding is proposed to meet the label sequence constraints. Attention focusing is proposed to avoid dilution of key words by redundant words. Finally, the semantic core extracted based on attention is used to determine relationships to generate relationship triples.
2. The method and system for enterprise intelligent consulting management based on big data analysis according to claim 1, characterized in that, The data collection module accesses enterprise internal system data through API, database connection, and log file reading methods, and obtains external network data through web crawlers, API interfaces, data subscriptions, and the like.
3. The method and system for enterprise intelligent consulting management based on big data analysis according to claim 1, characterized in that, The data preprocessing module cleans, converts, and integrates the collected data, extracts key information from unstructured text data, and stores it in a structured / semi-structured form, thereby preparing high-quality data for text understanding.
4. The method and system for enterprise intelligent consulting management based on big data analysis according to claim 1, characterized in that, The advisory question input and analysis module proposes an improved text concept semantic fusion algorithm to receive user-proposed advisory questions and perform text understanding and structuring.
5. The method and system for enterprise intelligence consulting management based on big data analysis according to claim 4, characterized in that, The improved text concept semantic fusion algorithm is as follows: first, receiving the text consultation question input by the user, then, text conceptualization is performed on the input consultation question, and the representation of the input text in the concept is: Wherein, c is the weight vector represented by the input text in the conceptual space, t is the term, P(c|t) is the conditional probability that the term belongs to the concept c in the case of the occurrence of the term t, c i is the traversal of all concepts, C is the concept set, the data source of C comes from the structured knowledge base and corpus, count(t, c) is the co-occurrence number of the term t and the concept c in the knowledge base, is the total co-occurrence number of the standardized term t and all concepts, the standardization is The co-occurrence number is counted in a sliding window, and the text is converted from the word vector space to the concept space after conceptualization, that is, Wherein, is the corresponding weight of the text under the concept c1, is the corresponding weight of the text under the concept c2, is the corresponding weight of the text under the concept c k In order to make the term t and the concept c have better distinguishing ability, the intra-concept inverse term frequency and the inverse concept frequency are proposed to improve the relevance between the concept and the text, so the association strength between the concept and the text is Wherein, is the corresponding weight of the text under the concept c i to the current text, ST is the current text, ST i is the i-th short text, t i is the i-th term in the text, T i is the probability that the term t i belongs to the concept c i , that is, T i =P(c i |t i ), is the all related terms of the concept c i in the knowledge base, idf c (t i ) is the intra-concept inverse term frequency, icf(c i ) is the inverse concept frequency, then, a score function is defined to enrich the semantic knowledge in the text, and a dynamic weight and a domain adaptive term are proposed to improve the score function to improve the accuracy of short text understanding, that is, Scoure(x|y,s)=α(s)P co-occur (x|y)+(1-α(s))P semantic ((x|y,s) + β · Domain(x|y), where Score(x|y,s) is the improved score function, x and y are terms, s is short text, a(s) is a dynamic weight coefficient, P co-occur (x|y) is the global co-occurrence probability of terms x and y, P semantic ((x|y,s) is the semantic association probability of terms x and y in text s, β is the domain knowledge weight coefficient, β is initialized as 0.5, and is adjusted according to user feedback as β ← β ± 0.1 · sign(accuracy delta ) where sign(·) is the sign function, accuracy delta is the accuracy change value of user feedback, Domain(x|y) is the association strength of terms x and y in the domain knowledge base, i.e. where P Ent (x,y) is the co-occurrence probability of terms x and y in the enterprise knowledge base, P Ent (x) is the independent occurrence probability of term x in the enterprise knowledge base, P Ent (y) is the independent occurrence probability of term y in the enterprise knowledge base; Then, the conceptualized information is embedded by Transformer with dimension d model , to further capture the contextual semantic information, assuming that the length of the text consultation question input by the user is n, the embedding matrix is M, and the correlation calculation between the query vector query and the key vector key in the self-attention mechanism of the Transformer is Similarity(query, key) = Q·K, wherein Similarity(query, key) is a similarity function, q is a numerical vector of query, K is a numerical vector of key, and Q·K is a dot product operation. In order to realize the normalization of attention weight, the correlation calculation result is determined by using a softmax function, i.e. r = softmax(Q·K) , wherein r is a weight coefficient of the correlation calculation result, the attention calculation is Attention(Q, K, V) = ∑r·V, wherein Attention(Q, K, V) is an attention output vector, V is a value vector, and multi-head attention mechanism is used for multi-subspace parallel calculation, i.e. MultiHeadAttention(ST i ) = Concat(head1, …, head k )W i , wherein MultiHeadAttention(ST i ) is the final output result after multi-head attention calculation on the i-th text ST i , Concat(·) is a concatenation operation, head1 is the first attention head result, head k is the k-th attention head result, and W i is a projection matrix for projecting the i-th text ST i of the embedding matrix M to the query, key and value vectors. The head calculation is , wherein head i is the i-th attention head result, is a projection matrix for projecting the i-th text ST i to the query vector, is a projection matrix for projecting the i-th text ST i to the key vector, is a projection matrix for projecting the i-th text ST i to the value vector, and thus the text representation vector rich in contextual information is T j = Norm(Norm(T j-1 + P j ) + Transition(Norm(T j-1 + P j ))), wherein T j is the j-th layer text representation vector, and T j-1 is the text representation vector for the j-1th layer, Norm(·) is the normalization operation, P j is the two-dimensional coordinate embedding, and Trannsition(·) is a depthwise separable convolution.
6. The method and system for enterprise intelligent consulting management based on big data analysis according to claim 1, characterized in that, The task scheduling and management unit determines the problem type based on the concept weight distribution and semantic density information in the vector, dynamically calculates the priority by fusing user metadata, business rules, and vector features, and dynamically allocates an execution path according to the problem type, priority, and resource occupation estimation.
7. The method and system for enterprise intelligent consulting management based on big data analysis according to claim 6, characterized in that, The improved bit sequence focusing-bidirectional collaborative entity information extraction algorithm is as follows: first, the text representation vector processed by the consultation question input and analysis module is taken as input, i.e. wherein T is a text representation vector sequence, is the first text representation vector, is the second text representation vector, is the nth text representation vector, n is the length of the text representation vector sequence, is the positioning word boundary of the set entity extraction, the boundary detection is activated when the position code is greater than 0.7, the absolute position code is proposed to add the position information of the text representation vector, p i =PositionEmbedding(pos i ), i is the position index of the word in the sequence, p i is the output position embedding vector, PositionEmbedding(·) is a position embedding function, pos i is the absolute position index of the current word in the text, i.e. pos i = i, is the fusion of semantic and position information, the vectors are spliced, i.e. wherein, is the spliced enhanced semantic vector, concat(·) is a splicing operation, is the ith text representation vector, p i is the position embedding vector, in order to better capture the influence of the entity boundary of the front and rear words, a bidirectional long short-term memory network is introduced to capture the context dependence, the time step is defined as t, and t = i, the forward LSTM calculation is wherein, is the output hidden state of the forward LSTM at time step t, LSTM(·) is a long short-term memory unit, is the input vector at time step t, is the output hidden state of the forward LSTM at time step t-1, the backward LSTM calculation is wherein, is the output hidden state of the backward LSTM at time step t, is the output hidden state of the backward LSTM at time step t+1, and the output of the bidirectional LSTM is A CRF (Conditional Random Field) label transition score coefficient λ (0≤λ≤1) is introduced to fuse with the features output by the LSTM, i.e. m t = λ·h t + (1-λ)s t , where m t is the mixed state vector, h t is the output vector of the bidirectional LSTM, s t is the hidden state vector of the CRF at time step t, and then the CRF decoding is performed to meet the label sequence constraint, i.e. where score(T,y) is the CRF sequence scoring function, y is the candidate label sequence, is the label confidence of being labeled as candidate label y i at position index i, is the label transition score from candidate label y i to y i+1 , and n is the length of the text representation vector sequence, and then the global optimal label after decoding is calculated as where y * is the optimal label sequence, and argmax(·) is the label sequence y when the value of score(T,y) is maximum, and then in the relationship extraction, attention focusing is proposed to highlight the relationship keywords to avoid dilution of the keywords by redundant words, i.e. where α i is the normalized attention weight, u is the attention vector, u T is the transpose of u, tanh(·) is the hyperbolic tangent activation function, W is the weight matrix of attention focusing, and W is initialized by the Xavier initialization method, and the diagonal elements are forced to be 1.0, is the relationship feature vector of the current focusing position index i, is the relationship feature vector of the traversal comparison position index j, and the weighted aggregated feature vector is and then the semantic core based on attention refinement is used to determine the relationship, i.e. where, is the relationship probability distribution vector, W rel is the relationship classification weight matrix, and b rel is the relationship classification bias vector, and finally the relationship triplets are generated.
8. The method and system for enterprise intelligent consulting management based on big data analysis according to claim 1, characterized in that, The intelligent consultation suggestion generation module performs reasoning according to the relation triple generated by the text engine unit, so as to generate structured and executable consultation suggestions, and the visualization module is used for visually displaying the analysis result and the consultation suggestion generated by the intelligent consultation suggestion generation module for the user. 9.The enterprise intelligent consulting management method and system based on big data analysis of claim 1, wherein, The feedback learning and system optimization module is used for performance optimization of the system model according to user feedback, the system re-answers the consultation question according to the feedback of the user to the system, so as to perform performance optimization on the system model.
Citation Information
Patent Citations
Text conceptualization and network representation fused opinion retrieval system and method
CN108399238A
Knowledge graph-based high-risk App detection and identification method
CN116910754A