Enterprise-level intelligent customer service interaction control system under multi-round dialogue scene

Through the combination of hybrid neural network model and dynamic decision-making engine, the enterprise-level intelligent customer service system has solved the problem of insufficient understanding of dialects and industry terms and data security in multiple rounds of dialogue, and achieved a high accuracy and security enterprise-level intelligent customer service system.

CN120372680AActive Publication Date: 2025-07-25HEBEI BITJUKE TECH CO LTD

Patent Information

Application Number
CN202510431548.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-08
Publication Date
2025-07-25
Estimated Expiration
2045-04-08

AI Technical Summary

Technical Problem

When existing enterprise-level intelligent customer service systems handle complex multi-round dialogue, they find it difficult to fully understand dialects and industry terms, lack the ability to identify sensitive information, and there are problems with data security and privacy protection.

Method used

Mixed neural network model (Transformer encoder and BiLSTM network) is used to perform diversified text analysis, combine the dynamic decision engine to judge sensitive information and select retrieval paths, correct errors in real time through closed-loop learning optimization module, and use the Guomi algorithm to achieve non-cloud data transmission.

Benefits of technology

It improves the accuracy and security of the system in multiple rounds of dialogue, ensures intelligent identification and secure processing of sensitive information, adapts to different contexts, reduces the risk of data leakage, and improves the quality and security of customer service.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120372680A_ABST
    Figure CN120372680A_ABST
Patent Text Reader

Abstract

The invention provides an enterprise-level intelligent customer service interaction control system under a multi-round dialogue scene. The enterprise-level intelligent customer service interaction control system comprises a multi-modal semantic analysis module, a dynamic decision engine module, a closed-loop learning optimization module and a multi-terminal safety delivery module, according to the system, semantic analysis of diversified texts is realized through a hybrid neural network model, text sensitivity is dynamically judged, a corresponding retrieval path is selected, the performance of the hybrid neural network model is optimized in real time, and a cryptographic algorithm is adopted for encryption to ensure secure transmission of data; according to the method, the defects of a traditional customer service system in the aspects of complex context understanding, sensitive information identification and data secure transmission are effectively overcome, the accuracy, safety and adaptability of the system are remarkably improved, and the method is particularly suitable for enterprise-level customer service scenes needing to process dialects, terminologies and sensitive information and has important application value.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of intelligent customer service systems, and more specifically to an enterprise-level intelligent customer service interaction control system in a multi-turn dialogue scenario. Background Art

[0002] Enterprise-level intelligent customer service interaction control systems are a type of technical solution widely used by enterprises in the customer service process, aiming to improve the quality and efficiency of customer service through automated and intelligent means. Such systems usually rely on technologies such as natural language processing, machine learning, and deep learning, and can handle customer requests from different channels and provide personalized services. With the continuous development of technology, intelligent customer service systems have gradually acquired the ability to handle multi-turn dialogues and can respond dynamically according to the user's context;

[0003] However, existing enterprise-level intelligent customer service systems have obvious deficiencies in many aspects; firstly, when dealing with complex multi-turn dialogues, existing systems are difficult to fully understand and parse different types of text content, especially when it comes to dialects, industry terms, or proper nouns, the recognition accuracy is relatively low, resulting in inaccurate responses; secondly, the existing technology has weak capabilities in judging and protecting sensitive information. The system usually relies on a single data retrieval path and cannot dynamically adjust the data processing strategy according to the sensitivity of the content, which is prone to the risk of data leakage; finally, many traditional systems rely on public cloud platforms for data storage and processing, facing privacy security and compliance issues. Summary of the Invention

[0004] In order to solve the technical problems mentioned in the current background art, the present invention effectively solves the problems of multi-turn dialogue understanding, sensitive information recognition, and data security transmission by introducing a hybrid neural network model, a dynamic decision-making engine, and a closed-loop learning optimization mechanism, improving the accuracy, security, and adaptability of the system.

[0005] To achieve the above objectives, the present invention adopts the following technical solutions:

[0006] An enterprise-level intelligent customer service interaction control system in a multi-turn dialogue scenario, comprising:

[0007] M1. A multi-modal semantic parsing module, deploying a hybrid neural network model composed of a Transformer encoder and a BiLSTM network; calculating scores of diverse text content through the hybrid neural network model, and further calculating score confidence; outputting structured semantic information including the diverse text content and the corresponding score confidence;

[0008] M2, the dynamic decision-making engine output module, which determines whether the diversified text content is sensitive information according to the structured semantic information to select a retrieval path. If it is sensitive information, the retrieval path is the core database; otherwise, the retrieval path is the extended knowledge base. Retrieve from the core database or the extended knowledge base based on the retrieval path to generate a retrieval result.

[0009] M3, the closed-loop learning optimization module, which captures the dialect recognition error, the emotional response deviation value, and the term library missing warning signal of the hybrid neural network model in real time and inputs them into the incremental model. The incremental model is used to optimize the hybrid neural network model. Synchronize the optimized hybrid neural network model to each business terminal through a secure channel. The incremental model is established in the enterprise intranet data sandbox.

[0010] M4, the multi-terminal dialogue secure delivery module, which generates and feeds back the user response text through the optimized neural network model, integrates the enterprise-level gateway and the private API communication protocol, and uses the national cryptographic algorithm to encrypt the channel to achieve full-link non-cloud data transmission.

[0011] Further, the diversified text content includes natural language, dialect features, and industry terms.

[0012] Convert each word in the diversified text content input by the user into a word vector representation X through word embedding i ; X i represents the word vector of the i-th word;

[0013] Subsequently, perform a linear transformation on each word vector X i to extract the query vector A i , the key vector B i and the value vector C i ;

[0014] The Transformer encoder optimizes the representation of each word vector by calculating the attention weights between the word vectors through the self-attention mechanism. The calculation formula for the attention weights is:

[0015]

[0016] where represents the self-attention mechanism; d is the dimension of the key vector; represents the Softmax function;

[0017] Through calculation, the Transformer encoder generates an optimized word vector sequence H, H = [h1, h2,..., h n , where h i is the optimized word vector of the i-th word in the text.

[0018] Furthermore, the optimized word vector sequence H is passed to the BiLSTM network to further model the time series dependencies in the word vector sequence. The calculation formula is as follows:

[0019]

[0020] where g t-1 is the hidden state at the previous moment; g t is the hidden state at the current moment, representing the context information of the current optimized word vector; h t is the optimized word vector sequence at the t-th moment, that is, h t comes from H; represents the LSTM function;

[0021] The hidden state g t output by the Transformer encoder is linearly transformed to generate scores corresponding to the natural language, dialect features, and industry terms;

[0022] The calculation formula for the score of the natural language is:

[0023] z1 = W1g t + b1

[0024] where z1 represents the natural language score; W1 represents the weight matrix related to the natural language; b1 represents the bias term related to the natural language;

[0025] The calculation formula for the score of the dialect features is:

[0026] z2 = W2g t + b2

[0027] where z2 represents the dialect feature score; W2 represents the weight matrix related to the dialect features; b2 represents the bias term related to the dialect features;

[0028] The calculation formula for the score of the industry terms is:

[0029] z3 = W3g t + b3

[0030] where z3 represents the industry term score; W3 represents the weight matrix related to the industry terms; b3 represents the bias term related to the industry terms.

[0031] Furthermore, the scores corresponding to the natural language, dialect features, and industry terms are passed to the Softmax function to calculate the score confidence of each natural language, dialect feature, and industry term;

[0032] The calculation formula for the confidence P1 of the natural language score is:

[0033] Among them is the normalization factor of all scores, ensuring that the sum of the final probabilities is 1;

[0034] The calculation formula for the confidence P2 of the dialect feature score is:

[0035] The calculation formula for the confidence P3 of the industry term score is:

[0036] Furthermore, a confidence weighted evaluation is performed on the structured semantic information, and the specific evaluation formula is extended to:

[0037]

[0038] Among them, is the confidence of the weighted natural language; The confidence of the weighted dialect feature; is the confidence of the weighted industry term; C1, C2, and C3 are context adjustment factors; is the factor for normalizing all scores, ensuring that the sum of the final probabilities is 1;

[0039] The initial structured semantic information is denoted as I; the weighted structured semantic information is denoted as Among them includes diverse text content and the confidences of the weighted natural language, dialect features, and industry terms;

[0040] Judge whether the diverse text content is sensitive information according to the weighted confidence value to determine the retrieval path: First, set the sensitivity threshold to 0.8; when the weighted confidence value is greater than or equal to the preset sensitivity threshold, it is marked as sensitive information, and the core database is selected as the retrieval path; when the weighted confidence value is lower than the set sensitivity threshold, it is marked as non-sensitive information, and the extended knowledge base is selected as the retrieval path.

[0041] Furthermore, the retrieval process of the core database is to input the sensitive information into the core database, encrypt the input sensitive information through an encrypted retrieval function, and then perform retrieval in the database through an indexing technique to obtain the retrieval result.

[0042] Furthermore, the dialect recognition error refers to the error that occurs when the hybrid neural network model processes accents, vocabulary, or grammar in different regions, and the calculation formula is:

[0043]

[0044] Among them, E dialectIndicates the dialect recognition error; Indicates the predicted value of the hybrid neural network system for the i-th dialect sample; y i Indicates the actual label value, that is, the correct dialect recognition result; n is the number of dialect samples;

[0045] The emotional response deviation value indicates the difference between the prediction result and the actual user emotion when the hybrid neural network model performs emotional analysis. The calculation formula is:

[0046]

[0047] Among them, E emotion Indicates the emotional response deviation value; Indicates the predicted emotion value of the hybrid neural network model for the i-th emotion sample; Indicates the actual emotion value; n is the number of emotion samples;

[0048] The term library missing alarm signal indicates that during the processing of the hybrid neural network system, missing industry terms are found. The calculation formula is:

[0049]

[0050] Among them, E term Indicates the term missing alarm signal; Term i Indicates the i-th term; T represents the current term library; 1 is an indicator function, which returns 1 when the term Term i is not in the term library T, otherwise returns 0; n is the number of terms.

[0051] Furthermore, the training steps of the incremental model include training data preprocessing, incremental model initialization, incremental training loop, real-time signal monitoring, and security auditing and rollback;

[0052] The incremental model initialization uses the hybrid neural network as the basic architecture for incremental training, and sets the weight and bias parameters through the Xavier initialization method. The calculation formula is:

[0053]

[0054] Among them, w ij is the weight parameter, indicating the weight from the i-th layer to the j-th layer; U(·,·) represents randomly extracting a value uniformly from the interval; n in is the number of neurons in the input layer; n out is the number of neurons in the output layer;

[0055] The steps of the incremental training loop include forward propagation, loss calculation, backpropagation, and gradient calculation;

[0056] The real-time signal monitoring evaluates the accuracy of the incremental model by monitoring the dialect recognition error, the emotional response deviation value, and the term library missing alarm signal of the hybrid neural network model, and continues the incremental training loop when the accuracy is lower than the set first threshold;

[0057] The security audit and rollback means that when the accuracy of the incremental model is lower than the second threshold, a rollback mechanism is started to restore the incremental model to the state before the previous incremental training loop.

[0058] Further, the forward propagation training is to perform forward propagation through the incremental model to calculate the prediction result. The calculation formula of the forward propagation is:

[0059]

[0060] where is the prediction output of the incremental model; X is the word vector; θ is the parameter of the incremental model;

[0061] The loss calculation calculates the difference between the predicted value and the actual label of the incremental model through the loss function and uses the cross-entropy loss function for calculation. The specific calculation formula is:

[0062] L = -(αE dialect + βE emotion + γE term )

[0063] where L is the loss function; α, β, and γ are weight coefficients; E dialect represents the dialect recognition error; E emotion represents the emotional response deviation value; E term represents the term library missing alarm signal;

[0064] The backpropagation and gradient calculation is to calculate the gradient of the loss function with respect to the parameters of the incremental model through backpropagation and update the parameters of the incremental model using gradient descent. The formula for the update process is:

[0065]

[0066] where θ new represents the updated parameters of the incremental model; θ old represents the parameters of the incremental model before update; η is the learning rate, which controls the step size of parameter update; is the gradient of the loss function with respect to the parameters of the incremental model.

[0067] Compared with the prior art, the advantages of the present invention are:

[0068] 1. The present invention adopts a hybrid neural network model, combining a Transformer encoder and a BiLSTM network, to precisely process diverse text content. The Transformer encoder can optimize the semantic representation of the text through the self-attention mechanism, while the BiLSTM network can effectively capture the dependencies in the time series, ensuring that the system can accurately understand the context of each round in a multi-round conversation. Through this technological innovation, the accuracy of the system in processing different contexts and complex conversations is significantly improved, making the interaction between the customer and the system smoother and more accurate.

[0069] 2. The present invention uses a dynamic decision-making engine to intelligently determine whether the text is sensitive information based on the confidence calculation of the text, and accordingly select an appropriate retrieval path. Specifically, the system will perform a weighted evaluation based on the confidence values of diverse text content. When the confidence value exceeds the set threshold, the sensitive information will be processed through encryption and the core database will be selected for retrieval; when the confidence value is lower than the threshold, the system will select the extended knowledge base for retrieval. This innovative mechanism ensures the intelligent identification and secure processing of sensitive information, avoiding the risk of data leakage caused by relying on a single retrieval path in traditional systems. At the same time, through non-cloud data transmission and a private API communication protocol, the data security is further enhanced.

[0070] 3. The present invention introduces a closed-loop learning optimization module, which can capture and correct problems such as dialect recognition errors, emotional response deviations, and missing term libraries in real time. When the system identifies an error in dialect or industry terms, it will input the error signal into the incremental model for optimization, and improve the accuracy of the hybrid neural network model through incremental learning. This mechanism enables the system to continuously adapt to new language environments and business scenarios, improving the accuracy and response ability during long-term operation, and ensuring high-quality customer service in different contexts. BRIEF DESCRIPTION OF THE DRAWINGS

[0071] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0072] Figure 1 It is a schematic diagram of the system working process of the present invention;

[0073] Figure 2 It is a schematic diagram of the working process of the multi-modal semantic parsing module of the present invention;

[0074] Figure 3 It is a schematic diagram of the incremental model training process of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0075] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the following will clearly and completely describe the technical solutions in the embodiments of this application with reference to the accompanying drawings in the embodiments of this application. Obviously, the described embodiments are some, but not all, of the embodiments of this application. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in this application without creative efforts fall within the scope of protection of this application.

[0076] To achieve the above objectives, the present invention is implemented through the following technical solutions. The present invention provides an enterprise-level intelligent customer service interaction control system in a multi-round dialogue scenario, as Figures 1 - 3 shown. The system includes:

[0077] M1. A multi-modal semantic parsing module that deploys a hybrid neural network model composed of a Transformer encoder and a BiLSTM network; calculates scores for diverse text contents through the hybrid neural network model, and further calculates score confidence levels; outputs structured semantic information including the diverse text contents and the corresponding score confidence levels.

[0078] The diverse text contents include natural language, dialect features, and industry terms;

[0079] 1) Convert each word in the diverse text content input by the user into a word vector representation X through word embedding i ; X i represents the word vector of the i-th word;

[0080] Subsequently, perform a linear transformation on each word vector X i to extract a query vector A i , a key vector B i , and a value vector C i

[0081] 2) The Transformer encoder calculates the attention weights between the word vectors through a self-attention mechanism to optimize the representation of each word vector. The calculation formula for the attention weights is:

[0082]

[0083] where represents the self-attention mechanism; d is the dimension of the key vector; represents the Softmax function;

[0084] Through calculation, the Transformer encoder generates an optimized word vector sequence H, H = [h1, h2,..., h n , where hi is the optimized word vector of the i-th word in the text;

[0085] 3) The optimized word vector sequence H is passed to the BiLSTM network to further model the time series dependence in the word vector sequence. The calculation formula is:

[0086]

[0087] where, g t-1 is the hidden state at the previous moment; g t is the hidden state at the current moment, representing the context information of the current optimized word vector; h t is the optimized word vector sequence at the t-th moment, that is, h t comes from H; represents the LSTM function;

[0088] 4) The hidden state g output by the Transformer encoder t generates scores corresponding to the natural language, dialect features, and industry terms through linear transformation;

[0089] The calculation formula for the score of the natural language is:

[0090] z1 = W1g t + b1

[0091] where, z1 represents the natural language score; W1 represents the weight matrix related to the natural language; b1 represents the bias term related to the natural language, used to adjust the natural language score;

[0092] The calculation formula for the score of the dialect features is:

[0093] z2 = W2g t + b2

[0094] where, z2 represents the dialect feature score; W2 represents the weight matrix related to the dialect features; b2 represents the bias term related to the dialect features, used to adjust the dialect feature score;

[0095] The calculation formula for the score of the industry terms is:

[0096] z3 = W3g t + b3

[0097] where, z3 represents the industry term score; W3 represents the weight matrix related to the industry terms; b3 represents the bias term related to the industry terms, used to adjust the industry term score;

[0098] 5) Pass the scores corresponding to natural language, dialect features, and industry terms to the Softmax function to calculate the score confidence of each natural language, dialect feature, and industry term;

[0099] The calculation formula for the confidence P1 of the natural language score is:

[0100] where is the normalization factor of all scores, ensuring that the sum of the final probabilities is 1;

[0101] The calculation formula for the confidence P2 of the dialect feature score is:

[0102] The calculation formula for the confidence P3 of the industry term score is:

[0103] M2, the dynamic decision engine output module, which determines whether the diverse text content is sensitive information according to the structured semantic information to select a retrieval path. If it is sensitive information, the retrieval path is the core database; otherwise, the retrieval path is the extended knowledge base. Retrieve from the core database or the extended knowledge base based on the retrieval path to generate a retrieval result;

[0104] 1) Perform a weighted evaluation on the confidence in the structured semantic information. Specifically, the evaluation formula for the confidence is extended to:

[0105]

[0106] where, is the weighted natural language confidence; the weighted dialect feature confidence; is the weighted industry term confidence; C1, C2, C3 are context adjustment factors; is the factor for normalizing all scores, ensuring that the sum of the final probabilities is 1;

[0107] Denote the initial structured semantic information as I; the weighted structured semantic information as where includes diverse text content and the weighted confidences of natural language, dialect features, and industry terms;

[0108] 2) Determine whether the diverse text content is sensitive information according to the weighted confidence values to determine the retrieval path;

[0109] Set the sensitive threshold to 0.8;

[0110] When the weighted confidence value is greater than or equal to the preset sensitivity threshold, it indicates that the text content contains more sensitive information, which is marked as sensitive information, and the core database is selected as the retrieval path;

[0111] When the weighted confidence value is lower than the set sensitivity threshold, it indicates that the text content contains less sensitive information, which is marked as non-sensitive information, and the extended knowledge base is selected as the retrieval path.

[0112] The retrieval process of the core database is to input the sensitive information into the core database. After encrypting the input sensitive information through an encryption retrieval function, it is retrieved in the database through an indexing technique to obtain a retrieval result;

[0113] The retrieval process of the extended knowledge base is to input the non-sensitive information into the extended knowledge base and retrieve it in the database through an indexing technique to obtain a retrieval result.

[0114] M3, a closed-loop learning optimization module, which captures the dialect recognition error, emotional response deviation value, and term library missing warning signal of the hybrid neural network model in real time and inputs them into the incremental model; the incremental model is used to optimize the hybrid neural network model; the optimized hybrid neural network model is synchronized to each business terminal through a secure channel; the incremental model is established in the enterprise intranet data sandbox;

[0115] The dialect recognition error refers to the error that occurs when the hybrid neural network model processes accents, vocabulary, or grammar in different regions. The calculation formula is:

[0116]

[0117] where E dialect represents the dialect recognition error; represents the predicted value of the hybrid neural network system for the i-th dialect sample; y i represents the actual label value, that is, the correct dialect recognition result; n is the number of dialect samples;

[0118] The emotional response deviation value represents the difference between the prediction result and the actual user emotion when the hybrid neural network model performs emotional analysis. The calculation formula is:

[0119]

[0120] where E emotion represents the emotional response deviation value; represents the predicted emotion value of the hybrid neural network model for the i-th emotion sample; represents the actual emotion value; n is the number of emotion samples;

[0121] The missing term library alarm signal indicates that during the processing of the hybrid neural network system, a missing industry term is found. The calculation formula is as follows:

[0122]

[0123] where E term represents the missing term alarm signal; Term i represents the i-th term; T represents the current term library (including all known industry terms); 1 is an indicator function, which returns 1 when the term Term i is not in the term library T, otherwise it returns 0; n is the number of terms;

[0124] Based on the calculation of the dialect recognition error, the emotional response deviation value, and the missing term library alarm signal, to ensure that the incremental model can process and store data more securely during the training phase, a security sandbox is further introduced. The security sandbox environment includes hardware configuration and data isolation:

[0125] The hardware configuration is an independent server cluster deployed in the enterprise intranet, equipped with a national secret second-level authentication encryption card;

[0126] The data isolation uses dual-channel memory isolation technology to achieve physical memory partitioning. The expression formula is as follows:

[0127] M train = M safe ⊕ M mask

[0128] where M train represents the data matrix finally used to train the incremental model; M safe is the original training data matrix in the security sandbox; M mask is a dynamically generated random mask matrix, which is destroyed after each training;

[0129] In the security sandbox environment that ensures data security, enter the training phase of the incremental model. The training steps of the incremental model include training data preprocessing, incremental model initialization, incremental training loop, real-time signal monitoring, and security audit and rollback:

[0130] The input x of the incremental model includes the structured semantic information output from M1; the structured semantic information weighted by the confidence value and the retrieval path output from M2; the specific retrieval result output from M3; the dialect recognition error, the emotional response deviation value, and the missing term library alarm signal captured by M4;

[0131] 1) The preprocessing of the training data means that the incremental model collects the input data and processes it through steps of cleaning, denoising, and standardization to ensure that the data can meet the needs of subsequent incremental model training;

[0132] 2) The initialization of the incremental model uses the hybrid neural network as the basic architecture for incremental training, and sets the weight and bias parameters through the Xavier initialization method. The calculation formula is:

[0133]

[0134] where, w ij is the weight parameter, representing the weight from the i-th layer to the j-th layer; U(·,·) represents randomly extracting a value uniformly from the interval; n in is the number of neurons in the input layer; n out is the number of neurons in the output layer;

[0135] 3) The steps of the incremental training loop include forward propagation, loss calculation, backpropagation, and gradient calculation;

[0136] ① The forward propagation means that the training performs forward propagation through the incremental model to calculate the prediction result. The calculation formula for the forward propagation is:

[0137]

[0138] where, is the predicted output of the incremental model; X is the input data; θ is the parameter (weight and bias) of the incremental model;

[0139] ② The loss calculation calculates the difference between the predicted value of the incremental model and the actual label through the loss function. In this embodiment, the cross-entropy loss function is used for calculation. The specific calculation formula is:

[0140] L = -(αE dialect + βE emotion + γE term )

[0141] where, L is the loss function; α, β, γ are weight coefficients; E dialect represents the dialect recognition error; E emotion represents the emotional response deviation value; E term represents the warning signal of the missing term library;

[0142] ③ The backpropagation and gradient calculation means calculating the gradient of the loss function with respect to the parameters of the incremental model through backpropagation, and updating the parameters of the incremental model using gradient descent. The formula for the update process is:

[0143]

[0144] Among them, θ new represents the parameters of the updated incremental model; θ old represents the parameters of the incremental model before update; η is the learning rate, which controls the step size of parameter update; is the gradient of the loss function with respect to the parameters of the incremental model;

[0145] 4) The real-time signal monitoring evaluates the accuracy of the incremental model by monitoring the dialect recognition error, emotional response deviation value, and term library missing alarm signal of the hybrid neural network model, and continues the incremental training loop when the accuracy is lower than the set first threshold of 0.7;

[0146] 5) The security audit and rollback means that when the accuracy of the incremental model is lower than the second threshold of 0.5, a rollback mechanism is started to restore the incremental model to the state before the previous incremental training loop;

[0147] The optimized hybrid neural network model refers to the hybrid neural network model after incremental training, which is synchronized to each business terminal through a national cryptographic algorithm encryption channel;

[0148] M4, the multi-terminal dialogue security delivery module, generates and feeds back the user response text through the optimized neural network model, integrates the enterprise-level gateway and the private API communication protocol, and uses the national cryptographic algorithm encryption channel to achieve full-link non-cloud data transmission.

[0149] The hybrid neural network model outputs the intermediate vector result of the response intention and converts it into a natural language response text through lexical mapping and feeds it back to the user;

[0150] The private API communication protocol ensures that all data transmissions do not rely on external public cloud services, but are completed in the enterprise's internal secure environment, providing encryption protection for communication, ensuring the confidentiality and integrity of data transmission, and preventing data from being leaked or tampered with during transmission;

[0151] The text data is encrypted and transmitted through the national cryptographic algorithm encryption channel. The national cryptographic algorithms (such as SM2, SM3, and SM4) are encryption standards issued by the China National Cryptography Administration, which are used to ensure that the text data in transmission is fully encrypted and protected, ensuring that the text data is not accessed or tampered with by unauthorized parties during transmission, and meeting the national information security standards;

[0152] The full-link non-cloud-based data transmission means that data is not transmitted through public cloud services, but is processed and transmitted through the enterprise's own network and infrastructure. This design reduces the dependence on external cloud platforms, avoids potential data security risks, and ensures that the enterprise has full control over its sensitive data.

[0153] As described above, it is only the specific implementation manner of this application, but the protection scope of this application is not limited thereto. Any person skilled in the art within the technical scope disclosed by this application can easily think of changes or substitutions, which should all be covered within the protection scope of this application. Therefore, the protection scope of this application should be subject to the protection scope of the claims.

[0154] Finally: The above description is only the preferred embodiment of the present invention and is not used to limit the present invention. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principle of the present invention should all be included within the protection scope of the present invention.

Claims

1. An enterprise-level intelligent customer service interaction control system in a multi-round dialogue scenario, characterized in that, Including: M1, a multi-modal semantic parsing module, deploying a hybrid neural network model composed of a Transformer encoder and a BiLSTM network; calculating scores of diverse text contents through the hybrid neural network model, and further calculating score confidence levels; outputting structured semantic information including the diverse text contents and corresponding score confidence levels; M2, a dynamic decision engine output module, which determines whether the diverse text contents are sensitive information according to the structured semantic information to select a retrieval path. If it is sensitive information, the retrieval path is the core database, otherwise the retrieval path is the extended knowledge base; retrieving from the core database or the extended knowledge base based on the retrieval path to generate a retrieval result; M3, a closed-loop learning optimization module, which captures the dialect recognition error, emotional response deviation value and term library missing warning signal of the hybrid neural network model in real time and inputs them into the incremental model; the incremental model is used to optimize the hybrid neural network model; synchronizing the optimized hybrid neural network model to each business terminal through a secure channel; the incremental model is established in the enterprise intranet data sandbox; M4, a multi-terminal dialogue security delivery module, generating and feeding back user response texts through the optimized neural network model, integrating enterprise-level gateways and private API communication protocols, and using national cryptography algorithms to encrypt channels to achieve full-link non-cloud data transmission.

2. The enterprise-level intelligent customer service interaction control system in a multi-round dialogue scenario according to claim 1, wherein The diverse text contents include natural language, dialect features and industry terms; Convert each word in the diversified text content input by the user into a word vector representation X through word embedding i ; X i represents the word vector of the i-th word; Subsequently, each of the word vectors X i is linearly transformed to extract a query vector A i , a key vector B i and a value vector C i ; The Transformer encoder optimizes the representation of each word vector by calculating the attention weights between the word vectors through the self-attention mechanism. The calculation formula of the attention weights is: Among them, represents the self-attention mechanism; d is the dimension of the key vector; represents the Softmax function; Through calculation, the Transformer encoder generates an optimized word vector sequence H, H = [h1, h2, …, h n , where h i is the optimized word vector of the i-th word in the text.

3. The enterprise-level intelligent customer service interaction control system in a multi-round dialogue scenario according to claim 2, wherein, The optimized word vector sequence H is passed to the BiLSTM network to further model the time series dependencies in the word vector sequence. The calculation formula is: where, g t-1 is the hidden state at the previous moment; g t is the hidden state at the current moment, representing the context information of the current optimized word vector; h t is the optimized word vector sequence at the t-th moment, that is, h t comes from H; represents the LSTM function; For the hidden state g output by the Transformer encoder t Generate scores corresponding to the natural language, dialect features, and industry terms through a linear transformation; The calculation formula of the score of the natural language is: z1 = W1g t + b1 where z1 represents the natural language score; W1 represents the weight matrix related to the natural language; b1 represents the bias term related to the natural language; The calculation formula of the score of the dialect features is: z2 = W2g t + b2 where z2 represents the dialect feature score; W2 represents the weight matrix related to the dialect features; b2 represents the bias term related to the dialect features; The calculation formula of the score of the industry terms is: z3 = W3g t + b3 where z3 represents the industry term score; W3 represents the weight matrix related to the industry terms; b3 represents the bias term related to the industry terms.

4. The enterprise-level intelligent customer service interaction control system in a multi-round dialogue scenario according to claim 3, characterized in that, Pass the scores corresponding to the natural language, dialect features and industry terms to the Softmax function to calculate the score confidence levels of each natural language, dialect feature and industry term; The calculation formula for the confidence level P1 of the natural language score is as follows: where is the normalization factor of all scores, ensuring that the sum of the final probabilities is 1; The calculation formula for the confidence level P2 of the dialect feature score is as follows: The calculation formula for the confidence level P3 of the industry term score is as follows:

5. The enterprise-level intelligent customer service interaction control system in a multi-round dialogue scenario according to claim 1, wherein Perform weighted evaluation on the confidence levels in the structured semantic information. The specific evaluation formula is extended to: Among them, is the confidence of the weighted natural language; is the confidence of the weighted dialect features; is the confidence of the weighted industry terms; C1, C2, and C3 are context adjustment factors; is the factor for normalizing all scores to ensure that the sum of the final probabilities is 1; Denote the initial structured semantic information as I; the weighted structured semantic information as wherein includes diverse text content, weighted natural language, dialect features, and industry term confidence levels; Determine whether the diversified text content is sensitive information based on the weighted confidence value to determine the retrieval path: First, set the sensitivity threshold to 0.8; when the weighted confidence value is greater than or equal to the preset sensitivity threshold, mark it as sensitive information and select the core database as the retrieval path; when the weighted confidence value is lower than the set sensitivity threshold, mark it as non-sensitive information and select the extended knowledge base as the retrieval path.

6. The enterprise-level intelligent customer service interaction control system in a multi-round dialogue scenario according to claim 1, characterized in that, The retrieval process of the core database is to input the sensitive information into the core database, encrypt the input sensitive information through an encryption retrieval function, and then perform retrieval in the database through indexing technology to obtain the retrieval result.

7. The enterprise-level intelligent customer service interaction control system in a multi-round dialogue scenario according to claim 1, characterized in that, The dialect recognition error refers to the error that occurs when the hybrid neural network model processes accents, vocabulary, or grammar in different regions. The calculation formula is: Among them, E dialect represents the dialect recognition error; represents the predicted value of the hybrid neural network system for the i-th dialect sample; y i represents the actual label value, that is, the correct dialect recognition result; n is the number of dialect samples; The emotional response deviation value represents the difference between the prediction result and the actual user emotion when the hybrid neural network model performs emotional analysis. The calculation formula is: Among them, E emotion represents the emotional response deviation value; represents the predicted emotional value of the hybrid neural network model for the i-th emotional sample; represents the actual emotional value; n is the number of emotional samples; The term library missing warning signal indicates that during the processing of the hybrid neural network system, a missing industry term is found. The calculation formula is: Among them, E term represents a term missing warning signal; Term i represents the i-th term; T represents the current term library; 1 is an indicator function, when the term Term i is not in the term library T, it returns 1, otherwise it returns 0; n is the number of terms.

8. The enterprise-level intelligent customer service interaction control system in a multi-round dialogue scenario according to claim 1, characterized in that, The training steps of the incremental model include training data preprocessing, incremental model initialization, incremental training loop, real-time signal monitoring, and security audit and rollback; The incremental model initialization uses the hybrid neural network as the basic architecture for incremental training, and sets the weight and bias parameters through the Xavier initialization method. The calculation formula is: where, w ij is a weight parameter, representing the weight from the i-th layer to the j-th layer; U(·,·) represents randomly extracting a value uniformly from an interval; n in is the number of neurons in the input layer; n out is the number of neurons in the output layer; The steps of the incremental training loop include forward propagation, loss calculation, backpropagation, and gradient calculation; The real-time signal monitoring evaluates the accuracy of the incremental model by monitoring the dialect recognition error, emotional response deviation value, and term library missing warning signal of the hybrid neural network model, and continues the incremental training loop when the accuracy is lower than the set first threshold; The security audit and rollback means that when the accuracy of the incremental model is lower than the second threshold, a rollback mechanism is started to restore the incremental model to the state before the previous incremental training loop.

9. The enterprise-level intelligent customer service interaction control system in a multi-round dialogue scenario according to claim 8, characterized in that, The forward propagation is that the training performs forward propagation through the incremental model to calculate the prediction result. The calculation formula of the forward propagation is: wherein, is the predicted output of the incremental model; X is the word vector; θ is the parameter of the incremental model; The loss calculation calculates the difference between the predicted value of the incremental model and the actual label through a loss function, and uses the cross-entropy loss function for calculation. The specific calculation formula is: L = -(αE dialect + βE emotion + γE term ) where L is the loss function; α, β, and γ are weight coefficients; E dialect represents the dialect recognition error; E emotion represents the emotional response deviation value; E term represents the term base missing warning signal; The backpropagation and gradient calculation is to calculate the gradient of the loss function with respect to the parameters of the incremental model through backpropagation, and update the parameters of the incremental model using gradient descent. The formula for the update process is: Among them, θ new represents the parameters of the updated incremental model; θ old represents the parameters of the incremental model before update; η is the learning rate, which controls the step size of parameter update; is the gradient of the loss function with respect to the parameters of the incremental model.

Citation Information

Patent Citations

  • Multi-round dialogue intelligent voice interaction system and device

    CN110209791A

  • Intelligent customer service method and system based on multiple rounds of dialogues

    CN116150338A

  • Database query generation method and system based on multi-selection optimization and storage medium

    CN119690987A

  • Speech interaction method and electronic device

    WO2023040658A1

Cited By

  • Data transmission method and device of cloud computer, electronic equipment and storage medium

    CN120979819A