Client intention recognition method, device and equipment and medium

By analyzing client operation records and relationship networks through large-scale model analysis, customer intent can be identified, solving the problem that insurance, fund, and medical institutions cannot accurately identify customer needs, thus reducing complaints and improving customer satisfaction.

CN121456818APending Publication Date: 2026-02-03CHINA PING AN PROPERTY INSURANCE CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511621564.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-11-06
Publication Date
2026-02-03

AI Technical Summary

Technical Problem

Insurance companies, fund managers, and medical institutions are unable to accurately identify customer intentions, leading to frequent contact with customers, which causes resentment and complaints.

Method used

By acquiring the client's operation records and relationship network, the multimodal feature fusion layer of the large model is used to perform text semantic understanding, relationship network analysis and temporal behavior modeling, outputting fused feature vectors, and using a dynamic intent prediction layer to determine the prediction probability distribution and execute the corresponding operation strategy.

Benefits of technology

Accurately identify customer intent, reduce complaints, improve customer satisfaction, and achieve precise marketing and service.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121456818A_ABST
    Figure CN121456818A_ABST
Patent Text Reader

Abstract

The invention relates to the field of finance and medical health, and discloses a customer intention recognition method, device, equipment and medium, and the method comprises the steps: carrying out text semantic understanding, relation network analysis and time sequence behavior modeling through employing a multi-modal feature fusion layer of a large model based on the operation record of a customer and a relation network used for describing the association relation of the customer; in this way, the speech, the social relation and the behavior of the customer can be comprehensively known. Based on the fusion feature vector, a dynamic intention prediction layer is utilized to accurately obtain prediction probability distribution and a corresponding intention recognition result, and a corresponding operation strategy is executed, so that complaints of customers can be reduced. The method comprises the steps of obtaining an operation record and a relational network of a client; performing conversion by using the multi-modal feature fusion layer to output a fusion feature vector; based on the fusion feature vector, a dynamic intention prediction layer is used for processing; and determining an intention recognition result, and executing a corresponding operation strategy.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the fields of finance and medical health, and in particular to a customer intention recognition method, device, equipment and medium. BACKGROUND

[0002] With the current growth of fuel vehicles gradually tending to saturation, the performance of many property insurance companies has gradually encountered a bottleneck. Therefore, it is urgent to find a new breakthrough to maintain the rapid growth of the company's performance, and various non-vehicle insurance products are particularly good breakthroughs. However, due to the current inability of insurance companies to understand the customer's intention, sometimes only a number is used to call the customer, which can easily cause the customer's aversion, thereby causing unnecessary complaints and the like. Similarly, in the financial field, fund managers cannot know whether the customer needs to buy funds, and regularly calling the customer can also cause the customer's aversion. For example, in the medical field, it is not possible to understand whether the customer needs to pay for medical services, and contacting the customer at will can also cause the customer's aversion. Therefore, how to accurately identify the customer's intention has become a technical problem to be solved by technical personnel in the field. SUMMARY

[0003] The present application provides a customer intention recognition method, device, equipment and medium to solve the technical problem of accurately understanding the customer's intention.

[0004] In a first aspect, a customer intention recognition method is provided, comprising: obtaining operation records of a customer on a client and a relationship network; wherein the operation records include input text and behavior records; the relationship network is composed of three nodes and edges for connecting the nodes; the three nodes include a customer node, a product node and a service institution node; converting the operation records and the relationship network using a multi-modal feature fusion layer in a large model to output a fusion feature vector; processing based on the fusion feature vector using a dynamic intention prediction layer in the large model to determine a prediction probability distribution; determining an intention recognition result based on the prediction probability distribution, and executing an operation strategy corresponding to the intention recognition result.

[0005] In a second aspect, a customer intention recognition device is provided, comprising: an obtaining unit configured to obtain operation records of a customer on a client and a relationship network; wherein the operation records include input text and behavior records; the relationship network is composed of three nodes and edges for connecting the nodes; the three nodes include a customer node, a product node and a service institution node; a conversion unit configured to convert the operation record and the relationship network using a multi-modal feature fusion layer in the large model to output a fusion feature vector; a processing unit configured to process, based on the fusion feature vector, using a dynamic intent prediction layer in the large model to determine a prediction probability distribution; an execution unit configured to determine an intent recognition result based on the prediction probability distribution, and execute an operation strategy corresponding to the intent recognition result.

[0006] In a third aspect, a computer device is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor, and the processor implements the steps of the customer intent recognition method when executing the computer program.

[0007] In a fourth aspect, a computer readable storage medium is provided, which stores a computer program, and the computer program implements the steps of the customer intent recognition method when executed by a processor.

[0008] In the scheme implemented by the above customer intent recognition method, device, equipment and medium, the operation record and the relationship network of the customer on the client can be obtained; the operation record includes input text and behavior record; the relationship network is composed of three nodes and edges for connecting the nodes; the three nodes include a customer node, a product node and a service institution node; the operation record and the relationship network are converted using a multi-modal feature fusion layer in the large model to output a fusion feature vector; based on the fusion feature vector, processing is performed using a dynamic intent prediction layer in the large model to determine a prediction probability distribution; based on the prediction probability distribution, an intent recognition result is determined, and an operation strategy corresponding to the intent recognition result is executed. In the present application, based on the operation record of the customer and the relationship network for describing the association relationship of the customer, text semantic understanding, relationship network analysis and time sequence behavior modeling are performed using the multi-modal feature fusion layer of the large model to obtain a fusion feature vector, so that the speech, social relationship and behavior of the customer can be comprehensively understood. Based on the fusion feature vector, the dynamic intent prediction layer can accurately obtain the prediction probability distribution and the corresponding intent recognition result, and based on the intent recognition result, the corresponding operation strategy can be executed, which can reduce the complaints of the customer, etc. BRIEF DESCRIPTION OF DRAWINGS

[0009] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the following will briefly introduce the drawings needed to be used in the description of the embodiments of the present application. Obviously, the drawings in the following description are only some embodiments of the present application, and other drawings can also be obtained by those skilled in the art without creative labor.

[0010] Figure 1 is an application environment schematic diagram of a customer intention recognition method in an embodiment of the present application; Figure 2 is a flow schematic diagram of a customer intention recognition method provided by an embodiment of the present application; Figure 3 exemplarily shows a structural schematic diagram of a customer intention recognition device provided according to some embodiments; Figure 4 is a structural schematic diagram of a computer device in an embodiment of the present application; Figure 5 is another structural schematic diagram of a computer device in an embodiment of the present application. DETAILED DESCRIPTION

[0011] The technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are some but not all of the embodiments of the present application. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative work fall within the protection scope of the present application.

[0012] With the current fuel vehicle growth gradually tending to saturation, the performance of many property insurance companies has gradually encountered a bottleneck. Therefore, it is urgent to find a new breakthrough to maintain the rapid growth of the company's performance, and various non-vehicle insurance products are particularly good ones. However, due to the current inability of insurance companies to understand the customer's intention, sometimes only a number is used to call the customer, which can easily cause the customer's aversion, thereby causing unnecessary complaints, etc. Similarly, in the financial field, fund managers cannot know whether the customer needs to buy funds, and regularly calling the customer can also cause the customer's aversion. For example, in the medical field, it is impossible to understand whether the customer needs to pay for medical services, and contacting the customer at will can also cause the customer's aversion. Therefore, how to accurately identify the customer's intention has become a technical problem to be solved by technical personnel in the field.

[0013] The customer intention recognition method provided by the embodiments of the present application can be applied to, for example, Figure 1In an application environment of the application, a customer operates at a client side of a client. A server obtains operation records of the customer at the client side and a relationship network; the operation records include input texts and behavior records; the relationship network is composed of three nodes and edges for connecting the nodes; the three nodes include a customer node, a product node and a service institution node; the operation records and the relationship network are converted by using a multi-modal feature fusion layer in a large model to output a fusion feature vector; based on the fusion feature vector, a dynamic intention prediction layer in the large model is used for processing to determine a prediction probability distribution; based on the prediction probability distribution, an intention recognition result is determined, and an operation strategy corresponding to the intention recognition result is executed. In the application, based on the operation records of the customer and the relationship network for describing the association relationship of the customer, the multi-modal feature fusion layer of the large model is used for text semantic understanding, relationship network analysis and time sequence behavior modeling to obtain the fusion feature vector, so that the speech, social relationship and behavior of the customer can be comprehensively understood. Based on the fusion feature vector, the dynamic intention prediction layer can accurately obtain the prediction probability distribution and the corresponding intention recognition result, and based on the intention recognition result, the corresponding operation strategy is executed, which can reduce the complaints of the customer and the like.

[0014] The client side can be, but is not limited to, various personal computers, notebook computers, smart phones, tablet computers and portable wearable devices. The server side can be implemented by an independent server or a server cluster composed of multiple servers. The application will be described in detail through specific embodiments.

[0015] Please refer to Figure 2 , as shown in the figure, Figure 2 A flowchart of a customer intention recognition method provided by an embodiment of the application includes the following steps S100-S400.

[0016] S100, obtaining operation records of a customer at a client side and a relationship network; the operation records include input texts and behavior records; the relationship network is composed of three nodes and edges for connecting the nodes; the three nodes include a customer node, a product node and a service institution node; It should be noted that the operation records of the customer at the client side in the embodiment of the application are obtained after the customer agrees. The operation records can be operation records of the customer in an insurance APP (application program) at the client side.

[0017] In the embodiment of the application, the operation records include input texts and behavior records. The input texts can include online customer service dialogues, search box keywords, product reviews, message feedbacks and the like.

[0018] The behavior record can include all behaviors of the customer in the past period of time, and exemplary can include all behaviors of the customer in the last 30 days. All behaviors in the behavior record are sorted by time. The behaviors can include clicking, scrolling a page, page staying, etc.

[0019] The relationship network in the embodiments of the application can be pre-established. The service agency corresponding to the service agency node can be a pet hospital, etc. The product corresponding to the product node can be insurance, fund, and paid medical service, etc.

[0020] S200, converting the operation record and the relationship network by using a multi-modal feature fusion layer in the large model to output a fusion feature vector.

[0021] In some embodiments, the step of converting the operation record and the relationship network by using a multi-modal feature fusion layer in the large model to output a fusion feature vector includes: extracting a corresponding text feature vector from a text input matrix corresponding to the input text by using a text feature extractor; In one example, the large model can analyze at most the first 128 words of the input text, so the dimension of the text input matrix corresponding to the input text is wherein = 128; = 96, that is, each word in the input text is converted into a 96-dimensional vector (word embedding).

[0022] In some embodiments, the text feature extractor can be a Transformer model (such as BERT), which can deeply understand the semantics and intention of the input text of the customer and extract a corresponding text feature vector.

[0023] determining a graph attention vector corresponding to the relationship network by using a graph attention network.

[0024] In the embodiments of the application, the graph attention network can analyze the strength of the relationship between different nodes. For example, the connection of the customer to a certain service agency (such as a certain pet hospital) is much closer than the connection to a certain ordinary advertisement. Specifically, the strength of the relationship between different nodes can be analyzed by using the following formula: ; wherein is the strength of the relationship between node and node ; is the feature vector of node ; is the feature vector of node .

[0025] In one example, the graph attention vector can be a 256-dimensional vector, representing the position of the customer in the relationship network he is in and the strength of the relationship with other nodes.

[0026] The time series encoder is used to determine a time series behavior feature vector from a time series behavior sequence corresponding to the behavior record.

[0027] In the embodiments of the present application, each behavior in the behavior record includes multi-dimensional features. In one example, the dimension of the time series behavior sequence is , wherein: = 30, i.e. the behavior in the last 30 days is analyzed; the behavior includes 5-dimensional features, specifically including timestamp, page type, stay seconds, scroll depth, and click count.

[0028] The time series encoder can use an LSTM network to learn and remember the behavior patterns of the customer. For example, it identifies that the customer browsed the pet insurance clause last week and started comparing prices this week, which is a gradually reinforced intention sequence. The time series encoder can output a 128-dimensional time series behavior feature vector, encoding the customer's historical behavior patterns and their evolution trends.

[0029] The fusion weight matrix and the first bias term are used to determine the fusion feature vector according to the text feature vector, the graph attention vector, and the time series behavior feature vector.

[0030] In the embodiments of the present application, the activation function introduces a nonlinear transformation for the large model, so that the large model can learn and represent the interaction between complex features. The activation function usually uses the ReLU function, which can make the large model learn the complex pattern that the intention strength is much greater than the sum of A (text question discount) and B (behavior comparison) when they appear at the same time.

[0031] In the embodiments of the present application, instead of simply concatenating the text feature vector, the graph attention vector, and the time series behavior feature vector, the fusion weight matrix is used to learn the information of different modalities, i.e. to weigh the text feature vector, the graph attention vector, and the time series behavior feature vector.

[0032] In one example, the dimension of the fusion weight matrix is The fusion weight matrix can learn a high weight for the behavior of "deeply reading the clause", such as 0.8, and a low weight for the behavior of "accidentally clicking on an advertisement", such as 0.1, through training.

[0033] The first bias term provides an offset on the basis of linear transformation, making the large model more flexible. In one example, the first bias term is a 384-dimensional vector.

[0034] In the embodiments of the present application, before steps S100-S400 are performed, the large model is trained and learned using a training set, and after training and learning, the fusion weight matrix and the first bias term can be determined. The large model after training is used to update the large model, and steps S100-S400 are performed.

[0035] In the embodiments of the present application, the fusion feature vector, which is equivalent to the customer panoramic portrait, is a 384-dimensional vector that condenses the customer's speech, social relations and behavior history.

[0036] In another way of description, the determination of the fusion feature vector using the activation function, the fusion weight matrix and the first bias term according to the text feature vector, the graph attention vector and the time series behavior feature vector can be represented by the following formula: ; wherein, is the fusion feature vector; is the activation function; is the fusion weight matrix; is the text input matrix corresponding to the input text; is the text feature vector extracted by the text feature extractor; is the relationship network; is the graph attention vector determined by the graph attention network; is the time series behavior sequence corresponding to the behavior record; is the time series behavior feature vector determined by the time series encoder; is the first bias term.

[0037] In the embodiments of the present application, the multi-modal feature fusion layer in the large model combines three core modules of text semantic understanding (TextBERT), relationship network analysis (GAT) and time series behavior modeling (TimeLSTM) with a learnable fusion weight matrix (W). In this way, the following can be achieved: comprehensive analysis of the customer: the large model is no longer one-sidedly analyzing the customer, but comprehensively listens to the input text of the customer, analyzes the relationship network and analyzes the behavior record, so as to make the most comprehensive judgment; and intelligent information weighting: the core is is not fixed, but is obtained through a large amount of data training and learning, which measures the value of different information, so as to ensure the accurate allocation of large model resources and make strong responses to high-value signals.

[0038] The multi-modal feature fusion layer in the embodiments of the present application can make the large model fully utilize the scattered multi-source heterogeneous data about the customer, i.e., operation records and relationship networks, to obtain a more accurate fusion feature vector.

[0039] S300, based on the fusion feature vector, processing is performed using the dynamic intent prediction layer in the large model to determine a predicted probability distribution.

[0040] In some embodiments, the step of based on the fusion feature vector, processing is performed using the dynamic intent prediction layer in the large model to determine a predicted probability distribution comprises: multiple times based on the fusion feature vector, processing is performed using the dynamic intent prediction layer to determine a predicted probability distribution. For example, based on the fusion feature vector, 100 times of prediction can be performed using the dynamic intent prediction layer to obtain 100 predicted probability distributions to be determined.

[0041] Based on the multiple predicted probability distributions obtained by multiple times of prediction, an average probability distribution is calculated, and the average probability distribution is determined as the predicted probability distribution.

[0042] For example, the predicted probability distribution indicates that the large model considers that the probability of the customer having an urgent purchase intention is 85%, the probability of the customer having a potential purchase intention is 12%, and the probability of the customer having no purchase intention is 3%.

[0043] In the embodiments of the present application, by determining the average probability distribution as the predicted probability distribution, the accuracy of prediction of the dynamic intent prediction layer can be improved.

[0044] In some embodiments, the step of based on the fusion feature vector, processing is performed using the dynamic intent prediction layer comprises: Based on the fusion feature vector and the time decay matrix, an intent signal sequence feature is output using a time-aware transformer encoding block.

[0045] In the embodiments of the present application, the time decay matrix is determined according to time and a decay coefficient. The time decay matrix can give higher weight to recent behavior and reduce the weight of long-term behavior, simulating the memory law of human beings. This is because the customer's behavior of comparing product prices yesterday is more valuable than an accidental click one month ago.

[0046] In some embodiments, the time decay matrix can be calculated according to the following formula: ; wherein, is the time decay matrix; is the decay coefficient; is the time, is the time interval (in days) from the time when the behavior occurs to the present time. For example, can be 30 days, is 0.05. Diag is used to convert A diagonal matrix is constructed, which is the time decay matrix.

[0047] In the embodiments of the present application, the time-aware transformer encoding block receives the fusion feature vector and is modulated by the time decay matrix to perform nonlinear transformation, and finally outputs the intent signal sequence feature. The time-aware transformer encoding block internally includes a self-attention mechanism, which can discover the internal relationship between behaviors, such as identifying that "viewing terms" is followed by "using insurance calculator".

[0048] Based on the intent signal sequence feature, the intent category weight matrix, and the second bias term, a normalized exponential function is used to determine the to-be-determined probability distribution.

[0049] In the embodiments of the present application, the intent category weight matrix is a linear transformation matrix that can map the intent signal sequence feature to the intent category. The dimension of the intent category weight matrix is , wherein is the dimension of the intent signal sequence feature output by the time-aware transformer encoding block, and 3 represents three intent categories, namely, emergency purchase intent, potential purchase intent, and no purchase intent.

[0050] The second bias term works with the intent classification weight matrix to adjust the boundary of the decision, so that the large model output is more flexible. The second bias term is a three-dimensional vector, and each dimension corresponds to a bias of an intent category.

[0051] The normalized exponential function (softmax) converts the output of the intent signal sequence feature, the intent category weight matrix, and the second bias term into the to-be-determined probability distribution, so that the sum of the probabilities of the three intent categories is 1, and each probability is between [0, 1], which intuitively reflects the possibility of each intent category.

[0052] In another way of description, the determination of the to-be-determined probability distribution based on the intent signal sequence feature, the intent category weight matrix, and the second bias term using the normalized exponential function can be represented by the following formula: ; , wherein y is the to-be-determined probability distribution; is the intent category weight matrix, is the fusion feature vector, is the second bias term; is the time decay matrix.

[0053] In some embodiments, the method further comprises: based on a plurality of the to-be-determined probability distributions and the predicted probability distribution, calculating a predicted uncertainty value of the large model using a dynamic intent prediction layer in the large model.

[0054] In this embodiment, the uncertainty quantification technology is used to determine the prediction uncertainty value of the large model.

[0055] Specifically, the prediction uncertainty value of the large model is calculated based on the plurality of to-be-determined probability distributions and the prediction probability distribution by using the dynamic intention prediction layer in the large model, which can be calculated by the following formula: ; wherein, is the prediction uncertainty value; represents the total prediction times; is the prediction probability distribution of the zth prediction; is the average value of the prediction probability distribution of the zth prediction.

[0056] In this embodiment, the prediction uncertainty value can be calculated by the Monte Carlo Dropout (MC-Dropout) technology. When reasoning, the Dropout is randomly disabled 100 times (n = 100), 100 prediction probability distributions are obtained, and then the standard deviation of the 100 prediction probability distributions is calculated. The standard deviation is determined as the prediction uncertainty value.

[0057] If the prediction uncertainty value is greater than the preset prediction uncertainty value, an alarm is sent to enable a business personnel to manually adjust the intention recognition result.

[0058] In this embodiment, if the prediction uncertainty value is greater than the preset prediction uncertainty value, it indicates that the accuracy of the prediction probability distribution output by the large model cannot be grasped, and manual review is triggered at this time, thereby avoiding the large model outputting an incorrect prediction probability distribution.

[0059] If the prediction uncertainty value is not greater than the preset prediction uncertainty value, step S400 is continuously executed.

[0060] In this embodiment, if the prediction uncertainty value is not greater than the preset prediction uncertainty value, it indicates that the large model considers that the accuracy of the prediction probability distribution output is high, and at this time, the consciousness recognition result determined according to the prediction probability distribution can be used to execute the corresponding operation strategy.

[0061] S400, based on the prediction probability distribution, determining an intention recognition result, and executing an operation strategy corresponding to the intention recognition result.

[0062] In the embodiment of the present application, the intention category with the highest probability in the prediction probability distribution is determined as the intention recognition result. In one example, the prediction probability distribution ​​At this time, it is determined that the intent category of the customer with an 85% probability of an emergency purchase intention is the intent recognition result.

[0063] In some embodiments, if the intent recognition result is an emergency purchase intention, the corresponding operation strategy is to automatically assign an agent to contact the customer. If the intent recognition result is a potential purchase intention, the corresponding operation strategy is to push precise marketing content. If the intent recognition result is no purchase intention, the corresponding operation strategy is to monitor the customer in a dormant state, and no longer obtain the operation records of the customer on the client and the relationship network, and no longer identify the potential purchase intention of the customer for non-vehicle insurance. For example, the preset time length is 30 days.

[0064] In the embodiments of the present application, the dynamic intent prediction layer can achieve accurate intent recognition by using the time-aware transformer encoding block and the uncertainty quantification technology. The large model not only considers the behavior of the customer, but also finely considers the time of behavior execution, making the judgment more consistent with the real logic. Safe intelligent decision-making can also be achieved. The large model is no longer a "black box", and it has self-evaluation capability (predicting uncertainty value). When the large model believes that the output result may not be accurate, it will actively trigger manual review, thereby combining automation efficiency with human judgment, and thus reducing complaints and improving customer satisfaction.

[0065] In some embodiments, the large model further comprises an online learning layer; and the method further comprises: receiving feedback information based on the operation strategy; For example, the operation strategy is to automatically assign an agent to contact the customer, and the corresponding feedback information includes agent follow-up results, whether the customer has made a transaction, and whether the customer has complained, etc. The operation strategy is to push precise marketing content, and the corresponding feedback information can also include whether the user has made a transaction and whether the customer has complained, etc.

[0066] Based on the intent recognition result and the feedback information, the model parameters in the large model are updated using the online learning layer.

[0067] In some embodiments, the step of updating the model parameters in the large model using the online learning layer based on the intent recognition result and the feedback information comprises: Based on the intent recognition result and the feedback information, the model parameters in the large model are updated using a focus loss function and an elastic weight constraint.

[0068] In the embodiments, the model parameters are parameters involved in the large model, including parameters involved in the text feature extractor and parameters involved in the graph attention network. The model parameters can include a vector of millions of weights.

[0069] In the embodiments of the present application, a powerful and suitable continuous learning framework for insurance business is formed by combining a focal loss function (FocalLoss) and an elastic weight consolidation constraint.

[0070] The elastic weight consolidation constraint punishes new parameters from deviating from the historical optimal value , thereby preventing the large model from forgetting previously learned important old knowledge (such as “cancer” being a strong correlation word for medical insurance) when learning new samples (such as new insurance data).

[0071] The focal loss function can enable the large model to continuously update the model parameters in the large model from the intent recognition result and the feedback information, and learn from failures and difficult cases.

[0072] In some embodiments, the model parameters in the large model are updated based on the intent recognition result and the feedback information using the focal loss function and the elastic weight consolidation constraint, which is completed using the following formula: ; wherein, is the model parameter of the a+1th step, i.e., the model parameter in the large model after updating the model parameter of the ath step; is the model parameter of the ath step; represents the focal loss function term; is the elastic weight consolidation constraint term; is the adaptive learning rate of the ath step; the a+1th step and the ath step both refer to the time step; is the gradient operator; is the batch size, i.e., the number of samples; is the sample weight; is the focal loss function value, wherein is the true label of the nth sample, is the prediction probability distribution of the nth sample; is the elastic weight consolidation strength; is a set of key parameters, which includes a plurality of key parameters in the model parameter; is the optimal value of the ith key parameter; is the historical optimal value of the ith key parameter.

[0073] The elastic weight consolidation constraint term is an L2 regularization term acting on a specific set of key parameters . This term can protect the historical optimal value of the key parameter, avoid the large model forgetting the verified key parameter about the customer intent while learning the sample, and ensure the stability of the large model performance.

[0074] The adaptive learning rate is the learning speed of the large model, which is not fixed but dynamically adjusted according to the current performance of the large model. For example, if the complaint rate increases recently, it means that the large model may have learned bias, and will automatically be reduced to make the update more cautious and smaller. The adaptive learning rate can be determined by the AdaGrad optimizer.

[0075] The gradient operator indicates the direction of the update of the parameters, and the partial derivative of the loss function with respect to each parameter is calculated to help fine-tune each parameter so that the large model can perform better overall. The gradient operator is automatically calculated by the backpropagation algorithm.

[0076] The batch size refers to the number of samples used to calculate the gradient operator in one iteration, which can be set to 256 or 512 samples, for example.

[0077] Different sample weights will be assigned to different samples, so that the large model can pay more attention to important samples whose feedback information is not very good. For example, customers who have an urgent purchase intention but ultimately do not complete the transaction, the sample weight should be set to a higher number. In addition, the sample weight of the customer who complains should also be set to a higher value, so that the large model can quickly correct the error and avoid further complaints.

[0078] The focal loss value is calculated by focusing on the loss function (FocalLoss). Here is a positive adjustment factor, in one example, =2. The focal loss function automatically reduces the contribution of simple samples by term, so that the large model can focus on learning difficult examples. Difficult examples can be understood as samples of customers who have consulted multiple times, have complex behavior (both looking at insurance products and searching for cancellation), and are difficult to judge their true intentions.

[0079] The larger the elastic weight consolidation strength, the stronger the protection of old knowledge by the large model, and the less likely it is to forget. This value needs to be carefully adjusted to balance between learning new knowledge and remembering old knowledge. Through cross-validation, it may be a small value, such as =0.01.

[0080] is the index set of key parameters, which are extremely important and cannot be forgotten in the large model. These parameters usually correspond to core and general business knowledge. The importance score of the parameter (such as the Fischer information matrix) is calculated to determine the key parameter.

[0081] is the optimal value of the i-th key parameter; is the historical optimal value of the i-th key parameter; is the parameter is the optimal value proved on the past task (i.e., “old knowledge”). An exemplary may represent the importance weight of the keyword “cancer” in medical insurance intent recognition. may also represent the weight of the behavior “price comparison” in prediction.

[0082] An exemplary ; wherein, is the focus loss function term, which is learned through the uncompleted transaction samples of Mr. X; is the elastic weight consolidation constraint term, for example, to protect the weight of the “medical insurance” keyword. 15.3 is the sample weight. is learned for the difficult cases that Mr. X consulted multiple times but did not purchase insurance. 0.01 is the elastic weight consolidation strength.

[0083] In one example, the large model is used to predict the intent recognition result at different times, and the corresponding operation strategy is executed, as shown in the following table:

[0084] As can be seen, the large model in the embodiments of the present application can execute different operation strategies according to the behaviors of the customer at different times.

[0085] In the embodiments of the present application, the consciousness recognition result predicted by the large model can enable the agent to better serve the customer, meet the customer's related needs, and avoid customer complaints and other problems.

[0086] As can be seen, in the above scheme, based on the operation record of the customer and the relationship network for describing the association relationship of the customer, the multi-modal feature fusion layer of the large model is used for text semantic understanding, relationship network analysis and time series behavior modeling to obtain a fusion feature vector. Based on the fusion feature vector, a dynamic intent prediction layer can accurately obtain a prediction probability distribution and a corresponding intent recognition result, and based on the intent recognition result, a corresponding operation strategy is executed, which can reduce customer complaints and the like.

[0087] It should be understood that the sequence number of each step in the above embodiments does not imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present invention.

[0088] In one embodiment, a customer intent recognition device is provided, which corresponds one-to-one with a customer intent recognition method described in the above embodiments. For example... Figure 3 As shown, the customer intent recognition device includes an acquisition unit 301, a conversion unit 302, a processing unit 303, and an execution unit 304. Detailed descriptions of each functional module are as follows: The acquisition unit 301 is used to acquire the customer's operation records and relationship network on the client; wherein, the operation records include input text and behavior records; the relationship network consists of three nodes and edges for connecting the nodes; the three nodes include a customer node, a product node, and a service organization node; The conversion unit 302 is used to convert the operation record and the relationship network using the multimodal feature fusion layer in the large model to output a fused feature vector. Processing unit 303 is used to process the fused feature vector using the dynamic intent prediction layer in the large model to determine the prediction probability distribution; The execution unit 304 is used to determine the intent recognition result based on the predicted probability distribution and execute the operation strategy corresponding to the intent recognition result.

[0089] In some embodiments, the conversion unit is specifically used for: Using a text feature extractor, the corresponding text feature vector is extracted from the text input matrix corresponding to the input text; Using a graph attention network, determine the graph attention vector corresponding to the relationship network; Using a time series encoder, a time series behavior feature vector is determined from the time series behavior sequence corresponding to the behavior record; Using the activation function, the fusion weight matrix, and the first bias term, the fused feature vector is determined based on the text feature vector, the graph attention vector, and the temporal behavior feature vector.

[0090] In some embodiments, the processing unit includes: The prediction unit is used to make multiple predictions based on the fused feature vector using the dynamic intent prediction layer. The calculation unit is used to calculate the average probability distribution based on the multiple probability distributions to be determined obtained from multiple predictions; and to determine the average probability distribution as the predicted probability distribution.

[0091] In some embodiments, the prediction unit is specifically configured to: based on the fusion feature vector and a time decay matrix, outputting an intent signal sequence feature by using a time-aware transformer encoding block; based on the intent signal sequence feature, an intent category weight matrix and a second bias term, determining a to-be-determined probability distribution by using a normalized exponential function; wherein the time decay matrix is determined according to time and a decay coefficient.

[0092] In some embodiments, the device is further configured to: based on a plurality of the to-be-determined probability distributions and the prediction probability distribution, calculating a prediction uncertainty value of the large model by using a dynamic intent prediction layer in the large model; if the prediction uncertainty value is greater than a preset prediction uncertainty value, sending an alarm to enable a business staff to manually adjust the intent recognition result.

[0093] In some embodiments, the large model further comprises an online learning layer; and the device further comprises: a receiving unit configured to receive feedback information based on the operation strategy; an updating unit configured to update model parameters in the large model by using the online learning layer based on the intent recognition result and the feedback information.

[0094] In some embodiments, the updating unit is specifically configured to: update the model parameters in the large model by using a focal loss function and an elastic weight consolidation constraint based on the intent recognition result and the feedback information.

[0095] The present application provides a customer intent recognition device, based on operation records of customers and a relationship network for describing customer association relationships, text semantic understanding, relationship network analysis and time sequence behavior modeling are performed by using a multi-modal feature fusion layer of a large model to obtain a fusion feature vector by fusion, so that the speech, social relationship and behavior of the customer can be comprehensively understood. Based on the fusion feature vector, a prediction probability distribution and a corresponding intent recognition result can be accurately obtained by using a dynamic intent prediction layer, and corresponding operation strategies can be executed based on the intent recognition result, which can reduce customer complaints and the like.

[0096] The specific limitations of the customer intent recognition device can be referred to the limitations of the customer intent recognition method in the above, which will not be repeated here. Each module in the above customer intent recognition device can be realized by software, hardware and combinations thereof, in whole or in part. The above modules can be embedded in or independent of the processor in the computer device in hardware form, or can be stored in the memory in the computer device in software form, so as to be called and executed by the processor to perform the operations corresponding to each module.

[0097] In an embodiment, a computer device is provided, which can be a server, and an internal structure diagram thereof can be as shown in Figure 4 The computer device includes a processor, a memory, a network interface and a database connected through a system bus. The processor of the computer device is configured to provide computing and control capabilities. The memory of the computer device includes a non-volatile and / or volatile storage medium, an internal memory. The non-volatile storage medium stores an operating system, a computer program and a database. The internal memory provides an environment for the operating system and the computer program in the non-volatile storage medium to run. The network interface of the computer device is configured to communicate with an external client through a network connection. The computer program, when executed by the processor, implements the functions or steps of the server side of a client intent recognition method.

[0098] In an embodiment, a computer device is provided, which can be a client, and an internal structure diagram thereof can be as shown in Figure 5 The computer device includes a processor, a memory, a network interface, a display screen and an input device connected through a system bus. The processor of the computer device is configured to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operating system and the computer program in the non-volatile storage medium to run. The network interface of the computer device is configured to communicate with an external server through a network connection. The computer program, when executed by the processor, implements the functions or steps of the client side of a client intent recognition method.

[0099] In an embodiment, a computer device is provided, which includes a memory, a processor and a computer program stored in the memory and executable on the processor, and the processor implements the following steps when executing the computer program: obtaining operation records of a client on a client side and a relationship network; wherein the operation records include input text and behavior records; the relationship network is composed of three nodes and edges connecting the nodes; the three nodes include a client node, a product node and a service institution node; converting the operation records and the relationship network using a multi-modal feature fusion layer in a large model to output a fusion feature vector; based on the fusion feature vector, processing using a dynamic intent prediction layer in the large model to determine a prediction probability distribution; based on the prediction probability distribution, determining an intent recognition result, and executing an operation strategy corresponding to the intent recognition result.

[0100] In one embodiment, a computer readable storage medium is provided, having stored thereon a computer program which, when executed by a processor, implements the following steps: obtaining operation records of a customer on a client and a relationship network; wherein the operation records include input text and behavior records; the relationship network is composed of three nodes and edges for connecting the nodes; the three nodes include a customer node, a product node and a service institution node; converting the operation records and the relationship network by using a multi-modal feature fusion layer in a large model to output a fusion feature vector; based on the fusion feature vector, processing by using a dynamic intention prediction layer in the large model to determine a prediction probability distribution; based on the prediction probability distribution, determining an intention recognition result, and executing an operation strategy corresponding to the intention recognition result.

[0101] It should be noted that the functions or steps that the above computer readable storage medium or computer device can implement can correspond to the related descriptions of the server side and the client side in the foregoing method embodiments. To avoid repetition, they will not be described one by one here.

[0102] Those of ordinary skill in the art can understand that all or part of the processes in the above-mentioned embodiment methods can be completed by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer readable storage medium. When the computer program is executed, it can include the processes of the above-mentioned embodiments. Any reference to memory, storage, database or other medium used in the embodiments provided by the present application can include non-volatile and / or volatile memory. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM) or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. As an illustration but not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchronous link (Synchlink) DRAM (SLDRAM), memory bus (Rambus) direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM) and memory bus dynamic RAM (RDRAM) and the like.

[0103] Those skilled in the art can clearly understand that, for the convenience and brevity of description, only the above-mentioned division of each functional unit and module is exemplified, and in actual application, the above-mentioned functions can be completed by different functional units and modules according to needs, that is, the internal structure of the device is divided into different functional units or modules to complete all or part of the functions described above.

[0104] The non-company software tools or components appearing in the embodiments of the present application are only exemplarily introduced, and do not represent actual use.

[0105] The above-described embodiments are only used to illustrate the technical solutions of the present application, rather than limit them; although the present application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that the technical solutions recorded in the foregoing embodiments can still be modified, or some technical features can be replaced by equivalents; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present application, and should be included in the protection scope of the present application.

Claims

1. A method for recognizing customer intent, characterized in that, include: The system acquires customer operation records and relationship networks on the client side; wherein the operation records include input text and behavior records; the relationship network consists of three nodes and edges connecting the nodes; the three nodes include a customer node, a product node, and a service organization node; The operation records and the relationship network are transformed using a multimodal feature fusion layer in a large model to output a fused feature vector. Based on the fused feature vector, the dynamic intent prediction layer in the large model is used for processing to determine the prediction probability distribution; Based on the predicted probability distribution, the intent recognition result is determined, and the operation strategy corresponding to the intent recognition result is executed.

2. The method according to claim 1, characterized in that, The step of transforming the operation record and the relationship network using a multimodal feature fusion layer in a large model to output a fused feature vector includes: Using a text feature extractor, the corresponding text feature vector is extracted from the text input matrix corresponding to the input text; Using a graph attention network, determine the graph attention vector corresponding to the relationship network; Using a time series encoder, a time series behavior feature vector is determined from the time series behavior sequence corresponding to the behavior record; Using the activation function, the fusion weight matrix, and the first bias term, the fused feature vector is determined based on the text feature vector, the graph attention vector, and the temporal behavior feature vector.

3. The method according to claim 1, characterized in that, The step of determining the prediction probability distribution based on the fused feature vector using the dynamic intent prediction layer in the large model includes: Multiple predictions are made using the dynamic intent prediction layer based on the fused feature vector; Based on the multiple probability distributions to be determined obtained from multiple predictions, the average probability distribution is calculated; and the average probability distribution is determined as the predicted probability distribution.

4. The method according to claim 3, characterized in that, The step of making predictions based on the fused feature vector using the dynamic intent prediction layer includes: Based on the fused feature vector and the time decay matrix, the intention signal sequence features are output using a time-aware transformer coding block. Based on the intent signal sequence features, intent category weight matrix, and second bias term, the probability distribution to be determined is determined using a normalized exponential function; wherein the time decay matrix is ​​determined based on time and decay coefficient.

5. The method according to claim 3, characterized in that, Also includes: Based on multiple probability distributions to be determined and the predicted probability distributions, the prediction uncertainty value of the large model is calculated using the dynamic intent prediction layer in the large model; If the predicted uncertainty value is greater than the preset predicted uncertainty value, an alarm is sent so that business personnel can manually adjust the intent recognition result.

6. The method according to claim 2, characterized in that, The large model also includes an online learning layer; the method further includes: Receive feedback information based on the operation strategy; Based on the intent recognition results and the feedback information, the model parameters in the large model are updated using an online learning layer.

7. The method according to claim 6, characterized in that, The step of updating the model parameters in the large model using an online learning layer based on the intent recognition result and the feedback information includes: Based on the intent recognition results and the feedback information, the model parameters in the large model are updated by using the focus loss function and elastic weights to consolidate the constraints.

8. A customer intent recognition device, characterized in that, include: The acquisition unit is used to acquire the customer's operation records and relationship network on the client; wherein, the operation records include input text and behavior records; the relationship network consists of three nodes and edges for connecting the nodes; the three nodes include a customer node, a product node, and a service organization node; The transformation unit is used to transform the operation record and the relationship network using the multimodal feature fusion layer in the large model to output a fused feature vector; The processing unit is used to process the fused feature vector using the dynamic intent prediction layer in the large model to determine the prediction probability distribution. An execution unit is used to determine the intent recognition result based on the predicted probability distribution and to execute the operation strategy corresponding to the intent recognition result.

9. A computer device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the steps of the customer intent recognition method as described in any one of claims 1 to 7.

10. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by a processor, it implements the steps of the customer intent recognition method as described in any one of claims 1 to 7.