Text recommendation method and device based on multi-head attention mechanism, equipment and medium
By adopting a text recommendation method based on multi-head attention mechanism in the speech navigation system, the problem of low accuracy in existing systems when dealing with the needs of fuzzy or implicit expression is solved, and higher recommended text accuracy and semantic comprehension capabilities are achieved.
Patent Information
- Application Number
- CN202510272071.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-07
- Publication Date
- 2025-06-06
AI Technical Summary
When existing speech navigation systems deal with customer vague or implicit expression needs, they cannot accurately identify user intentions and cannot capture dynamic changes in customer communication in real time, resulting in low accuracy of speech recommendations.
The text recommendation method based on the multi-head attention mechanism is adopted, and the text to be retrieved is processed and vectorized, and combined with text features and word participle position information is used to identify the dialogue intention, and the recommended text is generated.
The accuracy of the recommendation text is improved, and the accuracy and semantic understanding of the recommendation system are enhanced through clear semantic units and structured information capture, ensuring the generation of recommendation texts that are highly relevant to user needs.
Smart Images

Figure CN120104756A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of natural language processing technology, and in particular to a text recommendation method, device, equipment and medium based on a multi-head attention mechanism. Background Art
[0002] In the property and casualty insurance industry, sales personnel mainly communicate with customers through telephone or online platforms. They provide customers with professional insurance advice and recommend suitable insurance products through professional sales techniques.
[0003] At present, the telemarketing platform provides sales personnel with more professional marketing scripts through the script navigation function during the insurance product sales process, helping them improve the service quality and professionalism of sales services. By integrating the insurance knowledge base and application scenarios through AI technology, the script navigation function can provide sales personnel with script prompts and thought guidance in real time, helping them communicate with customers more efficiently.
[0004] However, when processing customer language, existing speech navigation systems cannot accurately identify user intentions if the customer expresses their needs in an ambiguous or implicit manner. They are also unable to capture the dynamic changes of customers during communication in real time and cannot adapt to the dynamic needs of customers, resulting in low accuracy of speech recommendations. Summary of the invention
[0005] The present invention provides a text recommendation method, device, computer equipment and medium based on a multi-head attention mechanism to solve the technical problem that the key feature recognition of the speech navigation system is inaccurate, resulting in low accuracy of speech recommendation.
[0006] First, a text recommendation method based on a multi-head attention mechanism is provided, including:
[0007] Get the first text to be retrieved;
[0008] Based on a text processing algorithm, segment the first text to obtain a first segmentation set, wherein the first segmentation set includes at least one first text segmentation;
[0009] Based on the text features of the first text and the position information of the first text segmentation in the first text, obtaining a first text vector corresponding to the first text segmentation;
[0010] Based on the multi-head attention matrix, feature recognition is performed on the first text vector to obtain target features;
[0011] Based on the target feature, the conversation intention corresponding to the first text is identified, and a recommended text is generated.
[0012] In a second aspect, a text recommendation device based on a multi-head attention mechanism is provided, comprising:
[0013] A first text acquisition module, used to acquire a first text to be retrieved;
[0014] A text segmentation processing module, configured to perform segmentation processing on the first text based on a text processing algorithm to obtain a first segmentation set, wherein the first segmentation set includes at least one first text segmentation;
[0015] A first text vector obtaining module, configured to obtain a first text vector corresponding to the first text segmentation based on text features of the first text and position information of the first text segmentation in the first text;
[0016] A feature extraction module, used to perform feature recognition on the first text vector based on a multi-head attention matrix to obtain target features;
[0017] The recommended text generation module is used to identify the conversation intention corresponding to the first text based on the target feature and generate a recommended text.
[0018] In a third aspect, a computer device is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the steps of the above-mentioned text recommendation method based on the multi-head attention mechanism when executing the computer program.
[0019] In a fourth aspect, a computer-readable storage medium is provided, which stores a computer program. When the computer program is executed by a processor, the steps of the above-mentioned text recommendation method based on the multi-head attention mechanism are implemented.
[0020] In the scheme implemented by the above-mentioned text recommendation method, device, computer equipment and storage medium based on the multi-head attention mechanism, by performing word segmentation processing on the first text, clear semantic units are provided for subsequent feature extraction and intent recognition, thereby improving the accuracy of the recommended text. By combining text features and the position information of word segments in the text, the generated vector can better capture the structured information of the text and more comprehensively reflect the semantic role and importance of each word segment in the context, thereby improving the accuracy and semantic understanding ability of the recommendation system. By extracting features from text vectors from multiple angles through a multi-attention matrix, features related to conversation intent can be more accurately identified, the accuracy of key feature recognition of the first text can be improved, and the conversation intent of the user input text can be accurately identified. Through clear intent recognition, recommended text that is highly relevant to user needs is generated, which improves the accuracy of the recommended text. BRIEF DESCRIPTION OF THE DRAWINGS
[0021] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings required for use in the description of the embodiments of the present invention will be briefly introduced below. Obviously, the accompanying drawings in the following description are only some embodiments of the present invention. For ordinary technicians in this field, other accompanying drawings can be obtained based on these accompanying drawings without paying creative labor.
[0022] Figure 1 Schematic diagram of an application environment of a text recommendation method based on a multi-head attention mechanism in one embodiment of the present invention;
[0023] Figure 2 A flowchart of a first embodiment of a text recommendation method based on a multi-head attention mechanism provided by an embodiment of the present invention;
[0024] Figure 3 A flowchart of a second embodiment of a text recommendation method based on a multi-head attention mechanism provided in an embodiment of the present invention;
[0025] Figure 4 is a structural schematic diagram of a text recommendation device based on a multi-head attention mechanism in one embodiment of the present invention;
[0026] Figure 5 is a schematic diagram of a structure of a computer device in one embodiment of the present invention;
[0027] Figure 6 It is another structural schematic diagram of a computer device in one embodiment of the present invention. DETAILED DESCRIPTION
[0028] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0029] The text recommendation method based on the multi-head attention mechanism provided by the embodiment of the present invention can be applied to Figure 1In an application environment, the client communicates with the server through a network. The server can receive voice data of communication between a business person and a user through the client, convert the voice data into text data, obtain a first text to be retrieved, perform word segmentation processing on the first text based on a text processing algorithm, and obtain a first word segmentation set, wherein the first word segmentation set includes at least one first text word segmentation; based on the text features of the first text and the position information of the first text word segmentation in the first text, obtain a first text vector corresponding to the first text word segmentation; based on a multi-head attention matrix, perform feature recognition on the first text vector to obtain a target feature; based on the target feature, identify the conversation intention corresponding to the first text, and generate a recommended text. In the present invention, for the low accuracy of speech recommendation in medical or financial dialogue service scenarios (including business consulting, sales and other service scenarios), the first text can be processed by word segmentation to provide a clear semantic unit for subsequent feature extraction and intention recognition, thereby improving the accuracy of the recommended text. By combining text features and the position information of the word segmentation in the text, the generated vector can better capture the structural information of the text and more comprehensively reflect the semantic role and importance of each word segmentation in the context, thereby improving the accuracy and semantic understanding ability of the recommendation system. By extracting features from text vectors from multiple angles through a multi-attention matrix, features related to the conversation intent can be more accurately identified, the accuracy of key feature recognition of the first text can be improved, and the conversation intent of the user input text can be accurately identified. Through clear intent recognition, recommended texts that are highly relevant to user needs are generated, which improves the accuracy of recommended texts. Among them, the client can be but is not limited to various personal computers, laptops, smart phones, tablet computers and portable wearable devices. The server can be implemented with an independent server or a server cluster consisting of multiple servers. The present invention is described in detail below through specific embodiments.
[0030] See also Figure 2 As shown, Figure 2 The flowchart of the first embodiment of the text recommendation method based on the multi-head attention mechanism provided by the embodiment of the present invention includes the following steps:
[0031] S101: Obtaining a first text to be searched;
[0032] In one embodiment, the first text may be text data obtained by converting voice data of a call operator (including a business person or an artificial intelligence) communicating with a customer. Specifically, the voice data may be converted into text information, and all voice data of each voice call may be converted into a first text, or each voice during each voice call may be converted into the first text.
[0033] For example, during a voice call, all the voice data of each complete voice call from the beginning to the end can be integrated and converted into a complete first text to comprehensively present the whole picture of the call; or more specifically, every time the operator or the customer outputs a piece of voice, it can be immediately converted into an independent first text.
[0034] In one embodiment, in addition to the original telephone voice data conversion to text information, the first text can also cover the online communication text data between the business personnel and the customers. Specifically, the online communication text data can include: text messages generated when the business personnel communicate with the customers through instant messaging tools, dialogue texts between the business personnel and the customers on the online customer service platform, interactions between the business personnel and the customers on the social media platform (including comment replies, private message exchanges, etc.), and email communication texts between the business personnel and the customers, etc.
[0035] Among them, the online communication text data can be text input text, semantic input text, or picture input text.
[0036] S102: Based on the text processing algorithm, perform word segmentation processing on the first text to obtain a first word segmentation set, where the first word segmentation set includes at least one first text word segmentation;
[0037] In one embodiment, the text processing algorithm can include processing processes such as data preprocessing, text word segmentation, and text vectorization of the first text.
[0038] Text preprocessing is the first step of text processing. The purpose is to clean, convert, and standardize the original unstructured text data into a format suitable for the input of the machine learning model, so as to improve the model performance and reduce the processing difficulty.
[0039] Specifically, it can include text standardization, such as converting all texts into a unified format to reduce the diversity of vocabulary; deleting common words that contribute little to the meaning of the text, such as "de", "shi", "zai", etc. These words appear frequently in most texts but rarely carry important semantic information; simplifying words to their basic forms (stemming), or restoring inflected words to their original forms (lemmatization), further reducing the complexity of vocabulary.
[0040] In one embodiment, word segmentation algorithms based on rules, statistics, or deep learning, etc. can be used to split the first text, split the first text into multiple word segmentations, and obtain the first word segmentation set. Each first text corresponds to a first word segmentation set, and each first word segmentation set includes one or more word segmentations. For example, for the first text: "I like to eat apples", the first text word segmentations can include: "I, like, eat, apples".
[0041] In one embodiment, the text processing of the first text may further include vectorization of first text word segments, where each first text word segment is vectorized to obtain a word segmentation vector of each first text word segment as a text feature of the first text.
[0042] Generally, the word segmentation vectorization of text can be implemented in a variety of ways, such as bag of words model, One-Hot encoding and word embedding. The embodiment of the present application does not specifically limit this, and only illustrates this process exemplarily. For example, word embedding models such as Word2Vec, GloVe, and FastText can be used to map the first text segmentation to a continuous vector space to capture the semantic and grammatical relationship between the first text segmentations. Specifically, taking Word2Vec as an example, it includes two modes: CBOW and skip-gram. The CBOW algorithm uses the words before and after the central word to calculate the word vector of the central word. First, the words before and after are one-hot encoded, and then the hidden layer is calculated by the weight matrix, and finally multiplied by the weight coefficient to obtain the word vector of the required length.
[0043] For example, the voice data of the operator communicating with the customer is converted into text information. One phone call corresponds to one text message. After word segmentation and word embedding operations are performed on each voice text, the word vector is obtained:
[0044] x i =[x i1 ,x i2 ,x i3 ,……,x in ]
[0045] X=[x 1 ,x 2 ,x 3 ,……,x N ]
[0046] Among them, x i is the vector expression of a voice text, n is the text length, X is the vector matrix of all voice texts, and N is the total number of voice texts.
[0047] S103: Obtaining a first text vector corresponding to the first text segmentation based on the text feature of the first text and the position information of the first text segmentation in the first text;
[0048] In one embodiment, the first text vector corresponding to the first text segmentation includes two feature vectors, one of which is the feature vector of the first text segmentation, and the other is the position feature vector of the first text segmentation in the first text.
[0049] Furthermore, based on the text features of the first text, the first text segmentation in the first segmentation set is encoded to obtain a first segmentation vector corresponding to the first text segmentation; based on the position information of the first text segmentation in the first text, the position information of the first text segmentation is encoded to obtain a first position vector corresponding to the first text segmentation; based on the first segmentation vector and the first position vector corresponding to the first text segmentation, the first text vector corresponding to the first text segmentation is obtained.
[0050] In one embodiment, position information is introduced into the text features of the first text segmentation itself, that is, the position information of the first text segmentation is introduced to improve the accuracy of feature extraction. For example, the feature expressions calculated by the previous steps for "I love you" and "You love me" are the same, but due to the different order of words, the semantics of their expressions are also different. Therefore, a position vector t is added to each input. 1 To make up for the missing position:
[0051] y 1 =x 1 +t 1
[0052] Among them, y 1 represents the first text vector, x 1 Represents the word vector of the first text segmentation, t 1 The position vector representing the first text token.
[0053] Understandably, t 1 Following a specific pattern of model learning (available in a variety of algorithms), that is, the way position information is introduced matches the model's architecture and learning mechanism. The representation and fusion of position information is not fixed, but can be implemented through a variety of algorithms. The specific choice depends on the model's goals and task requirements. For example: In Transformer, position encoding is usually introduced by adding it to word vectors (such as the sine and cosine function encoding mentioned earlier). In some RNN-based models, position information may be implicitly represented by the index of the time step without the need for explicit position encoding. In some models, position information may be implemented through learning parameters (such as trainable position encoding).
[0054] Specifically, the first text is first segmented to obtain a first segmentation set. Then, based on the text features of the first text, each first text segmentation in the first segmentation set is encoded. The encoding here can adopt the various vectorization methods mentioned above, such as the bag-of-words model, One-Hot encoding, word embedding, etc. Taking word embedding as an example, pre-trained Word2Vec, GloVe and other models can be used to map each first text segmentation to a high-dimensional continuous vector space, thereby obtaining the first segmentation vector corresponding to each first text segmentation.
[0055] The position information of the first text segmentation is encoded to obtain a first position vector corresponding to each first text segmentation. There are many ways to encode the position information, such as using one-hot encoding (One-Hot Encoding) of the position index, trigonometric function encoding, etc.
[0056] The first word segmentation vector and the first position vector corresponding to each first text segmentation are fused to obtain the final first text vector corresponding to each first text segmentation. The fusion method can be a simple concatenation, that is, directly concatenating two vectors together to form a longer vector. It can also be a more complex method such as weighted summation. The appropriate fusion strategy is selected according to the specific application scenario and requirements.
[0057] In a dialogue system, it is necessary to understand the semantic relationship between question and answer text. The text feature vector representation method that combines text vectorization and position vectorization can better represent the position information of words in the text, which helps to better understand the structure and semantics of the text, thereby improving the accuracy and relevance of question and answer.
[0058] S104: Based on the multi-head attention matrix, perform feature recognition on the first text vector to obtain target features;
[0059] Multi-Head Attention is one of the core components of the Transformer architecture. It enhances the model's attention capture ability by processing multiple attention distributions in parallel. Specifically, the multi-head attention mechanism processes the input features (usually queries, keys, and values) through multiple independent, parallel-running attention modules (or "heads"). Each head independently calculates the attention score and generates an attention-weighted output. These outputs are then merged (usually by concatenation or averaging) to form a final, more complex representation.
[0060] In one embodiment, three different linear transformation layers are first used to obtain query, key, and value matrices. The query, key, and value matrices are divided into multiple heads (i.e., multiple subspaces), each with different linear transformation parameters. For each head, a scaled dot product attention operation is performed, and the formula is as follows:
[0061]
[0062] Where Q, K, and V represent query, key, and value matrices, respectively. k is the dimension of the key vector, used to scale the dot product to stabilize the softmax function.
[0063] The outputs of all heads are concatenated together and then fused through a linear transformation layer to obtain the final output vector.
[0064] For example, assuming that the first text vector is X=[x 1 ,x 2 ,x 3 ,……,x N ], x i is the vector expression of the i-th first text segmentation, and N is the total number of first text segmentations. The first text vector X passes through three different linear transformation layers to obtain the query, key, and value matrices respectively:
[0065] Q=X×W Q
[0066] K=X×W K
[0067] V=X×W V
[0068] Among them, Q:query, to be queried, K:key, waiting to be queried, V:value, actual feature information. Q , W K , W V is a learnable weight matrix with dimensions [d model , d model ].
[0069] Divide Q, K, and V into h heads, and the dimension of each head is
[0070] Q split =split(Q,h)
[0071] K split =split(K,h)
[0072] V split=split(V,h)
[0073] For each head, compute the scaled dot product attention:
[0074]
[0075]
[0076] Concatenate the output of all headers together:
[0077] output_concat=concat(head 1 , head 2 , …, head h )
[0078] Through the linear transformation layer fusion, the final output vector, that is, the target feature, is obtained:
[0079] output=output_concatW 0
[0080] Among them, output is the vector representation of the target feature, W 0 As the model is trained, it is initialized from Xavier.
[0081] S105: Based on the target feature, identify the conversation intention corresponding to the first text and generate a recommended text.
[0082] In one embodiment, the target features corresponding to the first text can be collected multiple times through a multi-head attention mechanism to obtain multiple target features, and then the extracted multiple target features are fused, and the fused features are used as the final target features. Alternatively, the features extracted by the multi-head attention mechanism can be fused with the features obtained by other feature extraction methods (such as TF-IDF, Word2Vec, etc.). For example, the multi-head attention features can be spliced with the word vector features obtained based on Word2Vec to form a richer feature representation. In this way, the advantages of different feature extraction methods can be combined to more comprehensively capture the semantic information and key elements in the text, providing a more solid feature foundation for recommendation.
[0083] In one embodiment, in the encoder part, a multi-head attention mechanism is used to encode the text features of the input first text; in the decoder part, a multi-head attention mechanism is also used to combine the output of the encoder to gradually generate recommended texts. By stacking multiple layers of encoders and decoders, the model's semantic understanding of the first text and its ability to generate recommended texts can be further deepened.
[0084] For example, in the process of generating recommendation texts, some conditional variables can be introduced, such as customer profile information (age, gender, occupation, hobbies, etc.), product type (such as financial property insurance products, medical insurance products, etc.), sales scenarios, etc., to make the generated texts more targeted and personalized. For example, when the sales target is a young female customer and the product is an insurance product for women's health, the model can generate recommendation texts that are more in line with the characteristics and needs of the customer group based on these conditional variables, thereby improving the appeal and persuasiveness.
[0085] For example, the quality of generation can be further improved by combining pre-trained language models (such as GPT, BERT, etc.). Pre-trained language models are pre-trained on large-scale text data and have learned a wealth of language knowledge and semantic information. In text recommendation tasks, the pre-trained language model can be combined with a model based on a multi-head attention mechanism to utilize the semantic understanding and language generation capabilities of the pre-trained model to generate more natural, fluent, and accurate results. For example, the features extracted by the multi-head attention mechanism can be used as input conditions for the pre-trained language model, allowing the pre-trained model to generate corresponding results based on these features.
[0086] Furthermore, based on a preset text recommendation strategy, feature matching is performed between the target feature and the text feature of a preset text library to determine the preset text with the highest matching degree; based on the preset text, the dialogue intent corresponding to the first text is identified to generate the recommended text.
[0087] In one embodiment, a pre-established text recommendation strategy can be used to identify the conversational intent of the conversation partner in the first text based on the target features of the first text extracted by the multi-head attention mechanism matrix, thereby improving the content accuracy of the generated recommended text.
[0088] Exemplarily, the text recommendation strategy can be a recommendation based on similarity: calculate the feature similarity between the target feature and the text features in the pre-stored text library, and recommend the most matching recommended text based on the similarity ranking. Similarity calculation can use methods such as cosine similarity and Jaccard similarity. For example, the cosine similarity is calculated between the target feature of the first text and the feature vector of each text in the text library, and the texts with the highest similarity are selected as recommended texts. This method is simple and efficient, and can quickly find the most similar to the current sales scenario, providing instant reference for sales staff.
[0089] Exemplarily, the text recommendation strategy can be a rule-based recommendation: formulate recommendation rules according to business rules and sales logic. For example, formulate corresponding recommendation rules according to different sales stages (opening remarks, product introduction, objection handling, facilitation of transactions, etc.); or recommend different strategies according to the customer's purchase intention (high, medium, low). When the target feature of the first text triggers a rule, the corresponding one is recommended to ensure that the recommendation is in line with the business process and sales strategy, and improve sales efficiency and success rate.
[0090] Exemplarily, the text recommendation strategy can also be a recommendation based on reinforcement learning: the recommendation problem is modeled as a reinforcement learning task, and the recommendation strategy is continuously optimized through interaction with the environment (the interaction process between the salesperson and the customer). A reward function is defined, and corresponding rewards or penalties are given according to the customer's feedback after the salesperson uses the recommendation (such as whether to continue the conversation, whether a transaction is reached, etc.). The model adjusts the recommendation strategy based on the reward signal and learns to choose the best one in different situations to maximize long-term rewards. For example, if the salesperson successfully facilitates a transaction after using the recommendation, the model will receive a positive reward, thereby increasing the probability of selecting the recommendation strategy; conversely, if the recommendation effect is not good, the model will receive a negative reward and adjust the strategy to avoid choosing it again.
[0091] In one embodiment, result feedback information of the recommended text is obtained; based on the result feedback information, matrix parameters of the multi-head attention matrix are iteratively updated to optimize the multi-head attention matrix.
[0092] In one embodiment, the result feedback information may include multi-dimensional feedback information such as user behavior data, user display feedback, salesperson feedback, and system automatic feedback.
[0093] Specifically, the system log records the click rate, browsing time, conversion rate (such as purchase, consultation, etc.) of the recommended text by the user as user behavior data. A user rating module is set up on the recommendation result page to collect the user's satisfaction rating of the recommended text; at the same time, feedback channels are provided, such as comment boxes or opinion buttons, to collect users' specific opinions on the recommended text as user display feedback. Regarding the use of recommended texts by sales staff, data such as the frequency of their use of recommended texts, the guiding effect in actual sales conversations (such as whether it can effectively promote the deepening of conversations), and the transaction rate are collected; at the same time, feedback surveys of sales staff are regularly conducted to obtain their suggestions for improvement of recommended texts as feedback from sales staff. Based on the performance of recommended texts in the dialogue system, indicators such as the duration of the dialogue, the frequency of dialogue interruptions, and the number of negative customer feedback are automatically detected as feedback information at the system level.
[0094] In one embodiment, the collected result feedback information is preprocessed, including data cleaning, feature extraction, and data normalization. Then, a weighted multi-objective loss function is designed by comprehensively considering multiple feedback indicators such as user satisfaction, conversation duration, transaction rate, and frequency of use. For example, the negative log-likelihood loss of the user satisfaction score, the mean square error loss of the conversation duration, the binary cross entropy loss of the transaction rate, and the negative log-likelihood loss of the frequency of use are weighted and summed to obtain a comprehensive loss function, wherein the weights can be adjusted according to actual business needs and the importance of the indicators. The weights of each indicator in the loss function are dynamically adjusted according to changes in feedback information. For example, when it is found that user satisfaction has dropped significantly over a period of time, the weight of the user satisfaction indicator in the loss function is increased to prioritize user satisfaction. Parameter optimization strategies such as gradient descent optimization, learning rate adjustment strategy, and regularization technology application are used to optimize and update the matrix parameters of the multi-head attention matrix to improve the accuracy of the target features output by the multi-head attention matrix.
[0095] It can be seen that in the above scheme, by segmenting the first text, clear semantic units are provided for subsequent feature extraction and intent recognition, thereby improving the accuracy of the recommended text. By combining text features (such as the semantics of vocabulary, part of speech, etc.) and the position information of the segmented words in the text, the generated vector can better capture the structured information of the text and more comprehensively reflect the semantic role and importance of each segmented word in the context, thereby improving the accuracy and semantic understanding ability of the recommendation system. By extracting features from text vectors from multiple angles through a multi-attention matrix, features related to dialogue intent can be more accurately identified, the accuracy of key feature recognition of the first text can be improved, and the dialogue intent of the user input text can be accurately identified. Through clear intent recognition, recommended text that is highly relevant to user needs is generated, which improves the accuracy of the recommended text.
[0096] See also Figure 3 As shown, Figure 3 A flowchart of a second embodiment of a text recommendation method based on a multi-head attention mechanism provided in an embodiment of the present invention.
[0097] like Figure 3 As shown, based on the above Figure 2 In the illustrated embodiment, before step S104, the following steps are further included:
[0098] S201: Acquire a second text;
[0099] In one embodiment, the second text may be text data converted from stock voice data of communication between operators and customers, and used to train the matrix parameters of the multi-head attention matrix in the encoder. During the training process, by inputting these text data into the encoder, the weight matrix in the multi-head attention mechanism can be optimized, so that the model can better capture the key information in the communication between operators and customers.
[0100] After the multi-head attention matrix parameters of the encoder are trained, the real-time communication voice data between the operator and the customer can be directly analyzed through the encoder. Specifically, the real-time voice data will first be converted into text data and then input into the trained encoder. The encoder uses the trained multi-head attention matrix to quickly encode and analyze the text data to generate recommended text related to the current communication content. This can greatly improve the work efficiency of the operator and also improve the customer experience.
[0101] S202: Based on the text processing algorithm, perform text segmentation on the second text to obtain a second segmentation set corresponding to the second text, wherein the second segmentation set includes at least one second text segmentation;
[0102] As mentioned above, a word segmentation algorithm based on rules, statistics or deep learning can be used to split the second text, split the second text into multiple word segments, and obtain a second word segmentation set, which includes one or more second text word segments. A variety of methods can be used to implement, such as bag of words model, One-Hot encoding and word embedding to vectorize the second text word segmentation. For example, word embedding models such as Word2Vec, GloVe, and FastText can be used to map the first text word segmentation to a continuous vector space to capture the semantic and grammatical relationship between the first text word segmentations. Specifically, taking Word2Vec as an example, it includes two modes: CBOW and skip-gram. The CBOW algorithm uses the words before and after the central word to calculate the word vector of the central word. First, the words before and after are one-hot encoded, and then the hidden layer is calculated through the weight matrix, and finally multiplied by the weight coefficient to obtain the word vector of the required length.
[0103] The second text is segmented and vectorized by a text processing algorithm, thereby obtaining text features of the second text.
[0104] S203: Obtaining a second text vector corresponding to the second text segmentation based on the text feature of the second text and the position information of the second text segmentation in the second text;
[0105] In one embodiment, position features are introduced into the text features of the second text, and the position information of the second text segmentation is introduced to further improve the accuracy of feature extraction. A position vector c is added to each input 1 To make up for the missing position:
[0106] b 1 =a 1 +c 1
[0107] Among them, a 1 The word vector (text feature) representing the second text segmentation, b 1 represents the second text vector, c 1 Represents the position vector of the second text segmentation, c 1 Follow a specific pattern learned by the model (there are multiple algorithms available).
[0108] S204: Based on the second text vector, training the matrix parameters of the attention mechanism matrix in the attention mechanism algorithm to obtain the attention mechanism matrix;
[0109] Given the second text vector, the goal is to train the matrix parameters of the attention mechanism matrix in the attention mechanism algorithm to obtain an optimized attention mechanism matrix.
[0110] Furthermore, based on the contextual information of the second text segmentation in the second text and the attention mechanism matrix, the self-attention score of the second text segmentation is calculated to obtain the self-attention feature corresponding to the second text segmentation; based on the self-attention features corresponding to each of the second text segmentations in the second text and the pre-trained weight matrix, the multi-head attention matrix is constructed.
[0111] Specifically, calculate the attention mechanism matrix Q, K, V, W Q , W K and W V is a randomly generated matrix (long means there are multiple sets of W Q , W K and W V Parameter matrix), the parameters inside will change with training:
[0112] Q=X×W Q
[0113] K=X×W k
[0114] V=X×W V
[0115] Among them, Q:query, to be queried, K:key, waiting to be queried, V:value, actual feature information. For example, when searching for "white dress for vacation" in a search engine, the input content is the query. The search engine matches the corresponding key according to the query, such as the type, color, description of the product, etc., and then the matching content searched out according to the similarity between the query and the key is the value.
[0116] S205: Construct the multi-head attention matrix based on the context information of the second text segmentation in the second text and the attention mechanism matrix.
[0117] After calculating the attention mechanism matrix, the self-attention score is further calculated. The core of this is to calculate the global contextual association relationship of the second text segmentation, because in a complete conversation, the contextual information has a great impact on the word.
[0118] For example, the same sentence "What are you doing?", if it is said to a friend, then the implicit meaning is "Let's get together if you have nothing to do", but if it is said to a colleague, then the surface meaning may really be asking "What are you doing?", so it is very important to consider the impact of global context information on words.
[0119] for example:
[0120] h 1 =q 1 ·k 1 +q 1 ·k 2 +…+q 1 ·k n
[0121] h 2 =q 2 ·k 1 +q 2 ·k 2 +…+q 2 ·k n
[0122] …
[0123] Among them, h 1 is the relationship between the first word and all words (including the first word itself) in the first audio text, h 2 Similarly, q 1 =b 1 ×W Q , k 1 =b 1 ×W K , and so on.
[0124]
[0125] Among them, z 1 Indicates h 1 The feature expression vector of It is a parameter introduced to ensure gradient stability, that is, to avoid weight imbalance caused by different vector dimensions. The default value is 8, and it can also be customized. The function of softmax is to normalize the vector, that is, to normalize the similarity, and finally obtain a normalized weight matrix. The larger the weight of a value in the matrix, the closer the association.
[0126] From this, we can get the core formula of the attention mechanism as follows:
[0127]
[0128] The multi-head attention algorithm will correspond to z i Concatenate, because the feed-forward network only needs one matrix. It can be multiplied by a weight matrix W 0 (as the model is trained, initialized from Xavier) to ensure that the input and output dimensions are the same, and obtain the encoder's final multi-head attention matrix Z.
[0129] It should be understood that the order of execution of the steps in the above embodiment does not necessarily mean the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiment of the present invention.
[0130] In one embodiment, a text recommendation device based on a multi-head attention mechanism is provided, and the text recommendation device based on the multi-head attention mechanism corresponds one-to-one to the text recommendation method based on the multi-head attention mechanism in the above embodiment. Figure 4 As shown, the text recommendation device based on the multi-head attention mechanism includes a first text acquisition module 301, a text segmentation processing module 302, a first text vector acquisition module 303, a feature extraction module 304 and a recommended text generation module 305. The functional modules are described in detail as follows:
[0131] A first text acquisition module 301, used to acquire a first text to be searched;
[0132] A text segmentation processing module 302, configured to perform segmentation processing on the first text based on a text processing algorithm to obtain a first segmentation set, wherein the first segmentation set includes at least one first text segmentation;
[0133] A first text vector obtaining module 303 is used to obtain a first text vector corresponding to the first text segmentation based on the text features of the first text and the position information of the first text segmentation in the first text;
[0134] A feature extraction module 304 is used to perform feature recognition on the first text vector based on a multi-head attention matrix to obtain a target feature;
[0135] The recommended text generation module 305 is used to identify the conversation intention corresponding to the first text based on the target feature and generate a recommended text.
[0136] In one embodiment, the first text vector obtaining module 303 is specifically used to:
[0137] Based on the text feature of the first text, encode the first text segmentation in the first segmentation set to obtain a first segmentation vector corresponding to the first text segmentation;
[0138] Based on the position information of the first text segmentation in the first text, encoding the position information of the first text segmentation to obtain a first position vector corresponding to the first text segmentation;
[0139] Based on the first word segmentation vector and the first position vector corresponding to the first text word segmentation, the first text vector corresponding to the first text word segmentation is obtained.
[0140] In one embodiment, the text recommendation device based on the multi-head attention mechanism further includes a multi-head attention matrix construction module, which is used to:
[0141] Get the second text;
[0142] Based on the text processing algorithm, perform text segmentation on the second text to obtain a second segmentation set corresponding to the second text, wherein the second segmentation set includes at least one second text segmentation;
[0143] Based on the text features of the second text and the position information of the second text segmentation in the second text, obtaining a second text vector corresponding to the second text segmentation;
[0144] Based on the second text vector, training the matrix parameters of the attention mechanism matrix in the attention mechanism algorithm to obtain the attention mechanism matrix;
[0145] Based on the context information of the second text segmentation in the second text and the attention mechanism matrix, the multi-head attention matrix is constructed.
[0146] In one embodiment, the multi-head attention matrix construction module is also used to:
[0147] Based on the context information of the second text segmentation in the second text and the attention mechanism matrix, calculate the self-attention score of the second text segmentation to obtain the self-attention feature corresponding to the second text segmentation;
[0148] The multi-head attention matrix is constructed based on the self-attention features corresponding to each of the second text segmentations in the second text and the pre-trained weight matrix.
[0149] In one embodiment, the text recommendation device based on the multi-head attention mechanism further includes a matrix parameter optimization module for:
[0150] Obtaining result feedback information of the recommended text;
[0151] Based on the result feedback information, the matrix parameters of the multi-head attention matrix are iteratively updated to optimize the multi-head attention matrix.
[0152] In one embodiment, the result feedback information includes user behavior data, user display feedback information, salesperson feedback information, and system automatic feedback information.
[0153] In one embodiment, the recommendation text generation module 305 is specifically used to:
[0154] Based on a preset text recommendation strategy, feature matching is performed between the target feature and the text features of a preset text library to determine the preset text with the highest matching degree;
[0155] Based on the preset text, the conversation intention corresponding to the first text is identified, and the recommended text is generated.
[0156] The present invention provides a text recommendation device based on a multi-head attention mechanism, which performs word segmentation processing on a first text to provide clear semantic units for subsequent feature extraction and intent recognition, thereby improving the accuracy of the recommended text. By combining text features and the position information of word segments in the text, the generated vector can better capture the structural information of the text and more comprehensively reflect the semantic role and importance of each word segment in the context, thereby improving the accuracy and semantic understanding ability of the recommendation system. By extracting features from text vectors from multiple angles through a multi-attention matrix, features related to conversation intent can be more accurately identified, the accuracy of key feature recognition of the first text can be improved, and the conversation intent of the user input text can be accurately identified. Through clear intent recognition, a recommended text that is highly relevant to user needs is generated, which improves the accuracy of the recommended text.
[0157] For the specific definition of the text recommendation device based on the multi-head attention mechanism, please refer to the definition of the text recommendation method based on the multi-head attention mechanism in the above text recommendation method, which will not be repeated here. Each module in the above-mentioned text recommendation device based on the multi-head attention mechanism can be implemented in whole or in part by software, hardware and a combination thereof. The above-mentioned modules can be embedded in or independent of the processor in the computer device in the form of hardware, or can be stored in the memory of the computer device in the form of software, so that the processor can call and execute the operations corresponding to the above modules.
[0158] In one embodiment, a computer device is provided. The computer device may be a server, and its internal structure diagram may be as follows: Figure 5 As shown. The computer device includes a processor, a memory, a network interface and a database connected via a system bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile and / or volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The network interface of the computer device is used to communicate with an external client through a network connection. When the computer program is executed by the processor, it implements the functions or steps on the server side of a text recommendation method based on a multi-head attention mechanism.
[0159] In one embodiment, a computer device is provided. The computer device may be a client, and its internal structure diagram may be as follows: Figure 6 As shown. The computer device includes a processor, a memory, a network interface, a display screen and an input device connected via a system bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The network interface of the computer device is used to communicate with an external server via a network connection. When the computer program is executed by the processor, it implements the functions or steps on the client side of a text recommendation method based on a multi-head attention mechanism
[0160] In one embodiment, a computer device is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the following steps when executing the computer program:
[0161] Get the first text to be retrieved;
[0162] Based on a text processing algorithm, segment the first text to obtain a first segmentation set, wherein the first segmentation set includes at least one first text segmentation;
[0163] Based on the text features of the first text and the position information of the first text segmentation in the first text, obtaining a first text vector corresponding to the first text segmentation;
[0164] Based on the multi-head attention matrix, feature recognition is performed on the first text vector to obtain target features;
[0165] Based on the target feature, the conversation intention corresponding to the first text is identified, and a recommended text is generated.
[0166] In one embodiment, a computer readable storage medium is provided, on which a computer program is stored, and when the computer program is executed by a processor, the following steps are implemented:
[0167] Get the first text to be retrieved;
[0168] Based on a text processing algorithm, segment the first text to obtain a first segmentation set, wherein the first segmentation set includes at least one first text segmentation;
[0169] Based on the text features of the first text and the position information of the first text segmentation in the first text, obtaining a first text vector corresponding to the first text segmentation;
[0170] Based on the multi-head attention matrix, feature recognition is performed on the first text vector to obtain target features;
[0171] Based on the target feature, the conversation intention corresponding to the first text is identified, and a recommended text is generated.
[0172] It should be noted that the above functions or steps that can be implemented by the computer-readable storage medium or computer device can refer to the relevant descriptions on the server side and the client side in the aforementioned method embodiment. To avoid repetition, they will not be described one by one here.
[0173] Those skilled in the art can understand that all or part of the processes in the above-mentioned embodiment methods can be completed by instructing the relevant hardware through a computer program, and the computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to memory, storage, database or other media used in the embodiments provided in this application can include non-volatile and / or volatile memory. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM) or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. As an illustration and not limitation, RAM is available in many forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link (Synchlink) DRAM (SLDRAM), memory bus (Rambus) direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM).
[0174] Those skilled in the art can clearly understand that for the convenience and simplicity of description, only the division of the above-mentioned functional units and modules is used as an example. In actual applications, the above-mentioned functions can be distributed and completed by different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above.
[0175] The embodiments described above are only used to illustrate the technical solutions of the present invention, rather than to limit the same. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that the technical solutions described in the aforementioned embodiments may still be modified, or some of the technical features may be replaced by equivalents. Such modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present invention, and should all be included in the protection scope of the present invention.
Claims
1. A text recommendation method based on a multi-head attention mechanism, characterized in that: The method comprises: Get the first text to be retrieved; Based on a text processing algorithm, segment the first text to obtain a first segmentation set, wherein the first segmentation set includes at least one first text segmentation; Based on the text features of the first text and the position information of the first text segmentation in the first text, obtaining a first text vector corresponding to the first text segmentation; Based on the multi-head attention matrix, feature recognition is performed on the first text vector to obtain target features; Based on the target feature, the conversation intention corresponding to the first text is identified, and a recommended text is generated.
2. The text recommendation method based on the multi-head attention mechanism according to claim 1, characterized in that: The obtaining, based on the text feature of the first text and the position information of the first text segmentation in the first text, a first text vector corresponding to the first text segmentation includes: Based on the text feature of the first text, encode the first text segmentation in the first segmentation set to obtain a first segmentation vector corresponding to the first text segmentation; Based on the position information of the first text segmentation in the first text, encoding the position information of the first text segmentation to obtain a first position vector corresponding to the first text segmentation; Based on the first word segmentation vector and the first position vector corresponding to the first text word segmentation, the first text vector corresponding to the first text word segmentation is obtained.
3. The text recommendation method based on the multi-head attention mechanism according to claim 1, characterized in that: Before performing feature recognition on the first text vector based on the multi-head attention matrix to obtain the target feature, the method further includes: Get the second text; Based on the text processing algorithm, perform text segmentation on the second text to obtain a second segmentation set corresponding to the second text, wherein the second segmentation set includes at least one second text segmentation; Based on the text features of the second text and the position information of the second text segmentation in the second text, obtaining a second text vector corresponding to the second text segmentation; Based on the second text vector, training the matrix parameters of the attention mechanism matrix in the attention mechanism algorithm to obtain the attention mechanism matrix; Based on the context information of the second text segmentation in the second text and the attention mechanism matrix, the multi-head attention matrix is constructed.
4. The text recommendation method based on the multi-head attention mechanism according to claim 3, characterized in that: The constructing the multi-head attention matrix based on the context information of the second text segmentation in the second text and the attention mechanism matrix includes: Based on the context information of the second text segmentation in the second text and the attention mechanism matrix, calculate the self-attention score of the second text segmentation to obtain the self-attention feature corresponding to the second text segmentation; The multi-head attention matrix is constructed based on the self-attention features corresponding to each of the second text segmentations in the second text and the pre-trained weight matrix.
5. The text recommendation method based on the multi-head attention mechanism according to claim 1, characterized in that: After the dialog intention corresponding to the first text is identified based on the target feature and a recommended text is generated, the method further includes: Obtaining result feedback information of the recommended text; Based on the result feedback information, the matrix parameters of the multi-head attention matrix are iteratively updated to optimize the multi-head attention matrix.
6. The text recommendation method based on the multi-head attention mechanism according to claim 5, characterized in that: The result feedback information includes user behavior data, user display feedback information, salesperson feedback information and system automatic feedback information.
7. The text recommendation method based on the multi-head attention mechanism according to claim 1, characterized in that: The step of identifying the conversation intention corresponding to the first text based on the target feature and generating a recommended text includes: Based on a preset text recommendation strategy, feature matching is performed between the target feature and the text features of a preset text library to determine the preset text with the highest matching degree; Based on the preset text, the conversation intention corresponding to the first text is identified, and the recommended text is generated.
8. A text recommendation device based on a multi-head attention mechanism, characterized in that: The text recommendation device based on the multi-head attention mechanism includes: A first text acquisition module, used to acquire a first text to be retrieved; A text segmentation processing module, configured to perform segmentation processing on the first text based on a text processing algorithm to obtain a first segmentation set, wherein the first segmentation set includes at least one first text segmentation; A first text vector obtaining module, configured to obtain a first text vector corresponding to the first text segmentation based on text features of the first text and position information of the first text segmentation in the first text; A feature extraction module, used to perform feature recognition on the first text vector based on a multi-head attention matrix to obtain target features; The recommended text generation module is used to identify the conversation intention corresponding to the first text based on the target feature and generate a recommended text.
9. A computer device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the computer program, the steps of the text recommendation method based on the multi-head attention mechanism as described in any one of claims 1 to 7 are implemented.
10. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the steps of the text recommendation method based on the multi-head attention mechanism as described in any one of claims 1 to 7 are implemented.