5G message user dialogue neutral evaluation behavior evaluation method and device
Through a large language model and a multi-task learning framework, and by utilizing feature latent correlation analysis and single-modal feature fusion, the problem of identifying user neutral evaluations in the 5G messaging system was solved, and accurate evaluation of user feedback and optimization of task push strategies were achieved, thereby improving user experience and efficiency.
Patent Information
- Application Number
- CN202510710846.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-29
- Publication Date
- 2025-09-26
AI Technical Summary
The 5G messaging system has difficulty accurately identifying users' neutral evaluations in task-based conversations, resulting in a decrease in the accuracy of task push and user experience. In addition, the reinforcement learning model has difficulty optimizing decisions in multi-round conversations when rewards are sparse.
Using a large language model (LLM) and a multi-task learning framework, the latent feature analysis (LTF) module and the unimodal feature fusion (UFF) module are used to mine the multi-round dialogue features of user feedback. The emotional intensity, acceptance, and intent clarity are combined to determine the emotional score Et, solving the reward sparsity problem and achieving accurate evaluation of neutral evaluations.
The 5G messaging system has improved its ability to identify users' neutral needs, optimized task push strategies, and enhanced user experience and task completion efficiency.
Smart Images

Figure CN120705628A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of 5G message information processing technology, and in particular to a method and device for evaluating neutral evaluation behavior of 5G message user conversations. Background Art
[0002] In task-based conversations over 5G messaging, user feedback can be categorized into two types: explicit evaluations (satisfied, dissatisfied) and neutral evaluations. Explicit evaluations are clearly identifiable through direct user expression or behavior, such as when a user explicitly states, "I'm very satisfied" or "This service doesn't meet my needs." Neutral evaluations, however, lie somewhere between satisfaction and dissatisfaction, often manifesting as hesitation, ambiguity, or uncertainty. These neutral evaluations are difficult to accurately capture using simple keyword matching or predefined templates, as users often don't explicitly express whether they fully accept or reject a service. Instead, they may express an undecided attitude, such as, "I'd like to learn more, but I'm not sure if I need it" or "This option looks good, but I'm considering other options." Traditional rule-based matching or hand-coded finite state machine (FSM) methods can typically only effectively process explicit evaluations—feedback generated through explicit conversation or explicit behavioral responses. Existing methods often struggle to effectively identify and process neutral evaluations because they fail to accurately capture subtle changes in users' attitudes and uncertainty about their needs. This deficiency has led to the 5G messaging system's serious lack of ability to identify users' neutral needs in task-based conversations, which in turn affects the accuracy of task push and user experience.
[0003] Furthermore, in dialogue systems, reinforcement learning can continuously improve the system's response strategy through user interaction, enhancing user experience and system performance. The primary advantage of reinforcement learning in dialogue systems is its ability to continuously adjust system behavior based on user feedback to achieve the optimal dialogue strategy. Unlike traditional rule-based systems, reinforcement learning can dynamically adapt to changes in the environment, thus addressing more complex and dynamic dialogue scenarios. For example, intelligent customer service systems can use reinforcement learning to optimize question-and-answer strategies, select optimal responses, quickly and accurately answer user questions, and improve service quality. However, in task-based dialogue applications such as 5G messaging, reinforcement learning models face the problem of "sparse rewards." Specifically, task-based dialogues typically involve multiple rounds of conversation, with user intent gradually becoming clearer over the course of the conversation. However, the system often only rewards users when tasks are completed or cards are pushed. This results in the model being unable to provide timely feedback on effective interactions in the intervening phase, thus hindering dialogue strategy optimization. With sparse rewards, reinforcement learning struggles to efficiently optimize decision-making during multi-round dialogues, resulting in inefficient task completion or the system failing to accurately guide users toward their task goals. Therefore, addressing this sparse reward landscape is a technical issue addressed by this application. Summary of the Invention
[0004] The main purpose of the present invention is to provide a method and device for evaluating the neutral evaluation behavior of 5G message user conversations to address the deficiencies in the relevant technology.
[0005] In order to achieve the above-mentioned purpose, according to the first aspect of the present invention, a method for evaluating the neutral evaluation behavior of 5G message user conversations is provided, comprising obtaining multiple rounds of 5G conversation interaction messages in the current conversation fed back by the user, and obtaining a multi-round conversation set fed back by all users. Determine the set based on a pre-tuned large language model Each round of dialogue f t The corresponding feature representation f t =[s, a, ic] t , where s∈[0,1] represents the intensity of emotion; a∈[0,1] represents the degree of acceptance; ic∈[0,1] represents the clarity of intention; the feature potential correlation analysis LTF module and the unimodal feature fusion UFF module are used to mine the obtained multi-round feature representations, and the emotion score E is determined based on the multi-task learning framework. t .
[0006] Optionally, the feature potential correlation analysis LTF module uses Transformer to find potential correlations between different feature dimensions; and maps features containing potential correlations to higher dimensions through the soft mapping module SM to achieve full fusion of information from different feature dimensions.
[0007] Optionally, the Transformer includes a multi-head attention module and a forward propagation network. After obtaining the feature representation of emotion intensity s, acceptance a, and intention clarity ic, the multi-head attention module divides the feature representation into three vectors: current dialogue vector Q t , the previous and next dialogue vector K t and the true value vector V t ; Perform linear transformation on all vectors; Q t and K t Send it to the dot product function and softmax function to calculate the attention score, and use K t Dimension d k To restrict the calculation results; using attention score and V t Perform weighted summation to obtain the final calculation result The multiple head results are spliced together to obtain the final multi-head attention calculation result: head t =Attention(Q t W Q , K t W K , Vt W V ), H t =MHA(Q t , K t , V t )=Concat(head1, head2,..., head n )W, where W is a 2k×2k transformation matrix that maps the original vector to a higher dimension; the multi-head attention calculation results are subjected to residual connection and layer normalization: H′ t =LayerNorm(H t +Sublayer(H t )); Input the normalized calculation results into the forward propagation network to mine the nonlinear relationship of features: F t =Relu(H′ t W T +δ1)W+δ2, where δ1 and δ2 are random disturbance terms.
[0008] Optionally, the soft mapping module SM uses a set of vectors of size 1×2k The result F output by the previous forward propagation network t Calculating soft attention Then, the soft attention calculation results are weighted and summed to obtain the soft attention m i The calculation results Stack the calculation results Get the result of Soft Mapping of each feature, that is, each single-modal feature Y s 、Y a 、Y ic .
[0009] Optionally, for the unimodal features obtained by the LTF module, the unimodal feature fusion UFF module uses a one-dimensional tensor T s 、T a and T ic To learn the internal representation of each modality feature use Learn the interaction information between each two modalities; determine the fusion tensor in, represents the outer product between vectors, T s 、T a and T ic Represents the initial vectors of three unimodal features, Y m Represents the fused feature representation.
[0010] Optionally, the sentiment score E is determined based on a multi-task learning framework t Including: representing the fused feature Ym , each unimodal feature Y s 、Y a 、Y ic , and based on each unimodal feature Y s 、Y a 、Y ic The evaluation results E are output by the corresponding evaluation tasks. s 、E a 、E ic , and the evaluation results E s 、E a 、E ic Corresponding 5G card label L s 、L a 、L ic The input is sent to the network sharing layer for fusion; the fused features are classified into sentiment polarity and the sentiment score is determined.
[0011] Optionally, the emotion score E t Calculate the feedback reward to get the feedback reward R t Including: Determine the user feedback score U at time t based on the emotion score at time t and the sum of all historical emotion scores before time t t =E t +γ·E t-1 +γ 2 ·E t-2 +...+γ n ·E t-n , where γ is the reward coefficient with a value of (0,1); based on U t Determine feedback rewards: Among them, φ represents the gradient matrix of the reward function, Q t (s t , a t ) represents the system state S t Next, perform action A t When 5G message user feedback score comprehensive expectation, P π (a t |s t ·Φ) represents the state S t After updating the gradient matrix Φ, the strategy π will give the execution action A t The probability of Indicates that under the condition of determining the gradient matrix, the optimal feedback reward value is to perform action A t Probability P π and user feedback score expectation Q t The maximum value of the sum of products.
[0012] According to the second aspect of the present invention, a 5G message user conversation neutral evaluation behavior evaluation device is provided, including a behavior feedback acquisition unit for acquiring multiple rounds of 5G conversation interaction messages in the current conversation fed back by the user, and obtaining a multi-round conversation set fed back by all users. Large model processing unit, used to determine the set based on the pre-debugged large language model Each round of dialogue f t The corresponding feature representation f t =[s, a, ic]t, where s∈[0, 1] represents the intensity of emotion; s∈[0, 1] represents the degree of acceptance; ic∈[0, 1] represents the clarity of intention; the emotion score determination unit is used to perform feature mining on the obtained multi-round feature representations through the feature potential correlation analysis LTF module and the single modal feature fusion UFF module, and determine the emotion score E based on the multi-task learning framework. t .
[0013] According to a third aspect of the present invention, a computer-readable storage medium is provided, wherein the computer-readable storage medium stores computer instructions, and the computer instructions are used to enable the computer to execute any one of the methods described in the first aspect.
[0014] According to a fourth aspect of the present invention, an electronic device is provided, comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores a computer program executable by the at least one processor, and the computer program is executed by the at least one processor so that the at least one processor executes the method described in any one implementation of the first aspect.
[0015] This embodiment of the 5G message user dialogue neutral evaluation behavior evaluation method and device, wherein the method includes obtaining multiple rounds of 5G dialogue interaction messages in the current dialogue fed back by the user, and obtaining a multi-round dialogue set fed back by all users. Determine the set based on a pre-tuned large language model Each round of dialogue f t The corresponding feature representation f t =[s, a, ic] t , where s∈[0,1] represents the intensity of emotion; a∈[0,1] represents the degree of acceptance; ic∈[0,1] represents the clarity of intention; the feature potential correlation analysis LTF module and the unimodal feature fusion UFF module are used to mine the obtained multi-round feature representations, and the emotion score E is determined based on the multi-task learning framework. t Based on a large language model and multi-task learning, the system realizes the evaluation of user neutral behavior, overcoming the problem of the 5G messaging system's serious lack of ability to recognize user neutral needs in task-based conversations in related technologies, and solving the problem of sparse rewards. BRIEF DESCRIPTION OF THE DRAWINGS
[0016] In order to more clearly illustrate the specific embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the specific embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0017] Figure 1a This is a flow chart of a method for evaluating neutral evaluation behavior of 5G message user conversations according to an embodiment of the present invention;
[0018] Figure 1b This is a schematic diagram of the application of the neutral evaluation behavior evaluation method for 5G message users' conversations;
[0019] Figure 1c 2. This is a schematic diagram of a user neutral behavior evaluation framework based on LLM and multi-task learning according to an embodiment of the present invention;
[0020] Figure 2 Schematic diagram of the Transformer structure of LFT in an embodiment of the present invention;
[0021] Figure 3 1 is a schematic diagram of the soft mapping structure of LFT according to an embodiment of the present invention;
[0022] Figure 4 Schematic diagram of the structure of the single-modal feature fusion module UFF in an embodiment of the present invention;
[0023] Figure 5 This is a schematic diagram of the associated structure of the MTLF module according to an embodiment of the present invention;
[0024] Figure 6 is a schematic diagram of an electronic device according to an embodiment of the present invention. DETAILED DESCRIPTION
[0025] In order to enable those skilled in the art to better understand the solutions of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the embodiments described are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts should fall within the scope of protection of the present invention.
[0026] It should be noted that the terms "first," "second," and the like in the specification and claims of the present invention and the accompanying drawings are used to distinguish similar objects and are not necessarily used to describe a particular order or precedence. It should be understood that the terms used in this manner are interchangeable where appropriate for the embodiments of the present invention described herein. In addition, the terms "including," "having," and any variations thereof are intended to cover non-exclusive inclusions. For example, a process, method, system, product, or apparatus comprising a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units that are not explicitly listed or that are inherent to these processes, methods, products, or apparatuses.
[0027] It should be noted that, in the absence of conflict, the embodiments and features of the embodiments of the present invention can be combined with each other. The present invention will be described in detail below with reference to the accompanying drawings and in combination with the embodiments.
[0028] According to an embodiment of the present invention, a method for evaluating the neutral evaluation behavior of 5G message users' conversations is provided. Figure 1a include:
[0029] Step 101: Obtain multiple rounds of 5G dialogue interaction messages in the current dialogue fed back by the user, and obtain a collection of multiple rounds of dialogue fed back by all users
[0030] Step 102: Determine the set based on the pre-debugged large language model Each round of dialogue f t The corresponding feature representation f t =[s, a, ic] t , where s∈[0,1] represents the intensity of emotion; a∈[0,1] represents the acceptance; and ic∈[0,1] represents the clarity of intention.
[0031] When users receive a 5G push notification, their response goes beyond simple acceptance or rejection, often exhibiting intermediate emotional responses. This intermediate feedback reflects the user's interest in the content, the degree of fit with their needs, and the clarity of the feedback, thus requiring quantification through sophisticated analytical methods. This feedback not only provides the system with clues to user preferences but also helps it dynamically adjust its push strategy.
[0032] When users receive a 5G push notification, their response goes beyond simple acceptance or rejection, often exhibiting intermediate emotional responses. This intermediate feedback reflects the user's interest in the content, the degree of fit with their needs, and the clarity of the feedback, thus requiring quantification through sophisticated analytical methods. This feedback not only provides the system with clues to user preferences but also helps it dynamically adjust its push strategy.
[0033] refer to Figure 1b Schematic diagram of the application of the neutral evaluation behavior evaluation method for 5G message user conversations. Assume that the user's feedback on the 5G message push card is in is the set of all possible feedbacks from the user in the current dialogue round T. In order to quantify these feedbacks, we take each dialogue f in the current dialogue round as t It is represented as a vector f t = [s, a, ic]t, where s∈[0, 1] represents Sentiment Strength, which measures the intensity of the user's emotional response to the card in this conversation, such as positive, negative, or neutral. a∈[0, 1] represents Acceptance, which measures the extent to which the user's feedback card in this conversation meets their needs. ic∈[0, 1] represents Intent Clarity, which measures the clarity of the user's feedback in this conversation.
[0034] Each feedback is described by these three dimensions, and all values fall between the interval [0,1]. Below, we will describe three common feedback types:
[0035] (1) Users are very satisfied
[0036] If the user does not provide any text description but simply clicks on the card, we can assume that the user is very satisfied with the card. In this case, the emotional intensity s should be close to 1, indicating that the user's emotional feedback is very positive; the acceptance a should also be close to 1, indicating that the card content fully meets the user's needs; the intention clarity ic is 1, indicating that the user's feedback is very clear and positive. Therefore, in this case, the feedback vector is as follows: this feedback means that the user fully accepts the card content and the feedback is very clear, f t =[1.0, 1.0, 1.0].
[0037] (2) Users are very dissatisfied
[0038] When the user is very dissatisfied with the content of the card and explicitly rejects it, the feedback will show very negative emotions and low acceptance. For example, suppose the system pushes a card about "health insurance", but the user is not interested and even thinks the content of the card is offensive. The user may not leave text feedback, but can express a strong negative attitude by clicking the "Reject" or "Delete" button. In this case: the emotional intensity s will be close to 0, indicating that the user's emotional response is very negative; the acceptance a is also close to 0, indicating that the user believes that the content of the card does not meet their needs at all, and is even unpleasant; the intention clarity ic is close to 1, indicating that the user's feedback is very clear, that is, they clearly reject the card. For example, if the user's feedback is "I have no interest at all, the content is too messy" or "This is not the service I need", the feedback vector may be: f t =[0.1, 0.1, 1.0]
[0039] (3) Neutral user feedback
[0040] In 5G message push scenarios, user feedback often manifests as "intermediate feedback," meaning neither full acceptance nor complete rejection of the pushed content. This neutral feedback is particularly common in actual use, especially in subsequent conversations between users and the system. Users often do not directly express strong acceptance or rejection, but instead use a series of indirect actions or neutral responses to reflect their interest, attention, or concerns about the card content.
[0041] Unlike extreme feedback (complete acceptance or complete rejection), intermediate feedback expresses ambiguous and uncertain attitudes. This type of feedback is often reflected in further conversations between the user and the system. During 5G messaging, users express their interest in the content through the gradual development of subsequent conversations, or express their attitudes through vague and ambiguous responses. The system can extract key information from these conversations and adjust subsequent push strategies.
[0042] The following are some typical neutral feedback behaviors: Further questions around needs: Users often ask follow-up questions related to their own needs. These questions indicate that users are interested in the card content but are still further confirming whether the content meets their needs. For example, when the system pushes a card about a health product, the user may ask: "Is this product suitable for my situation?" or "How effective is this product?" These questions do not directly indicate whether they accept or reject the offer, but rather reflect that the user wants to know more specific information to help them make a decision. At this point, the user's emotional intensity is moderate, the acceptance level is low, and the clarity of intention is weak.
[0043] Supplementing information around goals: Users sometimes express interest in a card's content by describing their specific goals, while also indicating some doubts about the content's suitability. For example, after seeing a travel recommendation card, a user might respond, "I'm planning a trip to Europe soon, and I'm not sure if this recommendation matches the city I want to visit." This indicates that the user isn't directly rejecting the card's content but is looking for more details to confirm whether it meets their needs. At this point, the user's emotional intensity is likely at a moderate level, their acceptance is relatively low, and their intent is ambiguous. The system can extract the user's interests from these descriptions.
[0044] Not directly expressing a clear rejection: Unlike direct rejection feedback, users who receive neutral feedback may express some interest in the card content, but due to uncertainty about its applicability, they haven't made an immediate decision. For example, a user might reply, "This information looks good, but I'm not sure if I need it yet," or "I need to learn more about the features of this product." In this case, the user's feedback doesn't express strong aversion or clear acceptance, but rather demonstrates vague interest in the pushed content, hoping to obtain more information before making a decision.
[0045] In 5G message push scenarios, user feedback on push cards typically encompasses multiple dimensions, primarily sentiment strength, acceptance, and intent clarity. Sentiment strength measures the intensity of the user's emotional response to the card content, encompassing positive, negative, or neutral sentiment. Acceptance assesses the degree to which the user believes the card content meets their needs. Intent clarity measures the clarity and clarity of the intent expressed in the user's feedback. Traditional feedback evaluation methods often focus on handling extreme feedback, where users explicitly express acceptance or rejection. However, in practice, users more often express neutral feedback, which falls somewhere between explicit acceptance and explicit rejection and exhibits greater complexity and uncertainty. Neutral feedback is often expressed through further description or questions about the user's goals in subsequent conversations, rather than direct affirmation or denial. This type of feedback is not only difficult to quantify but also increases the challenges for the system in understanding the user's true needs and optimizing push strategies.
[0046] Specifically, the main difficulties of neutral feedback include its ambiguity, the multi-dimensional interactions involved, and the dynamic changes in user needs. Users' neutral feedback often manifests as partial interest or doubts about the content of the card. For example, users may ask specific questions related to their goals or describe the details of their needs, but do not make a clear decision to accept or reject. This type of feedback not only requires the system to have accurate natural language understanding capabilities, but also needs to be combined with dynamic analysis of user behavior to ensure the comprehensiveness and accuracy of the evaluation results. In addition, the diversity and context dependence of neutral feedback make it difficult for traditional rule-based or simple machine learning methods to effectively capture its intrinsic meaning. Therefore, how to effectively use large language models (LLMs) to evaluate users' neutral behavior has become a difficult problem that needs to be solved urgently.
[0047] Step 103: Perform feature mining on the obtained multi-round feature representations through the feature potential correlation analysis LTF module and the unimodal feature fusion UFF module, and determine the emotion score E based on the multi-task learning framework. t .
[0048] In this step, refer to Figure 1c Schematic diagram of the user neutral behavior evaluation framework based on LLM and multi-task learning. The LLM-based representation learning includes:
[0049] (1) Time series representation of current conversation data. In the 5G messaging system, the interaction between users and the system is usually carried out through multiple conversation rounds. Each conversation round contains multiple interactions between users and the system. In order to better understand the process of these interactions, these conversations are represented as a time series model. At a given moment, the current conversation round Includes all historical interactions from the first conversation to the current conversation. Assume that the user's current conversation turn is Then the round includes the conversation from the first conversation f1 to the tth conversation f t A series of interactive processes. Each dialogue f t represents the interaction content that occurred at the tth moment, where t represents the time sequence of the conversation. Each interaction f in the entire conversation process t All of these may involve different information transmission and intention expression, which can be expressed by the following mathematical formula:
[0050] This multi-round dialogue process constitutes a continuous interactive cycle between the user and the system, and the content of each dialogue is fed back and updated based on the user's input and the system's output. t It is the cumulative result of all previous conversations. The system can adjust subsequent response strategies based on the content of previous conversations.
[0051] (2) Prompt engineering design:
[0052] LLM demonstrates strong performance in sentiment analysis, particularly in extracting emotional features from neutral comments. Traditional sentiment analysis models can suffer from significant errors when processing neutral emotions, as neutrality is often difficult to define. However, LLM, through pre-training on large amounts of text data, can more precisely grasp the nuances of emotion during contextual understanding and reasoning, thus accurately identifying neutral user emotions. This enables LLM to more accurately capture the strength and bias of user emotions in 5G messaging conversations, particularly when analyzing user opinions and feedback.
[0053] The following is a project description for extracting features from 5G messaging user conversations designed in this embodiment: Role Setting. Assume you are a professional 5G messaging customer service representative with extensive experience communicating with 5G messaging customers. You can clearly and objectively assess the emotional intensity (s), acceptance of 5G cards (a), and clarity (ic) of user feedback.
[0054] Features to be extracted. You need to extract the above three key features from the feedback of 5G message users:
[0055] - Sentiment Strength(s): measures the strength of the user's sentiment, ranging from 0 to 1, where 0 represents very negative sentiment, 1 represents very positive sentiment, and neutral sentiment is close to 0.5.
[0056] - Acceptance of 5G cards (a): indicates the user's acceptance of the 5G card content, ranging from 0 to 1, where 0 indicates complete rejection, 1 indicates complete acceptance, and 0.5 indicates neutral.
[0057] - Feedback clarity (ic): measures the clarity of user feedback, ranging from 0 to 1, where 0 indicates completely unclear and 1 indicates completely clear.
[0058] Feature extraction examples. You can refer to the following feature extraction examples:
[0059] -Emotional intensity(s):
[0060] Conversation history: User: I am very angry, I don’t like this service!
[0061] System: I’m sorry, what problem did you encounter?
[0062] Current conversation: User: This service is really bad, I really can’t stand it anymore!
[0063] Feature extraction request: Please analyze the emotional intensity of the current conversation and output a sentiment value ranging from 0 to 1, where 0 represents very negative emotion, 1 represents very positive emotion, and 0.5 represents neutral emotion.
[0064] Output: Sentiment intensity (s): 0.1
[0065] - Acceptance of 5G cards (a):
[0066] Conversation history: User: I just received a 5G card, it looks complicated.
[0067] System: This card contains your 5G plan details, where you can find various service options.
[0068] Current conversation: User: Well, it seems that the information in this card is quite useful. I can learn more.
[0069] Feature extraction request: Please analyze the user's acceptance of 5G cards in the current conversation and output a value ranging from 0 to 1, where 0 indicates complete rejection, 1 indicates complete acceptance, and 0.5 indicates neutrality.
[0070] Output: Acceptance of 5G card (a): 0.7
[0071] - Clarity of feedback (ic):
[0072] Conversation history: User: I don’t quite understand what the service you provide is about.
[0073] System: No problem, please tell me what you are not clear about and I can explain it to you.
[0074] Current conversation: User: I want to know how to check my 5G data usage.
[0075] Feature extraction request: Please analyze the clarity of user feedback in the current conversation and output a value ranging from 0 to 1, where 0 indicates complete ambiguity and 1 indicates complete clarity.
[0076] Output: Clarity of feedback (ic): 0.9
[0077] Flexibility and generalization. Please note that during actual learning, you may need to adjust the details of the prompts based on the different contexts of the conversation. For example, sentiment analysis may need to be optimized for different ways users express emotions such as anger, joy, and doubt; and analysis of 5G card acceptance may need to be adjusted based on the specific content of the card (such as package discounts, data usage information, etc.).
[0078] (3) LLM’s current dialogue feature representation
[0079] After completing the above-mentioned 5G message conversation data information collation and prompt engineering design, we can input the conversation time sequence within the current round into the pre-prepared LLM, and the LLM will complete the feature representation of the current round of conversation. The following is a specific example:
[0080] Current round of dialogue:
[0081] -User Ask1: I recently received a 5G card, but I'm not quite sure what it means.
[0082] -System Ans1: Hello! Thank you for using our 5G service. This card contains your 5G package information. Please let me know if you have any questions and I will answer them for you.
[0083] -User Ask2: I don’t understand what the package includes. Can you explain it to me in detail?
[0084] -System Ans2: Of course, the package includes unlimited 5G data package, 5G high-speed service and some additional benefits.
[0085] -User Ask3: I still don’t quite understand what unlimited traffic means. Can you explain it more clearly?
[0086] -System Ans3: Unlimited traffic means you can surf the Internet freely on the 5G network without being restricted by traffic until the end of the package period.
[0087] -User Ask4: I see. Does the 5G network cover my area?
[0088] -System Ans4: For 5G network coverage, please refer to our official coverage map. You can check the network signal conditions in your area on the map.
[0089] -User Ask5: OK, I'll check it out. Thanks!
[0090] -System Ans5: You’re welcome.
[0091] LLM feature representation output:
[0092] -Sentence intensity feature representation: s5 = [0.4, 0.5, 0.5, 0.6, 0.8]. The user's emotions gradually improve during the conversation, and the emotions tend to be positive when they go from confusion to confirmation of understanding.
[0093] -Acceptance of 5G cards: a5 = [0.4, 0.5, 0.6, 0.7, 0.8]. Users' acceptance of 5G cards gradually increases with the explanation, and ultimately shows a high level of acceptance.
[0094] - Clarity of feedback: ic5 = [0.3, 0.5, 0.6, 0.8, 1.0], where users’ feedback progresses from vague to clear, and finally to completely clear and understood.
[0095] As an optional implementation method of this embodiment, the feature potential correlation analysis LTF module uses Transformer to find potential correlations between different feature dimensions; and uses the soft mapping module SM to map features containing potential correlations to higher dimensions to achieve full fusion of information from different feature dimensions.
[0096] In this optional implementation, the feature representation of the current dialogue turn is completed based on the large language model LLM, and feature mining is completed through the LFT and UFF modules.
[0097] The feature latent association analysis module (Latent-Features Transformer, LFT) mainly uses Transformer to find the latent associations between different feature dimensions, and uses the soft mapping module (Soft Mapping, SM) to map the features containing latent associations to higher dimensions to achieve full fusion of information from different feature dimensions, ultimately achieving more accurate feature representation.
[0098] refer to Figure 2 The figure shows the Transformer structure of LFT. The Transformer structure includes a multi-head attention module (Multi-Head Attention, MHA) and a forward propagation network (Forward Neural Networks, FNN). The parameters between all modules are independent, and the output of the previous module serves as the input of the next module.
[0099] As an optional implementation of this embodiment, the Transformer includes a multi-head attention module and a forward propagation network. After obtaining the feature representation of emotion intensity s, acceptance a, and intention clarity ic, the multi-head attention module divides the feature representation into three vectors: the current dialogue vector Q t , the previous and next dialogue vector K t and the true value vector V t ; Perform linear transformation on all vectors; Q t and K t Send it to the dot product function and softmax function to calculate the attention score, and use K t Dimension d k To restrict the calculation results; using attention score and V t Perform weighted summation to obtain the final calculation result The multiple head results are spliced together to obtain the final multi-head attention calculation result: head t =Attention(Q t W Q , K t W K , V t W V ), H t =MHA(Q t , K t , V t )=Concat(head1, head2,..., head n )W, where W is a 2k×2k transformation matrix that maps the original vector to a higher dimension; the multi-head attention calculation results are subjected to residual connection and layer normalization: H′ t =LayerNorm(H t +Sublayer(H t )); Input the normalized calculation results into the forward propagation network to mine the nonlinear relationship of features: F t =Relu(H′ t W T +δ1)W+δ2, where δ1 and δ2 are random disturbance terms.
[0100] In this optional implementation, the emotion intensity s, acceptance a, and intention clarity ic are input into the multi-head attention module respectively, and the multi-head attention mechanism is used to mine the potential feature correlation information. The specific calculation process is as follows: The features contained in the conversation are divided into three vectors: the current conversation vector Q t , the previous and next dialogue vector K t and the true value vector V t , and perform linear transformation on all vectors; then, Q t and K t Send it to the dot product function and softmax function to calculate the attention score, and use K t Dimension d k To limit the calculation results to ensure that the inner product is not too large; Finally, use the attention score and V t The final calculation result is obtained by weighted summation. The specific formula is as follows:
[0101]
[0102] The above calculation process is performed multiple times, and each calculation is regarded as a head. The final multi-head attention calculation result can be obtained by splicing the results of multiple heads. The specific formula is shown below.
[0103] headt =Attention(Q t W Q , K t W K , V t W V )
[0104] H t =MHA(Q t , K t , V t )=Concat(head1, head2,..., head n )W
[0105] Among them, W is a 2k×2k transformation matrix that maps the original vector to a higher dimension.
[0106] After obtaining the attention calculation results, the output of each layer in LGT needs to be processed through residual connection and layer normalization (R&L). The specific formula is shown below.
[0107] H′ t =LayerNorm(H t +Sublayer(H t ))Finally, the normalized calculation results are passed into FFN to mine the nonlinear relationship of features to enhance the expressiveness of the features. The specific formula is shown below.
[0108] F t =Relu(H′ t W T +δ1)W+δ2
[0109] Among them, δ1 and δ2 are two random perturbation terms that improve the robustness of FFN.
[0110] As an optional implementation of this embodiment, the soft mapping module SM uses a set of vectors with a size of 1×2k The result F output by the previous forward propagation network t Calculating soft attention Then, the soft attention calculation results are weighted and summed to obtain the soft attention m i The calculation results Stack the calculation results Get the result of Soft Mapping of each feature, that is, each single-modal feature Y s 、Y a 、Y ic .
[0111] In this optional implementation, refer to Figure 3 Schematic diagram of the soft mapping structure of LFT. The model has learned the potential correlation information between features. It is necessary to project the learning results of each feature into a new performance space in the soft mapping module for fusion and then score.
[0112] Specifically, using a set of vectors of size 1×2k The result F output by the previous forward propagation network t Calculate the soft attention and then perform weighted summation of the results to get the soft attention m i The calculation results are shown as follows:
[0113] By stacking these results, we can get the soft mapping results of each feature, as shown below
[0114] As an optional implementation method of this embodiment, for the single-modal features obtained by the LTF module, the single-modal feature fusion UFF module uses a one-dimensional tensor T s 、T a and T ic To learn the internal representation of each modality feature use Learn the interaction information between each two modalities; determine the fusion tensor in, represents the outer product between vectors, T s 、T a and T ic Represents the initial vectors of three unimodal features, Y m Represents the fused feature representation.
[0115] In this optional implementation, in the latent-features transformer (LFT), transformers are used to learn the latent correlation feature information between features separately. This learning method is similar to the data splicing mode, which will cause the interaction relationship between features to be lost in the low-dimensional space. Therefore, in order to capture the interaction between two or three features, this method designs a unimodal feature fusion module (UFF) to feed the unimodal features into a higher-dimensional space for fusion. Figure 4 ,Schematic diagram of the structure of the unimodal feature fusion module (UFF),UFF considers different factors in different fusion stages. In the first stage, a one-dimensional tensor T is used s 、T a and T icTo learn the internal representation of each modality. In the second stage, use Learn the interaction information between each two modalities. In the third stage, the final fusion tensor is obtained by combining the results of the previous steps. The specific calculation process is defined as:
[0116]
[0117] in, represents the outer product between vectors, T s 、T a and T ic Represents the initial vectors of three features, Y m Represents the fused feature representation. In this way, an additional dimension of size 1 can be added to the three single-modal tensors, and then their outer products are calculated to obtain the fused features.
[0118] As an optional implementation of this embodiment, determining the emotion score Et based on the multi-task learning framework includes: m , each unimodal feature Y s 、Y a 、Y ic , and based on each unimodal feature Y s 、Y a 、Y ic The evaluation results E are output by the corresponding evaluation tasks. s 、E a 、E ic , and the evaluation results E s 、E a 、E ic Corresponding 5G card label L s 、L a 、L ic The input is sent to the network sharing layer for fusion; the fused features are classified into sentiment polarity and the sentiment score is determined.
[0119] In this optional implementation, the LFT module and the UFF module respectively mine potential correlation features and fusion features in 5G message user evaluation information. Next, a multi-task learning framework (MTLF) is designed to integrate the output of the individual evaluation tasks and ultimately obtain the final feature description of the conversation.
[0120] refer to Figure 5 The schematic diagram of the MTLF module association structure is shown, regarding:
[0121] (1) Network sharing layer:
[0122] The core of MTLF is to construct shared multi-task feature representations, neurons and weights in the low-level network through a sharing mechanism, and complete the final feature representation output through neuron and weight fusion in the high-level network. The input information of the low-level network includes: the fusion feature representation Y obtained by the UFF module m The correlation feature representation Y of the emotional intensity, acceptance, and intention clarity obtained by the LTF module s 、Y a 、Y ic ; In addition, through the evaluation functions of separate emotion intensity, acceptance, and intention clarity, the single task output result E can be obtained s 、E a 、E ic , and the corresponding 5G card label L s 、L a 、L ic Here, we set the 5G card tag L s 、L a 、L ic And the output result E s 、E a 、E ic There is a one-to-one correspondence.
[0123] (2) Polarity classifier
[0124] At this point, the feature representations are fused based on the common neurons and weights. This method designs a model classifier that ultimately expresses the three features of the historical conversation as three neutral evaluation polarities. Combined with the initial two extreme evaluations, the user's current conversation can ultimately be expressed as five final evaluations.
[0125] Among them, "evaluation polarity" is used to describe the emotional tendency of evaluation information. The "evaluation polarity" described here in this method refers to the representation of the emotional tendency of evaluation information after characterization. "3 neutral evaluation polarities" refers to the associated feature representation of the emotional intensity, acceptance, and intention clarity of the neutral evaluation obtained by the LTF module. "2 extreme evaluations" refer to "positive evaluation features" and "negative evaluation features" respectively. In this embodiment, if the evaluation is identified as a "positive evaluation", its feature representation can be regarded as "approaching 1 (→1)", conversely, if the evaluation is identified as a "negative evaluation", its feature representation can be regarded as "approaching 0 (→0)".
[0126] The degree of deviation of the current dialogue representation from the three neutral evaluation polarity centers can be calculated using the Bhattacharyya coefficient, which is calculated as follows:
[0127]
[0128] Where K represents the number of elements in the modality representation, j represents the index of the j-th dimension (from 1 to K), F(j) represents the feature value of the current conversation representation on the j-th dimension (indicating the degree of activation of the dimension), and C p (j) and C n (j) represents the positive / negative eigenvalue of the neutral evaluation polarity center on the jth dimension, S p and S n The Bhattacharyya coefficient (i.e., similarity) represents the degree of positive (p) and negative (n) deviation of the current dialogue representation from the neutral evaluation polarity center.
[0129] The emotional polarity of a 5G messaging user’s neutral evaluation in the current conversation can be determined by the degree of deviation in three situations:
[0130] If S p >2S n , the sample is closer to the center, the emotional polarity is relatively satisfied, and MTLF selects S p The value is taken as the final 5G message user sentiment score E t .
[0131] If S n >2S p , the sample is closer to the negative center, the sentiment polarity is relatively dissatisfied, and MTLF selects S n The value is taken as the final 5G message user sentiment score E t .
[0132] If S p ≤2S n And S n ≤2S p , it means that the sample is located between the positive center and the negative center, the sentiment polarity is neutral, and MTLF selects The value is taken as the final 5G message user sentiment score E t .
[0133] As an optional implementation of this embodiment, the emotion score E t Calculate the feedback reward to get the feedback reward R t Including: Determine the user feedback score U at time t based on the emotion score at time t and the sum of all historical emotion scores before time t t =E t +γ·E t-1 +γ 2 ·E t-2 +...+γ n ·E t-n , where γ is the reward coefficient with a value of (0,1); based on U tDetermine feedback rewards: Among them, Φ represents the gradient matrix of the reward function, Q t (s t , a t ) represents the system state S t Next, perform action A t When 5G message user feedback score comprehensive expectation, P π (a t |s t ·Φ) indicates that in the state S After updating the gradient matrix Φ, the strategy π will give the execution action A t The probability of Indicates that under the condition of determining the gradient matrix, the optimal feedback reward value is to perform action A t Probability P π and user feedback score expectation Q t The maximum value of the sum of products.
[0134] In this optional implementation, in reinforcement learning theory, the user feedback score U at time t is t The current 5G message user sentiment score and the sum of all historical sentiment scores are used to adjust the score downwards using a discount rate as time progresses from recent to distant, as shown below: t =E t +γ·E t-1 +γ 2 ·E t-2 +...+γ n ·E t-n , where γ is the reward coefficient with a value of (0,1). From the perspective of the entire system, solve the strategy π and state s generated by the system under the intention recognition module at time t t Under the condition, the action a of pushing the kth 5G message card t The optimal feedback reward value, that is, the maximum expected return at time t, can be expressed as:
[0135]
[0136] Where Φ represents the gradient matrix of the reward function; Q t (s t , a t ) means that when the system is in state s t Next, perform action a t The comprehensive expectation of 5G message user feedback score when P π (a t |s t ·Φ) represents the execution strategy a given by strategy π after the state and the gradient matrix Φ is updatedt probability; Indicates that under the condition of determining the gradient matrix, the optimal feedback reward value is the execution strategy a t Probability P π and user feedback score expectation Q t The maximum value of the sum of products; in solving During the process, the gradient update method of φ can be expressed as: β is a hyperparameter that controls the speed of gradient learning. It can be further expressed as
[0137]
[0138] At this point, we can see that is the reward function in the Policy-Based method, denoted as r(·). Then the gradient matrix update can be further simplified as: Φ * ←Φ+β·r(Φ) a,s,t In each iteration, the Adam optimizer used in this embodiment is based on the cumulative reward expectation Q t and gradient matrix Φ, dynamically adjust the learning rate β, thereby improving the learning efficiency and stability of the strategy. Through the above steps, the system can gradually optimize the strategy parameters θ so that in a given state s t Next, select action a t Push kth i 5G message card strategy π θ (a t |s t ) gradually tends to maximize the user's global feedback reward R t . R t The conversation content Ct at time t and the state of the intention recognition module s t , Push 5G message card action A t 、User behavioral intention I t , sentiment score E t , feedback reward R t Together with the array [C, S, A, I, E, R] t Stored in the historical conversation database, it provides a basis for decision making for the t+1 conversation.
[0139] The conversation content Ct and sentiment score E targeted by the feedback t , feedback reward R t , Push action A corresponding to the currently pushed 5G message card after intent recognition t , the corresponding system state S t , and user behavior feedback intention I t Composed of array [C, S, A, I, E, R] tStore into the historical conversation database to utilize the array [C, S, A, I, E, R] in the next conversation t Intent recognition can further improve the accuracy of intent judgment and help push more accurate 5G message cards.
[0140] This embodiment addresses the challenge of identifying neutral evaluations, caused by the ambiguity and elusiveness of user feedback. It also addresses the inability of existing methods to accurately capture subtle changes in user attitudes and uncertain needs, effectively identifying and addressing them. This further addresses the aforementioned shortcomings that have led to the 5G messaging system's severely inadequate ability to identify neutral user needs in task-based conversations, which in turn impacts the accuracy of task push and user experience.
[0141] This approach also addresses the "sparse rewards" issue faced by traditional reinforcement learning models in task-based conversations over 5G messaging. User feedback often appears at the end of a conversation, making it difficult for the system to optimize the interactions in between. By continuously updating and learning the state and rewards of multiple rounds of conversation, combined with historical conversation experience and the capabilities of a large language model, this effectively alleviates the reward sparsity issue and improves the optimization of task completion paths.
[0142] It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer executable instructions, and although a logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in a different order than here. According to an embodiment of the present invention, there is also provided a device for evaluating user neutral behavior based on 5G message dialogue interaction, which is characterized by comprising: a behavior feedback acquisition unit for acquiring multiple rounds of 5G dialogue interaction messages in the current dialogue of user feedback, and obtaining a multi-round dialogue set of all user feedback Large model processing unit, used to determine the set based on the pre-debugged large language model Each round of dialogue f t The corresponding feature representation f t =[s, a, ic] t , where s∈[0,1] represents the intensity of emotion; a∈[0,1] represents the degree of acceptance; ic∈[0,1] represents the clarity of intention; the emotion score determination unit is used to perform feature mining on the obtained multi-round feature representations through the feature potential correlation analysis LTF module and the unimodal feature fusion UFF module, and determine the emotion score E based on the multi-task learning framework. t .
[0143] According to an embodiment of the present invention, the present invention also provides an electronic device, which includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can implement the method described in any of the above embodiments when executing.
[0144] According to an embodiment of the present invention, the present invention further provides a readable storage medium, which stores computer instructions, and the computer instructions are used to enable a computer to implement the method described in any of the above embodiments when executed.
[0145] According to an embodiment of the present invention, the present invention further provides a computer program product, which can implement the method described in any of the above embodiments when executed by a processor.
[0146] Figure 6 A schematic block diagram of an example electronic device 300 that can be used to implement embodiments of the present invention is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital assistants, cellular phones, smartphones, wearable devices, and other similar computing devices.
[0147] like Figure 6 As shown, the electronic device 300 includes a computing unit 301, which can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 302 or a computer program loaded from a storage unit 308 into a random access memory (RAM) 303. Various programs and data required for the operation of the electronic device 300 can also be stored in the RAM 303. The computing unit 301, the ROM 302, and the RAM 303 are connected to each other via a bus 304. An input / output (I / O) interface 305 is also connected to the bus 304.
[0148] Multiple components in the electronic device 300 are connected to the I / O interface 305, including an input unit 306, such as a keyboard, a mouse, etc.; an output unit 307, such as various types of displays, speakers, etc.; a storage unit 308, such as a magnetic disk, an optical disk, etc.; and a communication unit 309, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 309 allows the electronic device 300 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.
[0149] The computing unit 301 can be a variety of general and / or special processing components with processing and computing capabilities. Some examples of the computing unit 301 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units that run machine learning model algorithms, digital signal processors (DSPs), and any appropriate processors, controllers, microcontrollers, etc. The computing unit 301 performs the various methods and processes described above, such as the object matching method. For example, in some embodiments, the object matching method can be implemented as a computer software program that is tangibly contained in a machine-readable medium, such as the storage unit 308. In some embodiments, part or all of the computer program can be loaded and / or installed on the electronic device 300 via the ROM 302 and / or the communication unit 309. When the computer program is loaded into the RAM 303 and executed by the computing unit 301, one or more steps of the method described above can be performed.
[0150] Various embodiments of the systems and techniques described herein can be implemented in digital electronic circuit systems, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), system-on-chip systems (SOCs), programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include being implemented in one or more computer programs that are executable and / or interpreted on a programmable system that includes at least one programmable processor, which can be a special purpose or general purpose programmable processor that can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit data and instructions to the storage system, the at least one input device, and the at least one output device.
[0151] The program code for implementing the method of the present invention can be written in any combination of one or more programming languages. Such program code can be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing device so that when the program code is executed by the processor or controller, the functions / operations specified in the flow chart and / or block diagram are implemented. The program code can be executed entirely on the machine, partially on the machine, as a stand-alone software package, partially on the machine and partially on a remote machine, or entirely on a remote machine or server.
[0152] In the context of the present invention, machine-readable medium can be a tangible medium that can contain or store a program for use with an instruction execution system, device or equipment or used in combination with an instruction execution system, device or equipment. Machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. Machine-readable medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared or semiconductor systems, devices or equipment, or any suitable combination of the foregoing. More specific examples of machine-readable storage media can include electrical connections based on one or more lines, portable computer disks, hard disks, random access memories (RAM), read-only memories (ROM), erasable programmable read-only memories (EPROM or flash memory), optical fibers, portable compact disk read-only memories (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.
Claims
1. A method for evaluating the neutral evaluation behavior of 5G message users' conversations, characterized by: Based on the LLM and multi-task learning framework, the evaluation of neutral evaluation behavior of 5G message users in conversation is realized. The method includes: Get the multi-round 5G dialogue interaction messages in the current dialogue fed back by the user, and get the multi-round dialogue collection fed back by all users Determine the set based on a pre-tuned large language model Each round of dialogue f t The corresponding feature representation f t =[s, a, ic] t , where s∈[0,1] represents the intensity of emotion; a∈[0,1] represents the acceptance; ic∈[0,1] represents the clarity of intention; The feature potential correlation analysis LTF module and the unimodal feature fusion UFF module are used to mine the obtained multi-round feature representations, and the emotion score E is determined based on the multi-task learning framework. t .
2. The method for evaluating neutral evaluation behavior of 5G message user conversations according to claim 1 is characterized in that: The feature potential correlation analysis LTF module uses Transformer to find the potential correlation between different feature dimensions; the soft mapping module SM maps the features containing potential correlation to a higher dimension to achieve full fusion of information from different feature dimensions.
3. The method for evaluating neutral evaluation behavior of 5G message user conversations according to claim 2 is characterized in that: The Transformer includes a multi-head attention module and a forward propagation network, wherein the multi-head attention module obtains the feature representation of emotion intensity s, acceptance a, and intention clarity i c Finally, the feature representation is divided into three vectors: current dialogue vector Q t , the previous and next dialogue vector K t and the true value vector V t ; Perform linear transformation on all vectors; Q t and K t Send it to the dot product function and softmax function to calculate the attention score, and use K t Dimension d k To restrict the calculation results; Using attention score and V t Perform weighted summation to obtain the final calculation result The multiple head results are spliced together to obtain the final multi-head attention calculation result: head t =Attention(Q t W Q , K t W K , V t W V ) head t =Attention(Q t W Q ,K t W K ,V t W V ) H t =MHA(Q t , K t , V t )=Concat(head1,head2,...,head n )W where W is a 2k×2k transformation matrix that maps the original vector to a higher dimension; Perform residual connection and layer normalization on the multi-head attention calculation results: H′ t =LayerNorm(H t +Sublayer(H t )); The normalized calculation results are input into the forward propagation network to mine the nonlinear relationship of features: F t =Relu(H′ t W T +δ1)W+δ2, where δ1 and δ2 are random disturbance terms.
4. The method for evaluating neutral evaluation behavior of 5G message user conversations according to claim 2 is characterized in that: The soft mapping module SM uses a set of vectors of size 1×2k The result F output by the previous forward propagation network t Calculating soft attention Then, the soft attention calculation results are weighted and summed to obtain the soft attention m i The calculation results Stack the calculation results Get the result of SoftMapping of each feature, that is, each single-modal feature Y s 、Y a 、Y ic .
5. The method for evaluating neutral evaluation behavior of 5G message user conversations according to claim 4 is characterized in that: For the unimodal features obtained by the LTF module, the unimodal feature fusion UFF module uses a one-dimensional tensor T s 、T a and T ic To learn the internal representation of each modality feature use Learn the interaction information between each two modalities; Determine the fused tensor in, represents the outer product between vectors, T s , T a , T ic Represents the initial vectors of three unimodal features, Y m Represents the fused feature representation.
6. The method for evaluating neutral evaluation behavior of 5G message user conversations according to claim 5, characterized in that: Determine the sentiment score E based on a multi-task learning framework t include: The fused feature is represented by Y m , each unimodal feature Y s 、Y a 、Y ic , and based on each unimodal feature Y s 、Y a 、Y ic The evaluation results E are output by the corresponding evaluation tasks. s 、E a 、E ic , and the evaluation results E s 、E a 、E ic Corresponding 5G card label L s 、L a 、L ic Input to the network sharing layer for fusion; The fused features are classified into sentiment polarity and the sentiment score is determined.
7. The method for evaluating neutral evaluation behavior of 5G message user conversations according to claim 6, characterized in that: The emotion score E t Calculate the feedback reward to get the feedback reward R t include: The user feedback score U at time t is determined based on the sentiment score at time t and the sum of all historical sentiment scores before time t. t =E t +γ·E t-1 +γ 2 ·E t-2 +...+γ n ·E t-n , where γ is the reward coefficient with value (0,1); Based on U t Determine feedback rewards: Among them, Φ represents the gradient matrix of the reward function, Q t (s t , a t ) represents the system state S t Next, perform action A t When 5G message user feedback score comprehensive expectation, P π (a t |s t ·Φ) represents the state S t After updating the gradient matrix Φ, the strategy π will give the execution action A t The probability, R t (Φ)=Max∑ t (·) indicates that under the condition of determining the gradient matrix, the optimal feedback reward value is to perform action A t Probability P π and user feedback score expectation Q t The maximum value of the sum of products.
8. A device for evaluating neutral evaluation behavior of 5G message user conversations, characterized in that: The evaluation of neutral evaluation behavior of 5G message users in conversation is realized based on LLM and multi-task learning framework. The device includes: The behavior feedback acquisition unit is used to obtain the multi-round 5G dialogue interaction messages in the current dialogue of the user feedback, and obtain the multi-round dialogue set of all user feedback Large model processing unit, used to determine the set based on the pre-debugged large language model Each round of dialogue f t The corresponding feature representation f t =[s, a, ic] t , where s∈[0,1] represents the intensity of emotion; a∈[0,1] represents the acceptance; ic∈[0,1] represents the clarity of intention; The emotion score determination unit is used to perform feature mining on the obtained multi-round feature representations through the feature potential correlation analysis LTF module and the unimodal feature fusion UFF module, and determine the emotion score E based on the multi-task learning framework. t .
9. A computer-readable storage medium, characterized in that The computer-readable storage medium stores computer instructions, and the computer instructions are used to enable the computer to execute the method according to any one of claims 1 to 7.
10. An electronic device, characterized in that: include: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores a computer program executable by the at least one processor, and the computer program is executed by the at least one processor to enable the at least one processor to perform the method according to any one of claims 1 to 7.
Citation Information
Patent Citations
5G message session method and 5G message session system
CN115344683A
Emotion evaluation method and device, equipment and medium
CN119721035A
Intention recognition method for voice question answering based on large-model multi-agent
CN119831043A