A message interaction method and device, a storage medium and an electronic device

CN122554427APending Publication Date: 2026-08-11YIYUNYING (SHANDONG) NETWORK TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-04-30
Publication Date
2026-08-11

AI Technical Summary

Technical Problem

但现有方案大多面向单一平台构建,缺乏针对不同交互环境差异的统一适配机制,难以实现跨平台的稳定交互处理

Benefits of technology

[0017]通过实施上述一种消息交互方法、系统、存储介质及电子设备,该方法通过获取并分析目标交互环境中与会话对象对应的会话数据,得到反映当前轮次交互特征的会话分析信息,再结合会话对象的用户画像,并利用会话分析信息以及前次答复处理数据对用户画像进行更新,之后基于更新后的用户画像以及会话分析信息生成待选答复,用于辅助形成答复交互数据的交互答复。通过上述处理过程,能够将当前轮次会话内容、用户既有特征以及前次答复反馈情况进行结合,使用户画像能够随交互过程动态调整,从而提高对会话对象当前需求、表达习惯及关注内容的识别准确性;在更新后的用户画像基础上生成待选答复,使生成的答复更加贴合当前会话语境和用户特征,提高交互答复的针对性、适配性和连续性,进而提升消息交互的准确程度和交互效果,能够适用于不同交互环境下的会话交互场景,并使交互过程中形成的数据得到延续利用,从而提高后续回复内容生成的连贯性和贴合性。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122554427A_ABST
    Figure CN122554427A_ABST
Patent Text Reader

Abstract

This application relates to a message interaction method, apparatus, storage medium, and electronic device. The method includes: acquiring and analyzing session data corresponding to a session object in a target interaction environment to obtain session analysis information; wherein the session data includes the interaction data of the session object in the current round; acquiring a user profile of the session object; updating the user profile using the session analysis information and previous response processing data; generating candidate responses adapted to the updated user profile and session analysis information; wherein the candidate responses are used to assist in generating interactive responses for reply interaction data. This method can achieve cross-platform session interaction and generate subsequent reply content.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the fields of artificial intelligence and computer communication technology, and in particular to a message interaction method, apparatus, storage medium and electronic device. Background Technology

[0002] With the increasing number of instant messaging platforms, social platforms, and collaborative office platforms, users are exchanging messages more frequently across different platforms. Automatic analysis of conversation data, response suggestions, and interactive assistance are gradually becoming important application areas in the field of intelligent interaction. Existing solutions typically improve message processing efficiency by semantically recognizing single-turn message content to generate corresponding response suggestions or by outputting response content based on preset templates. Some solutions enhance the relevance of responses by combining historical conversation records to provide contextual supplementation. However, most existing solutions are designed for single platforms and lack a unified adaptation mechanism for different interactive environments, making it difficult to achieve stable cross-platform interactive processing. Furthermore, existing solutions do not adequately utilize user interaction data during candidate response selection, modification, and subsequent conversation feedback, making it difficult to make targeted adjustments to response generation based on continuous interaction processes. Summary of the Invention

[0003] Therefore, it is necessary to provide a message interaction method, device, storage medium, and electronic device that can realize cross-platform conversational interaction and generate subsequent reply content to address the above-mentioned technical problems.

[0004] Firstly, a message interaction method is provided, which includes: Acquire and analyze the session data corresponding to the session object in the target interaction environment to obtain session analysis information; wherein, the session data includes the interaction data of the session object in the current round; Retrieve the user profile of the session object; Update user profiles using conversation analysis information and previous response processing data; Generate candidate responses that adapt and update the user profile and conversation analysis information; among them, candidate responses are used to assist in generating interactive responses that generate response interaction data.

[0005] In one embodiment, obtaining and analyzing session data corresponding to the session object in the target interaction environment to obtain session analysis information includes: Extract the interaction data of the current round and the historical interaction data from the session data; The interaction data and historical interaction data are arranged and combined according to the time sequence of the conversation to obtain the content of the conversation to be analyzed; Semantic parsing is performed on the content of the conversation to be analyzed to determine one or more of the intent information, sentiment information, and conversation topic information of the content of the conversation to be analyzed; Based on one or more of the intent information, emotion information, and conversation topic information, determine the corresponding intent information, emotion information, and product keywords, and use them as conversation analysis information.

[0006] In one embodiment, generating candidate responses that adapt and update the user profile and conversation analysis information includes: Based on intent information, emotional information, and product keywords, determine the style of the response template, the composition of the response, and the organization of the information for the current round; Based on the style of the response template and the information organization method, at least two matching response expression structures are determined from the preset response expression structures; Obtain the product keywords corresponding to the information that constitutes the response, extract the content elements corresponding to the product keywords from each conversation analysis information, and organize and fill the content elements according to at least two determined response expression structures to generate at least two candidate responses. Among them, at least two candidate responses differ in one or more aspects such as response length, semantic focus, and guiding expression.

[0007] In one embodiment, the candidate responses used to assist in generating response interaction data include: If there are multiple candidate responses, select one of them as the target candidate response. In response to the modification operation of the target candidate response, record the target candidate response, the modified response after modification, and the modification content; In response to the completion of sending the target candidate response or modified response, the actual target candidate response or modified response sent is recorded as an interactive response. Generate response processing data for the current round based on one or more of the target candidate responses, modified content, and interactive responses.

[0008] In one embodiment, updating the user profile using session analysis information and previous response processing data includes: Get the conversation analysis information for the current round, and get the previous response processing data corresponding to the conversation object; Based on the object identifier and session identifier of the session object, the session analysis information of the current round is associated with the previous response processing data to obtain profile update data; Based on the profile update data, update at least one of the following in the user profile: user intent preference, emotional tendency, and product attention preference.

[0009] In one embodiment, after generating candidate responses that adapt and update the user profile and conversation analysis information, the method further includes: Extract one or more of the user preference tags, communication style tags, and product interest tags from the updated user profile that correspond to the current round of conversation analysis information; Based on one or more of the user preference tags, communication style tags, and product focus tags, determine the filtering criteria for multiple candidate responses. The filtering criteria include one or more of the following: response length, tone and style, product information presentation order, and guidance expression method. The multiple candidate responses are matched with the filtering criteria to obtain the matching degree of each candidate response; At least one candidate response that meets the preset matching criteria is selected as the recommended candidate response.

[0010] In one embodiment, before obtaining and analyzing the session data corresponding to the session object in the target interaction environment to obtain session analysis information, the method further includes: Obtain platform characteristic information of the target interaction environment; The platform feature information is matched with the preset platform feature template, and the platform adaptation parameters corresponding to the target interaction environment are determined based on the matching results. Based on platform adaptation parameters, the system controls the reading of session data, input of reply content, and control of message sending in the target interactive environment.

[0011] In one embodiment, obtaining the user profile of the session object includes: In response to a user profile that does not have a session object, retrieve the historical interaction data of the session object; In response to the existence of historical interaction data, an initial user profile of the session object is generated based on the historical interaction data and the preset profile initialization template, and the initial user profile is used as the user profile. In response to the absence of historical interaction data, an initial user profile of the session object is generated based on the interaction data of the current round and the preset profile initialization template, and the initial user profile is used as the user profile.

[0012] In one embodiment, the output of candidate responses includes: Based on the response processing data, determine the response preference characteristics corresponding to the conversation object; Based on the response preference characteristics, the matching degree is calculated for each of the at least two candidate responses generated; Based on the matching score calculation results, sort at least two candidate responses; Output at least two candidate responses based on the sorting results.

[0013] In one embodiment, updating the user profile corresponding to the session object based on interaction samples includes: extracting information on the expression adjustment of the target candidate response content from the modified content; Based on the expression adjustment information, generate expression correction rules corresponding to the session object; Link expression correction rules to user profiles; When generating alternative or new alternative responses, the structure of the response is adjusted based on the expression correction rules.

[0014] Secondly, this application provides a message interaction method applied to the interaction environment side, including: Obtain the interaction data for the current round and send the interaction data to the message interaction device; Receive the candidate responses returned by the message interaction device, and generate an interactive response based on the candidate responses.

[0015] Thirdly, this application provides a message interaction device, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, it implements the message interaction method described in the first aspect.

[0016] Fourthly, this application provides a computer-readable storage medium having a computer program stored thereon, wherein the computer program, when executed by a processor, implements the message interaction method described in the first aspect.

[0017] By implementing the aforementioned message interaction method, system, storage medium, and electronic device, this method acquires and analyzes conversation data corresponding to the conversation object in the target interaction environment to obtain conversation analysis information reflecting the characteristics of the current round of interaction. This information is then combined with the user profile of the conversation object, and updated using the conversation analysis information and previous response processing data. Subsequently, candidate responses are generated based on the updated user profile and conversation analysis information to assist in forming interactive responses. Through this process, the current round of conversation content, existing user characteristics, and previous response feedback can be combined, allowing the user profile to dynamically adjust with the interaction process. This improves the accuracy of identifying the conversation object's current needs, expression habits, and areas of interest. Generating candidate responses based on the updated user profile makes the generated responses more relevant to the current conversation context and user characteristics, improving the relevance, adaptability, and continuity of the interactive responses. This enhances the accuracy and effectiveness of message interaction, making it applicable to conversation interaction scenarios in different interaction environments. Furthermore, it allows for the continued use of data generated during the interaction process, thereby improving the coherence and relevance of subsequent response content generation. Attached Figure Description

[0018] To more clearly illustrate the embodiments of this application, the accompanying drawings used in the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be determined based on these drawings without creative effort.

[0019] Figure 1 A flowchart illustrating a message interaction method provided in an embodiment of this application; Figure 2 This is an overall architecture diagram of a message interaction method in an embodiment of this application; Figure 3 This is a flowchart illustrating the platform adaptation of a message interaction method according to an embodiment of this application. Figure 4 This is a flowchart illustrating the data acquisition and intent analysis process of a message interaction method according to an embodiment of this application. Figure 5 This is a flowchart illustrating the reply generation and customer profile update process of a message interaction method in an embodiment of this application. Figure 6 This is a diagram of the internal structure of the electronic device in an embodiment of this application. Detailed Implementation

[0020] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments determined by those skilled in the art without creative effort are within the protection scope of this application.

[0021] It should be noted that, in the description of this application, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or treatment method that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or treatment method. The terms "first," "second," etc., in this application are used to distinguish similar objects and are not used to describe a specific order or sequence.

[0022] To enable those skilled in the art to better understand the present application, the present application will be further described in detail below with reference to the accompanying drawings and specific embodiments.

[0023] In one embodiment, such as Figure 1 As shown, a message interaction method is provided, including: S100: Obtain and analyze the session data corresponding to the session object in the target interaction environment to obtain session analysis information; wherein, the session data includes the interaction data of the session object in the current round; S200: Obtain the user profile of the session object; S300: Update user profiles using conversation analysis information and previous response processing data; S400: Generate candidate responses that adapt and update the user profile and conversation analysis information; among them, candidate responses are used to assist in generating interactive responses that generate response interaction data.

[0024] In this context, a session object refers to an object participating in message interaction within the target interaction environment; a session object can specifically be a user. A user profile refers to user characteristic descriptions formed based on the session object's historical interaction history, feature information, or preference information. Previous response processing data refers to the processing results data formed before the current round, based on generated, modified, or actually sent responses. Pending responses refer to selectable response content generated based on session analysis information and the updated user profile. Interactive responses refer to the content actually generated and used to respond to the current round's interaction data.

[0025] Specifically, the process involves acquiring and analyzing session data corresponding to the session object within the target interaction environment to obtain session analysis information that reflects the session characteristics embodied in the current round of interaction data. Subsequently, a user profile of the session object is acquired as the foundational information describing its existing characteristics. Combining the currently acquired session analysis information with previous response processing data, the user profile is updated to reflect changes in the session object during continuous interaction. Then, candidate responses are generated based on the updated user profile and session analysis information, and these candidate responses are used to assist in generating interactive responses. This process integrates the current round of interaction content, the existing characteristics of the session object, and the results of previous response processing, allowing the user profile to continuously adjust during interaction, thereby improving its ability to reflect the current needs and interaction characteristics of the session object. Generating candidate responses based on the updated user profile and session analysis information ensures that subsequent interactive responses are more closely aligned with the current round of interaction content and the characteristics of the session object, thus improving the targeting and adaptability of message interaction and enhancing the continuity, targeting, and relevance of subsequent response generation.

[0026] In one embodiment, obtaining and analyzing session data corresponding to the session object in the target interaction environment to obtain session analysis information includes: Extract the interaction data of the current round and the historical interaction data from the session data; The interaction data and historical interaction data are arranged and combined according to the time sequence of the conversation to obtain the content of the conversation to be analyzed; Semantic parsing is performed on the content of the conversation to be analyzed to determine one or more of the intent information, sentiment information, and conversation topic information of the content of the conversation to be analyzed; Based on one or more of the intent information, emotion information, and conversation topic information, determine the corresponding intent information, emotion information, and product keywords, and use them as conversation analysis information.

[0027] Among these, intent information refers to information used to characterize the purpose or need expressed by the conversation participants in the current conversation. Emotional information refers to information used to characterize the emotional state exhibited by the conversation participants in the current conversation. Product keywords refer to keyword information extracted from the conversation content that can characterize the product content that the conversation participants are interested in.

[0028] Specifically, by extracting current and historical interaction data from the conversation data, and then arranging and combining them according to the conversation's chronological order, the conversation content to be analyzed is formed, creating a continuous semantic link between the current expression and the historical context. Subsequently, semantic parsing is performed on the conversation content to identify one or more of the intent, emotion, and conversation topic information. Based on this, the corresponding intent, emotion, and product keywords are further determined and used as conversation analysis information. This process avoids the problem of incomplete information caused by analyzing only a single interaction from the current round. The resulting conversation analysis information not only reflects the current needs and emotions of the conversation participants but also accurately identifies the conversation topic and focus by combining historical interaction content. Furthermore, extracting product keywords after identifying the conversation topic information helps make the conversation analysis information more relevant to the current communication content, thereby improving the completeness of conversation understanding, the accuracy of semantic recognition, and the targeted nature of subsequent message processing.

[0029] In one embodiment, generating candidate responses that adapt and update the user profile and conversation analysis information includes: Based on intent information, emotional information, and product keywords, determine the style of the response template, the composition of the response, and the organization of the information for the current round; Based on the style of the response template and the information organization method, at least two matching response expression structures are determined from the preset response expression structures; Obtain the product keywords corresponding to the information that constitutes the response, extract the content elements corresponding to the product keywords from each conversation analysis information, and organize and fill the content elements according to at least two determined response expression structures to generate at least two candidate responses. Among them, at least two candidate responses differ in one or more aspects such as response length, semantic focus, and guiding expression.

[0030] Specifically, based on intent information, sentiment information, and product keywords, the style of the response template, the structure of the response, and the organization of information for the current round are determined to clarify the expressive characteristics, key elements, and organizational logic of the response in this round. Then, based on the response template style and information organization, at least two matching response expression structures are determined from the preset response expression structures to provide different expression frameworks for generating candidate responses. On this basis, product keywords corresponding to the response structure information are obtained, and content elements corresponding to the product keywords are extracted from the conversation analysis information. These content elements are then organized and filled according to the determined at least two response expression structures to generate at least two candidate responses. These candidate responses differ in one or more aspects, such as response length, semantic focus, and guiding expression. Through this process, multiple candidate responses with differentiated expressive characteristics can be generated based on the intent information, sentiment information, and product keywords corresponding to the current round. This avoids the problem of insufficient adaptation caused by generating only a single form of response, making the candidate responses more aligned with the current conversation analysis information in terms of expression and semantic presentation, and improving the relevance, flexibility, and selectivity of response generation.

[0031] In one embodiment, the candidate responses used to assist in generating response interaction data include: If there are multiple candidate responses, select one of them as the target candidate response. In response to the modification operation of the target candidate response, record the target candidate response, the modified response after modification, and the modification content; In response to the completion of sending the target candidate response or modified response, the actual target candidate response or modified response sent is recorded as an interactive response. Generate response processing data for the current round based on one or more of the target candidate responses, modified content, and interactive responses.

[0032] Here, "target candidate response" refers to the selected response from multiple candidate responses. "Modified response" refers to the response content formed after the target candidate response has been modified. "Modified content" refers to the content changes that have occurred relative to the target candidate response. "Response processing data" refers to the data generated during the selection, modification, and transmission of responses in the current round.

[0033] Specifically, when using candidate responses to assist in generating interactive response data, if there are multiple candidate responses, a target candidate response is first selected. When modifying the target candidate response, the target candidate response, the modified response, and the modified content are recorded. After the target candidate response or modified response is sent, the actually sent target candidate response or modified response is recorded as an interactive response. Based on one or more of the target candidate response, modified content, and interactive responses, the response processing data for the current round is generated. This system can correlate and record the selection, modification, and actual sending of responses in the current round, forming a complete data link in the response processing process. This provides a basis for subsequent analysis of the adoption, modification, and final output of responses in the current round, improving the traceability and data integrity of the response processing process, and facilitating more accurate data support for subsequent user profile updates or response optimization.

[0034] In one embodiment, updating the user profile using session analysis information and previous response processing data includes: Get the conversation analysis information for the current round, and get the previous response processing data corresponding to the conversation object; Based on the object identifier and session identifier of the session object, the session analysis information of the current round is associated with the previous response processing data to obtain profile update data; Based on the profile update data, update at least one of the following in the user profile: user intent preference, emotional tendency, and product attention preference.

[0035] Among these, object identifier refers to information used to identify the identity of a conversation object. Conversation identifier refers to information used to identify the current conversation. Profile update data refers to data used to update the user profile after associating the current round of conversation analysis information with the previous response processing data. User intent preference refers to information in the user profile reflecting the relatively stable or continuous intent tendencies of the conversation object during the interaction. Emotional tendency refers to information in the user profile reflecting the emotional expression trends exhibited by the conversation object during the interaction. Product attention preference refers to information in the user profile reflecting the degree or direction of attention the conversation object pays to different product content.

[0036] Specifically, the process involves acquiring conversation analysis information for the current round and obtaining previous response processing data for the corresponding conversation object. Then, based on the object identifier and conversation identifier of the conversation object, the conversation analysis information for the current round is correlated with the previous response processing data to obtain profile update data. Based on this, at least one of the user intent preferences, emotional tendencies, and product attention preferences in the user profile is updated according to the profile update data. This allows for the mapping of conversation characteristics reflected in the current round with the previous response processing, enabling targeted adjustments to the user profile based on existing response processing, thereby improving the user profile's ability to reflect changes in the interaction characteristics of the conversation object. By updating at least one of the user intent preferences, emotional tendencies, and product attention preferences, the user profile can more closely reflect the actual performance of the conversation object in continuous interaction, improving the accuracy and dynamic adaptability of the user profile.

[0037] In one embodiment, after generating candidate responses that adapt and update the user profile and conversation analysis information, the method further includes: Extract one or more of the user preference tags, communication style tags, and product interest tags from the updated user profile that correspond to the current round of conversation analysis information; Based on one or more of the user preference tags, communication style tags, and product focus tags, determine the filtering criteria for multiple candidate responses. The filtering criteria include one or more of the following: response length, tone and style, product information presentation order, and guidance expression method. The multiple candidate responses are matched with the filtering criteria to obtain the matching degree of each candidate response; At least one candidate response that meets the preset matching criteria is selected as the recommended candidate response.

[0038] Among them, user preference tags refer to label information used to characterize the preferred characteristics of conversation participants in message interaction. Communication style tags refer to label information used to characterize the preferred communication methods or expression methods of conversation participants. Product interest tags refer to label information used to characterize the focus of conversation participants on different product content. Filtering conditions refer to the conditions used to determine the degree of fit between multiple candidate responses and the user characteristics of the current round. Matching degree refers to the degree of conformity between the candidate responses and the filtering conditions. Recommended candidate responses refer to the candidate responses that are more suitable for use in the current round after filtering from multiple candidate responses.

[0039] Specifically, after generating multiple candidate responses, the system first extracts user preference tags, communication style tags, and product interest tags corresponding to the current round from the updated user profile, and uses these to form filtering conditions. This allows the selection of candidate responses to be combined with the current user characteristics, thereby improving the targeting of the filtering direction. Then, each candidate response is matched with the filtering conditions to distinguish the degree of suitability of each candidate response to the current round. This helps to identify more suitable responses from multiple results. Finally, candidate responses that meet the preset matching conditions are determined as recommended candidate responses. This ensures that the final recommendation results not only come from the response generation process but also further align with the updated user profile, thereby improving the accuracy and suitability of the recommended candidate responses.

[0040] In one embodiment, before acquiring session data in the target interactive environment, the method further includes: Obtain platform characteristic information of the target interaction environment; The platform feature information is matched with the preset platform feature template, and the platform adaptation parameters corresponding to the target interaction environment are determined based on the matching results. Based on platform adaptation parameters, the system controls the reading of session data, input of reply content, and control of message sending in the target interactive environment.

[0041] Among them, platform feature information refers to information used to characterize the characteristics of the target interaction environment in terms of session data structure, interaction method, interface form or message control rules; platform feature template refers to pre-set template information used to describe the corresponding feature patterns of different types of interaction environments; platform adaptation parameters refer to parameters such as data reading rules, content input rules and message control rules determined according to the matching results of platform feature information and platform feature template, used to adapt to the target interaction environment. Specifically, before acquiring and analyzing session data, the platform characteristic information of the target interaction environment is first obtained and matched with a preset platform characteristic template to determine the platform adaptation parameters corresponding to the target interaction environment. This ensures that subsequent processing is based on compatibility with the current platform. Then, based on the platform adaptation parameters, session data reading, reply content input, and message sending control are executed, enabling the session processing to adapt to the interaction methods of different target interaction environments. This improves the adaptability of session data acquisition and message interaction control. Before entering session analysis, platform-side adaptation of the target interaction environment is completed, which helps improve the accuracy and stability of subsequent session processing.

[0042] In one embodiment, obtaining the user profile of the session object includes: In response to a user profile that does not have a session object, retrieve the historical interaction data of the session object; In response to the existence of historical interaction data, an initial user profile of the session object is generated based on the historical interaction data and the preset profile initialization template, and the initial user profile is used as the user profile. In response to the absence of historical interaction data, an initial user profile of the session object is generated based on the interaction data of the current round and the preset profile initialization template, and the initial user profile is used as the user profile.

[0043] The profile initialization template refers to the pre-set template information used to generate the initial user profile. The initial user profile refers to the basic user profile generated when the session object does not yet have a corresponding user profile.

[0044] Specifically, when acquiring user profiles for conversation objects, if no user profile exists, the system first checks for historical interaction data. If historical interaction data exists, an initial user profile is generated based on this data and a pre-defined profile initialization template. This ensures the generated user profile reflects the conversation object's past interaction characteristics more effectively, improving the completeness of the user profile initialization result. If historical interaction data is unavailable, an initial user profile is generated based on the interaction data of the current round and the pre-defined profile initialization template. This allows the system to complete user profile initialization even when historical data is lacking, ensuring the continuity of the user profile generation process. Through this approach, the appropriate initialization path can be selected based on different data foundations, enabling user profiles to be established in various scenarios and improving the flexibility and applicability of user profile initialization.

[0045] In one embodiment, for example, the system employs a combination of a unified adaptation interface specification and a plug-in extension architecture to achieve compatible integration with WhatsApp (an instant messaging platform). The system predefines a unified adaptation interface specification, which includes message reading interfaces, message sending interfaces, interface element location interfaces, event listening interfaces, and exception feedback interfaces. The WhatsApp plugin is developed according to the unified adaptation interface specification and encapsulates the platform's corresponding message reading logic, input invocation logic, sending trigger logic, and exception handling logic. After system startup, the plugin registration information is loaded, and the WhatsApp plugin is initialized, making it callable.

[0046] During the adaptation process, the system first acquires features from the WhatsApp Web page, including the conversation list area, message display area, input editing area, send icon, and page event response characteristics. These page features are then matched against preset platform feature templates. After matching, a platform adaptation configuration corresponding to WhatsApp is generated. This configuration includes DOM (Document Object Model) element positioning parameters, message extraction rules, input call parameters, send control parameters, and exception fallback parameters. The WhatsApp plugin then uses this platform adaptation configuration to read messages, locate interfaces, call input, and control send, thus adapting the message interaction process.

[0047] When a WhatsApp Web page version update is detected, causing existing element paths to become invalid, the system re-extracts page structure features and updates the corresponding parameter mapping relationships to regenerate the platform adaptation configuration. This allows plugins to regain their adaptability without rewriting the overall logic. The system can identify, match, and generate adaptation configurations for the target interaction environment under a unified interface specification, enabling upper-layer session processing logic to call the corresponding plugins to complete message interactions. This improves platform access efficiency and the stability of the adaptation process.

[0048] In the communication scenario of foreign trade in the construction machinery industry, users communicate with customers via WeChat or WhatsApp. The platform adaptation module loads the corresponding plugins according to the current interaction environment and completes the adaptation of conversation reading, reply input, and message sending based on the platform interface features. The data acquisition module acquires chat content, organizes and filters it, and sends the processing results to the intent analysis module. The intent analysis module performs semantic analysis on the conversation text to identify the customer's current purchasing needs and emotional tendencies. The first reply generation module calls the knowledge base in the construction machinery field based on the analysis results to generate candidate replies that match the current communication content. After selecting a candidate reply, the user can edit and send it according to the actual communication needs. The profile management module updates the customer profile by combining keyword information, reply processing data, and subsequent feedback content. The updated customer profile continues to participate in subsequent rounds of intent analysis and reply generation to improve the recognition accuracy and reply matching effect in subsequent communications.

[0049] In one embodiment, within a financial services scenario, users communicate with customers through a target interactive environment. The platform adaptation module acquires current platform characteristics and loads corresponding plugins, completing message reading, content input, and sending control adaptation. The data acquisition module acquires the chat content sent by the customer, cleans and structures the conversation data, and then sends the processing results to the intent analysis module. The intent analysis module analyzes the conversation content, identifying the types of financial products the customer is interested in, their risk preferences, and their current consultation focus. Based on the analysis results, the first response generation module extracts corresponding information from the financial product database and preset response templates, generating candidate responses that match the customer's needs; users can edit these responses before sending. The profile management module dynamically updates the customer profile based on keyword information and feedback from multiple rounds of communication. The updated user profile participates in subsequent rounds of intent recognition and response generation, enabling the system to continuously adapt to changes in customer needs and improve the relevance and continuity of financial service communication.

[0050] In one embodiment, such as Figure 2 As shown, the system receives conversation data from WeChat Work, WhatsApp (an instant messaging platform), and other interactive environments, and first enters the platform access layer. The platform access layer performs platform identification, anomaly monitoring, plugin loading, protocol translation, field mapping, interface positioning, and sending control on the target platform, enabling message reading, input calls, and sending operations from different platforms to be converted into a unified processing form. The data processed by the platform access layer enters the standardization interface layer, which organizes data and control information from different sources into a unified event structure, a unified control command structure, a unified anomaly feedback structure, and a unified message structure. After standardization, the data enters the data acquisition and preprocessing layer, where behavior acquisition, text acquisition, and voice acquisition units acquire corresponding types of data, and combine this with time series processing, encoding standardization, round-context concatenation, and structured message generation to form analyzable structured conversation content. Subsequently, the structured conversation content enters the intelligent analysis layer, where text recognition, intent extraction, topic recognition, emotion recognition, customer stage judgment, and risk signal recognition units complete the analysis and processing. The analysis results are transmitted to the knowledge and generation layer, where units such as knowledge base retrieval, terminology constraints, multi-style generation, candidate response sorting, and response result output generate potential answers. Simultaneously, they are transmitted to the profiling and feedback layer, where processing units such as preference change detection, silence time analysis, profiling tag updating, response strategy correction, and identification strategy correction update customer profiling and feedback optimization. The generated potential answers are finally output to the interactive display layer, where units such as response suggestion display, user editing, user selection and sending, and result recording complete the human-computer interaction, thus forming a complete closed-loop processing chain covering platform access, data analysis, response generation, profiling updates, and feedback optimization.

[0051] In one embodiment, such as Figure 3 As shown, after the system initiates the platform adaptation process, it first receives the platform type selected by the user and loads the corresponding platform recognition mechanism. Subsequently, the system collects environmental information of the target platform, including window features, window orientation features, message area layout features, input area position features, control hierarchy features, message extraction method features, interaction response features, and resolution and scaling ratio features. Based on the collected environmental information, the system forms a target platform feature dataset and matches it with pre-saved platform templates. Upon successful matching, the system generates a platform recognition result and further generates a platform adaptation configuration. The platform adaptation configuration includes field mapping rules, chat area positioning parameters, contact area positioning parameters, input box positioning parameters, send control trigger parameters, exception fallback parameters, and version compatibility parameters. Based on the platform adaptation configuration, the system establishes a message reading module, a send control module, an interface positioning module, and an input call module, and obtains the contact area recognition result, send control area recognition result, chat message area recognition result, and input box area recognition result to form a comprehensive platform adaptation result. During operation, the system continues to monitor changes in the platform interface. When it detects control malfunction, field malfunction, resolution change, version change, or layout change, it triggers an exception rollback and re-identification process. This process re-acquires platform feature data, updates platform adaptation configuration, restores the control positioning and message reading structure, and continues to perform adaptation operations. Through this process, compatible integration across different platforms, versions, interface layouts, and resolutions can be achieved without rewriting the overall plugin logic.

[0052] In one embodiment, such as Figure 4As shown, after the system starts the data acquisition module, it triggers data acquisition tasks according to a preset cycle to capture raw data in the chat area. The raw data includes text data, voice data, behavioral data, timestamp data, and conversation switching records. After acquisition, the system performs data preprocessing, including noise reduction, invalid segment filtering, encoding standardization, message segmentation, multi-turn context concatenation, and temporal alignment. The preprocessed data enters the text recognition and structured conversion stage. The system performs speech-to-text processing on the voice content, standardization processing on the text content, and further completes message role recognition, conversation object recognition, and structured message generation. Afterward, the system loads the industry corpus fine-tuning model and performs intent analysis and sentiment recognition on the structured messages. Intent analysis includes identifying inquiry intent, logistics consultation intent, after-sales communication intent, general consultation intent, and purchasing intent; sentiment recognition includes identifying positive, neutral, urgent, hesitant, and negative emotional states. Simultaneously, the system also performs keyword and topic extraction and customer stage judgment, and summarizes the various recognition results to form a comprehensive analysis result. The comprehensive analysis results further generate intent information, sentiment information, topic tags, and stage tags, which are then output to the response generation module and the profile management module, respectively. After each round of analysis, the system records the analysis results; when the initial data collection and analysis is completed, subsequent analysis processes are automatically triggered, thereby achieving continuous analysis and dynamic tracking of conversation content.

[0053] In one embodiment, such as Figure 5As shown, after receiving the analysis results, the system extracts sentiment information, customer stage tags, intent information, and topic tags. These are then combined with industry preference tags, historical response preferences, historical interest tags, and initiative frequency tags from the customer profile, as well as silent time records, to construct response generation conditions. Based on these conditions, the system accesses resources such as the consultation knowledge base, product parameter retrieval, business terminology retrieval, scenario-based dialogue retrieval, and frequently asked questions retrieval. It then generates candidate responses by combining friendly style parameters, length control parameters, risk constraint parameters, and professional style parameters. Candidate responses can include friendly style responses, concise style responses, in-depth explanation style responses, and professional style responses. After generating candidate responses, the system sorts them, considering style matching scores, user consistency scores, business constraint verification, and relevance scores. The sorting results are then output to the interactive interface. Users can browse candidate responses in the interactive interface, edit them, or directly select a response. The system then sends the response and records the sending result. After sending, the system updates the customer profile, including updated topics of interest, industry preferences, activity level, risk status, and response preferences. After the profile is updated, the system generates further warning results, including warnings of communication stagnation, weakened intent, and decreased interest. Simultaneously, the system generates feedback information, including identification correction information, strategy adjustment rules, response style adjustment parameters, and knowledge base retrieval priority adjustment information. Part of this feedback is sent to the intent analysis module, and the other part is sent to the response generation module, thus initiating the next round of closed-loop interaction to continuously optimize subsequent conversation analysis and response generation.

[0054] In one embodiment, the output of candidate responses includes: Based on the response processing data, determine the response preference characteristics corresponding to the conversation object; Based on the response preference characteristics, the matching degree is calculated for each of the at least two candidate responses generated; Based on the matching score calculation results, sort at least two candidate responses; Output at least two candidate responses based on the sorting results.

[0055] Among them, response preference features refer to the feature information extracted from historical response processing to characterize the conversation object's preference for response style, expression, content structure, or response tendency; matching degree calculation refers to the process of analyzing the correspondence between each candidate response and the response preference features to determine the degree of fit between each candidate response and the response preference features; ranking result refers to the order result formed according to the matching degree of each candidate response.

[0056] Specifically, by analyzing response processing data, response preference features corresponding to the conversation partner can be extracted from the actual response selection, modification, and sending. This allows response preferences to move beyond preset levels and establish a correspondence with the actual communication process. After obtaining the response preference features, the matching degree of multiple candidate responses is calculated separately, and the results are sorted. This ensures that the candidate responses output at the top are closer to the communication characteristics and usage needs of the conversation partner. This processing method can improve the targeting of candidate responses, reduce the cost for users to filter and adjust among multiple responses, and also improve the adaptability of candidate responses to the actual communication scenario.

[0057] In one embodiment, updating the user profile corresponding to the session object based on interaction samples includes: extracting information on the expression adjustment of the target candidate response content from the modified content; Based on the expression adjustment information, generate expression correction rules corresponding to the session object; Link expression correction rules to user profiles; When generating alternative or new alternative responses, the structure of the response is adjusted based on the expression correction rules.

[0058] Among them, expression adjustment information refers to information that reflects changes in expression when users modify the content of the target candidate response, including changes in wording, tone, content additions or deletions, or structural adjustments; expression correction rules refer to rule information summarized based on expression adjustment information, used to guide the adjustment of subsequent response expression; response expression structure refers to the structured expression form of the response content in terms of word choice, tone organization, sentence arrangement, content hierarchy, or expression order.

[0059] Specifically, by extracting expression adjustment information from the modified content, the user's expression habits and revision tendencies during the actual editing process can be preserved, enabling the user's modification behavior of the candidate response to be transformed into usable data. Based on this, expression correction rules are generated and linked to the user profile, allowing the user profile to not only reflect the needs of the conversation partner, but also to further reflect the user's modification habits in response expression. When subsequent candidate responses or new candidate responses are generated, the response expression structure is adjusted based on the expression correction rules, making the generated results closer to the user's actual expression style and reducing repeated modifications.

[0060] Get the input text.

[0061] In this embodiment, the input text can be actively collected or passively received.

[0062] Taking message interaction scenarios as an example, the raw text content input by the user can be collected to complete the acquisition of input text. The input text can be the text content directly entered by the user, or the text content converted from the user's voice message, video message, etc., or the text content obtained by converting the user's multimodal input data. There are no restrictions here.

[0063] The semantic analysis system is used to identify and filter redundant text in the input text to obtain key text, and the intent of the key text is determined to obtain the interaction intent; among them, redundant text represents text content that has semantic discontinuity with the target industry scenario.

[0064] In this embodiment, a pre-set or pre-trained semantic analysis system can be used to perform semantic screening on the complete input text. This identifies redundant text that may be semantically disconnected from the target industry scenario and has no practical interactive value. Redundant text is then filtered out, and relatively core and effective key text is extracted from the input text. It is easy to understand that redundant text can be considered as content that is detached from the business scenario, semantically fragmented, and invalid, which may pose a risk of interfering with the accuracy of subsequent interaction recognition. By relying on a semantic analysis system to perform semantic parsing and intent determination on the simplified key text, the accuracy of identifying the user's true interaction needs can be improved, and the reliability of the output interaction intent can be enhanced.

[0065] The emotion analysis model is used to identify the emotion of the target text to obtain the interactive emotion; the target text includes the input text and / or key text.

[0066] In this embodiment, either the input text or the key text, or a combination of both, can be selected as the input to the sentiment analysis model. At least one of the input text or the key text is used as the target text for sentiment analysis. The model leverages its sentiment recognition logic to capture emotional expression features within the text, thereby determining the user's emotional tendency and quantifying the user's current interactive emotions.

[0067] Generate intent tags that match the interaction intent, and generate emotion tags that match the interaction emotion.

[0068] In this embodiment, based on the identified interaction intent, corresponding intent tags can be automatically generated according to preset tagging specifications; simultaneously, based on the parsed interaction emotion, corresponding emotion tags are generated. This dual-tag annotation method distinguishes between the demand attributes and emotional attributes of the input text.

[0069] Therefore, this embodiment, at least when performing intent analysis on the input text, can identify text content within the input text that has semantic gaps with the target industry scenario, treating it as redundant text. It can be assumed that redundant text in the target industry scenario typically lacks substantial content, thus filtering it from the input text to obtain key text. This simplifies the text content to be identified and extracts effective information, thereby improving the reliability of intent analysis. Specifically, it can identify and filter interfering information within the input text that is inconsistent with the actual intent, thus enhancing the reliability of content analysis. Furthermore, using key text for sentiment analysis can also reduce noise interference. Simultaneously, sentiment analysis can be performed using the input text to reduce the risk of filtering out emotionally expressive words.

[0070] In one embodiment, input text can be obtained.

[0071] Taking the input text as an example of image recognition, the following is an example of the principle behind its acquisition.

[0072] It can acquire an image containing the input text to be recognized. For example, an image acquisition unit can capture a scene containing the text to be extracted, forming a complete image to be recognized, within which the input text can be contained.

[0073] This method can detect text-generating regions within an image that match regional features. These regional features include at least one of text box features, font features, language arrangement features, and background features. In other words, this embodiment can pre-configure various regional features adapted to the target recognition scenario, and rely on feature matching algorithms to perform a full-domain search and comparison of the image to be recognized, accurately locating text-generating regions containing text content within the image. It is easy to understand that features such as text box outlines, font styles, text layout, and background can all distinguish text regions from blank background regions, thereby achieving rapid and accurate text region locking and narrowing the processing scope for subsequent text recognition.

[0074] Furthermore, a pre-defined corpus of the target industry scenario can be used to assist in character recognition of the text region to obtain the initial text. The pre-defined corpus includes at least one of industry terminology, product information, and multilingual mixed text. Text metrics of the initial text can be evaluated, and these metrics can be used to filter the initial text to obtain the input text. These text metrics include at least one of the following: matching degree with the target industry scenario, semantic continuity of the initial text, and character confidence.

[0075] In other words, collaborative text recognition is performed on the identified text generation areas using a pre-built corpus corresponding to the target industry scenario. By leveraging industry-specific terminology, product-related materials, and multilingual text samples, the adaptability of text recognition is optimized, reducing errors in recognizing technical terms and special expressions, thereby extracting the original initial text content from the text generation areas. The generated initial text is then comprehensively evaluated using multi-dimensional metrics, verifying the text's relevance to the target industry scenario, the semantic coherence of the sentence context, and the reliability of single-character recognition. Based on the evaluation results of these text metrics, invalid content with recognition errors, semantic breaks, and poor industry adaptability can be relatively effectively eliminated, purifying the initial text and ultimately yielding standardized input text suitable for subsequent analysis and processing.

[0076] In response to the acquisition of input text, a semantic analysis system can be used to identify whether the input text contains redundant text.

[0077] In response to the semantic analysis system failing to identify redundant text, the semantic analysis system is used to determine the intent of the input text to obtain the interaction intent. The principle of determining the intent of the input text is similar to that of determining the intent of the key text, which can be found in the later section on the determination principle of the intent of the key text, and will not be repeated here.

[0078] In response to the semantic analysis system identifying redundant text, the redundant text is filtered out.

[0079] Optionally, the semantic analysis system includes a semantic analysis model. Input text can be fed into the semantic analysis model. The model identifies industry keywords related to the target industry scenario within the input text. Semantic analysis is performed on the input text to determine its semantic continuity. Text content with semantic breaks from the industry keywords is considered redundant text. Redundant text is then removed from the input text to obtain the key text. Here, redundant text refers to text content with semantic breaks from the target industry scenario.

[0080] In other words, a semantic analysis system can use a trained semantic analysis model as a processing unit. Input text is fed into a sentiment analysis model, which then identifies industry-specific keywords relevant to the target industry scenario. Based on keyword relationships, the system assesses the contextual coherence of the entire text, identifying redundant text fragments that are logically disconnected from industry keywords or semantically fragmented. Redundant text is then removed to retain only the valid information fragments in the input text, thereby improving the conciseness and semantic integrity of the key text used for intent recognition.

[0081] Optionally, the semantic analysis system may also include an intent output layer and a joint discriminator.

[0082] Key text can be transmitted to the intent output layer and the joint discriminator. The intent output layer, combined with industry keywords within the key text, evaluates and outputs the probability distribution of the key text belonging to various preset intents. The joint discriminator extracts industry keywords from the key text to match keyword information associated with those keywords in the target industry scenario. Finally, the preset intent, based on a combination of keyword information and probability distribution, is selected as the interaction intent.

[0083] In layman's terms, the semantic analysis system in this embodiment integrates an intent output layer and a joint discriminator. It synchronously distributes the extracted key text to both the intent output layer and the joint discriminator for collaborative discrimination. The intent output layer combines industry keywords from the key text with preset interactive intents (i.e., preset intents), quantifying the probability of the key text matching each intent to form a probability distribution of the input text containing each preset intent. The joint discriminator extracts industry keywords from the key text, retrieves industry databases to match keyword association information, and clarifies the scenario and business attributes corresponding to the keywords. Combining keyword association information and intent probability distribution as references, the system selects the preset intent with the highest matching degree from multiple preset intents and determines it as the interactive intent of the input text and the key text.

[0084] Specifically, the interaction categories represented by keyword information can be identified, and the confidence level of the key text belonging to each preset intent can be evaluated by combining the interaction categories. The interaction categories include at least one of the following: purchasing, inquiry, and order constraint categories. Based on the combined probability distribution and category confidence level, the preset intent is selected as the interaction intent. Furthermore, the matched keyword information can be categorized and analyzed to classify different interaction categories such as purchasing, inquiry, and order constraints. Based on the interaction category to which the key text belongs, the category confidence level of its corresponding preset intent is specifically calculated, supplementing the criteria for category dimension determination. Combining intent probability distribution and category confidence level enables multi-dimensional verification to improve the comprehensiveness and accuracy of verification, effectively reducing intent recognition bias and thus improving the accuracy of identifying the user's true needs through interaction intent.

[0085] Furthermore, in this embodiment, a sentiment analysis model can be used to identify the emotion of the target text to obtain the interactive emotion. The target text includes the input text and / or key text.

[0086] Specifically, the target text can be input into a sentiment analysis model. The model then identifies the sentiment features of the target text to obtain the current sentiment represented by the text. Finally, a smoothing process is applied to the current sentiment to obtain the interactive sentiment.

[0087] The target text is input into a preset or pre-trained sentiment analysis model. The model captures multi-dimensional information such as word usage features, emotional vocabulary, and sentence tone to extract the emotional tendency and characteristics within the target text, thus initially identifying the current emotion corresponding to the target text. Simultaneously, to reduce the risk of incomplete emotional expression in a single sentence or excessive fluctuations in emotion recognition affecting the reliability of emotion determination, this embodiment can also perform smoothing correction and optimization calibration on the initially acquired current emotion. This weakens the interference of accidental and fragmented emotions, improves the stability and reliability of the determined interactive emotions, and thus helps to improve the completeness of restoring the user's true emotional state during instant messaging interactions.

[0088] The following provides a detailed explanation of the principles behind smoothing correction processing.

[0089] The current emotion can be corrected by incorporating textual factors. And / or, it can be corrected by incorporating auditory features. That is, smooth correction can adopt a single correction method or a multi-dimensional composite correction method. It can complete emotion correction based on the text's own correlation features, or it can combine the acoustic features of the original speech to assist in emotion correction. The two correction methods can be implemented independently or used in combination, and it can adapt to target texts from different source types.

[0090] In detail, the sentiment type represented by the textual factors of the historical input text and the input text can be analyzed to obtain text correction factors. These textual factors include at least one of the following: the time interval between the input text and the historical input text, the trend of historical sentiment notes, semantic transition structures, and the trend of text content changes. The current sentiment is then corrected using these text correction factors.

[0091] In this embodiment, historical input text corresponding to the current text interaction is retrieved and compared with the historical text, comprehensively considering the emotional change patterns reflected by multiple text factors. Specifically, the interaction time interval reflects the duration of the user's emotions, the trend of historical emotion tag changes reflects the user's long-term emotional tendencies, semantic transition structures are used to identify sudden emotional shifts caused by sentence transitions, and the trend of text content changes reflects emotional fluctuations caused by changes in interaction requests. Through comprehensive analysis and quantitative calculation of various text factors, a text correction factor adapted to the current interaction scenario is generated. Based on this correction factor, the initial output of the model's current emotion is weighted and smoothed, weakening occasional extreme emotional expressions in single sentences, which helps improve the fit between the emotion recognition results and the user's overall interaction state.

[0092] This method can respond to input text converted from speech messages and analyze the emotion type represented by the sound features of the speech message to obtain a speech correction factor. The sound features include at least one of speech rate features, pause features, and pitch features. The speech correction factor is then used to correct the current emotion.

[0093] In this embodiment, when the input text originates from the transcription of a voice message, the original voice data can be traced back to extract multi-dimensional sound features such as speech rate, pause rhythm, and pitch fluctuations. These features are then combined with the emotional expression patterns corresponding to various acoustic features to determine the true emotional type carried at the speech level, and a voice correction factor is calculated accordingly. While text recognition only reflects the literal emotion, the voice correction factor can supplement the emotional differences brought about by tone and intonation. In this embodiment, the voice correction factor is applied to the current emotion for secondary calibration, which can compensate for the limitations of pure text emotion recognition and further improve the comprehensiveness and realism of interactive emotion recognition.

[0094] Generate intent tags that match the interaction intent, and generate emotion tags that match the interaction emotion.

[0095] Interaction intent and / or interaction emotion can be used as intent tags and / or emotion tags themselves, or they can be obtained through secondary processing such as formatting. No limitation is made here.

[0096] The training principle of a pre-trained semantic analysis model will be illustrated below with examples.

[0097] Optionally, training samples labeled with industry keywords for at least one interaction stage can be obtained to train a semantic analysis model. Industry keywords include at least one of product keywords, industry terminology keywords, and business keywords; interaction stages include consultation stages, transaction stages, after-sales stages, and other stages.

[0098] In other words, actual interaction data can be collected and organized in advance, and training samples can be selected to form multiple interaction stages, and industry keywords can be labeled for them. It's easy to understand that the text expression habits and content emphasis differ significantly across different interaction stages. Coupled with refined labeling of product keywords, industry terminology keywords, and business keywords, the learning dimensions of the model can be enriched. Inputting well-labeled training samples into the model training process allows the semantic analysis model to fully learn the text semantic rules, keyword association logic, and contextual expression features of different interaction stages in the target industry scenario. This strengthens the model's ability to understand industry-specific content and improves its performance in recognizing industry keywords, judging semantic continuity, and distinguishing redundant text.

[0099] Optionally, at least one type of perturbation sample can be obtained as validation samples to be used for model validation and iterative optimization of the trained semantic analysis model. The perturbation sample types include sample confusion category types, industry-specific near-synonymous expression sample types for the target industry scenario, sample classification bias types, and questionable sample types, where the questionable sample type indicates that the sample confidence level is below a confidence threshold.

[0100] In other words, after completing the basic training, multiple types of validation sample sets can be further constructed, including validation samples of at least one type of perturbation sample. Perturbation samples with high difficulty and susceptibility to interference are specifically introduced. Among them, confusion category samples can easily cause confusion in the model's intent judgment; industry-specific synonymous expression samples can test the model's ability to distinguish between synonymous but different expressions; classification bias samples are used to identify common recognition vulnerabilities in the model; and questionable samples with insufficient confidence can cover special text scenarios with ambiguous boundaries. By using the above-mentioned multi-dimensional perturbation validation samples to test and verify the trained semantic analysis model, it is beneficial to discover the defects and shortcomings of the semantic analysis system / model in complex contexts, similar expressions, and boundary text recognition, and to iteratively update and optimize them. Adjusting model parameters and optimizing semantic discrimination rules based on the validation feedback results can enhance the semantic analysis model's anti-interference ability and generalization ability in complex real-world interaction scenarios, and also improve the long-term stability and accuracy of the semantic analysis model and system in completing text semantic parsing and redundant content recognition.

[0101] The training principle of a pre-trained sentiment analysis model will be illustrated below with examples.

[0102] We can acquire text corpora carrying pre-defined emotion labels as emotion samples. We then convert these emotion samples into word vectors and extract local emotion features from these word vectors.

[0103] Local emotion features are pooled to obtain discriminative features for emotion classification using predefined emotion labels. A fully connected layer evaluates the probability distribution of these discriminative features belonging to each predefined emotion, and the predefined emotion to which the emotion sample belongs is determined by adapting the probability distribution to the training emotion. The emotion analysis model is updated using the training loss corresponding to the training emotion and the predefined emotion labels.

[0104] In layman's terms, this approach allows for the batch collection of massive amounts of text corpora from instant messaging interactions and target industry scenarios. Standardized sentiment annotation is then applied to these collected text corpora, assigning each piece of text a unique or composite pre-defined sentiment label. This constructs a rich and well-annotated dataset of sentiment samples, serving as a supervised training data source for the sentiment analysis model. The selected sentiment samples undergo text digitization, and a pre-trained word embedding algorithm maps the natural language text content into uniformly dimensional word vectors to quantify the semantic information of words and sentences. Subsequently, the sentiment analysis model can analyze the continuously arranged word vectors layer by layer, extracting key information such as emotional tendencies and tone expressions contained in word combinations and short sentence contexts, thereby improving the accuracy of extracting local sentiment features scattered throughout text fragments.

[0105] For the extracted fragmented local emotion features, pooling operations are used to aggregate and denoise the features, which helps to remove redundant feature information and compress feature dimensions. This facilitates the integration of semantically complete and feature-focused discriminative features, making them suitable for subsequent multi-category emotion classification. The integrated discriminative features are then input into a fully connected layer for feature mapping and weight calculation, quantifying the probability distribution of the discriminative features matching various preset emotions. Based on the numerical proportion of the probability distribution and the matching priority, the preset emotion with the highest confidence is selected as the training emotion for the model's inference output.

[0106] Furthermore, the training emotions inferred by the sentiment analysis model are compared with the pre-labeled sentiment tags of the emotion samples. The deviation between the two is calculated using loss functions such as cross-entropy loss, generating the corresponding training loss. Using the training loss as the basis for backpropagation, the model parameters of the sentiment analysis model are updated. This helps to reduce the error between the model's predictions and the standard labeled data, thereby strengthening the model's ability to capture text sentiment features and improve classification accuracy. Ultimately, this enhances the stability and generalization performance of the sentiment analysis model in complex interactive text scenarios.

[0107] In one embodiment, image updating includes: Identify the user's interaction data in the current round, and obtain the single-round preset words and single-round tags for the current round. The single-round preset words are used to represent the semantic units extracted from the interaction data.

[0108] This involves acquiring real-time chat data between users and staff based on communication platforms, such as social media or instant messaging software. The system can also obtain raw chat logs generated on the platform, i.e., interaction data, which includes text and images. Images can be converted into analyzable text using Optical Character Recognition (OCR). The interaction data from the current round with the user represents the information generated between the user and the trade support tool in this dialogue, including: the content entered by the user, such as text, commands, or selected options; and the content of responses to the user, such as answers, hints, follow-up questions, or guiding statements. This interaction data serves as the foundation for subsequent backend analysis. The specific communication platform can be determined based on the actual usage scenario or user habits.

[0109] The pre-set words for the current round are used to represent the basic semantic units extracted from the interaction data with the user in the current round. They can be the recognition results obtained after recognizing pre-set keywords, fixed feature words, and rule words during the interaction with the user in the current round. The single-round label is used to represent the feature identifier for classifying users. It can be temporary user labels identified based on the current single-round interaction data.

[0110] Specifically, if keywords such as "excavator" and "overseas promotion" appear multiple times in the user's interaction data in the current round, it can be determined that the preset words for the single round include "excavator" and "overseas promotion". If the preset words for the single round (excavator, overseas promotion) have a strong semantic association with the export trade tags of construction machinery products in the tag knowledge graph, then the user's single round tags include the export trade tags of construction machinery products.

[0111] Based on single-turn preset words, single-turn tags, and historical interaction data with users, the state data of the dialogue state machine is determined. The dialogue state machine is determined based on a finite state machine model, and the state data is used to represent the state switching logic.

[0112] The user's historical interaction data represents past dialogues or interactions with that user, including historical keywords, historical tags, historical state records, and historical profile data. The dialogue state machine is determined based on a Finite State Machine (FSM) model and includes multiple states, with pre-defined state transition conditions between each state. Specifically, the dialogue state machine can be configured to include an initial state, intermediate states, and a termination state. The initial state represents the state at the beginning of the session; the intermediate states represent the state from the beginning of the session until its end; and the termination state represents the state at the end of the session. Pre-defined single-turn words, single-turn tags, and historical interaction data can be matched with the state transition conditions of the dialogue state machine, and the state switching logic of the dialogue state machine is determined based on the matching results. The state data of the dialogue state machine includes: the current state of the state machine, the historical state transition trajectory, and the state transition trigger conditions.

[0113] Determining single-turn tags solely based on current user interaction data is susceptible to user input errors or single-item ambiguities; relying solely on historical interaction data may fail to reflect changes in user intent. A better approach is to integrate current and historical interaction data. Based on this integrated data, determine whether the single-turn tag continues the previous dialogue theme, whether the core requirement has changed, or whether previously unspoken information has been added. Matching these results with the state transition rules of the dialogue state machine determines the corresponding state switching logic for the current dialogue turn.

[0114] Specifically, the current conversation state can be determined based on pre-set single-turn words, single-turn tags, and historical interaction data with the user. Based on the historical and current conversation states, the state data of the dialogue state machine can then be determined. For example, if a user first asks, "Hello, I'd like to learn about excavator promotions," the state data of the dialogue state machine is determined to be in the initial state. If the user then asks again, "Promotions for a certain excavator," the state data of the dialogue state machine is determined to be in the initial state, transitioning to an intermediate state. If the user asks, "Okay, I'll contact you again if I need anything later," the state data of the dialogue state machine is determined to be in the intermediate state, transitioning to the terminated state.

[0115] Identify single-wheel labels that have semantic conflicts and / or that do not conform to the state switching logic represented by the state data, and treat them as anomalous labels.

[0116] Semantic conflict is used to characterize logically mutually exclusive combinations between single-round labels, such as a refund intention label and a satisfaction label, or a high purchase intention label and a abandonment of cooperation label. Single-round labels that do not conform to the state transition logic represented by the state data are used to characterize labels that contradict the state transition logic of the dialogue state machine. Anomaly labels are used to characterize temporary labels within a single round that contain internal semantic contradictions or contradict the state transition logic of the state machine.

[0117] It can determine whether there are semantically conflicting tags, tags that contradict the state transition logic of the dialogue state machine, or tags that are both semantically conflicting and contradict the state transition logic of the dialogue state machine. If there are tags that are semantically conflicting or contradict the state transition logic of the dialogue state machine, these tags are judged as abnormal tags.

[0118] Specifically, if the user sends "Hello" for the first time, the state data of the dialogue state machine is determined to be the initial state. If there is a session end label in the single-round label of the current round, the session end label does not conform to the state switching logic represented by the state data, and the session end label is regarded as an abnormal label.

[0119] Corrected labels are obtained by correcting abnormal labels based on historical interaction data. Based on single-round preset words and corrected labels, the user profile and the state of the dialogue state machine are updated.

[0120] Among them, the corrected label is used to represent the compliant label obtained after adjusting or replacing the confidence level of abnormal labels based on historical interaction data. The user profile is used to represent the user feature model including labels, keywords, and weights.

[0121] For anomalous tags, confidence adjustments, conflicting tag replacements, and invalid tag removals can be performed based on historical interaction data to resolve misjudgments and contradictions in single-round recognition. Based on reliable single-round preset words and corrected tags, user profile tags, intention weights, and behavioral preferences can be iterated synchronously. Simultaneously, based on the single-round preset words and corrected tags, the dialogue state machine is driven to update its state, awaiting the next round of interaction recognition.

[0122] The user profile update method provided in this embodiment determines keywords and tags to be determined based on single-round interactions, determines the state constraints of the dialogue state machine based on single-round and historical interaction data, filters abnormal tags based on tag semantic verification and state machine state constraints, corrects abnormal tags based on historical interaction data, and updates the user profile and dialogue state machine state data based on keywords and corrected tags. Using each round of user interaction data as a trigger condition, it identifies single-round preset words and tags in real time, no longer relying on long-term fixed static tags. Each round of user dialogue initiates a completely new round of semantic capture, responding to real-time expressions such as budget adjustments, interest switching, and intent changes, solving the problem of untimely user profile updates. This reduces contradictory tags identified in single-round interactions, lowers intent misjudgment, and improves the accuracy of tag recognition. Simultaneously, based on the state constraints of the dialogue state machine, it ensures that user tags are consistent with the conversation evolution logic, conforming to the real interaction process, and enabling intelligent interaction strategies to dynamically adapt to user needs, thus improving the user experience.

[0123] In some optional implementations, obtaining the single-round preset words and single-round tags for the current round includes: identifying at least one preset word contained in the interaction data between the user and the current round as the single-round preset word; querying all target tags corresponding to the preset words in the tag knowledge graph and linking the preset words with the target tags; extracting a first semantic vector based on the preset words and extracting a second semantic vector based on the target tags, and determining the similarity between the first semantic vector and the second semantic vector; if the similarity is higher than a preset similarity threshold, adding the target tags to the single-round tags, wherein the number of target tags contained in the single-round tags is at least one.

[0124] The tag knowledge graph can be a pre-constructed knowledge system containing hierarchical tags and semantic relationships, used to map identified preset words (keywords) to tags in user profiles. Entities represent semantic objects with clear business meanings corresponding to preset words. Entity links match and bind the identified preset words (keywords) with corresponding standard tag nodes in the tag knowledge graph, unifying diverse colloquial user expressions to semantic entities within the tag knowledge graph and avoiding tag matching errors caused by different meanings of the same word or different words with the same meaning. The first semantic vector represents the vector representation obtained after semantically vectorizing the preset words. The second semantic vector represents the vector representation obtained after semantically vectorizing the corresponding target tag in the tag knowledge graph. Similarity represents the semantic closeness between word vectors and tag vectors. The preset similarity threshold can be a pre-set minimum score for judging whether a preset word matches a tag. The target tag represents the tag in the tag knowledge graph that semantically matches the preset words.

[0125] In this implementation, a tag knowledge graph can be created based on a pre-defined multi-dimensional tag system. User tags can be hierarchically linked according to core tags, sub-tags, associated attributes, and business entities, and the functional semantic edges of synonyms, near-synonyms, hierarchical relationships, causal relationships, and subordinate relationships between user tags can be labeled. Pre-defined words are mapped to the standard tags already defined in the tag knowledge graph, completing the link from the pre-defined words to the tag system. Here, the pre-defined words are keywords, which are the core basis for generating tags, and the tags are the abstraction and classification of the semantic meaning of the keywords. Subsequently, through the repeated occurrence of keywords and semantic association verification, the accurate labeling of user profiles is completed.

[0126] Specifically, if the user's interaction data in the current round is "I want to inquire about the device's delivery time," the identified preset words are "device," "delivery," and "delivery time." The retrieved target tags in the tag knowledge graph are: for "device," "product inquiry" and "goods intention"; for "delivery," "logistics inquiry" and "delivery timeliness inquiry"; and for "delivery time," "timeliness inquiry." The device is then linked to the device entity in the tag knowledge graph; the delivery time is linked to the delivery timeliness entity in the tag knowledge graph.

[0127] Both the preset words and their corresponding target tags are converted into semantic vectors. The preset words are then projected onto tag knowledge graph nodes, and their similarity is calculated to determine whether their semantics are truly consistent. Tag nodes with a semantic relevance higher than a preset similarity threshold are identified as target tags, thus completing the accurate matching between the preset words and existing user tags.

[0128] The interaction data between the user and the current user session can also be input into the intent analysis model to obtain single-round labels. The intent analysis model matches the corresponding category in a multi-dimensional labeling system based on the semantics of the chat text, outputting labels related to the user's current needs, emotions, and behaviors. The intent analysis model can use a Bidirectional Encoder Representations from Transformers (BERT) model, an optimized version of the original BERT model, pre-trained and fine-tuned for specialized vocabulary in trade scenarios (such as overseas promotion and cross-border logistics) and industry terminology to improve the accuracy of intent recognition and keyword extraction. Alternatively, an improved pre-trained language model can be used, such as the Robustly Optimized BERT Pretraining Approach (RoBERTa), or a lightweight BERT model for self-supervised learning of language representations (ALBERT). In this way, by pre-selecting words, linking entities in the tag knowledge graph, and verifying semantic vector similarity, the system can accurately and automatically generate standard user tags from user interaction text. Compared with related technologies that directly map keywords, the user tags are more standardized, more accurate, and more resistant to interference, and can provide a high-quality tag source for dynamic user profiles in real time and stably.

[0129] In some optional implementations, the dialogue state machine includes an initial state, an intermediate state, and a termination state; single-turn labels include intent labels and / or emotion labels. Identifying single-turn labels with semantic conflicts and / or those that do not conform to the state switching logic represented by the state data as abnormal labels includes: if a single-turn label does not conform to the state switching logic represented by the state data, determining that the single-turn label is an abnormal label, wherein the state switching logic includes switching from the initial state to the intermediate state, switching from the intermediate state to the termination state, the initial state being used to represent the state of establishing a conversation with the user for the first time, the intermediate state being used to represent the state from the establishment of the conversation to the end of the conversation, and the termination state being used to represent the state of the end of the conversation; and / or combining the intent labels and emotion labels in the single-turn labels to obtain at least one combination, if the semantic similarity between the intent labels and emotion labels in the target combination is lower than a first similarity threshold, determining that the intent labels and emotion labels in the target combination have semantic conflicts, and identifying the intent labels and emotion labels in the target combination as abnormal labels.

[0130] Specifically, if a conversation with a user has just begun, and the user has not clearly expressed a valid need, nor exhibited a fixed intention or stable emotion, then the dialogue state machine is in the initial state. If the conversation is ongoing, and the user continues to inquire, their intention or emotion may iteratively change, then the dialogue state machine is in an intermediate state. If the conversation ends, such as when a need is confirmed, a request is fulfilled, or communication concludes and the intention no longer changes, then the dialogue state machine is in the terminated state. The state transition logic is used to represent the preset, legal state transition paths.

[0131] In this implementation, the states of the dialogue state machine are defined as an initial state, an intermediate state, and a termination state. The switching path of the dialogue state machine is set: from the start of the interaction process to its end. Based on a specific tag in the current single round, the existing states of the dialogue state machine are compared with the allowed transition logic. If the intent corresponding to that tag does not conform to the content that can occur in the current stage (e.g., a conversation end tag appearing in the initial state), the tag is determined to be an abnormal tag.

[0132] The intent tags and emotion tags identified in this round are paired, and the semantic vectors of the intent tags and emotion tags within each pair are extracted. The semantic correlation between the intent tag vector and the emotion tag vector within the pair is calculated using a cosine similarity algorithm to obtain their semantic similarity values. A first similarity threshold is pre-configured. If the semantic similarity between the intent tag vector and the emotion tag vector in the target pair is lower than this threshold, the intent tag and emotion tag in the target pair are determined to be abnormal tags.

[0133] In this way, abnormal tags with temporal anomalies or logical conflicts can be accurately identified, ensuring that the user profile updated based on tags always matches the user's real and latest needs, which facilitates the execution of subsequent intelligent recommendation and interactive response strategies.

[0134] In some optional implementations, intermediate states may include a main intermediate state and sub-intermediate states. The main intermediate state is used to characterize the global session process state, maintaining a forward flow logic with the initial and final states. Sub-intermediate states can be determined based on business intent. For example, sub-intermediate states may include consultation states, intention states, and decision states; consultation states may be product consultation states, price consultation states, order consultation states, or logistics consultation states; intention states may be order placement intention states or negotiation states; decision states may be order confirmation states, refund request states, or order cancellation states.

[0135] If a single-round label includes multiple labels, the switching or parallel execution of the corresponding sub-intermediate states is triggered based on these multiple labels. Specifically, if a single-round label includes multiple labels, the sub-intermediate states corresponding to these multiple labels can be triggered in parallel. For example, if a single-round label includes a "Consult Excavator Export Price" label and a "Inquire about Refund Policy" label, the "Price Inquiry" sub-state corresponding to the "Consult Excavator Export Price" label and the "Refund Inquiry" sub-state corresponding to the "Inquire about Refund Policy" label can be triggered simultaneously. The aforementioned state switching logic also includes switching from the first sub-intermediate state to the second sub-intermediate state, or from the second sub-intermediate state to the first sub-intermediate state. Sub-intermediate state switching does not affect the main intermediate state. Sub-intermediate states support free switching and do not need to follow a fixed order.

[0136] If the dialogue state machine is in the main intermediate state, and if the single-round label is a normal label, combine the intent label and emotion label in the single-round label. If there is a semantic conflict in the combination result, determine that the intent label and emotion label in the combination result are abnormal labels.

[0137] Specifically, if the main intermediate state of the dialogue state machine is in the engineering machinery interaction, and the sub-intermediate state is the procurement consultation sub-state, the single-round labels include: refund intention label, consultation export customs declaration process label, and concern emotion label; if the refund consultation sub-state corresponding to the refund intention label and the customs declaration consultation sub-state corresponding to the consultation export customs declaration process label are triggered, then the sub-intermediate state of the dialogue state machine will be updated to include: procurement consultation sub-state, refund consultation sub-state, and customs declaration consultation sub-state.

[0138] In this way, by using a hierarchical design of main intermediate states and sub-intermediate states, the rigid constraints of the original single intermediate state are broken, allowing users to express multiple intentions in parallel and switch between intentions freely when the dialogue state machine is in an intermediate state. This aligns with users' real communication habits and reduces the probability of interaction interruption caused by system judgment anomalies.

[0139] In some optional implementations, correcting abnormal labels based on historical interaction data to obtain corrected labels includes: evaluating a set of confidence weights for abnormal labels; wherein the set of confidence weights includes at least one of historical label confidence weights, dialogue round confidence weights, and semantic matching weights; the historical label confidence weights are determined based on the average confidence of labels of the same type as the abnormal labels in the historical interaction data; within the maximum range of the dialogue round confidence weights, the dialogue round confidence weights increase with the increase of the dialogue rounds; the semantic matching weights are determined based on the semantic similarity between the abnormal labels and historical labels; the current confidence of the abnormal labels is determined based on the semantic similarity between the abnormal labels and the single-round preset words corresponding to the abnormal labels; the corrected confidence is obtained by multiplying the weighted sum of the historical label confidence weights, dialogue round confidence weights, and semantic matching weights with the current confidence; and the abnormal labels with corrected confidence are used as corrected labels.

[0140] The current confidence score (ranging from 0 to 1) characterizes the current confidence score of the anomalous label, reflecting the initial reliability of the single-round identification result. It can be determined based on the semantic similarity between the anomalous label and the corresponding pre-defined words in the single round. The confidence weight set is used to adjust the initial confidence score of the anomalous label and calculate the weighted parameters for the corrected confidence score. It can optimize label confidence by combining historical data, ensuring that the corrected result aligns with the user's true intent.

[0141] In this embodiment, the initial confidence level of the abnormal label can be used as a basis, combined with at least one confidence level weight (historical label confidence level weight, dialogue round confidence level weight, semantic matching degree weight), and a preset weighting algorithm can be used to calculate the corrected confidence level by calculating the weighted sum of the corrected confidence level = initial confidence level × confidence level weight, thereby completing the optimization and adjustment of the credibility of the abnormal label.

[0142] Specifically, you can retrieve tags of the same type as the current abnormal tag from the user's historical interaction data and calculate the average confidence score of these historical tags. For example, if the confidence scores of three historical intent tags of the same type are 0.8, 0.75, and 0.85 respectively, and the average is 0.8, this average is the historical tag confidence score weight, which ranges from 0 to 1. The higher the average, the higher the credibility of the user's historical tags of the same type, and the greater the reference value for correcting the current abnormal tag.

[0143] The maximum value of this weight can be preset, such as 0.5. Following the rule of increasing weight with each round, the weight increases as the current dialogue round progresses, but the weight value never exceeds the preset maximum value. In this way, the user's intent becomes clearer as the dialogue round progresses, and the reliability of tag recognition increases.

[0144] The semantic vectorization algorithm can be used to convert abnormal tags and user's historical tags of the same type into semantic vectors respectively, and calculate the semantic similarity between the two. This similarity is the semantic matching weight, with a value range of 0-1. The higher the similarity, the more the abnormal tag matches the user's historical intent and behavioral habits, and the more likely the tag will be retained when making corrections.

[0145] In this way, by using historical tag confidence weights and semantic matching weights, the correction of abnormal tags is deeply bound to users' long-term interaction habits, avoiding correction bias caused by relying solely on single-round recognition results, and ensuring that the corrected tags conform to the user's past behavioral logic. At the same time, the confidence weight of each dialogue round follows the logic that the weight increases as the round progresses, which is consistent with the pattern that user intent gradually becomes clearer as the rounds progress in real interactions, avoiding the miscorrection of tags with clear intent in later rounds, and improving the temporal rationality of the correction results.

[0146] In some optional implementations, the user profile is updated based on a single round of preset words and corrected labels, including: if the corrected confidence level is greater than or equal to a preset confidence threshold, the confidence level of the label corresponding to the corrected confidence level in the user profile is adjusted to the corrected confidence level; if the corrected confidence level is less than the preset confidence threshold, the abnormal label is deleted from the user profile.

[0147] In this embodiment, a preset confidence threshold is used to characterize the preset label confidence judgment standard, such as 0.6. When an abnormal label is corrected and its corrected confidence level reaches or exceeds this threshold, it indicates that the corrected label has high credibility and matches the user's true intention. At this time, based on the corrected confidence level, the corresponding type of label and its weight in the user profile are updated synchronously. In this way, the higher the corrected confidence level, the higher the weight of the corresponding label in the user profile, strengthening the influence of the label on the user profile and achieving accurate iteration of the profile.

[0148] When the corrected confidence level of an anomaly label is lower than the preset confidence threshold, it means that even after correction, the anomaly label still lacks sufficient confidence and cannot be used to update the user profile; the anomaly label can be deleted.

[0149] In this way, by setting a pre-set confidence threshold, the corrected tags are filtered a second time. Only high-confidence tags are used to update the tags in the profile, while low-confidence tags are deleted. This reduces the number of abnormal or low-confidence tags entering the user profile, avoids profile distortion, and ensures that the profile always matches the user's true intentions.

[0150] User profiling in related technologies does not consider the chat time dimension and ignores the impact of the silent period on customer interest, which can lead to inaccurate timing of follow-up.

[0151] In some optional implementations, the aforementioned user profile update method further includes: obtaining the last interaction time with the user, determining the interaction termination duration based on the last interaction time; determining relevant tags based on tags whose tag frequency is greater than or equal to a preset frequency threshold and / or whose tag weight is greater than or equal to a weight threshold in a preset number of rounds prior to the last interaction time; and if the interaction termination duration exceeds a preset feedback cycle threshold, reducing the weight of the relevant tags and / or adding a tag to be contacted for the user.

[0152] In this embodiment, the last valid interaction time with the user recorded in the trade support tool can be obtained, such as the timestamp of the last conversation sent by the user that triggered the interaction operation; then the time difference between the current system time and the last interaction time is calculated, and this time difference is the duration of the user's interaction interruption, which is used to measure the duration of the user's silent interaction.

[0153] The preset rounds are a pre-defined range of historical interaction rounds used to filter relevant tags, such as the five rounds of dialogue before the last interaction. User tag data within this preset round is retrieved, and relevant tags are determined using at least one of the following methods: First, the frequency of each tag is counted; tags with higher frequencies are more relevant to the user's core needs, for example, tags with a frequency greater than or equal to a preset tag threshold. Second, the weight of each tag is considered; tags with higher weights better reflect the user's core intent, for example, tags with a weight greater than or equal to a weight threshold. The tags ultimately selected that are related to the user's core needs or intent are the relevant tags that need adjustment.

[0154] If the interaction pause duration exceeds a preset feedback cycle threshold, the weight of relevant tags is reduced, or a "to be contacted" tag is added to the user, or both are done simultaneously. The preset feedback cycle threshold is a pre-defined critical duration for determining user inactivity, such as 7 days. When a user's interaction pause duration exceeds this threshold, it indicates a decrease in user engagement and a potential change in needs. At this point, at least one of two adjustment operations is performed: first, the weight of relevant tags is reduced to weaken their impact on the user profile, such as lowering the tag weight from 0.8 to 0.4, preventing outdated tags from dominating the user profile; second, a "to be contacted" tag is added to the user, such as a "silent user" tag or a "needs confirmation pending" tag, to mark the user's current status and provide a clear indicator for subsequent interaction strategy adjustments. For example, if a user's inactivity exceeds 48 hours, sales personnel are automatically alerted to the risk of customer churn.

[0155] In one embodiment, response generation includes: Acquire and analyze user input text to obtain intent information, sentiment information, and product keywords.

[0156] User input text can be obtained through the chat window of communication tools (such as WhatsApp, WeChat Work, etc.).

[0157] To ensure the accuracy and integrity of the text, the screen can be periodically captured, and the captured images can be used for text recognition to obtain the user's input text. For example, the screen image of the chat area can be captured at a preset period (such as 5 seconds), and irrelevant characters in the image (such as meaningless text generated by input errors, advertising logos, etc.) can be filtered out through a multi-threshold noise reduction algorithm. Then, a high-precision text recognition engine can be used to extract the user's input text.

[0158] After obtaining user input text, the user input text is fed into a language model to determine product keywords and intent information; the user input text is then fed into a convolutional neural network to perform emotion recognition and obtain emotion information.

[0159] For example, the language model could be the BERT model, specifically bert-base-chinese. The language model includes embedding layers, encoding layers, classification layers, and sequence labeling layers. This model can contain a 12-layer Transformer encoder, 768 hidden layers, 12 attention heads, and approximately 110 million parameters. Based on pre-training on general Chinese corpora (such as Chinese Wikipedia), it can be further fine-tuned using business data from trade communication scenarios to adapt to the intent classification task in this field.

[0160] User input text can be fed into the input embedding layer of the BERT model. The model first segments the user input text, converting each word into a corresponding word embedding vector with a dimension of 768. Simultaneously, a positional embedding vector is added to each word to represent its position within the input sequence, preventing the self-attention mechanism from losing sequence order. The positional embeddings are learned during model pre-training and have the same dimension as the word embeddings. A segmental embedding vector is then added to each word to distinguish different segments in the input sequence; for single-sentence input scenarios, all words share the same segmental embedding. The word embedding vectors, positional embedding vectors, and segmental embedding vectors are element-wise summed to form the model's input vector, with a dimension of [sequence length × 768], which serves as the input to the subsequent Transformer encoder.

[0161] The input vector is sequentially passed through a 12-layer Transformer encoder. Each encoder layer contains a multi-head self-attention mechanism and a feedforward neural network. The multi-head self-attention mechanism uses 12 attention heads to compute the association weights between words in the text in parallel, enabling the model to capture long-distance dependencies; the feedforward neural network performs a non-linear transformation on the attention output to extract higher-level semantic features. After layer-by-layer computation by the 12 encoder layers, the model outputs a contextual semantic representation vector of the text, with dimensions of [sequence length × 768].

[0162] The vector corresponding to the [CLS] marker in the context semantic representation vector output by the encoder (this vector aggregates the semantic information of the entire sequence) is used as the overall representation of the text and input into the classification layer. The classification layer consists of a fully connected layer and a softmax function, mapping the 768-dimensional semantic representation to a multi-dimensional intent category space and outputting the probability distribution of each intent category. The category with the highest probability is selected as the final intent information output. For example, when a user inputs "I want to buy an excavator", the model calculates that the probability of the "inquiry" category is 0.92, so the output intent information is "inquiry".

[0163] The contextual semantic representation vector can be input into the sequence labeling layer, which outputs a label for each word position in the input sequence. The label identifies whether the word is part of a product keyword. Continuously labeled entity words are extracted as product keywords. Specifically, the model outputs information for each word in the input sequence, with information types including B-product, I-product, B-model, I-model, B-price, I-price, and O (irrelevant). By extracting continuously labeled entity words, a list of product keywords is obtained. For example, from "I want to buy a SY215C excavator," the model labels "SY215C" as "B-model" and "excavator" as "B-product," extracting the product keywords as "SY215C" (model) and "excavator" (product category).

[0164] Convolutional neural networks (CNNs) can be used to perform emotion recognition on user-input text to obtain emotional information. For example, a TextCNN model (a deep learning model that applies CNNs to text classification tasks) can be used. Its embedding dimension can be 300 (using Word2Vec pre-trained word vectors), the convolutional kernel size can be [2,3,4,5]-gram, the number of each type of convolutional kernel is 100, the total number of features is 400, the dropout rates are 0.5 and 0.3 respectively, the hidden layer activation function can be ReLU, the output layer activation function can be Sigmoid, and the loss function can be BinaryCrossEntropy.

[0165] The user input text can be segmented into a word sequence [w1, w2, ..., wn]. Each word is mapped to a low-dimensional, dense vector representation using a pre-trained word embedding layer (e.g., 300-dimensional word vectors trained on a large corpus using Word2Vec). After processing, the text of length n is transformed into an input matrix of dimension n × 300, where each row corresponds to the semantic vector of a word.

[0166] The obtained input matrix is ​​fed into the convolutional layer. To capture local semantic patterns of different lengths (granularities), convolutional kernels of various sizes can be used for parallel convolution operations. Specifically, convolutional kernels of sizes 2, 3, 4, and 5 are selected to extract 2-gram, 3-gram, 4-gram, and 5-gram continuous word sequence features, respectively. For example, a convolutional kernel of size 3 covers the word vectors of three consecutive words at a time, and the semantics of the combination of these three words are represented through convolution operations. To fully capture diverse emotional expressions, 100 convolutional kernels of each size are set. Each convolutional kernel performs sliding window calculation on the input matrix to extract local n-gram features (local semantic features). For example, a 3-gram convolutional kernel covers the word vectors of three consecutive words at a time, and the local semantic features of the combination of these three words are extracted through convolution operations. After the convolution operation, each kernel size produces 100 feature maps. Each feature map represents a specific local pattern in the text that is related to a certain emotion (such as very good, not so good, eagerly want, etc.). The dimension of each feature map is the sequence length minus the kernel size plus 1.

[0167] To extract the most critical signals from each feature map and address the issue of inconsistent input text lengths, max pooling is performed on each feature map, which involves taking the maximum value from all values ​​in that feature map. This maximum value represents the strongest emotional feature captured by the convolutional kernel throughout the entire text. After pooling, the 400 feature maps become 400 pooling values. Concatenating these values ​​forms a 400-dimensional global feature vector, which represents the sentiment representation of the entire input text.

[0168] Before inputting the feature vector into the fully connected layer, a Dropout layer can be used to randomly discard some neurons with a certain probability (e.g., 0.5 and 0.3). This effectively prevents the model from over-relying on certain specific features and improves generalization ability. The 400-dimensional feature vector is then input into the fully connected layer and, with the help of activation functions such as ReLU, undergoes a non-linear transformation to map it to a predefined sentiment category space.

[0169] For multi-label emotion recognition tasks (where a text may contain both excitement and urgency), the output layer uses the sigmoid activation function to output an independent probability value P, ∈ [0, 1], for each emotion category, representing the confidence level of that emotion. During training, the convolutional neural network model can use binary cross-entropy as the loss function to optimize the model parameters. Since this task involves multi-label classification, binary cross-entropy can independently measure the prediction error of each emotion category, making it more suitable than multi-class cross-entropy.

[0170] For example, the output of the Sigmoid layer can be binarized according to a preset threshold (such as 0.5), and at least one emotion category with a probability higher than the threshold can be used as the output emotion information. For instance, when a user inputs "I want to buy an excavator!!", the model outputs "excitement" with a probability of 0.85, "eagerness" with a probability of 0.72, and "neutrality" with a probability of 0.12, so the output emotion information is "excitement" and "eagerness".

[0171] It is understood that intent information may include awareness tags obtained from user input text and / or intent text contained within user input text, and emotion information may also include emotion tags obtained from user input text and / or emotion text contained within user input text, without limitation.

[0172] In one embodiment, the response generation method provided in this application further includes: determining whether a user profile exists in the cache; in response to the existence of a user profile in the cache, loading the user profile from the cache; in response to the absence of a user profile in the cache, loading the user profile from the database and writing the user profile into the cache, and determining a response strategy that matches the user profile.

[0173] Specifically, when retrieving a user's profile, the system first checks if the user's profile data exists in the Redis cache. If it doesn't exist (i.e., it's the first time accessing the database or the cache has expired), it queries the user's profile information from the MySQL database, writes the query result to the Redis cache, and sets the TTL to 24 hours. Within the next 24 hours, when accessing the user's profile again, it is read directly from the Redis cache without needing to query the database again.

[0174] The determination of response strategies matching user profiles includes: responding to user profiles including historical intents, determining the activity level and dispersion of various historical intents, determining the tone of the response based on the activity level of historical intents, determining the guidance method of the response based on the dispersion of historical intents, and incorporating tone and guidance method into the response strategy; and / or, responding to user profiles including product preferences, determining the preferred product types based on the distribution characteristics of product preferences, and incorporating the preferred product types into the response strategy; responding to user profiles including language preferences, determining the language style of the response based on the type of language preference, and incorporating the language style into the response strategy; and / or, responding to user profiles including interaction frequency, determining the guidance strength and information density of the response based on the interaction frequency index, and incorporating the guidance strength and information density into the response strategy; and / or, responding to user profiles including user levels, determining the priority of response generation methods based on user levels, and incorporating the priority of generation methods into the response strategy, where the priority of generation methods is the order in which different generation methods are selected when generating a response.

[0175] User profiles can be obtained by analyzing users' historical interaction behavior. User profiles include information from multiple dimensions such as historical intent, product preferences, language preferences, interaction frequency, and user value level.

[0176] Among them, historical intent is a record of the types of intents that users have shown in past interactions, including inquiries, technical consultations, complaints, order guidance, casual conversations, etc., with an accompanying time decay weight.

[0177] Product preferences are the product categories and their weight distribution that a user focuses on in their historical interactions, reflecting the degree of interest a user has in specific product types. Product preferences can be obtained by extracting entities from a user's historical chat logs.

[0178] Language preference refers to the tendency of a user's language expression in their historical interactions, reflecting their preferred language style and expression methods. Language preferences can be obtained by analyzing the language style of a user's historical messages. Specifically, language preferences can include concise (users prefer short sentences and direct expression), detailed (users prefer long sentences and complete expressions), technical (users prefer technical jargon and technical details), and colloquial (users prefer everyday expressions and conversational language).

[0179] Interaction frequency refers to the number of times a user interacts with the system and the speed of their response within a certain time window, reflecting the user's activity level and responsiveness. The interaction frequency index can be obtained by statistically analyzing user interaction behavior within a preset window (such as the last 30 days).

[0180] User value rating refers to a user's overall rating based on their historical value, interactive behavior, and potential value. It is used to guide the selection of response strategies and can be specifically divided into high-value users, ordinary users, and new users.

[0181] In one embodiment, based on the message acquisition rules corresponding to the target interaction environment, the original message data in the target interaction environment is acquired, and the original message data is converted into standard message data; Based on the interface area features and interactive operation location features of the target interactive environment, the message interactive operation mapping relationship is determined; Obtain the business scenario corresponding to the standard message data, determine the message generation branch corresponding to the business scenario in the pre-trained message generation agent, and generate reply message data corresponding to the standard message data through the message generation branch. Generate an interactive response adapted to the target interactive environment based on the reply message data, and send the interactive response to the target interactive environment based on the message interaction operation mapping relationship.

[0182] The aforementioned message interaction method unifies the conversion and processing of message content in different message interaction environments, enabling message data from different sources and with varying formats to enter a more suitable processing flow, reducing the impact of environmental differences on message recognition and processing. Furthermore, by combining the interface state and interaction position relationships of the target interaction environment, the output position and interaction path of the reply message are adapted, ensuring that the reply content can be accurately sent in the corresponding environment. Additionally, by generating corresponding reply content based on the business scenario to which the message belongs, the matching degree between the reply result and the actual message semantics and application requirements is improved. This achieves the technical effects of enhancing cross-environment message processing compatibility, improving reply sending accuracy and interaction execution stability, and enhancing the relevance and accuracy of reply content, thereby solving the problems of poor cross-environment compatibility, high adaptation costs, and insufficient interaction stability in existing technologies.

[0183] Embodiments of this application also provide a computer-readable storage medium storing a computer program, wherein the computer program is configured to execute the steps in any of the above-described message interaction method embodiments at runtime.

[0184] In one exemplary embodiment, the aforementioned computer-readable storage medium may include, but is not limited to, various media capable of storing computer programs, such as a USB flash drive, read-only memory (ROM), random access memory (RAM), portable hard disk, magnetic disk, or optical disk.

[0185] Embodiments of this application also provide a computer program product, which includes a computer program that, when executed by a processor, implements the steps in any of the above-described message interaction method embodiments.

[0186] Embodiments of this application also provide an electronic device. Memory, used to store computer programs; A processor is a process used to implement message exchange methods when executing computer programs.

[0187] Those skilled in the art will further recognize that the units and algorithm steps of the various examples described in connection with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of both, such as Figure 6 As shown, to clearly illustrate the interchangeability of hardware and software, the components and steps of each example have been generally described in terms of functionality in the above description. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.

[0188] The above provides a detailed description of a message interaction method provided in this application. Specific examples have been used to illustrate the principles and implementation methods of this application. The descriptions of the embodiments above are only intended to help understand the method and its core ideas. It should be noted that those skilled in the art can make various improvements and modifications to this application without departing from its principles, and these improvements and modifications also fall within the protection scope of the claims of this application.

Claims

1. A message interaction method, characterized in that, The method includes: Session analysis information is obtained by acquiring and analyzing session data corresponding to the session object in the target interaction environment; wherein, the session data includes the interaction data of the session object in the current round; Obtain the user profile of the session object; The user profile is updated using the conversation analysis information and previous response processing data; Generate an updated user profile and candidate responses to the session analysis information; wherein the candidate responses are used to assist in generating interactive responses to the interaction data.

2. The message interaction method according to claim 1, characterized in that, The process of acquiring and analyzing session data corresponding to the session object in the target interaction environment to obtain session analysis information includes: Extract the interaction data of the current round and the historical interaction data from the session data; The interaction data and the historical interaction data are arranged and combined according to the session time order to obtain the session content to be analyzed; Semantic parsing is performed on the content of the conversation to be analyzed to determine one or more of the intent information, sentiment information, and conversation topic information of the content of the conversation to be analyzed; Based on one or more of the intent information, emotion information, and conversation topic information, determine one or more of the corresponding intent information, emotion information, and product keywords, and use them as the conversation analysis information.

3. The message interaction method according to claim 2, characterized in that, The generation of the adapted and updated user profile and the candidate responses based on the session analysis information includes: Based on the intent information, the emotion information, and the product keywords, determine the response template style, response composition information, and information organization method for the current round; Based on the style of the response template and the information organization method, at least two matching response expression structures are determined from the preset response expression structures; Obtain the product keywords corresponding to the information constituting the response, extract the content elements corresponding to the product keywords from each of the conversation analysis information, and organize and fill the content elements according to at least two determined response expression structures to generate at least two candidate responses; Among them, the at least two candidate responses differ in one or more aspects such as response length, semantic focus, and guiding expression.

4. The message interaction method according to claim 1, characterized in that, The candidate responses are used to assist in generating interactive responses that address the interactive data, including: In response to the fact that there are multiple candidate responses, one of the candidate responses is selected as the target candidate response; In response to obtaining a modification operation on the target candidate response, the target candidate response, the modified response to the target candidate response, and the modification content are recorded; In response to the completion of sending the target candidate response or the modified response, the actual sent target candidate response or modified response is recorded as the interaction response; The response processing data for the current round is generated based on one or more of the target candidate responses, the modified content, and the interactive responses.

5. A message interaction method according to claim 4, characterized in that, The step of updating the user profile using the session analysis information and previous response processing data includes: Obtain the session analysis information for the current round, and obtain the previous response processing data corresponding to the session object; Based on the object identifier and session identifier of the session object, the session analysis information of the current round is associated with the previous response processing data to obtain profile update data; Based on the profile update data, update at least one of the user intent preferences, emotional tendencies, and product attention preferences in the user profile.

6. A message interaction method according to claim 3, characterized in that, After generating the adapted and updated user profile and the candidate responses based on the session analysis information, the process further includes: Extract one or more of the user preference tags, communication style tags, and product interest tags from the updated user profile that correspond to the current round of conversation analysis information; Based on one or more of the user preference tags, the communication style tags, and the product attention tags, filter conditions are determined for multiple candidate responses. The filter conditions include one or more of the following: response length, tone style, product information presentation order, and guiding expression method. Each of the candidate responses is matched with the filtering conditions to obtain the matching degree of each candidate response. At least one candidate response that meets the preset matching criteria is selected as the recommended candidate response.

7. The message interaction method according to claim 1, characterized in that, Before obtaining and analyzing the session data corresponding to the session object in the target interaction environment to obtain session analysis information, the method further includes: Obtain platform characteristic information of the target interactive environment; The platform feature information is matched with a preset platform feature template, and the platform adaptation parameters corresponding to the target interaction environment are determined based on the matching results. Based on the platform adaptation parameters, the system performs session data reading, reply content input, and message sending control for the target interactive environment.

8. The message interaction method according to claim 1, characterized in that, The step of obtaining the user profile of the session object includes: In response to the absence of a user profile for the session object, retrieve the historical interaction data of the session object; In response to the existence of the historical interaction data, an initial user profile of the session object is generated based on the historical interaction data and the preset profile initialization template, and the initial user profile is used as the user profile; In response to the absence of the historical interaction data, an initial user profile of the session object is generated based on the interaction data of the current round and the preset profile initialization template, and the initial user profile is used as the user profile.

9. A message interaction device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the steps of the method according to any one of claims 1 to 8.

10. A message interaction method, characterized in that, Applied to the interactive environment side, including: Obtain the interaction data of the current round and send the interaction data to the message interaction device according to claim 9; Receive the candidate responses returned by the message interaction device, and generate an interactive response based on the candidate responses.

11. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 8, or the steps of the method according to claim 10.