User intention recognition and autonomous intelligent shopping guide method of multi-agent architecture
Patent Information
- Application Number
- CN202610890821.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-06-18
- Publication Date
- 2026-09-11
AI Technical Summary
[0005]本发明提供多智能体架构的用户意图识别与自主智能导购方法,解决相关技术中现有无法显式保留文本语义与行为意图之间的维度差异,导致关键会话状态识别不足;策略输出缺乏与用户跨轮次及跨会话约束的一致性校验机制;以及交互数据未回写CRM导致用户画像与跨会话记忆无法形成闭环迭代的技术问题
通过在各意图维度上对文本语义向量与行为意图向量分别设置独立线性投影层,取两侧投影值之差构成冲突向量,以二者余弦相似度作为一致性评分,连同意图分类标签一并写入意图图谱,并将历史画像高权重标签映射为历史偏好参考节点。冲突向量各分量的符号与幅度分别表征模态信号强弱与偏离程度,策略网络对冲突向量提取符号特征与幅度特征拼接为冲突结构向量后构建状态输入,结合对话状态向量表达会话进展与策略使用历史,可在价格认知冲突、场景歧义等状态下选择价值阐释式或澄清提问式策略;当最高策略概率未超过预设单一策略置信阈值时执行联合策略,用户回复经参与深度等级判定驱动在线参数更新,避免融合方法将差值抹去而误判偏好方向;
Smart Images

Figure CN122736729A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of intelligent recommendation technology, and more specifically, to a method for user intent recognition and autonomous intelligent shopping guidance using a multi-agent architecture. Background Technology
[0002] With the increasing popularity of official mini-programs and online service channels of premium retail brands, intelligent shopping assistants are gradually taking on the task of providing purchase decision consultation in scenarios with high SKU complexity. Users typically complete their independent browsing before initiating a conversation, and their consultation language is often brief, vague, and even deviates from the browsing direction. The system needs to understand both textual expression and behavioral trajectory to provide effective shopping guidance. These scenarios place higher demands on the accuracy of intent recognition and the matching degree between shopping guidance content and user constraints.
[0003] Existing intelligent shopping guide solutions generally employ a multimodal fusion approach, encoding consultation text and browsing behavior separately and then integrating them into a single vector through weighted summation or attention mechanisms. Based on this vector, user preferences are inferred and recommended content is generated. Some solutions incorporate CRM historical profiles to assist in personalization, with the strategy type determined once based on the fusion vector. However, the generation process does not compare the results with user constraints across different rounds, nor does it explicitly model the dimensional differences between text and behavior.
[0004] The above solutions have the following technical problems: the fusion operation masks modal differences in dimensions such as price; when users express price constraints but browse high-priced products, it is difficult to identify price perception conflicts and tends to give mid-priced recommendations; the strategy parameter package is generated based on the current round of intent information and lacks consistency verification with cross-session and cross-round constraints; constraints confirmed in multi-round dialogues are not used as mandatory verification items; intent tags and user response signals are not structured and written back to the CRM, user profiles and cross-session memories cannot be continuously iterated, and strategy selection is difficult to dynamically adjust with the user's level of participation in the session. Summary of the Invention
[0005] This invention provides a user intent recognition and autonomous intelligent shopping guide method based on a multi-agent architecture, which solves the technical problems in related technologies such as the inability to explicitly retain the dimensional differences between text semantics and behavioral intent, resulting in insufficient recognition of key session states; the lack of a consistency verification mechanism between policy output and user cross-round and cross-session constraints; and the failure to write interaction data back to CRM, resulting in the inability of user profiles and cross-session memories to form a closed-loop iteration.
[0006] This invention provides a method for user intent recognition and autonomous intelligent shopping guidance using a multi-agent architecture, comprising the following steps: S1, collect the current conversation behavior event sequence and CRM historical profile, encapsulate the consultation text, behavior event sequence and historical profile into input triples, and merge the consultation text into the dialogue history text; S2, the dialogue history text is encoded into a text semantic vector, the behavioral event sequence is encoded into a behavioral intent vector, the difference between the independent projection values of the two for multiple intent dimensions is used to form a conflict vector, the similarity between the two is used as a consistency score, and the same graph classification label is written into the intent graph. S3 extracts symbolic and amplitude features from the conflict vectors in the intent graph and concatenates them into a conflict structure vector. Combined with consistency score, intent classification label, historical profile category distribution vector and dialogue state vector, it constructs a policy network state input and output policy parameter package. S4 initializes the confirmed constraint set with cross-session constraints in the historical profile and adds constraints extracted from the consultation text. After generating candidate sales guide dialogues guided by the strategy parameter package, each item is verified. If there are contradictions, the strategy parameter package is revised through generation and strategy loop negotiation and regeneration is performed until it passes or the question is switched to clarify. Finally, the sales guide content is output. S5 uses the intent tags, strategy types, and user response signals generated in this round of interaction as data sources, and writes them back to the CRM system in a structured manner to complete the closed-loop iterative update of user profiles and preference models.
[0007] Preferably, in S1, the behavioral event sequence collection extracts page interaction events with the time window from the start of the conversation to the arrival of the consultation text. The filtering threshold is dynamically determined by the median duration of the user's historical valid browsing events. Newly registered users are initially referenced by the historical median of users in the same category. During the cold start phase, the median is degraded to the global median of the entire category. When the number of valid browsing samples in a category exceeds the preset cold start switching threshold, the median is switched to the category-specific median. Historical profiles include historical order categories and price ranges, interest tags, cross-session strong confidence constraint items, and historical category interest weight vectors.
[0008] Preferably, in step S2, the text semantic encoding concatenates all user statements and system replies from the start of the session to the current round in chronological order and sends them to the pre-trained language model, taking the sequence pooling output as the text semantic vector; The behavior sequence encoding constructs type embedding, product attribute embedding and duration features for each event, and superimposes relative time position encoding with the last event of the behavior event sequence as the reference point and the time difference between each event and the reference point discretized into time distance level. After processing by a multi-head self-attention layer, the average pooling output of the whole sequence is taken as the behavior intent vector.
[0009] Preferably, in step S2, two independent linear projection layers are set for each intent dimension of the text semantic vector and the behavioral intent vector, respectively mapping to the corresponding one-dimensional scalar space. The difference between the text-side projection value and the behavioral-side projection value is taken as the conflict component of the corresponding dimension, and the conflict components of each dimension together constitute the conflict vector. The projection layer and encoder are trained together. The objective function is composed of a weighted sum of policy classification cross-entropy loss and projection orthogonal regularization term. The gradient only updates the encoder and projection layer parameters. The cosine similarity between the text semantic vector and the behavior intent vector is used as the consistency score. The two vectors are each processed by an independent intent classification head to output intent classification labels, which are written into the intent graph along with the conflict vector and the consistency score.
[0010] Preferably, in step S3, symbolic features and amplitude features are extracted from each dimension of the conflict vector in the intent graph, and each is concatenated into a conflict structure vector after undergoing independent linear transformation; the conflict structure vector is concatenated with the consistency score, the one-hot encoding corresponding to the intent classification label, the historical profile category distribution vector, and the dialogue state vector to construct the policy network state input; The dialogue state vector is obtained by compressing the sequence of historical policy types in the session, the sequence of participation depth levels in each round, and the number of the current round through a fully connected coding layer.
[0011] Preferably, in S3, the policy network consists of two fully connected layers and a Softmax output layer connected in series. The first layer is activated by ReLU, and the second layer outputs the probability distributions of five policy types: narrative guidance, direct recommendation, comparison and display, value interpretation and clarification questioning. The low-rank matrix is connected in parallel to the input side of the second layer to generate an adaptation increment. When the highest probability exceeds the preset single-strategy confidence threshold, the corresponding strategy is adopted; otherwise, the two strategy types with the highest probability are selected and a joint strategy is executed. The target SKU set is taken as the union of the two strategy candidates and rearranged according to the matching degree. The narrative dimension tags are combined to form a dual-clue guide.
[0012] Preferably, in S3, user replies are judged into four levels according to the depth of participation. The reward weight of each level increases in order of level number. When the same reply meets the conditions of multiple levels, the highest level is taken. When the dialogue ends and there is no positive conversion signal, a negative reward is applied. Offline pre-training uses the historical session conversion results as the basis weights multiplied by the participation depth level coefficient to obtain pseudo-rewards, and uses the weighted sum of policy gradient loss and KL divergence constraint terms as the objective function to update all parameters in batches. The online update uses the previous round's state and action pairs as the update objects, and drives the gradient update with the instant reward for participating in the depth level conversion, adjusting only the parameters of the second fully connected layer and the low-rank matrix parameters.
[0013] Preferably, in step S4, the constraint extraction module detects constraints on the consultation text in four categories of slots: budget, scenario, ingredient taboos, and audience, based on the slot filling model, and the fuzzy price expression is quantified into a numerical threshold through category price quantile mapping; Constraints containing explicit quantification values or negative qualifiers are classified as hard constraints, while the rest are classified as soft constraints. Cross-session strong confidence constraint entries in the historical profile are immediately filled into the initial content of the confirmed constraint set after the input triple is available. Constraints extracted in the current round are appended to the confirmed constraint set. Only hard constraints trigger subsequent validation loops.
[0014] Preferably, in step S4, after the candidate sales pitch is divided into sentence units according to the statement boundaries, it is verified one by one with the hard constraints in the confirmed constraint set: price hard constraints are determined by numerical comparison, and semantic hard constraints are sent to a lightweight natural language inference model to determine the three types of relationships: implication, neutrality or contradiction. When any sentence is determined to be contradictory to the hard constraint, the process enters the generation to policy loop negotiation, generates an agent to build a conflict report, and the policy agent regenerates the policy parameter package after partially revising the main conflict intention dimension. When the cumulative number of times the same hard constraint is triggered exceeds the preset loop threshold, the strategy is switched to clarification questioning.
[0015] Preferably, in step S5, the intent tags in the intent graph are merged with the existing interest tags in the CRM using a sliding weighting method. Existing tags of the same type are updated with weights and then normalized and written back. New tags are added directly with the current confidence level. The newly added hard constraint entries in the confirmed constraint set in this round are merged with the CRM cross-session constraint memory. If the number of times the same constraint appears in different sessions exceeds the preset strong confidence threshold, it will be upgraded to a strong confidence constraint and pre-populated into the confirmed constraint set in the next session. The category interest distribution vectors in the behavior event sequence of this session are weighted and merged with the historical category interest weight vectors, then normalized and written back to the CRM.
[0016] The beneficial effects of this invention are as follows: By setting independent linear projection layers for text semantic vectors and behavioral intent vectors at each intent dimension, the difference between the two projection values is used to form a conflict vector. The cosine similarity between the two is used as a consistency score, and the intent graph classification label is written into the intent graph along with the cosine similarity. The high-weight labels of the historical profile are mapped to historical preference reference nodes. The sign and amplitude of each component of the conflict vector represent the strength and deviation of the modal signal, respectively. The policy network extracts the sign and amplitude features of the conflict vector and concatenates them into a conflict structure vector to construct the state input. Combined with the dialogue state vector, it expresses the conversation progress and policy usage history. It can select value interpretation or clarification questioning strategies in states such as price perception conflict and scenario ambiguity. When the probability of the highest strategy does not exceed the preset single strategy confidence threshold, the joint strategy is executed. The user's response is driven by the participation depth level to update the online parameters, avoiding the fusion method from erasing the difference and misjudging the preference direction. The system initializes the confirmed constraint set with cross-session constraints from historical profiles and adds constraints extracted from consultation text. After generating candidate sales guide dialogues guided by a strategy parameter package, each dialogue is validated. Price-related hard constraints are prioritized for numerical comparison, while semantic-related hard constraints are determined by a natural language inference model to indicate implied, neutral, or contradictory relationships. When contradictions exist, an agent is generated to construct a conflict report, the strategy agent partially revises the strategy parameter package, and then regenerates the dialogue. If a preset loopback threshold is exceeded, a clarification question is switched. Intent tags, strategy types, and user response signals are asynchronously written back to the CRM after the sales guide content is output. Interest tags are merged using a sliding weighting system, the cumulative occurrence count of constraints is used to upgrade strong confidence constraints, and category interest weights are weighted and merged. Strong confidence constraints are pre-populated into the confirmed constraint set in the next session, and category interest weights are normalized and written back for loading into the historical profile in the next session, forming a closed-loop iteration of the user profile. Attached Figure Description
[0017] Figure 1 This is a flowchart of the user intent recognition and autonomous intelligent shopping guide method based on the multi-agent architecture of the present invention; Figure 2 This is a flowchart of the multi-agent architecture user intent recognition and autonomous intelligent shopping guide method of the present invention. Figure 1 ; Figure 3 This is a flowchart of the multi-agent architecture user intent recognition and autonomous intelligent shopping guide method of the present invention. Figure 2 . Detailed Implementation
[0018] The subject matter described herein will now be discussed with reference to exemplary embodiments. It should be understood that these embodiments are discussed only to enable those skilled in the art to better understand and implement the subject matter described herein, and changes may be made to the function and arrangement of the elements discussed without departing from the scope of this specification. Various processes or components may be omitted, substituted, or added as needed in the examples. Furthermore, some features described in the examples may be combined in other examples.
[0019] At least one embodiment of the present invention discloses a multi-agent architecture method for user intent recognition and autonomous intelligent shopping guidance, consisting of five sequentially connected functional steps. S1 completes the acquisition of multi-source heterogeneous signals from the current user session, integrating consultation scripts, real-time browsing behavior sequences, and CRM historical profiles into a structured input triplet; S2 uses this triplet as input, and through a cross-modal tension-aware encoding mechanism, explicitly quantifies the dimensional deviation between text semantics and behavioral intent into a conflict vector, constructing a structured intent graph carrying conflict information; S3 uses the conflict vector of the intent graph and the dialogue state as decision input, and infers the current optimal shopping guidance strategy type and parameter package through a strategy network driven by a multi-level reward mechanism for dialogue participation depth; S4 performs consistency verification at the generation level between the strategy parameter package and the continuously accumulated confirmed constraint set in the session, ensuring global consistency between the output content and the user's historical constraints through a generation-strategy loop negotiation mechanism; S5 structurally writes back the intent recognition results, strategy paths, and user response signals generated in this round of interaction to the CRM, driving continuous iteration of the user profile and cross-session memory. Figures 1 to 3 As shown, it includes the following steps: S1, collect the current conversation behavior event sequence and CRM historical profile, encapsulate the consultation text, behavior event sequence and historical profile into input triples, and merge the consultation text into the dialogue history text; When a user sends an inquiry message in the shopping assistant's chat interface, the system marks this round of interaction with the session identifier (session_id) assigned when the session started, records the original text of the message (T_raw) and its arriving Unix timestamp (t_query), and uses t_query as the anchor point to simultaneously trigger two parallel operations: behavior sequence collection and CRM query.
[0020] Behavioral sequence collection starts at the session start time t_start and ends at t_query, extracting all valid page interaction events within the (t_start, t_query) time window from the event tracking data stream. Each behavioral event record contains five fields: event type (enumerated types, including page browsing start, product image click, scrolling to the bottom of details, dwelling in the price area, favorites / add to cart, etc.), product object ID (associated with the product node in the product knowledge graph), duration of the event (in seconds), page position coordinates at the time of the event, and event timestamp. In the original event stream, some page dwell times are too short (e.g., users quickly return after accidental touches). These events do not carry valid browsing intent and need to be filtered before serialization. The filtering threshold is dynamically determined by the median duration of valid browsing events in the user's historical sessions. For newly registered users, the historical median of users of the same category within the system is used as the initial reference value. When the system is in the cold start phase of a new product category (where there are no historical browsing event records for this category within the system), the filtering threshold degenerates into the global median of the duration of valid browsing events across all categories in the system. As the number of valid sessions for this category accumulates, when the sample size of valid browsing events specific to the category exceeds the cold start switching threshold (the default value is 50 valid browsing events), the system automatically switches to the category-specific median and no longer uses the global replacement value. The filtered events are arranged in ascending order of timestamps, forming a behavioral event sequence B_seq.
[0021] The CRM query uses the user ID as the search key and extracts four types of fields from the CRM system: the user's historical order categories and price range records in brand channels (within the last 12 months), system-labeled interest tags (including tag names and tag weights), strong confidence constraint entries in cross-session constraint memory (historical high confidence constraints written in step S5, such as scenario preferences or allergy information that the user has confirmed multiple times), and the historical category interest weight vector I_history (continuously maintained by the browsing interest weight update mechanism in step S5, reflecting the user's category preference distribution across sessions). These four types of fields are encapsulated into a historical profile object H, which, together with T_raw and B_seq, constitutes the input triplet for this round of interaction.
[0022] The core content of the triple consists of three parts: the consultation text T_raw, the behavioral event sequence B_seq, and the historical profile H. To support cross-agent message routing, these three parts, along with the session identifier session_id and the timestamp t_query, are encapsulated into a structured message.
[0023] S2, the dialogue history text is encoded into a text semantic vector, the behavioral event sequence is encoded into a behavioral intent vector, the difference between the independent projection values of the two for multiple intent dimensions is used to form a conflict vector, the similarity between the two is used as a consistency score, and the same graph classification label is written into the intent graph. In one embodiment of the present invention, the dimensional difference information between text signals and behavioral signals is preserved in a structured form, rather than merging the two into a single optimized vector, so that the subsequent policy agent can directly read and utilize the difference itself. Specifically, text semantic encoding is performed within the complete context of a multi-turn dialogue, rather than encoding only the utterance of the current turn. Taking the question "Do you have any gift recommendations?" as an example: if the user has mentioned "for elders" in a previous dialogue, the word "gift" in this question carries a specific audience focus; if this is the first turn of the dialogue, the same words need to be interpreted within a broader semantic space. Therefore, all user statements and system responses from the start of the current conversation to the current turn are concatenated chronologically into a dialogue history text sequence, which is then fed into a pre-trained language model fine-tuned from dialogue corpora in the brand boutique retail domain. The sequence pooling output of the final encoding layer is taken as the context-aware text semantic vector V_T. This vector semantically carries both the utterance content of the current turn and the trajectory of intent evolution in the dialogue history.
[0024] Behavioral sequence encoding deals with heterogeneous event data of a time series nature, employing a different encoding architecture than text modality. Each event record in B_seq first undergoes embedding layer processing: the event type is converted into a fixed-dimensional type vector through the event type embedding layer; the product object ID is associated with the product knowledge graph, and the primary category, secondary category, and price range to which the product belongs are extracted as product attribute embeddings (the sum of the three embedding vectors); the event duration is logarithmically transformed and added as a numerical feature. The type vector and product attribute embeddings are summed and concatenated with the duration feature to form the complete embedding representation of the event. The embedding at each position in the sequence is coupled with a relative time position encoding. Using the last event in B_seq as a reference point, the time difference between each event and the reference point is discretized into several time distance levels. The time-distance levels correspond to the actual intent intensity distribution of user browsing behavior in boutique retail scenarios: within 1 minute corresponds to rapid scanning behavior, where the user has not yet formed an intention to stay; 1 to 3 minutes corresponds to browsing with initial interest, where the user begins to pay attention to product details; 3 to 8 minutes corresponds to in-depth reading, where the user is actively acquiring product information; and more than 8 minutes corresponds to repeated viewing or comparison, representing the browsing state with the highest intent intensity. The default boundary values for these four intervals are 1 minute, 3 minutes, and 8 minutes, respectively, and can be adjusted based on user behavior data for specific product categories. Each level corresponds to a learnable position embedding vector, which is added together and used in the self-attention calculation. Different users have significantly different browsing rhythms and session durations, making it difficult to generalize the numerical distribution of absolute timestamps across users; in contrast, the relative time interval between events (such as continuously and rapidly clicking on multiple products vs. staying on a single product for a long time before switching to the next) is a more stable signal of behavioral preferences, therefore, relative time position encoding is used instead of absolute timestamps as the carrier of position information. The entire sequence is processed by a multi-head self-attention layer, with the average pooling output of the last layer serving as the behavioral intent vector V_B.
[0025] After obtaining V_T and V_B, the core computational step of cross-modal tension-aware encoding is performed. The intent dimension set consists of two parts: general dimensions and category-specific dimensions. Price sensitivity and exploration / cognition stage are cross-category general dimensions, applicable in any high-SKU complexity boutique retail scenario. Quality orientation, gift scenario suitability, and origin traceability interest are category-specific dimensions for specialty coffee. Other categories can be replaced using the same principle; for example, for customized skincare products, "origin traceability interest" can be replaced with "ingredient traceability interest," and for selected digital products, it can be replaced with "technical parameter depth interest." The projection layer parameters of the general dimensions can be transferred and reused across categories, while the projection layers of category-specific dimensions need to be relearned using the training data of the corresponding category. Taking the specialty coffee category as an example, K takes a default value of 5, i.e., the above 5 dimensions. Two independent linear projection layers are trained for each intent dimension: the text-side projection layer receives V_T and maps V_T to a one-dimensional scalar space of that dimension by performing an inner product operation with a set of learnable weight vectors, obtaining the text-side projection value p_T_k; the behavior-side projection layer performs the same operation on V_B, obtaining p_B_k; the difference between the two projection values (p_T_k minus p_B_k) is the conflict component C_k of the k-th intent dimension, and the conflict components of the K dimensions together constitute the K-dimensional conflict vector C.
[0026] The parameters of the K projection layers are determined through end-to-end joint training, using the final policy type labels adopted in historical session data as the supervision target. They are jointly optimized with the text encoder and behavior encoder. The objective function consists of a weighted sum of the policy classification cross-entropy loss and the projection orthogonal regularization term. The orthogonal regularization term constrains the cosine similarity between each pair of the K text-side projection weight vectors to approach zero, and the same applies to the K behavior-side projection weight vectors. The aim is to ensure that the K intent dimensions capture independent semantic directions in their respective feature spaces, avoiding the reduction of diagnostic granularity of conflict vectors due to multiple dimensions learning highly correlated projections. The default value of the orthogonal regularization term weight coefficient λ is 0.1. The gradients of the cross-entropy loss and the orthogonal regularization term only update the parameters of the text encoder, behavior encoder, and the K projection layers, and are not backpropagated to the policy network in step S3. The policy network in step S3 is an independent module, trained offline and online using the policy gradient loss based on participation depth reward described below. Although both stages use policy type information from historical sessions as a reference, they belong to different network modules, and their parameter spaces are isolated. After training, by analyzing the activation distribution of the projected weight vectors of each dimension on the category text corpus and behavioral event labels, posterior semantic annotations were performed on the K dimensions. The naming of dimensions such as "price sensitivity" and "gift scenario suitability" is based on semantic induction from the training results, rather than prior constraints imposed during the training phase. Simultaneously, the cosine similarity between V_T and V_B was calculated as a global consistency score S_cons. A value closer to 1 indicates a greater alignment in the overall semantic direction between the two modalities, while a low value indicates a significant deviation between the text expression and browsing behavior in the overall direction.
[0027] Multimodal fusion methods integrate V_T and V_B into a single vector through weighted summation or attention mechanisms, in most "meaningful" contexts. Figure 1 While sufficient for general sales guidance in typical conversations, this approach has structural limitations in handling a specific type of conversation within premium retail scenarios. This type of conversation, comprising approximately 5% to 15% of all conversations, possesses the highest sales guidance value: users explicitly express price constraints in their speech, yet their browsing behavior focuses on higher-priced items exceeding these constraints. These users are in a state of "price perception conflict," attracted by the quality of higher-priced products but subjectively concerned about the price. Their correct sales guidance direction (value interpretation) is neither following the verbal signal (recommending lower-priced products) nor following the behavioral signal (recommending higher-priced products). When the fusion method handles this type of conversation, the weighted average of the signals on both sides of the price dimension causes the projected value to fall into the middle range. The policy network receives a "mid-price preference" signal, tending to recommend mid-priced products. This error is not due to insufficient model capability, but rather because the fusion operation erases the difference between the two signals (i.e., the "price perception conflict" itself). Furthermore, because this type of conversation accounts for a low percentage in the training data, the policy network of the fusion solution suffers from significant limitations in handling large amounts of "intentional" or "price-related" conversations. Figure 1Under the gradient-dominated model, the model tends to underfit a small number of important conflict state samples, and its identification ability is marginalized during model optimization. The design of the conflict vector C, by placing the difference operation before the feature calculation stage, ensures that the difference component of the price dimension is explicitly preserved and independently input into the policy network under any training distribution. This allows the model to autonomously discover this low-frequency but high-value feature pattern from the high-dimensional fusion space without relying on the network.
[0028] The sign of each component of the conflict vector indicates which modality's signal is stronger in that dimension (a positive value indicates that the projection value of the text signal in that dimension is higher than that of the behavioral signal), and the absolute value indicates the degree of deviation between the two modalities in that dimension. The multidimensional structure of the conflict vector also enables the policy network to distinguish between three different psychological states, rather than uniformly grouping all sessions with price conflict components into the same processing path. In State 1 (the user is genuinely attracted to the high-priced item but has price concerns), after dynamic threshold filtering, the number of effective dwell events is sufficient and concentrated on high-priced items. The projection value of the behavioral intent vector V_B in the price dimension is significantly higher, resulting in a larger absolute amplitude of the price-dimensional conflict component, sufficient to trigger the amplitude threshold of the value interpretation strategy. In State 2 (the user clicks on a high-priced page casually but does not develop substantial interest), after B_seq filtering, the number of effective dwell events is sparse, and the overall signal strength of V_B is weak. Even if a positive price conflict component is generated, its absolute amplitude is constrained by insufficient behavioral signal quality, making it difficult to meet the triggering condition of the value interpretation strategy. State 3 (the user is buying for someone else but has not directly expressed the intention to give a gift) has independent conflict components in the gift scenario dimension. The wording does not mention the gift, but the behavior has already browsed the gift box page. The policy network identifies the state corresponding to this combination signal by receiving the joint distribution of price conflict and gift dimension conflict, and prioritizes triggering the scenario clarification strategy rather than the value interpretation strategy.
[0029] V_T and V_B each come with an independent intent classification head to generate intent type labels in the intent graph. The V_T classification head maps the text semantic vector V_T to 8 explicit intent categories: product recommendation request, price inquiry, functional attribute inquiry, product comparison, gift selection inquiry, scenario adaptation inquiry, delivery and inventory inquiry, and generalized browsing (no explicit intent expression). The V_B classification head maps the behavioral intent vector V_B to 5 implicit intent categories, corresponding to the user's purchase decision stages: shallow category scanning, product detail exploration, in-depth product research, multi-product horizontal comparison, and gift selection exploration. Both classification heads are lightweight structures with a single-layer linear transformation plus Softmax, and are trained end-to-end in conjunction with the encoder and projection layer. The supervision labels come from intent type records in historical sessions, which are manually or rule-based. Each historical session corresponds to a set of explicit intent labels and implicit intent labels. The intent type labels and confidence scores output by the classification heads, along with the conflict vector C and consistency score S_cons, are written into the intent graph G.
[0030] The intent graph G is constructed using the above calculation results as its core fields. Product-type entities (such as "ear loop bag," "blended beans," etc.), attribute words (such as "not bitter," "floral fragrance," "portable"), and scenario words (such as "gift," "office") are extracted from T_raw using a named entity recognition model. A basic graph structure is constructed using semantic relationships as edges. The following fields are written into the graph's attribute area: explicit statement intent (intent type and confidence level output by the V_T classification head), implicit behavioral intent (intent type and confidence level output by the V_B classification head), conflict vector C (a K-dimensional floating-point array, retaining two decimal places of precision), overall consistency score S_cons, and high-weight labels mapped from the historical profile H (labels with weights higher than the first confidence threshold in the historical profile are included in the graph as historical preference reference nodes).
[0031] S3 extracts symbolic and amplitude features from the conflict vectors in the intent graph and concatenates them into a conflict structure vector. Combined with consistency score, intent classification label, historical profile category distribution vector and dialogue state vector, it constructs a policy network state input and output policy parameter package. The conflict vector, consistency score, and intent classification label are read from the intent graph G. The historical preference information obtained by mapping the historical profile high-weight labels carried by the intent graph and the dialogue state vector accumulated in the session are combined to construct the policy network state input S_policy. The policy network in the policy agent infers the optimal shopping guide strategy type for the current round and outputs the corresponding parameter package.
[0032] The construction process of the dialogue state vector D_state is as follows: The sequence of all historical system responses up to the current round (each strategy type corresponds to one-hot encoding; if the history is less than N rounds, it is zero-padding to a length of N rounds), the sequence of engagement depth levels corresponding to each round's user responses (integer labels of 4 levels), and the number of dialogue rounds already performed are compressed into a fixed-dimensional dialogue state vector D_state through a fully connected encoding layer. The default value for the maximum number of rounds N for zero-padding the historical strategy sequence is 6, and the default value for the output dimension of D_state after compression by the fully connected encoding layer is 32. D_state represents the current stage of the session's progress and the history of strategy usage, enabling the policy network to avoid repeatedly using the same type of strategy within a short period when inferring the current strategy.
[0033] The construction of the policy network's state input S_policy is divided into two levels. The first level is the structured decomposition of the conflict vector C: for each dimension of C, two types of derived features are extracted. The sign feature takes the sign of the conflict component in that dimension and encodes it as a scalar of +1 or -1, indicating which mode's signal is stronger in that dimension; the amplitude feature takes the absolute value of the conflict component in that dimension, indicating the degree of deviation between the two modes in that dimension. The sequence of sign features and amplitude features in K dimensions are mapped to two fixed-dimensional vectors through independent linear transformation layers. The default value of the output dimension of the linear transformation of sign features is 16, and the default value of the output dimension of the linear transformation of amplitude features is 16. After concatenation, the dimension of the conflict structure vector C_struct is 32. This decomposition method forces the policy network to structurally distinguish between the two different decision criteria of "which mode's signal is stronger" and "how far the two modes deviate", rather than directly learning the original value of C as a normal feature. The second level is the concatenation of the complete state vector: the conflict structure vector C_struct, the overall consistency score S_cons (1-dimensional), the one-hot encoding of the explicit intent type (converted from the explicit intent label in G, with a dimension equal to the number of explicit intent categories, 8), the one-hot encoding of the implicit intent type (converted from the implicit intent label in G, with a dimension equal to the number of implicit intent categories, 5), the category distribution vector of high-weight labels in the historical profile (representing the user's historical preference direction, with a default dimension of 12), and the dialogue state vector D_state are concatenated into the complete state input S_policy, with a default total input dimension of 90 for S_policy.
[0034] In one embodiment of the present invention, the policy network consists of two fully connected layers and a Softmax output layer connected in series. The data flow between the layers is as follows: The first fully connected layer receives S_policy and linearly maps it to the first hidden layer vector h1 through the weight matrix W1 and the bias b1. The dimension is 64 (the default value of the first hidden layer output dimension is 64). The ReLU activation function is applied to h1 element by element to obtain the intermediate representation h1'. The second fully connected layer receives h1' and linearly maps it to the 5-dimensional logits vector z through the weight matrix W2 and the bias b2. The 5-dimensional output corresponds to five policy types in sequence: narrative guidance, direct recommendation, comparison and display, value explanation, and clarification questioning. The Softmax layer normalizes z element by element to obtain the selection probability distribution π of the five policy types. The low-rank adaptation module is connected in parallel to the input side of the second fully connected layer: the low-rank matrix A has a dimension of r×64, the low-rank matrix B has a dimension of 5×r, the adaptation increment term is B·A·h1' (dimension 5), and the sum of B·A·h1' and W2·h1' constitutes z. The default value of the rank r of the low-rank matrix is 8. During the inference phase, only the above forward propagation is performed; the parameter update phase distinguishes between offline full training and online local fine-tuning modes, each using a different loss function.
[0035] All five shopping guide strategies are designed to address the real psychological states of users in boutique retail shopping guides, with each strategy corresponding to a decision-making scenario with clearly distinguishable criteria. The narrative-guided strategy targets users in a vague exploratory state where S_cons are low and their statements lack effective intent keywords (e.g., simply saying "What do you recommend?"). Traditional keyword matching and intent classification models cannot provide reliable judgments under such sparse input. The narrative-guided strategy evokes user interest and resonance through the product's origin story, craftsmanship, or cultural background, propelling them from the vague exploratory stage to a stage where they can describe specific preferences. The direct recommendation strategy targets users with high S_cons and high intent classification confidence, directly outputting 2 to 3 products with the highest matching degree from the product knowledge graph, without requiring additional intent clarification steps.
[0036] The comparative display strategy targets the state where multiple similar products are browsed in parallel depth within the behavioral sequence (the V_B classification header outputs a significant probability of "multi-product horizontal comparison"). It proactively provides a structured comparison of two representative products across core attribute dimensions, aligning with the user's ongoing comparison behavior. The value explanation strategy targets the state where price sensitivity conflict is prominent, meaning the user's language expresses price constraints but their browsing behavior is concentrated in the high-price range. For these users, existing shopping guide systems typically offer only two options: recommending lower-priced products based on language (making the user feel their preference is ignored), or recommending higher-priced products based on behavior (making the user feel pressured to buy). The value explanation strategy identifies the psychological state unique to the boutique retail scenario of "attracted by high-priced items but with price concerns," helping users establish a reasonable price perception by explaining the value composition of products at the corresponding price point, rather than compromising on price ranges.
[0037] The clarifying questioning strategy targets highly ambiguous states where S_cons is low and the conflict vector lacks a clear principal direction dimension. It generates a directional question for the dimension with the largest absolute value of the conflict component, obtaining the highest-value disambiguation information with the fewest query rounds, thus establishing a more reliable intent foundation for subsequent strategy decisions. The highest conflict dimension is chosen as the query target rather than simultaneously querying multiple dimensions because this dimension represents the direction of greatest divergence between the current two modal signals; focusing on this dimension significantly improves disambiguation efficiency compared to concurrent queries.
[0038] In the probability distribution output by the policy network, when the highest probability exceeds the single policy confidence threshold (default value is 0.60), the policy type corresponding to the highest probability is adopted as the sole policy for the current round; when the highest probability does not exceed the threshold, the system determines that the current user state cannot be fully covered by a single policy type, and executes a joint policy using the two policy types with the highest probabilities. The parameter package fusion rules of the joint policy are as follows: the target SKU set is the union of the candidate products retrieved from the product knowledge graph by each of the two policies, and after re-sorting according to the comprehensive matching degree with the current intent, the top 5 to 10 products with the highest comprehensive ranking are retained; the narrative dimension label is the narrative label corresponding to each of the two policies, forming a combined narrative label as a dual-clue guide for content generation in step S4; the presentation priority sequence uses the main policy logic with higher probability as the main line, and the narrative content of the secondary policy with lower probability as a supplementary embedding, which is uniformly integrated by the generating agent in step S4 during the content generation stage, rather than mechanically splicing together two separate texts. The strategy parameter package (SP) is generated from the product knowledge graph after the strategy type is determined. It includes a target SKU set (5 to 10 candidate products, sorted by their matching degree with the current intent), narrative dimension tags (used for rich media retrieval in step S4), and a presentation priority sequence.
[0039] The parameter training of the policy network is separated into two stages: pre-training and in-session online fine-tuning. Both stages share the same set of participation depth level reward signals. Upon each user reply, the system performs a participation depth level determination on the reply content. The determination result serves as an immediate reward signal for online updates and also as a source of adjustment coefficients for offline pseudo-rewards. Level 1 (continuous dialogue but unrefined question granularity) has a base positive value; Level 2 (the user's next question has significantly refined granularity, determined by the intent recognition model detecting the specificity of entities in the question, such as refining from category-level queries to single-item queries, or from functional descriptions to process parameter descriptions) has a reward weight twice that of Level 1; Level 3 (users actively share personal scenario information, detected by the co-occurrence relationship of first-person pronouns with scenario words, time words, and audience words) has a reward weight three times that of Level 1; Level 4 (users inquire about delivery, inventory, price calculations, or package details, detected by a specific intent type classifier) has the highest reward weight. When a user's response in the same round simultaneously meets the criteria for multiple levels, the highest level is taken as the final recorded value for that round. This integer level is stored in the D_state history sequence. Both online updates and offline pseudo-rewards use the weight of the corresponding highest level, and multi-level rewards are not superimposed. A negative reward is applied when the dialogue terminates and there is no positive conversion signal.
[0040] Offline pre-training of the policy network uses session data with clear transformation results from historical dialogue logs to perform batch gradient updates on all parameters of W1, b1, W2, b2, and low-rank matrices A and B. For each historical session, the S_policy state vector s_t for each round is reconstructed in chronological order. The index a_t of the actual policy type used in that round is read as the action label (when the joint policy is executed in that round, a_t takes the index of the primary policy type with higher probability). The pseudo-reward r_t for that round is calculated according to the following pseudo-reward rules. The objective function for offline training is a weighted sum of the policy gradient loss and the distribution constraint term:
[0041] Where T represents the number of valid rounds in a single historical session. Let a be the predicted probability of the policy network outputting policy type a_t for state s_t in round t under the current parameters θ. This is the pseudo-reward value for this round. Let KL divergence be the probability distribution between two strategies. The strategy distribution is used as a reference (a uniform distribution is used in the cross-brand migration pre-training stage, and the strategy distribution after the migration pre-training converges is used in the local brand fine-tuning stage). The offline KL constraint strength coefficient has a default value of 0.5. The estimation of the pseudo-reward r_t involves two steps: First, the conversion of the session into a base weight is used. Sessions that undergo conversion receive positive base weights for all rounds, while sessions that do not convert receive zero base weights. Sessions where the user actively leaves mid-session receive a negative base weight for the last round before leaving. Second, the base weight is multiplied by an adjustment coefficient for the user's response type in that round. The adjustment coefficient's levels are consistent with the online participation depth level determination rules (level 1 coefficient is 1.0, level 2 coefficient is 2.0, level 3 coefficient is 3.0, and level 4 coefficient is 4.0). The two steps are multiplied together to obtain r_t. The default learning rate for offline batch training is 1×10⁻³, and mini-batch stochastic gradient descent is used for iteration until convergence of the average pseudo-reward on the validation set.
[0042] When deploying for the first time and the brand's historical conversation data is insufficient to support effective offline pre-training, cross-brand migration pre-training is used as the initialization source: a pre-training corpus is constructed from the historical dialogue logs of boutique retail brands with similar category attributes, and pre-trained using the same pseudo-reward annotation logic and the above L_off objective function to obtain initial weights with general boutique retail guide strategy knowledge; after the brand goes live, as the number of effective conversations of the brand increases, the strategy network is fine-tuned in batches using the brand's data. During the fine-tuning phase, the default initial value of β_off is adjusted to 0.7 to protect the basic strategy capabilities of cross-brand migration. After the accumulated amount of effective conversations of the brand exceeds the cold start switching batch (the default value is 500 effective conversations), it is adjusted to the standard default value of 0.5.
[0043] The online update of the policy network is triggered upon each user response, using the state-action pair from the previous round of policy decisions as the update object: s_{t-1} is the reconstructed S_policy in round t-1, a_{t-1} is the policy type index of the actual output in round t-1 (when executing a joint policy in this round, a_{t-1} takes the index of the primary policy type with higher probability), and r is the immediate reward value converted from the user's response engagement depth level determination result (levels 1 to 4 correspond to a base positive value and their 2x, 3x, and 4x weights, respectively; a negative value is taken when the dialogue terminates and there is no positive conversion signal). The objective function for the online update is:
[0044] in, Let $\beta$ be the output probability distribution of the policy network in state $s_{t-1} before this update (serving as a reference anchor to prevent drastic policy distribution shifts due to limited samples within a single session), and $β$ be the online KL constraint strength coefficient. Online updates only perform gradient descent on $W2$, $b2$, and low-rank matrices $A$ and $B$. $W1$ and $b1$ remain unchanged from offline training results to protect the cross-user general feature extraction capability established through offline pre-training. The default learning rate for gradient updates in the online phase is 2 × 10⁻ 4 The KL constraint strength β is adaptively set within the session: the default initial value of β for the first 3 rounds of the session is 0.5; after the cumulative number of valid rewards (reward judgments of level 2 and above) exceeds 2 times within the session, β gradually decays at a rate of 0.85 per round, allowing the policy distribution to converge towards the actual response pattern of the current user, until β drops to the default lower limit of 0.1 and remains unchanged. Within a single session of 5 to 15 rounds, the parameter size of the second fully connected layer and the low-rank adaptation module is small enough that a limited number of online samples can produce a statistically significant directional shift; if the update range is expanded to all network parameters, the gradient update of a limited number of samples has no practical effect.
[0045] S4 initializes the confirmed constraint set with cross-session constraints in the historical profile and adds constraints extracted from the consultation text. After generating candidate sales guide dialogues guided by the strategy parameter package, each item is verified. If there are contradictions, the strategy parameter package is revised through generation-strategy loop negotiation and regenerated until it passes or the question is switched to clarify. Finally, the sales guide content is output. In one embodiment of the present invention, when the policy agent generates policy parameters, it references an intent graph based on cross-sectional information generated in the current round. However, the constraints expressed by the user throughout the multi-round dialogue constitute a complete global context. There may be contradictions between these two, where the policy level is not explicitly modeled. The independent verification mechanism at the generation level is introduced precisely to address this structural inconsistency.
[0046] The session-level initialization of the constraint set CCS is triggered immediately after the output triples and historical profile H in step S1 are available: the session management layer reads the strong confidence constraint entries from the cross-session constraint memory field of H and fills them into the initial content of the CCS for this session. This initialization operation starts in parallel with the encoding process in step S2, ensuring that the CCS has the user's historical strong confidence constraint knowledge when it runs for the first time in step S4. At the same time, T_raw output in S1 is synchronously broadcast to the constraint extraction module via the inter-agent message bus, and constraint extraction runs in parallel with the encoding inference in steps S2 / S3. The constraint extraction module performs constraint entry detection on T_raw based on a slot-filling model. This model is based on a pre-trained language model and, after domain adaptation fine-tuning using BIO sequence labeling format, performs detection on four types of slots: budget slots (slot value type is a numerical range or fuzzy descriptor; expressions with explicit numerical values directly extract the upper limit of the numerical value, while fuzzy expressions like "don't be too expensive" trigger category price quantile mapping), scenario slots (slot value enumeration: gift / personal use / office / other), ingredient taboo slots (slot value is a list of ingredient names), and audience slots (slot value enumeration: elderly / children / friends / partners / self, etc.). The domain adaptation fine-tuning dataset was constructed by the brand operations team using BIO annotation based on constraint expressions that appeared in historical user language. The default lower limit for the number of labeled samples for each slot category is 100. Constraint expressions with clear structures (such as budget expressions with explicit numerical values) are processed with rule-based methods to reduce reliance on neural network models. The computational cost of the slot filling model is much less than that of the transformer encoding in step S2. Constraint extraction can be completed before step S3 is finished. When step S4 starts, the updated CCS is read directly without having to re-access T_raw inside S4.
[0047] Each successfully extracted constraint is appended to the CCS in the format {Constraint ID, Type Marker, Constraint Content Semantic Embedding, Original Text Fragment, Round Number, Number of Repeated Confirmations}. For ambiguous price expressions in the wording (such as "Don't be too expensive," "Don't consider expensive ones," etc.), the constraint extraction module performs quantification processing through category price distribution mapping before entering them into the CCS: the 25th percentile of the price distribution of all SKUs in the current category is used as the default quantification upper limit for the ambiguous low-price expression "Don't be too expensive," and the 50th percentile is used as the default quantification upper limit for the expression "Moderate price." The quantified numerical thresholds are stored together with the original expression text in the price field of the constraint entry. This quantification processing allows the subsequent NLI verification stage to perform deterministic numerical comparisons when dealing with specific product prices, rather than relying on semantic reasoning to interpret relative price expressions, thereby avoiding the inherent limitations of semantic models in judging numerical magnitude. When a user mentions content that is semantically highly similar to an existing item in CCS in subsequent rounds (determined by similarity calculation), the number of repeated confirmations for the corresponding item increments, indicating that the constraint is a persistent concern of the user rather than a one-time expression. The constraint type determination rules are as follows: constraints containing explicit quantifiable values or negative qualifiers are classified as hard constraints; other constraints that express a tendency rather than exclusivity are classified as soft constraints. Hard constraints in CCS are mandatory checks in the consistency verification step S4, while soft constraints are only used as a reference for content priority and do not trigger loops.
[0048] The content generation phase primarily uses the strategy parameter package SP output from step S3 as the guiding parameter. The narrative dimension labels in SP determine the narrative framework of the generated content (e.g., when the narrative dimension label for a narrative-guided strategy is "origin story," the generated content should focus on the origin background; when the narrative dimension label for a value-interpretation strategy is "craft value," the generated content should focus on the professionalism of the processing techniques and their flavor association). The target SKU set determines the specific range of products involved in the content; the presentation priority sequence determines the order in which multiple products appear in the dialogue content. Under the constraints of these parameters, the generative language model, combined with the current conversation history (as contextual background, ensuring consistency in tone and style with previous rounds of responses) and the brand's standardized language template (injected as system prompts, specifying language style, taboos, and brand tone), generates candidate sales guide texts. The length and structure of the candidate texts are determined by the strategy type (narrative-guided strategies allow for longer narrative paragraphs, while clarifying question strategies only generate a concise question sentence).
[0049] After the candidate text is generated, it enters the consistency verification phase. The text is segmented into sentence-level units according to statement boundaries (periods, question marks, exclamation marks), and each sentence unit is paired with all hard constraint entries in CCS for verification. For price-related hard constraints, the verification phase prioritizes numerical comparison: the product price value is extracted from the sentence unit and directly compared with the quantification threshold stored in the constraint entry; if it exceeds the threshold, it is directly judged as contradictory without going through NLI semantic inference. For semantic constraints such as audience, scenario, and category, the sentence unit is paired with the corresponding constraint content text and fed into a lightweight NLI (Natural Language Inference) model to perform semantic relationship judgment. The initial capabilities of the NLI model are built upon a pre-trained general natural language inference model. Domain-adapted fine-tuning is used to customize constraint verification for product recommendation scenarios. The fine-tuning dataset is manually labeled by the brand operations team for three semantic constraint scenarios: audience constraint (recommended content does not match the user's stated audience age group, such as recommending products "suitable for young people" while the user's constraint is "for elders"), ingredient constraint (recommended ingredients overlap with the user's stated allergens or contraindications), and scenario constraint (recommended scenario positioning does not match the user's stated usage scenario). The default minimum number of labeled data points for each category is 200, with implicit contradiction samples (contradictions that require category knowledge to identify) accounting for at least 40% of each category to ensure the initial model's baseline ability to identify implicit contradictions. The NLI model undergoes knowledge distillation and compression, with a default target upper limit of 50 milliseconds for single inference latency. For each input pair (generated sentence, constraint text), the NLI model outputs three types of relation labels: implication (generated content and constraint direction are consistent), neutral (generated content and constraint do not involve the same aspect), and contradiction (generated content and constraint have directional conflicts). If any sentence returns a contradiction label with any hard constraint, the candidate text in the current batch fails validation and enters the loop-closure negotiation process. The lightweight NLI model may output "neutral" instead of "contradictory" for implicit contradictions that require multiple steps of reasoning to discover (e.g., generated content recommends "products suitable for young people," while the user's constraint is "gift for elders"—this type of contradiction cannot be directly judged from the literal meaning), resulting in missed detections. These missed detections do not trigger loop closures but will manifest as low engagement levels or negative feedback in subsequent user responses. Step S5 records the corresponding strategy effect in the log database, and through periodic batch retraining, gradually corrects the NLI model's judgment boundary for implicit contradictions, achieving continuous iteration of validation capabilities.
[0050] The negotiation and communication between the generative agent and the policy agent is the core of the multi-dedicated agent architecture in this step. The generative agent possesses independent constraint awareness: after generating candidate content, it proactively performs NLI verification and identifies contradictory states, autonomously constructs a structured conflict report, and decides whether to initiate a negotiation request to the collaborative control module, without relying on external commands. Upon receiving the conflict report, the policy agent autonomously determines which specific parameter fields in the SP need revision based on the conflict dimension, completes local revisions within its own decision-making authority, and sends the revised SP' back to the generative agent, without relying on item-by-item commands from the collaborative control module. This bidirectional, autonomous negotiation design ensures that each loop is not a mechanical execution of preset rules, but rather that the two dedicated agents work independently within a complete closed loop of perception-decision-output, collaboratively locating and correcting constraint conflicts through a structured message protocol.
[0051] The triggering and execution process of the loop negotiation procedure is as follows. The system constructs a structured conflict report for conflicting constraint entries. The report fields include the generated sentence text that caused the conflict judgment, the corresponding constraint ID and constraint content text, the intent dimension of the conflict (mapped through the constraint type field), and the cumulative number of loop triggers for that constraint entry. When a generated sentence simultaneously generates conflict judgments with multiple hard constraint entries, the constraint entry with the highest conflict confidence output by the NLI model is taken as the primary conflict source for this loop, and that constraint dimension is processed first. If the revised and regenerated content still triggers conflict judgments for other constraints, secondary conflicts are processed in the next loop. A single loop only processes the primary conflict; revising multiple SP fields simultaneously can cause an uncontrollable shift in the overall direction of the policy parameter package. Gradual revision keeps the impact of each revision within a predictable range, resulting in better policy coherence. After receiving the conflict report, the collaborative control module forwards the intent dimension and constraint content of the primary conflict to the policy agent. The strategy agent performs precise local revisions: only parameter fields in the strategy parameter package SP associated with the main conflict intent dimension are adjusted. If the conflict dimension is price-related (the hard constraint is a price ceiling, and all products in the target SKU set exceed this ceiling), the target SKU set is revised, replacing or adding product options that meet the price constraint, while adjusting the narrative dimension labels accordingly. If the conflict dimension is scenario-related, the scenario adaptation filtering conditions of the target SKU set are adjusted; other strategy parameters remain unchanged. The revised SP' re-triggers the content generation and consistency verification in step S4, forming the second attempt at looping. If the cumulative number of loop triggers for the same hard constraint reaches 2 in the current session, the system determines that it cannot generate content that meets the strategy requirements without violating the constraint within the current intent graph information framework, and switches to a clarification question strategy: the generation agent generates a targeted clarification question based on the constraint dimension that triggered the conflict (e.g., when there are consecutive price-related constraints, generating an open-ended confirmation question such as "Is this for personal use or as a gift? Could you please tell me your approximate budget range?"), returning the initiative to resolve the question to the user and terminating the current loop. Clarifying questions do not involve product recommendations, thus bypassing the validation of product-related hard constraints.
[0052] Candidate texts that pass verification enter the rich media content matching stage. Each piece of material in the brand's rich media resource library (including product images, origin video clips, process illustrations, gift box packaging display images, etc.) carries standardized tags: {Material Type, Associated SKU Set, Narrative Dimension Tag List, Scene Tag List}. The matching algorithm uses the narrative dimension tags in the SP as the main index, retrieves all candidate materials whose narrative dimension tags contain the SP narrative dimension, sorts the candidate materials according to the size of their intersection with the target SKU set, and selects 1 to 2 materials with the largest intersection to be added to the text, forming a combined text + rich media content package.
[0053] Output the final shopping guide content (including the text and rich media combination) that has passed the global hard constraint validation, as well as an updated version of CCS with the new constraint items added in the current round.
[0054] S5 uses the intent tags, strategy types, and user response signals generated in this round of interaction as data sources, and writes them back to the CRM system in a structured manner to complete the closed-loop iterative update of user profiles and preference models; This process is executed after the shopping guide content is delivered to the user, completing the structured extraction and CRM persistence of the interaction data for this round, providing a more accurate historical reference for subsequent conversations, and forming a closed loop of business data.
[0055] During the execution of steps S1 to S4, this round of interaction generates several types of data with cross-session value, corresponding to different dimensions of the user profile. These data are written back to the CRM in a structured manner, providing a foundation for accurate historical reference of subsequent sessions.
[0056] In the intent graph G, explicit and implicit intent tags carry intent type and confidence level, respectively, and need to be merged with the user's existing interest tag system in CRM for calculation: For category or scenario tags appearing in the intent graph, if similar tags already exist in CRM, a sliding weighted update strategy is adopted, with signals generated by recent interactions given higher update weights, and the weights of historical signals decreasing according to a decay strategy as the time distance from the current session increases. The merged weights are normalized and written back to the CRM tag library; For new intent tags that do not exist in CRM, they are directly added to the tag library with the current confidence level, and the session ID and timestamp of the first appearance are marked.
[0057] The strategy effect records generated in each round of interaction are appended to the system strategy effect log library in the format of {Session ID, Current Round Number, Strategy Type Label, Core Fields of Strategy Parameter Package SP (Narrative Dimension Label and SKU Quantity Range), User Response Level Judgment Result, Whether the Session Continues After This Round}. This serves as an offline training data source in two directions: First, it is used as training samples in the offline training phase of the strategy network through periodic batch processing, forming a dual-track iterative mechanism of "online reward-driven real-time optimization + batch retraining of historical effect data" to continuously improve the strategy network's coverage of boutique retail shopping guide scenarios. Second, for implicit contradictions that the NLI model failed to detect in step S4, their impact will be reflected in the user response in subsequent rounds as low participation levels or negative feedback. These records are labeled and included in the periodic retraining dataset of the NLI model to gradually correct the model's judgment boundary for implicit contradictions and achieve continuous iteration of verification capabilities.
[0058] After step S4, the newly added constraint entries in this round of CCS are merged with the CRM cross-session constraint memory: for the newly added hard constraint entries, a corresponding constraint memory entry is created in the CRM and its occurrence count is initialized; if the same or highly semantically similar constraint memory entry already exists in the CRM, its occurrence count is incremented and the most recent occurrence timestamp is updated. Entries whose constraint occurrence count exceeds the cross-session strong confidence threshold are upgraded to strong confidence constraints. The default value of the cross-session strong confidence threshold is 3 times (i.e., the same user mentions the same type of constraint in 3 or more different sessions). After the upgraded strong confidence constraints are output as triples in step S1 of the next session, the session management layer reads H and pre-fills the initial content of CCS, so that the system has the user's high-confidence historical constraint knowledge when a new session starts, avoiding the system having to re-learn the stable preferences that the user has confirmed multiple times at the beginning of each new conversation.
[0059] The behavior sequence B_seq collected in step S1 contains the dwell time distribution for each category. The category interest distribution vector I_new for the current session is calculated based on the proportion of each category's effective dwell time to the total effective dwell time. I_new is then weighted and merged with the user's historical category interest weight vector I_history in the CRM: the default value for recent session weights is 0.3, and the default value for historical weights is 0.7. The merged result is I_new multiplied by 0.3 plus I_history multiplied by 0.7. After merging, the result is normalized and written back to the CRM, completing the sliding update of the interest weights. This weight ratio ensures that the historically accumulated preference distribution remains dominant, while allowing new signals from a single session to moderately influence the overall distribution, avoiding excessive overwriting of historical preferences due to a single abnormal browsing behavior.
[0060] The CRM write-back operations for the above-mentioned data are performed asynchronously after the shopping guide content is sent, without affecting the response speed of the current dialogue interface. Step S5 outputs the updated CRM user records, including the refreshed interest tag system, newly added strategy effect records, updated cross-session constraint memory, and updated category interest weight vectors. Together, these constitute the data source for the historical profile H in step S1 during the next session initialization, enabling the user profile to be continuously refined and improved across multiple sessions.
[0061] Within the complete lifecycle of a dialogue, the processing chain from S1 to S5 runs once for each round of user messages. When a user message arrives in the Nth round (N>1), S1 updates T_raw with the new message text and re-extracts B_seq using (t_start, t_query_N) as the window. This window covers all valid browsing events from the start of the session to the current message time, including any new browsing behavior generated by the user while waiting for a system response. The historical profile H uses the CRM snapshot loaded at the start of the session, and the CRM is not re-queried within this session. The dialogue history text sequence (used for V_T encoding), the confirmed constraint set CCS, and the dialogue state history D_state are continuously accumulated and maintained across rounds within the session. Each new system response and user dialogue are added to the dialogue history, the CCS is updated with the constraint extraction results, and the D_state includes the strategy type record and participation depth level of the current round. When processing user messages in round N, S3 executes two ordered phases: first, it determines the participation depth level of the message, using the result as an immediate reward signal for the strategy decision in round N-1 and updating online parameters; then, it generates the strategy decision for round N based on the updated strategy network. In round 1, there is no previous round strategy, and the participation depth determination result is only used to initialize the historical sequence of D_state, without triggering online updates. S5's CRM write-back operation is performed asynchronously after each round of shopping guide content output. The write-back result is used for loading in the next session step S1 and does not affect the content of the current session H.
[0062] The embodiments of the present invention have been described above. However, the embodiments are not limited to the specific implementation methods described above. The specific implementation methods described above are merely illustrative and not restrictive. Those skilled in the art can make more equivalent embodiments under the guidance of the present embodiments, and all of them are within the protection scope of the present embodiments.
Claims
1. A method for user intent recognition and autonomous intelligent shopping guidance using a multi-agent architecture, characterized in that, Includes the following steps: S1, collect the current conversation behavior event sequence and CRM historical profile, encapsulate the consultation text, behavior event sequence and historical profile into input triples, and merge the consultation text into the dialogue history text; S2, the dialogue history text is encoded into a text semantic vector, the behavioral event sequence is encoded into a behavioral intent vector, the difference between the independent projection values of the two for multiple intent dimensions is used to form a conflict vector, the similarity between the two is used as a consistency score, and the same graph classification label is written into the intent graph. S3 extracts symbolic and amplitude features from the conflict vectors in the intent graph and concatenates them into a conflict structure vector. Combined with consistency score, intent classification label, historical profile category distribution vector and dialogue state vector, it constructs a policy network state input and output policy parameter package. S4 initializes the confirmed constraint set with cross-session constraints in the historical profile and adds constraints extracted from the consultation text. After generating candidate sales guide dialogues guided by the strategy parameter package, each item is verified. If there are contradictions, the strategy parameter package is revised through generation and strategy loop negotiation and regeneration is performed until it passes or the question is switched to clarify. Finally, the sales guide content is output. S5 uses the intent tags, strategy types, and user response signals generated in this round of interaction as data sources, and writes them back to the CRM system in a structured manner to complete the closed-loop iterative update of user profiles and preference models.
2. The user intent recognition and autonomous intelligent shopping guide method with a multi-agent architecture according to claim 1, characterized in that, In S1, the behavioral event sequence collection extracts page interaction events with the time window from the start of the conversation to the arrival of the consultation text. The filtering threshold is dynamically determined by the median duration of the user's historical valid browsing events. New registered users use the historical median of users in the same category as the initial reference. During the cold start phase, it degenerates to the global median of the entire category. When the effective browsing sample size of a category exceeds the preset cold start switching threshold, it switches to the category-specific median. Historical profiles include historical order categories and price ranges, interest tags, cross-session strong confidence constraint items, and historical category interest weight vectors.
3. The user intent recognition and autonomous intelligent shopping guide method based on a multi-agent architecture according to claim 1, characterized in that, In step S2, the text semantic encoding concatenates all user messages and system replies from the start of the session to the current round in chronological order and sends them to the pre-trained language model, taking the sequence pooling output as the text semantic vector; The behavior sequence encoding constructs type embedding, product attribute embedding and duration features for each event, and superimposes relative time position encoding with the last event of the behavior event sequence as the reference point and the time difference between each event and the reference point discretized into time distance level. After processing by a multi-head self-attention layer, the average pooling output of the whole sequence is taken as the behavior intent vector.
4. The user intent recognition and autonomous intelligent shopping guide method based on a multi-agent architecture according to claim 1, characterized in that, In S2, two independent linear projection layers are set for each intent dimension of the text semantic vector and the behavior intent vector, respectively mapping to the corresponding one-dimensional scalar space. The difference between the text-side projection value and the behavior-side projection value is taken as the conflict component of the corresponding dimension. The conflict components of each dimension together constitute the conflict vector. The projection layer and encoder are trained together. The objective function is composed of a weighted sum of policy classification cross-entropy loss and projection orthogonality regularization term. The gradient only updates the parameters of the encoder and projection layer. The cosine similarity between the text semantic vector and the behavior intent vector is used as the consistency score. Each of the two side vectors is processed by an independent intent classification header to output an intent classification label, which is then written into the intent graph along with the conflict vector and the consistency score.
5. The user intent recognition and autonomous intelligent shopping guide method based on a multi-agent architecture according to claim 1, characterized in that, In step S3, symbolic features and amplitude features are extracted from each dimension of the conflict vector in the intent graph, and each is concatenated into a conflict structure vector after independent linear transformation. The conflict structure vector is then concatenated with the consistency score, the one-hot encoding corresponding to the intent classification label, the historical profile category distribution vector, and the dialogue state vector to construct the policy network state input. The dialogue state vector is obtained by compressing the sequence of historical policy types in the session, the sequence of participation depth levels in each round, and the number of the current round through a fully connected coding layer.
6. The user intent recognition and autonomous intelligent shopping guide method based on a multi-agent architecture according to claim 1, characterized in that, In S3, the policy network consists of two fully connected layers and a Softmax output layer connected in series. The first layer is activated by ReLU, and the second layer outputs the probability distributions of five policy types: narrative guidance, direct recommendation, comparison and display, value interpretation and clarification questioning. The low-rank matrix is connected in parallel to the input side of the second layer to generate the adaptation increment. When the highest probability exceeds the preset single-strategy confidence threshold, the corresponding strategy is adopted; otherwise, the two strategy types with the highest probability are selected and a joint strategy is executed. The target SKU set is taken as the union of the two strategy candidates and rearranged according to the matching degree. The narrative dimension tags are combined to form a dual-clue guide.
7. The user intent recognition and autonomous intelligent shopping guide method based on a multi-agent architecture according to claim 1, characterized in that, In S3, user replies are judged into four levels based on the depth of participation. The reward weight of each level increases sequentially according to the level number. When the same reply meets the conditions of multiple levels, the highest level is taken. When the dialogue ends and there is no positive conversion signal, a negative reward is applied. Offline pre-training uses the historical session conversion results as the basis weights multiplied by the participation depth level coefficient to obtain pseudo-rewards, and uses the weighted sum of policy gradient loss and KL divergence constraint terms as the objective function to update all parameters in batches. The online update uses the previous round's state and action pairs as the update objects, and drives the gradient update with the instant reward for participating in the depth level conversion, adjusting only the parameters of the second fully connected layer and the low-rank matrix parameters.
8. The user intent recognition and autonomous intelligent shopping guide method based on a multi-agent architecture according to claim 1, characterized in that, In S4, the constraint extraction module detects constraints on the consultation text in four categories of slots: budget, scenario, ingredient taboos, and audience, based on the slot filling model. The fuzzy price expression is quantized into a numerical threshold by the category price quantile mapping. Constraints containing explicit quantification values or negative qualifiers are classified as hard constraints, while the rest are classified as soft constraints. Cross-session strong confidence constraint entries in the historical profile are immediately filled into the initial content of the confirmed constraint set after the input triple is available. Constraints extracted in the current round are appended to the confirmed constraint set. Only hard constraints trigger subsequent validation loops.
9. The user intent recognition and autonomous intelligent shopping guide method based on a multi-agent architecture according to claim 1, characterized in that, In S4, the candidate sales scripts are divided into sentence units according to the statement boundaries and then verified one by one with the hard constraints in the confirmed constraint set: price hard constraints are judged by numerical comparison, and semantic hard constraints are sent to a lightweight natural language inference model to determine the three types of relationships: implication, neutrality or contradiction. When any sentence is determined to be contradictory to the hard constraint, the process enters the generation to policy loop negotiation, generates an agent to build a conflict report, and the policy agent regenerates the policy parameter package after partially revising the main conflict intention dimension. When the cumulative number of times the same hard constraint is triggered exceeds the preset loop threshold, the strategy is switched to clarification questioning.
10. The user intent recognition and autonomous intelligent shopping guide method based on a multi-agent architecture according to claim 1, characterized in that, In S5, intent tags in the intent graph are merged with existing interest tags in the CRM using a sliding weighting method. Existing tags of the same type are updated with weights, normalized, and written back. New tags are added directly with the current confidence level. The newly added hard constraint entries in the confirmed constraint set in this round are merged with the CRM cross-session constraint memory. If the number of times the same constraint appears in different sessions exceeds the preset strong confidence threshold, it will be upgraded to a strong confidence constraint and pre-populated into the confirmed constraint set in the next session. The category interest distribution vectors in the behavior event sequence of this session are weighted and merged with the historical category interest weight vectors, then normalized and written back to the CRM.