A method and system for seamless switching of full-channel conversations based on reinforcement learning
By fusing and semantically aligning multimodal information using reinforcement learning techniques, the problem of semantic inconsistency and interruption in omnichannel session switching is solved, enabling seamless switching and highly reliable migration of sessions across devices.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- SHANGHAI WICRESOFT
- Filing Date
- 2026-05-28
- Publication Date
- 2026-07-24
AI Technical Summary
Existing technologies lack the ability to collaboratively model and adaptively decide on multimodal dialogue semantics and dynamic environmental factors, resulting in insufficient continuity and reliability of omnichannel session switching. In particular, semantic inconsistencies, switching delays, or session interruptions are prone to occur in complex scenarios.
By employing a reinforcement learning-based approach, multimodal information is acquired and fused for semantic alignment and expression difference analysis to determine the optimal migration path. Adaptive filtering and information completion are then performed to generate a complete migration data packet. Finally, the migration effectiveness is evaluated on the target device to achieve seamless session migration.
It achieves accurate characterization and dynamic representation of semantic consistency across devices, optimizes the determination of migration paths, improves the semantic integrity and consistency during session migration, and enhances the reliability of session migration and the user's seamless experience.
Smart Images

Figure CN122285853B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of multimodal information processing and cross-device session migration technology, and in particular to a method and system for seamless omnichannel session switching based on reinforcement learning. Background Technology
[0002] Omnichannel conversation mechanisms enable cross-device information connection and context continuity. In practical applications, different channels differ in terms of interaction forms, data structures, and device status. Moreover, user interaction is often accompanied by dynamic changes such as network fluctuations and device switching, which makes the dialogue context prone to semantic shifts or information loss during transmission and migration. This places higher demands on conversation continuity and consistency. Reinforcement learning has received widespread attention due to its adaptive decision-making capabilities in dynamic environments, providing a potential technical direction for seamless switching of omnichannel conversations.
[0003] In existing technologies, omnichannel session switching is typically achieved through session data synchronization or context caching mechanisms, which involves transferring historical conversation records between different devices to maintain session continuity. While this method can achieve basic session migration in simple scenarios, its implementation relies heavily on static rules or single-modal data processing, lacking a unified modeling capability for semantic relationships between multimodal information (such as text, voice, and images). Furthermore, during the switching decision-making process, most methods fail to adaptively optimize for dynamic changes in device status and network environment, leading to semantic inconsistencies, switching delays, or session interruptions in complex scenarios, thus impacting the overall user experience.
[0004] In summary, existing technologies lack the ability to collaboratively model and adaptively decide on multimodal dialogue semantics and dynamic environmental factors, resulting in insufficient continuity and reliability of omnichannel session switching. Summary of the Invention
[0005] This invention provides a method and system for seamless switching of omnichannel conversations based on reinforcement learning, in order to solve the problem that existing technologies lack the ability to collaboratively model and adaptively decide on multimodal dialogue semantics and dynamic environmental factors, resulting in insufficient continuity and reliability of omnichannel conversation switching.
[0006] Firstly, to address the aforementioned technical problems, this invention provides a method for seamless omnichannel conversation switching based on reinforcement learning, comprising: Obtain multimodal information from the current device and perform fusion processing on the multimodal information to obtain the initial dialogue context; Based on the initial dialogue context, semantic alignment is performed to obtain the adaptation vector for the target device. Based on the adaptation vector, expression difference analysis is performed to determine the degree of deviation in dialogue semantics. The degree of deviation is compared with a preset deviation threshold to obtain a comparison result. Based on the comparison result, a deep Q-network is used to determine an optimized migration path for session handover. Based on the optimized migration path, the multimodal information is vector-extracted to obtain a semantic representation vector. Based on the semantic representation vector, the semantic consistency index between devices is calculated, and the integrity analysis is performed based on the semantic consistency index to obtain the consistency evaluation result. Based on the consistency evaluation results, adaptive filtering is performed to obtain a corrected semantic representation vector. Based on the corrected semantic representation vector, dialogue state recovery and continuity reconstruction are performed to obtain the corrected dialogue state. Based on the corrected dialogue state, information completion is performed on the initial dialogue context to obtain a completed dialogue context. Based on the completed dialogue context, cyclic interference compensation is performed to obtain a compensation result. Based on the compensation result and the completed dialogue context, a complete migration data packet is generated. Based on the migration data packet, a migration effectiveness assessment is performed on the target device to obtain the effectiveness assessment result, and a state generation process is performed to determine the seamless session migration state.
[0007] Secondly, the present invention provides a seamless omnichannel conversation switching system based on reinforcement learning, comprising: The data acquisition module is used to acquire multimodal information on the current device and perform fusion processing on the multimodal information to obtain the initial dialogue context; The deviation degree determination module is used to perform semantic alignment based on the initial dialogue context to obtain the adaptation vector of the target device, and perform expression difference analysis based on the adaptation vector to determine the degree of deviation of the dialogue semantics. An optimized migration path generation module is used to compare the degree of deviation with a preset deviation threshold to obtain a comparison result, and to determine an optimized migration path for session switching using a deep Q network based on the comparison result. The consistency assessment result determination module is used to extract semantic representation vectors from the multimodal information according to the optimized migration path, calculate semantic consistency index between devices according to the semantic representation vectors, and perform integrity analysis according to the semantic consistency index to obtain the consistency assessment result. The corrected dialogue state generation module is used to perform adaptive filtering based on the consistency evaluation result to obtain a corrected semantic representation vector, and to perform dialogue state recovery and continuity reconstruction based on the corrected semantic representation vector to obtain the corrected dialogue state. The migration data packet generation module is used to perform information completion operation on the initial dialogue context according to the corrected dialogue state to obtain a completed dialogue context, perform cyclic interference compensation based on the completed dialogue context to obtain a compensation result, and generate a complete migration data packet according to the compensation result and the completed dialogue context. The final migration status determination module is used to perform a migration validity assessment on the target device based on the migration data packet, obtain the validity assessment result, perform status generation processing, and determine the seamless session migration status.
[0008] Compared with the prior art, the present invention has the following beneficial effects: (1) This invention performs semantic alignment processing on the initial dialogue context and obtains the target device adaptation vector by combining cross-device semantic mapping. At the same time, it introduces expression difference analysis to quantify and evaluate the deviation of dialogue semantics, thereby achieving accurate characterization and dynamic representation of cross-device semantic consistency. This method breaks through the limitations of existing methods that rely solely on surface semantic matching or simple rule mapping for migration, enabling the implicit semantic shifts caused by differences in expression methods and semantic structures between different devices to be effectively identified and quantified, thus providing a reliable and precise basis for subsequent conversation migration path optimization and semantic deviation correction. (2) This invention determines the degree of semantic deviation in the dialogue by comparing it with a preset deviation threshold, and performs session switching strategy selection and migration path optimization based on the comparison result. Combined with the dynamic decision-making mechanism in the migration process, it determines the optimal migration path and outputs a stable session migration control scheme. This technology establishes a dynamic correlation between the semantic deviation state and the session migration behavior, enabling orderly switching and smooth transition of cross-device sessions even when there are semantic inconsistencies or differences in expression. This effectively avoids the problems of session interruption or semantic confusion caused by semantic mismatch or improper path selection in the prior art. (3) Based on the optimized migration path, this invention extracts key semantic elements from multimodal information by vector and performs integrity analysis by calculating the semantic consistency index between devices, thereby achieving accurate determination and quantitative evaluation of cross-device session semantic consistency. This method introduces semantic feature extraction and integrity analysis into the consistency evaluation mechanism, overcoming the information loss and semantic distortion problems caused by traditional methods that rely solely on a single semantic representation or coarse-grained matching, improving the semantic integrity and consistency during session migration, and providing a reliable basis for subsequent semantic deviation correction and dialogue state recovery; (4) Based on the migration data packet, the present invention performs migration effectiveness evaluation on the target device and combines state generation processing to make a comprehensive judgment on the evaluation results, thereby realizing seamless verification and accurate quantification of cross-device session migration state. This method integrates migration effect evaluation with multi-dimensional indicators into the final state judgment process, overcoming the evaluation bias and migration awareness problems caused by traditional methods that rely solely on a single verification indicator or static threshold judgment, and significantly improving the reliability, stability and seamless user experience of session migration. Attached Figure Description
[0009] Figure 1 This is a schematic diagram of a method for seamless switching of omnichannel conversations based on reinforcement learning provided in the first embodiment of the present invention; Figure 2 This is a schematic diagram of a seamless omnichannel conversation switching system structure based on reinforcement learning, provided in the second embodiment of the present invention. Detailed Implementation
[0010] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0011] Reference Figure 1 The first embodiment of the present invention provides a method for seamless switching of omnichannel conversations based on reinforcement learning, including the following steps: S1, acquire multimodal information on the current device, and perform fusion processing on the multimodal information to obtain the initial dialogue context; S2, based on the initial dialogue context, perform semantic alignment to obtain the adaptation vector for the target device, and perform expression difference analysis based on the adaptation vector to determine the degree of deviation in dialogue semantics; S3, compare the degree of deviation with a preset deviation threshold to obtain a comparison result, and use a deep Q network to determine the optimized migration path for session switching based on the comparison result; S4. Based on the optimized migration path, the multimodal information is vector-extracted to obtain a semantic representation vector. The semantic consistency index between devices is calculated based on the semantic representation vector. Integrity analysis is performed based on the semantic consistency index to obtain a consistency evaluation result. S5. Based on the consistency evaluation result, adaptive filtering is performed to obtain the corrected semantic representation vector. Based on the corrected semantic representation vector, dialogue state recovery and continuity reconstruction are performed to obtain the corrected dialogue state. S6, according to the corrected dialogue state, perform information completion operation on the initial dialogue context to obtain a completed dialogue context, perform cyclic interference compensation based on the completed dialogue context to obtain a compensation result, and generate a complete migration data packet based on the compensation result and the completed dialogue context; S7. Based on the migration data packet, perform a migration validity assessment on the target device, obtain the validity assessment result, and perform state generation processing to determine the seamless session migration state.
[0012] In step S1, acquiring multimodal information on the current device and fusing the multimodal information to obtain an initial dialogue context includes: The text, voice, and visual data collected from the current device interface are acquired as the multimodal information, and the multimodal information is encoded to obtain the original multimodal feature vector. Based on the original multimodal feature vectors, the semantic associations between different modalities are modeled using a preset multimodal fusion model to obtain a multimodal joint representation vector; Based on the multimodal joint representation vector, the intermodal similarity between different modalities is calculated, and the modal features with intermodal similarity below a preset similarity threshold are subjected to weight adjustment processing to obtain a weighted fusion vector; When the similarity between the modalities is greater than the preset similarity threshold, it is determined that the semantic association between the modalities is reliable, and the current fusion result has good semantic consistency, so no additional weight adjustment is required. Based on the weighted fusion vector, a difference analysis is performed on the semantic features expressed by different devices to obtain semantic difference measurement results; Based on the semantic difference measurement results, the multimodal joint representation vector is corrected to obtain a multimodal corrected vector, and the multimodal corrected vector is then subjected to context fusion processing to generate the initial dialogue context.
[0013] In one implementation, to determine the multimodal original feature vector, this embodiment first acquires text, speech, and visual data collected by the current device interface, and encodes the different modal data respectively. Specifically, semantic encoding is performed on the text data to obtain a text feature vector, acoustic feature extraction is performed on the speech data to obtain a speech feature vector, and image feature extraction is performed on the visual data to obtain a visual feature vector. Subsequently, the modal features are subjected to unified dimensionality mapping and alignment processing to obtain the multimodal original feature vector.
[0014] For example, this embodiment automatically collects multimodal information on the current device through the built-in sensors and API interfaces of the device. For example, it uses the MediaRecorder API of the Android system to collect voice signals at a sampling rate of 16kHz, and calls the Camera2 API to collect visual frame sequences at a resolution of 1920x1080. It also scans the screen text content through the Accessibility Service API to form an initial dataset. The voice data is converted into Mel-frequency cepstral coefficient features, the visual data is extracted into an RGB image matrix, and the text data is parsed into a UTF-8 encoded string.
[0015] In one implementation, to generate the multimodal joint representation vector, this embodiment constructs a multimodal fusion model based on the original multimodal feature vectors to model the semantic relationships between different modalities. Specifically, the text feature vector, speech feature vector, and visual feature vector are input into the corresponding linear mapping layer to obtain the query vector Q, key vector K, and value vector V for each modality. Then, the query vector Q and key vector K between each modality are multiplied by a dot product and divided by a dimension scaling factor before normalization to obtain the attention weight coefficients between modalities. The attention weight coefficients are then weighted and summed with the corresponding value vector V to obtain the fused feature representation of each modality. Finally, the fused feature representations of each modality are concatenated and input into a fully connected layer for mapping, outputting a multimodal joint representation vector of uniform dimension.
[0016] It should be noted that the multimodal fusion model is a Transformer-based cross-modal attention network, which is trained using 100,000 sets of aligned text, speech, and visual triplet data. The goal is to minimize the contrast loss between features of each modality. The Adam optimizer is used with a learning rate of 0.001, and the training is conducted for fifty rounds until the loss converges.
[0017] In one implementation, regarding the intermodal similarity calculation and weight adjustment, this embodiment further evaluates and optimizes the internal consistency of the obtained multimodal joint representation vector. The system calculates the semantic similarity between the original feature vectors of different modalities, for example, using cosine similarity to measure the degree of association between multiple modal combinations such as text vectors and speech vectors, text vectors and visual vectors. The calculated similarity values are compared with a preset similarity threshold. If the similarity of a certain modal combination is lower than the threshold, it is determined that the contribution of that modality in the current context may be unreliable due to noise, ambiguity, or missing information. Subsequently, the contribution weight of the modality's features in the fusion vector is dynamically increased to strengthen the consistency of the overall semantic expression, thereby obtaining an optimized weighted fusion vector.
[0018] It is worth noting that the dynamic enhancement refers to multiplying the original feature vector of a certain modality by an increment coefficient greater than one when the modality similarity of a certain modality is lower than a threshold. The increment coefficient is the square of the quotient of the threshold divided by the modality similarity, but the upper limit of the increment coefficient does not exceed 2.0.
[0019] It should be noted that the preset similarity threshold is used to determine whether the semantic association between feature vectors of different modalities reaches the standard for reliable fusion. Specifically, in this embodiment, during the offline training phase, a large number of labeled, semantically aligned multimodal dialogue data samples are used to calculate the similarity between each pair of modal feature vectors and to statistically analyze their distribution. For example, in samples where the speech and text modalities are correctly aligned, the cosine similarity between their encoding vectors is calculated to obtain their historical data distribution, and a specific quantile of this distribution, such as the fifth percentile, is used as the reference threshold for determining the similarity of the modal pair. Simultaneously, similarity data for corresponding modal pairs is collected from abnormal samples containing significant noise, missing information, or semantic ambiguity to obtain their outlier distribution range. Finally, by comparing the similarity data distributions of normal samples and abnormal samples, a critical value that can effectively distinguish between the two is selected and determined as the final similarity threshold.
[0020] In one implementation, for the semantic difference analysis between the devices, this embodiment aims to identify and quantify the underlying feature differences introduced by differences in acquisition devices, environments, or their own expression methods. The system performs dimensionality reduction on the weighted fusion vector to focus on the most significant semantic features. In the dimensionality-reduced feature space, the distance or divergence between feature distributions from different modalities or different data streams is calculated and compared, for example, Euclidean distance or KL divergence is calculated, as a quantitative indicator of semantic differences. By performing statistical analysis on these difference indicators, the system can locate modalities or feature dimensions with significant differences and output a comprehensive semantic difference measurement result.
[0021] In one implementation, for the generation of the initial dialogue context, this embodiment generates a correction matrix or compensation vector based on the semantic difference measurement result, and performs reverse correction on the multimodal joint representation vector to eliminate or reduce the bias introduced by the identified device expression differences. The corrected multimodal semantic vector is input into a preset sequence modeling network, such as a long short-term memory network. This network takes the preceding dialogue history (if any) and the current corrected multimodal vector as input, models its semantic dependencies and evolution relationships over time, and finally outputs a vector containing the current user intent, dialogue topic, and contextual information, i.e., the initial dialogue context; this sequence modeling network is trained using 50,000 multi-turn dialogue sequences, with the goal of predicting the dialogue state at the next moment, using the cross-entropy loss function and a stochastic gradient descent optimizer, and is trained for thirty rounds.
[0022] It should be noted that the preset sequence modeling network is used to model the semantic dependencies of the corrected multimodal vectors across time. This network can adopt a recursive architecture such as a Long Short-Term Memory (LSTM) network or a Gated Recurrent Unit (GRU). Its input includes the previous dialogue history state vector and the current corrected multimodal vector, and its output is a semantic vector containing the current user intent, dialogue topic, and contextual information.
[0023] In step S2, the process of obtaining the target device's adaptation vector through semantic alignment based on the initial dialogue context, and performing expression difference analysis based on the adaptation vector to determine the degree of deviation in dialogue semantics, includes: Obtain the preset semantic template of the target device, perform semantic alignment based on the initial dialogue context, and obtain the adapted semantic vector of the target device; The semantic alignment similarity is calculated based on the adapted semantic vector and the preset semantic template, and the semantic alignment similarity is compared with the preset semantic alignment threshold. When the semantic alignment similarity is less than the semantic alignment threshold, the adapted semantic vector is dimensionality reduced and mapped to obtain the aligned semantic vector. When the semantic alignment similarity is greater than the semantic alignment threshold, it is determined that the current adapted semantic vector and the preset semantic template of the target device have sufficient semantic alignment, and the current adapted semantic vector is directly output as the aligned semantic vector. Based on the alignment semantic vector, a difference analysis is performed on the semantic distribution difference between the current device and the target device to obtain a semantic difference metric. The semantic difference metric is compared with a preset difference threshold. When the semantic difference metric is greater than the difference threshold, the aligned semantic vector is corrected to obtain an updated semantic vector. The semantic drift degree is obtained by performing temporal analysis based on the updated semantic vector, and then weighted and fused with the semantic alignment similarity and the semantic difference metric to determine the degree of deviation in the dialogue semantics.
[0024] In one implementation, to obtain the preset semantic template of the target device, this embodiment first loads the preset semantic template from the local configuration of the target device. The preset semantic template is a high-dimensional vector set, which is pre-constructed by collecting a large amount of standard interactive corpus on the target device and extracting its deep semantic features, representing the typical semantic expression space of the device; specifically, corpus of all typical interactive scenarios supported by the device is collected, with at least two hundred user statements collected for each scenario. A vector is generated for each statement using a pre-trained semantic encoder, and the average value of all vectors in the same scenario is taken to obtain the template vector for that scenario. The template vectors of all scenarios are combined into a set, which is the preset semantic template.
[0025] In one implementation, the formula for calculating semantic alignment similarity can be expressed as:
[0026] Among them, the adaptation semantic vector is It is a set of one or more semantic template vectors pre-defined for the target device. ,in . For vector dimensions. The Euclidean norm of a vector. The value range is [-1, 1].
[0027] Furthermore, since there are multiple template vectors, the calculated set of similarities is aggregated to generate a comprehensive semantic similarity. This embodiment uses a maximum value aggregation strategy, namely:
[0028] It is worth noting that maximum value aggregation can effectively avoid misjudgments caused by mismatch of a single atypical template, thus more robustly characterizing alignment quality.
[0029] In one implementation, to generate the aligned semantic vector, this embodiment performs dimensionality reduction mapping on the adapted semantic vector. Specifically, the adapted semantic vector is dimensionality reduced by decomposing its high-dimensional semantic features, calculating the contribution of each feature component, and sorting the features according to their contribution. Principal components are selected based on a preset principal component retention ratio, and redundant features with low contribution are removed. The retained principal components are then linearly combined and reconstructed to obtain a low-dimensional semantic representation, which is then output as the aligned semantic vector.
[0030] For example, suppose the preset semantic alignment threshold is 0.75, and the calculated similarity is 0.68. Since 0.68 < 0.75, the initial alignment is deemed insufficient, and a dimensionality reduction mapping process is triggered. This process uses techniques such as principal component analysis to project the adapted semantic vector from a high-dimensional space of, for example, 1024 dimensions, to a low-dimensional subspace of, for example, 128 dimensions. This process filters out device-specific noise that may exist in the high-dimensional space and is irrelevant to the core semantics, thereby outputting an aligned semantic vector that is more focused on the essential semantic information.
[0031] It should be noted that this embodiment determines the semantic alignment threshold through offline statistical methods. First, a large number of cross-device dialogue sample pairs with correct semantic alignment are collected, and the semantic alignment similarity between the adaptation vector and the corresponding target device semantic template is calculated to obtain the positive sample similarity distribution. Simultaneously, similarity data is collected on abnormal sample pairs with semantic alignment deviations or mapping distortions to obtain the negative sample similarity distribution. Then, by comparing the similarity distribution curves of positive and negative samples, a critical value that can effectively distinguish between reliable and unreliable alignment is selected, thereby determining the semantic alignment threshold.
[0032] In one implementation, for generating the semantic difference metric, this embodiment uses KL divergence as the difference metric algorithm. The semantic probability distribution represented by the aligned semantic vector is compared with the baseline distribution represented by the preset template of the target device, and the KL divergence value between the two is calculated. This value is quantified as the semantic difference metric.
[0033] In one implementation, to determine the updated semantic vector, this embodiment compares the semantic difference metric with a preset difference threshold. When the semantic difference metric is greater than the difference threshold, it is determined that there is a significant difference, and the aligned semantic vector is then corrected. Specifically, based on the excess difference metric, a non-linear correction factor is generated, and the amplitude and direction of the aligned semantic vector are adjusted to obtain the updated semantic vector. The difference threshold is determined by collecting one thousand samples with known semantic difference metric values from cross-device sessions, of which 30% are labeled as having significant differences. The samples are then sorted in ascending order of difference metric value, and the value corresponding to the 80th percentile is taken as the difference threshold.
[0034] Furthermore, regarding the correction process, its core is to map the degree of excess difference to a correction factor using an adaptive nonlinear function, and then use this factor to scale the original aligned semantic vector to compensate for the identified systematic expression biases. Specifically, the relative difference intensity exceeding the threshold is first calculated, and then this intensity value is converted into a smooth, bounded scaling coefficient through a nonlinear mapping function based on the hyperbolic tangent function. This coefficient is directly multiplied by the aligned semantic vector, thereby enhancing or weakening its amplitude while maintaining the semantic direction essentially unchanged, to counteract the inherent attenuation or distortion caused by cross-device mapping.
[0035] In one implementation, for generating the semantic drift, this embodiment takes the sequence of the updated semantic vector in the time dimension as input and quantifies the dynamic shift of the semantic representation through time series analysis. Specifically, it obtains the updated semantic vectors for multiple consecutive time steps within the current time window, calculates the change in vectors between adjacent time steps, for example, using the L2 norm or cosine distance to measure the difference between adjacent vectors. Subsequently, it statistically aggregates the changes of all adjacent vectors within the window, such as by taking the mean or weighted summation, to obtain a scalar value reflecting the overall semantic shift trend, i.e., the semantic drift.
[0036] In step S3, the deviation level is compared with a preset deviation threshold to obtain a comparison result. Based on the comparison result, a deep Q-network is used to determine an optimized migration path for session handover, including: Obtain the source dialogue semantic vector from the initial dialogue context; The degree of deviation is compared with a preset deviation threshold. When the degree of deviation is greater than the deviation threshold, a deviation metric is calculated between the source dialogue semantic vector and the updated semantic vector to obtain a semantic distance value. When the deviation is less than the deviation threshold, the preset migration path is maintained as the migration path for the session switching.
[0037] Based on the semantic distance value, the dynamic changes during the session migration process are monitored in real time to obtain the change trend parameters; Based on the change trend parameter and the semantic distance value, multiple candidate migration paths are generated; A deep Q-network is used to evaluate the multiple candidate migration paths to obtain a value estimate for each candidate migration path, and the candidate migration path with the highest value estimate is selected as the optimized migration path for session switching.
[0038] In one implementation, for calculating the semantic distance value, this embodiment first obtains the initial dialogue context vector representing the dialogue state of the source device, and the updated semantic vector representing the state of the target device after correction in step S2. The semantic distance value is obtained by calculating the weighted Euclidean distance between these two high-dimensional vectors.
[0039] It is worth noting that the weighting coefficients are determined during the model training phase. Their role is to amplify the difference contribution of core semantic dimensions (such as user intent and dialogue focus) that have a significant impact on conversation coherence, so that the calculated semantic distance value can more accurately reflect the key semantic gaps that affect the success of transfer.
[0040] It should be noted that the deviation threshold is used to characterize the acceptable range of semantic offset. Specifically, it is set by statistically analyzing the degree of semantic deviation based on historical cross-device session migration data, extracting the deviation distribution interval under normal migration conditions, and determining the boundary value by combining abnormal migration samples. The deviation threshold is then determined by weighted statistics or percentile methods. For example, two thousand sets of historical cross-device session migration data are collected, the degree of semantic deviation before each migration is calculated, and the user satisfaction after migration is marked. The 90th percentile of the degree of deviation among all satisfactory samples is taken as the deviation threshold.
[0041] It should be noted that the preset migration path is a full data synchronization path, which means that the conversation migration is performed step by step according to the default order of the device at the factory, such as first synchronizing the text dialogue history, then synchronizing the voice cache, and finally synchronizing the visual context, without any dynamic optimization or path selection.
[0042] In one implementation, regarding the determination of the change trend parameter, this embodiment takes the semantic distance value sequence of continuous time steps as input, and uses a sliding window mechanism to statistically analyze the change characteristics of the distance values, such as calculating the rate of change, fluctuation variance, or trend slope of the semantic distance values within the window, to capture the dynamic evolution of semantic deviation on the session migration path.
[0043] For example, suppose the semantic distance value sequence for m consecutive time steps within the current time window is [d_1, d_2, ..., d_m], then the rate of change can be expressed as (d_m-d_1) / (m-1), and the variance of the fluctuation can reflect the degree of dispersion of the distance values. Combining the above statistical characteristics forms a set of multidimensional parameters, namely the trend parameters.
[0044] In one implementation, for generating the multiple candidate migration paths, this embodiment abstracts the possible migration channels between the current device and the target device into a path graph structure, where nodes represent intermediate states in the migration process, and edges represent allowed migration operations between states. Using the semantic distance value as the cost benchmark for the current state, and the rate of change and fluctuation characteristics in the trend parameters as dynamic constraints, the path graph is expanded and searched to retain all reachable paths that meet preset feasibility conditions. Each path starts from the current state, goes through several intermediate states, and arrives at the target device state, forming a complete candidate migration path. By traversing the path graph, all paths that meet the conditions are output, thus generating the multiple candidate migration paths.
[0045] In one implementation, to determine the optimized migration path, this embodiment utilizes a Deep Q-Network (DQN) to evaluate and select strategies from the generated multiple candidate migration paths. Specifically, the semantic distance value at the current moment is combined with a trend parameter to construct the state vector of the DQN. The semantic distance value reflects the degree of core semantic difference between the source and target devices, while the trend parameter captures the dynamic evolution of the semantic deviation, such as the rate of change and variance. The multiple candidate migration paths constitute the action space of the DQN, with each action corresponding to a specific candidate migration path.
[0046] It should be noted that this embodiment pre-trains a deep Q-network, whose input is a state vector and output is a value estimate for each action. During the execution phase, the current state vector is input into the DQN to obtain the value estimates of all candidate paths, and the action with the highest value estimate is selected as the optimal migration path for the session switching. Specifically, the state space of the deep Q-network is a continuous vector composed of semantic distance values and trend parameters, and the action space is a discrete set of all possible candidate migration paths. The reward function is a comprehensive reward value constructed after a migration is completed, based on the reconstructed dialogue semantic consistency index on the target device, the end-to-end migration latency, and whether the user perceives the switch. This reward value can be generated through manual annotation or a post-hoc questionnaire. For example, the reward for high semantic consistency, low latency, and imperceptible migration is +1, while the reward for significant semantic deviation or interruption is -1. The training data consists of 20,000 migration samples collected continuously over one month from real user interaction data. Each sample records the state, action, reward, and next state. The network is trained using the standard experience replay and target network mechanism of DQN, such as random sampling from the experience replay pool, until the network converges.
[0047] To ensure robustness of the migration, this embodiment also sets a minimum value threshold. If the value estimates of all candidate paths are lower than the minimum value threshold, the current DQN decision is deemed unreliable, and the process reverts to a preset default migration path, such as a full data synchronization path, to ensure that the session migration process is not interrupted or erroneous. The minimum value threshold is determined through offline simulation experiments. In a simulated session migration environment, a large number of semantic deviation states and candidate path sets are randomly generated. The optimal path value estimate for each state is calculated, and the distribution of the minimum value estimate in the successful migration samples is statistically analyzed. The lower quartile of this distribution is taken as the minimum value threshold to ensure that the path output by the deep Q network can be selected in most reliable scenarios, and the process reverts only in cases of extreme uncertainty.
[0048] In step S4, the process of extracting semantic representation vectors from the multimodal information based on the optimized migration path, calculating semantic consistency indices between devices based on the semantic representation vectors, and performing integrity analysis based on the semantic consistency indices to obtain consistency evaluation results includes: Based on the optimized migration path, multimodal key semantic elements are extracted from the multimodal information, and the key semantic elements are encoded to obtain a semantic representation vector. Based on the semantic representation vector, calculate the semantic consistency index between the current device and the target device; When the semantic consistency index is lower than the preset integrity threshold, the semantic information of the current device is determined to be incomplete, and the semantic deviation is calculated based on the semantic representation vector to obtain the information loss value. When the semantic consistency index is higher than the preset integrity threshold, the semantic migration from the source device to the target device is determined to be successful and complete. At this time, the optimized migration path will be confirmed to be effective, and the conversation will be seamlessly restored or continued based on the semantic context reconstructed on the target device, without any retransmission or additional user confirmation.
[0049] Based on the information loss value, the data packets during transmission are parsed to identify abnormal data packets and determine the range of recoverable data. Based on the recoverable data range, a recovery data packet sequence is constructed, and the consistency of multimodal information is weighted and calculated in conjunction with the semantic consistency index to obtain a consistency evaluation result.
[0050] In one implementation, for generating the semantic representation vector, this embodiment extracts core content from the multimodal information flow with a focus based on the optimized transfer path determined in step S3. The key semantic elements refer to information units crucial for maintaining dialogue coherence and understanding user intent, such as intent keywords and named entities in speech-to-text results, and dialogue-related target objects and their attributes identified in visual information. A feature extractor matching the optimized transfer path is invoked to process the text, speech, and visual modal data separately, extracting their high-level semantic features. These features are then fused and compressed to ultimately generate a fixed-dimensional semantic representation vector.
[0051] In one implementation, the semantic consistency index is calculated using a multi-step quantization process. First, the semantic core vector extracted by the source device and the corresponding semantic vector reconstructed by the target device are acquired and normalized. Then, the weighted cosine similarity between the two is calculated, where the weights of each semantic dimension are obtained through offline learning to reflect the differences in importance of different dimensions to dialogue understanding. If the vector includes dimension-level confidence information, the lowest confidence scores of both parties are further integrated to weight and correct the basic similarity, thereby improving the index's robustness to noise. Finally, the corrected similarity value is converted into a final semantic consistency index ranging from 0 to 1 using a preset S-shaped mapping function. .
[0052] For example, the semantic deviation is obtained by subtracting the final semantic consistency index from 1; based on this, a preset integrity threshold is also considered. The semantic information loss is normalized to obtain the information loss value, which is expressed as:
[0053] Where L represents the information loss value, used to characterize the degree of missing semantic information relative to the integrity requirements; when hour, This indicates that the semantic information is complete. When hour, Furthermore, the larger the value, the more severe the semantic loss. The integrity threshold is calculated by collecting one thousand sets of session migration data, calculating the semantic consistency index after migration, and marking whether there is information loss in the dialogue after migration. Using grid search, candidate thresholds are traversed between 0.5 and 0.95 with a step size of 0.01, and the value that makes the missing detection accuracy the highest is selected as the integrity threshold.
[0054] In one implementation, regarding the determination of the recoverable data range, this embodiment first dynamically selects the data packet parsing depth mode based on the magnitude of the information loss value. Further, under the selected parsing mode, the system decodes the data packets in the transmission link, identifies abnormal data packets that have failed verification, been lost, or timed out through checksum verification, sequence number continuity analysis, and timestamp comparison, and organizes them into an ordered list or a continuous interval according to their sequence numbers. This list or interval is the determined recoverable data range.
[0055] In one implementation, regarding the construction of the recovered data packet sequence, this embodiment reassembles and sorts the effective data segments according to the recoverable data range to construct the recovered data packet sequence. The recovered data packet sequence is then fused with the current semantic representation vector, and the consistency of the multimodal information is weighted and calculated in conjunction with the semantic consistency index to obtain a consistency evaluation result.
[0056] In step S5, the process of performing adaptive filtering to obtain a corrected semantic representation vector based on the consistency evaluation result, and then performing dialogue state recovery and continuity reconstruction based on the corrected semantic representation vector to obtain the corrected dialogue state includes: Based on the consistency evaluation results, a semantic deviation state vector is constructed, and the semantic deviation state vector is subjected to adaptive filtering to obtain a corrected semantic representation vector. The contextual relationships are extracted based on the modified semantic representation vector, and the recovered data packet sequence is parsed. Based on the contextual relationships and the parsing results, the dialogue state is reconstructed to obtain the preliminary dialogue state. Based on the initial dialogue state, a dialogue state graph model is constructed, and node feature propagation optimization is performed on the dialogue state graph model to output an enhanced dialogue state. Based on the enhanced dialogue state, the dialogue sequence between the current device and the target device is aligned and analyzed to obtain a dialogue continuity index. When the dialogue continuity index is less than a preset continuity threshold, the missing dialogue information is predicted and completed to obtain a complete dialogue state. Based on the complete dialogue state, the modified semantic representation vector, the enhanced dialogue state, and the dialogue continuity index are fused and calculated to determine the modified dialogue state.
[0057] In one implementation, to generate the corrected semantic representation vector, this embodiment first constructs a semantic deviation state vector containing the current semantic deviation value and its changing trend based on the consistency evaluation result. Specifically, the vector is recursively estimated using an adaptive Kalman filter algorithm. In the prediction step of the algorithm, the deviation state at the current time step is predicted based on the corrected state and state transition model of the previous time step. In the update step, the predicted value is fused with the new deviation value directly observed from the current consistency evaluation result, and the prediction is corrected by calculating the optimal Kalman gain, thereby outputting a more stable and accurate deviation estimation vector.
[0058] It should be noted that the process noise covariance matrix and measurement noise covariance matrix of the adaptive Kalman filter are both obtained by maximum likelihood estimation of the deviation states in 500 sets of real migration data; the process noise covariance matrix is taken as the covariance of the state prediction error, and the measurement noise covariance matrix is taken as the error covariance between the observed value and the true value.
[0059] Furthermore, this bias estimation vector is used to perform reverse compensation on the original semantic representation vector. For example, if the estimated bias value is positive (indicating that the semantic expression is too strong), the vector is scaled down proportionally. If it is negative (indicating that the semantic expression is too weak), it is scaled up proportionally, ultimately resulting in a corrected semantic representation vector with the bias dynamically suppressed.
[0060] In one implementation, for the dialogue state reconstruction process, this embodiment does not directly concatenate the recovery data packets, but rather reassembles them by combining semantic structures. Specifically, a correlation matrix between semantic features is calculated based on the modified semantic representation vector to characterize the dependencies between semantic units in the context; simultaneously, the recovery data packet sequence is parsed to extract semantic fragments and temporal sequence information. The semantic fragments are mapped to corresponding positions in the correlation matrix, and then reordered and combined according to the semantic association strength to restore the dialogue logic structure and obtain the initial dialogue state.
[0061] For example, when the parsed semantic fragment includes "query weather", "current city" and "temperature information", and the correlation weight between "query weather" and "current city" is calculated to be 0.92 through the correlation matrix, which is higher than other combinations, they are prioritized for combination to form the complete semantic unit "query current city weather", thereby improving the accuracy of semantic reconstruction.
[0062] In one implementation, for constructing the dialogue state graph model, this embodiment constructs the semantic units in the initial dialogue state as a graph structure. Specifically, each semantic fragment is treated as a graph node, semantic dependencies as edges, and weight values are assigned to the edges to represent the semantic association strength. The nodes in the graph are iteratively updated through a node feature propagation mechanism, enabling each node to integrate the semantic information of its neighboring nodes, thereby obtaining an enhanced dialogue state.
[0063] Furthermore, the node feature propagation mechanism collects semantic feature information of its neighboring nodes centered on each node during the feature propagation process, and performs weighted processing on the neighbor features according to the semantic association strength between nodes. The weighted neighbor features are then fused with the current node's own features to obtain the updated node features.
[0064] It should be noted that after multiple rounds of such message passing and feature aggregation, the final feature representation of each node not only includes its own initial information but also incorporates the contextual information of the entire local subgraph connected to it. This makes logically related nodes in the graph closer in the feature space, thereby improving the internal consistency and coherence of the entire dialogue state graph.
[0065] For example, in the ordering dialogue, there is an edge between the "dish" node and the "spiciness" node. After optimization, the features of the "dish-boiled fish" node will be incorporated into the features of the "spiciness-medium spicy" node, so that the representation of "boiled fish" includes an implicit confirmation of spiciness, which enhances the internal consistency of the state.
[0066] In one implementation, the dialogue continuity index is calculated using alignment analysis. Specifically, the dialogue sequences between the current device and the target device are mapped to a unified timeline, and the continuity index is obtained by calculating the matching degree between corresponding semantic units. This index can be expressed as the ratio of the number of matched semantic units to the total number of semantic units.
[0067] Furthermore, when the continuity index is less than the continuity threshold, a significant continuity gap is identified, and information completion is performed. Specifically, a sequence prediction model trained from historical dialogue sequences is used to perform semantic inference on the missing positions and insert corresponding semantic units to restore the complete dialogue state. When the continuity index is greater than or equal to the threshold, the current dialogue state is kept unchanged, and the current dialogue state is directly output as the complete dialogue state for subsequent fusion calculation steps.
[0068] It should be noted that the continuity threshold is obtained through statistical analysis of the dialogue integrity during the historical conversation migration process. Specifically, eight hundred sets of continuous dialogue sequences are collected, half of which are complete sequences and the other half are incomplete sequences after some semantic units have been manually deleted. The dialogue continuity index of all sequences is calculated, and the optimal dividing point that distinguishes between complete and incomplete sequences, i.e., the value that minimizes the classification error rate through binary search, is taken as the continuity threshold.
[0069] In one implementation, after obtaining the complete dialogue state, this embodiment determines the final correction result through fusion calculation. Specifically, the corrected semantic representation vector, the enhanced dialogue state, and the dialogue continuity index are used as inputs to construct a weighted fusion function for comprehensive calculation to obtain the final corrected dialogue state.
[0070] For example, when the corrected semantic similarity is 0.91, structural integrity is 0.94, and continuity index is 0.92, the weighted fusion calculation yields a comprehensive state score of 0.925, indicating that the current dialogue has recovered to a stable state and can proceed to the next stage of migration processing. It should be noted that for the weighted fusion of the three indicators—semantic alignment similarity, semantic difference measure, and semantic drift—in step S2, a linear regression method is used. One thousand samples with known deviation levels are collected, with the three indicators as independent variables and the manually labeled deviation level as the dependent variable. The least squares regression method is used to obtain three weight coefficients, which are then normalized to make their sum equal to one. For the fusion calculation of the corrected semantic representation vector, enhanced dialogue state, and dialogue continuity index in this embodiment, the same regression method is used, with the final migration success rate as the objective to determine the weights.
[0071] In step S6, the process of performing information completion on the initial dialogue context based on the corrected dialogue state to obtain a completed dialogue context, performing cyclic interference compensation based on the completed dialogue context to obtain a compensation result, and generating a complete migration data packet based on the compensation result and the completed dialogue context includes: Based on the corrected dialogue state, semantic alignment and completion operations are performed on the missing information in the initial dialogue context to obtain the completed dialogue context. Based on the completed dialogue context, loop detection is performed on the optimized migration path to obtain cyclic interference parameters; When the cyclic interference parameter is greater than the interference threshold, the optimized migration path is de-looped to obtain a loop-free migration path. Data integration processing is performed based on the completed dialogue context and the loop-free migration path to generate a complete migration data package.
[0072] In one implementation, for the semantic alignment and completion operation, this embodiment uses the corrected dialogue state as a reference to recover missing information in the initial dialogue context caused by transmission loss or device switching. Specifically, the corrected dialogue state and the initial dialogue context are semantically aligned to identify the semantic gap between them, i.e., to determine the missing semantic segments in the initial dialogue context. Subsequently, a contextual reasoning mechanism is used to predict and generate the missing segments, and the generated completion content is semantically fused with the initial dialogue context to form a semantically coherent and information-complete completed dialogue context.
[0073] It should be noted that the context reasoning mechanism takes the complete semantic information contained in the corrected dialogue state as a condition, combines the semantic relationship between the missing segments in the initial dialogue context and the context, and infers and generates complete content that is semantically coherent with the context through an attention mechanism.
[0074] For example, suppose the initial dialogue context is missing the user's preference for "coffee temperature" (e.g., "without ice"), while the corrected dialogue state includes this information inferred from subsequent interactions. The completion model identifies this missing information and generates a feature vector fragment representing "preference: without ice". The system fuses this fragment with the initial context vector along the corresponding dimensions to output the completed dialogue context.
[0075] In one implementation, regarding the determination of the cyclic interference parameters, this embodiment represents the optimized migration path as a directed graph structure, where nodes correspond to states during the migration process, and edges correspond to migration operations between states. By traversing this directed graph structure, it is detected whether a closed loop exists, i.e., whether it is possible to return to a node after traveling along a directed edge from a certain node. If a loop is detected, features such as the number of state nodes involved in the loop, the loop length, and the frequency of loop occurrence are extracted, and these features are quantified as the cyclic interference parameters.
[0076] Furthermore, when the cyclic interference parameter is greater than the interference threshold, the optimized migration path undergoes delooping. Specifically, the loop portion in the directed graph structure corresponding to the optimized migration path is identified, and closed loops are broken by removing redundant back edges or merging loop nodes, transforming the path structure into a directed acyclic graph. The path structure obtained after delooping is the acyclic migration path. When the cyclic interference parameter is less than the interference threshold, the migration path structure is determined to be stable, requiring no structural adjustment, and the current path structure is directly retained for subsequent processing.
[0077] It should be noted that the interference threshold is used to determine the degree of cyclic interference. It is set by statistically analyzing historical session migration path data, extracting the structural feature differences between normal migration paths and abnormal cyclic paths, and determining the critical boundary value based on statistical distribution, thereby achieving effective identification of cyclic interference. For example, five hundred sets of migration path graph structure samples are collected, their cyclic interference parameters are calculated, and it is marked whether the path causes a dead loop or significant delay in the migration process. The 95th percentile of the cyclic interference parameters in all samples without dead loops is taken as the interference threshold.
[0078] In one implementation, to generate a complete migration data packet, this embodiment serializes and encodes the dialogue context. Simultaneously, key path metadata is extracted from the optimized migration path structure, including the final determined sequence of migration steps, the target device or status identifier for each step, and the overall quality assessment of the path.
[0079] Further, data integration processing is performed. The serialized and completed dialogue context is used as the core data payload, the path metadata is used as control header information, and necessary fields such as timestamps, session IDs, and integrity check codes are added and encapsulated according to a preset communication protocol format.
[0080] It should be noted that the integrity check code (such as a SHA-256 hash) is calculated based on the core data payload and critical control header information, and is used to ensure the integrity of the data packet during the final transmission process. The generated complete structure containing data, control information, and the check code constitutes the complete migration data packet.
[0081] It should be noted that when generating the migration data packet, the semantic consistency index calculated in step S4, the intermediate state parameters recorded in step S5 during the adaptive filtering process, and the cyclic interference parameters detected in step S6 are also encapsulated into the extended fields of the migration data packet as channel state parameters and cyclic interference parameters required for subsequent migration effect evaluation.
[0082] In step S7, the migration effectiveness assessment is performed on the target device based on the migration data packet to obtain the effectiveness assessment result, and a state generation process is performed to determine the seamless session migration state, including: Based on the migration data packet, the dialogue state sequence is reconstructed on the target device to obtain the reconstructed dialogue sequence; Based on the reconstructed dialogue sequence, a sequence reconstruction similarity is calculated by combining it with a preset reference dialogue sequence, and the sequence reconstruction similarity is compared with a preset verification threshold to determine the unconscious state verification result. The migration effect is evaluated based on the channel status parameters and cyclic interference parameters recorded in the migration data packet, and the effectiveness evaluation result is obtained. Based on the effectiveness evaluation results and the final state information in the reconstructed dialogue sequence, a final state generation process is performed to obtain a final dialogue state adapted to the target device environment. Based on the seamless state verification results, the validity evaluation results, and the final dialogue state, the final seamless session transition state is determined.
[0083] In one implementation, for the reconstruction of the dialogue state sequence, this embodiment decodes and deserializes the received complete migration data packet on the target device. Specifically, the protocol header of the data packet is first parsed to extract the core data payload of the completed dialogue context and the embedded optimized migration path metadata. Based on the step sequence indicated by the path metadata, each migration step is instantiated sequentially. The instantiation process of each step includes loading the target device's local resources (such as application interface templates and service interfaces) corresponding to that step, and filling the corresponding resources with semantic information related to that step from the completed dialogue context (such as user commands and entity parameters), thereby gradually reconstructing the complete dialogue state sequence from beginning to end.
[0084] Furthermore, this reconstruction process is incremental. The system maintains a dialogue state stack. After each step of instantiation is completed, the specific dialogue state generated (e.g., "Main interface displayed", "Destination filled as 'Beijing'", "Navigation started") is pushed onto the stack, ultimately forming a chronologically ordered, executable, or presentable sequence of reconstructed dialogues.
[0085] It should be noted that the reconstruction process strictly follows the acyclic path structure embedded in the migration data packet and optimized in step S6, ensuring that the logical order of reconstruction is completely consistent with the design intent and avoiding reconstruction errors caused by path ambiguity.
[0086] In one implementation, when the sequence reconstruction similarity is greater than the preset verification threshold, it is determined that the reconstructed dialogue sequence is highly consistent with the reference dialogue sequence, and the non-perceptual state verification result is determined to be passed. When the sequence reconstruction similarity is less than the preset verification threshold, it is determined that there is a significant difference between the two, and the non-perceptual state verification result is determined to be failed, requiring further migration correction or alarm processing; the reference dialogue sequence consists of one hundred standard interaction flows pre-collected by the target device under normal use, and each sequence contains typical dialogue state nodes from start to end.
[0087] It should be noted that the preset verification threshold is determined through offline statistical methods. Specifically, a large number of labeled positive sample dialogue sequence pairs that have successfully migrated and are imperceptible to the user are collected, and their sequence reconstruction similarity is calculated to obtain the positive sample distribution. At the same time, negative sample sequence pairs with perceptible or obvious migration anomalies are collected, and their similarity is calculated to obtain the negative sample distribution. By comparing the similarity distributions of positive and negative samples, a critical value that can effectively distinguish between imperceptible and perceptible migration is selected and determined as the preset verification threshold.
[0088] In one implementation, the migration effect evaluation in this embodiment is based on the auxiliary parameters carried in the migration data packets for quantitative evaluation. Specifically, key channel status parameters (such as average transmission delay and final semantic consistency index) and cyclic interference parameters (such as the maximum interference value before elimination) recorded in previous steps are extracted from the data packets. These parameters reflect the "transmission quality" and "logical health" of the migration process, respectively.
[0089] Furthermore, these parameters are input into a pre-trained evaluation model. This model is a simple regression or classification model, whose input is the aforementioned multi-dimensional parameters, and whose output is a comprehensive transfer performance evaluation score, typically ranging from 0 to 1 or 0 to 100. The learning objective of this model is to make its output score highly correlated with subjective human ratings of transfer performance. Specifically, the evaluation model is a random forest regression model, trained using 3,000 labeled transfer quality scores, where features include channel state parameters and cyclic interference parameters, with the objective of minimizing the mean squared error between the predicted score and the human-labeled score. The number of trees is set to 100, and the maximum depth is set to 10.
[0090] It should be noted that this assessment provides an internal quantitative evaluation of the implementation of this migration technology, which complements the front-end user-perceived verification (non-perceived verification).
[0091] In one implementation, regarding the generation of the final dialogue state, this embodiment adaptively optimizes the dialogue state on the target device side based on the validity evaluation result and the final state information in the reconstructed dialogue sequence. Specifically, the evaluation result is used as an adjustment factor to guide the fine-tuning of the semantic expression of the final state of the reconstructed dialogue sequence, making it more consistent with the operating environment and interaction characteristics of the target device, thereby obtaining a final dialogue state adapted to the target device environment.
[0092] In one implementation, this embodiment makes a final decision on determining the final seamless conversation transition state by integrating all verification and evaluation information. Specifically, the system integrates three decisive inputs: the seamless state verification result ("pass" / "fail"), the transition effect evaluation score E, and the final dialogue state vector.
[0093] Furthermore, the system's decision-making logic is as follows: If the verification result is "passed" and the evaluation score E is higher than the preset passing threshold, such as 0.8, the system determines that the migration is completely successful. At this point, the system officially determines the final dialogue state as the "final seamless session migration state" and immediately sets it as the currently active dialogue state on the target device. All subsequent user interactions will continue based on this state. Simultaneously, the system clears all migration-related temporary caches and intermediate states, completing the entire migration process.
[0094] If the verification result is "failed" or the evaluation score E is too low, the system determines that the migration has not met expectations. In this case, the system will not set the final dialogue state to active, but may initiate a rollback mechanism to try to restore the previous state, or provide the user with a gentle prompt, but it will never cause the application to crash or the dialogue to freeze, thus ensuring robustness at the system level.
[0095] In summary, this invention discloses a seamless omnichannel conversation switching method based on reinforcement learning. It generates an initial dialogue context by constructing a multimodal information acquisition and fusion mechanism, introduces cross-device semantic alignment and expression difference analysis for deviation quantification evaluation, and combines multi-level processing methods such as optimized migration path search and semantic integrity analysis, adaptive filtering correction and dialogue state continuity reconstruction, missing information completion, and cyclic interference compensation. Finally, on the target device, through migration effectiveness evaluation and weighted fusion judgment, it achieves seamless cross-device conversation switching, significantly improving semantic consistency, state continuity, and the seamlessness of user experience during conversation migration.
[0096] Reference Figure 2 The second embodiment of the present invention provides a seamless omnichannel conversation switching system based on reinforcement learning, comprising: The data acquisition module is used to acquire multimodal information on the current device and perform fusion processing on the multimodal information to obtain the initial dialogue context; The deviation degree determination module is used to perform semantic alignment based on the initial dialogue context to obtain the adaptation vector of the target device, and perform expression difference analysis based on the adaptation vector to determine the degree of deviation of the dialogue semantics. An optimized migration path generation module is used to compare the degree of deviation with a preset deviation threshold to obtain a comparison result, and to determine an optimized migration path for session switching using a deep Q network based on the comparison result. The consistency assessment result determination module is used to extract semantic representation vectors from the multimodal information according to the optimized migration path, calculate semantic consistency index between devices according to the semantic representation vectors, and perform integrity analysis according to the semantic consistency index to obtain the consistency assessment result. The corrected dialogue state generation module is used to perform adaptive filtering based on the consistency evaluation result to obtain a corrected semantic representation vector, and to perform dialogue state recovery and continuity reconstruction based on the corrected semantic representation vector to obtain the corrected dialogue state. The migration data packet generation module is used to perform information completion operation on the initial dialogue context according to the corrected dialogue state to obtain a completed dialogue context, perform cyclic interference compensation based on the completed dialogue context to obtain a compensation result, and generate a complete migration data packet according to the compensation result and the completed dialogue context. The final migration status determination module is used to perform a migration validity assessment on the target device based on the migration data packet, obtain the validity assessment result, perform status generation processing, and determine the seamless session migration status.
[0097] It should be noted that the reinforcement learning-based omnichannel seamless session switching system provided in this embodiment of the invention is used to execute all the process steps of the reinforcement learning-based omnichannel seamless session switching method in the above embodiment. The working principles and beneficial effects of the two are one-to-one, so they will not be described again.
[0098] It should be noted that the system embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. Furthermore, in the accompanying drawings of the system embodiments provided by this invention, the connection relationships between modules indicate that they have communication connections, which can be specifically implemented as one or more communication buses or signal lines. Those skilled in the art can understand and implement this without any creative effort.
[0099] The specific embodiments described above further illustrate the purpose, technical solution, and beneficial effects of the present invention. It should be understood that the above descriptions are merely specific embodiments of the present invention and are not intended to limit the scope of protection of the present invention. In particular, it should be noted that any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the scope of protection of the present invention for those skilled in the art.
Claims
1. A method for seamless omnichannel conversation switching based on reinforcement learning, characterized in that, include: Obtain multimodal information from the current device and perform fusion processing on the multimodal information to obtain the initial dialogue context; Based on the initial dialogue context, semantic alignment is performed to obtain the adaptation vector for the target device. Based on the adaptation vector, expression difference analysis is performed to determine the degree of deviation in dialogue semantics. The degree of deviation is compared with a preset deviation threshold to obtain a comparison result. Based on the comparison result, a deep Q-network is used to determine an optimized migration path for session handover. Based on the optimized migration path, the multimodal information is vector-extracted to obtain a semantic representation vector. Based on the semantic representation vector, the semantic consistency index between devices is calculated, and the integrity analysis is performed based on the semantic consistency index to obtain the consistency evaluation result. Based on the consistency evaluation results, adaptive filtering is performed to obtain a corrected semantic representation vector. Based on the corrected semantic representation vector, dialogue state recovery and continuity reconstruction are performed to obtain the corrected dialogue state. Based on the corrected dialogue state, information completion is performed on the initial dialogue context to obtain a completed dialogue context. Based on the completed dialogue context, cyclic interference compensation is performed to obtain a compensation result. Based on the compensation result and the completed dialogue context, a complete migration data packet is generated. Based on the migration data packet, a migration effectiveness assessment is performed on the target device to obtain the effectiveness assessment result, and a state generation process is performed to determine the seamless session migration state.
2. The method for seamless omnichannel conversation switching based on reinforcement learning according to claim 1, characterized in that, The step of acquiring multimodal information on the current device and fusing the multimodal information to obtain the initial dialogue context includes: The text, voice, and visual data collected from the current device interface are acquired as the multimodal information, and the multimodal information is encoded to obtain the original multimodal feature vector. Based on the original multimodal feature vectors, the semantic associations between different modalities are modeled using a preset multimodal fusion model to obtain a multimodal joint representation vector; Based on the multimodal joint representation vector, the intermodal similarity between different modalities is calculated, and the modal features with intermodal similarity below a preset similarity threshold are subjected to weight adjustment processing to obtain a weighted fusion vector; Based on the weighted fusion vector, a difference analysis is performed on the semantic features expressed by different devices to obtain semantic difference measurement results; Based on the semantic difference measurement results, the multimodal joint representation vector is corrected to obtain a multimodal corrected vector, and the multimodal corrected vector is then subjected to context fusion processing to generate the initial dialogue context.
3. The seamless omnichannel conversation switching method based on reinforcement learning according to claim 2, characterized in that, The step of obtaining an adaptation vector for the target device by performing semantic alignment based on the initial dialogue context, and performing expression difference analysis based on the adaptation vector to determine the degree of deviation in dialogue semantics includes: Obtain the preset semantic template of the target device, perform semantic alignment based on the initial dialogue context, and obtain the adapted semantic vector of the target device; The semantic alignment similarity is calculated based on the adapted semantic vector and the preset semantic template, and the semantic alignment similarity is compared with the preset semantic alignment threshold. When the semantic alignment similarity is less than the semantic alignment threshold, the adapted semantic vector is dimensionality reduced and mapped to obtain the aligned semantic vector. Based on the alignment semantic vector, a difference analysis is performed on the semantic distribution difference between the current device and the target device to obtain a semantic difference metric. The semantic difference metric is compared with a preset difference threshold. When the semantic difference metric is greater than the difference threshold, the aligned semantic vector is corrected to obtain an updated semantic vector. The semantic drift degree is obtained by performing temporal analysis based on the updated semantic vector, and then weighted and fused with the semantic alignment similarity and the semantic difference metric to determine the degree of deviation in the dialogue semantics.
4. The method for seamless omnichannel conversation switching based on reinforcement learning according to claim 3, characterized in that, The step of comparing the degree of deviation with a preset deviation threshold to obtain a comparison result, and then using a deep Q-network to determine an optimized migration path for session handover based on the comparison result, includes: Obtain the source dialogue semantic vector from the initial dialogue context; The degree of deviation is compared with a preset deviation threshold. When the degree of deviation is greater than the deviation threshold, the deviation degree between the source dialogue semantic vector and the updated semantic vector is calculated to obtain a semantic distance value. Based on the semantic distance value, the dynamic changes during the session migration process are monitored in real time to obtain the change trend parameters; Based on the change trend parameter and the semantic distance value, multiple candidate migration paths are generated; A deep Q-network is used to evaluate the multiple candidate migration paths to obtain a value estimate for each candidate migration path, and the candidate migration path with the highest value estimate is selected as the optimized migration path for session switching.
5. The method for seamless omnichannel conversation switching based on reinforcement learning according to claim 4, characterized in that, The process involves extracting semantic representation vectors from the multimodal information based on the optimized migration path, calculating semantic consistency indices between devices based on these semantic representation vectors, and performing integrity analysis based on these indices to obtain consistency evaluation results. This includes: Based on the optimized migration path, multimodal key semantic elements are extracted from the multimodal information, and the key semantic elements are encoded to obtain a semantic representation vector. Based on the semantic representation vector, calculate the semantic consistency index between the current device and the target device; When the semantic consistency index is lower than the preset integrity threshold, the semantic information of the current device is determined to be incomplete, and the semantic deviation is calculated based on the semantic representation vector to obtain the information loss value. Based on the information loss value, the data packets in the transmission process are parsed to identify abnormal data packets and determine the range of recoverable data. Based on the recoverable data range, a recovery data packet sequence is constructed, and the consistency of multimodal information is weighted and calculated in conjunction with the semantic consistency index to obtain the consistency evaluation result.
6. The method for seamless omnichannel conversation switching based on reinforcement learning according to claim 5, characterized in that, The process of obtaining a corrected semantic representation vector by performing adaptive filtering based on the consistency evaluation result, and then performing dialogue state recovery and continuity reconstruction based on the corrected semantic representation vector to obtain a corrected dialogue state includes: Based on the consistency evaluation results, a semantic deviation state vector is constructed, and the semantic deviation state vector is subjected to adaptive filtering to obtain a corrected semantic representation vector. The contextual relationships are extracted based on the modified semantic representation vector, and the recovered data packet sequence is parsed. Based on the contextual relationships and the parsing results, the dialogue state is reconstructed to obtain the preliminary dialogue state. Based on the initial dialogue state, a dialogue state graph model is constructed, and node feature propagation optimization is performed on the dialogue state graph model to output an enhanced dialogue state. Based on the enhanced dialogue state, the dialogue sequence between the current device and the target device is aligned and analyzed to obtain a dialogue continuity index. When the dialogue continuity index is less than a preset continuity threshold, the missing dialogue information is predicted and completed to obtain a complete dialogue state. Based on the complete dialogue state, the modified semantic representation vector, the enhanced dialogue state, and the dialogue continuity index are fused and calculated to determine the modified dialogue state.
7. The method for seamless omnichannel conversation switching based on reinforcement learning according to claim 6, characterized in that, The process includes: performing information completion on the initial dialogue context based on the corrected dialogue state to obtain a completed dialogue context; performing cyclic interference compensation based on the completed dialogue context to obtain a compensation result; and generating a complete migration data packet based on the compensation result and the completed dialogue context, including: Based on the corrected dialogue state, semantic alignment and completion operations are performed on the missing information in the initial dialogue context to obtain the completed dialogue context. Based on the completed dialogue context, loop detection is performed on the optimized migration path to obtain cyclic interference parameters; When the cyclic interference parameter is greater than the preset interference threshold, the optimized migration path is de-looped to obtain a loop-free migration path as the compensation result. Data integration processing is performed based on the completed dialogue context and the loop-free migration path to generate a complete migration data package.
8. The method for seamless omnichannel conversation switching based on reinforcement learning according to claim 1, characterized in that, The step of performing a migration effectiveness assessment on the target device based on the migration data packet, obtaining the effectiveness assessment result, and performing state generation processing to determine the seamless session migration state includes: Based on the migration data packet, the dialogue state sequence is reconstructed on the target device to obtain the reconstructed dialogue sequence; Based on the reconstructed dialogue sequence, a sequence reconstruction similarity is calculated by combining it with a preset reference dialogue sequence, and the sequence reconstruction similarity is compared with a preset verification threshold to determine the unconscious state verification result. The migration effect is evaluated based on the channel status parameters and cyclic interference parameters recorded in the migration data packet, and the effectiveness evaluation result is obtained. Based on the effectiveness evaluation results and the final state information in the reconstructed dialogue sequence, a final state generation process is performed to obtain a final dialogue state adapted to the target device environment. Based on the seamless state verification results, the validity evaluation results, and the final dialogue state, the final seamless session transition state is determined.
9. A seamless omnichannel conversation switching system based on reinforcement learning, characterized in that, include: The data acquisition module is used to acquire multimodal information on the current device and perform fusion processing on the multimodal information to obtain the initial dialogue context; The deviation degree determination module is used to perform semantic alignment based on the initial dialogue context to obtain the adaptation vector of the target device, and perform expression difference analysis based on the adaptation vector to determine the degree of deviation of the dialogue semantics. An optimized migration path generation module is used to compare the degree of deviation with a preset deviation threshold to obtain a comparison result, and to determine an optimized migration path for session switching using a deep Q network based on the comparison result. The consistency assessment result determination module is used to extract semantic representation vectors from the multimodal information according to the optimized migration path, calculate semantic consistency index between devices according to the semantic representation vectors, and perform integrity analysis according to the semantic consistency index to obtain the consistency assessment result. The corrected dialogue state generation module is used to perform adaptive filtering based on the consistency evaluation result to obtain a corrected semantic representation vector, and to perform dialogue state recovery and continuity reconstruction based on the corrected semantic representation vector to obtain the corrected dialogue state. The migration data packet generation module is used to perform information completion operation on the initial dialogue context according to the corrected dialogue state to obtain a completed dialogue context, perform cyclic interference compensation based on the completed dialogue context to obtain a compensation result, and generate a complete migration data packet according to the compensation result and the completed dialogue context. The final migration status determination module is used to perform a migration validity assessment on the target device based on the migration data packet, obtain the validity assessment result, perform status generation processing, and determine the seamless session migration status.