A method and system for dual media session continuity orchestration
By deploying a converged acquisition unit, a collaborative orchestration decision engine, a linked switching control unit, and a session anchor synchronization unit in a dual-media session, the problem of real-time media interruption and non-real-time media resource waste during cross-device or cross-network switching in existing technologies is solved. This achieves efficient seamless continuation and systematic consistency verification, improving user experience and resource utilization efficiency.
Patent Information
- Application Number
- CN202610660497.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2026-05-14
- Publication Date
- 2026-08-25
- Estimated Expiration
- 2046-05-14
AI Technical Summary
Existing technologies cannot integrate and uniformly perceive the network transmission quality and status of real-time and non-real-time media when switching or migrating between devices or networks in dual-media sessions. This leads to interruption of real-time media or waste of non-real-time media resources. The lack of systematic consistency verification and adaptive adjustment makes it difficult to achieve seamless connection and affects user experience.
By deploying dual-media convergence acquisition units to collect network transmission quality and status data of real-time and non-real-time media, a collaborative orchestration decision engine is built to dynamically allocate priority weights, a linkage switching control unit is configured to realize media status binding and synchronization, a session anchor point synchronization unit is built to perform status synchronization, and a status collaborative verification unit is deployed to perform system consistency verification. Combined with lightweight context management and adaptive adjustment, a closed-loop guarantee system is formed.
It achieves priority protection for real-time media in highly sensitive scenarios and reasonable yielding of non-real-time media when bandwidth is limited, improving session continuity and user experience, significantly improving handover success rate and continuity reliability, and reducing resource utilization inefficiency and central cloud load.
Smart Images

Figure CN122204832B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of media session management technology, specifically to a method and system for continuous arrangement of dual-media sessions. Background Technology
[0002] With the widespread adoption of 5G, cloud communication, and multi-device smart terminals, dual-media sessions have become the mainstream form for scenarios such as remote work, online education, video collaboration, and enterprise instant messaging. A dual-media session refers to a hybrid communication mode where real-time media (such as voice and video calls) and non-real-time media (such as file transfer, document sharing, and data reporting) coexist within the same session. Real-time media is highly sensitive to network latency and jitter, requiring instant transmission and playback; interruptions or excessive latency severely impact user experience. Non-real-time media, on the other hand, allows for caching and resumeable interruptions, has lower real-time requirements, and can be flexibly paused or resumed based on network conditions.
[0003] Existing technologies mainly rely on media switching mechanisms under the IMS architecture, the WebRTC real-time communication framework, and some protocol extension schemes to achieve session continuity. For example, media migration across access networks or across devices can be supported through SIP renegotiation, SRVCC voice call continuity, or media stream context carrying.
[0004] However, existing technologies still have the following major drawbacks when handling cross-device or cross-network switching and migration of dual-media sessions: 1. The inability to integrate and uniformly perceive the network transmission quality of real-time media and the transmission status of non-real-time media results in the decision engine lacking comprehensive context and being unable to dynamically and reasonably allocate the priority weights of the two media according to the session scenario and the current task intent, causing problems such as interruption or stuttering of real-time media and waste or priority mismatch of non-real-time media resources during the switching process. 2. The context migration and state synchronization mechanisms are relatively crude, limited to parameter renegotiation or simple information carrying at the protocol level. They lack binding caching of session context summaries and the current transmission state of non-real-time media, pause control, locking of real-time media encoding parameters and playback progress, as well as progress monitoring and threshold triggering synchronization based on session anchors. This leads to disconnection of the target end state, deviation of transmission progress or context breakage, making it difficult to achieve truly seamless continuation. 3. After switching and migration, there is a lack of systematic consistency verification from three dimensions: media transmission, state synchronization, and context coherence, as well as corresponding hierarchical compensation mechanisms. Once a deviation occurs, it often has to rely on retry or manual intervention. It cannot automatically identify and correct transmission deviations, state disconnections, or context breaks. At the same time, it lacks lightweight edge caching and scenario-intention-based dual-drive adaptive adjustment capabilities, resulting in low resource utilization efficiency and insufficient continuity guarantee in dynamic network environments.
[0005] These shortcomings make it difficult for existing dual-media conversation systems to provide a reliable and consistent user experience in complex cross-device and cross-network scenarios, limiting their further promotion in collaborative applications with high continuity requirements.
[0006] Therefore, a dual-media session continuity orchestration method and system are needed to solve the above problems. Summary of the Invention
[0007] To address the shortcomings of existing technologies, this invention provides a dual-media session continuity orchestration method and system, which solves the problems mentioned in the background section above.
[0008] To achieve the above objectives, the present invention provides a method for arranging the continuity of a dual-media session, comprising the following steps: S1. Deploy a dual-media fusion acquisition unit in a dual-media session scenario to collect network transmission quality data for real-time media, transmission status data for non-real-time media, and session scenario identifiers and session context data. It should be noted that real-time media refers to media types that require real-time transmission and playback, are sensitive to transmission latency and jitter, and cannot be cached for extended periods. Their transmission process must ensure "instant transmission and playback," and any transmission interruption or excessive latency will directly impact the session experience. Non-real-time media refers to media types that do not require real-time transmission, can be cached, and can resume interrupted transmissions. They have lower real-time requirements, and their transmission can be paused and resumed based on network conditions without affecting the core session flow. Session context data refers to relevant data that runs throughout the entire dual-media session, affects session coherence, and can be used to identify session intent. Combined with S1.4, this specifically includes participant identification, the current task node in the session flow, and the most recent key operation record affecting session coherence. The session scenario identifier is a category label generated based on the initial configuration information at the time of session initiation; its core function is to distinguish the session's sensitivity to media real-time performance. S2. Build a dual-media collaborative orchestration decision engine with a built-in scene-strategy mapping model. Receive various data collected in step S1, identify the intent of the session task through a context-aware algorithm, and dynamically allocate the priority weights of the two media in combination with the session scene identifier to generate an integrated strategy for dual-media collaborative linkage switching and context continuation. S3. Configure and start the linkage switching control unit, extract the session context summary, bind the cache with the current transmission status of non-real-time media, control the non-real-time media to pause transmission, and lock the real-time media encoding protocol type, bitrate parameters and playback progress. S4. Establish a session anchor synchronization unit to monitor the switching and migration progress. Synchronize the context summary and non-real-time media status to the target device or network node through the session anchor. After the switching and migration are completed, trigger the non-real-time media to resume playback and synchronize the real-time media playback progress and context summary. The session anchor is a specific data structure used to realize the synchronization of session status. Combined with S4.1, it includes a unique session identifier, a switching timestamp, a context summary index, and a non-real-time media status pointer. S5. Deploy a dual-media state collaborative verification unit. After switching and migration, verify the state consistency from three dimensions: media transmission, state synchronization, and context coherence. Determine if there is any transmission deviation, state disconnection, or context breakage. If the state is consistent, maintain the orchestration state. If there is a deviation, trigger a hierarchical compensation mechanism to correct the deviation and continue the context.
[0009] Preferably, the method for deploying the dual-media fusion acquisition unit and acquiring related data in step S1 is as follows: S1.1 The dual-media fusion acquisition unit collects real-time media network transmission quality data through a preset monitoring interface. The network transmission quality data includes latency, jitter, and bandwidth utilization. While acquiring the network transmission quality data, the dual-media fusion acquisition unit parses the real-time media data stream to extract the encoding protocol type and bitrate parameters. S1.2 The dual-media fusion acquisition unit establishes a data channel with the storage service of non-real-time media to obtain the transmission status data of non-real-time media in real time. The transmission status data of non-real-time media includes the amount of data transmitted, the amount of data remaining, the data retention time in the current cache, the preset file transmission priority, and the file format type. S1.3 The dual-media fusion acquisition unit extracts the scene identification code carried by parsing a specific field in the session initialization signaling, and uses it as the current session scene identifier. The session scene identifier is generated based on the initial configuration information when the session is initiated, and is used to distinguish the level of sensitivity of the session to media real-time performance. S1.4 The dual-media fusion acquisition unit synchronously collects session context data by listening to session control signaling and recording interaction logs. The session context data includes the identity identifiers of the participants, the current task nodes in the session flow, and the record of the most recent key operation that affects the continuity of the session. S1.5 The dual-media fusion acquisition unit has a built-in data preprocessing module to perform noise reduction and standardization processing on the acquired data.
[0010] Preferably, the specific steps for building the dual-media collaborative orchestration decision engine and related operations in step S2 are as follows: S2.1 Process the conversation context data collected in step S1.4 using a context-aware algorithm to extract keywords, task node markers, and media type features from the context data. Calculate the similarity between the extracted features and a preset intent feature library. When the similarity is higher than a preset threshold, the matched intent category is determined as the conversation task intent. Here, the preset threshold is an intent recognition similarity threshold used to determine the degree of matching between the conversation context features and the intent categories in the preset intent feature library. The preset threshold is 85%, which can be adjusted according to the actual application scenario. When the calculated similarity is higher than 85%, the corresponding conversation task intent is determined to be matched.
[0011] S2.2. Determine the baseline weights of real-time media and non-real-time media in the scenario based on the session scenario identifiers collected in S1.3. Then, dynamically adjust the baseline weights based on the session task intent identified in S2.1 to obtain the final priority weights of real-time media and non-real-time media. S2.3. Based on the final priority weight allocated in S2.2, generate an integrated strategy for dual-media collaborative switching and context continuation. The integrated strategy includes triggering conditions for cross-device or cross-network switching and migration, caching rules for context summary and non-real-time media status, and dual-media status synchronization methods.
[0012] Preferably, the linkage switching control unit comprises a switching trigger monitoring module, a context extraction module, a cache control module, a media state locking module, and a temporary cache unit. The switching trigger monitoring module monitors the real-time media transmission environment in real time and identifies trigger signals for cross-device switching and cross-network migration. The linkage switching control unit and related operation methods include the following steps: S3.1 When the switching trigger monitoring module detects a real-time media trigger for a cross-device or cross-network switching or migration, it sends a trigger signal to the context extraction module. The context extraction module extracts a session context summary directly related to the current task node from the session context data collected in step S1. The context summary is the core concise content of the context data. S3.2 The cache control module binds the extracted session context summary with the current transmission status of the non-real-time media and caches it in the temporary cache unit. At the same time, it sends a pause command to the non-real-time media transmission node to control the non-real-time media to pause transmission. The current transmission status of the non-real-time media includes transmission progress, cache position, and incomplete segment identifier. S3.3 The media status locking module locks the current encoding protocol type and bitrate parameters of the real-time media, generates a status lock identifier, and ensures that the real-time media can be seamlessly resumed after switching or migration.
[0013] Preferably, the session anchor synchronization unit consists of an anchor generation module, a handover progress monitoring module, a data synchronization module, and a resume triggering module. The operation steps of the session anchor synchronization unit are as follows: S4.1. Session anchors are generated through the anchor generation module, and the real-time media switching and migration progress is monitored through the switching progress monitoring module. The session anchor is a data structure containing a unique session identifier, a switching timestamp, a context digest index, and a non-real-time media status pointer. The unique session identifier is used to distinguish different sessions, the switching timestamp is used to mark the switching and migration trigger time, the context digest index is used to quickly locate the cached context digest, and the non-real-time media status pointer is used to locate the current transmission status of the non-real-time media. S4.2 The switching progress monitoring module collects real-time media switching and migration progress data in real time. When the switching and migration progress reaches the threshold, a synchronization command is sent to the data synchronization module. The progress data includes the switching completion rate and the connection status of the target device or network node. The threshold here is the switching and migration progress threshold, which is used to judge the degree of completion of real-time media switching and migration. The preset threshold is 90%, which can be adjusted according to the actual application scenario. When the switching completion rate reaches 90% and the connection of the target device or network node is stable, the session status synchronization operation is triggered. S4.3 The data synchronization module synchronizes the context summary and non-real-time media state in the temporary cache unit to the corresponding cache unit of the target device or network node through the session anchor point, ensuring that the target node obtains complete session state data. S4.4 After the switching and migration are completed, the resume trigger module sends a resume instruction to the non-real-time media transmission node, triggering the non-real-time media to resume transmission based on the synchronous transmission status. At the same time, the encoding protocol type, bitrate parameters and playback progress of the real-time media are synchronized to ensure that the dual media status is consistent.
[0014] Preferably, the method further includes a lightweight context-associated state management step, as follows: S5.1. Filter core status data based on session task intent and discard redundant data. The core status data includes the encoding and decoding parameters of real-time media, the unique session identifier, the current playback progress and task association marker, the transmission progress of non-real-time media, the incomplete task identifier, the priority and data anonymization marker, and the session context summary. Redundant data includes historical chat records, non-core files that have been transmitted, configuration parameters that are not associated with tasks, and expired status data. S5.2 Cache the core state data and context summary to the edge computing server, while the central cloud server only caches the core data index. The two interact through a dedicated network channel. S5.3 During session switching and migration, only the changed core state data and context update content are transmitted, and the core state and context are kept consistent through session anchors. Preferably, the method further includes a scene-intention dual-driven adaptive adjustment step, as follows: S6.1 Build a scenario-intention dual-driven orchestration template library, and pre-store orchestration templates that adapt to different scenarios and different task intentions; S6.2 Deploy template adaptive learning units to collect user operation behavior, network bandwidth, latency, packet loss rate and changes in session task intent, automatically select suitable orchestration templates, and adjust media priority weights, switching trigger thresholds and caching strategy parameters through reinforcement learning algorithms; set up hierarchical manual intervention interfaces to support administrators to manually adjust orchestration strategies and parameters, and retain operation records for traceability.
[0015] Preferably, the system includes a dual-media fusion acquisition unit, a dual-media collaborative orchestration decision engine, a linkage switching control unit, a session anchor point synchronization unit, and a dual-media status collaborative verification unit. Each unit realizes data interaction and command transmission through a dedicated data channel, and collaboratively completes the continuous orchestration of dual-media sessions. The dual-media fusion acquisition unit is used to deploy and execute the data acquisition and preprocessing operations of step S1 in a dual-media session scenario; The dual-media collaborative orchestration decision engine has a built-in scene-policy mapping model, which is used to receive data collected by the dual-media fusion acquisition unit and perform the S2 step of session task intent recognition, priority weight allocation and integrated policy generation operation. The linkage switching control unit consists of a switching trigger monitoring module, a context extraction module, a cache control module, a media state locking module, and a temporary cache unit, and is used to execute the linkage switching control related operations in step S3; The session anchor point synchronization unit consists of an anchor point generation module, a switching progress monitoring module, a data synchronization module, and a resume triggering module, and is used to perform the session anchor point generation, state synchronization, and resume triggering related operations in step S4. The dual-media state collaborative verification unit is used to perform dual-media state consistency verification and deviation compensation triggering operations after the switching and migration of step S5 is completed.
[0016] Preferably, the system further includes a lightweight context-related state management unit and a scene-intent dual-driven adaptive adjustment unit, both of which interact with the dual-media collaborative orchestration decision engine. The lightweight context-associated state management unit is used to execute the lightweight context-associated state management steps S5.1 to S5.3; The scenario-intention dual-drive adaptive adjustment unit is used to execute the scenario-intention dual-drive adaptive adjustment steps S6.1 to S6.2.
[0017] This invention provides a method and system for arranging continuous dual-media sessions. It has the following beneficial effects: 1. This invention uses a dual-media fusion acquisition unit to simultaneously collect data on the transmission quality of real-time media networks (latency, jitter, bandwidth utilization) and the transmission status of non-real-time media (transmitted / remaining data volume, buffer latency, etc.). Combined with session scenario identifiers and contextual data, it achieves unified perception and preprocessing of all elements of a dual-media session, solving the problems of isolated information between the two media states and the lack of fusion perception leading to imbalanced priority allocation in existing technologies. Compared to solutions relying solely on protocol layer parameter renegotiation, this invention provides a comprehensive and real-time contextual basis for the subsequent decision engine, enabling more accurate dynamic priority weight allocation. Real-time media is prioritized in highly sensitive scenarios, while non-real-time media is given reasonable priority when bandwidth is limited, thereby significantly improving session continuity and user experience during handover and migration processes.
[0018] 2. This invention utilizes a dual-media collaborative orchestration decision engine with a built-in scene-policy mapping model. Combined with a context-aware algorithm, it identifies the intent of the session task and dynamically adjusts the priority weights of real-time and non-real-time media based on scene identifiers. This generates an integrated linkage switching and context continuation strategy, achieving intent-driven intelligent orchestration decision-making. This overcomes the problems of existing technologies with single decision rules and the inability to flexibly adjust media priorities according to the current task intent. Simultaneously, the linkage switching control unit completes the binding and caching of context summaries and non-real-time media states, pauses non-real-time media transmission, locks the real-time media encoding protocol / bitrate / playback progress, and monitors and accurately synchronizes the progress threshold of the session anchor synchronization unit. This ensures seamless alignment of the target end state, overcoming the state disconnection, transmission deviation, and context break defects easily caused by existing coarse migration mechanisms, achieving truly seamless dual-media continuation.
[0019] 3. This invention deploys a dual-media state collaborative verification unit. After switching and migration, it performs systematic consistency verification from three dimensions: media transmission, state synchronization, and contextual coherence. It also triggers a tiered compensation mechanism to automatically correct deviations. Combined with lightweight context-related state management (edge caching of core data, central cloud storing only indexes) and scenario-intent dual-driven adaptive adjustment (reinforcement learning dynamically optimizes weights, thresholds, and caching strategies), a complete closed-loop guarantee system is formed. Compared to existing technologies that lack verification compensation, rely on manual intervention or restarts, and have inefficient resource utilization, this invention significantly improves switching success rate and continuity reliability, reduces central cloud load and transmission overhead, and enhances scalability and stability in dynamic networks and high-concurrency scenarios. Simultaneously, through adaptive learning, it continuously optimizes orchestration strategies to adapt to different user behaviors and network fluctuations, further shortening interruption time, improving overall session experience, and system robustness. Attached Figure Description
[0020] Figure 1 This is the main flowchart of the present invention; Figure 2 This is a flowchart of step S1 of the present invention; Figure 3 This is a flowchart of step S2 of the present invention; Figure 4 This is a flowchart of step S3 of the present invention; Figure 5 This is a flowchart of step S4 of the present invention; Figure 6 This is a system framework diagram of the present invention. Detailed Implementation
[0021] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention. Specific Implementation Example 1:
[0023] like Figures 1 to 6As shown, the entire dual-media conversation continuity orchestration system consists of a dual-media fusion acquisition unit, a dual-media collaborative orchestration decision engine, a linkage switching control unit, a conversation anchor synchronization unit, a dual-media state collaborative verification unit, a lightweight context-related state management unit, and a scene-intent dual-driven adaptive adjustment unit. Each unit interacts and transmits commands through a dedicated data channel to collaboratively complete the dual-media conversation continuity orchestration. Specifically, the lightweight context-related state management unit and the scene-intent dual-driven adaptive adjustment unit interact with the dual-media collaborative orchestration decision engine to execute the dual-media conversation continuity orchestration method. The specific working principle is as follows: Step 1: Data Acquisition. Step S1 is executed by the dual-media fusion acquisition unit. Deployed in a dual-media session scenario, the dual-media fusion acquisition unit's core function is to comprehensively collect various key data during the session and preprocess them to provide accurate data support for subsequent decision-making. The specific operations are as follows: S1.1: Collect network transmission quality data of real-time media through a preset monitoring interface, including latency, jitter, and bandwidth utilization. Simultaneously, parse the real-time media data stream to extract the encoding protocol type and bitrate parameters. S1.2: Establish a data channel with the storage service of non-real-time media to obtain real-time transmission status data of non-real-time media, including the amount of data transmitted, the amount of data remaining, the current data retention time in the cache, the preset file transmission priority, and the file format type. S1.3: Extract the scene identifier code carried in the specific field of the session initialization signaling by parsing it, as the current session identifier. The conversation scene identifier, generated from the initial configuration information at the start of the conversation, serves to differentiate the conversation's sensitivity to media real-time performance. S1.4: By monitoring conversation control signaling and recording interaction logs, conversation context data is synchronously collected, including participant identification, current task nodes in the conversation flow, and the most recent operation record affecting conversation continuity, ensuring the context data fully reflects the current state of the conversation. S1.5: Through the data preprocessing module built into the dual-media fusion acquisition unit, all collected data undergoes denoising and standardization. Preprocessing first denoises various types of collected data, removing abnormal data, and then standardizes the denoised data to a format compatible with the subsequent decision engine. After preprocessing, the data is transmitted to the dual-media collaborative orchestration decision engine in real time.
[0024] The second step, S2, is executed by the dual-media collaborative orchestration decision engine. This engine has a built-in scene-policy mapping model. Its core function is to receive the preprocessed data collected in step S1, identify the session task intent through a context-aware algorithm, dynamically allocate priority weights for the two media, and generate an integrated strategy for dual-media collaborative switching and context-based continuation. The specific operation is as follows: S2.1, The conversation context data collected in S1.4 is processed using a context-aware algorithm to extract keywords, task node markers, and media type features. These extracted features are then compared with various intent features in a pre-defined intent feature library to calculate similarity. When the calculated similarity exceeds a pre-defined threshold, the matched intent category is determined as the conversation task intent. The complete calculation method is as follows: extract keywords, task node markers, and media type features from the conversation context data collected in S1.4, and construct an image vector. Each dimension of the feature vector corresponds to a key extracted feature, and the value of each dimension represents the weight of that feature; the feature vector... Compared with the standard feature vector of a certain type of conversation intent in the preset intent feature library Substituting the values into the cosine similarity algorithm formula, we calculate the similarity. The specific algorithm is as follows: in This represents the similarity between the session context feature vector and a certain type of intent feature vector in the preset intent feature library. The value ranges from 0 to 1. The closer the value is to 1, the higher the matching degree between the two. This represents a feature vector extracted from the current session context data. Each dimension of the vector corresponds to a key extracted feature, and the value of each dimension represents the weight of that feature. This represents a standard feature vector representing a certain type of conversation intent from a predefined intent feature library, with vector dimensions equal to... Completely consistent, with the value of each dimension representing the standard weight percentage of the corresponding feature in that type of intent; Representative eigenvector With feature vectors The dot product is calculated by multiplying the corresponding dimensions of the two vectors and then summing the results. Representative eigenvector The magnitude is calculated as a vector. The square root of the sum of the squares of all dimension values. Representative eigenvector The modulus, the calculation method is the same as Consistent.
[0025] Calculate the similarity for each type of intent. Then, when the calculated similarity is higher than the preset threshold of 85%, the matched intent category is determined as the session task intent.
[0026] S2.2, The scenario-policy mapping model determines the media priority weights. Based on the baseline weights pre-stored in the conversation scenario identifier matching model collected in S1.3, the baseline weights of real-time media and non-real-time media are obtained. Then, based on the conversation task intent identified in S2.1, an intent adjustment coefficient is introduced to linearly correct the baseline weights, resulting in the final priority weights of real-time media and non-real-time media.
[0027] S2.3 Generate an integrated strategy for dual-media collaborative switching and context continuation. Based on the final priority weights determined in S2.2, an integrated strategy comprising three parts is generated: first, the triggering conditions for cross-device or cross-network switching and migration, specifying the specific circumstances triggering switching and migration, such as real-time media transmission anomalies and user manual operations; second, the caching rules for context summaries and non-real-time media states, specifying the specific content to be cached, the cache duration, and the cache update frequency; and third, the dual-media state synchronization method, specifying the synchronization frequency, the synchronized data content, and the synchronization implementation path. The generated integrated strategy is transmitted in real time to the collaborative switching control unit, the session anchor synchronization unit, and the dual-media state collaborative verification unit as the basis for subsequent operations.
[0028] The third step, handover control, is executed in step S3 and is completed by the linkage handover control unit. This unit consists of a handover trigger monitoring module, a context extraction module, a cache control module, a media state locking module, and a temporary cache unit. Its core function is to complete context extraction, cache binding, and media state locking after detecting a handover migration trigger signal, preparing for subsequent state synchronization. The specific operation is as follows: S3.1, the handover trigger monitoring module monitors the real-time media transmission environment in real time, identifies trigger signals for cross-device handover and cross-network migration, and immediately sends the trigger signal to the context extraction module when a trigger signal is detected; after receiving the trigger signal, the context extraction module extracts the relevant information from the session context data collected in step S1. The current task node directly relates to the session context summary, where the context summary is the core concise content of the context data; S3.2, the cache control module binds the extracted session context summary with the current transmission status of the non-real-time media and caches it in the temporary cache unit, and at the same time sends a pause command to the non-real-time media transmission node to control the non-real-time media to pause transmission, where the current transmission status of the non-real-time media includes transmission progress, cache position, and incomplete segment identifier; S3.3, the media status locking module locks the current encoding protocol type and bitrate parameters of the real-time media, generates a status lock identifier, and ensures that the real-time media can be seamlessly resumed after switching or migration.
[0029] The fourth step is state synchronization, executed in step S4, completed by the session anchor synchronization unit. This unit consists of an anchor generation module, a handover progress monitoring module, a data synchronization module, and a resume trigger module. Its core function is to monitor the handover and migration progress, synchronize session states through session anchors, and trigger non-real-time media resume and real-time media progress synchronization. The specific operations are as follows: S4.1, the anchor generation module generates session anchors, and the handover progress monitoring module monitors the real-time media handover and migration progress. The session anchor is a data structure containing a unique session identifier, a handover timestamp, a context summary index, and a non-real-time media state pointer; S4.2, the handover progress monitoring module collects real-time progress data for real-time media handover and migration. When the handover and migration progress reaches... When the threshold is reached, a synchronization command is sent to the data synchronization module, where the progress data includes the switching completion rate and the connection status of the target device or network node; S4.3, the data synchronization module synchronizes the context digest and non-real-time media status in the temporary cache unit to the corresponding cache unit of the target device or network node through the session anchor point, ensuring that the target node obtains complete session status data; S4.4, when the switching and migration are completely completed, the resume triggering module sends a resume command to the non-real-time media transmission node, triggering the non-real-time media to resume transmission based on the synchronized transmission status, while synchronizing the encoding protocol type, bitrate parameters and playback progress of the real-time media, ensuring that the dual media status is consistent.
[0030] Step 5, Status Verification: Step S5 is executed by the dual-media status collaborative verification unit. Its core function is to verify the consistency of the dual-media status and handle deviations after switching and migration. The specific operation is as follows: After switching and migration, the dual-media status collaborative verification unit verifies status consistency from three dimensions: media transmission, status synchronization, and context coherence, determining whether there are transmission deviations, status disconnects, or context breaks. If all three dimensions pass verification, the status is considered consistent, the current orchestration status is maintained, and the dual-media session continues. If any dimension fails verification, a deviation is determined, triggering a tiered compensation mechanism to correct the deviation and reconnect the context. The tiered compensation mechanism is divided into three levels based on the severity of the deviation: Level 1 deviation automatically supplements missing parameters without interrupting the session; Level 2 deviation adjusts the non-real-time media transmission rate to quickly catch up and synchronize the context; Level 3 deviation restarts the linkage switching control and session anchor synchronization process, resynchronizing all status data to ensure session continuity is restored.
[0031] Step 6: Lightweight Management. This step involves implementing lightweight context-associated state management, completed by the lightweight context-associated state management unit. Its core function is to optimize state data storage and transmission efficiency, reducing resource consumption. Specific operations are as follows: Core state data is filtered based on session task intent, discarding redundant data. Core state data includes encoding / decoding parameters for real-time media, a unique session identifier, current playback progress, and task association markers; transmission progress for non-real-time media, incomplete task identifiers, priority, and data anonymization markers; and a session context summary. Redundant data includes historical chat logs, completed non-core files, non-task-associated configuration parameters, and expired state data. The core state data and context summary are cached on the edge computing server, while the central cloud server only caches the core data index. The two interact through a dedicated network channel. During session switching and migration, only the changed core state data and updated context content are transmitted, ensuring consistency between core state and context through session anchor association.
[0032] Step 7: Adaptive Adjustment. This step, driven by both scenario and intent, is performed by the scenario-intent dual-drive adaptive adjustment unit. Its core function is to optimize orchestration strategies and improve adaptability. Specific operations are as follows: A scenario-intent dual-drive orchestration template library is built, pre-storing orchestration templates adapted to different scenarios and task intents; a template adaptive learning unit is deployed to collect user operation behavior, network bandwidth, latency, packet loss rate, and changes in session task intent. Through reinforcement learning algorithms, it automatically selects the appropriate orchestration template, adjusts media priority weights, switching trigger thresholds, and caching strategy parameters; and a tiered manual intervention interface is set up to support administrators in manually adjusting orchestration strategies and parameters, with operation records retained for traceability.
[0033] The reinforcement learning algorithm achieves dynamic policy optimization. The specific algorithm is as follows:
[0034] in, Representing the current moment, the updated Q-value of the state-action pair corresponding to "Scene-Intent-Orchestration Parameters" in this solution reflects the expected long-term adaptation benefit of performing the corresponding adjustment action in this state after the update, and is directly used to guide the selection of orchestration templates and parameter adjustments; The original Q-value of the state-action pair corresponding to "scene-intent-orchestration parameter" in this scheme at the current moment represents the initial long-term adaptation benefit expectation of performing an orchestration parameter adjustment action (such as adjusting media priority weight or switching trigger threshold) under the current session scene (such as high / medium / low sensitivity scene) and task intent. It is the basis for algorithm iteration and update. This represents the learning rate, ranging from 0 to 1. It is used to adapt to the template adaptive learning requirements of this solution and controls the update weight of the original orchestration strategy parameters based on "newly collected user operation and network status data". The closer it is to 1, the greater the impact of newly collected session-related data (such as users frequently adjusting media priority and network latency fluctuations) on the adjustment of orchestration parameters, and the faster it can adapt to the current session state; This represents the execution of the current adjustment action. The instant reward value obtained from the session environment of this solution directly reflects the adaptation effect of the adjustment action. For example, if adjusting the media priority weight results in smooth real-time media transmission and seamless non-real-time media retransmission, then... The value is relatively high; if transmission deviation or context breakage occurs, then... The value is relatively low; This represents a discount factor, ranging from 0 to 1, used to weigh the importance of current immediate rewards against future long-term rewards. The closer it is to 1, the more importance is attached to the long-term adaptability in future conversations, so as to avoid neglecting the long-term smoothness of conversations due to short-term adaptability effects. Represents the execution of actions The next session state after entering the next moment, in this scheme, specifically refers to the comprehensive state corresponding to the new session scenario, task intent, and network status (such as changes in network bandwidth, and switching of session intent from file sharing to real-time communication) after adjusting the orchestration parameters; Represents the state at the next moment. The maximum Q value corresponding to the execution of all possible orchestration parameter adjustment actions (such as adjusting cache duration, switching trigger threshold, and media priority weight), that is, the long-term expected benefit of achieving the best session adaptation effect in the future, is used to guide the selection of subsequent adjustment actions. This represents any orchestration parameter adjustment action that can be executed at the next moment. In conjunction with this solution, this specifically includes all actions that can optimize the orchestration strategy, such as adjusting media priority weights, switching trigger thresholds, caching strategy parameters (caching duration, update frequency), and selecting suitable scenario-intent orchestration templates.
[0035] The template adaptive learning unit continuously updates the Q value through this algorithm and optimizes the action selection strategy, thereby realizing the automatic adaptation and parameter adjustment of the orchestration template. This ensures that the orchestration strategy always fits the current session scenario, task intent and network status, and guarantees the continuity of the dual-media session.
[0036] In summary, the entire dual-media session continuity orchestration method and system achieves continuity of dual-media sessions during cross-device and cross-network switching migrations through a complete process of "data acquisition - strategy generation - handover control - state synchronization - verification compensation - lightweight management - adaptive adjustment".
[0037] Also included in the following attached figures Figure 1 The main flowchart clearly illustrates the core steps of the method of this invention, with appendices. Figure 2 -Appendix Figure 5 The core steps S1-S4 are shown in detail below. Figure 6 This constitutes the system framework. Specific Implementation Example 2: like Figures 1 to 6 As shown, the units and modules in Embodiment 1 will be further described below: The dual-media fusion acquisition unit's hardware components include an industrial-grade acquisition controller, a multi-interface data acquisition card, a signal conditioning module, a preprocessing chip, and a high-speed cache module. The industrial-grade acquisition controller adopts an embedded ARM architecture, supports multi-protocol data interaction, and is responsible for coordinating the collaborative work of all hardware within the acquisition unit. The multi-interface data acquisition card integrates Ethernet, USB 3.0, and fiber optic interfaces, corresponding to the acquisition of real-time media network data, non-real-time media storage data, and session signaling data, respectively, supporting simultaneous acquisition of multiple data streams without interference. The signal conditioning module filters and amplifies the acquired analog signals to ensure data transmission stability. The preprocessing chip uses a dedicated digital signal processing chip to perform data denoising and standardization, ensuring preprocessing efficiency. The high-speed cache module temporarily stores the acquired raw data and preprocessed data to prevent data loss. The hardware specifications of this unit focus on meeting the synchronous acquisition and rapid preprocessing of multiple data types, adapting to the real-time data acquisition requirements during sessions, and ensuring that all hardware interfaces support hot-swapping for easy maintenance and expansion.
[0039] The dual-media collaborative orchestration decision engine's hardware components include a main control chip, a memory module, a solid-state drive (SSD), a dedicated algorithm acceleration module, and a network communication module. The main control chip, employing a high-performance CPU, is responsible for parsing received data, calling built-in models and algorithms, and generating orchestration strategies; it is the core hardware of the engine. The memory module uses DDR4 specifications to cache intermediate data, model parameters, and algorithm results during operation, ensuring smooth engine operation. The SSD stores preset intent feature libraries, scene-policy mapping model parameters, orchestration strategy templates, and operation logs. The dedicated algorithm acceleration module uses a GPU acceleration chip to accelerate the computation speed of context-aware algorithms and subsequent reinforcement learning algorithms, reducing computational latency. The network communication module integrates a gigabit Ethernet interface for data interaction and command transmission with other functional units, ensuring real-time command issuance and data feedback. The core hardware specifications of this engine are designed to meet the requirements of high computing speed and multi-task parallel processing, ensuring timely policy generation and adapting to the needs of dynamic adjustments during the session.
[0040] The hardware components of the linkage handover control unit include a handover control chip, a trigger monitoring module, a context extraction module, a cache control module, a media state locking module, and a temporary cache unit. The handover control chip uses an embedded microcontroller to coordinate the collaborative work of all modules within the unit, receiving and executing integrated strategies issued by the decision engine. The trigger monitoring module, composed of a signal monitoring chip and an interface adapter module, monitors the real-time media transmission environment and user operation commands in real time, capturing handover migration trigger signals. The context extraction module uses a dedicated data extraction chip to extract the core digest from the received session context data. The cache control module, composed of a cache controller and an interface module, controls the binding and caching of context digests with non-real-time media states. The media state locking module uses a state monitoring chip to lock the encoding protocol and bitrate parameters of real-time media and generate a locking identifier. The temporary cache unit has a capacity of no less than 8GB and is used to temporarily store the bound context digest and non-real-time media state data. The hardware description of this unit focuses on ensuring the sensitivity of handover trigger monitoring and the security of cached data, ensuring the smooth and efficient operation of the handover control process.
[0041] The session anchor synchronization unit hardware comprises an anchor generation chip, a handover progress monitoring module, a data synchronization module, a resume triggering module, and a synchronization interface module. The anchor generation chip generates session anchor data based on session information, with a processing speed of at least 50 MIPS. The handover progress monitoring module, consisting of a progress monitoring chip and a network status detection module, collects handover migration progress data and target node connection status in real time. The data synchronization module uses a high-speed synchronization chip to support multi-node data synchronization, ensuring rapid synchronization of context summaries and non-real-time media status. The resume triggering module, consisting of a trigger chip and a command issuance module, issues resume commands after handover migration is complete. The synchronization interface module integrates multiple types of interfaces for data interaction with the linked handover control unit, target devices, and non-real-time media transmission nodes. The core hardware specification of this unit is to ensure the timeliness and accuracy of status synchronization, avoiding data loss or deviation during the synchronization process.
[0042] The hardware components of the dual-media state collaborative verification unit include a verification control chip, a multi-dimensional verification module, a deviation detection module, and a compensation control module. The verification control chip coordinates all operations of the verification unit and receives synchronized dual-media state data. The multi-dimensional verification module consists of three independent verification chips, corresponding to the verification of media transmission, state synchronization, and contextual coherence, respectively. The deviation detection module detects deviations occurring during the verification process and identifies the deviation level. The compensation control module triggers the corresponding graded compensation mechanism based on the deviation level and issues compensation commands. The hardware description of this unit focuses on ensuring the comprehensiveness of the verification and the timeliness of deviation handling to guarantee the consistency of the dual-media state.
[0043] The lightweight context-sensitive state management unit's hardware comprises a management control chip, a data filtering module, a cache management module, and a data interaction module. The management control chip coordinates the work of each module within the unit and receives session task intent data. The data filtering module uses a dedicated filtering chip to filter core state data and discard redundant data based on session task intent. The cache management module controls the cache allocation of core state data and manages cache synchronization between the edge computing server and the central cloud server. The data interaction module interacts with the dual-media collaborative orchestration decision engine, transmitting core state data and cache indexes. The hardware description of this unit focuses on optimizing cache resource allocation, reducing resource consumption, and improving data transmission efficiency.
[0044] The hardware components of the scene-intention dual-driven adaptive adjustment unit include an adjustment control chip, a template storage module, an adaptive learning module, a manual intervention interface module, and a parameter adjustment module. The adjustment control chip coordinates the work of all modules within the unit; the template storage module has a capacity of no less than 64GB and stores the scene-intention dual-driven orchestration template library; the adaptive learning module integrates a reinforcement learning algorithm processing chip, responsible for executing reinforcement learning algorithms to optimize orchestration template selection and parameter adjustment; the manual intervention interface module supports administrators in manually adjusting parameters and is equipped with a dedicated interaction chip; the parameter adjustment module is responsible for adjusting parameters such as media priority weights and switching trigger thresholds based on algorithm results and manual instructions. The core of this unit's hardware description is to ensure the flexibility and accuracy of adaptive adjustment, adapting to the needs of different scenes and task intentions.
[0045] The data flow and input / output relationships between modules are as follows: The dual-media fusion acquisition unit receives real-time media network transmission quality data, non-real-time media transmission status data, session initialization signaling, and session control signaling as inputs. After preprocessing, it outputs the preprocessed data to the dual-media collaborative orchestration decision engine. The dual-media collaborative orchestration decision engine receives the preprocessed data transmitted by the dual-media fusion acquisition unit and outputs a dual-media collaborative linkage switching and context continuation integrated strategy, which is transmitted to the linkage switching control unit, the session anchor synchronization unit, and the dual-media status collaborative verification unit, respectively. It also receives core status data transmitted by the lightweight context association status management unit and optimized parameters transmitted by the scene-intent dual-drive adaptive adjustment unit. The linkage switching control unit receives the integrated strategy issued by the decision engine as input and outputs trigger signals, context summaries and non-real-time media status binding data, media status lock flags, and pause commands. The trigger signals are transmitted to the context extraction module, the binding data is cached in the temporary cache unit, and the pause commands are transmitted to the non-real-time media transmission nodes. The session anchor synchronization unit receives the linkage switching control unit as input. The binding data and locking identifiers of the meta-transmission, the integrated strategy issued by the decision engine, and the output are session anchors, synchronization instructions, resume instructions, and synchronized status data. The synchronization instructions are transmitted to the data synchronization module, the resume instructions are transmitted to the non-real-time media transmission node, and the synchronized status data is transmitted to the dual-media status collaborative verification unit. The input of the dual-media status collaborative verification unit is the synchronized status data transmitted by the session anchor synchronization unit and the integrated strategy issued by the decision engine. The output is the verification result and compensation instruction. The verification result is transmitted to the decision engine, and the compensation instruction is transmitted to the corresponding related unit. The input of the lightweight context-related status management unit is the session task intent and various status data transmitted by the decision engine. The output is the filtered core status data and cache index, which are transmitted to the decision engine, edge computing server, and central cloud server. The input of the scenario-intent dual-driven adaptive adjustment unit is user operation behavior, network status data, and session task intent change data. The output is the optimized orchestration template and parameter adjustment instructions, which are transmitted to the dual-media collaborative orchestration decision engine. At the same time, it receives adjustment instructions input by the administrator through the manual intervention interface. All data streams are transmitted through a dedicated data channel to ensure the security and real-time performance of data transmission. The input-output relationship strictly corresponds to the workflow of Implementation Example 1, with no logical conflicts. Specific Implementation Example 3:
[0047] like Figures 1 to 6 As shown, the algorithm in Example 1 will be further explained below: The input data for the context-aware algorithm primarily comes from the preprocessed session context data of the dual-media fusion acquisition unit, and also includes keywords, task node tags, and media type features extracted from this data. This input data undergoes denoising and standardization to ensure accuracy and consistency, providing reliable support for the algorithm's operation. The input data uses a fixed format for rapid parsing and processing, and its update frequency matches the acquisition frequency of the session context data, ensuring the algorithm can acquire the latest session information in real time.
[0048] The output structure of the context-aware algorithm mainly consists of two parts. One part is the conversation context feature vector, which is composed of extracted keywords, task node tags, and media type features. Each feature corresponds to one dimension of the vector, and the value of each dimension represents the weight of that feature. The vector dimension is dynamically adjusted according to the number of extracted features to ensure that it can comprehensively reflect the core information of the conversation context. The other part is the conversation task intent recognition result, which is the specific intent category, along with the similarity value of that intent category, used to judge the reliability of the recognition result. If the similarity of the recognition result is lower than a preset threshold, a recognition failure prompt is output to trigger subsequent re-recognition operations.
[0049] The input data for the reinforcement learning algorithm primarily comes from various data collected by the scene-intent dual-driven adaptive adjustment unit. Specifically, this includes user action behavior data, network status data, and session task intent change data. User action behavior data includes user adjustments to media priority and switching / migrating operations; network status data includes network bandwidth, latency, and packet loss rate; and session task intent change data includes information on task intent switching during the session. This input data is collected and updated in real time, ensuring the algorithm can promptly grasp changes in system operating status and the session environment, providing data support for iterative optimization.
[0050] The output structure of the reinforcement learning algorithm mainly consists of two parts. One part is the updated Q-value of the state-action pair. This Q-value reflects the expected long-term adaptation benefit of performing a certain orchestration parameter adjustment action in the current session state. The Q-value ranges from 0 to 1. The higher the value, the better the adaptation effect of the adjustment action. The other part is the optimized orchestration strategy adjustment instruction. This instruction includes the adaptation scene-intent orchestration template selection result, media priority weight adjustment parameters, switching trigger threshold adjustment parameters, and caching strategy adjustment parameters. The instruction format is uniform, which is convenient for the dual-media collaborative orchestration decision engine to parse and execute. Specific Implementation Example 4:
[0052] like Figures 1 to 6 As shown, the following are use cases in different scenarios: Case 1: Indoor Office Dual-Media Conversation Scenario. Employees of a company conduct remote collaborative work, simultaneously engaging in real-time video communication and non-real-time file transfer, forming a dual-media conversation. After system startup, the dual-media fusion acquisition unit begins operation, collecting network transmission quality data such as latency, jitter, and bandwidth utilization of real-time video, while simultaneously collecting transmission status data such as the amount of data transmitted, the amount of data remaining, and the cache latency of non-real-time files. The conversation scenario identifier is extracted as an office scenario, and conversation context data such as participant identities, current collaborative task nodes, and the most recent file transfer operation record are collected synchronously. After collection, the preprocessing module performs noise reduction and standardization on all data, transmitting the processed data to the dual-media collaborative orchestration decision engine. The decision engine extracts keywords and features from the data using a context-aware algorithm, identifies the conversation task intent as remote collaborative office work, matches the baseline weight with the office scenario identifier, introduces an intent adjustment coefficient for correction, determines the final priority weight of real-time and non-real-time media, and generates an integrated strategy. This strategy clarifies the triggering conditions for cross-device switching, context caching rules, and dual-media status synchronization methods, and is distributed to the linkage switching control unit, the conversation anchor synchronization unit, and the dual-media status collaborative verification unit, respectively. When an employee needs to switch from their office computer to a tablet to continue the conversation, the linkage switching control unit detects the user's switching command, extracts the core context summary, binds it to the non-real-time file transfer status and caches it, and locks the encoding protocol and bitrate parameters of the real-time video. The session anchor synchronization unit generates a session anchor point, monitors the switching progress, and once the switching completion rate reaches the target and the tablet connection is stable, it synchronizes the context summary and file transfer status, triggers file resuming, and synchronizes the video playback progress. The dual-media status collaborative verification unit verifies from three dimensions: media transmission, status synchronization, and context coherence, confirming that the status is consistent. The lightweight management unit filters core data to reduce resource consumption. The adaptive adjustment unit optimizes the orchestration strategy parameters based on the current network status and user operations. Ultimately, this achieves seamless cross-device switching, smooth real-time video communication, and normal non-real-time file resuming, ensuring the continuity of remote collaborative work.
[0053] Case Study 2: Outdoor Mobile Dual-Media Conversation Scenario. A worker is working outdoors, simultaneously conducting real-time voice communication and transmitting non-real-time data reports, constituting a dual-media conversation. After system startup, the dual-media fusion acquisition unit adapts to the fluctuating characteristics of outdoor mobile networks, frequently collecting network transmission quality data such as latency and jitter for real-time voice, while simultaneously collecting data report transmission status data. It extracts the conversation scenario identifier as an outdoor mobile scenario and collects conversation context data such as worker identity, current work task node, and the most recent report transmission operation record. After preprocessing, the data is transmitted to the dual-media collaborative orchestration decision engine. The decision engine uses a context-aware algorithm to identify the conversation task intent as outdoor work communication and data reporting. Based on the outdoor scenario baseline weight, it prioritizes real-time voice transmission, generates an integrated strategy, clarifies network migration trigger conditions, caching rules, and synchronization methods, and distributes it to all relevant units. When staff switch from an outdoor mobile network to an indoor WiFi network, the linkage switching control unit detects the network status change as a trigger signal, extracts the context summary, binds and caches the report transmission status, and locks the voice encoding protocol and bitrate parameters. The session anchor synchronization unit monitors the network migration progress, synchronizes the context and report transmission status, triggers report continuation, and synchronizes the voice progress. After the status verification unit verifies that there are no deviations, the lightweight management unit only transmits the changed core data to adapt to outdoor network bandwidth fluctuations. The adaptive adjustment unit adjusts the caching strategy and switching trigger threshold according to network status changes. Ultimately, this achieves continuous real-time voice communication and uninterrupted transmission of non-real-time data reports during network migration, adapting to the work requirements of outdoor mobile scenarios. Specific Implementation Example 5:
[0055] like Figures 1 to 6 As shown, the following experimental data compares this solution with traditional techniques:
[0056] This experiment used the same experimental environment, session scenario, and task requirements to conduct multiple parallel tests on both the proposed system and a traditional dual-media transmission system. The traditional dual-media transmission system refers to the commonly used dual-media processing systems in the industry. Its core characteristics are independent transmission of real-time and non-real-time media, lack of collaborative orchestration mechanisms, and no unified context synchronization and status verification process during handover. It can only achieve basic media transmission functions and is prone to handover interruptions, context breaks, and excessive resource consumption. This test covers core indicators related to dual-media session continuity, and all data are averages from multiple experiments.
[0057] The experimental data shows that the proposed solution outperforms traditional dual-media transmission systems in all test indicators. The data fully demonstrates that the proposed solution, through the complete process of "data acquisition - strategy generation - switching control - state synchronization - verification compensation - lightweight management - adaptive adjustment", effectively improves the continuity, fluency and stability of dual-media sessions, reduces resource consumption and operating costs, and enhances the adaptability and practicality of the system.
[0058] It should be noted that, in this document, relational terms such as "first" and "second" are used merely to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising a reference structure" does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes the element.
[0059] Although embodiments of the invention have been shown and described, it will be understood by those skilled in the art that various changes, modifications, substitutions and alterations can be made to these embodiments without departing from the principles and spirit of the invention, the scope of which is defined by the appended claims and their equivalents.
Claims
1. A method for arranging the continuity of a dual-media conversation, characterized in that, Includes the following steps: S1. Deploy a dual-media fusion acquisition unit in a dual-media session scenario to collect network transmission quality data of real-time media, transmission status data of non-real-time media, and session scenario identifier and session context data respectively. S2. Build a dual-media collaborative orchestration decision engine with a built-in scene-strategy mapping model. Receive various data collected in step S1, identify the intent of the session task through a context-aware algorithm, and dynamically allocate the priority weights of the two media in combination with the session scene identifier to generate an integrated strategy for dual-media collaborative linkage switching and context continuation. S3. Configure and start the linkage switching control unit, extract the session context summary, bind the cache with the current transmission status of non-real-time media, control the non-real-time media to pause transmission, and lock the real-time media encoding protocol type, bitrate parameters and playback progress. S4. Set up a session anchor point synchronization unit to monitor the switching and migration progress. Synchronize the context summary and non-real-time media status to the target device or network node through the session anchor point. After the switching and migration is completed, trigger the non-real-time media to resume playback and synchronize the real-time media playback progress and context summary. S5. Deploy a dual-media state collaborative verification unit. After switching and migration, verify the state consistency from three dimensions: media transmission, state synchronization, and context coherence. Determine if there is any transmission deviation, state disconnection, or context breakage. If the state is consistent, maintain the orchestration state. If there is a deviation, trigger a hierarchical compensation mechanism to correct the deviation and continue the context.
2. The dual-media session continuity orchestration method according to claim 1, characterized in that: The method for deploying the dual-media fusion acquisition unit and acquiring related data in S1 is as follows: S1.1 The dual-media fusion acquisition unit collects real-time media network transmission quality data through a preset monitoring interface. The network transmission quality data includes latency, jitter, and bandwidth utilization. While acquiring the network transmission quality data, the dual-media fusion acquisition unit parses the real-time media data stream to extract the encoding protocol type and bitrate parameters. S1.2 The dual-media fusion acquisition unit establishes a data channel with the storage service of non-real-time media to obtain the transmission status data of non-real-time media in real time. The transmission status data of non-real-time media includes the amount of data transmitted, the amount of data remaining, the data retention time in the current cache, the preset file transmission priority, and the file format type. S1.3 The dual-media fusion acquisition unit extracts the scene identification code carried in the field by parsing the field in the session initialization signaling, and uses it as the current session scene identifier. The session scene identifier is generated based on the initial configuration information when the session is initiated. S1.4 The dual-media fusion acquisition unit synchronously collects session context data by listening to session control signaling and recording interaction logs. The session context data includes the identity identifiers of the participants, the current task nodes in the session flow, and the record of the most recent key operation that affects the continuity of the session. S1.5 The dual-media fusion acquisition unit has a built-in data preprocessing module to perform noise reduction and standardization processing on the acquired data.
3. The dual-media session continuity arrangement method according to claim 2, characterized in that: The specific steps for building the dual-media collaborative orchestration decision engine and related operations in step S2 are as follows: S2.1 Process the conversation context data collected in step S1.4 using a context-aware algorithm, extract keywords, task node markers and media type features from the context data, calculate the similarity between the extracted features and the preset intent feature library, and determine the matched intent category as the conversation task intent when the similarity is higher than the preset threshold. S2.
2. Determine the baseline weights of real-time media and non-real-time media in the scenario based on the session scenario identifiers collected in S1.
3. Then, dynamically adjust the baseline weights based on the session task intent identified in S2.1 to obtain the final priority weights of real-time media and non-real-time media. S2.
3. Based on the final priority weight allocated in S2.2, generate an integrated strategy for dual-media collaborative switching and context continuation. The integrated strategy includes triggering conditions for cross-device or cross-network switching and migration, caching rules for context summary and non-real-time media status, and dual-media status synchronization methods.
4. The dual-media session continuity orchestration method according to claim 1, characterized in that: The linkage switching control unit consists of a switching trigger monitoring module, a context extraction module, a cache control module, a media state locking module, and a temporary cache unit. The switching trigger monitoring module monitors the real-time media transmission environment in real time and identifies trigger signals for cross-device switching and cross-network migration. The linkage switching control unit and related operation methods include the following steps: S3.1 When the switching trigger monitoring module detects a real-time media trigger for a cross-device or cross-network switching or migration, it sends a trigger signal to the context extraction module. The context extraction module extracts a session context summary directly related to the current task node from the session context data collected in step S1. S3.2 The cache control module binds the extracted session context summary with the current transmission status of the non-real-time media and caches it in the temporary cache unit. At the same time, it sends a pause command to the non-real-time media transmission node to control the non-real-time media to pause transmission. S3.3 The media status locking module locks the current encoding protocol type and bitrate parameters of the real-time media and generates a status lock identifier.
5. The dual-media session continuity orchestration method according to claim 1, characterized in that: The session anchor point synchronization unit consists of an anchor point generation module, a handover progress monitoring module, a data synchronization module, and a resume triggering module. The operation steps of the session anchor point synchronization unit are as follows: S4.
1. Generate session anchor points through the anchor point generation module, and monitor real-time media switching and migration progress through the switching progress monitoring module; S4.
2. The switching progress monitoring module collects real-time media switching and migration progress data in real time. When the switching and migration progress reaches the threshold, a synchronization command is sent to the data synchronization module. S4.3 The data synchronization module synchronizes the context summary and non-real-time media state in the temporary cache unit to the corresponding cache unit of the target device or network node through the session anchor point. S4.4 After the switching and migration are completed, the resume trigger module sends a resume instruction to the non-real-time media transmission node, triggering the non-real-time media to resume transmission based on the synchronous transmission status, while synchronizing the encoding protocol type, bitrate parameters and playback progress of the real-time media.
6. The dual-media session continuity orchestration method according to claim 1, characterized in that: The method also includes a lightweight context-associated state management step, as follows: Filter core state data based on the session task intent and discard redundant data; Core state data and context summary are cached on edge computing servers, while the central cloud server only caches the core data index. The two interact through a network channel. During session switching and migration, only the changed core state data and context update content are transmitted.
7. The dual-media session continuity orchestration method according to claim 1, characterized in that: The method further includes a scene-intention dual-driven adaptive adjustment step, as detailed below: Build a scenario-intention dual-driven orchestration template library, pre-store orchestration templates adapted to different scenarios and task intentions; Deploy template adaptive learning units to collect user operation behavior, network bandwidth, latency, packet loss rate, and changes in session task intent, automatically select suitable orchestration templates, and adjust media priority weights, switching trigger thresholds, and caching strategy parameters through reinforcement learning algorithms; set up tiered manual intervention interfaces to support administrators in manually adjusting orchestration strategies and parameters, and retain operation records for traceability.
8. A dual-media session continuity orchestration system, used to implement the dual-media session continuity orchestration method as described in claim 1, characterized in that: It includes a dual-media fusion acquisition unit, a dual-media collaborative orchestration decision engine, a linkage switching control unit, a session anchor point synchronization unit, and a dual-media status collaborative verification unit. Each unit realizes data interaction and command transmission through a data channel, and collaboratively completes the continuous orchestration of dual-media sessions. The dual-media fusion acquisition unit is used to deploy and execute the data acquisition and preprocessing operations of step S1 in a dual-media session scenario; The dual-media collaborative orchestration decision engine has a built-in scene-policy mapping model, which is used to receive data collected by the dual-media fusion acquisition unit and perform the S2 step of session task intent recognition, priority weight allocation and integrated policy generation operation. The linkage switching control unit consists of a switching trigger monitoring module, a context extraction module, a cache control module, a media state locking module, and a temporary cache unit, and is used to execute the linkage switching control operation in step S3. The session anchor point synchronization unit consists of an anchor point generation module, a switching progress monitoring module, a data synchronization module, and a resume triggering module, and is used to perform the session anchor point generation, state synchronization, and resume triggering related operations in step S4. The dual-media state collaborative verification unit is used to perform dual-media state consistency verification and deviation compensation triggering operations after the switching and migration of step S5 is completed.
9. A dual-media session continuity orchestration system according to claim 8, characterized in that: The system also includes a lightweight context-related state management unit and a scene-intent dual-driven adaptive adjustment unit, both of which interact with the dual-media collaborative orchestration decision engine.
Citation Information
Patent Citations
Automatic seamless context sharing across multiple devices
CN104782220A
Method and system for managing session across multiple electronic devices in network system
CN108781360A