Semantic parsing and mapping method for cross-platform touch instruction of same-screen device

By introducing semantic analysis and mapping methods of cross-platform touch instructions in the same-screen device, the compatibility and adaptation problems in cross-platform interactions in the existing technology are solved, efficient cross-platform semantic adaptation and multi-source instruction collaboration are achieved, and the system flexibility and user experience are improved.

CN120104032AActive Publication Date: 2025-06-06SHENZHEN XINZHENGYU TECH

Patent Information

Application Number
CN202510579575.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-07
Publication Date
2025-06-06
Estimated Expiration
2045-05-07

AI Technical Summary

Technical Problem

In the cross-platform interaction, existing same-screen devices have problems such as insufficient compatibility of touch commands, rigid adaptation caused by static mapping rules, and missing multi-source command conflicts and processing.

Method used

A semantic analysis and mapping method for cross-platform touch instructions of the same-screen device is proposed. By obtaining the touch instructions of the source device, semantic analysis is performed to generate intermediate semantic layer instructions, and dynamically mapped into native instructions according to the platform protocol and interaction rules of the target device, supporting multi-source instruction integration and conflict processing.

Benefits of technology

It realizes cross-platform semantic adaptation capabilities, improves the abstract expression of touch instructions, solves cross-platform gesture compatibility problems, supports real-time response to changes in the target device interface, reduces development and maintenance costs, and provides multi-source instruction coordination and conflict resolution capabilities.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120104032A_ABST
    Figure CN120104032A_ABST
Patent Text Reader

Abstract

The invention discloses a semantic parsing and mapping method for cross-platform touch instructions of a same-screen device, and the method comprises the steps: S1, obtaining a touch instruction of a source device through the same-screen device, the touch instruction comprising a touch position, a touch gesture type and touch time sequence information; s2, based on an operating system type and an application scene of a target device, semantic analysis is conducted on the touch instruction, an intermediate semantic layer instruction is generated, and the intermediate semantic layer instruction comprises operation logic, an interaction intention and a target control identifier; s3, dynamically mapping the intermediate semantic layer instruction into a native instruction of the target device according to a platform protocol and an interaction rule of the target device, the native instruction being adapted to a touch event processing mechanism of the target device; and S4, sending the native instruction to the target equipment for execution, and synchronously updating the display interface of the same-screen device and the display content of the target equipment.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of on-screen devices, and in particular to a method for semantic parsing and mapping cross-platform touch instructions of on-screen devices. Background Art

[0002] With the diversification of smart devices and the growing demand for cross-platform interaction, screen sharing devices, as the core tool for connecting different devices, have gradually become a key technology in home entertainment, business meetings and education scenarios. However, existing technologies still have the following limitations: Insufficient compatibility of touch commands: Traditional screen mirroring devices usually only support one-way content projection (such as video and picture transmission), and cannot achieve cross-platform interaction of touch commands. For example, there is a semantic difference between the Force Touch gesture of iOS devices and the long press event of Android systems. Directly forwarding touch signals will cause operation failure or logical confusion; Static mapping rules lead to rigid adaptation: Existing solutions (such as the wireless screen projection device mentioned in page 4) mostly use fixed protocol conversion rules and cannot dynamically adapt to the interaction logic of different operating systems. For example, when the target device interface is updated or the control type changes, the original mapping rules may become invalid and need to be manually reconfigured; Multi-source command conflicts and lack of processing: In a multi-device collaboration scenario, when multiple source devices send touch commands at the same time, the existing technology lacks an effective conflict resolution mechanism, resulting in operation delays or execution errors. Summary of the invention

[0003] The present invention aims to solve at least one of the technical problems in the related art to a certain extent. To this end, one object of the present invention is to propose a semantic parsing and mapping method for cross-platform touch commands of a same-screen device, comprising: S1: Acquire a touch command of a source device through a screen sharing device, where the touch command includes touch position, touch gesture type and touch timing information; S2: Based on the operating system type and application scenario of the target device, semantically parse the touch control instruction to generate an intermediate semantic layer instruction, where the intermediate semantic layer instruction includes operation logic, interaction intent, and target control identifier; S3: dynamically mapping the intermediate semantic layer instructions to native instructions of the target device according to the platform protocol and interaction rules of the target device, wherein the native instructions are adapted to the touch event processing mechanism of the target device; S4: Send the native command to the target device for execution, and synchronously update the display interface of the same-screen device and the display content of the target device.

[0004] Preferably, the semantic analysis in S2 further includes: Matching the interaction logic of the target device through a dynamic adaptation rule base, wherein the rule base includes operating system protocols, application interface control types, and touch event priorities of different platforms; When it is detected that the control type of the target device is a dynamically generated control, the intermediate semantic layer instructions are optimized based on the control hierarchy and context semantics.

[0005] Preferably, step S2 further includes multi-level semantic analysis: The first level analyzes the physical parameters of touch gestures, including sliding direction, pressure value, and multi-touch coordinates; The second level combines the current display interface layout of the target device to identify the functional modules corresponding to the touch operation; The third level dynamically adjusts the semantic weights according to the application scenarios, giving priority to mapping high-frequency operation instructions.

[0006] Preferably, the mapping method supports the integration of multi-source touch commands, including: Receive touch commands from multiple source devices and sort the commands by timestamp and event queue; When multiple command conflicts are detected, the execution command is selected based on user identity permissions or operation priority.

[0007] Preferably, the dynamic mapping in S3 further includes: Monitor the interface status changes of the target device in real time. If the interface update causes the control identifier to become invalid, it triggers the re-parsing of the semantic layer instructions. When the target device is a virtual reality device, the touch commands are mapped to three-dimensional space interaction events, and the display content of the virtual scene and the physical screen is synchronized.

[0008] Preferably, the method further comprises a conflict handling mechanism: When the touch protocols of the source device and the target device are incompatible, simulated touch events are generated, including virtual clicks, sliding track simulation, and time compensation for long press events; The execution results are returned to the source device through the feedback channel to optimize the semantic parsing rules of subsequent instructions.

[0009] Preferably, the mapping method supports cross-platform protocol conversion, including: Convert mouse events of Windows system to MotionEvent events of Android system; Map the Force Touch gesture of the iOS system to a long press event of the Android system, and dynamically adjust the mapping parameters according to the pressure value.

[0010] Preferably, the dynamic adaptation rule base is optimized by machine learning: Collect users' correction operations on mapping results and generate training data sets; Based on the neural network model, the mapping strategy in the rule base is updated to improve the accuracy of semantic parsing.

[0011] Preferably, the mapping method supports multi-window collaborative control: When the target device displays multiple application windows in split screen, the command is mapped to the corresponding application process according to the window area to which the touch position belongs; When the display interface of the same screen device is a multi-window layout, cross-window command distribution is achieved by dividing the sub-interaction area.

[0012] Preferably, the mapping method further includes a device switching mechanism: When the target device is detected to be switched, the context state of the current intermediate semantic layer instruction is retained; Remaps commands and restores continuity of operation flow based on the platform protocol of the new target device.

[0013] The above solution of the present invention includes at least the following beneficial effects: Cross-platform semantic adaptation capabilities Abstract expression of touch commands is achieved through intermediate semantic layer commands (such as operation logic, interaction intent, and target control identification), effectively isolating protocol differences between different platforms. For example, the 3D Touch pressure value of iOS can be dynamically mapped to the duration parameter of the long press event of Android, solving the problem of cross-platform gesture compatibility; Combined with the dynamic adaptation rule library (including operating system protocols, control type priorities, etc.), it supports real-time response to changes in the target device interface to avoid mapping failures caused by control updates; Multi-level semantic analysis and dynamic optimization A three-level semantic analysis architecture (physical parameters → functional modules → scene weights) is used to accurately identify touch intent. For example, the sliding direction and pressure value combined with the current interface layout can distinguish between "page turning" and "zooming" operations, reducing the misjudgment rate; Introducing a machine learning optimization mechanism to iteratively update the rule base through user feedback data to improve the mapping efficiency and accuracy of high-frequency operations (such as game touch); Multi-source instruction coordination and conflict resolution Through timestamp sorting and permission management, the problem of multi-device command conflicts can be solved. For example, in a conference scenario, the speaker's device has operation priority to avoid interference from multiple users' operations; Supports three-dimensional mapping of virtual reality devices, converts plane touch commands into spatial interaction events (such as mapping gesture trajectories to object rotation in VR scenes), and expands application scenarios; Reduce development and maintenance costs Through the protocol conversion module (such as Windows mouse event → Android MotionEvent), developers can reduce the repeated adaptation work for different platforms; Seamless device switching: The context state of the intermediate semantic layer is retained to ensure the continuity of the operation process when switching the target device, improving user efficiency.

[0014] Additional aspects and advantages of the present invention will be given in part in the following description and in part will be obvious from the following description, or will be learned through practice of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS

[0015] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on the structures shown in these drawings without paying creative work.

[0016] Figure 1 It is a flow chart of a semantic parsing and mapping method of cross-platform touch control instructions of a same-screen device provided in an embodiment of the present invention; Figure 2 is a flow chart of acquiring touch control instructions provided in an embodiment of the present invention; Figure 3 is a flowchart providing semantic parsing in an embodiment of the present invention; Figure 4 is a flow chart providing dynamic mapping in an embodiment of the present invention; Figure 5 is a flowchart of performing synchronization in an embodiment of the present invention; Figure 6 is a semantic weight table provided in an embodiment of the present invention; Figure 7 is a diagram describing the types of conflict scenarios and resolution strategies provided in an embodiment of the present invention; Figure 8 is a corresponding relationship diagram between touch gestures and three-dimensional interactive actions provided in an embodiment of the present invention; Fig. 9 is a pressure level diagram provided in an embodiment of the present invention; Fig.10 It is a protocol compatibility matrix provided in an embodiment of the present invention.

[0017] The realization of the purpose, functional features and advantages of the present invention will be further explained in conjunction with embodiments and with reference to the accompanying drawings. DETAILED DESCRIPTION

[0018] The embodiments of the present invention are described in detail below, and examples of the embodiments are shown in the accompanying drawings, wherein the same or similar reference numerals throughout represent the same or similar elements or elements having the same or similar functions. The embodiments described below with reference to the accompanying drawings are exemplary and are intended to be used to explain the present invention, but are not to be construed as limiting the present invention. All other embodiments obtained by ordinary technicians in the field without creative work based on the embodiments of the present invention are within the scope of protection of the present invention.

[0019] In the description of the present invention, it should be understood that the terms "center", "longitudinal", "lateral", "length", "width", "thickness", "up", "down", "front", "back", "left", "right", "vertical", "horizontal", "top", "bottom", "inside", "outside", "clockwise", "counterclockwise", "axial", "circumferential", "radial" and the like indicate orientations or positional relationships based on the orientations or positional relationships shown in the accompanying drawings, and are only for the convenience of describing the present invention and simplifying the description, and do not indicate or imply that the referred device or element must have a specific orientation, be constructed and operated in a specific orientation, and therefore should not be understood as limiting the present invention.

[0020] In addition, the terms "first" and "second" are used for descriptive purposes only and should not be understood as indicating or implying relative importance or implicitly indicating the number of the indicated technical features. Therefore, the features defined as "first" and "second" may explicitly or implicitly include one or more of the features. In the description of the present invention, the meaning of "plurality" is two or more, unless otherwise clearly and specifically defined.

[0021] In the present invention, unless otherwise clearly specified and limited, the terms "installed", "connected", "connected", "fixed" and the like should be understood in a broad sense, for example, it can be a fixed connection, a detachable connection, or an integral connection; it can be a mechanical connection or an electrical connection; it can be a direct connection, or it can be an indirect connection through an intermediate medium, or it can be the internal communication of two components. For ordinary technicians in this field, the specific meanings of the above terms in the present invention can be understood according to specific circumstances.

[0022] In the present invention, unless otherwise clearly specified and limited, a first feature being "above" or "below" a second feature may include that the first and second features are in direct contact, or may include that the first and second features are not in direct contact but are in contact through another feature between them. Moreover, a first feature being "above", "above" and "above" a second feature includes that the first feature is directly above and obliquely above the second feature, or simply indicates that the first feature is higher in level than the second feature. A first feature being "below", "below" and "below" a second feature includes that the first feature is directly below and obliquely below the second feature, or simply indicates that the first feature is lower in level than the second feature.

[0023] The following describes in detail the semantic parsing and mapping method of cross-platform touch commands of a same-screen device according to an embodiment of the present invention with reference to the accompanying drawings.

[0024] See also Figures 1 to 5 , in this embodiment, includes: S1: Acquire the touch command of the source device through the screen sharing device, where the touch command includes the touch position, touch gesture type and touch timing information; S2: Based on the operating system type and application scenario of the target device, the touch command is semantically parsed to generate an intermediate semantic layer command, which includes the operation logic, interaction intent and target control identifier; S3: According to the platform protocol and interaction rules of the target device, the intermediate semantic layer instructions are dynamically mapped to the native instructions of the target device, and the native instructions are adapted to the touch event processing mechanism of the target device; S4: Send the native command to the target device for execution, and synchronously update the display interface of the same-screen device and the display content of the target device; Cross-platform semantic adaptation capabilities Abstract expression of touch commands is achieved through intermediate semantic layer commands (such as operation logic, interaction intent, and target control identification), effectively isolating protocol differences between different platforms. For example, the 3D Touch pressure value of iOS can be dynamically mapped to the duration parameter of the long press event of Android, solving the problem of cross-platform gesture compatibility; Combined with the dynamic adaptation rule library (including operating system protocols, control type priorities, etc.), it supports real-time response to changes in the target device interface to avoid mapping failures caused by control updates; Multi-level semantic analysis and dynamic optimization A three-level semantic analysis architecture (physical parameters → functional modules → scene weights) is used to accurately identify touch intent. For example, the sliding direction and pressure value combined with the current interface layout can distinguish between "page turning" and "zooming" operations, reducing the misjudgment rate; Introducing a machine learning optimization mechanism to iteratively update the rule base through user feedback data to improve the mapping efficiency and accuracy of high-frequency operations (such as game touch); Multi-source instruction coordination and conflict resolution Through timestamp sorting and permission management, the problem of multi-device command conflicts can be solved. For example, in a conference scenario, the speaker's device has operation priority to avoid interference from multiple users' operations; Supports three-dimensional mapping of virtual reality devices, converts plane touch commands into spatial interaction events (such as mapping gesture trajectories to object rotation in VR scenes), and expands application scenarios; Reduce development and maintenance costs Through the protocol conversion module (such as Windows mouse event → Android MotionEvent), developers can reduce the repeated adaptation work for different platforms; Seamless device switching: The context state of the intermediate semantic layer is retained to ensure the continuity of the operation process when switching the target device, thereby improving user efficiency.

[0025] In this embodiment, the semantic parsing in S2 further includes: matching the interaction logic of the target device through a dynamic adaptation rule base, the rule base includes operating system protocols, application interface control types and touch event priorities of different platforms; when it is detected that the control type of the target device is a dynamically generated control, optimizing the intermediate semantic layer instructions based on the control hierarchy structure and context semantics; Dynamic Adaptation Rule Base Construction Operating system protocol sub-library: records the touch event definitions of different platforms (Android / iOS / Windows), including Android's MotionEvent event type, iOS's UITouch properties, and Windows's WM_TOUCH message structure; Control type mapping table: defines the equivalent relationship between cross-platform controls. For example, Android's RecyclerView and iOS's UICollectionView are considered as the same type of scrolling containers, and their sliding event thresholds are associated. Priority strategy table: Set the execution priority of touch events in different scenarios. For example, in the video playback interface, "sliding the progress bar" takes precedence over "clicking the pause button".

[0026] Identification and processing of dynamically generated controls When the target device interface contains dynamically generated controls (such as web controls dynamically loaded through JavaScript or custom components rendered by Flutter), perform the following operations: Hierarchical structure analysis: obtain the control tree of the current interface of the target device through the screen sharing device, and analyze the parent-child hierarchical relationship and properties of the control (such as ID, class name, visibility); Contextual semantic analysis: Infer the expected function of the control based on the interaction logic of the container where the control is located (for example, the sliding direction and data loading method of a dynamic list); Optimized instruction generation: If a control does not have a unique identifier, a temporary control identifier is generated based on the hierarchical path (such as Root / FrameLayout[2] / ListView[0] / Button) and a semantic annotation is added (such as "the third item on the shopping cart page Add to cart button").

[0027] Rule matching and instruction correction According to the operating system type of the target device, extract the matching mapping policy from the rule base: Example 1: When a dynamically generated button in a WebView is detected on an Android device, the touch coordinates are converted to relative coordinates in the web page and a simulated click event is injected through JavaScript. Example 2: In an AR application on an iOS device, if a 3D control (such as a virtual button) is dynamically generated, the mapping parameters of the touch coordinates are corrected based on the projection position of the control in the three-dimensional space.

[0028] Machine Learning Optimization Rule Base Introducing a feedback learning mechanism to continuously optimize the dynamic adaptation rule base: Collect the user's correction operations on the mapping results (such as manually reselecting the target control or adjusting the sliding distance) to generate a labeled dataset; Train a convolutional neural network (CNN) model to identify the association pattern between the control hierarchy and touch intent; When there is no matching strategy in the rule base, the trained model is called to predict the best mapping solution and automatically add new rules to the database.

[0029] This embodiment achieves the following technical advantages by refining the application logic of the dynamic adaptation rule library and the processing flow of dynamically generating controls: Improve cross-platform control adaptation efficiency Solve the problem of identifying dynamic controls in hybrid development frameworks such as WebView and Flutter through control type mapping tables and hierarchical structure analysis; Semantic annotation generation enables controls without IDs to be accurately located, avoiding mapping failures caused by interface revisions; Enhanced robustness in complex scenarios 3D space coordinate correction supports dynamic control interaction of AR / VR applications; Priority strategy table ensures that critical operations are responded to first; Reduce manual maintenance costs The machine learning optimization mechanism can automatically learn user operation habits. For example, for combo operations in game scenarios, the model automatically optimizes the mapping parameters of click intervals, reducing the workload of developers' manual debugging. The dynamic expansion capability of the rule base supports rapid adaptation to new operating system versions; Improve user experience In dynamic lists (such as infinite scrolling feeds in social apps), the sliding distance mapping algorithm based on contextual semantics makes the page turning operation smoother across platforms and consistent with native devices; Continuously optimize the rule base through feedback learning to reduce the user error rate.

[0030] In this embodiment, step S2 also includes multi-level semantic analysis: The first level analyzes the physical parameters of touch gestures, including sliding direction, pressure value, and multi-touch coordinates; Touch gestures go digital: The touch signal acquisition module of the same screen device can capture the original touch data of the source device in real time, including: Sliding direction (8-way angle quantization, accuracy ±5°); Pressure value (graded sampling, 0-1N range is divided into 256 levels); Multi-touch coordinates (normalized to relative coordinates between 0 and 1 based on the target device screen resolution); Perform noise filtering and normalization on the raw data, for example: Use Kalman filtering to eliminate touch track jitter; Converts absolute coordinates of different devices (such as iPad and Surface) to relative coordinates of the target device screen.

[0031] Gesture type classification: Establish touch gesture feature vectors, including sliding speed, acceleration, and contact area change rate; A lightweight convolutional neural network (MobileNetV3) is used to classify gestures into preset types (single click, double click, long press, slide, pinch) in real time.

[0032] The second level combines the current display interface layout of the target device to identify the functional modules corresponding to the touch operation; Dynamic modeling of interface layout Get a screenshot of the current display interface of the target device and extract UI elements (buttons, lists, input boxes) through OCR and image recognition technology; Build a topological diagram of the interface layout, marking the function type of each control (such as "play / pause button", "scroll list", "text input area"); Touch intent association: Spatially match the touch coordinates to the interface layout: If the touch point falls on the "play button" area, the operation intention is marked as "media control"; If the touch track spans multiple list items, the mark intent is "scroll through"; Combined with the application process information (such as whether the currently running app is a video player or a game), further refine the intent (such as "video progress adjustment" or "character movement").

[0033] The third level dynamically adjusts the semantic weight according to the application scenario, giving priority to mapping high-frequency operation instructions; Scene perception and weight allocation Define multi-dimensional scene labels (such as "game mode", "document editing", "video conferencing") and determine the scene in real time through the following methods: Monitor the application window focus changes of the target device; Analyze the spatiotemporal distribution characteristics of touch command sequences (e.g., high-frequency clicks may be game operations); Set semantic weight tables for different scenarios (such as Figure 6 shown): In the weight calculation module, high-frequency instructions in the current scene (such as the "jump" combo in the game) are dynamically accelerated: Establish a command history queue and count the triggering frequency of each command type in the past 5 seconds; When it is detected that the frequency of a certain instruction exceeds the threshold (such as 3 times per second), its mapping priority is automatically increased; Through the instruction preloading mechanism, the native instruction cache of the target platform is generated in advance to reduce transmission delay.

[0034] This embodiment achieves the following technical improvements by refining the multi-level semantic parsing process: The accuracy of touch intention recognition is significantly improved Refined processing of physical parameters: Through Kalman filtering and coordinate normalization, the touch track error of different devices can be reduced to ±2 pixels. For example, the tilted writing track of Surface Pen can be accurately mapped to the pressure sensing event of Apple Pencil on iPad. Functional module association accuracy: Interface layout modeling reduces the misrecognition rate of "play button click" from 18% in traditional solutions to 3%, especially for complex UIs (such as Photoshop toolbar); Enhanced dynamic scene adaptation capabilities Game mode optimization: In MOBA games, the command response delay of high-frequency skill combos can be shortened from 120ms to 45ms, meeting e-sports level operation requirements; Improved document editing efficiency: The mapping accuracy of pinch-to-zoom and long-press selection has been increased to 98%, supporting seamless cross-platform editing (e.g. Windows trackpad → iPad version of WPS); Resource utilization optimization Instruction preloading mechanism: By caching the native code of high-frequency instructions, the CPU usage rate is reduced by 22% (from 35% to 27%), especially on low-end devices (such as Android thousand-yuan phones). Dynamic weight allocation: In video conferencing scenarios, the response priority of key operations (such as the "Share Screen" button) is increased, and the probability of accidentally touching other controls is reduced by 60%; Consistent experience across platforms In hybrid office scenarios, the three-finger swipe on the Windows device’s touchpad can be accurately mapped to the MacBook’s Mission Control gesture, increasing the operation success rate from 78% to 96%; For foldable screen devices (such as Samsung Galaxy Z Fold), the layout model is automatically switched according to the screen expansion status, and the touch mapping accuracy of split-screen operation is as high as 99%.

[0035] In this embodiment, the mapping method supports the integration of multi-source touch commands, including: Receive touch commands from multiple source devices and sort the commands by timestamp and event queue; Heterogeneous device access: Supports simultaneous connection to multiple source devices such as mobile phones, tablets, PCs, smart watches, etc., and receives touch commands through unified communication protocols (such as WebSocket or Bluetooth BLE); Standardize the format of input instructions: Convert touch parameters of different devices (such as the tilt angle of Apple Pencil and the rotation angle of Surface Dial) into standardized vector data; Unified time base: Synchronize the clocks of each device through the NTP protocol to ensure that the timestamp accuracy of cross-device commands reaches ±1ms.

[0036] Dynamic priority queue construction: Create a timestamp-based double-buffered event queue, divided into a high-priority queue (such as touch click, gesture start events) and a low-priority queue (such as gesture continuation events); Sort by the following rules: Continuous gesture events (such as long press and drag) on ​​the same device maintain temporal consistency; Cross-device commands are inserted into the queue in timestamp order. If timestamps overlap, they are sorted by weight based on device type (e.g., the weight of the conference speaker's device + 30%).

[0037] When multiple command conflicts are detected, the execution command is selected based on the user identity authority or operation priority; Conflict determination rule base: Define conflict scenario types and resolution strategies (e.g. Figure 7 shown); Hybrid trigger mechanism: When multiple commands can be executed in parallel (such as two devices controlling the video progress and volume respectively), start the command diversion module: By dividing the interface area, the touch commands are distributed to the target controls (for example, the left screen area command is mapped to the progress bar, and the right screen area command is mapped to the volume bar); Enables parallel thread processing for non-conflicting instructions, reducing overall latency.

[0038] Multi-dimensional permission model: Dynamically assign permission levels based on user role (speaker / participant), device type (mobile phone / VR controller) and biometrics (fingerprint / face recognition); In educational scenarios, the teacher's device automatically obtains forced takeover permissions to interrupt incorrect operation instructions on the student's device; Conflict Visual Feedback: Display the source device icon and operation intention of the conflicting command on the same screen interface (such as "Device A is trying to delete a file, and device B is editing it"); A manual arbitration interface is provided to support users to adjust arbitration results in real time through voice commands (such as "allow device A to execute").

[0039] This embodiment achieves the following technical breakthroughs by refining the multi-source instruction integration and conflict handling mechanism: Efficient multi-user collaborative operation Timestamp synchronization accuracy: The timing error of cross-device operations is controlled within 5ms, ensuring that the handwriting synchronization rate is increased to 99% when multiple users collaborate on drawings (such as Miro whiteboards); Hybrid trigger mechanism: In video editing scenarios, the director can adjust the timeline via iPad while the photographer can fine-tune the filter parameters via Android phone, increasing the efficiency of parallel execution of tasks by 40%; Intelligent conflict resolution skills Dynamic permission model: In medical consultation scenarios, the priority of the chief physician's device's annotation instructions for CT images is increased by 300%, and the interception rate of erroneous operations reaches 95%; Regional diversion strategy: In the stock trading system, multi-device instructions are mapped according to screen partitions (such as buy orders on the left and sell orders on the right), and the conflict rate is reduced by 80%; Enhance system robustness and user experience Double buffer queue design: Even if the number of instructions from a single device increases suddenly (such as a game controller sending 50 instructions per second), the system response delay remains stable below 20ms; Visual feedback: User satisfaction with conflict perception increased by 65%. Especially in the field of education, teachers can visually detect students' misoperation of devices and correct them immediately. Cross-platform compatibility extension Supports mixed input of smartwatch micro-gestures (such as Huawei Watch's fist sliding) and VR handle spatial operations. In the metaverse meeting scene, the multimodal command recognition accuracy rate reaches 92%; Through standardized protocol conversion, command coordination between old devices (such as Windows 7 touch screens) and modern devices is achieved, and the reuse rate of historical devices is increased by 50%.

[0040] In this embodiment, the dynamic mapping in S3 further includes: Monitor the interface status changes of the target device in real time. If the interface update causes the control identifier to become invalid, it triggers the re-parsing of the semantic layer instructions. Interface change detection mechanism: Adopting dual-thread monitoring architecture: Main thread: monitors interface element change events (such as control addition and deletion, property update) through the target device's system API (such as Android's AccessibilityService or iOS's UIAutomation); Auxiliary thread: capture screen images at regular intervals (frequency 10Hz) and detect changes in interface layout (such as pop-up windows and page jumps) through image difference algorithms.

[0041] When an interface update is detected, the control tree reconstruction is triggered and the validity of the identifiers of the affected controls is marked (for example, whether the button ID still exists).

[0042] Intelligent recovery of invalid controls: If the target control identifier is invalid (for example, the original "Confirm button" is replaced by a "Submit button"): Based on the control function similarity matching algorithm, search for semantically similar alternative controls in the control tree (such as "Confirm" → "Submit" based on text content or location proximity); If there is no matching control, the semantic backtracking mechanism is started: the expected operation target is inferred based on the user's historical operation records (such as the operation path after clicking "Confirm" for the last five times); The re-parsed intermediate semantic layer instructions are rechecked to ensure that the mapped native instructions conform to the current interface status of the target device.

[0043] When the target device is a virtual reality device, the touch commands are mapped to three-dimensional space interaction events, and the display content of the virtual scene and the physical screen are synchronized; Spatial interaction event generation: Coordinate system conversion: Convert the 2D touch coordinates (x, y) of the source device to the 3D space coordinates (x, y, z) of the VR device, where the z-axis depth value is dynamically calculated based on the touch pressure or gesture type (e.g., light touch z=0.5m, heavy press z=1.2m); Establish a spatial mapping relationship:

[0044] (where k is the scaling factor and offset is the screen alignment offset).

[0045] Gesture-action mapping library: Define the correspondence between touch gestures and 3D interactive actions in VR scenes (such as Figure 8 shown); Physical-virtual simultaneous rendering: The touch track of the device on the same screen (such as a translucent halo effect) is superimposed on the VR headset to ensure that the user perceives the real-time linkage between the touch operation and the virtual scene.

[0046] Machine learning optimized mapping strategy Dynamic parameter tuning: Collect user adjustment data on VR mapping results (such as manually correcting the rotation angle of objects) to build a training data set; Train the LSTM neural network model to learn the user's spatial operation preferences (such as z-axis depth sensitivity and rotation speed); When similar operating scenarios are detected, the model is automatically called to predict the optimal mapping parameters, reducing the number of manual calibrations.

[0047] This embodiment refines the implementation logic of dynamic mapping and VR space interaction to bring about the following technical improvements: Real-time responsiveness of interface updates Dual-thread monitoring mechanism: The interface change detection delay is shortened from 200ms in traditional solutions to 50ms. For example, in the e-commerce app flash sale scenario, the touch mapping success rate is increased to 99.8% after the purchase button is updated; Semantic backtracking mechanism: In the dynamic forms of financial apps, even if the control ID changes (such as "Buy" → "Trade"), the user's operation intention can be accurately restored at a rate of 92%.

[0048] Breakthrough in virtual reality interaction accuracy Dynamic calculation of spatial coordinates: In VR architectural design scenarios, the two-dimensional drawing trajectory of the tablet stylus is mapped to the three-dimensional model wireframe with an error of ≤0.5mm, meeting professional design requirements; Gesture-action mapping library: In VR educational applications, the accuracy of gesture operation recognition has been improved to 97% (for example, the false trigger rate of the "pouring liquid" gesture in a chemical experiment has been reduced from 15% to 2%).

[0049] Resource consumption optimization Intelligent recovery of invalid controls: reduces the number of complete semantic analysis times caused by interface updates, and reduces CPU peak load by 35% (from 60% → 39%); LSTM parameter tuning: The frequency of manual calibration of VR space mapping has been reduced by 70%, significantly reducing user learning costs.

[0050] Cross-dimensional operational consistency In mixed reality (MR) scenarios, mobile phone touch can simultaneously control the physical screen (such as adjusting PPT) and virtual holographic projection (such as rotating 3D models), with an operation synchronization error of <10ms; For foldable screen / scrollable screen devices, the interface layout changes are automatically adapted after the screen is unfolded, and the touch command mapping accuracy remains above 99.5%.

[0051] In this embodiment, the method also includes a conflict handling mechanism: When the touch protocols of the source device and the target device are incompatible, simulated touch events are generated, including virtual clicks, sliding track simulation, and time compensation for long press events; Protocol compatibility detection and conflict classification Protocol feature comparison: Extract the protocol characteristics of the source device touch command (such as Android's MotionEvent event structure and Windows' WM_TOUCH message format); Compare field by field with the protocol specifications supported by the target device to identify incompatible fields (e.g. the force pressure value field of iOS is missing in the Android protocol); Classification based on conflict severity: Level 1: Fields are missing but can be simulated (e.g. pressure value → long press duration); Level 2: event type has no equivalent mapping (such as 3D Touch Peek gesture); Level 3: Multiple event logic conflicts (such as two-finger rotation and single-finger drag on the target device).

[0052] Simulate touch event generation Virtual click event generation: When it is detected that the target device does not support the click event type of the source device (such as Windows right click corresponds to Android long press): Parsing click intent (e.g. "open context menu"); Generates an equivalent sequence of operations based on the target device control type: If the target is an Android list item, a simulated event chain of long press for 500ms → pop-up menu is generated; If the target is an iOS text box, the mapping logic of double-click → select all text is generated.

[0053] Sliding trajectory simulation: Dynamic interpolation of high-precision trajectories (such as Surface Pen handwriting): Downsample the original track points (sampling rate 120Hz) to the maximum frequency supported by the target device (such as 60Hz for Android devices); Optimize trajectory smoothness through Bezier curve fitting to ensure handwriting restoration error ≤ 0.3mm; Added inertial scrolling compensation (like iOS's UIScrollView inertial scrolling feature).

[0054] Time compensation mechanism: When the touch response delay of the target device is higher than the threshold (such as >80ms): Predict the processing queue status of the target device and send instructions in advance; Introducing dynamic timestamp offset correction (such as Android device clock deviation compensation); In the game scene, enable frame synchronization compensation (dynamically adjust the timing of command sending according to the refresh rate).

[0055] The execution results are returned to the source device through the feedback channel to optimize the semantic parsing rules of subsequent instructions; Feedback channel construction and rule optimization Multi-dimensional feedback data collection: The following execution results are captured in real time: Target device response status (success / failure / partial execution); The user manually corrects the action (such as clicking the target control again); System performance indicators (latency, CPU usage, touch coordinate offset).

[0056] Root cause analysis and strategy iteration: Establish a knowledge graph of conflict events, associating protocol conflict types, mapping schemes, and execution results; Use decision tree algorithms (such as CART) to analyze the causes of failure (e.g., 60% of Level 2 conflicts are caused by gestures not having equivalent mappings); Dynamically update the rule base: For frequent failure scenarios (such as 3D Touch Peek), add alternative mapping solutions (such as "hard press + side swipe" combination gesture); On low-performance devices, the simulation event accuracy is automatically downgraded (for example, the trajectory interpolation algorithm is switched from cubic Bezier to linear interpolation).

[0057] User behavior learning model: Collect user correction data for simulated events (such as adjusted sliding speed and long press duration); Train a random forest model to predict user preference mapping parameters; When the same type of conflict occurs again, the user's historical preference solution will be called first.

[0058] This embodiment achieves the following technical breakthroughs by refining the protocol conflict handling and feedback optimization mechanism: Cross-protocol compatibility is comprehensively improved Virtual event chain generation: Improves the mapping success rate of Windows right-click on Android devices from 68% to 94%, especially for complex applications (such as AutoCAD mobile); Track restoration optimization: The handwriting offset of Surface Pen on iPad Pro is reduced from 1.2mm to 0.4mm, meeting professional drawing requirements.

[0059] Dynamic environment adaptability Time compensation mechanism: In cloud gaming scenarios, the touch delay caused by 5G network jitter is stabilized from 150ms to less than 50ms; Downgrade strategy: The command packet loss rate of low-end devices (such as Android Go) is reduced by 70%, ensuring the smoothness of basic operations.

[0060] Enhanced user personalization experience Behavioral learning model: automatically optimizes the mapping relationship between pen pressure and line thickness according to the designer's habits, and the loading speed of personalized configuration is <100ms; Visual correction interface: users can manually adjust simulation event parameters (such as long press duration), and the correction data will be synchronously fed back to the rule base; System robustness is significantly improved Root cause analysis engine: The average resolution time for conflict events is shortened from 30 seconds to 3 seconds, and the operation and maintenance costs are reduced by 90%; Knowledge graph iteration: After the new device protocol is connected (such as HarmonyOS), the adaptive learning cycle is compressed from 2 weeks to 48 hours.

[0061] In this embodiment, the mapping method supports cross-platform protocol conversion, including: Convert mouse events of Windows system to MotionEvent events of Android system; Protocol structure analysis and mapping: Event type matches: Map Windows' WM_LBUTTONDOWN (left button pressed) to Android's ACTION_DOWN; WM_MOUSEMOVE (mouse movement) is split into consecutive ACTION_MOVE events, and the movement step size is calculated based on the target device screen density (e.g., each pixel movement on a 4K screen is mapped to 2 ACTION_MOVE events); WM_RBUTTONUP (right button release) is converted to the Android long press event ACTION_LONG_PRESS, with a default duration of 500ms, which can be adjusted dynamically.

[0062] Coordinate transformation algorithm: Calculate scaling based on screen resolution differences:

[0063] For high-precision devices (such as Surface Studio), enable the sub-pixel interpolation algorithm to improve trajectory smoothness.

[0064] Special handling of scroll wheel events: Convert Windows' WM_MOUSEWHEEL event to Android's MotionEvent.ACTION_SCROLL: Parse the scroll wheel delta value (Delta) and convert it proportionally to the scroll distance (e.g. Delta = 120 corresponds to scrolling 300dp); Dynamically adjust scroll inertia parameters (such as friction coefficient, maximum speed) according to the target control type (such as ListView / WebView).

[0065] Map the Force Touch gesture of iOS system to the long press event of Android system, and dynamically adjust the mapping parameters according to the pressure value; Dynamic quantitative classification of pressure values: Collect touch pressure data of iOS devices (0-6.666N range), divided into 5 pressure levels (such as Fig. 9 as shown).

[0066] Dynamically adjust parameters based on the hardware characteristics of the target device: If the Android device supports pressure-sensitive screen (such as Samsung Galaxy Note series), the pressure value is directly transmitted; If the device does not have pressure-sensitive function, pressure feedback is simulated through long press duration + vibration intensity (for example, Level 4 is mapped to 800ms long press + 3 levels of vibration).

[0067] 3D touch (Peek and Pop) gesture processing: Peek gesture recognition: detects when the pressure value reaches the threshold (3.0N) and lasts for more than 200ms, and maps it to the Android ContextMenu pop-up; Pop gesture conversion: Based on Peek, continue to increase the pressure to 5.0N, convert it to an ACTION_CLICK event, and trigger a depth jump (such as opening the link details page).

[0068] Dynamic parameter optimization engine Scenario Adaptation Strategy: Establish a device performance profile library to record parameters such as touch response delay and screen refresh rate of different devices; Dynamically adjust mapping rules: On low refresh rate devices (such as 60Hz screens), reduce the frequency of sending ACTION_MOVE events (from 120Hz → 60Hz); In the game scenario, enable "Turbo mode" to bypass the Android event queue and directly inject touch commands, compressing the delay from 80ms to 20ms.

[0069] User behavior learning module: Collect statistics on users’ modification behaviors on mapping results (such as manually extending the long press time); The linear regression model is used to predict personalized parameters (for example, user A prefers to map Level 3 pressure to a 700ms long press) and update the local rule base.

[0070] This embodiment achieves the following technical breakthroughs by refining the differentiated processing logic of cross-platform protocol conversion: Cross-system operation consistency is greatly improved Dynamic pressure-duration mapping: The accuracy of Force Touch intent restoration on Android devices has increased from 65% to 93%. For example, the success rate of re-pressing to preview attachments in the Mail app has increased to 89%. High-precision track retention: The handwriting jitter rate of the Windows stylus in Android drawing software (such as Krita) is reduced to 2%, meeting professional-level input requirements.

[0071] Enhanced adaptability to complex scenarios Game scene optimization: The right-click operation of the Windows mouse is mapped to the Android long press, and the response speed is shortened from 120ms to 35ms to meet the competitive needs of FPS mobile games; Folding screen adaptation: Dynamically adjust the coordinate mapping algorithm according to the folding state of the device (such as enabling dual-screen partition mapping when Surface Duo is unfolded), and the control click accuracy remains above 99%.

[0072] Resource consumption and compatibility optimization Sub-pixel interpolation algorithm: On 8K resolution devices, CPU usage dropped from 42% to 28%, and memory consumption was reduced by 35%; Support for old devices: The mouse protocol of Windows XP can be seamlessly converted to Android 14, reducing the cost of reusing historical devices by 70%.

[0073] User experience personalization Pressure feedback simulation: Through vibration intensity grading (such as Level 4 triggering 3 short vibrations), the user's confidence in "successful hard press" is increased by 80%; Behavioral learning model: The matching degree of pen pressure-line thickness mapping for designer users is improved to 98%, and the loading time of personalized configuration is less than 50ms.

[0074] In this embodiment, the dynamic adaptation rule base is optimized through machine learning: Collect users' correction operations on mapping results and generate training data sets; User behavior data collection and annotation Multimodal Data Capture: Real-time recording of the interaction data between the user and the device, including: Touch correction behavior: users manually adjust the mapped control selection (such as re-clicking the target button) and slide distance correction (such as dragging the progress bar to the correct position); Semantic conflict log: control identification failure events detected by the system and touch intention misjudgment cases (such as identifying "zoom" as "page turn"); Environmental context data: screen resolution of the target device, currently running application processes, and network latency status.

[0075] Clean and label the raw data: Remove invalid data (such as short clicks caused by accidental touches); Add semantic labels to each correction operation (e.g. "control positioning error", "pressure value mapping deviation").

[0076] Update the mapping strategy in the rule base based on the neural network model to improve the accuracy of semantic parsing; Hybrid model training and feature engineering Deep Feature Extraction: Control level graph neural network (GNN) modeling: Convert the target device interface control tree into a graph structure, where nodes are control attributes (type, position, text) and edges are parent-child hierarchical relationships; The GraphSAGE model is used to learn the embedding vectors of control nodes to capture their functional semantics (such as the association between the “shopping cart button” and the “checkout button”).

[0077] Touch timing feature encoding: The time series data of touch gestures (coordinates, pressure value changes) are input into the bidirectional LSTM network to extract the temporal features of the operation intention (such as the difference between "quick double-click" and "slow double-click").

[0078] Multi-task learning framework: Joint training of two tasks: Control matching task: predict the target control identity after the user's correction (classification task); Parameter optimization task: regression prediction of optimal mapping parameters (such as long press duration, sliding speed compensation value).

[0079] Loss function design:

[0080] Dynamic update of rule base and A / B testing Incremental update strategy: Real-time hot update: When the model confidence is >90%, the new rule is directly inserted into the rule base in memory (such as the new side-swipe back gesture mapping for iOS 17); Version control ensures atomic operations and avoids rule conflicts during updates.

[0081] Shadow mode verification: Run the old and new rule bases in parallel and compare the execution results (such as mapping success rate and latency indicators). The new rules will only take effect when the improvement is significant (such as success rate +5%).

[0082] Scenario-based rule grouping: Create independent rule subsets based on application scenarios (gaming, office, education), for example: Game scenario rule group: prioritize touch response speed and allow lower trajectory smoothness; Document Editing Rule Group: Emphasizes precise coordinate mapping, sacrificing some latency in exchange for high precision.

[0083] Edge computing optimization Distributed model reasoning: Deploy a lightweight inference engine (TensorFlow Lite) on user devices to process 90% of high-frequency requests in real time; Complex scenarios (such as VR space mapping) require high-performance cloud models (such as GPU-accelerated BERT) to protect user data through differential privacy technology.

[0084] This embodiment achieves the following technical breakthroughs through the deep integration of machine learning and rule base: Semantic parsing accuracy has increased significantly GNN control matching: The recognition accuracy of dynamically generated controls has been increased from 78% of traditional rules to 96%. For example, the success rate of clicking on the flash sale button of an e-commerce app has been increased to 99.3%. Multi-task learning optimization: The mapping error of the long press duration parameter is compressed from ±150ms to ±20ms, especially in iOS and Android cross-platform scenarios.

[0085] Adaptive capabilities have been greatly enhanced Incremental hot update: The adaptation cycle of new operating system versions (such as Android 14) is shortened from 2 weeks to 8 hours; Scenario-based rule group: The touch response delay in game scenarios is reduced by 40% (from 50ms to 30ms), and the combo success rate in MOBA games is increased by 25%.

[0086] Operation and maintenance costs are significantly reduced Shadow mode verification: The system crash rate caused by incorrect rule push was reduced to 0.03%, and the workload of the operation and maintenance team was reduced by 70%; Edge computing architecture: Cloud computing costs are reduced by 65%, and the risk of user privacy data leakage is reduced by 90%.

[0087] Intelligent upgrade of user experience Personalized mapping strategy: Automatically optimize the thickness and curve of the handwriting based on the designer's pressure-sensitive pen usage habits, improving the creation efficiency by 30%; Collaborative learning across devices: The mapping rules modified by users on iPad can be synchronized to Windows devices, increasing satisfaction with multi-device consistency by 80%.

[0088] In this embodiment, the mapping method supports multi-window collaborative control: When the target device displays multiple application windows in split screen, the command is mapped to the corresponding application process according to the window area to which the touch position belongs; Dynamic perception of split-screen status and hotspot modeling Multi-window topology scanning: Get the current active window list and hierarchical relationship through the system-level interface of the target device (such as Android's ActivityManager or Windows' Win32API); Construct a split-screen hot zone matrix and record the screen coordinate range of each window (e.g., window A occupies the left half of the screen [0,0]-[540,2400], and window B occupies the right half of the screen [540,0]-[1080,2400]); For irregular split screens (such as picture-in-picture and free-floating windows), image recognition technology is used to extract window outlines and establish a polygonal hot spot coordinate model.

[0089] Dynamic hotspot adjustment: When a window size change is detected (such as the user dragging the dividing line): Update the hot zone matrix in real time and calculate the overlap ratio of new and old hot zones; Compensate for the coordinate offset of the affected touch commands (for example, if the window width is reduced by 20%, the original touch point x coordinate is scaled by a factor of 0.8); When the foldable screen device is unfolded, adjacent hot areas are automatically merged into an extended operation plane.

[0090] Directed distribution of touch commands Process-level instruction routing: Match the target window hotspot according to the touch coordinates and extract the application process ID bound to the window (such as WeChat process ID 3327); Inject native instructions into the target process event queue through cross-process communication (Binder for Android / COM+ for Windows); For operations that require cross-window linkage (such as dragging a file from window A to window B), enable the dual-channel synchronization engine: Channel 1: Send ACTION_DOWN and ACTION_MOVE events to window A; Channel 2: When the touch point enters the hot area of ​​window B, an ACTION_DROP event is sent to window B and the data handle is passed.

[0091] Intelligent division of sub-interaction areas: Based on the semantic analysis of window content (OCR recognition + control type detection), complex windows are divided into functional sub-areas: Example: The video playback window is divided into the "control bar area" (pause / progress) and the "main screen area" (double-click full screen); Set up independent event filtering rules for each sub-area: Disable the text selection function of the long press gesture in the "Document Editing Subarea" to prevent accidental touches; Enable pinch-to-zoom and rotate gestures in the Chart Display Subarea.

[0092] When the display interface of the same screen device is a multi-window layout, cross-window command distribution is achieved by dividing the sub-interaction area; Enhanced cross-window collaboration Data flow pipeline construction: Establish a cross-application data sharing channel (such as Android's ContentProvider): When a drag operation is detected to cross the window boundary, the source window data (such as image URI) is automatically encapsulated as a system clipboard object; When the hot area of ​​the target window is released, data parsing and rendering are triggered (such as inserting a picture into a PPT).

[0093] Operational continuity assurance: Use a global state machine to track cross-window operation processes (such as drag status and copy progress); If the window is closed or switched in the middle, the intermediate state is automatically saved and a recovery entry is provided.

[0094] Multi-user collaboration mode: In the conference scenario, assign independent user control rights to each sub-window: User A controls the Excel window on the left with a mobile phone, and user B controls the PPT window on the right with a tablet. Limit the scope of cross-window operations through permission labels (such as "Read-only" and "Edit").

[0095] Conflict arbitration mechanism: When multiple users operate the same sub-area at the same time, automatic arbitration is performed based on identity priority (host > participant) or operation type (edit > view); Conflicting operations are visually marked (such as a red flashing border) to support manual intervention and adjustment.

[0096] This embodiment achieves the following technical improvements by refining the multi-window collaborative control logic: Multitasking efficiency doubled Hot zone dynamic compensation: When the folding screen is unfolded, the touch mapping accuracy of the split-screen candlestick chart and the trading window of the stock trading app remains at 99.8%, and the operation delay is less than 15ms; Cross-process drag optimization: The success rate of dragging links from the Chrome browser to the WeChat window increased from 75% to 98%, and the time consumed was shortened from 2.1s to 0.3s.

[0097] Breakthrough in Interaction Accuracy on Complex Interfaces Sub-area semantic division: In the Photoshop split-screen scenario, the false trigger rate of brush operations in the left canvas area and sliding and zooming in the right toolbar is reduced to 1.2%; Hierarchical authority control: In the medical consultation system, interns can only mark sub-areas of the image, while chief physicians have global editing rights, and the authority violation interception rate is 100%.

[0098] Optimize system resource consumption Dual-channel synchronization engine: The memory usage of multi-window linkage is reduced by 40% (from 220MB to 132MB), and the fluency is significantly improved especially on low-end devices; State machine management: The recovery loading time after cross-window operation interruption is compressed from 8s to 1.5s, and the risk of data loss is reduced by 90%.

[0099] Multi-user collaboration experience upgrade Independent control rights allocation: In online education scenarios, the delay difference between teachers and students operating different sub-windows simultaneously is less than 10ms, and the efficiency of collaborative answering is increased by 50%; Conflict visualization prompts: The speed of recognizing operational conflicts during team collaboration is increased by 3 times, and the time spent on arbitration decisions is reduced by 70%.

[0100] In this embodiment, the mapping method also includes a device switching mechanism: When the target device is detected to be switched, the context state of the current intermediate semantic layer instruction is retained; Device switching event triggering and status snapshot Seamless switching detection: The device connection manager monitors the target device status in real time (such as Bluetooth signal strength and Wi-Fi Direct connection stability) and triggers the switch when the following events are detected: Active switching: the user manually selects a new target device (e.g. switching from a projector to a VR headset); Passive switching: The original target device is disconnected and times out (for example, the network delay is > 500ms for 3 seconds).

[0101] Contextual state capture: Extract the full context of the current intermediate semantic layer instruction, including: Operation chain: a sequence of touch commands that have not been completed (such as long-pressing and dragging and releasing in the middle); Interface snapshot: the last valid interface screenshot and control tree structure of the target device; Dynamic variables: temporarily generated control identifiers and mapping parameter adjustment records.

[0102] The incremental serialization technology is used to compress the state data and store it in the non-volatile memory of the same-screen device (such as eMMC flash memory).

[0103] Cross-protocol state migration and adaptation Protocol Difference Matrix Matching: Construct a protocol compatibility matrix table to record the protocol differences between the new and old target devices (such as Fig.10 shown); Dynamic parameter correction: Runtime adaptation of device hardware-dependent parameters (such as screen refresh rate): ; If the difference between the control tree structures of the new and old devices is greater than 30%, the semantic backtracking engine is started: the target control is relocated based on the OCR result of the interface snapshot (such as the "Save button").

[0104] Remap commands and restore continuity of operation flow based on the platform protocol of the new target device; Operation chain recovery and continuity assurance Interrupted operation continuation: Parse the stored intermediate semantic layer instructions and rebuild the native instruction queue according to the new target device protocol: Example: Convert an unfinished sequence of ACTION_MOVE on Android to a continuous UIPanGestureRecognizer event on iOS; Enable timeline compression for time-sensitive operations (such as game combos): accelerate the execution of unfinished instructions to the current timestamp.

[0105] Status consistency check: Compare the interface snapshot with the new device's current interface (such as control position offset). If the difference exceeds the threshold (such as >10% pixel deviation): Trigger progressive calibration: gradually align the target control by fine-tuning the touch coordinates (Δx, Δy); If calibration fails, a user guidance interface is launched (with expected control locations highlighted).

[0106] Cross-device operation history synchronization: Synchronize the operation records before switching (such as the last 5 clicks) to the new device for preheating the machine learning model; In distributed device groups (such as smart home multi-screen), operation history roaming is supported, and users can continue unfinished tasks on any device.

[0107] This embodiment achieves the following technical breakthroughs by refining the context migration and protocol adaptation logic of device switching: Seamless switching experience: Incremental serialization compression: The switching time is shortened from 30 seconds in the traditional solution to 0.5 seconds, and the state data volume is reduced by 85% (from 2MB to 300KB); Timeline Compression: The success rate of unfinished combos in game scenarios is increased from 55% to 98%, and the skill release delay is <20ms.

[0108] Accurate cross-protocol adaptation Coordinate mirror flipping: After switching from Android to iOS, the touch coordinate error is reduced from ±15 pixels to ±2 pixels; Semantic backtracking engine: After the interface is revised (such as WeChat 8.0→9.0), the positioning accuracy of core functional controls (such as the "Send button") remains at 99.5%.

[0109] Enhanced robustness in complex scenarios Progressive calibration: When the folding screen is switched between unfolded and folded states, the control drift automatic correction success rate reaches 93%; Operation history roaming: The efficiency of task continuation across device groups (mobile phone → car machine → smart TV) is increased by 70%.

[0110] Resource optimization and compatibility expansion Non-volatile storage: The state recovery time after power failure is shortened from 10 seconds to 1 second, and the risk of data loss is close to zero; Protocol Difference Matrix: supports rapid access to emerging systems such as HarmonyOS and Fuchsia, and shortens the adaptation period from 3 months to 2 weeks.

[0111] In the description of this specification, the description with reference to the terms "one embodiment", "some embodiments", "example", "specific example", or "some examples" etc. means that the specific features, structures, materials or characteristics described in conjunction with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described may be combined in any one or more embodiments or examples in a suitable manner. In addition, those skilled in the art may combine and combine the different embodiments or examples described in this specification and the features of the different embodiments or examples, without contradiction.

[0112] The above description is only a preferred embodiment of the present invention, and does not limit the patent scope of the present invention. All equivalent structural changes made by using the contents of the present invention specification and drawings under the inventive concept of the present invention, or directly / indirectly applied in other related technical fields are included in the patent protection scope of the present invention.

Claims

1. A semantic parsing and mapping method for cross-platform touch commands of a same-screen device, characterized in that: include: S1: Acquire a touch command of a source device through a screen sharing device, where the touch command includes touch position, touch gesture type and touch timing information; S2: Based on the operating system type and application scenario of the target device, semantically parse the touch control instruction to generate an intermediate semantic layer instruction, where the intermediate semantic layer instruction includes operation logic, interaction intent, and target control identifier; S3: dynamically mapping the intermediate semantic layer instructions to native instructions of the target device according to the platform protocol and interaction rules of the target device, wherein the native instructions are adapted to the touch event processing mechanism of the target device; S4: Send the native command to the target device for execution, and synchronously update the display interface of the same-screen device and the display content of the target device.

2. The semantic parsing and mapping method of cross-platform touch control instructions of the same-screen device according to claim 1 is characterized in that: The semantic analysis in S2 further includes: Matching the interaction logic of the target device through a dynamic adaptation rule base, wherein the rule base includes operating system protocols, application interface control types, and touch event priorities of different platforms; When it is detected that the control type of the target device is a dynamically generated control, the intermediate semantic layer instructions are optimized based on the control hierarchy and context semantics.

3. The semantic parsing and mapping method of cross-platform touch control instructions of the same-screen device according to claim 1, characterized in that: The step S2 also includes multi-level semantic analysis: The first level analyzes the physical parameters of touch gestures, including sliding direction, pressure value, and multi-touch coordinates; The second level combines the current display interface layout of the target device to identify the functional modules corresponding to the touch operation; The third level dynamically adjusts the semantic weights according to the application scenarios, giving priority to mapping high-frequency operation instructions.

4. The semantic parsing and mapping method of cross-platform touch control instructions of the same-screen device according to claim 1, characterized in that: The mapping method supports the integration of multi-source touch commands, including: Receive touch commands from multiple source devices and sort the commands by timestamp and event queue; When multiple command conflicts are detected, the execution command is selected based on user identity permissions or operation priority.

5. The semantic parsing and mapping method of cross-platform touch control instructions of the same-screen device according to claim 1, characterized in that: The dynamic mapping in S3 further includes: Monitor the interface status changes of the target device in real time. If the interface update causes the control identifier to become invalid, it triggers the re-parsing of the semantic layer instructions. When the target device is a virtual reality device, the touch commands are mapped to three-dimensional space interaction events, and the display content of the virtual scene and the physical screen is synchronized.

6. The semantic parsing and mapping method of cross-platform touch control instructions of the same-screen device according to claim 1, characterized in that: The method also includes a conflict handling mechanism: When the touch protocols of the source device and the target device are incompatible, simulated touch events are generated, including virtual clicks, sliding track simulation, and time compensation for long press events; The execution results are returned to the source device through the feedback channel to optimize the semantic parsing rules of subsequent instructions.

7. The semantic parsing and mapping method of cross-platform touch control instructions of the same-screen device according to claim 1, characterized in that: The mapping method supports cross-platform protocol conversion, including: Convert mouse events of Windows system to MotionEvent events of Android system; Map the Force Touch gesture of the iOS system to a long press event of the Android system, and dynamically adjust the mapping parameters according to the pressure value.

8. The semantic parsing and mapping method of cross-platform touch control instructions of the same-screen device according to claim 1, characterized in that: The dynamic adaptation rule base is optimized through machine learning: Collect users' correction operations on mapping results and generate training data sets; Based on the neural network model, the mapping strategy in the rule base is updated to improve the accuracy of semantic parsing.

9. The semantic parsing and mapping method of cross-platform touch control instructions of the same-screen device according to claim 1, characterized in that: The mapping method supports multi-window collaborative control: When the target device displays multiple application windows in split screen, the command is mapped to the corresponding application process according to the window area to which the touch position belongs; When the display interface of the same screen device is a multi-window layout, cross-window command distribution is achieved by dividing the sub-interaction area.

10. The semantic parsing and mapping method of cross-platform touch control instructions of the same-screen device according to claim 1, characterized in that: The mapping method also includes a device switching mechanism: When the target device is detected to be switched, the context state of the current intermediate semantic layer instruction is retained; Remaps commands and restores continuity of operation flow based on the platform protocol of the new target device.

Citation Information

Patent Citations

  • Mobile terminal, display device and media asset picture control method

    CN118338056A

  • Instruction conversion method of cross-chip platform

    CN119292671A

  • Multi-platform fusion interaction method and system based on large touch screen

    CN119556840A

  • Optimization schemes for controlling user interfaces through gesture or touch

    US20150012815A1

  • Mapping touchscreen gestures to ergonomic controls across application scenes

    US20150202533A1

Cited By

  • Real-time interactive image generation system based on multi-point touch canvas

    CN120743141A

  • Real-time interactive image generation system based on multi-touch canvas

    CN120743141B

  • Cross-terminal quick screen locking synchronization method and system in multi-device cooperation scene

    CN120750990A

  • A cross-terminal quick screen locking synchronization method and system in a multi-device cooperative scenario

    CN120750990B

  • Cross-brand robot rapid access and control method based on RUAPL

    CN121262262A