Vision-based cross-network interaction method and system
By jointly encoding the spatial morphological features and temporal evolution features of visual interaction sequences in a heterogeneous network environment, a multimodal feature representation is constructed, and cross-domain semantic alignment and security classification are performed. This solves the problems of interaction accuracy and security in heterogeneous network environments and achieves efficient, flexible and reliable cross-network control.
Patent Information
- Application Number
- CN202511736928.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-11-25
- Publication Date
- 2026-02-24
- Estimated Expiration
- 2045-11-25
AI Technical Summary
Existing technologies struggle to effectively capture the rich semantic information and temporal correlations in an operator's intent within a heterogeneous network environment. Furthermore, the lack of cross-domain semantic alignment mechanisms and security classification management results in insufficient interaction accuracy and inadequate security assurance.
By jointly encoding the spatial morphological features and temporal evolution features of visual interaction sequences, a multimodal feature representation is constructed, and cross-domain semantic alignment is performed. A multi-level mapping relationship between the visual semantic space and the target network control space is established, and a security classification and temporary trust credential mechanism are implemented. The cross-domain semantic alignment is optimized using a bidirectional mapping relationship graph.
It enables efficient, flexible and secure cross-network control in heterogeneous network environments, improves interaction accuracy and system adaptability, and ensures operational reliability and stability.
Smart Images

Figure CN121209719B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to intelligent interaction technology, and more particularly to a vision-based cross-network interaction method and system. Background Technology
[0002] With the development of network technology, the demand for interaction between different network environments is increasing. Traditional cross-network interaction usually relies on physical connections or data conversion of specific protocols, which face many limitations in secure and isolated heterogeneous network environments. Vision-based cross-network interaction technology, as an emerging solution, enables information transmission and control in physically isolated environments by capturing, parsing, and converting visual signals.
[0003] Visual interaction technology has been initially applied in fields such as industrial control, safety supervision, and operation in special environments. These systems typically use computer vision technology to recognize visual commands such as operator gestures and movements, parse them into control signals using specific algorithms, and then execute corresponding operations on the target network. With the development of deep learning and multimodal fusion technologies, the recognition accuracy and semantic understanding capabilities of visual interaction have been significantly improved, providing a new technical path for cross-network interaction.
[0004] Existing technologies for encoding visual interaction sequences are often limited to a single feature dimension, making it difficult to effectively capture the rich semantic information and temporal correlation contained in the operator's intentions, resulting in insufficient interaction accuracy in complex operation scenarios.
[0005] Significant semantic and control pattern differences exist between heterogeneous network environments. Existing technologies lack effective cross-domain semantic alignment mechanisms, making it difficult to establish a precise mapping relationship between the visual semantic space and the target network control space, resulting in instruction parsing deviations and execution errors.
[0006] Existing cross-network interaction solutions do not adequately consider security isolation and trust verification, lack dynamic trust mechanisms based on behavioral patterns and security-level management of execution permissions, making it difficult to obtain sufficient security guarantees in high-security network environments and limiting the application scope of the technology. Summary of the Invention
[0007] The embodiments of the present invention provide a vision-based cross-network interaction method and system that can solve the problems in the prior art.
[0008] A first aspect of this invention provides a vision-based cross-network interaction method, comprising:
[0009] In the source network environment, the operator's visual interaction sequence is acquired through an image acquisition device. The spatial morphological features and temporal evolution features in the visual interaction sequence are jointly encoded to construct a multimodal feature representation.
[0010] Based on the heterogeneity of the source network environment and the target network environment, cross-domain semantic alignment is performed on the multimodal feature representation. By establishing a multi-level mapping relationship between the visual semantic space and the target network control space, the multimodal feature representation is deconstructed into a hierarchical control instruction set.
[0011] The instructions are classified for security according to the execution priority of the hierarchical control instruction set. The hierarchical control instruction set is transmitted to the target network environment through a non-physical connection transmission medium. Temporary trust credentials are constructed at the security domain boundary node based on behavior pattern verification.
[0012] In the target network environment, the execution conditions are determined and control operations are implemented based on the temporary trust credential. At the same time, state semantic features containing operation response timing and resource occupancy status are extracted and encoded into abstract state descriptors.
[0013] The abstract state descriptor is converted into a visual representation perceptible to the source network environment through reverse semantic mapping. After spatiotemporal alignment with the multimodal feature representation, a bidirectional mapping relationship graph of operation intention and execution result is constructed. The semantic deviation evolution law captured by the bidirectional mapping relationship graph is used to incrementally optimize the mapping parameters of cross-domain semantic alignment.
[0014] The spatial morphological features and temporal evolution features in the visual interaction sequence are jointly encoded to construct a multimodal feature representation, including:
[0015] Spatial domain analysis is performed on each frame of the visual interaction sequence to extract spatial morphological feature vectors containing the topological structure of the operator's limb posture, the fine movement pattern of the hand, and the orientation angle of the head; temporal correlation analysis is performed on the spatial morphological feature vectors corresponding to multiple consecutive frames to capture the change trajectory and evolution pattern of the spatial morphological feature vectors in the time dimension, and a temporal evolution feature sequence is constructed.
[0016] The spatial morphological feature vector is semantically aligned with the temporal evolution feature sequence. By establishing the association constraint between the spatial static state and the temporal dynamic process, a multimodal feature representation that integrates spatial semantics and temporal semantics is generated.
[0017] Based on the heterogeneity between the source and target network environments, cross-domain semantic alignment is performed on the multimodal feature representations. By establishing a multi-level mapping relationship between the visual semantic space and the target network control space, the multimodal feature representations are deconstructed into a hierarchical control instruction set, including:
[0018] Obtain the visual interaction capability description and control execution capability description of the source network environment, analyze the differences between the two in terms of operation granularity, response latency constraints and security policy restrictions, and construct a heterogeneous characteristic model;
[0019] Based on the heterogeneous characteristic model, the visual semantic space is divided into multiple semantic levels. Each semantic level corresponds to a control capability with a different level of abstraction in the target network control space. A mapping function is established between each semantic level and the corresponding control capability. According to the mapping function, the operational intent semantics in the multimodal feature representation is decomposed into control semantic units that match the execution capability of the target network.
[0020] Based on the multi-level mapping relationship, the control semantic units are organized into a hierarchical control instruction set according to execution dependencies and resource scheduling priorities.
[0021] The instructions are classified for security based on their execution priority within the hierarchical control instruction set. The hierarchical control instruction set is then transmitted to the target network environment via a non-physical connection transmission medium. Temporary trust credentials are constructed at the security domain boundary node based on behavioral pattern verification, including:
[0022] The execution priority sequence of the hierarchical control instruction set is analyzed and modeled in depth. The expected intensity of resource consumption and sensitivity to state changes of each instruction are dynamically evaluated through machine learning algorithms. The instruction risk score is calculated based on the expected intensity of resource consumption and the sensitivity to state changes. The instructions are divided into different security access levels according to the instruction risk score.
[0023] Differentiated transmission encapsulation structures are constructed for instructions at different security access levels. The differentiated transmission encapsulation structures include adaptive redundancy check codes. The layered control instruction sets encapsulated in differentiated manner are transmitted to the target network environment through a non-physical connection transmission medium. A dynamic timing feature library for instruction transmission is established based on the non-physical connection transmission medium.
[0024] The received instruction sequence is parsed at the security domain boundary node of the target network environment. Based on the dynamic temporal feature library, the verification result of the adaptive redundancy check code and the logical dependency relationship between instructions are used to generate a behavior pattern feature vector. The behavior pattern feature vector is then adaptively matched with a pre-established historical normal behavior benchmark library. The deviation metric between the behavior pattern feature vector and each benchmark pattern in the historical normal behavior benchmark library is calculated.
[0025] When the minimum value of the deviation metric is lower than the trust threshold, dynamic optimization modeling is performed based on the permission configuration template in the corresponding benchmark mode, and a temporary trust credential is generated by combining the adaptive redundancy check code verification result in the current instruction sequence with the instruction risk score.
[0026] The expected resource consumption intensity and state change sensitivity of each instruction are dynamically evaluated using machine learning algorithms. Based on the expected resource consumption intensity and the state change sensitivity, an instruction risk score is calculated, including:
[0027] An instruction behavior feature vector is constructed from the historical execution records of the hierarchical control instruction set. The instruction behavior feature vector contains the actual resource consumption data and the state change impact data of each instruction. The actual resource consumption data and the state change impact data are combined into a training sample set.
[0028] A deep neural network model is constructed based on the training sample set. The deep neural network model learns the correlation between the actual resource occupancy data and the state change impact data to obtain the resource occupancy assessment weight and the state change assessment weight.
[0029] A multi-layer feature mapping matrix is established for the instruction to be evaluated in the hierarchical control instruction set. The multi-layer feature mapping matrix is input into the deep neural network model. The expected intensity of resource consumption of the instruction is calculated using the resource consumption evaluation weight, and the sensitivity of state change of the instruction is calculated using the state change evaluation weight.
[0030] The expected intensity of resource consumption and the sensitivity to state changes are input into the deep neural network model to generate an instruction risk score.
[0031] In the target network environment, the execution conditions are determined and control operations are implemented based on the temporary trust credential. Simultaneously, state semantic features containing operation response timing and resource occupancy status are extracted, and these state semantic features are encoded into an abstract state descriptor, including:
[0032] The temporary trust credential is parsed in the execution engine of the target network environment to determine the execution conditions; the execution of control instructions is triggered based on the execution conditions; a dual-stream feature network is constructed during the instruction execution process; and parallel feature extraction is performed on the operation response timing and resource occupancy status based on the dual-stream feature network.
[0033] State semantic features are generated based on the features output by the dual-stream feature network, and the state semantic features are converted into abstract state descriptors.
[0034] The abstract state descriptor is converted into a visual representation perceptible to the source network environment through inverse semantic mapping. After spatiotemporal alignment with the multimodal feature representation, a bidirectional mapping relationship graph between operation intent and execution result is constructed, including:
[0035] The abstract state descriptor is hierarchically decomposed, and the decomposed information is organized into state attribute identifiers and state value descriptions. The state attribute identifiers and state value descriptions are used as input data for reverse mapping.
[0036] Based on the state attribute identifier and the state value description, a reverse semantic mapping transformation is performed to map the state attribute identifier to the corresponding visual element type label in the source network environment, and the state value description to the corresponding visual element appearance feature parameter in the source network environment. A perceptible visual representation of the source network environment is generated through the visual element type label and the visual element appearance feature parameter.
[0037] Based on the visual representation combined with the operation time marker and operation area coordinates in the multimodal feature representation, bidirectional alignment processing is performed. The spatial position coordinates in the visual representation are spatially aligned with the operation area coordinates, and the execution completion time corresponding to the abstract state descriptor is temporally aligned with the operation time marker, generating spatiotemporally aligned visual representation and spatiotemporally aligned operation feature pairing data.
[0038] A bidirectional mapping relationship map between operation intention and execution result is established by pairing the spatiotemporally aligned visual representation with the spatiotemporally aligned operation feature data.
[0039] A second aspect of the present invention provides a vision-based cross-network interaction system, comprising:
[0040] The first unit is used to acquire the operator's visual interaction sequence through an image acquisition device in the source network environment, and to jointly encode the spatial morphological features and temporal evolution features in the visual interaction sequence to construct a multimodal feature representation.
[0041] The second unit is used to perform cross-domain semantic alignment of the multimodal feature representation based on the heterogeneity of the source network environment and the target network environment. By establishing a multi-level mapping relationship between the visual semantic space and the target network control space, the multimodal feature representation is deconstructed into a hierarchical control instruction set.
[0042] The third unit is used to classify the instructions for security according to the execution priority of the hierarchical control instruction set, transmit the hierarchical control instruction set to the target network environment through a non-physical connection transmission medium, and construct temporary trust credentials based on behavior pattern verification at the security domain boundary node.
[0043] The fourth unit is used to determine the execution conditions and implement control operations in the target network environment based on the temporary trust credential, and at the same time extract the state semantic features containing the operation response timing and resource occupancy status, and encode the state semantic features into an abstract state descriptor.
[0044] The fifth unit is used to convert the abstract state descriptor into a visual representation that can be perceived by the source network environment through reverse semantic mapping, and after spatiotemporally aligning it with the multimodal feature representation, construct a bidirectional mapping relationship graph between operation intention and execution result, and use the semantic deviation evolution law captured by the bidirectional mapping relationship graph to perform incremental optimization of the mapping parameters of cross-domain semantic alignment.
[0045] A third aspect of the present invention provides an electronic device, comprising:
[0046] processor;
[0047] Memory used to store processor-executable instructions;
[0048] The processor is configured to invoke instructions stored in the memory to execute the aforementioned method.
[0049] A fourth aspect of the present invention provides a computer-readable storage medium having stored thereon computer program instructions that, when executed by a processor, implement the aforementioned method.
[0050] The beneficial effects of this application are as follows:
[0051] The vision-based cross-network interaction method provided by this invention achieves efficient representation of multimodal features by jointly encoding the spatial morphological features and temporal evolution features of the visual interaction sequence. This enables operators to achieve cross-network control in a natural and intuitive visual interaction manner, reducing operational complexity and improving the human-computer interaction experience.
[0052] This invention establishes a multi-level mapping relationship between the visual semantic space and the target network control space. By using cross-domain semantic alignment technology and constructing a hierarchical control instruction set, it solves the semantic gap problem between heterogeneous network environments, enhances the system's adaptability and compatibility, and makes the interaction between different network environments more flexible and efficient.
[0053] This invention achieves effective control under non-physical connection conditions while ensuring interaction security through a security classification and temporary trust credential mechanism. It also captures the semantic deviation evolution law by using a bidirectional mapping relationship graph and incrementally optimizes the mapping parameters for cross-domain semantic alignment, thereby improving the system's adaptability and operational stability and providing reliable protection for network interaction in isolated environments. Attached Figure Description
[0054] Figure 1 This is a flowchart illustrating the vision-based cross-network interaction method according to an embodiment of the present invention.
[0055] Figure 2This is a flowchart illustrating the security control operation and status assessment based on temporary trust credentials in an embodiment of the present invention. Detailed Implementation
[0056] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0057] The technical solution of the present invention will be described in detail below with reference to specific embodiments. These specific embodiments can be combined with each other, and the same or similar concepts or processes may not be described again in some embodiments.
[0058] Figure 1 This is a flowchart illustrating a vision-based cross-network interaction method according to an embodiment of the present invention, as shown below. Figure 1 As shown, the method includes:
[0059] In the source network environment, the operator's visual interaction sequence is acquired through an image acquisition device. The spatial morphological features and temporal evolution features in the visual interaction sequence are jointly encoded to construct a multimodal feature representation.
[0060] Based on the heterogeneity of the source network environment and the target network environment, cross-domain semantic alignment is performed on the multimodal feature representation. By establishing a multi-level mapping relationship between the visual semantic space and the target network control space, the multimodal feature representation is deconstructed into a hierarchical control instruction set.
[0061] The instructions are classified for security according to the execution priority of the hierarchical control instruction set. The hierarchical control instruction set is transmitted to the target network environment through a non-physical connection transmission medium. Temporary trust credentials are constructed at the security domain boundary node based on behavior pattern verification.
[0062] In the target network environment, the execution conditions are determined and control operations are implemented based on the temporary trust credential. At the same time, state semantic features containing operation response timing and resource occupancy status are extracted and encoded into abstract state descriptors.
[0063] The abstract state descriptor is converted into a visual representation perceptible to the source network environment through reverse semantic mapping. After spatiotemporal alignment with the multimodal feature representation, a bidirectional mapping relationship graph of operation intention and execution result is constructed. The semantic deviation evolution law captured by the bidirectional mapping relationship graph is used to incrementally optimize the mapping parameters of cross-domain semantic alignment.
[0064] In one optional implementation, the spatial morphological features and temporal evolution features in the visual interaction sequence are jointly encoded to construct a multimodal feature representation, including:
[0065] Spatial domain analysis is performed on each frame of the visual interaction sequence to extract spatial morphological feature vectors containing the topological structure of the operator's limb posture, the fine movement pattern of the hand, and the orientation angle of the head; temporal correlation analysis is performed on the spatial morphological feature vectors corresponding to multiple consecutive frames to capture the change trajectory and evolution pattern of the spatial morphological feature vectors in the time dimension, and a temporal evolution feature sequence is constructed.
[0066] The spatial morphological feature vector is semantically aligned with the temporal evolution feature sequence. By establishing the association constraint between the spatial static state and the temporal dynamic process, a multimodal feature representation that integrates spatial semantics and temporal semantics is generated.
[0067] The system acquires a visual interaction sequence consisting of consecutive image frames, recording the interaction between the operator and the environment or device. For each frame in the sequence, the system performs spatial domain analysis to extract spatial morphological feature vectors. Specifically, the system uses human keypoint detection technology to locate the operator's main joints, including 17 keypoints such as the shoulder, elbow, wrist, hip, knee, and ankle. The system uses an HRNet-based keypoint detection model, which achieves a detection accuracy of 76.3% AP on the COCO dataset. For each detected keypoint, the system records its two-dimensional position (x, y) in the image coordinate system, forming a 34-dimensional vector representing the operator's overall limb posture.
[0068] To capture the topological structure of the operator's limb posture, a connection matrix between key points is established based on the skeletal connections of the human body. This matrix describes the spatial topological relationships between body parts such as "left shoulder-left elbow" and "right hip-right knee". The system calculates the relative distances and angles between adjacent key points, generating a 28-dimensional topological feature vector to represent the spatial configuration features of the operator's limbs.
[0069] For the extraction of fine hand movements, a specialized hand keypoint detection model was used to locate 21 key points on the hand, including the center of the palm and the knuckles. The system employs the MediaPipe Hand model, achieving a detection accuracy of 95.7% on a custom gesture dataset. For each detected hand keypoint, the system records its two-dimensional coordinates and calculates the relative position and bending angle between adjacent keypoints, paying particular attention to the bending state of the finger joints, generating an 84-dimensional hand morphology feature vector. This vector effectively describes the flexion and extension states of the fingers, the rotation angle of the palm, and the gesture type.
[0070] To analyze the operator's attention direction, the head orientation angle was also detected, and five key points on the face (both eyes, the tip of the nose, and the corners of the mouth) were located. Then, the pitch, yaw, and roll angles of the head were estimated to form a 3D head orientation feature vector. The average angle error of the head orientation detection module was controlled within 5.2 degrees.
[0071] The extracted limb posture topological features (28-dimensional), fine hand movement morphological features (84-dimensional), and head orientation angle features (3-dimensional) are concatenated to form a 115-dimensional comprehensive spatial morphological feature vector, which is then analyzed using S... t This represents the spatial morphological feature vector at time t.
[0072] After extracting the spatial morphological features of a single frame image, temporal correlation analysis is performed on the spatial morphological feature vectors corresponding to multiple consecutive frames. The system selects a sliding window with a length of 16 frames (approximately 0.5 seconds of video length) and performs temporal correlation analysis on the spatial morphological feature vector sequence {S} within the window. t-15 , S t-14 , ..., S t The data is processed to capture the trajectory and evolution pattern of feature changes over time.
[0073] Calculate the difference values of each dimension of features between adjacent frames to obtain the velocity features. For the i-th dimension of features at time t, its velocity feature is calculated as V. t ,i = S t ,i - S t-1 The system further calculates the difference values of the velocity characteristics to obtain the acceleration characteristic A. t i = V t ,i - V t-1 These differential features can effectively describe the speed and acceleration patterns of the operator's movements.
[0074] The system analyzes the overall trend of features within the window, calculates the mean, variance, maximum, minimum, and number of peaks for each feature dimension, and constructs a statistical feature vector. For key action features, the system additionally calculates their autocorrelation coefficients to detect periodic action patterns. These temporal statistical features, together with the difference features, constitute a temporal evolution feature sequence, denoted by T. t This represents the temporal evolution characteristics at time t.
[0075] The spatial morphological feature vector S t With time-series evolution feature sequence T tSemantic-level feature alignment and fusion were performed. A feature fusion network was designed, consisting of two branches: a spatial feature processing branch and a temporal feature processing branch. Both branches map the original features to a latent space of the same dimension (256) using a multilayer perceptron. The spatial feature processing branch uses three fully connected layers with 128, 192, and 256 neurons respectively; the temporal feature processing branch uses four fully connected layers with 256, 256, 256, and 256 neurons respectively.
[0076] In the latent space, an attention mechanism is used to establish the association constraints between the spatial static state and the temporal dynamic process. Specifically, the system calculates the attention weights between the spatial feature vector and the temporal feature vector, and generates a comprehensive feature representation through weighted summation. The calculation of the attention weights is based on the similarity between the two feature vectors, using the cosine similarity calculation method. In this way, the system can adaptively adjust the importance of spatial and temporal features in different interaction scenarios.
[0077] A two-layer fully connected network (with 512 and 1024 neurons respectively) is used to further map the fused features into a final multimodal feature representation with a dimension of 1024. This multimodal feature representation includes both the static spatial information of the operator's current state and the temporal dynamic information of action changes, which can comprehensively describe the semantic content of the visual interaction sequence and provide effective feature support for subsequent tasks such as interaction intent recognition and behavior prediction.
[0078] In practical applications, this method achieved an accuracy rate of 93.7% in action intent recognition in intelligent interactive systems, which is 8.5% and 6.2% higher than methods that only use spatial features or temporal features, respectively, verifying the effectiveness of joint encoding of spatial morphological features and temporal evolution features.
[0079] In one optional implementation, based on the heterogeneity of the source network environment and the target network environment, cross-domain semantic alignment is performed on the multimodal feature representation. By establishing a multi-level mapping relationship between the visual semantic space and the target network control space, the multimodal feature representation is deconstructed into a hierarchical control instruction set, including:
[0080] Obtain the visual interaction capability description and control execution capability description of the source network environment, analyze the differences between the two in terms of operation granularity, response latency constraints and security policy restrictions, and construct a heterogeneous characteristic model;
[0081] Based on the heterogeneous characteristic model, the visual semantic space is divided into multiple semantic levels. Each semantic level corresponds to a control capability with a different level of abstraction in the target network control space. A mapping function is established between each semantic level and the corresponding control capability. According to the mapping function, the operational intent semantics in the multimodal feature representation is decomposed into control semantic units that match the execution capability of the target network.
[0082] Based on the multi-level mapping relationship, the control semantic units are organized into a hierarchical control instruction set according to execution dependencies and resource scheduling priorities.
[0083] The heterogeneous characteristic model is constructed by simultaneously acquiring descriptions of the visual interaction capabilities of the source network environment and the control execution capabilities of the target network environment. The visual interaction capability description module collects hardware parameters from the source network, such as resolution range, frame rate limit, color space support, and latency characteristics of the image acquisition devices. It also records software capability indicators such as gesture recognition accuracy, eye tracking range, voice command vocabulary, and multimodal fusion latency. The control execution capability description module acquires key parameters from the target network, such as device response capabilities, instruction execution granularity, concurrency limit, and resource scheduling strategies, through interface probing, protocol analysis, and performance benchmarking. The operation granularity difference analysis compares the continuous characteristics of the source network's visual interaction with the granularity of the target network's discrete control commands, calculating a granularity mapping ratio factor. When the sampling interval of the source network's gesture trajectory is ten milliseconds while the minimum execution interval of the target network's control commands is one hundred milliseconds, the granularity mapping ratio is set to ten to one. The response latency constraint difference is addressed by measuring the end-to-end latency of the source network's visual processing link and the target network's control command execution latency to establish a latency compensation model. When the source network's visual processing latency is 50 milliseconds and the target network's command execution latency is 200 milliseconds, the system sets an overall latency tolerance of 300 milliseconds and reserves a 50-millisecond buffer. The security policy restriction difference analysis covers dimensions such as access control levels, permission verification methods, and audit log requirements. A security level mapping table is constructed to correlate user authentication and operation behavior records in the source network with device access permissions and operation approval processes in the target network.
[0084] The multi-level partitioning of the visual semantic space is implemented based on the matching degree between the operation granularity and control capability in the heterogeneous characteristic model. The first semantic level corresponds to the device-level control capability of the target network, handling visual semantics involving basic operations such as single device switching and state transitions. The mapping function adopts a direct correspondence strategy, where the gesture "swipe right" is directly mapped to the "device start" command. The second semantic level corresponds to the system-level control capability, handling visual semantics involving complex operations such as multi-device coordination and process control. The mapping function adopts a combinatorial decomposition strategy, where complex gesture sequences are decomposed into multiple parallel or serial device control commands. The third semantic level corresponds to the business-level control capability, handling visual interactions involving high-level semantics such as scene switching and mode adjustment. The mapping function adopts a semantic reasoning strategy, converting abstract operational intentions into specific control flows through context analysis. Each semantic level maintains an independent feature space: the first level uses pixel-level features and geometric shape descriptors, the second level uses temporal patterns and action sequence features, and the third level uses semantic labels and intent classification features. The mapping function is established by optimizing the parameters using the training dataset. The training set contains 10,000 sets of labeled visual interaction samples and corresponding control commands, while the validation set contains 2,000 sets of samples for model evaluation. The mapping accuracy is required to reach over 95%.
[0085] The decomposition of operational intent semantics in multimodal feature representation is achieved through a semantic parsing engine. This engine receives feature vectors containing spatial morphological features, temporal evolution features, and contextual information as input. Spatial morphological features include geometric information such as the coordinate sequence of gesture contour points, the distribution of eye-tracking gaze points, and the positions of key facial expression points, with a feature dimension set to a 512-dimensional vector. Temporal evolution features include time-related information such as action duration, velocity change curves, and acceleration distribution, extracted using a sliding window method with a window length of two seconds and an overlap rate of 50%. The semantic parsing engine employs a multi-layer neural network architecture, comprising three main components: a feature fusion layer, a semantic extraction layer, and an intent classification layer. The feature fusion layer uses an attention mechanism to weightedly fuse multimodal features, with attention weights automatically learned through end-to-end training. The semantic extraction layer uses a recurrent neural network structure to process temporal features, with a hidden layer dimension of 256 and a learning rate of 0.001. The intent classification layer outputs control semantic units that match the control capabilities of the target network. Each semantic unit contains structured information such as operation type, parameter range, and execution conditions.
[0086] The organization of control semantic units employs a strategy combining dependency analysis and resource scheduling priority calculation. Dependency analysis uses a directed acyclic graph (DAG) to represent the sequence and constraints between control semantic units. Nodes in the graph represent individual control semantic units, and edges represent execution dependencies. Dependency identification is based on factors such as the operation object, resource requirements, and preconditions. A serial dependency is established when two control semantic units operate on the same device, and a data dependency is established when a control semantic unit requires the output of a preceding operation as input. Resource scheduling priority is calculated by comprehensively considering factors such as operation urgency, resource consumption, and execution time. Urgency is scored from one to ten, resource consumption is expressed as a percentage, and execution time is quantified in milliseconds. The priority calculation formula uses a weighted summation method, with urgency weighted at 0.5, resource consumption weighted at 0.3, and execution time weighted at 0.2. The hierarchical control instruction set is organized using a tree structure. The root node represents the complete operation intent, leaf nodes represent the most basic device control instructions, and intermediate nodes represent control subtasks at different abstraction levels.
[0087] The hierarchical control instruction set uses a nested dictionary format for its data structure. Each level of instruction includes fields such as instruction identifier, execution parameters, dependencies, timeout settings, and error handling strategies. Instruction identifiers follow a hierarchical naming convention: first-level instructions are prefixed with L1, second-level instructions with L2, and so on. Execution parameters are stored as key-value pairs, supporting data types such as integers, floating-point numbers, strings, and booleans. Parameter validation is constrained by predefined value ranges and format rules. Dependencies are described using logical expressions, supporting logical operators such as AND, OR, and NOT. Condition judgments are based on information such as system status, resource availability, and the execution results of preceding instructions. Timeout settings include two mechanisms: instruction-level timeout and hierarchical timeout. Instruction-level timeouts are set for individual control instructions, with a default value of five seconds. Hierarchical timeouts are set for the entire instruction set at each level, with a default value of thirty seconds. Error handling strategies include retry mechanisms, degraded execution, and exception reporting options. The default number of retries is three, and the retry interval uses an exponential backoff strategy with an initial interval of one second.
[0088] In one optional implementation, the instructions are classified for security based on their execution priority within the hierarchical control instruction set. The hierarchical control instruction set is then transmitted to the target network environment via a non-physical connection transmission medium. Temporary trust credentials are constructed at the security domain boundary node based on behavioral pattern verification, including:
[0089] The execution priority sequence of the hierarchical control instruction set is analyzed and modeled in depth. The expected intensity of resource consumption and sensitivity to state changes of each instruction are dynamically evaluated through machine learning algorithms. The instruction risk score is calculated based on the expected intensity of resource consumption and the sensitivity to state changes. The instructions are divided into different security access levels according to the instruction risk score.
[0090] Differentiated transmission encapsulation structures are constructed for instructions at different security access levels. The differentiated transmission encapsulation structures include adaptive redundancy check codes. The layered control instruction sets encapsulated in differentiated manner are transmitted to the target network environment through a non-physical connection transmission medium. A dynamic timing feature library for instruction transmission is established based on the non-physical connection transmission medium.
[0091] The received instruction sequence is parsed at the security domain boundary node of the target network environment. Based on the dynamic temporal feature library, the verification result of the adaptive redundancy check code and the logical dependency relationship between instructions are used to generate a behavior pattern feature vector. The behavior pattern feature vector is then adaptively matched with a pre-established historical normal behavior benchmark library. The deviation metric between the behavior pattern feature vector and each benchmark pattern in the historical normal behavior benchmark library is calculated.
[0092] When the minimum value of the deviation metric is lower than the trust threshold, dynamic optimization modeling is performed based on the permission configuration template in the corresponding benchmark mode, and a temporary trust credential is generated by combining the adaptive redundancy check code verification result in the current instruction sequence with the instruction risk score.
[0093] like Figure 2 As shown, the method includes:
[0094] The deep analysis and modeling of execution priority sequences in a hierarchical control instruction set is implemented through a priority parsing module. This module receives a hierarchical control instruction set containing instruction identifiers, execution parameters, and dependencies as input data. The priority sequence extraction algorithm traverses the tree structure of the instruction set, extracting priority values for each level of instruction according to a breadth-first search strategy, forming a one-dimensional priority sequence vector. The sequence length equals the total number of instructions, and the values range from one to ten. The machine learning algorithm uses a support vector regression model to model the priority sequences. The training dataset contains 5,000 sets of historical instruction sequences and their corresponding measured resource usage data. The model input features include a 12-dimensional feature vector encompassing instruction type encoding, parameter complexity, dependency depth, and execution time prediction. The training iterations are set to 1,000, and the learning rate is set to 0.01.
[0095] The assessment of expected resource consumption intensity is based on a quantitative calculation of the consumption of computing, storage, and network resources during instruction execution. Computational resource intensity is estimated by multiplying instruction complexity by processor utilization. Instruction complexity is comprehensively evaluated based on factors such as algorithm time complexity, data size, and concurrency requirements, with a value ranging from 0.1 to 10.0. Processor utilization is predicted using the moving average of historical execution data, with a sliding window length set to twenty execution records. Storage resource intensity is calculated by a weighted sum of peak memory usage and persistent storage requirements, with weighting coefficients set to 0.7 and 0.3, respectively. Network resource intensity is quantified by multiplying the amount of data transmitted by the duration of bandwidth usage. Sensitivity to state changes is calculated by analyzing the scope and degree of impact of instruction execution on the system state. The scope of impact is quantified based on dimensions such as the number of operating devices, involved service modules, and affected user groups. The degree of impact is calculated based on factors such as the reversibility of state changes, recovery costs, and business interruption risks.
[0096] The instruction risk score is calculated using a weighted fusion strategy of expected resource consumption intensity and state change sensitivity. The weighting coefficient for expected resource consumption intensity is set to 0.6, and the weighting coefficient for state change sensitivity is set to 0.4. The risk score calculation results are normalized to the range of 0 to 100. Instructions with scores below 20 are classified as low-risk and corresponding to the basic access level; instructions with scores between 20 and 60 are classified as medium-risk and corresponding to the restricted access level; and instructions with scores above 60 are classified as high-risk and corresponding to the strict access level.
[0097] The differentiated transmission encapsulation structure employs corresponding encapsulation strategies for different security access levels. The basic access level uses a standard encapsulation format, comprising a command header, payload data, and a checksum tail. The command header is 32 bytes long and includes fields such as version information, command type, data length, and timestamp. The payload data is encoded in JSON format, and the checksum tail uses the CRC32 algorithm to generate a four-byte checksum. The restricted access level adds digital signature and timeliness verification to the standard encapsulation. The digital signature uses the RSA algorithm, with a key length of 2048 bits and a default validity period of 60 seconds. The strict access level uses double-encryption encapsulation: the inner encryption uses the AES algorithm with a key length of 256 bits, and the outer encryption uses elliptic curve cryptography.
[0098] The adaptive redundancy check (CDR) code generation is dynamically adjusted based on the characteristics of the instruction content and the transmission environment. Instruction content characteristics include parameters such as data complexity, importance level, and execution frequency, while transmission environment characteristics include network quality, latency jitter, and packet loss rate. The CDR generation algorithm selects the appropriate check strength based on a comprehensive score of the content and environment characteristics: CRC16 checksum is used when the score is below 30, CRC32 checksum is used when the score is between 30 and 70, and MD5 hash checksum is used when the score is above 70. Non-physical connection transmission media include various methods such as wireless networks, Bluetooth communication, infrared communication, and acoustic communication. The system automatically selects the optimal transmission method based on the environmental detection results and instruction characteristics.
[0099] The dynamic time-series feature database is established by continuously monitoring and updating the time-series characteristics during command transmission. These characteristics include network performance metrics such as transmission delay, jitter variance, bandwidth utilization, and retransmission count. The database uses a sliding window mechanism for data collection, with a window length of 1,000 transmission records and a sliding step size of 100 records. Feature data is stored in a time-series database with a data retention period of 30 days, and the database is updated every 10 minutes.
[0100] The instruction sequence parsing of security domain boundary nodes is implemented through a multi-layered parsing engine, comprising four layers: physical layer parsing, data link layer parsing, network layer parsing, and application layer parsing. The parsing engine adopts a pipelined architecture, with asynchronous communication between layers via message queues. The queue depth is set to 1,000 messages, and the message processing timeout is set to 5 seconds. The generation of behavioral pattern feature vectors is based on a comprehensive analysis of a dynamic temporal feature library, adaptive redundancy check code verification results, and logical dependencies between instructions. The feature vector dimension contributed by temporal features is set to 64 dimensions, the feature vector dimension contributed by check code verification results is set to 32 dimensions, and the feature vector dimension contributed by logical dependencies is set to 128 dimensions.
[0101] The historical normal behavior benchmark database contains 10,000 normal behavior samples. K-means clustering is used, with 100 clusters, each representing a typical normal behavior pattern. An adaptive similarity matching algorithm uses cosine similarity to calculate the similarity between the behavior pattern's feature vector and each benchmark pattern in the database, employing an approximate nearest neighbor search algorithm to improve computational efficiency. The deviation metric is calculated by subtracting the maximum similarity from one, ranging from zero to two. A confidence threshold is set at 0.3; when the minimum deviation metric is below 0.3, the current behavior pattern is considered to be within the confidence range.
[0102] The generation of temporary trust credentials employs a strategy combining dynamic optimization modeling of permission configuration templates with current instruction characteristics. The permission configuration template is extracted from the matched baseline pattern and includes configuration items such as access permission lists, operation scope restrictions, time validity periods, and resource quotas. The adaptive redundancy checksum verification result serves as a reference factor for permission adjustment. When the verification success rate is higher than 98%, the permission configuration retains the original template settings; when the verification success rate is lower than 95%, permissions are further restricted and the validity period is shortened. Instruction risk scores serve as another permission adjustment factor; when the risk score is lower than 30, permission configuration can be appropriately relaxed; when the risk score is higher than 60, permission configuration needs to be strictly restricted.
[0103] In one optional implementation, the expected resource consumption intensity and state change sensitivity of each instruction are dynamically evaluated using a machine learning algorithm. The instruction risk score is calculated based on the expected resource consumption intensity and the state change sensitivity, including:
[0104] An instruction behavior feature vector is constructed from the historical execution records of the hierarchical control instruction set. The instruction behavior feature vector contains the actual resource consumption data and the state change impact data of each instruction. The actual resource consumption data and the state change impact data are combined into a training sample set.
[0105] A deep neural network model is constructed based on the training sample set. The deep neural network model learns the correlation between the actual resource occupancy data and the state change impact data to obtain the resource occupancy assessment weight and the state change assessment weight.
[0106] A multi-layer feature mapping matrix is established for the instruction to be evaluated in the hierarchical control instruction set. The multi-layer feature mapping matrix is input into the deep neural network model. The expected intensity of resource consumption of the instruction is calculated using the resource consumption evaluation weight, and the sensitivity of state change of the instruction is calculated using the state change evaluation weight.
[0107] The expected intensity of resource consumption and the sensitivity to state changes are input into the deep neural network model to generate an instruction risk score.
[0108] Historical execution records of the hierarchical control instruction set are collected through an execution monitoring module deployed on various execution nodes in the target network environment to monitor resource consumption and state changes during instruction execution in real time. The historical execution record data structure includes four main fields: instruction identifier, execution timestamp, resource usage details, and state change details. Data storage uses a time-series database with a retention period of six months and a data sampling frequency of ten times per second. Actual resource usage data covers four dimensions: CPU utilization, memory consumption, network bandwidth usage, and disk I / O operations. CPU utilization is recorded as a percentage, with precision to two decimal places; memory consumption is recorded in megabytes; network bandwidth usage is recorded in kilobits per second; and disk I / O operations are recorded in read / write operations per second. State change impact data includes four evaluation dimensions: change scope, change depth, impact duration, and recovery complexity. The change scope is quantified by the number of affected modules; the change depth is assessed by the degree of change in state hierarchy; the impact duration is recorded in seconds; and the recovery complexity is quantified by the number of rollback steps required.
[0109] The construction of instruction behavior feature vectors is achieved through a feature engineering module, which preprocesses and extracts features from the original execution records. The feature vectors are set to 128 dimensions, including 64 static features and 64 dynamic features. Static features cover attributes that do not change with the execution environment, such as instruction type encoding, parameter complexity, dependency depth, and target device type. Instruction type encoding uses one-hot encoding, supporting 32 different instruction types. Parameter complexity is calculated by multiplying the number of parameters by the nesting level, and dependency depth is quantified by the maximum path length of the instruction dependency graph. Dynamic features cover attributes that change with the execution environment, such as historical execution success rate, average execution time, resource consumption volatility, and state change frequency. Historical execution success rate is calculated by dividing the number of successful executions in the last 100 executions by the total number of executions. Average execution time uses a sliding window average with a window length of 50 execution records. Resource consumption volatility is quantified using the variance coefficient, and state change frequency is calculated by the number of changes per unit time.
[0110] The training sample set is constructed using a sample generator, which extracts instruction execution instances that meet certain criteria from historical execution records. Sample selection criteria include execution integrity, data integrity, and time validity. Execution integrity requires that the complete execution process of an instruction from start to finish be recorded. Data integrity requires that resource usage data and state change impact data be fully collected. Time validity requires that the execution record be generated within the last three months. The training sample set is set to 100,000 samples, with 80% being normal execution samples and 20% being abnormal execution samples. Sample labels are generated based on expert experience and automated rules. Data preprocessing includes three steps: outlier detection, missing value imputation, and data normalization. Outlier detection uses the three-standard-deviation principle, truncating values outside the range. Missing value imputation uses a forward imputation strategy. Data normalization uses a maximum-minimum scaling method to scale all feature values to the range of zero to one.
[0111] The deep neural network model employs a multilayer perceptron architecture, comprising a five-layer structure: an input layer, three hidden layers, and an output layer. The number of nodes in the input layer is equal to the dimension of the instruction behavior feature vector, set to 128 nodes. The number of nodes in the hidden layers are set to 256, 128, and 64 respectively, with the ReLU activation function used to avoid the vanishing gradient problem. The output layer contains two sub-outputs: a resource occupancy evaluation output and a state change evaluation output, each containing a corresponding evaluation weight vector. Network training uses the backpropagation algorithm, with mean squared error loss as the loss function, the Adam algorithm as the optimizer, a learning rate of 0.001, a batch size of 64, and 1000 training epochs. Model regularization uses Dropout with a dropout rate of 0.5 to prevent overfitting. An early stopping strategy is employed during training: training stops when the validation set loss fails to improve for 20 consecutive epochs, and the model parameters with the minimum validation set loss are saved.
[0112] Resource occupancy assessment weights and state change assessment weights are obtained through weight learning during network training. The resource occupancy assessment weight matrix has a dimension of 128 x 64, with each weight value representing the contribution of the corresponding feature to resource occupancy prediction. Weight values range from -1 to +1, with larger absolute values indicating higher feature importance. The state change assessment weight matrix also has a dimension of 128 x 64. Weight learning is optimized using the gradient descent algorithm, and the weight update frequency is synchronized with the overall network training. The weight matrix is initialized using the Xavier initialization method to ensure gradient stability in the early stages of training. The interpretability of the weight vectors is evaluated through feature importance analysis, calculating the average absolute value of the weight for each feature dimension, and then ranking and identifying the top twenty features with the greatest impact on the prediction results.
[0113] The multi-layer feature mapping matrix of the instruction to be evaluated is constructed through a feature mapping module. This module receives the basic attributes of the instruction as input and generates a feature representation consistent with that of the training phase. The first layer of the multi-layer feature mapping matrix is the original feature layer, directly corresponding to the basic attributes of the instruction. The second layer is the combined feature layer, generated through feature cross-multiplication and polynomial expansion. The third layer is the abstract feature layer, generated through nonlinear transformation and dimensionality reduction operations. The feature mapping process adopts the same preprocessing procedure as the training phase to ensure the consistency of feature distribution. The mapping matrix has a dimension of 3 x 128, with each layer containing 128 feature values. The data type of the matrix elements is 32-bit floating-point numbers, and the value range is normalized and limited to between zero and one.
[0114] The calculation of expected resource consumption intensity is achieved through matrix multiplication of the multi-layer feature mapping matrix and resource consumption evaluation weights. The calculation process includes three steps: weighted summation of feature weights, nonlinear activation, and output normalization. Weighted summation of feature weights uses a dot product operation, nonlinear activation uses the sigmoid function to limit the output to the range of zero to one, and output normalization ensures the comparability of expected intensities for different instructions. The output value of expected resource consumption intensity represents the resource consumption level of the instruction relative to the historical average level; the closer the value is to one, the higher the resource consumption, and the closer the value is to zero, the lower the resource consumption. The calculation precision is maintained to four decimal places, the calculation time is controlled within ten milliseconds, and concurrent computation is supported to improve processing efficiency.
[0115] The calculation of state change sensitivity employs a matrix operation method similar to that used for expected resource consumption intensity, replacing resource consumption assessment weights with state change evaluation weights. State change sensitivity reflects the degree to which instruction execution affects the system state; a higher value indicates stronger sensitivity to state changes and a greater impact on system stability. Sensitivity calculation considers factors such as the instruction's operational scope, depth of impact, and propagation, automatically capturing the complex relationships between these factors through weight learning. The validity of the calculation results is verified by comparison with expert evaluation results, with a correlation coefficient required to be above 0.9.
[0116] The risk score is generated through a risk scoring network. This network receives the expected intensity of resource occupancy and the sensitivity to state changes as input, and outputs a comprehensive risk score. The risk scoring network employs a shallow neural network structure, including an input layer, one hidden layer, and an output layer. The hidden layer has 32 nodes, and the Tanh function is used as the activation function. The risk score calculation considers the interaction effect of resource occupancy and state changes; when both are at high levels simultaneously, the risk score exhibits a non-linear growth trend. The output layer uses a linear activation function, and the risk score ranges from zero to one hundred, with higher values indicating greater risk. The risk score is calibrated using historical security event data to ensure consistency between the score and the actual risk level.
[0117] Performance optimization for model inference is achieved through multiple strategies. A feature caching mechanism stores frequently accessed feature vectors in memory, reducing redundant computation overhead. The cache capacity is set at 10,000 feature records, employing an LRU eviction policy. Batch inference supports simultaneous processing of risk assessments for multiple instructions, with a batch size of 128, fully utilizing the parallel computing capabilities of the GPU. Model quantization technology compresses 32-bit floating-point weights into 16-bit integers, reducing storage space and computation time while maintaining accuracy. Inference acceleration is achieved through model pruning, removing connections where the absolute value of the weights is less than a threshold set to 0.01. Pruning reduces the model size by 30% and improves inference speed by 50%.
[0118] In one optional implementation, in the target network environment, the execution conditions are determined and control operations are implemented based on the temporary trust credential. Simultaneously, state semantic features containing operation response timing and resource occupancy status are extracted, and these state semantic features are encoded into an abstract state descriptor, including:
[0119] The temporary trust credential is parsed in the execution engine of the target network environment to determine the execution conditions; the execution of control instructions is triggered based on the execution conditions; a dual-stream feature network is constructed during the instruction execution process; and parallel feature extraction is performed on the operation response timing and resource occupancy status based on the dual-stream feature network.
[0120] State semantic features are generated based on the features output by the dual-stream feature network, and the state semantic features are converted into abstract state descriptors.
[0121] During the process of determining execution conditions and implementing control operations based on temporary trust credentials in the target network environment, a temporary trust credential containing access permissions and expiration information is received. The execution engine verifies the credential using a key verification algorithm to confirm its legitimacy. The verification process includes checking whether the digital signature, timestamp, and permission scope meet preset conditions. For example, when a temporary credential in the format "TTC-R3-AC5-20240515-1200-SYS_ADMIN" is received, the system parses it and determines that the credential has R3 level permissions, is valid until 12:00 on May 15, 2024, and has system administrator privileges.
[0122] After the temporary trust credential is verified, the execution engine further determines the execution conditions, including resource availability checks, security policy compliance analysis, and operational risk assessment. Resource availability checks ensure the target system has sufficient computing resources to execute the requested operation, such as memory usage below 80% and CPU load below 75%. Security policy compliance analysis ensures the requested operation complies with network security policy requirements. Operational risk assessment calculates the risk value of the operation, and execution is only permitted if the risk value is below a threshold (e.g., 0.65).
[0123] Once the execution conditions are met, control commands are triggered based on these conditions. The execution of control commands employs a phased execution strategy, decomposing the complete operation sequence into three phases: initialization, main execution, and status verification. In the initialization phase, the system prepares the execution environment and allocates necessary resources; in the main execution phase, the core operation logic is executed; and in the status verification phase, the operation execution results are verified, and relevant log information is recorded.
[0124] A dual-stream feature network is constructed during instruction execution to achieve parallel feature extraction of operation response timing and resource occupancy status. The dual-stream feature network consists of a timing feature extraction stream and a resource feature extraction stream. The timing feature extraction stream employs a sliding time window technique to sample the time-series data of the operation response. For example, a sliding window with a size of 200 milliseconds and a step size of 50 milliseconds is used to extract time features from the operation response. By analyzing the time intervals between discrete sampling points, the timing patterns of the operation response can be identified, such as typical patterns like "fast response - delayed processing - fast completion".
[0125] The resource feature extraction stream focuses on the dynamic changes of resource indicators such as CPU utilization, memory usage, and network traffic. The system collects resource usage data every 100 milliseconds, forming a resource usage time series. Through combined analysis of multi-dimensional resource indicators, characteristic patterns of resource usage can be identified, such as "CPU-intensive," "memory-intensive," or "network-intensive." In practical applications, a certain operation exhibits a characteristic pattern where CPU utilization sharply increases from 30% to 85%, remains at that level for 300 milliseconds, and then gradually decreases to 40%, while memory usage fluctuates around 45%.
[0126] The dual-stream feature network extracts features from the temporal stream and resource stream separately, and then integrates the two features through a feature fusion module to generate a comprehensive state semantic feature. Feature fusion employs a multi-level fusion strategy, first reducing the dimensionality of each feature, and then fusing them through a weighted combination. The fusion weights are dynamically adjusted according to different operation types. For example, for computationally intensive operations, the weight of resource features is 0.7, while the weight of temporal features is 0.3; for interactive operations, the weight of temporal features is increased to 0.6, and the weight of resource features is reduced to 0.4.
[0127] The fused state semantic features are converted into abstract state descriptors by an encoder. The encoder employs a multi-level mapping structure to compress and map the high-dimensional feature space to a fixed-length descriptor space. The conversion process includes three steps: feature normalization, principal component extraction (PCE), and semantic encoding. Feature normalization scales the features of each dimension to the [-1, 1] interval; PCE retains more than 95% of the feature information; and semantic encoding maps the principal component features to fixed-length binary or hexadecimal codes. The final generated abstract state descriptor has the format "SD-0xA7F9B3D2-R85-T42", where 0xA7F9B3D2 is the state feature hash value, R85 indicates a resource utilization intensity of 85%, and T42 indicates a time complexity of 42.
[0128] After state descriptors are generated, they are stored in a state database for subsequent anomaly detection and behavior analysis. The descriptors in the state database are periodically clustered to identify typical operational patterns and abnormal behavior patterns. When a newly generated state descriptor has a similarity of more than 85% to a known abnormal pattern, the system will trigger a security alert.
[0129] In practical applications, taking system configuration update operations as an example, the complete process is as follows: Receive the configuration update request and the attached temporary trust credential "TTC-R4-CF8-20240510-1500-CFG_UPDATE"; verify the validity of the credential and determine that the execution conditions are met; trigger the configuration update command; during the execution process, extract the temporal features (configuration file loading time 75 ms, verification time 120 ms, application time 350 ms) and resource features (CPU peak utilization 62%, memory increase 8%, network traffic increase 2.5MB) through a dual-stream feature network; fuse the features to generate state semantic features; encode them into an abstract state descriptor "SD-0xB2C4D6E8-R62-T54"; and store them in the state database to complete the entire processing flow.
[0130] In one optional implementation, the abstract state descriptor is converted into a visual representation perceptible to the source network environment through inverse semantic mapping. After spatiotemporal alignment with the multimodal feature representation, a bidirectional mapping relationship graph between operation intent and execution result is constructed, including:
[0131] The abstract state descriptor is hierarchically decomposed, and the decomposed information is organized into state attribute identifiers and state value descriptions. The state attribute identifiers and state value descriptions are used as input data for reverse mapping.
[0132] Based on the state attribute identifier and the state value description, a reverse semantic mapping transformation is performed to map the state attribute identifier to the corresponding visual element type label in the source network environment, and the state value description to the corresponding visual element appearance feature parameter in the source network environment. A perceptible visual representation of the source network environment is generated through the visual element type label and the visual element appearance feature parameter.
[0133] Based on the visual representation combined with the operation time marker and operation area coordinates in the multimodal feature representation, bidirectional alignment processing is performed. The spatial position coordinates in the visual representation are spatially aligned with the operation area coordinates, and the execution completion time corresponding to the abstract state descriptor is temporally aligned with the operation time marker, generating spatiotemporally aligned visual representation and spatiotemporally aligned operation feature pairing data.
[0134] A bidirectional mapping relationship map between operation intention and execution result is established by pairing the spatiotemporally aligned visual representation with the spatiotemporally aligned operation feature data.
[0135] The hierarchical decomposition of abstract state descriptors is implemented through a state parsing module. This module receives abstract state descriptors containing device state, system state, and environment state as input data. The state parsing module employs a recursive descent parsing algorithm, decomposing the descriptors layer by layer according to a predefined state hierarchy. The state hierarchy consists of four levels: the first level is the system level state, the second level is the module level state, the third level is the device level state, and the fourth level is the parameter level state. The decomposition process determines the level by matching the prefixes of the state identifiers: system level state identifiers are prefixed with SYS, module level state identifiers with MOD, device level state identifiers with DEV, and parameter level state identifiers with PAR.
[0136] The data structure for status attribute identifiers includes three fields: hierarchy tag, category code, and entity identifier. The hierarchy tag uses a two-digit numeric code, the category code uses a four-digit alphanumeric code, and the entity identifier uses an eight-digit numeric code. The data structure for status value descriptions includes four fields: numeric type, value range, precision requirement, and unit information. The numeric type supports four basic types: integer, floating-point, Boolean, and string. The value range is represented by a minimum and maximum value interval. The precision requirement is indicated by the number of decimal places. The unit information is encoded using the International System of Units (SI).
[0137] The reverse semantic mapping transformation is implemented through a mapping engine, which maintains a mapping table from state attribute identifiers to visual element type tags and conversion rules from state value descriptions to visual element appearance feature parameters. Visual element type tags include three main categories: graphic type, interactive type, and dynamic type. Graphic types cover geometric shapes such as circles, rectangles, triangles, and polygons; interactive types cover control elements such as buttons, sliders, switches, and indicators; and dynamic types cover motion effects such as animation, gradients, blinking, and rotation. The mapping table is built based on training with a large number of historical mapping instances, containing 20,000 sets of correspondences between state identifiers and visual elements, with a required mapping accuracy of over 95%. The conversion from state value descriptions to visual element appearance feature parameters is implemented through a parameter mapping function, which converts abstract state values into specific visual parameters. Appearance feature parameters include visual attributes such as color, size, position, transparency, border, and fill. Color parameters use the RGB color model, with values ranging from zero to 255 integers; size parameters are in pixels; position parameters use screen coordinates; and transparency parameters are represented as floating-point numbers from zero to one.
[0138] The combination of visual element type tags and visual element appearance feature parameters is achieved through a visual compositor. This compositor selects the corresponding visual template based on the type tag and parameterizes the template according to the appearance feature parameters. The visual template library contains fifty basic templates, each defining the structure and default attributes of a specific visual element. Templates are stored in SVG format and support scaling and deformation of vector graphics. The parameterization process is implemented through a template engine, which parses the parameter placeholders in the template, fills in the appearance feature parameters in the corresponding positions, and generates a complete visual element description. The generated visual representation uses a structured data format, including fields such as element identifier, type information, attribute parameters, position coordinates, and timestamps. The data format uses JSON encoding and supports nested structures and array types. The verification of the visual representation is performed through a rendering engine, which converts the visual representation into actual image output to verify the correctness and completeness of the generated results.
[0139] The bidirectional alignment process is based on a spatiotemporal synchronization algorithm, which simultaneously handles the synchronization issues of spatial location alignment and temporal sequence alignment. Spatial alignment is achieved through a coordinate transformation matrix, which converts the spatial location coordinates in the visual representation into a coordinate system identical to the coordinates of the operating area. The coordinate system transformation considers three basic transformations: translation, rotation, and scaling. Translation parameters are determined by the offset between the origins of the two coordinate systems, rotation parameters by the angle between the coordinate axes, and scaling parameters by the proportional relationship of the coordinate units. The transformation matrix is calculated using the least squares method, calculating the optimal transformation parameters based on multiple sets of corresponding points. The number of corresponding points must be no less than four, and the calculation accuracy must reach the pixel level. The error assessment of spatial alignment is quantified by the average distance error, and the error value must be controlled within five pixels.
[0140] Time alignment is achieved through a timestamp synchronization mechanism, which calibrates the execution completion time and operation time marker corresponding to the abstract state descriptor. Time calibration considers the impact of network transmission latency, processing latency, and synchronization errors on the timestamp, and corrects for these factors using a time offset compensation algorithm. This algorithm uses a network time protocol for clock synchronization, requiring millisecond-level accuracy. Time offset is measured using round-trip time measurement, with a measurement frequency set to once per minute. The execution completion time is determined based on the completion flag of the state change; the completion time is recorded when all state parameters in the state descriptor reach the target value. Operation time marker extraction is based on time-series data from multimodal feature representations, using a peak detection algorithm to identify key moments in the operation.
[0141] The generation of paired data between spatiotemporally aligned visual representations and spatiotemporally aligned operational features is achieved through a pairing algorithm. This algorithm associates visual representations with operational features based on the principle of spatiotemporal proximity. Spatiotemporal proximity is determined based on Euclidean distance calculations, with a weighting coefficient of 0.6 for spatial distance and 0.4 for temporal distance. A pairing threshold is set at a comprehensive distance of less than ten units; pairings exceeding this threshold are filtered out. The paired data structure includes five fields: visual representation identifier, operational feature identifier, spatial alignment parameter, temporal alignment parameter, and confidence score. The confidence score is calculated based on alignment accuracy and feature similarity, and its value ranges from zero to one (floating-point number). The quality of the paired data is assessed using cross-validation. The validation set contains one thousand manually labeled paired samples, and the accuracy of automatic pairing is required to reach over 90%.
[0142] The bidirectional mapping relationship graph is established using a graph construction algorithm that organizes spatiotemporally aligned paired data into a directed graph structure. Nodes in the graph are divided into two types: operation intention nodes and execution result nodes. Operation intention nodes correspond to operational behaviors in multimodal feature representation, while execution result nodes correspond to state feedback in visual representation. Node attributes include fields such as node identifier, node type, feature vector, timestamp, and confidence level. The feature vector has a dimension of 512, and a deep learning model is used for feature extraction. Edges in the graph represent the causal relationship between operation intentions and execution results. Edge attributes include fields such as relationship type, relationship strength, delay time, and success probability. Relationship strength is quantified through correlation analysis, delay time is calculated using time difference, and success probability is estimated using historical statistical data.
[0143] The graph is stored using a graph database, supporting efficient graph traversal and query operations. The graph database's indexing strategy includes two types: node indexes and edge indexes. Node indexes are built based on node identifiers and timestamps, while edge indexes are built based on relation types and relation strengths. The graph is updated using an incremental update mechanism; new pairings are dynamically added with corresponding nodes and edges. Expired historical data is deleted through periodic cleanup tasks, with a data retention period set at one month. Graph queries support various query modes, including path queries, neighbor queries, and subgraph queries. Query performance is optimized through a caching mechanism, with a cache capacity of 10,000 query results and a cache hit rate requirement of over 80%.
[0144] The verification of the bidirectional mapping relationship is achieved through a loop closure detection algorithm, which detects the complete mapping path from the operation intention to the execution result and back to the operation intention. Loop closure detection is based on a graph traversal algorithm, starting from the operation intention node and performing a depth-first search along the directed edges. A complete loop is formed when the search path returns to the starting node. The quality evaluation of loop closures is based on a comprehensive score using indicators such as path length, number of nodes, and edge weights. High-quality loop closures indicate a stable bidirectional mapping relationship between the operation intention and the execution result. Statistical analysis of the mapping relationship is achieved through a graph mining algorithm, which identifies frequent and abnormal patterns in the graph, providing a basis for system optimization decisions.
[0145] The method further includes:
[0146] The vision-based cross-network bidirectional interactive mouse control system acquires user operation commands through a vision acquisition layer. This layer is equipped with a standard RGB camera module with a working resolution of 1920×1080 pixels and a frame rate maintained at 30 frames per second to ensure real-time and clear command capture. The camera connects to the control terminal via a USB 3.0 interface and achieves stable acquisition of image data streams through DirectShow or the V4L2 driver framework. The vision acquisition module continuously monitors a preset gesture recognition area, which occupies a 640×480 pixel area in the center of the screen. When a hand is detected entering this area, the command recognition process is automatically activated.
[0147] The visual analysis module receives raw image frame data transmitted from the acquisition layer and performs preprocessing operations including image denoising, brightness equalization, and contrast enhancement. The gesture recognition algorithm is based on a combination of contour detection and feature point matching. The Canny edge detection algorithm is used to extract the hand contour, with edge detection thresholds set to a low threshold of 50 and a high threshold of 150. The system predefines five basic gesture patterns corresponding to different mouse operations: single finger extension corresponds to a left click, two fingers spread correspond to a right click, a fist clench corresponds to the start of a drag, an open palm corresponds to the end of a drag, and rapid finger waving corresponds to a double-click. The feature matching process uses a template matching algorithm with a similarity threshold set to 0.85. Gesture validity is confirmed when the matching degree exceeds this threshold.
[0148] After command parsing, a structured control data packet is generated. The data packet format includes a command type field, a coordinate information field, a timestamp field, and a checksum field. The command type field uses 8-bit encoding, with a value range of 0-255, where 1 represents a left click, 2 represents a right click, 3 represents the start of dragging, 4 represents the end of dragging, and 5 represents a double-click. The coordinate information field uses 16-bit encoding to record the mouse target position, with the X and Y coordinates each occupying 2 bytes, supporting a maximum resolution range of 65535×65535. The timestamp field uses a 32-bit Unix timestamp format, accurate to the millisecond level, for timing synchronization during cross-network transmission. The checksum field calculates the checksum value of all the aforementioned fields using the CRC16 algorithm to ensure data transmission integrity.
[0149] The cross-network relay system is deployed at network boundary nodes, responsible for receiving control commands from the source network and forwarding them to the target network. The relay system employs a dual-NIC configuration: one NIC connects to the source network, and the other connects to the target network. Command relay between the two network interfaces is achieved through software-level data transfer. To ensure security isolation requirements, the relay system internally operates a whitelist verification mechanism, allowing only pre-configured command types to be forwarded. The whitelist configuration file is stored in JSON format, containing a list of allowed command types, a range of source network IP addresses, a range of target network IP addresses, and a time window limit. During command forwarding, the system records detailed audit logs, including the source IP address, target IP address, command type, execution time, and processing result.
[0150] Non-physical connection transmission utilizes optical signal modulation technology to achieve cross-network data transfer. The transmitting end modulates digital command data into an optical signal of a specific frequency, and the receiving end uses a photoelectric sensor to restore the optical signal to digital data. The optical signal carrier frequency is set to 38kHz, and the modulation method is pulse width modulation, with a pulse width of 0.56 milliseconds for logic 0 and 1.69 milliseconds for logic 1. The transmitting end uses an infrared LED as the light source, with the transmission power controlled within 5mW, and an effective transmission distance of up to 3 meters. The receiving end is equipped with an infrared receiver tube with a response wavelength range of 870-970 nanometers and a receiving sensitivity of not less than 0.1mW / cm². 2 To improve transmission reliability, the system employs differential Manchester encoding, where each data bit is represented by a level transition, effectively reducing the impact of ambient light interference.
[0151] The execution layer module in the target network receives and parses the instruction data transmitted across the network. It verifies the integrity of the data packets through CRC16 checksum comparison; packets failing the check are discarded and an error log is recorded. Valid instructions are parsed and converted into underlying mouse events on the operating system. This is implemented using the SendInput API function on Windows, the uinput device interface on Linux, and the CGEvent series of functions on macOS. Mouse coordinate mapping takes into account the target screen resolution difference, employing a scaling algorithm to convert source coordinates to target coordinates. The scaling formula is: target coordinates equal to source coordinates multiplied by target resolution divided by source resolution.
[0152] The permission verification mechanism is based on a multi-level verification strategy. The first level verifies the legitimacy of the source network identity using an HMAC-SHA256 digest verification via a pre-shared key. The key length is set to 256 bits and is automatically updated every 24 hours. The second level verifies command type permissions, configuring different operation permission matrices based on user roles. Administrators have full command execution permissions, ordinary users are only allowed basic mouse operations, and visitors are restricted to read-only viewing permissions. The third level verifies target resource access permissions, limiting the scope of operable applications and file directories through resource access control lists.
[0153] In a real-world application, a user makes a single-finger outstretched gesture pointing to the upper right corner of the screen in front of the source network terminal. The visual acquisition layer captures this gesture image and identifies it as a left-click operation. The parsing module generates a command data packet containing the following information: type field 0x01, X coordinate field 0x0780, Y coordinate field 0x0168, timestamp field 0x61A2B3C4, and checksum field 0x5F2A. This data packet is forwarded to the target network via a cross-network relay system. After receiving the packet, the target network's execution layer maps the coordinates 1920,360 to the corresponding position in the target screen resolution, ultimately executing a left-click operation in the target system. The latency from gesture recognition to target execution is controlled within 200 milliseconds, with a transmission error rate of less than 0.1%, meeting the real-time and reliability requirements of cross-network bidirectional interaction.
[0154] A second aspect of the present invention provides a vision-based cross-network interaction system, comprising:
[0155] The first unit is used to acquire the operator's visual interaction sequence through an image acquisition device in the source network environment, and to jointly encode the spatial morphological features and temporal evolution features in the visual interaction sequence to construct a multimodal feature representation.
[0156] The second unit is used to perform cross-domain semantic alignment of the multimodal feature representation based on the heterogeneity of the source network environment and the target network environment. By establishing a multi-level mapping relationship between the visual semantic space and the target network control space, the multimodal feature representation is deconstructed into a hierarchical control instruction set.
[0157] The third unit is used to classify the instructions for security according to the execution priority of the hierarchical control instruction set, transmit the hierarchical control instruction set to the target network environment through a non-physical connection transmission medium, and construct temporary trust credentials based on behavior pattern verification at the security domain boundary node.
[0158] The fourth unit is used to determine the execution conditions and implement control operations in the target network environment based on the temporary trust credential, and at the same time extract the state semantic features containing the operation response timing and resource occupancy status, and encode the state semantic features into an abstract state descriptor.
[0159] The fifth unit is used to convert the abstract state descriptor into a visual representation that can be perceived by the source network environment through reverse semantic mapping, and after spatiotemporally aligning it with the multimodal feature representation, construct a bidirectional mapping relationship graph between operation intention and execution result, and use the semantic deviation evolution law captured by the bidirectional mapping relationship graph to perform incremental optimization of the mapping parameters of cross-domain semantic alignment.
[0160] A third aspect of the present invention provides an electronic device, comprising:
[0161] processor;
[0162] Memory used to store processor-executable instructions;
[0163] The processor is configured to invoke instructions stored in the memory to execute the aforementioned method.
[0164] A fourth aspect of the present invention provides a computer-readable storage medium having stored thereon computer program instructions that, when executed by a processor, implement the aforementioned method.
[0165] This invention can be a method, apparatus, system, and / or computer program product. The computer program product may include a computer-readable storage medium having computer-readable program instructions loaded thereon for performing various aspects of the invention.
[0166] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some or all of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.
Claims
1. A vision-based cross-network interaction method, characterized in that, include: In the source network environment, the operator's visual interaction sequence is acquired through an image acquisition device. The spatial morphological features and temporal evolution features in the visual interaction sequence are jointly encoded to construct a multimodal feature representation. Based on the heterogeneity of the source network environment and the target network environment, cross-domain semantic alignment is performed on the multimodal feature representation. By establishing a multi-level mapping relationship between the visual semantic space and the target network control space, the multimodal feature representation is deconstructed into a hierarchical control instruction set. The instructions are classified for security based on their execution priority within the hierarchical control instruction set. The hierarchical control instruction set is then transmitted to the target network environment via a non-physical connection transmission medium. Temporary trust credentials are constructed at the security domain boundary node based on behavioral pattern verification, including: The execution priority sequence of the hierarchical control instruction set is analyzed and modeled in depth. The expected intensity of resource consumption and sensitivity to state changes of each instruction are dynamically evaluated through machine learning algorithms. The instruction risk score is calculated based on the expected intensity of resource consumption and the sensitivity to state changes. The instructions are divided into different security access levels according to the instruction risk score. Differentiated transmission encapsulation structures are constructed for instructions at different security access levels. The differentiated transmission encapsulation structures include adaptive redundancy check codes. The layered control instruction sets encapsulated in differentiated manner are transmitted to the target network environment through a non-physical connection transmission medium. A dynamic timing feature library for instruction transmission is established based on the non-physical connection transmission medium. The received instruction sequence is parsed at the security domain boundary node of the target network environment. Based on the dynamic temporal feature library, the verification result of the adaptive redundancy check code and the logical dependency relationship between instructions are used to generate a behavior pattern feature vector. The behavior pattern feature vector is then adaptively matched with a pre-established historical normal behavior benchmark library. The deviation metric between the behavior pattern feature vector and each benchmark pattern in the historical normal behavior benchmark library is calculated. When the minimum value of the deviation metric is lower than the trust threshold, dynamic optimization modeling is performed based on the permission configuration template in the corresponding benchmark mode, and temporary trust credentials are generated by combining the adaptive redundancy check code verification result in the current instruction sequence with the instruction risk score. In the target network environment, the execution conditions are determined and control operations are implemented based on the temporary trust credential. At the same time, state semantic features containing operation response timing and resource occupancy status are extracted and encoded into abstract state descriptors. The abstract state descriptor is converted into a visual representation perceptible to the source network environment through reverse semantic mapping. After spatiotemporal alignment with the multimodal feature representation, a bidirectional mapping relationship graph of operation intention and execution result is constructed. The semantic deviation evolution law captured by the bidirectional mapping relationship graph is used to incrementally optimize the mapping parameters of cross-domain semantic alignment.
2. The method according to claim 1, characterized in that, The spatial morphological features and temporal evolution features in the visual interaction sequence are jointly encoded to construct a multimodal feature representation, including: Spatial domain analysis is performed on each frame of the visual interaction sequence to extract spatial morphological feature vectors containing the topological structure of the operator's limb posture, the fine movement form of the hand, and the orientation angle of the head; temporal correlation analysis is performed on the spatial morphological feature vectors corresponding to multiple consecutive frames to capture the change trajectory and evolution pattern of the spatial morphological feature vectors in the time dimension, and a temporal evolution feature sequence is constructed. The spatial morphological feature vector is semantically aligned with the temporal evolution feature sequence. By establishing the association constraint between the spatial static state and the temporal dynamic process, a multimodal feature representation that integrates spatial semantics and temporal semantics is generated.
3. The method according to claim 1, characterized in that, Based on the heterogeneity between the source and target network environments, cross-domain semantic alignment is performed on the multimodal feature representations. By establishing a multi-level mapping relationship between the visual semantic space and the target network control space, the multimodal feature representations are deconstructed into a hierarchical control instruction set, including: Obtain the visual interaction capability description and control execution capability description of the source network environment, analyze the differences between the two in terms of operation granularity, response latency constraints and security policy restrictions, and construct a heterogeneous characteristic model; Based on the heterogeneous characteristic model, the visual semantic space is divided into multiple semantic levels. Each semantic level corresponds to a control capability with a different level of abstraction in the target network control space. A mapping function is established between each semantic level and the corresponding control capability. According to the mapping function, the operational intent semantics in the multimodal feature representation is decomposed into control semantic units that match the execution capability of the target network. Based on the multi-level mapping relationship, the control semantic units are organized into a hierarchical control instruction set according to execution dependencies and resource scheduling priorities.
4. The method according to claim 1, characterized in that, The expected resource consumption intensity and state change sensitivity of each instruction are dynamically evaluated using machine learning algorithms. Based on the expected resource consumption intensity and the state change sensitivity, an instruction risk score is calculated, including: An instruction behavior feature vector is constructed from the historical execution records of the hierarchical control instruction set. The instruction behavior feature vector contains the actual resource consumption data and the state change impact data of each instruction. The actual resource consumption data and the state change impact data are combined into a training sample set. A deep neural network model is constructed based on the training sample set. The deep neural network model learns the correlation between the actual resource occupancy data and the state change impact data to obtain the resource occupancy assessment weight and the state change assessment weight. A multi-layer feature mapping matrix is established for the instruction to be evaluated in the hierarchical control instruction set. The multi-layer feature mapping matrix is input into the deep neural network model. The expected intensity of resource consumption of the instruction is calculated using the resource consumption evaluation weight, and the sensitivity of state change of the instruction is calculated using the state change evaluation weight. The expected intensity of resource consumption and the sensitivity to state changes are input into the deep neural network model to generate an instruction risk score.
5. The method according to claim 1, characterized in that, In the target network environment, the execution conditions are determined and control operations are implemented based on the temporary trust credential. Simultaneously, state semantic features containing operation response timing and resource occupancy status are extracted, and these state semantic features are encoded into an abstract state descriptor, including: The temporary trust credential is parsed in the execution engine of the target network environment to determine the execution conditions; the execution of control instructions is triggered based on the execution conditions; a dual-stream feature network is constructed during the instruction execution process; and parallel feature extraction is performed on the operation response timing and resource occupancy status based on the dual-stream feature network. State semantic features are generated based on the features output by the dual-stream feature network, and the state semantic features are converted into abstract state descriptors.
6. The method according to claim 1, characterized in that, The abstract state descriptor is converted into a visual representation perceptible to the source network environment through inverse semantic mapping. After spatiotemporal alignment with the multimodal feature representation, a bidirectional mapping relationship graph between operation intent and execution result is constructed, including: The abstract state descriptor is hierarchically decomposed, and the decomposed information is organized into state attribute identifiers and state value descriptions. The state attribute identifiers and state value descriptions are used as input data for reverse mapping. Based on the state attribute identifier and the state value description, a reverse semantic mapping transformation is performed to map the state attribute identifier to the corresponding visual element type label in the source network environment, and the state value description to the corresponding visual element appearance feature parameter in the source network environment. A perceptible visual representation of the source network environment is generated through the visual element type label and the visual element appearance feature parameter. Based on the visual representation combined with the operation time marker and operation area coordinates in the multimodal feature representation, bidirectional alignment processing is performed. The spatial position coordinates in the visual representation are spatially aligned with the operation area coordinates, and the execution completion time corresponding to the abstract state descriptor is temporally aligned with the operation time marker, generating spatiotemporally aligned visual representation and spatiotemporally aligned operation feature pairing data. A bidirectional mapping relationship map between operation intention and execution result is established by pairing the spatiotemporally aligned visual representation with the spatiotemporally aligned operation feature data.
7. A vision-based cross-network interaction system for implementing the method of any one of claims 1-6, characterized in that, include: The first unit is used to acquire the operator's visual interaction sequence through an image acquisition device in the source network environment, and to jointly encode the spatial morphological features and temporal evolution features in the visual interaction sequence to construct a multimodal feature representation. The second unit is used to perform cross-domain semantic alignment of the multimodal feature representation based on the heterogeneity of the source network environment and the target network environment. By establishing a multi-level mapping relationship between the visual semantic space and the target network control space, the multimodal feature representation is deconstructed into a hierarchical control instruction set. The third unit is used to classify the instructions for security according to the execution priority of the hierarchical control instruction set, transmit the hierarchical control instruction set to the target network environment through a non-physical connection transmission medium, and construct temporary trust credentials based on behavior pattern verification at the security domain boundary node. The fourth unit is used to determine the execution conditions and implement control operations in the target network environment based on the temporary trust credential, and at the same time extract the state semantic features containing the operation response timing and resource occupancy status, and encode the state semantic features into an abstract state descriptor. The fifth unit is used to convert the abstract state descriptor into a visual representation that can be perceived by the source network environment through reverse semantic mapping, and after spatiotemporally aligning it with the multimodal feature representation, construct a bidirectional mapping relationship graph between operation intention and execution result, and use the semantic deviation evolution law captured by the bidirectional mapping relationship graph to perform incremental optimization of the mapping parameters of cross-domain semantic alignment.
8. An electronic device, characterized in that, include: processor; Memory used to store processor-executable instructions; The processor is configured to invoke instructions stored in the memory to execute the method according to any one of claims 1 to 6.
9. A computer-readable storage medium having computer program instructions stored thereon, characterized in that, When the computer program instructions are executed by the processor, they implement the method described in any one of claims 1 to 6.
Citation Information
Patent Citations
Cross-network safe pushing method and system
CN116996562A
Wireless network interaction control method and system based on video analysis
CN118555159A