A cloud computing engine access control method based on a remote desktop protocol
Patent Information
- Application Number
- CN202611130749.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-07-29
- Publication Date
- 2026-09-29
AI Technical Summary
[0004]但这类基于静态规则的管控方法难以适配云端资源高度动态的运行环境,当云端资源实例发生创建、热迁移或销毁等状态变更时,已建立的授权会话无法实时感知底层资源的状态变化,容易因资源销毁后会话权限未及时同步回收形成权限越界风险,也易因资源跨节点迁移后权限上下文丢失产生访问孤岛问题,难以实现访问控制策略与资源状态的动态联动调整
1.本发明通过对远程桌面协议传输的屏幕图像帧进行多层次视觉特征提取和语义解析,实现了从协议代理层无法感知的像素级数据向应用级语义信息的跨越转化,将无法识别应用类型的协议转发代理升级为具备应用感知能力的智能访问控制节点;
Smart Images

Figure CN122845266A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of computer technology, specifically a cloud computing engine access control method based on a remote desktop protocol. Background Technology
[0002] Access control for cloud computing engines based on remote desktop protocols is a core technology supporting secure and compliant access to cloud computing resources. With the widespread adoption of elastic multi-tenant cloud computing models, cloud computing resources exhibit typical operating characteristics such as dynamic scaling of instances, frequent tenant changes, and rapid start-up and shutdown of instances. The ability to synchronize access control policies with the underlying resource status in real time has become a key indicator for measuring the security management level of cloud platforms.
[0003] Existing cloud computing engine remote access control generally adopts a predefined permission matrix or a role-based static permission model. It completes the determination of user access permissions and session admission control through pre-configured identity authentication rules and permission mapping relationships. In static operation and maintenance scenarios where the resource status is relatively stable, it has the advantages of clear logic, simple deployment, and convenient audit traceability.
[0004] However, this type of static rule-based management method is difficult to adapt to the highly dynamic operating environment of cloud resources. When cloud resource instances undergo state changes such as creation, hot migration, or destruction, the established authorized sessions cannot perceive the state changes of the underlying resources in real time. This can easily lead to the risk of permission overstepping due to the failure to promptly reclaim session permissions after resource destruction. It can also easily cause access silos due to the loss of permission context after resource migration across nodes. It is difficult to achieve dynamic linkage adjustment between access control policies and resource status.
[0005] Therefore, the present invention provides a cloud computing engine access control method based on a remote desktop protocol. Summary of the Invention
[0006] In order to overcome the shortcomings of the prior art, at least one technical problem raised in the background art is solved.
[0007] The technical solution adopted by this invention to solve its technical problem is: a cloud computing engine access control method based on a remote desktop protocol, comprising the following steps: Step 1: Protocol parsing and image frame acquisition. The remote desktop protocol is parsed and processed, and screen image frame data during the remote desktop session is captured as samples to be analyzed. The acquired image frames are preprocessed by color space conversion and geometric parameter normalization. Step 2: Visual feature extraction and semantic parsing, extracting multi-level visual feature vectors from the preprocessed image frames and performing feature-level fusion processing; Step 3: Application type identification and window hierarchy determination. Based on the fused comprehensive feature vector, an application type classifier is constructed, and the application type determination result and spatial location bounding box coordinates are output. Step 4: Loading session-level permission policies and building contexts. Load the session-level permission policy file and build the permission control context tree structure based on the user identity and tenant identity. Step 5: Application-level access control decision generation and execution, matching the application type identification results with the access control context and executing the corresponding control instructions; Step Six: Dynamic feedback and policy optimization for access control decisions, real-time monitoring of decision execution status, and triggering automatic update process of permission policies based on evaluation results.
[0008] Preferably, the protocol parsing includes protocol header parsing, security layer parsing, and instruction stream parsing; Image frame acquisition employs a dual-buffering mechanism. The main buffer area continuously receives protocol data streams to update the pixel matrix, while the auxiliary buffer area saves a pixel snapshot of the previous complete frame for incremental detection.
[0009] Preferably, the color space conversion transforms the pixel matrix from the original color space to the target analysis color space; The conversion process includes calculating the saturation component, calculating the hue component, and determining the luminance component; The geometric parameter normalization standardizes the resolution, pixel depth, and color channel mapping relationship of the image frame.
[0010] Preferably, the multi-layered visual feature vector includes three dimensions: low-level texture features, mid-level shape features, and high-level semantic features; The underlying texture features are calculated using a local binary pattern operator, and the gray-level change patterns of local regions of the image are statistically analyzed to form texture descriptors. The mid-level shape features extract contour boundary information from the image through edge detection operators, and perform multi-scale geometric analysis on the contour boundaries to obtain shape feature vectors; High-level semantic features are predicted based on a pre-trained multi-label image classification network.
[0011] Preferably, the application type classifier adopts a deep convolutional neural network architecture, which includes a feature encoding layer, a spatial attention layer, and a classification output layer; The feature encoding layer uses multiple convolutional operations to abstract and represent features layer by layer. The spatial attention layer generates an attention mask based on the importance distribution of spatial locations in the image, which enhances the features of key regions while suppressing background interference regions. The classification output layer outputs the application type determination result and the spatial positioning bounding box coordinates.
[0012] Preferably, the session-level permission policy file includes four policy elements: user role attributes, resource access scope, application usage whitelist, and application operation blacklist. User role attributes include the user's security level and functional positioning in the cloud platform; resource access scope defines the list of cloud computing resource instances that the user can access; The application whitelist includes a set of application types that the user is authorized to use; The application operation blacklist includes a set of high-risk operation types that users are prohibited from performing.
[0013] Preferably, the access control context is organized in a tree structure, with the root node being the user identity identifier, and the child nodes being the resource scope node, the application whitelist node, and the operation blacklist node in sequence.
[0014] Preferably, the access control decision matching operation follows the following order: Determine if the application type exists in the application whitelist set; if not, generate an access denial decision. If it exists, determine whether the operation type triggers the operation blacklist rule; if it does, generate a decision to restrict the operation. If the blacklist rule is not triggered, an access permission decision is generated. The control command package includes command types such as session termination command, window minimize command, application freeze command, and operation interception command.
[0015] Preferably, the dynamic feedback and strategy optimization constructs a decision-making effect evaluation model by collecting user behavior response data and session state change data; The decision-making effectiveness evaluation model adopts the reward function design in the reinforcement learning framework, taking the improvement of user compliance rate and resource security indicators as positive rewards, and the frequency of security incidents and false blocking rate as negative penalties; The automatic policy updates include adding or deleting application whitelist types, optimizing rules for operating blacklists, and adjusting user role permissions.
[0016] Preferably, an application recognition inference engine is deployed as a pre-processing unit for access control at the protocol proxy layer of the remote desktop protocol; The application recognition and reasoning engine runs as an independent process and interacts with the protocol agent main process through inter-process communication mechanisms; The inference engine uses a batch processing mode for inference calculations, accumulating continuously acquired image frames into a fixed batch size before performing inference calculations uniformly. The generation of application-level access control decisions adopts a hybrid decision architecture that combines a rule engine and a machine learning model. The rule engine is responsible for handling deterministic policy terms, while the machine learning model is responsible for handling boundary scenarios and ambiguous scenarios. When the rule engine and the machine learning model make conflicting decisions, the decision result of the rule engine is given priority.
[0017] The beneficial effects of this invention are as follows: 1. This invention achieves a leapfrog transformation from pixel-level data that cannot be perceived by the protocol proxy layer to application-level semantic information by performing multi-level visual feature extraction and semantic parsing on screen image frames transmitted by remote desktop protocol, and upgrades the protocol forwarding proxy that cannot identify application type into an intelligent access control node with application perception capability. By constructing an access control context that includes user role attributes, resource access scope, application usage whitelist, and application operation blacklist, the granularity of access control on the cloud platform side is refined from coarse-grained session-level permissions to application-level operation-level permissions. By intelligently matching application type identification results with the access control context and executing differentiated control instructions, proactive defense capabilities are achieved, enabling prediction and real-time blocking based on application type before user operations occur. Furthermore, by establishing a dynamic feedback and policy optimization mechanism for access control decisions, access control policies can be continuously optimized and evolved based on user behavior data and environmental status data.
[0018] 2. This invention, through the generation and execution mechanism of application-level access control decisions, can effectively distinguish access requests from different application types such as file managers, system terminals, and download tools, and implement differentiated permission control strategies for each application type. This breaks the limitations of the coarse-grained permission mode of the traditional method of all-allow or all-prohibit, enabling the cloud platform to achieve fine-grained control over the application layer operations within the remote desktop session without intruding on the virtual machine operating system, thereby reducing the security and compliance risks of enterprise cloud operations and maintenance.
[0019] 3. This invention optimizes GPU resource utilization by deploying an independently running application identification and inference engine at the protocol proxy layer, adopts a batch processing inference calculation mode, optimizes the loading and query efficiency of permission policies through a hierarchical indexing mechanism, balances policy determinism and boundary scenario adaptability through a hybrid decision architecture of rule engine and machine learning model, and achieves automated policy tuning and updating through a reinforcement learning framework. The overall system completes the technical upgrade from coarse-grained control to fine-grained control while ensuring the real-time performance of access control. Attached Figure Description
[0020] The invention will now be further described with reference to the accompanying drawings.
[0021] Figure 1This is a flowchart of a cloud computing engine access control method based on a remote desktop protocol in this invention. Figure 2 This is a schematic diagram of the interaction relationship and data flow between the construction of session-level permission policy context and the generation of application-level access control decisions in this invention. Detailed Implementation
[0022] To make the technical means, creative features, objectives and effects of this invention easier to understand, the invention will be further described below in conjunction with specific embodiments.
[0023] Example 1: like Figure 1 and Figure 2 As shown, this embodiment of the invention provides a cloud computing engine access control method based on a remote desktop protocol. This method deploys an application recognition and inference engine at the protocol proxy layer of the cloud platform. By performing multi-level visual feature extraction and semantic parsing on the screen image frames transmitted by the remote desktop protocol, it achieves a leapfrog transformation from pixel-level data that cannot be perceived by the protocol proxy layer to application-level semantic information. Then, based on the permission control context, it generates and executes application-level access control decisions and establishes a dynamic feedback and policy optimization mechanism to achieve continuous optimization and evolution of access control policies.
[0024] The method includes:
[0025] Step 1: Perform remote desktop session protocol parsing and image frame acquisition. After receiving protocol data packets from remote clients or cloud virtual machines, the remote desktop protocol proxy module performs protocol layer parsing. Taking the RDP protocol as an example, the protocol proxy module performs multi-stage parsing on the received data stream: the first stage is to parse the protocol header and extract the protocol version identifier, channel identifier, timestamp field and sequence number information; The second stage involves security layer parsing. For RDP sessions with TLS encryption enabled, the decrypted data contains a composite instruction stream of Fast-Path or Slow-Path packets. The third stage involves instruction stream parsing, which decomposes different types of protocol data units, such as drawing command sequences, bitmap transmission instructions, raster operation instructions, and window state update messages.
[0026] Based on protocol parsing, the image frame acquisition unit extracts screen image frame data from the protocol data stream as samples to be analyzed. The image frame data originates from the graphics output channel of the remote desktop protocol, and its data structure contains three dimensions: The first dimension is pixel matrix data, which records the complete pixel value information of the current screen. The pixel format may be 24-bit RGB, 32-bit ARGB or 16-bit 565 format depending on the session configuration. The resolution of the pixel matrix is consistent with the display resolution of the remote desktop. The second dimension is the drawing instruction sequence, which records all drawing operation commands from the previous frame to the current frame, including rectangle fill commands, bitmap block transfer commands, text rendering commands, and curve drawing commands, etc. The third dimension is window state information, which records attribute data such as the position coordinates, size parameters, Z-order arrangement, activation status, and window class name of all windows on the current screen.
[0027] The image frame acquisition unit uses a dual-buffering mechanism to achieve complete image frame capture: the main buffer area continuously receives the protocol data stream and updates the pixel matrix; the auxiliary buffer area saves a pixel snapshot of the previous complete frame for incremental detection. When a screen update signal is detected, the data in the auxiliary buffer is copied as the target frame data, and the current data in the main buffer is pushed to the analysis queue. For RDP sessions that support Fast-Path compression, the image frame acquisition unit also integrates an RDP decompression module, which decompresses the compressed bitmap block data in real time and merges it into the pixel matrix.
[0028] The acquired image frames are preprocessed, specifically including two processing steps: color space conversion and geometric parameter normalization. The color space conversion module converts the pixel matrix from the original color space to the target analysis color space. When the original data is in RGB color space, it is converted to HSV color space to facilitate subsequent texture feature extraction. When the original data is in YUV color space, it is first converted to RGB and then to HSV.
[0029] Color space conversion uses the following calculation process: calculating the three color components. maximum value and minimum value Calculate the three components in sequence: Luminance component Take the maximum value directly (Value range is 0~1 or 0~255, depending on the quantization depth); Saturation component according to Calculate (when) When the saturation is 0, it reflects the purity of the color; Tone Components Calculate segmented based on the location of the largest channel: when the red channel is at its maximum... , When the green channel is at its maximum , When the blue channel is at its maximum .
[0030] Calculated Multiplying the value by 60° will give you the hue angle within the range of 0° to 360°.
[0031] The geometric parameter normalization calibration module standardizes the resolution, pixel depth, and color channel mapping of image frames. Resolution normalization uses a bilinear interpolation algorithm to uniformly scale input images of different resolutions to a preset analysis resolution, which is configured as 384×256 pixels. Pixel depth normalization converts input data of different pixel depths into 8-bit grayscale or 24-bit true color format. When color information needs to be retained for subsequent analysis, 24-bit RGB format is used. When only texture analysis is needed, it is converted to 8-bit grayscale format to reduce computational complexity. Color channel mapping calibration compensates for color deviations caused by different monitor profiles. The calibration parameters include the gamma curve parameters and white point color temperature parameters of each channel, which are obtained through standard color chart calibration.
[0032] The preprocessed image frame data is encapsulated into standardized analysis data units, which contain five fields: preprocessed pixel matrix, original resolution parameters, color space identifier, acquisition timestamp, and session context identifier. The analysis data unit transmits the data to the application recognition module for analysis and processing via a message queue. The message queue adopts an asynchronous non-blocking mode to ensure that the real-time performance of image frame acquisition is not affected by the processing speed of downstream modules. The maximum buffer depth of the queue is configured to 32 frames to prevent memory overflow.
[0033] Step 2: Perform visual feature extraction and semantic parsing of image frames. Extract multi-level visual feature vectors from the preprocessed image frames. These feature vectors contain three dimensions: low-level texture features, mid-level shape features, and high-level semantic features. The three dimensions of features describe the visual attributes of an image at different levels, and together they constitute the feature representation space required for application type recognition.
[0034] The underlying texture feature extraction module uses the local binary mode operator for computation. The local binary pattern operator takes each pixel in the image as the center and selects the pixels in its 8-neighborhood as the comparison object. It compares the value of the center pixel with the value of the neighboring pixels. If the value of the neighboring pixels is greater than or equal to the value of the center pixel, it is marked as 1; otherwise, it is marked as 0. This generates an 8-bit binary number as the local binary pattern code of that pixel.
[0035] The distribution of local binary pattern encoding histograms across the entire image is statistically analyzed to form a texture descriptor vector. The formula for calculating the local binary pattern encoding value is as follows: ,in The local binary pattern encoded value (decimal integer) of the center pixel is used to characterize the local texture features of the pixel and its surrounding neighborhood. Summation symbol This indicates that the weighted values of the 8 neighboring pixels are accumulated; For the first The binarization comparison result of the neighboring pixels, when the gray value of the neighboring pixels... Greater than or equal to the gray value of the center pixel hour ,otherwise ; For the first The binary weights corresponding to each neighboring pixel are encoded into 8-bit binary numbers according to their positions, and finally the LBP encoded values in the range of 0 to 255 are obtained. Neighboring pixels are numbered counterclockwise, starting from the top right corner. Corresponding to the top right corner, (Corresponding to the top, and so on), this encoded value serves as the texture feature descriptor for that pixel. By statistically analyzing the distribution of the LBP encoded histogram of the entire image, the texture feature vector of the image can be constructed.
[0036] The distribution histogram of LBP mode values for all pixels is statistically analyzed. This histogram vector is the texture descriptor, which has a dimension of 256. Each dimension represents the frequency of occurrence of the corresponding mode value. To enhance the adaptability of texture features to multi-scale texture changes, this invention adopts a multi-radius sampling strategy, calculating local binary mode codes at three scales with radii of 1, 2, and 3 respectively. The histogram vectors of the three scales are concatenated to form the final texture descriptor, which has a dimension of 768.
[0037] The mid-level shape feature extraction module extracts contour boundary information from the image using edge detection operators; The edge detection operator is implemented using the Canny algorithm, and the specific processing flow includes four steps: Gaussian smoothing filtering, gradient calculation, non-maximum suppression, and double-threshold edge connection. Gaussian smoothing filtering uses a 5×5 Gaussian convolution kernel with a sigma parameter configured to 1.4 to eliminate interference from image noise on edge detection.
[0038] In the gradient calculation stage, the horizontal gradient is obtained by convolving the image with the Sobel horizontal gradient operator and the Sobel vertical gradient operator respectively. and vertical gradient .
[0039] gradient magnitude The calculation formula is: This represents the total intensity of the grayscale change at that pixel; gradient direction. The calculation formula is: ,in It is a two-parameter arctangent function, which can be determined according to... and The sign of the direction angle is automatically determined to determine the quadrant in which it is located. The range of values is or The unit is radians, representing the direction of the fastest change in grayscale value for that pixel. If angles are required, [the following can be used]: Multiply Convert to degree.
[0040] The non-maximum suppression stage refines the gradient magnitude image. For each pixel, only the pixel with the largest local gradient magnitude in its gradient direction is retained as a candidate edge point, while other candidate points are suppressed. In the dual-threshold edge connection stage, two thresholds are set: a high threshold of 0.15 and a low threshold of 0.05. Pixels with gradient magnitudes higher than the high threshold are marked as strong edge points, while pixels between the high and low thresholds are marked as weak edge points. If a weak edge point has an 8-neighborhood connectivity with a strong edge point, it is retained as a true edge point; otherwise, it is discarded.
[0041] Multi-scale geometric analysis is performed on the edge detection results to obtain shape feature vectors; A multi-scale edge detection strategy is adopted, performing edge detection at three scales with sigma values of 1.0, 2.0, and 3.0 respectively. The binary edge images at the three scales are then weighted and fused to obtain a multi-scale edge map, with the weight ratio configured as 0.5:0.3:0.2. Contour extraction processing is performed on the multi-scale edge map, and the Suzuki algorithm is used to trace the boundaries of connected regions, extracting the coordinate sequences of all closed contours. For each contour, calculate the following shape descriptors: contour length (total number of pixels along the contour boundary), contour bounding area (total number of pixels inside the contour), contour compactness (4π × area / perimeter squared), and contour roundness (4π × area / (equivalent perimeter)). ^2 1. Outline rectangularity (area / area of minimum circumscribed rectangle) 2. Outline eccentricity (minor axis length / major axis length); The shape descriptors of all contours are statistically summarized, and the mean, variance, and extreme values of each descriptor are calculated as a global shape feature vector with a dimension of 48.
[0042] The high-level semantic feature extraction module predicts outputs based on a pre-trained multi-label image classification network. This network uses ResNet-50 as its backbone architecture, and its structure includes a 7×7 convolutional layer, 16 residual blocks, and a global average pooling layer. The residual blocks employ a bottleneck structure design, with each residual block consisting of a three-layer structure: a 1×1 convolutional layer, a 3×3 convolutional layer, and a 1×1 convolutional layer. The last layer of the network is the classification output layer, which uses the sigmoid activation function to support multi-label classification tasks; The network is pre-trained on the ImageNet dataset to obtain general image feature representation capabilities, and then fine-tuned using transfer learning on the desktop screenshot dataset collected in this invention. The desktop screenshot dataset contains over 50,000 labeled images, covering seven main application types: system desktop, file manager, command line terminal, web browser, database client, remote connection tool, and download tool. Each application type contains multiple subcategories and various interface layout variations.
[0043] The specific process of high-level semantic feature extraction is as follows: the preprocessed image frame data is scaled to 224×224 pixels and normalized before being input into the network; the forward propagation process of the network sequentially passes through each convolutional layer and residual block for feature extraction, and after the feature map is output from the last residual block, a global average pooling layer is connected to obtain a 2048-dimensional feature vector; the feature vector is reduced to 512 dimensions through a fully connected layer and then connected to two classification heads respectively; The first classification head outputs a probability distribution vector for the application type, with a dimension of 7, corresponding to the 7 main application types; The second classification header outputs a probability distribution vector of the window interface type, with an output dimension of 12, corresponding to 12 interface categories; Each component of the probability distribution vector has a value range of 0, 1, representing the confidence level of the existence of that category.
[0044] The visual feature vectors from the three dimensions mentioned above are fused at the feature level. The fusion method employs a cross-dimensional attention weight mechanism, dynamically allocating fusion weights based on the discriminative power of each feature channel. The implementation architecture of the cross-dimensional attention weight mechanism includes the following components: The texture feature attention module receives a 768-dimensional texture descriptor vector and outputs a 512-dimensional attention weight vector. This module consists of two fully connected layers. The first layer maps the 768-dimensional vector to 128-dimensional vectors, and the second layer maps the 128-dimensional vectors to 512-dimensional vectors. The ReLU activation function is used between the two layers, and the output layer uses the sigmoid activation function. The shape feature attention module receives a 48-dimensional shape feature vector and outputs a 256-dimensional attention weight vector. This module also adopts a two-layer fully connected structure. The semantic feature attention module receives a 512-dimensional semantic feature vector and outputs a 512-dimensional attention weight vector. The weight vectors of the three attention modules are normalized using a softmax operation to ensure that the sum of the weights in the three dimensions is 1. The fusion calculation process is as follows: the original feature vectors of the three dimensions are weighted and summed to obtain three weighted feature vectors. The three weighted feature vectors are then concatenated to form the final fused feature vector. The fused feature vector has a dimension of 1772 (768+512+512=1792 minus the attention dimension bias correction to 1772 dimensions).
[0045] Step 3: Identify application types and determine window levels, and construct an application type classifier based on the fused comprehensive feature vector. This classifier adopts a deep convolutional neural network architecture, which includes three functional modules: a feature encoding layer, a spatial attention layer, and a classification output layer. The feature encoding layer receives the fusion results of multi-level visual features and performs layer-by-layer abstraction of the features through multi-level convolution operations; The feature encoding layer is designed to contain four convolutional blocks, each containing two 3×3 convolutional layers, one BatchNorm layer, and one ReLU activation layer. The number of output channels for the four convolutional blocks are 64, 128, 256, and 512, respectively. Each convolutional block is downsampled using a 2×2 max-pooling layer to expand the receptive field; The feature encoding layer receives a 1772-dimensional fused feature vector, reshapes it into a multi-channel feature map for processing, and outputs a 512-channel abstract feature map. The spatial size of the feature map is 1 / 8 of the input image analysis resolution.
[0046] The spatial attention layer generates an attention mask based on the importance distribution of spatial locations in the image, which enhances the features of key regions while suppressing background interference regions. The spatial attention layer is implemented using the Self-Attention mechanism, and the computation process includes three steps: The first step involves projecting the 512-channel feature map into the three embedding spaces of Query, Key, and Value using 1×1 convolutions to obtain three feature matrices of Q, K, and V, with a projection dimension of 64. The second step is to calculate the dot product similarity matrix between the Query and the Key. The dimension of the similarity matrix is the number of spatial locations multiplied by the number of spatial locations. The third step involves scaling and normalizing the similarity matrix, then multiplying it by the Value matrix to obtain a self-attention weighted feature map. After residual connection and LayerNorm processing between the attention-weighted feature map and the original feature map, a nonlinear transformation is performed through a feedforward neural network containing two fully connected layers with a hidden dimension of 2048. The output of the spatial attention layer is a position-weighted feature representation, in which spatial regions highly relevant to the current application type receive greater response weights.
[0047] The classification output layer outputs the application type determination result based on the enhanced feature representation, and also outputs the spatial location bounding box coordinates of the application in the current screen image; The network structure of the classification output layer contains two branches: The first branch is the category classification branch, which contains a global average pooling layer and a fully connected classification layer. The output dimension of the classification layer is 7, corresponding to 7 application types: system desktop, file manager, command line terminal, web browser, database client, remote connection tool, and download tool. The classification layer uses the softmax activation function to output the probability distribution of each category. The second branch is the bounding box regression branch, which contains two parallel fully connected layers that predict the center coordinates and width and height dimensions of the bounding box, respectively. The bounding box uses (cx, cy, w, h) The format indicates that the coordinate values are normalized scale values relative to the image size.
[0048] Independent recognition confidence thresholds are set for each category to ensure accuracy in judgment.
[0049] To further improve the reliability of application type recognition in complex scenarios, this invention incorporates a confidence assessment and hierarchical processing mechanism after the classification output layer. This mechanism comprises the following three levels: Level 1: Multimodal Assisted Verification. When the difference between the highest and second-highest probability values output by the application type classifier is less than a preset discrimination margin threshold (default setting is 0.15), it indicates that the current image frame has ambiguity in visual features, and the system automatically activates the auxiliary verification channel. The auxiliary verification channel includes: The OCR text recognition channel extracts text information from the window title bar, status bar, and menu bar through optical character recognition and performs string matching with a preset application name feature library. The window attribute analysis channel extracts the window class name, process name, and window style attributes from the window status update message of the remote desktop protocol and compares them with the preset application feature library. When the visual classification result matches the recognition result of the auxiliary verification channel, the recognition result is adopted; when the two do not match, the system adopts the recognition result with higher confidence and marks the sample as a sample to be reviewed.
[0050] The second level: Unknown category detection and incremental learning mechanism. When the probability values of all categories output by the application type classifier are lower than their respective confidence thresholds, the system determines that the current application is an unknown category. For unknown categories, the system does not directly reject access, but instead: Record the feature vector, window attribute information and user operation context of the image frame, and store them in the unknown sample cache pool; Temporarily grant restricted access to the session (allowing only basic operations and prohibiting high-risk operations), and mark the session as an audit watch state; When the number of samples for the same application type in the unknown sample cache pool exceeds the preset threshold (50 samples by default), the system automatically triggers the incremental learning process. After performing cluster analysis and manual annotation on the accumulated samples in the background, the classification model is fine-tuned and updated to gradually include the application type in the recognition scope.
[0051] The third level: multi-frame temporal verification. To avoid misidentification caused by noise in a single frame image or instantaneous interface changes, the system performs temporal consistency verification on the recognition results of multiple consecutive frames (default 5 frames). When the recognition results are consistent in 5 consecutive frames, the result is used as the final judgment output. When the recognition result jumps between consecutive frames, the system does not output the judgment result temporarily and continues to collect subsequent frames until a stable recognition result is obtained or a timeout occurs (default 2 seconds). If a stable result is not obtained after the timeout, the system handles it according to the "unknown category" processing procedure.
[0052] The confidence thresholds for each application type are configured as follows: the confidence threshold for system desktop is set to 0.75, the confidence threshold for file manager is set to 0.80, the confidence threshold for command line terminal is set to 0.82, the confidence threshold for web browser is set to 0.78, the confidence threshold for database client is set to 0.85, the confidence threshold for remote connection tool is set to 0.80, and the confidence threshold for download tool is set to 0.78. When the probability values of each category output by the classification output layer do not exceed their corresponding thresholds, the determination result is an unknown application type. At this time, the access control decision adopts the default conservative strategy, i.e., denial of access.
[0053] The application type identification result also includes the spatial location bounding box coordinates of the application in the current screen image; The bounding box coordinates are output by the bounding box regression branch. When multiple application windows are detected to exist in the screen image at the same time, the bounding box regression branch outputs multiple bounding boxes and their corresponding application type labels. This invention employs a non-maximum suppression algorithm to filter duplicate detection results. The overlap IoU threshold for the filtering threshold is set to 0.5, and only detection results with an overlap lower than this threshold are retained as the final output.
[0054] Step 4: Load session-level permission policies and build context. Based on the user identity and associated tenant identity of the current remote desktop session, load the corresponding session-level permission policy file from the permission policy repository. The permission policy repository is organized using a hierarchical indexing mechanism. The policy repository is divided into primary partitions according to the tenant dimension, with each tenant having an independent data partition. The tenant partition identifier is a unique identifier of the tenant as the partition key value. Within each tenant partition, there are secondary partitions according to user role type. Common role types include administrators, developers, operations and maintenance personnel, business users, and visitors. Within each role partition, three levels of partitioning are created based on the policy's effective time period. The policy within each time segment can be configured independently to meet temporary permission adjustment needs.
[0055] With the hierarchical indexing mechanism, the loading and querying of permission policies can skip irrelevant data partitions and directly locate the target policy file; The query process is as follows: First, the first-level partition is determined based on the tenant identifier carried in the session, locating the data storage area of that tenant; second, the second-level partition is determined based on the role attributes of the session user, locating the policy storage space of that role; finally, the third-level partition is determined based on the current timestamp, locating the policy file for the currently effective time period. This three-level index query strategy reduces the time complexity of policy loading from O(n^2) to O(n^2). Reduced to O(1) This reduces the strategy loading latency.
[0056] The session-level permission policy file contains four policy elements: user role attributes, resource access scope, application whitelist, and application operation blacklist. The user role attributes include the current user's security level and job function in the cloud platform. This attribute includes a security level field (with a value range of 1 to 5, where 5 is the highest security level) and a job function type field (with values of administrator, developer, operations and maintenance personnel, business user, or visitor). The resource access scope defines the list of cloud computing resource instances that the user can access. This list is stored in the form of resource instance identifiers, with each identifier corresponding to a virtual machine or a container instance in the cloud platform.
[0057] The application whitelist includes the set of application types that the user is authorized to use. The whitelist is stored in the form of an application type code list, and the code corresponds to the application type number defined in step three. For example, a whitelist configuration of 2,4,5 indicates that users are allowed to use file managers, web browsers, and remote connection tools; The application operation blacklist includes a set of high-risk operation types that are prohibited from being performed by the user. Blacklist entries contain three fields: application type, operation behavior type, and trigger condition. For example, a blacklist configuration of {application type: 4, operation behavior: file download, trigger condition: file extension is .exe / dll / bat} means that the user is prohibited from downloading files with specific file extensions in a web browser.
[0058] The internal data structure of the policy file is organized in key-value pairs for easy retrieval. Keys include a combination of four dimensions: policy element type identifier, application type identifier, operation behavior identifier, and resource instance identifier. Key values correspond to the permission determination result or constraint parameters. The policy file is persistently stored using a binary serialization format to reduce storage space usage and parsing computational overhead. The binary serialization format uses Protocol Buffers for encoding, and the average size of the encoded policy file is about 2KB to 5KB, saving about 70% of storage space and about 60% of parsing time compared to the JSON text format.
[0059] Based on the four policy elements mentioned above, a tree structure for the access control context of the current session is constructed. The root node of this tree structure stores the user's identity and related attribute information. Under the root node, three child nodes are attached sequentially: the first-level child node is the resource scope node, storing a list of resource instances that the current user can access; the second-level child node is the application whitelist node, storing a set of application type codes authorized for use by the user; and the third-level child node is the operation blacklist node, storing a set of operation rules that the user is prohibited from executing. Each node in the tree structure is also associated with an effective timestamp and version number to support dynamic policy updates.
[0060] This tree structure serves as input parameters for the access control decision engine in the subsequent permission determination process. When performing permission determination, the decision engine accesses the tree structure using a depth-first traversal, retrieving user role attributes, resource scope, whitelist, and blacklist information sequentially from top to bottom, and then performing permission matching calculations. The tree structure is stored in memory in a compact binary format to improve traversal efficiency, and the pointer offsets between nodes are calculated and cached during policy loading to avoid runtime lookup overhead.
[0061] Step 5: Generate and execute application-level access control decisions, matching the application type identification results output in Step 3 with the access control context constructed in Step 4. The matching logic follows a specific priority order for decision-making. The decision-making process includes the following three decision levels: The first level of judgment is application type whitelist verification. The system determines whether the currently identified application type exists in the user's application usage whitelist set. Whitelist verification is implemented by iterating through the application type code list stored in the whitelist node, comparing the identified application type number with each code in the list one by one. If the comparison result is that the application type does not exist, an access denial decision is immediately generated and a blocking instruction is output. This blocking instruction triggers the protocol proxy layer to interrupt the current session and return an access denial message to the client. The denial decision has the highest priority and can quickly block unauthorized application access attempts.
[0062] The second level of judgment is the operation behavior blacklist verification. After the application type passes the whitelist verification, the system continues to determine whether the application's operation type has triggered an operation blacklist rule. The trigger detection of blacklist rules is implemented based on a rule engine. The rule engine receives real-time operation behavior monitoring data, which comes from the instruction parsing channel of the remote desktop protocol. It parses and extracts information such as mouse click coordinates, keyboard key events, and file path access requests input by the user. The rule engine performs matching operations on the operation behavior data against the rules stored in the blacklist node. The matching operation is performed using conditional expression evaluation. When the conditional expression of any blacklist rule evaluates to true, it indicates that the operation has triggered the blacklist restriction. The system generates a restriction operation decision and outputs an operation blocking instruction.
[0063] The third decision level is the default allow processing. When the application type passes the whitelist verification and the operation does not trigger the blacklist rules, the system generates an allow access decision and maintains the normal session state. The allow decision allows the current operation to continue, without sending any intervention instructions to the protocol proxy layer.
[0064] The above three-level judgment process enables fine-grained control from coarse-grained session-level permissions to application-level operation-level permissions. The degree of fine-grainedness is reflected in the following three levels: Fine-grained control: Traditional Role-Based Access Control (RBAC) only performs a permission check once when a session is established. Once the check is successful, all operations during the session are allowed, representing a coarse-grained "all-allow or all-deny" mode. This invention refines the control granularity to the application and operation levels: within the same session, different application windows (such as file managers and command-line terminals) can be assigned different permission states; within the same application, different operational behaviors (such as file browsing and file downloading, command viewing and system configuration modification) can be controlled separately. This fine-grained control allows users to use authorized applications normally within the same session while unauthorized operations are blocked in real time.
[0065] Shifting the timing of permission determination forward: Traditional solutions only perform auditing and tracing after the operation is executed, which is a reactive response. This invention parses screen image frames in real time at the proxy layer of the remote desktop protocol, completing application type identification and permission matching before user operation occurs. When unauthorized application startup or high-risk operation attempts are detected, they are intercepted by the protocol proxy layer before the operation instructions reach the target virtual machine, achieving proactive defense.
[0066] Context sensitivity of permission policies: The permission control context tree structure constructed in this invention not only includes static user roles and resource scopes, but also incorporates dynamic application whitelists and operation blacklists. When performing permission determination, the decision engine considers four dimensions simultaneously: "who is operating" (user identity), "what is being operated" (application type), "how is the operation performed" (operation behavior), and "in what environment is the operation performed" (session context), thus achieving multi-dimensional and refined permission control.
[0067] After a decision is generated, the access control decision is written to the access control decision cache queue. The decision cache queue is implemented using a FIFO queue data structure, with newly generated decisions enqueued sequentially according to their generation order. The maximum queue length is configured to 1024 decision records; if this length is exceeded, the earliest enqueued decision record will be removed to free up memory space. The decision cache queue supports persistent storage of decision records; when a decision is executed or times out, the corresponding record is moved to the historical log storage area.
[0068] Simultaneously, the protocol proxy layer sends control command packets to the other end of the remote desktop session. The control command packets include four types of commands: Session termination commands, used to immediately terminate the entire remote desktop session, suitable for scenarios where a serious security threat is detected; the command packet includes a termination reason code and a session termination timestamp; Window minimize commands, used to force the window of a specified application to be minimized to the taskbar; the command packet includes the target window handle and the minimize duration; Application freeze commands, used to pause the response of a specified application; the command packet includes the target process identifier and the freeze duration; and Operation interception commands, used to block specified types of operation requests and return an operation prohibition prompt to the user; the command packet includes the type of operation intercepted and the user prompt text content. Based on the decision result, the appropriate command type is selected for transmission. The command transmission between the protocol proxy layer and the remote desktop session terminal is implemented using the virtual channel mechanism of the RDP protocol.
[0069] Step Six: Perform dynamic feedback and policy optimization for access control decisions, monitor the execution status of the access control decisions generated in Step Five in real time, and collect user behavior response data and session state change data after the decision is executed. User behavior response data includes three types of indicators: The first type is the frequency of secondary operation attempts, which counts the number of repeated operation attempts made by the user after receiving the blocking prompt. This indicator reflects the inhibitory effect of the blocking decision on user behavior. The second type is the frequency of switching application types, which counts the number of times the user switches between different application windows per unit time. This indicator reflects whether the user's work mode meets expectations. The third type is the session exit timing, which records whether the user exits the session after completing their work normally or exits prematurely due to the blocking prompt. This indicator reflects the degree of impact of the blocking decision on user experience.
[0070] The session state change data includes three categories of metrics: The first category is the remote desktop connection lifespan, which records the duration from session establishment to session termination. This metric is related to user experience and system resource consumption. The second category is the resource utilization fluctuation curve, which collects time-series data on CPU utilization, memory utilization, and network bandwidth utilization during the session. This metric reflects the application's demand characteristics on system resources. The third category is protocol interaction error code statistics, which records various protocol error codes that occur during the session and their frequency. This metric reflects the health status and potential problems of the session.
[0071] A decision-making effectiveness evaluation model is constructed based on the aforementioned multi-dimensional monitoring data. This evaluation model employs a reward function design within a reinforcement learning framework. The reward function includes two types of calculation factors: positive reward terms and negative penalty terms. Positive reward terms include two dimensions: improvement in user compliance rate and improvement in resource security indicators. The improvement in user compliance rate is calculated by comparing the number of violations within the evaluation period with the previous period; if the number of violations decreases, a positive reward value is calculated, with the reward value being proportional to the decrease. The improvement in resource security indicators is calculated by checking whether any security incidents occurred within the evaluation period; if no security incidents occurred, a positive reward value is calculated.
[0072] The comprehensive calculation formula for positive rewards is as follows: ,in This represents the overall positive reward value (dimensionless), used to quantify the overall performance of the current control strategy in terms of both compliance and security. The change in compliance rate (dimensionless, usually expressed as a percentage or decimal) reflects the improvement in compliance achieved by the current strategy compared to the previous moment. The security score (dimensionless) represents the level of security of the current system state; a higher score indicates better security. and The weighting coefficients are configured as 0.6 and 0.4 respectively, used to adjust the contribution ratio of compliance improvement and security to the total reward, with the sum of the two being 1. This reward function, through a linear weighted combination, guides the strategy towards optimizing for higher compliance rates while also considering system security, avoiding the sacrifice of one metric for the other.
[0073] The negative penalties include two dimensions: frequency of security incidents and false blocking rate. The frequency of security incidents is calculated by the number of malicious operation attempts and the number of sensitive data leakage incidents within the statistical assessment period, and a corresponding penalty value is calculated for each security incident. The false blocking rate is calculated by the proportion of blocked operations that are confirmed as legitimate by manual review. The higher the false blocking rate, the greater the penalty value.
[0074] The comprehensive calculation formula for negative penalties is as follows: ,in This represents the overall negative penalty value (dimensionless), used to quantify the adverse effects of the current control strategy on both security incidents and false alarms; The number of security incidents (dimensionless integer) reflects the actual number of security violations or abnormal events that occurred during the policy execution process; False alarm rate (dimensionless, ranging from 0 to 1) represents the proportion of false alarms generated by the strategy during the detection or judgment process; and The weighting coefficients are configured as 0.7 and 0.3 respectively, used to adjust the contribution ratio of the number of security incidents and the false alarm rate to the total penalty, with the sum of the two being 1. This penalty function, through a linear weighted combination, guides the strategy to reduce the occurrence of security incidents while also controlling the false alarm rate, avoiding ignoring actual security risks in an excessive pursuit of low false alarms, or introducing too many false alarms in order to reduce security incidents.
[0075] The formula for calculating the marginal benefit of strategy adjustment is as follows: ,in This represents the dimensionless marginal benefit resulting from strategy adjustments, used to quantify the change in net return of the current strategy compared to the previous strategy. This indicates that the strategy adjustment brought positive returns. This indicates that the strategy adjustment resulted in a net loss. This indicates that the strategy adjustment had no impact; This is the cumulative value of positive rewards, calculated using the comprehensive formula for positive rewards. The calculations are used to evaluate the overall performance of the strategy in terms of compliance and security. The cumulative value of negative penalties is calculated using the comprehensive formula for negative penalties. The calculated value is used to quantify the adverse effects of the strategy on security incidents and false alarms. By using the difference between positive rewards and negative penalties, the system can quantitatively evaluate the overall effect of each strategy adjustment, providing a clear direction for subsequent strategy iterations: when... Retain the current adjustment direction when The policy will revert to the previous policy. When the marginal benefit ΔU exceeds a preset threshold, the automatic update process of the permission policy will be triggered. The default configuration for the marginal benefit threshold is 0.3, and administrators can adjust this threshold based on actual operational data to control the sensitivity of policy updates.
[0076] The automatic policy update process includes three types of adjustments: The first category is adding and deleting application types from the whitelist. When the usage frequency of a specific application type decreases or remains at zero, the system recommends removing it from the whitelist to reduce security risks. When a user frequently applies to use a certain application type due to business needs and the application is approved, the system adds that application type to the whitelist. The second category is the optimization of blacklist rules. When the triggering frequency of a specific blacklist rule is too high, resulting in a decline in user experience, the system analyzes whether the triggering conditions are too strict and suggests adjusting the triggering condition parameters. When a new high-risk operation mode is discovered, the system suggests adding a new blacklist rule entry. The third category is the promotion and demotion of user role attributes. When a user's frequency of violations remains high, the system recommends downgrading the user's security level. When a user maintains a record of compliant operations, the system recommends upgrading their permissions.
[0077] The adjusted policy version is synchronized to the policy repository for use in subsequent sessions. Policy version management uses an incrementing version number mechanism, with the version number incremented by 1 after each policy adjustment. The policy repository also retains the policy files of the 10 most recent historical versions to support version rollback operations; The policy synchronization process employs a distributed consistency protocol to ensure policy data consistency among multiple replica storage nodes.
[0078] The system architecture includes a protocol proxy module, an application identification and inference engine, a permission policy management module, an access control decision engine, and a monitoring and feedback module. The protocol proxy module is deployed in the network forwarding layer of the cloud platform and is responsible for parsing, forwarding and issuing commands for the remote desktop protocol. This module interacts with the application recognition and inference engine through a virtual channel mechanism. The application identification inference engine is deployed as an independent process and interacts with the protocol proxy main process through an inter-process communication mechanism. This architecture design ensures the decoupling of the application identification function and the protocol forwarding function. When the application identification inference engine malfunctions, the protocol proxy main process can still maintain basic remote desktop forwarding capabilities.
[0079] To address the computational overhead issue in scenarios with concurrent multi-path remote desktop sessions, this invention designs a multi-level computational scheduling and resource management mechanism within the application recognition and inference engine, specifically including: Inference request queue management and dynamic batch processing optimization: The inference engine maintains a global inference request queue, receiving image frame inference requests from multiple remote desktop sessions. The queue manager sorts requests according to their arrival timestamps and session priorities. Priorities are determined by session security levels by default (sessions with higher security levels receive higher inference priorities). The batch processing scheduler aggregates requests in the queue according to a dynamic batch size. The batch size is dynamically adjusted based on the current GPU memory utilization and queue backlog length: when GPU memory utilization is below 60%, the batch size is configured to 16 frames; when memory utilization is between 60% and 80%, the batch size is reduced to 8 frames; when memory utilization exceeds 80%, the batch size is reduced to 4 frames to ensure stable execution of inference tasks. The maximum queue buffer depth is configured to 128 frames; requests exceeding this depth are discarded according to a priority strategy, and alarm logs are recorded.
[0080] Multi-GPU load balancing and computing power pooling: When the system is configured with multiple GPUs, the inference engine uses a round-robin plus load-aware scheduling strategy to distribute batch processing tasks to different GPU devices. The scheduler monitors the memory utilization and computing power of each GPU in real time, prioritizing the allocation of tasks to GPU devices with lower loads. When the memory utilization of a single GPU continuously exceeds 85%, the system automatically switches subsequent inference tasks to other GPUs or falls back to CPU inference mode (CPU inference mode uses a lightweight model, which reduces recognition accuracy but ensures continuous service availability).
[0081] Session-level computing power quota management: The system allocates an independent computing power quota to each remote desktop session. The quota parameters include: the maximum number of inference frames per unit time (default 60 frames), the maximum waiting time for single frame inference (default 100 milliseconds), and the maximum GPU memory usage limit for a single session (default 256MB).
[0082] When a session's inference requests exceed its quota, the excess requests are downgraded: for non-critical sessions, the acquisition frequency of its image frames is reduced (from 30fps to 5fps); for critical sessions, inference resources are prioritized, and computing power quotas are allocated from other low-priority sessions.
[0083] The system employs a tiered inference accuracy strategy, dynamically adjusting inference accuracy based on the session's security level and the current system load.
[0084] Under high load, the system automatically enables model quantization (quantizing floating-point model parameters into 8-bit integers) and model pruning strategies, which increases the model inference speed by 2 to 4 times and keeps the recognition accuracy decrease within 3%. Under low load, the system restores full-precision model inference to ensure the highest recognition accuracy.
[0085] Inference result caching and reuse mechanism. For image frames acquired consecutively in the same session, if the pixel difference between consecutive frames is less than a preset threshold (default 5%), the inference result of the previous frame is directly reused, and the inference calculation for this frame is skipped, thus avoiding repeated inference for static or slowly changing images.
[0086] Pixel differences are quantified by calculating the peak signal-to-noise ratio (PSNR) of the preceding and following frames. When the PSNR is greater than 35dB, it is considered highly similar, triggering result reuse.
[0087] The inference computation of the application recognition inference engine adopts a batch processing mode to improve GPU utilization; The continuously acquired image frames are accumulated to a fixed batch size and then used for inference calculations. The batch size parameter is configured to be 12 frames, which achieves a balance between inference latency and computation throughput. The maximum wait time for the batch processing queue is configured to 50 milliseconds. When the number of image frames in the queue has not reached the batch size but the wait time has exceeded this threshold, forced inference computation is triggered to ensure real-time response. The GPU memory usage limit for the inference engine is configured to 4GB to avoid affecting other inference tasks running on the same GPU.
[0088] The permission policy management module is responsible for three functions: policy storage, policy loading, and policy updating. Policy storage is implemented using a distributed key-value database, and data is stored in isolation according to tenant partitions to ensure data security in a multi-tenant environment; Policy loading employs a combination of preloading and on-demand loading. Frequently accessed policy files are loaded into the memory cache upon system startup, while infrequently accessed policy files are read from the database on demand when a session is established. Policy updates support atomic transaction operations, ensuring policy integrity and consistency.
[0089] The access control decision engine is responsible for policy matching calculations and decision generation and execution. Internally, the decision engine maintains a policy matching state machine, whose state transitions are driven by application type identification results and operational behavior monitoring data. The decision generation latency of the decision engine is controlled within 10 milliseconds to meet real-time requirements; The decision engine also integrates two types of decision components: a rules engine and a machine learning model. The rules engine is responsible for handling deterministic policy terms, while the machine learning model is responsible for handling boundary scenarios and ambiguous scenarios.
[0090] The monitoring and feedback module is responsible for three functions: data collection, effect evaluation, and strategy optimization. The data acquisition component collects session state data once per second and user behavior data on a per-decision-event basis. The performance evaluation component uses a sliding window algorithm to calculate evaluation metrics, with the window length configured to be 1 hour. The strategy tuning component performs strategy evaluation and automatic adjustment processes once a day, and this frequency parameter can be adjusted according to the actual operating load.
[0091] For example, an enterprise cloud platform deploys the access control method of the present invention to perform application-level control over developers' remote desktop access; The enterprise's access control policy is configured as follows: Developer roles are authorized to use three applications: file manager, web browser, and command-line terminal; developers are prohibited from downloading executable files in web browsers; developers are prohibited from executing sensitive commands that modify system configurations in command-line terminals; developers are prohibited from using database clients to access production environment databases.
[0092] When developers connect to the development and testing virtual machine in the cloud platform via the remote desktop protocol, the system executes the following access control process: the protocol proxy module receives and parses the RDP protocol data packets, and the image frame acquisition unit captures screen image frames from the protocol data stream with a resolution of 1920×1080 pixels and a frame rate of 30fps. The application recognition inference engine preprocesses the acquired image frames, extracts visual features, and identifies the application type. It identifies two application windows, a web browser window and a command-line terminal window, on the current screen and outputs the bounding box coordinates of the two windows. The permission policy management module loads the policy file corresponding to the developer, constructs a permission control context tree structure, and parses out the application whitelist as 2, 4, 5 and the blacklist rules as 3 rules. The access control decision engine matches the identification results with the permission context, determines that both application windows are in the whitelist, and continues to perform blacklist rule detection. Finally, it detects that there is an operation to execute a sensitive configuration command in the command line terminal, triggers the blacklist rule, generates a restriction operation decision, executes the operation interception instruction to block the execution of the command, and returns an operation prohibition prompt to the user.
[0093] Example 2: The generation of application-level access control decisions adopts a hybrid decision architecture that combines rule engines and machine learning models; The rule engine is responsible for processing the deterministic policy terms explicitly specified in the policy file. The rule engine is implemented using the Rete algorithm. The core data structure of the Rete algorithm is the inference network, which contains three types of nodes: type nodes, fact nodes, and rule nodes. Type nodes correspond to the application types and operation behavior types defined in the policy terms; fact nodes correspond to real-time application identification results and operation behavior monitoring data; rule nodes correspond to the condition-action rules in the policy file. When new fact data is input into the rule engine, the Rete algorithm quickly determines the triggering rule and executes the corresponding action through node matching and propagation mechanisms.
[0094] Machine learning models are responsible for handling boundary and ambiguous scenarios that rule engines cannot cover. The input features of machine learning models include four dimensions: The first dimension is the application type identification confidence score, which is the maximum value in the application type probability distribution vector output in step three; The second dimension is the user's historical behavior sequence, which records the user's operational behavior in the past hour and uses a unique encoding format to represent the frequency of each type of operation. The third dimension is the session duration, which is the length of time since the current session was established. The fourth dimension is the real-time resource load rate, which is the weighted average of the current virtual machine's CPU utilization and memory utilization. The machine learning model is implemented using the XGBoost gradient boosting decision tree algorithm. The model structure contains 100 decision trees, with a maximum depth of 6 layers for each tree. The model output is a risk score for the scenario, with a value ranging from 0 to 1.
[0095] The risk score output by the machine learning model is mapped to three decision categories: deny, restrict, and allow. The configuration parameters for the mapping threshold are stored in the policy configuration file, and the mapping threshold uses a dual-threshold design. The first threshold is 0.2. When the risk score is less than or equal to this threshold, it is mapped to an allowed decision. The second threshold is 0.6. When the risk score is greater than or equal to this threshold, it is mapped to a rejection decision; when the risk score is between the two thresholds, it is mapped to a restriction decision. The mapping thresholds can be visually adjusted by administrators through the management interface, and the adjusted threshold parameters are synchronized to the decision engine in real time.
[0096] When the rule engine and the machine learning model make conflicting decisions, the decision result of the rule engine is given priority to ensure the principle of certainty in the strategy. The conflict detection mechanism is executed during the decision generation phase. When the decision output by the rule engine is inconsistent with the decision output by the machine learning model, the system records the conflict event and executes the rule engine's decision. Simultaneously, the feature data of the conflict event is logged for subsequent analysis. When the rule engine fails to output a decision, the system uses the decision from the machine learning model as the final decision.
[0097] Example 3: This invention provides an application type recognition method based on a multi-task learning framework, which further optimizes the performance of visual feature extraction and application type classification based on steps two and three.
[0098] The visual feature extraction of image frames adopts a multi-task learning framework, which simultaneously extracts application type recognition features and outputs auxiliary prediction results of window interface type in parallel. The auxiliary prediction results for window interface types include 12 interface categories such as normal windows, modal dialog boxes, file selectors, and system notification bars; The network structure of the multi-task learning framework is an extension of the high-level semantic feature extraction network in step two. The extended network contains two task heads: The first task header is the application type classification header, with an output dimension of 7, corresponding to 7 main application types; The second task head is a window interface classification head, with an output dimension of 12, corresponding to 12 interface categories. The two task heads share the network parameters of the feature encoding layer and semantic feature extraction layer, and only perform task-specific parameter learning in the final fully connected layer.
[0099] The loss function of the multi-task learning framework is designed as a weighted sum of the application type recognition loss and the window interface type recognition loss; The formula for calculating the weighted loss is: ; in This represents the combined loss value for multi-task learning, used to jointly optimize the two tasks of application type recognition and window interface type recognition; The application type identification loss is calculated using the cross-entropy loss function, which measures the difference between the application category predicted by the model and the true label. The loss for window interface type recognition is also calculated using the cross-entropy loss function, which measures the difference between the interface category predicted by the model and the true label. and This is the loss weighting coefficient, used to adjust the contribution ratio of the two tasks to the total loss, with a configured ratio of [value missing]. ,Right now , This configuration ensures that application type recognition, as the primary task, has priority in learning weight, allowing the model to focus more on optimizing the recognition accuracy of the primary task during training, while also taking into account the learning of auxiliary tasks.
[0100] The context-awareness enhancement mechanism for auxiliary prediction results is achieved in the following way: when both the main application window and auxiliary interface elements exist in the image, the type information of the auxiliary interface can help eliminate ambiguity in the determination of the main application type. For example, when a command-line terminal window and a file selector dialog box appear in the recognition results, the system can determine, based on the context of the file selector dialog box's appearance, that the command-line terminal is likely in a file operation-related command input state, and thus adjust the access control policy for that terminal. When only the file selector dialog box is recognized and the main application window is not detected, the system can use the type of the file selector dialog box as the basis for inferring the type of the main application.
[0101] The foregoing has shown and described the basic principles, main features, and advantages of the present invention. Those skilled in the art should understand that the present invention is not limited to the above embodiments. The embodiments and descriptions in the specification are merely illustrative of the principles of the invention. Various changes and modifications can be made to the invention without departing from its spirit and scope, and all such changes and modifications fall within the scope of the present invention as claimed. The scope of protection of the present invention is defined by the appended claims and their equivalents.
Claims
1. A cloud computing engine access control method based on a remote desktop protocol, characterized in that, Includes the following steps: Step 1: Protocol parsing and image frame acquisition. The remote desktop protocol is parsed and processed, and screen image frame data during the remote desktop session is captured as samples to be analyzed. The acquired image frames are preprocessed by color space conversion and geometric parameter normalization. Step 2: Visual feature extraction and semantic parsing, extracting multi-level visual feature vectors from the preprocessed image frames and performing feature-level fusion processing; Step 3: Application type identification and window hierarchy determination. Based on the fused comprehensive feature vector, an application type classifier is constructed, outputting the application type determination result and spatial positioning bounding box coordinates. After the application type classifier outputs, confidence evaluation and hierarchical processing are performed: when the difference between the highest probability value and the second highest probability value is less than the discrimination margin threshold, the auxiliary verification channels of OCR text recognition and window attribute analysis are activated for cross-validation; when the probabilities of all categories are lower than the confidence threshold, it is determined to be an unknown category and the incremental learning process is triggered. Perform a temporal consistency check on the recognition results of multiple consecutive frames, and output the final judgment when the recognition results of consecutive frames are consistent. Step 4: Loading session-level permission policies and building contexts. Load the session-level permission policy file and build the permission control context tree structure based on the user identity and tenant identity. Step 5: Application-level access control decision generation and execution, matching the application type identification results with the access control context and executing the corresponding control instructions; Step Six: Dynamic feedback and policy optimization of access control decisions, real-time monitoring of decision execution status, and triggering automatic update process of permission policies based on evaluation results; The application recognition inference engine employs a multi-level computing power scheduling mechanism: maintaining a global inference request queue, sorting and dynamically batching requests according to session priority, with the batch size dynamically adjusted based on GPU memory usage; when multiple GPUs are configured, a round-robin plus load-aware scheduling strategy is used for task distribution; an independent computing power quota is allocated to each session, and the image frame acquisition frequency is reduced for non-critical sessions when the quota is exceeded; and quantization inference and result reuse mechanisms are automatically enabled under high load conditions.
2. The cloud computing engine access control method based on remote desktop protocol according to claim 1, characterized in that, The protocol parsing includes protocol header parsing, security layer parsing, and instruction stream parsing; Image frame acquisition employs a dual-buffering mechanism. The main buffer area continuously receives protocol data streams to update the pixel matrix, while the auxiliary buffer area saves a pixel snapshot of the previous complete frame for incremental detection.
3. The cloud computing engine access control method based on remote desktop protocol according to claim 1, characterized in that, The color space conversion transforms the pixel matrix from the original color space to the target analysis color space; The conversion process includes calculating the saturation component, calculating the hue component, and determining the luminance component; The geometric parameter normalization standardizes the resolution, pixel depth, and color channel mapping relationship of the image frame.
4. The cloud computing engine access control method based on remote desktop protocol according to claim 1, characterized in that, The multi-layered visual feature vector includes three dimensions: low-level texture features, mid-level shape features, and high-level semantic features; The underlying texture features are calculated using a local binary pattern operator, and the gray-level change patterns of local regions of the image are statistically analyzed to form texture descriptors. The mid-level shape features extract contour boundary information from the image through edge detection operators, and perform multi-scale geometric analysis on the contour boundaries to obtain shape feature vectors; High-level semantic features are predicted based on a pre-trained multi-label image classification network.
5. The cloud computing engine access control method based on remote desktop protocol according to claim 1, characterized in that, The application type classifier adopts a deep convolutional neural network architecture, which includes a feature encoding layer, a spatial attention layer, and a classification output layer. The feature encoding layer uses multiple convolutional operations to abstract and represent features layer by layer. The spatial attention layer generates an attention mask based on the importance distribution of spatial locations in the image, which enhances the features of key regions while suppressing background interference regions. The classification output layer outputs the application type determination result and the spatial positioning bounding box coordinates.
6. The cloud computing engine access control method based on remote desktop protocol according to claim 1, characterized in that, The session-level permission policy file contains four policy elements: user role attributes, resource access scope, application usage whitelist, and application operation blacklist. User role attributes include the user's security level and functional positioning in the cloud platform; resource access scope defines the list of cloud computing resource instances that the user can access; The application whitelist includes a set of application types that the user is authorized to use; The application operation blacklist includes a set of high-risk operation types that users are prohibited from performing.
7. The cloud computing engine access control method based on remote desktop protocol according to claim 1, characterized in that, The access control context is organized in a tree structure, with the root node being the user identity identifier, and the child nodes being the resource scope node, the application whitelist node, and the operation blacklist node in sequence.
8. The cloud computing engine access control method based on remote desktop protocol according to claim 1, characterized in that, The access control decision matching operation follows the following order: Determine if the application type exists in the application whitelist set; if not, generate an access denial decision. If it exists, determine whether the operation type triggers the operation blacklist rule; if it does, generate a decision to restrict the operation. If the blacklist rule is not triggered, an access permission decision is generated. The control command package includes command types such as session termination command, window minimize command, application freeze command, and operation interception command.
9. The cloud computing engine access control method based on remote desktop protocol according to claim 1, characterized in that, The dynamic feedback and strategy optimization constructs a decision-making effectiveness evaluation model by collecting user behavior response data and session state change data. The decision-making effectiveness evaluation model adopts the reward function design in the reinforcement learning framework, taking the improvement of user compliance rate and resource security indicators as positive rewards, and the frequency of security incidents and false blocking rate as negative penalties; The automatic policy updates include adding or deleting application whitelist types, optimizing rules for operating blacklists, and adjusting user role permissions.
10. The cloud computing engine access control method based on remote desktop protocol according to claim 1, characterized in that, Deploy an application recognition inference engine as a front-end processing unit for access control at the protocol proxy layer of the remote desktop protocol; The application recognition and reasoning engine runs as an independent process and interacts with the protocol agent main process through inter-process communication mechanisms; The inference engine uses a batch processing mode for inference calculations, accumulating continuously acquired image frames into a fixed batch size before performing inference calculations uniformly. The generation of application-level access control decisions adopts a hybrid decision architecture that combines a rule engine and a machine learning model. The rule engine is responsible for handling deterministic policy terms, while the machine learning model is responsible for handling boundary scenarios and ambiguous scenarios. When the rule engine and the machine learning model make conflicting decisions, the decision result of the rule engine is given priority.