Out-of-band perception-free cloud desktop content auditing system
By designing an out-of-band, no perception cloud desktop content audit system, including user video acquisition, application recognition, word processing and video traceability modules, the problems of unperceived acquisition, insufficient output refinement and imperfect behavior sequence recording in cloud desktop content audit are solved, and efficient and accurate audit and analysis are achieved.
Patent Information
- Application Number
- CN202510206518.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-25
- Publication Date
- 2025-05-30
AI Technical Summary
The existing cloud desktop content audit system has problems in obtaining cloud desktop videos without perception, insufficient output refinement, and incomplete behavior sequence recording and traceability, which affects the objectivity and analysis efficiency of audit results.
A cloud desktop content audit system with no sense out-of-band out-of-band is designed, including user video acquisition module, application recognition module, word processing and keyword matching module based on PaddleOCR, and video traceability and playback module. Through these modules, user video acquisition, application recognition, operation details analysis and behavior traceability are realized.
It realizes a meticulous audit of cloud desktop operation content, ensures the authenticity and integrity of data, builds the training set and functional keyword database of the target software, and improves audit efficiency and analysis accuracy.
Smart Images

Figure CN120075479A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of information security and behavior analysis, and in particular to an out-of-band, non-perceptual cloud desktop content auditing system. Background Art
[0002] With the rapid development of information technology, the network environment is becoming increasingly complex, and information security is facing huge challenges. The demand for desktop operation content auditing of cloud computing host users has increased accordingly. This demand is not only related to the internal information security and business secrets protection of the enterprise, but also an important measure to meet compliance requirements.
[0003] From a technical perspective, it is difficult to achieve seamless monitoring of different applications while ensuring the high quality and integrity of the video. Existing audit methods have the following technical problems:
[0004] 1. No perceived loss. In the cloud computing host desktop environment, especially in scenarios such as safe production that have strict requirements on user behavior auditing, it is difficult to obtain cloud desktop videos, which may affect the objectivity of the audit results.
[0005] 2. The output is not detailed enough. The acquired videos present user operations in a general manner and lack structure, making it impossible to deeply analyze specific operation details. The audit results are presented in a rough manner and lack effective organization and arrangement. A clear structural system has not yet been formed, which increases the difficulty of subsequent analysis and utilization and makes it difficult to quickly extract key information.
[0006] 3. The record and traceability of behavior sequences are incomplete, making it impossible to restore the full picture of behavior and difficult to understand potential risks, which is particularly unfavorable for deep mining and analysis of user behavior patterns. Summary of the invention
[0007] In view of this, the present invention discloses an out-of-band, non-aware cloud desktop content auditing system to solve the above problems; the cloud desktop content auditing system includes four modules connected in sequence: a user video acquisition module, an application identification module, a text processing and keyword matching module based on PaddleOCR, and a video tracing and replaying module; wherein, the user video acquisition module is used to realize the acquisition of user videos, the application identification module is used to solve the problem that there are multiple target software at the same time in the image and have overlapping areas, the text processing and keyword matching module based on PaddleOCR is used to identify the specific functions performed by the user in the application, and the video tracing and replaying module is used to realize the review and playback of user operation behaviors.
[0008] The beneficial effects of the present invention include: By applying recognition to match keywords, constructing a training set of target software to identify target applications and software function keywords, realizing the specific matching from software to functions, solving the problem of insufficient refinement degree of output and avoiding potential risks; By obtaining the cloud platform key, simulating user login to record videos, and then realizing non-intrusive auditing, ensuring the authenticity and integrity of data; By forming a JSON file with the recognized application time series and custom exceptions, and forming an exception timeline through breakpoint transmission for quick viewing, solving the problems of behavior sequence recording and traceability, and then accurately positioning problems and improving the auditing efficiency. BRIEF DESCRIPTION OF THE DRAWINGS
[0009] Figure 1 It is a schematic diagram of the working process of the user video acquisition module in an embodiment of the present invention;
[0010] Figure 2 It is a schematic diagram of the working process of the application recognition module in an embodiment of the present invention;
[0011] Figure 3 It is a schematic diagram of the working process of the text processing and keyword matching module based on PaddleOCR in an embodiment of the present invention;
[0012] Figure 4 It is a schematic diagram of the working process of the video traceability and replay module in an embodiment of the present invention;
[0013] Figure 5 It is a structural diagram of the out-of-band non-intrusive cloud desktop content auditing system in the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0014] In order to make the objectives, technical solutions, features, and advantages of the present invention clearer and more understandable, the present invention will be further described below in conjunction with the accompanying drawings and embodiments.
[0015] This embodiment includes an out-of-band non-intrusive cloud desktop content auditing system, and the out-of-band non-intrusive cloud desktop content auditing system includes four modules connected in sequence: a user video acquisition module, an application recognition module, a text processing and keyword matching module based on PaddleOCR, and a video traceability and replay module. The structure of the system is as Figure 5 shown.
[0016] Furthermore, Figure 1 It is a schematic diagram of the user video acquisition module in the present invention. The user video acquisition module acquires and saves user videos, including:
[0017] S101. Obtain the interactive link information of the cloud computing host through the management interface or API.
[0018] Specifically, the link information related to the user's use of the cloud host is obtained through the cloud platform's management interface or specific API calls. The link information is a key identifier for the user to interact with the cloud host, and includes various parameters and path information for the user to access the cloud host.
[0019] S102: parse the content of the link information to obtain basic user information and platform key.
[0020] Specifically, the corresponding TLS / SSL keys are obtained from the configuration file of the cloud host. These keys play a vital role in the secure communication of the cloud host. They are used to encrypt and decrypt the transmitted data. The SSL / TLS keys are used to decrypt the previously obtained link information and the encrypted information that may be involved in the communication process. Encryption technology is used to protect the privacy and integrity of data in network communications, and decryption is a necessary step in processing this data. The encrypted information is restored to a readable plain text form using keys and decryption algorithms for subsequent analysis and processing.
[0021] S103, mirroring the user video based on the WebSocket protocol and Selenium tool according to the user's basic information and platform key.
[0022] Specifically, the obtained user information and keys obtained by the cloud host platform are parsed, the user video is mirrored, and with the help of WebSocket technology, it is disguised as the same link as the user accessing the cloud host. WebSocket is a protocol for full-duplex communication over a single TCP connection, which allows real-time, two-way communication between the server and the client.
[0023] Furthermore, a connection with the cloud host is established through WebSocket-client, and the parsed user information and key are added to the request message to simulate the user's real request. The request header of the user's real request, such as User-Agent, Cookie, etc., is added. According to the user's real access mode, the frequency and content of the message are controlled, and the operation is disguised as the user's real access request, so that the communication link between the user and the cloud host can be smoothly accessed. After the link is successfully disguised, the received data is converted into video stream data. The data generated during the interaction between the user and the cloud host are in various forms and may contain various information such as text, images, instructions, etc. These data are reorganized and encoded according to the rules of video encoding to generate video stream data that meets the video format standard.
[0024] Furthermore, use the tool of Selenium to automate the Web browser to simulate user behaviors, including simulating network requests. Create a simulated browser environment and simulate the process of a user logging in to a cloud host in it. In this simulated environment, we can play and record the video stream data obtained from the previous conversion, so as to realize the screen recording operation of the process of the user using the cloud host.
[0025] S104. Record the user video, and save the recorded video in the local storage. Stop recording when the interactive link information becomes invalid.
[0026] Specifically, after completing the video recording task of a specific operation or process, save the recorded video in the local storage server. The naming rule is the combination of the relevant IP address and the first ten digits of the timestamp at the end moment of the video recording. For example, if the IP address of the device for recording the video is 192.168.1.100, tshark reads the network card IP address, and the timestamp of the start time of the recording is 1673456789.123456, then according to the rule, the naming of this video file will be 192.168.1.100_1673456789, which is used to save the user address and infer the specific event occurrence time later. Use local storage for the original video. Centered on data, completely separate the storage device from the server, centrally manage the data, thereby releasing bandwidth, improving performance, and reducing the total cost of ownership. The cost of local storage of the original video is much lower than using server storage, while the efficiency is much higher than the latter. We can use its powerful functions to achieve centralized management of video files, including but not limited to operations such as file storage, retrieval, backup, and permission control. When it is detected that the user performs the logout operation, the system will automatically determine that the associated link becomes invalid. Based on this logic, to avoid invalid data collection and resource waste, the system will synchronously trigger the stop recording mechanism and immediately terminate the currently ongoing video recording task to ensure that the recording work is accurate and efficient.
[0027] Furthermore, the recording standard is as follows: Obtain the original video frame rate and resolution through the open-source tool FFmpeg. If the original video frame rate reaches 30 frames per second and the resolution is higher than 1920*1080, then record according to the standard of 1920*1080 resolution and 30 frames. If it is lower than this standard, record according to the original video standard. Exceeding the standard will cause a large collection and storage of the original data set, and being lower than the standard will make the picture relatively blurred, which is not conducive to the edges of the target becoming blurred in the low-resolution image.
[0028] Furthermore, Figure 2 is a schematic diagram of the application recognition module in the present invention. The application recognition module is used to solve the problem that there are multiple target software in the image and there are overlapping regions. The application recognition module processes the user video and generates a time series including:
[0029] S201. Read the user video from local storage, and use the pre-trained YOLO model to divide the overlapping parts of different target software in the user video by color features, and output the detected target software name, time point, and the divided area of the target software.
[0030] Specifically, in the complex cloud desktop image analysis scenario, there are often overlapping parts between different target software in the picture, which increases the difficulty of accurately identifying and distinguishing each software. To effectively solve this problem, a method based on the average color difference is used to finely divide the overlapping area. Calculate the average color difference between the common area and each overlapping software. Compare the average color difference of the common area with the average color difference of the overlapping software. If the average color difference of the common area is closer to the average color difference of the non-overlapping area software, that is, the difference is smaller, then it is determined that the common area belongs to the software with a closer average color difference, and the area division prepares for the subsequent text recognition and matching.
[0031] Specifically, in the CIE Lab color space, L represents the brightness of the color, a represents the axis from green to red, and b represents the axis from blue to yellow. The calculation formula of the color difference calculation formula ΔE is as follows:
[0032]
[0033] Among them, ΔL * , Δa * , Δb * respectively represent the differences between two colors on the three axes of L, a, and b. Its meaning is to combine the differences between the two colors in three dimensions into one value, so as to obtain a single number used to describe the degree of color difference.
[0034] Furthermore, the method for training the YOLO model is as follows:
[0035] Step 1. Read the user video from local storage and construct a standardized original dataset.
[0036] Specifically, read specific video resources from the local storage system through the API, and sample the video at a fixed time interval of once every 10 seconds. Organize and store the extracted frames according to specific naming rules and directory structures, so as to construct a standardized and normalized original dataset.
[0037] Step 2. Label the target software and function interface in the original dataset, and perform data augmentation on the original dataset.
[0038] Specifically, use the overall annotation method for the annotation target, and only annotate the affiliated software. In the two-dimensional image space, for each target object to be annotated, use (x min , y min,x max ,y max ), representing its position and size, where (x min ,y min ) are the coordinates of the upper left corner of the bounding box, and (x max ,y max ) are the coordinates of the lower right corner of the bounding box. This annotation method can accurately locate the position information of the target in the image. After completing the annotation work, in order to further improve the performance and robustness of the model, comprehensive data augmentation operations are performed on these annotated data. Data augmentation aims to increase the diversity of data through various transformation methods, enabling the model to learn richer features.
[0039] In terms of geometric transformation, operations such as rotating, flipping, translating, and scaling the data can enable the model to learn the features of the target at different angles. Taking the rotation operation as an example, various positions that may appear in the cloud desktop video are imitated. Assuming that the coordinates of a pixel point in the image are (x, y), after rotating by an angle θ around the center of the image (C x ,C y ), the new coordinates can be calculated by the following formula:
[0040]
[0041] In terms of color transformation, through color transformation, the model can better adapt to the target features in different host color environments. Operations such as adjusting the brightness, contrast, saturation, and hue of the data are performed. Taking the brightness adjustment as an example, let the color value of the pixel point in the original image be I(x, y), and the adjusted brightness value be I ′ (x, y), and the brightness adjustment factor be α. Then I ′ (x, y) = α·I(x, y), α ∈ [0.2, 2]. When α > 1, the image brightness increases; when α < 1, the image brightness decreases. The range of α is limited to avoid invalid training data in extremely bright or dark situations.
[0042] S202. Statistically record the time points when the target software appears each time, and arrange the target software in regions to form an ordered time series.
[0043] Specifically, for the detection process of a specific object, record and organize the time points that appear in each detection, so as to form a strictly ordered time series where t n represents the time point of the nth successful detection of the object, that is, t 1 < t 2 < t 3 … < t n < … < t ∞ , and the time series It provides basic data support for subsequent in-depth analysis of the behavior sequence and usage duration of the object. To accurately define whether the object is in an unused state, a time interval threshold T with clear physical meaning is introduced. In this embodiment, T is preferably 10 seconds. If t n+1 -t n >T, it is considered that t 1 …t n is a continuous behavior sequence.
[0044] S203. Transmit the recognized time series to the next module.
[0045] Specifically, convert the frame picture into the BASE64 encoding format, re-divide the bounding box of the target object in the frame picture, the initially recognized target software, and a count. Taking the remainder of the frame rate by count can obtain the current second, and thus infer the specific time when the current event occurs through the timestamp on the video file name. Transmit the time series in JSON format to the next step.
[0046] Furthermore, Figure 3 is a schematic diagram of the text processing and keyword matching module based on PaddleOCR in the present invention. This module processes the time series to obtain a text file including the specific operations performed by the user in the application program, including:
[0047] S301. Perform an optical character recognition task on the time series using the PaddleOCR framework to obtain an initial text sequence.
[0048] Specifically, restore the BASE64-encoded picture and perform optical character recognition on the divided area text using the PaddleOCR framework. Convert the text information in the image into a recognizable feature vector. To accurately recognize text in different directions, use a text direction classifier and adjust the text direction through affine transformation technology to achieve effective classification and recognition of text in the 0-degree and 90-degree directions, and then convert the feature vector into the corresponding text sequence.
[0049] S302. Remove the noise in the initial text sequence through regular expressions.
[0050] Specifically, perform feature extraction on the corrected text image through a convolutional neural network. Input the feature map into a recurrent neural network to capture the sequence information of the text. Process the output of the recurrent neural network through a connectionist temporal classification layer to map the feature sequence to the corresponding text sequence, thereby realizing the recognition of the text.
[0051] Furthermore, the recognized text may contain various noises, such as punctuation marks, spaces, special characters, or garbled codes. These noises are located and removed through the regular expression [^a-zA-Z0-9\u4e00-\u9fa5\x00-\x7F]. Among them, the caret (^) inside the square brackets represents negation, that is, it matches all characters except those specified inside the square brackets; a-zA-Z matches all English letters (both uppercase and lowercase), 0-9 matches all digits;
[0052] \u4e00-\u9fa5 matches all Chinese characters; \x00-\x7F matches characters within the ASCII code range (from 0 to 127). Common punctuation marks, spaces, special characters, and garbled codes outside the ASCII range are removed through this regular expression.
[0053] S303. Construct a software-oriented keyword library.
[0054] Specifically, to construct a software-oriented keyword library, a multi-level structure is established. The construction method is as follows: Store the keyword library in the form of text files (such as JSON, XML formats) locally or on a network server. Through operations on the file system, it is convenient to manage and control the version of the keyword library file.
[0055] S304. Based on the improved BM25 algorithm, match the text sequence after removing noises with the keyword library.
[0056] Specifically, the BM25 algorithm can score the relevance between a document and a query based on factors such as term frequency and document length, and is used to match and retrieve the cleaned data text with the multi-level software keywords, so as to quickly find relevant matching information. To further improve the ability to capture and process key information, an adaptive weight parameter w is introduced into the BM25 algorithm to obtain the improved BM25 algorithm.
[0057] Furthermore, in practical applications, the importance degrees of different software keywords vary under different business scenarios and user requirements. By assigning an adaptive weight parameter w to each software keyword, the weight can be dynamically adjusted according to the actual importance of the keyword. For each software keyword, its weight w is dynamically calculated according to the hierarchical structure it belongs to and the characteristics of the business scenario. The weight w will be combined with other parameters in the BM25 algorithm to jointly score the retrieval results and output the results. The following is the BM25 formula with the introduction of the weight w:
[0058]
[0059] Among them, q represents the query term, d represents the document, n represents the number of terms in the query, idf(q i) represents the inverse document frequency, f(q i , d) represents the frequency of the query term in the document, k 1 and b represent the adjustment parameters of the BM25 algorithm, |d| represents the document length, avgdl represents the average document length, w i represents the weight of field i.
[0060] S305. Output the matching results and store them in the database.
[0061] Specifically, the matching results are sorted according to the scores generated during the retrieval process, and the sorted results are hierarchically output according to the pre-constructed keyword thesaurus of the multi-level software. During the output process, starting from the top-level classification and gradually delving into each sub-level, the specific content matching the retrieval query under each level is displayed, that is, the specific behavior review from the software to the function is realized.
[0062] Furthermore, Figure 4 is a schematic diagram of the video traceability replay module in the present invention. The video traceability replay module is used to implement the review and replay of the user's operation behavior. The video traceability replay module replays the specific operations performed by the user in the application program according to the text file, including:
[0063] S401. Define the abnormal features and retrieve the corresponding time series from the database according to the abnormal features.
[0064] Specifically, define the abnormal features, such as operating irrelevant software within a limited time range, operating special content within a limited time range, and query relevant normalized behavior sequences and abnormal nodes from the database by inputting specified query conditions.
[0065] S402. Output the abnormal log in JSON format according to the time series.
[0066] S403. Play the corresponding user video according to the abnormal log.
[0067] Specifically, when the video traceability replay module plays the user video corresponding to the abnormal features, the abnormal log in JSON format is loaded as the event timeline information. When the user clicks on a specific position on the video progress bar or enters a specific time point, the video traceability replay module locates the corresponding video frame and starts playing, using breakpoint resumption to transmit the video to reduce network resource consumption and improve the resource utilization efficiency of the network, presenting a clear and complete behavior sequence.
[0068] Finally, it should be noted that the above description only depicts some embodiments of the present invention. For those skilled in the art, various changes, modifications, substitutions, and variations to these embodiments can be conceived without departing from the principles and spirit of the present invention. The protection scope of the present invention is defined by the appended claims and their equivalents, and the above actions should all be covered within the protection scope of the present invention.
Claims
1. An out-of-band, non-aware cloud desktop content audit system, applied to a cloud computing host system, characterized in that: It includes four modules connected in sequence: user video acquisition module, application identification module, text processing and keyword matching module based on PaddleOCR, and video tracing and replay module; Among them, the user video acquisition module acquires and saves user videos; the application identification module processes the user video to generate a time series; the text processing and keyword matching module based on PaddleOCR processes the time series to obtain a text file including the specific operations performed by the user in the application; the video tracing and replay module is used to replay the specific operations performed by the user in the application according to the text file.
2. The out-of-band, non-aware cloud desktop content audit system according to claim 1 is characterized in that: The user video acquisition module acquires and saves user videos including: S101, obtaining interactive link information of a cloud computing host through a management interface or API; S102, parsing the content of the link information to obtain basic user information and platform key; S103, mirroring the user video based on the WebSocket protocol and Selenium tool according to the user's basic information and the platform key; S104: Record the user video, save the recorded video in local storage, and stop recording when the interactive link information becomes invalid.
3. The out-of-band, non-aware cloud desktop content audit system according to claim 1 is characterized in that: The application identification module processes user videos including: S201, reading a user video from local storage, using a pre-trained YOLO model to divide overlapping parts of different target software in the user video by color features, and outputting the name of the detected target software, time point, and the target software division area; S202, counting the time points of each detection of the target software, and arranging the target software into regions to form an orderly time series; S203: Pass the identified time series to the next module.
4. The out-of-band, non-aware cloud desktop content auditing system according to claim 3 is characterized in that: The data set used for pre-training the YOLO model is an enhanced data set. The enhanced data set is obtained by reading user videos from local storage and constructing a standardized original data set. Mark the target software and functional interface in the original data set, and perform data enhancement on the original data set to obtain an enhanced data set; Among them, data enhancement includes geometric transformation and color transformation. Geometric transformation includes rotating, flipping, translating and scaling operations on data, and color transformation includes adjusting the brightness, contrast, saturation and hue of data.
5. The out-of-band non-aware cloud desktop content audit system according to claim 1 is characterized in that: The data processing based on PaddleOCR's text processing and keyword matching module includes: S301, using the PaddleOCR framework to perform a text recognition task on the time series to obtain an initial text sequence; S302, removing noise in the initial text sequence by using a regular expression; S303, constructing a software-oriented keyword thesaurus; S304, matching the text sequence after noise removal with the keyword vocabulary based on the improved BM25 algorithm; S305: Output the matching results and store them in the database.
6. The out-of-band non-aware cloud desktop content audit system according to claim 5 is characterized in that: The method for constructing a software-oriented keyword thesaurus is to establish a multi-level structure and store the thesaurus in the form of a text file locally or on a network server, wherein the text file includes JSON and XML formats.
7. The out-of-band, non-aware cloud desktop content audit system according to claim 5 is characterized in that: The formula of the improved BM25 algorithm is: Among them, q represents the query term, d represents the document, n represents the number of terms in the query, and idf(q i ) represents the inverse document frequency, f(q i ,d) represents the frequency of the query word in the document, k1 and b represent the adjustment parameters of the BM25 algorithm, |d| represents the document length, avgdl represents the average document length, and w i Represents the weight of field i.
8. The out-of-band, non-aware cloud desktop content audit system according to claim 1 is characterized in that: The video tracing and replay module processes data in the following ways: S401, defining abnormal features, and retrieving corresponding time series in the database according to the abnormal features; S402, outputting an exception log in JSON format according to the time series; S403: Play the corresponding user video according to the abnormal log.
9. The out-of-band, non-aware cloud desktop content audit system according to claim 8 is characterized in that: The predefined abnormal features include: operating limited software within a limited time range, and operating limited content within a limited time range.