Operation behavior recognition method and system, and computing device and storage medium

By calculating the difference between the current operation and the previous operation in the page image and using a risk identification model for multimodal data analysis, the problems of high cost and low accuracy in risk identification in existing technologies are solved, achieving more efficient and accurate user access security.

WO2026021095A1PCT designated stage Publication Date: 2026-01-29CLOUD INTELLIGENCE ASSETS HOLDING (SINGAPORE) PTE LTD +1
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2025/103100
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-07-22
Filing Date
2025-06-24
Publication Date
2026-01-29

AI Technical Summary

Technical Problem

Current technologies are costly to identify risks in user access content and are prone to missed detections and false positives, which affect user access security and data security.

Method used

The target object server determines the first page image of the current operation and the second page image of the previous operation, calculates the difference page image, and sends it to the risk identification server. The risk identification model is used to identify risks in the multimodal data in the difference page image to obtain the behavior risk identification result.

Benefits of technology

It reduces the cost of using risk identification models, improves the accuracy and comprehensiveness of risk identification, avoids missed detections and false judgments, and ensures the security of user access.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2025103100_29012026_PF_FP_ABST
    Figure CN2025103100_29012026_PF_FP_ABST
Patent Text Reader

Abstract

An operation behavior recognition method and system, and a computing device and a storage medium. The method comprises: in response to a current operation behavior sent by a target object client, a target object server determining a first page image corresponding to the current operation behavior, determining a second page image corresponding to an operation behavior previous to the current operation behavior, determining, on the basis of the first and second page images, a difference page image indicating a difference between the first and second page images, and sending the difference page image to a risk recognition server; and on the basis of a risk recognition model, the risk recognition server performing risk recognition on data of at least two modalities included in the difference page image, so as to obtain an image risk recognition result, and determining a behavior risk recognition result of the current operation behavior on the basis of the image risk recognition result.
Need to check novelty before this filing date? Find Prior Art

Description

Operation behavior recognition methods and systems, computing devices and storage media Cross-references to related applications

[0001] This disclosure claims priority to Chinese patent application No. 202410986179.6, filed on July 22, 2024, the entire contents of which are incorporated herein by reference. Technical Field

[0002] This disclosure relates to the field of computer technology, and in particular to methods and systems for recognizing operational behaviors, computing devices, and storage media. Background Technology

[0003] With the development of computer technology, users can typically access resources and services provided by servers through client applications, such as browsing web pages and accessing applications. To ensure user access security, it is usually necessary to identify the risks associated with the content accessed by users.

[0004] However, risk identification of user access content is costly because it requires risk assessment for every single access, leading to high demands on server-side computing and storage resources. Furthermore, it is prone to missed detections and false positives, impacting user access security and data security on both the client and server sides. Therefore, an effective technical solution is urgently needed to address these issues. Summary of the Invention

[0005] In view of this, the present disclosure provides three methods for identifying user behavior. One or more embodiments of this disclosure also relate to two types of user behavior identification devices, a user behavior identification system, a computing device, a computer-readable storage medium, and a computer program product, to address the technical shortcomings of existing technologies, such as the high cost of risk identification of user access content.

[0006] According to a first aspect of the present disclosure, an operation behavior recognition method is provided, applied to an operation behavior recognition system, the operation behavior recognition system including a target object server and a risk recognition server, wherein the method includes: the target object server, in response to a current operation behavior sent by a target object client, determining a first page image corresponding to the current operation behavior, and determining a second page image corresponding to the previous operation behavior of the current operation behavior; determining a difference page image between the first page image and the second page image based on the first page image and the second page image, and sending the difference page image to the risk recognition server; the risk recognition server, according to a risk recognition model, performing risk recognition on data of at least two modalities included in the difference page image to obtain an image risk recognition result of the difference page image; and determining a behavior risk recognition result of the current operation behavior based on the image risk recognition result.

[0007] According to a second aspect of the present disclosure, an operation behavior recognition system is provided, the operation behavior recognition system including a target object server and a risk recognition server, wherein the target object server is configured to, in response to a current operation behavior sent by a target object client, determine a first page image corresponding to the current operation behavior, and determine a second page image corresponding to the previous operation behavior of the current operation behavior; determine a difference page image between the first page image and the second page image based on the first page image and the second page image, and send the difference page image to the risk recognition server; the risk recognition server is configured to, according to a risk recognition model, perform risk recognition on data of at least two modalities included in the difference page image to obtain an image risk recognition result of the difference page image, and determine a behavior risk recognition result of the current operation behavior based on the image risk recognition result.

[0008] According to a third aspect of the present disclosure, an operation behavior recognition method is provided, applied to a target object server, wherein the method includes: responding to a current operation behavior sent by a target object client, determining a first page image corresponding to the current operation behavior, and determining a second page image corresponding to the previous operation behavior of the current operation behavior; determining a difference page image between the first page image and the second page image based on the first page image and the second page image, and sending the difference page image to a risk recognition server; receiving a behavior risk recognition result of the current operation behavior sent by the risk recognition server, wherein the behavior risk recognition result is determined based on an image risk recognition result, and the image risk recognition result is obtained by the risk recognition server performing risk recognition on data of at least two modalities included in the difference page image according to a risk recognition model.

[0009] According to a fourth aspect of the present disclosure, an operation behavior recognition device is provided, applied to a target object server, wherein the method includes: a first determining module configured to, in response to a current operation behavior sent by a target object client, determine a first page image corresponding to the current operation behavior, and determine a second page image corresponding to the previous operation behavior of the current operation behavior; a second determining module configured to, based on the first page image and the second page image, determine a difference page image between the first page image and the second page image, and send the difference page image to a risk recognition server; and a receiving module configured to receive a behavior risk recognition result of the current operation behavior sent by the risk recognition server, wherein the behavior risk recognition result is determined based on an image risk recognition result, the image risk recognition result being obtained by the risk recognition server performing risk recognition on data of at least two modalities included in the difference page image according to a risk recognition model.

[0010] According to a fifth aspect of the present disclosure, an operation behavior recognition method is provided, applied to a risk recognition server, wherein the method includes: receiving a current operation behavior and a differential page image sent by a target object server in response to a target object client, wherein the differential page image is determined by the target object server based on a first page image corresponding to the current operation behavior and a second page image corresponding to the previous operation behavior; performing risk recognition on data of at least two modalities included in the differential page image according to a risk recognition model to obtain an image risk recognition result of the differential page image; and determining a behavior risk recognition result of the current operation behavior based on the image risk recognition result.

[0011] According to a sixth aspect of the present disclosure, an operation behavior recognition device is provided, applied to a risk recognition server, wherein the method includes: a receiving module configured to receive a current operation behavior and a differential page image sent by a target object server in response to a target object client, wherein the differential page image is determined by the target object server based on a first page image corresponding to the current operation behavior and a second page image corresponding to the previous operation behavior; an recognition module configured to perform risk recognition on data of at least two modalities included in the differential page image according to a risk recognition model, to obtain an image risk recognition result of the differential page image; and a determining module configured to determine a behavior risk recognition result of the current operation behavior based on the image risk recognition result.

[0012] According to a seventh aspect of the present disclosure, a computing device is provided, comprising: a memory and a processor, wherein the memory is used to store computer programs / instructions, and the processor is used to execute the computer programs / instructions, wherein the computer programs / instructions, when executed by the processor, implement the steps of the above-described operation behavior recognition method.

[0013] According to an eighth aspect of the present disclosure, a computer-readable storage medium is provided that stores a computer program / instructions that, when executed by a processor, implement the steps of the above-described operation behavior recognition method.

[0014] According to a ninth aspect of the present disclosure, a computer program product is provided, including a computer program / instructions that, when executed by a processor, implement the steps of the above-described operation behavior recognition method.

[0015] In one embodiment of the operation behavior recognition method disclosed herein, the target object server, in response to a current operation behavior sent by the target object client, determines a first page image corresponding to the current operation behavior and a second page image corresponding to the previous operation behavior. Based on the first and second page images, a difference page image is determined between the first and second page images, thereby obtaining a difference page image generated based on the target object client's current operation behavior that differs from the previous operation behavior. This difference page image is then sent to a risk recognition server. The risk recognition server performs risk recognition on the difference page image according to a risk recognition model, eliminating the need to perform risk recognition on the entire first page image, thus reducing the cost of using the risk recognition model and saving resources on the risk recognition server. Furthermore, the risk recognition model can also perform risk recognition on data from at least two modalities included in the difference page image, further ensuring the comprehensiveness and accuracy of risk recognition. This makes the subsequent behavior risk recognition results determined based on the image risk recognition results more accurate, avoiding missed detections and misjudgments of risky operation behaviors. Attached Figure Description

[0016] Figure 1 is a schematic diagram of an application scenario of an operation behavior recognition method provided in an embodiment of this disclosure;

[0017] Figure 2 is a flowchart of an operation behavior recognition method provided in an embodiment of this disclosure;

[0018] Figure 3 is a schematic diagram of the structure of an operation behavior recognition system in an operation behavior recognition method provided in an embodiment of this disclosure;

[0019] Figure 4 is a flowchart of the processing procedure of an operation behavior recognition method provided in an embodiment of this disclosure;

[0020] Figure 5 is a schematic diagram of the structure of an operation behavior recognition system provided in an embodiment of this disclosure;

[0021] Figure 6 is a flowchart of a second operation behavior recognition method provided in an embodiment of this disclosure;

[0022] Figure 7 is a schematic diagram of the structure of a second type of operation behavior recognition device provided in an embodiment of the present disclosure;

[0023] Figure 8 is a flowchart of a third operation behavior recognition method provided in an embodiment of this disclosure;

[0024] Figure 9 is a schematic diagram of the structure of a third type of operation behavior recognition device provided in an embodiment of this disclosure;

[0025] Figure 10 is a structural block diagram of a computing device provided in an embodiment of this disclosure. Detailed Implementation

[0026] Numerous specific details are set forth in the following description to provide a full understanding of this disclosure. However, this disclosure can be implemented in many other ways than those described herein, and those skilled in the art can make similar extensions without departing from the spirit of this disclosure. Therefore, this disclosure is not limited to the specific implementations disclosed below.

[0027] The terminology used in one or more embodiments of this disclosure is for the purpose of describing particular embodiments only and is not intended to be limiting of the one or more embodiments of this disclosure. The singular forms “a,” “the,” and “the” as used in one or more embodiments of this disclosure and the appended claims are also intended to include the plural forms unless the context clearly indicates otherwise. It should also be understood that the term “and / or” as used in one or more embodiments of this disclosure refers to and includes any or all possible combinations of one or more associated listed items.

[0028] It should be understood that although the terms first, second, etc., may be used to describe various information in one or more embodiments of this disclosure, such information should not be limited to these terms. These terms are only used to distinguish information of the same type from one another. For example, first may also be referred to as second without departing from the scope of one or more embodiments of this disclosure, and similarly, second may also be referred to as first. Depending on the context, the word “if” as used herein may be interpreted as “when”, “in response to a determination”, or “when…”.

[0029] Furthermore, it should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in one or more embodiments of this disclosure are all information and data authorized by the user or fully authorized by all parties. Moreover, the collection, use and processing of related data must comply with the relevant laws, regulations and standards of the relevant countries and regions, and corresponding operation entry points are provided for users to choose to authorize or refuse.

[0030] In one or more embodiments of this disclosure, a large model refers to a deep learning model with a large number of model parameters, typically containing hundreds of millions, tens of billions, hundreds of billions, trillions, or even tens of trillions of model parameters. A large model can also be called a foundation model. It is pre-trained using large-scale unlabeled corpora to produce a pre-trained model with hundreds of millions of parameters. Such models can adapt to a wide range of downstream tasks and have good generalization ability. Examples include Large Language Models (LLMs) and multi-modal pre-training models.

[0031] In practical applications, large models only require a small number of samples to fine-tune the pre-trained model before they can be applied to different tasks. Large models can be widely used in fields such as Natural Language Processing (NLP) and Computer Vision. Specifically, they can be applied to computer vision tasks such as Visual Question Answering (VQA), Image Captioning (IC), and Image Generation, as well as NLP tasks such as text-based sentiment classification, text summarization, and machine translation. The main application scenarios for large models include digital assistants, intelligent robots, search, online education, office software, e-commerce, and intelligent design.

[0032] First, the terms and concepts involved in one or more embodiments of this disclosure will be explained.

[0033] Cloud PC: An operating system environment that runs on the cloud, which clients connect to and use via a streaming protocol.

[0034] Multimodal large models are a machine learning technique based on deep learning that integrates different media data (such as text, images, audio, and video) and learns the relationships between different modalities to achieve more intelligent information processing and understanding.

[0035] Screen recording audit: Conduct behavioral analysis and monitoring of employees' screen recording using cloud computers, and audit the compliance of their computer use.

[0036] Session: Each time a user connects to and uses a cloud computer, the activities and operations that occur constitute a session.

[0037] This disclosure provides three methods for recognizing operational behaviors, and also relates to two devices for recognizing operational behaviors, an operational behavior recognition system, a computing device, a computer-readable storage medium, and a computer program product, which will be described in detail in the following embodiments.

[0038] Referring to Figure 1, Figure 1 illustrates an application scenario of an operation behavior recognition method according to an embodiment of this disclosure. The method specifically includes the following steps: The target object server, in response to a current operation behavior sent by the target object client, determines a first page image corresponding to the current operation behavior and a second page image corresponding to the previous operation behavior. Based on the first page image and the second page image, it determines a difference page image between the first page image and the second page image and sends the difference page image to the risk recognition server. The risk recognition server, based on a risk recognition model, performs risk recognition on the data of at least two modalities included in the difference page image to obtain an image risk recognition result for the difference page image. Based on the image risk recognition result, it determines the behavior risk recognition result for the current operation behavior.

[0039] In one embodiment of this disclosure, the target object server, in response to a current operation sent by the target object client, determines a first page image corresponding to the current operation and a second page image corresponding to the previous operation. Based on the first and second page images, a difference page image is determined between the first and second page images, thereby obtaining a difference page image generated based on the target object client's current operation that differs from the previous operation. This difference page image is then sent to a risk identification server. The risk identification server performs risk identification on the difference page image according to a risk identification model, eliminating the need to perform risk identification on the entire first page image, thus reducing the cost of using the risk identification model and saving resources on the risk identification server. Furthermore, the risk identification model can also perform risk identification on data from at least two modalities included in the difference page image, further ensuring the comprehensiveness and accuracy of risk identification. This makes the subsequent behavioral risk identification results determined based on the image risk identification results more accurate, avoiding missed detections and misjudgments of risky operations.

[0040] As shown in Figure 1, Figure 1 includes a target object client 102, a target object server 104, and a risk identification server 106.

[0041] In practice, the user can send the current operation behavior to the target object server 104 through the target object client 102. The target object server can determine the first page image corresponding to the current operation behavior, determine the second page image corresponding to the previous operation behavior, and determine the difference page image between the first page image and the second page image. The target object server 104 sends the difference page image to the risk identification server 106. The risk identification server 106 can perform risk identification on the difference page image according to the risk identification model, obtain the image risk identification result of the difference page image, and determine the behavior risk identification result of the current operation behavior according to the image risk identification result. The target object server 104 can determine whether to execute the control command of the current operation behavior according to the behavior risk identification result, and perform interactive control on the target object client 102 according to the control command to ensure the security of the user's access to the target object server.

[0042] Referring to Figure 2, which shows a flowchart of an operation behavior recognition method according to an embodiment of the present disclosure, applied to an operation behavior recognition system, the operation behavior recognition system includes a target object server and a risk recognition server, specifically including the following steps.

[0043] Step 202: The target object server, in response to the current operation sent by the target object client, determines the first page image corresponding to the current operation and determines the second page image corresponding to the previous operation.

[0044] Specifically, the operation behavior recognition method provided in this disclosure can be applied to the field of cloud computing. In the cloud computing field, a cloud computer is an operating system environment running in the cloud. Clients can connect to and use the cloud computer through a streaming protocol. The cloud computer can transfer the functions and data of a personal computer to a cloud server for operation. Users can connect to the virtual computer in the cloud through a network and use the computing resources and storage services provided by the cloud computer. Users are not limited by the hardware configuration of their local computer and only need simple terminal devices (such as laptops, tablets, smartphones, etc.) to access cloud computer services. Alternatively, the operation behavior recognition method provided in this disclosure can also be applied to the field of remote control. In this field, users can remotely control another client through one client, and the operation behavior recognition method can be used to identify risks in the user's remote control operation behavior.

[0045] For ease of understanding, this disclosure uses the application of the operation behavior recognition method in the field of cloud computing as an example for illustration.

[0046] In this context, the target object can be understood as a cloud computer. Therefore, the target object server is the cloud computer's server-side, and the target object client is the cloud computer's client-side. The current operation sent by the target object client can be understood as the operation performed by the user on the target object client at the current moment. This operation could be, for example, web browsing, game login, video download, or application launch. In practical applications, the current operation sent by the target object client could be a session initiated by the user. The previous operation can be understood as the operation performed by the user on the target object client at the previous moment. For example, if the user entered a webpage and then turned a page, the page turning is the current operation, and the webpage entering is the previous operation.

[0047] The first page image corresponding to the current operation can be understood as the display image composed of the page data corresponding to the current operation. For example, if the current operation is web browsing, then the page data corresponding to this web browsing operation is the web page data that the web browsing operation wants to browse, which can include text data, video data, and image data, etc. Therefore, the first page image corresponding to this web browsing operation can be understood as the display image rendered based on this web page data. Similarly, the second page image corresponding to the previous operation can be understood as the display image composed of the page data corresponding to the previous operation.

[0048] Based on this, the server side of the cloud computer can respond to the current operation sent by the user through the client of the cloud computer, determine the page data corresponding to the current operation, thereby obtaining the first page image corresponding to the current operation, and obtaining the second page image corresponding to the previous operation.

[0049] In specific implementation, determining the first page image corresponding to the current operation behavior in response to the current operation behavior sent by the target object client includes: determining the page data corresponding to the current operation behavior in response to the current operation behavior sent by the target object client; rendering the page data to obtain the first page image corresponding to the current operation behavior.

[0050] Specifically, the target client responds to the user's click command, determines the user's current operation, and sends the current operation to the target server. The target server responds to the current operation sent by the target client, retrieves the page data corresponding to the current operation from the database, renders the page data, and obtains the first page image corresponding to the current operation.

[0051] Referring to Figure 3, Figure 3 shows a schematic diagram of the structure of an operation behavior recognition system in an operation behavior recognition method according to an embodiment of the present disclosure. As shown in Figure 3, the operation behavior recognition system includes a target object server and a target object client. The target object client is the client of the cloud computer, which can be understood as a client connected to the cloud computer through a streaming protocol, and may include a soft terminal or a hardware terminal. The target object server can be understood as the environment running the cloud computer. The target object server may include an operating system environment, a streaming server, and an interactive control process. The target object client can send the current operation behavior to the target object server. The target object server responds to the current operation behavior, determines the page data corresponding to the current operation behavior, renders the page data in the operating system environment, and obtains the first page image corresponding to the current operation behavior by taking a screenshot of the operating system environment. The streaming server can send the first page image to the target object client according to the streaming protocol, so as to facilitate the subsequent display of the first page image on the display interface of the target object client.

[0052] In summary, by rendering the page data corresponding to the current operation on the target object server, the first page image corresponding to the current operation can be obtained, which facilitates subsequent risk identification and client-side streaming.

[0053] In one embodiment of this disclosure, after obtaining the first page image corresponding to the current operation behavior, the method further includes: sending the first page image to the target object client and displaying it through the display interface of the target object client.

[0054] Specifically, after obtaining the first page image, it can be transmitted to the target client according to the network transmission protocol so that it can be displayed through the target client's display interface.

[0055] In practice, before sending the first page image to the target client, the first page image can be analyzed, encoded, and compressed to reduce the amount of data transmitted.

[0056] In summary, by sending the first page image to the target client, a push streaming service to the target client is achieved, ensuring the user's experience when using the cloud computer.

[0057] In addition, after determining the first page image corresponding to the current operation behavior, the method further includes: storing the first page image in a database so that, in response to the next operation behavior, the first page image can be retrieved from the database to identify the behavior risk of the next operation behavior.

[0058] Specifically, after obtaining the first page image corresponding to the current operation, the first page image can be stored in the database. This allows the target object server to retrieve the first page image from the database in response to the next operation, and to perform behavioral risk identification for that next operation. It is understood that the process of identifying behavioral risks for the next operation is similar to the process of identifying behavioral risks for the current operation, and this embodiment will not repeat the details.

[0059] Therefore, when determining the second page image corresponding to the previous operation of the current operation, the second page image corresponding to the previous operation can be obtained from the database.

[0060] For example, when a user sends a webpage browsing action to a target server through a target client, the first page image corresponding to this webpage browsing action can be determined, stored in a database, and the second page image corresponding to the previous action (e.g., webpage login) can be retrieved from the database. Subsequent behavioral risk identification is then performed based on the first and second page images. Similarly, when the target server responds to the webpage browsing action sent by the target client with its next action (e.g., page turning), it can determine the page image corresponding to that page turning action and retrieve the first page image corresponding to the previous action (webpage browsing) from the database, enabling subsequent behavioral risk identification for that page turning action.

[0061] In practical applications, referring to Figure 3 above, the operation behavior recognition system also includes a bypass server. The bypass server can be understood as a server within the cloud computer that receives real-time desktop screen streams. The risk recognition server can run as a risk recognition process within this bypass server. The target object server can push the first page image to the bypass stream via the streaming service, and then store the first page image in the database through the bypass stream. The first page image can be understood as a screenshot of the operating system environment within the cloud computer, used to represent real-time changes to the desktop screen. Furthermore, the database can be understood as a cloud-based object storage system. The database can be based on object storage and can be used to store screen recording files and screen recording audit files in specific formats, such as extended formats. Correspondingly, the first page image can be a screen recording video frame, and the first page image can be stored in the database as a screen recording file. In practical applications, this operation behavior recognition method can be applied to enterprise monitoring of employee operations. The database can also store an enterprise knowledge base, which can be stored in the form of a vector database.

[0062] In summary, by storing the page image corresponding to each operation in a database, it is easier to retrieve page images from the database for behavioral risk identification, further improving the efficiency of risk identification. Furthermore, the risk identification server runs on a bypass server, i.e., a streaming server in the cloud, preventing malicious damage or bypassing of monitoring in the user environment and ensuring the stability of the user environment and behavioral risk identification.

[0063] Step 204: Based on the first page image and the second page image, determine the difference page image between the first page image and the second page image, and send the difference page image to the risk identification server.

[0064] Here, a difference page image can be understood as an image that differs from the first page image and the second page image. For example, if the first page image includes text 1, text 2 and image 1, and the second page image includes text 1, text 2 and image 2, then the page image containing image 2 in the second page image is the difference page image.

[0065] In practical applications, the difference page image between the first page image and the second page image refers to the difference points that exist between the two images. These difference points can be at the pixel level. The pixel-level differences between the first page image and the second page image can be calculated according to the streaming protocol, and these pixel-level differences can be used as the difference page image. By determining the difference page image, it is possible to meet the risk identification requirements of the multimodal data contained in the difference page image, and it is also possible to save the number of training labels for the subsequent multimodal risk identification model.

[0066] In a specific implementation, determining the difference page image between the first page image and the second page image based on the first page image and the second page image includes: calculating the image similarity between the first page image and the second page image; and determining the difference page image between the first page image and the second page image based on the image similarity.

[0067] Specifically, the image similarity between the first page image and the second page image can be calculated using an image similarity calculation algorithm. Based on the image similarity, the images in the first page image and the second page image whose image similarity meets a preset similarity threshold are determined as the difference page images between the first page image and the second page image.

[0068] In summary, by determining the differences in page images based on image similarity, subsequent risk identification does not require submitting the entire screen content; only the screen content with differences needs to be submitted. This reduces the portion that the risk identification model needs to identify, thereby saving on the number of times the risk identification model needs to be called and the number of training labels, thus reducing the cost of use.

[0069] In specific implementation, the difference page image between the first page image and the second page image can be determined using hash algorithms, histogram algorithms, or pixel comparison methods. For example, the first and second page images can be reduced in size and converted to grayscale. The difference between each pixel and the average value in the first and second page images can be calculated separately. The difference can be converted into a binary value using a hash function. The difference page image can be determined based on the comparison between the hash values ​​of the first and second page images. Alternatively, the difference between adjacent pixels in the first and second page images can be calculated separately. A hash value can be generated based on the difference. The difference page image can be determined based on the comparison between the hash values ​​of the first and second page images. Alternatively, the difference page image can also be determined based on parameters such as the structural similarity index, cosine similarity, or Euclidean distance between the first and second page images. This disclosure does not limit the specific methods used.

[0070] In one embodiment of this disclosure, calculating the image similarity between the first page image and the second page image includes: determining a plurality of first pixels in the first page image and determining a plurality of second pixels in the second page image at the same positions as the plurality of first pixels; calculating the pixel similarity between the first pixels and the second pixels; and determining the difference page image between the first page image and the second page image based on the image similarity includes: determining the second pixels with a pixel similarity less than a preset similarity threshold as target pixels, and determining the difference page image based on the target pixels.

[0071] The preset similarity threshold can be understood as a pre-set similarity threshold. If the pixel similarity is less than the preset similarity threshold, it means that the first pixel and the second pixel are not similar and there is a difference.

[0072] Based on this, multiple first pixels in the first page image can be identified, and multiple second pixels in the second page image that are in the same position as the multiple first pixels can be identified. According to the similarity calculation algorithm, the pixel similarity between the first pixels and the second pixels in the same position is calculated. The second pixels with a pixel similarity less than a preset similarity threshold are identified as target pixels. The difference page image is constructed based on the target pixels.

[0073] For example, we can determine that the first pixel A corresponds to position 1 in the first page image and the second pixel B corresponds to position 1 in the second page image. Since the image positions of the first pixel A and the second pixel B are the same, we can calculate the pixel similarity between the first pixel A and the second pixel B. If the pixel similarity is less than a preset similarity threshold, the second pixel B is determined as the target pixel. Similarly, we perform the same calculation on other pixels in the same position in the first page image and the second page image, and construct the difference page image based on the target pixel.

[0074] In summary, by calculating pixel similarity, the differences in page images can be determined, ensuring the accuracy of the determined differences in page images.

[0075] In practical applications, before sending the difference page image to the risk identification server, the method further includes: determining the data of at least two modalities included in the difference page image; sending the difference page image to the risk identification server includes: encapsulating the difference page image and the data of the at least two modalities into a data frame and sending it to the risk identification server.

[0076] At least two modalities can include text, image, and video modalities. The difference page image and the data from at least two modalities can be understood as the content updated in real-time on the target client's screen. The data frame can be a streaming data frame.

[0077] Specifically, the data of at least two modalities included in the difference page image can be identified, and the difference page data and the data of the at least two modalities can be encapsulated into a data frame and sent to the risk identification server.

[0078] In practical applications, referring to Figure 3 above, the differential page image and data from at least two modalities can be encapsulated into streaming data frames using a bypass stream and sent to the risk identification server. Specifically, the differential page image and data from at least two modalities can be converted to the target transmission format, and the encapsulation format can be determined according to the network transmission protocol. The encapsulation format could be, for example, an Ethernet frame format. Frame headers and trailers are added before and after the differential page image and data from at least two modalities. The frame header can include control information, such as synchronization and address information, while the frame trailer can include checksum information to detect whether errors have occurred during data transmission, thus completing the encapsulation of the differential page image and data from at least two modalities.

[0079] In summary, by encapsulating the data into data frames and sending them to the risk identification server, the amount of data transmitted each time is reduced, the integrity of each data transmission is ensured, and subsequent behavioral risk identification of the data frames is facilitated.

[0080] Step 206: The risk identification server performs risk identification on the data of at least two modalities included in the difference page image according to the risk identification model, and obtains the image risk identification result of the difference page image.

[0081] The image risk identification result can be understood as the result of whether there is a risk in the differing page image. For example, if there is abnormal information in the differing page image, the image risk identification result is that the differing page image poses a risk. The image risk identification result can also include the type of risk present in the differing page image, which may include data leakage risk, browsing abnormal websites risk, accessing prohibited applications risk, etc. The risk identification model can be understood as a multimodal risk identification model. This multimodal risk identification model can identify risks in multiple modalities of data, such as text data, image data, and video data. This multimodal risk identification model has text understanding, image understanding, and video understanding capabilities.

[0082] Specifically, the risk identification server can perform risk identification on the difference page image and the data of at least two modalities included in the difference page image based on the risk identification model, and obtain the image risk identification result of the difference page.

[0083] For example, if a discrepancy page image includes text data, image data, and video data, then the discrepancy page image can be input into a risk identification model to obtain the image risk identification result for the discrepancy page image. Alternatively, the text data, image data, and video data included in the discrepancy page image can be input into the risk identification model to obtain the image risk identification result for the discrepancy page image. Or, the discrepancy page image itself, along with the text data, image data, and video data included in the discrepancy page image, can be input into the risk identification model to obtain the image risk identification result for the discrepancy page image.

[0084] In one embodiment of this disclosure, there can be multiple risk identification models, such as text risk identification models, image risk identification models, and video risk identification models. These multiple risk identification models can be used to perform risk identification on data from at least two modalities to obtain the image risk identification result for the difference page image. For example, if it is determined that the difference page image includes both text data and image data, the text risk identification model can be used to identify the risk of the text data, and the image risk identification model can be used to identify the risk of the image data, thereby obtaining the image risk identification result for the difference page image.

[0085] In one embodiment of this disclosure, the risk identification model is a multimodal machine learning model. In another embodiment of this disclosure, the risk identification model can be a large multimodal model. It is understood that the risk identification model can be subjected to supervised or unsupervised training before risk identification is performed based on it.

[0086] In specific implementation, the risk identification server, based on the risk identification model, performs risk identification on the data of at least two modalities included in the difference page image to obtain the image risk identification result of the difference page image, including:

[0087] The risk identification server invokes the risk identification model through the model invocation interface, and performs risk identification on the data of at least two modalities included in the difference page image according to the invoked risk identification model, so as to obtain the image risk identification result of the difference page image.

[0088] Specifically, the risk identification server can invoke the risk identification model through its deployed model call interface. The model inputs the difference page image and data from at least two modalities it contains to perform risk identification, obtaining the image risk identification result of the difference page image as output by the model. Furthermore, the risk identification model can also be deployed on the risk identification server, allowing the server to directly use the model.

[0089] In summary, by introducing a risk identification and processing process into the streaming server and utilizing multimodal large model capabilities, the understanding and identification of user session content is enhanced, without relying on user environment, operation sequences, or application software tracking points for data collection.

[0090] Step 208: Based on the image risk identification result, determine the behavior risk identification result of the current operation behavior.

[0091] Among them, the behavioral risk identification result can be understood as the result of whether there is a risk in the current operation behavior. The behavioral risk identification result may also include the risk type of the current operation behavior.

[0092] In specific implementation, determining the behavioral risk identification result of the current operation based on the image risk identification result includes: matching the image risk identification result with a reference behavioral rule based on the image risk identification result to obtain a matching result, and determining the behavioral risk identification result of the current operation based on the matching result; wherein, the reference behavioral rule is a predefined risk behavior rule.

[0093] Specifically, the image risk identification results can be matched with predefined risk behavior rules to obtain matching results. Based on the matching results, it can be determined whether the image risk identification results match the predefined risk behavior rules, thereby determining the behavioral risk identification results of the current operation.

[0094] In practical applications, as shown in Figure 3, the risk identification server can be accessed through a third-party configuration system to configure reference behavior rules. Specifically, reference behavior rules can include content auditing, action rules, and an enterprise knowledge base. Enterprises can customize configurations for content auditing, prompts and actions, and the enterprise knowledge base according to project needs. Content auditing can include monitoring whether employees install prohibited applications, send data files to external systems, or detect abnormal files. Action rules can include enterprise-defined risk behaviors, such as deleting files.

[0095] In summary, by integrating with a third-party configuration system, enterprises are allowed to customize configuration actions and update their enterprise knowledge base. This enables risk identification of employee misconduct, data leaks, and unauthorized software installations. It integrates an understanding of enterprise knowledge into the general knowledge understanding of the large model, enhancing the retrieval capabilities of external knowledge bases within the large model and demonstrating strong scalability. Furthermore, direct connection to the large model and enterprise knowledge base on the cloud server minimizes the computational power and desktop performance overhead required by the target client, reducing the impact on user experience.

[0096] In practical applications, after determining the behavioral risk identification result of the current operation, the method further includes: the risk identification server determining a control instruction for the current operation based on the behavioral risk identification result of the current operation, and sending the control instruction to the target object server; and the target object server processing the current operation based on the control instruction.

[0097] Control commands can include alarm commands, execution commands, and interception commands. When the control command is an alarm command, the current operation can be skipped, an alarm message can be generated, and sent to the target client. When the control command is an execution command, the current operation can be executed directly. When the control command is an interception command, the current operation can be intercepted.

[0098] In practical applications, referring to Figure 3 above, the risk identification server can determine control instructions for the current operational behavior based on the behavioral risk identification results, and send the control instructions to the target object server. The interactive control process within the target object server then processes the current operational behavior according to the control instructions. Furthermore, the risk identification server can also connect to a monitoring and alarm system. This system can generate and send alarm information for high-risk behaviors. This monitoring and alarm system can be integrated into the project system, and upon detecting high-risk behaviors, it can send notifications to project personnel via SMS or email.

[0099] In summary, by intervening and controlling current operational behaviors, the security of users' use of cloud computers can be ensured.

[0100] In one embodiment of this disclosure, after determining the first page image corresponding to the current operation behavior sent by the target object client, the method further includes: the target object server, if it determines that there is no previous operation behavior of the current operation behavior, sending the first page image to the risk identification server; the risk identification server, according to a risk identification model, performing risk identification on the data of at least two modalities included in the first page image to obtain a first image risk identification result of the first page image, and determining the behavior risk identification result of the current operation behavior based on the first image risk identification result.

[0101] The first image risk identification result of the first page image can be understood as the image risk identification result of the first page image.

[0102] Specifically, if the target object server determines that there is no previous operation behavior, it indicates that the current operation behavior is the first operation behavior initiated by the user. At this time, the first page image corresponding to the current operation behavior can be directly sent to the risk identification server. The risk identification server can perform risk identification on the data of at least two modalities included in the first page image according to the risk identification model, obtain the first image risk identification result of the first page image, and determine the behavior risk identification result of the current operation behavior based on the first image risk identification result.

[0103] In one embodiment of this disclosure, the target object server, in response to a current operation sent by the target object client, determines a first page image corresponding to the current operation and a second page image corresponding to the previous operation. Based on the first and second page images, a difference page image is determined between the first and second page images, thereby obtaining a difference page image generated based on the target object client's current operation that differs from the previous operation. This difference page image is then sent to a risk identification server. The risk identification server performs risk identification on the difference page image according to a risk identification model, eliminating the need to perform risk identification on the entire first page image, thus reducing the cost of using the risk identification model and saving resources on the risk identification server. Furthermore, the risk identification model can also perform risk identification on data from at least two modalities included in the difference page image, further ensuring the comprehensiveness and accuracy of risk identification. This makes the subsequent behavioral risk identification results determined based on the image risk identification results more accurate, avoiding missed detections and misjudgments of risky operations.

[0104] The following description, in conjunction with Figure 4, uses the application of the operation behavior recognition method provided in this disclosure in a cloud computer as an example to further illustrate the operation behavior recognition method. Figure 4 shows a flowchart of the processing procedure of an operation behavior recognition method according to an embodiment of this disclosure, specifically including the following steps.

[0105] Step 402: The target object server, in response to the current operation sent by the target object client, determines the first page image corresponding to the current operation and determines the second page image corresponding to the previous operation.

[0106] Specifically, the target client is the client of the cloud computer. This client can be understood as a client connected to the cloud computer via a streaming protocol, and can include a soft terminal or a hardware terminal. The target server can be understood as the environment running the cloud computer. The target server can include an operating system environment, a streaming server, and an interactive control process. The target client can send its current operation to the target server. The target server responds to the current operation by determining the page data corresponding to the current operation, rendering the page data in the operating system environment, and obtaining the first page image corresponding to the current operation by taking a screenshot of the operating system environment. The streaming server can then send the first page image to the target client according to the streaming protocol, facilitating its subsequent display on the target client's interface.

[0107] The cloud computer's server can respond to the current operation sent by the user through the cloud computer's client, determine the page data corresponding to the current operation, thereby obtaining the first page image corresponding to the current operation, and obtaining the second page image corresponding to the previous operation.

[0108] Step 404: The target object server determines the difference page image between the first page image and the second page image based on the first page image and the second page image.

[0109] Specifically, multiple first pixels in the first page image can be identified, and multiple second pixels in the second page image that are in the same position as the multiple first pixels can be identified. The pixel similarity between the first pixels and the second pixels in the same position can be calculated. The second pixels with a pixel similarity less than a preset similarity threshold can be identified as target pixels. The difference page image can be constructed based on the target pixels.

[0110] For example, we can determine that the first pixel A corresponds to position 1 in the first page image and the second pixel B corresponds to position 1 in the second page image. Since the image positions of the first pixel A and the second pixel B are the same, we can calculate the pixel similarity between the first pixel A and the second pixel B. If the pixel similarity is less than a preset similarity threshold, the second pixel B is determined as the target pixel. Similarly, we perform the same calculation on other pixels in the same position in the first page image and the second page image, and construct the difference page image based on the target pixel.

[0111] Step 406: The target object server encapsulates the difference page image and the data of at least two modalities included in the difference page image into a data frame and sends it to the risk identification server.

[0112] Step 408: The risk identification server performs risk identification on the data of at least two modalities included in the difference page image according to the risk identification model, and obtains the image risk identification result of the difference page image.

[0113] Step 410: The risk identification server determines the behavior risk identification result of the current operation based on the image risk identification result.

[0114] Step 412: The risk identification server determines the control instructions for the current operation behavior based on the behavior risk identification results, and sends the control instructions to the target object server.

[0115] In one embodiment of this disclosure, the target object server, in response to a current operation sent by the target object client, determines a first page image corresponding to the current operation and a second page image corresponding to the previous operation. Based on the first and second page images, a difference page image is determined between the first and second page images, thereby obtaining a difference page image generated based on the target object client's current operation that differs from the previous operation. This difference page image is then sent to a risk identification server. The risk identification server performs risk identification on the difference page image according to a risk identification model, eliminating the need to perform risk identification on the entire first page image, thus reducing the cost of using the risk identification model and saving resources on the risk identification server. Furthermore, the risk identification model can also perform risk identification on data from at least two modalities included in the difference page image, further ensuring the comprehensiveness and accuracy of risk identification. This makes the subsequent behavioral risk identification results determined based on the image risk identification results more accurate, avoiding missed detections and misjudgments of risky operations.

[0116] Corresponding to the above method embodiments, this disclosure also provides an embodiment of an operation behavior recognition system. Figure 5 shows a schematic diagram of the structure of an operation behavior recognition system provided by an embodiment of this disclosure. As shown in Figure 5, the system 500 includes: a target object server 502 and a risk recognition server 504. The target object server 502 is configured to, in response to a current operation behavior sent by a target object client, determine a first page image corresponding to the current operation behavior, and determine a second page image corresponding to the previous operation behavior of the current operation behavior. Based on the first page image and the second page image, it determines a difference page image between the first page image and the second page image, and sends the difference page image to the risk recognition server 504. The risk recognition server 504 is configured to, based on a risk recognition model, perform risk recognition on data of at least two modalities included in the difference page image to obtain an image risk recognition result of the difference page image, and determine a behavior risk recognition result of the current operation behavior based on the image risk recognition result.

[0117] In an optional embodiment, the target object server is further configured to send the first page image to the risk identification server when it is determined that there is no previous operation behavior before the current operation behavior; the risk identification server is further configured to perform risk identification on the data of at least two modalities included in the first page image according to the risk identification model, obtain a first image risk identification result of the first page image, and determine the behavior risk identification result of the current operation behavior according to the first image risk identification result.

[0118] In an optional embodiment, the risk identification server is further configured to: match the image risk identification result with a reference behavior rule based on the image risk identification result, obtain a matching result, and determine the behavior risk identification result of the current operation behavior based on the matching result; wherein the reference behavior rule is a predefined risk behavior rule.

[0119] In an optional embodiment, the risk identification server is further configured to: determine a control instruction for the current operation based on the behavior risk identification result of the current operation, and send the control instruction to the target object server; the target object server is further configured to: process the current operation based on the control instruction.

[0120] In an optional embodiment, the target object server is further configured to: calculate the image similarity between the first page image and the second page image; and determine the difference page image between the first page image and the second page image based on the image similarity.

[0121] In an optional embodiment, the target object server is further configured to: determine a plurality of first pixels in the first page image, and determine a plurality of second pixels in the second page image at the same position as the plurality of first pixels; calculate the pixel similarity between the first pixels and the second pixels; and determine the difference page image between the first page image and the second page image based on the image similarity, comprising: determining the second pixels with a pixel similarity less than a preset similarity threshold as target pixels, and determining the difference page image based on the target pixels.

[0122] In an optional embodiment, the target object server is further configured to: determine data of at least two modalities included in the difference page image; encapsulate the difference page image and the data of the at least two modalities into a data frame and send it to the risk identification server.

[0123] In an optional embodiment, the risk identification server is further configured to invoke the risk identification model through a model invocation interface, and perform risk identification on the data of at least two modalities included in the difference page image according to the invoked risk identification model, so as to obtain the image risk identification result of the difference page image.

[0124] In an optional embodiment, the target object server is further configured to: respond to the current operation behavior sent by the target object client, determine the page data corresponding to the current operation behavior, render the page data, and obtain the first page image corresponding to the current operation behavior.

[0125] In an optional embodiment, the target object server is further configured to: send the first page image to the target object client and display it through the display interface of the target object client.

[0126] In an optional embodiment, the target object server is further configured to: store the first page image in a database, so that in response to the next operation behavior, it retrieves the first page image from the database to perform behavioral risk identification for the next operation behavior. In an optional embodiment, the risk identification model is a multimodal machine learning model.

[0127] In one optional embodiment, the target object server is the server of the cloud computer, and the target object client is the client of the cloud computer.

[0128] In one embodiment of this disclosure, the target object server, in response to a current operation sent by the target object client, determines a first page image corresponding to the current operation and a second page image corresponding to the previous operation. Based on the first and second page images, a difference page image is determined between the first and second page images, thereby obtaining a difference page image generated based on the target object client's current operation that differs from the previous operation. This difference page image is then sent to a risk identification server. The risk identification server performs risk identification on the difference page image according to a risk identification model, eliminating the need to perform risk identification on the entire first page image, thus reducing the cost of using the risk identification model and saving resources on the risk identification server. Furthermore, the risk identification model can also perform risk identification on data from at least two modalities included in the difference page image, further ensuring the comprehensiveness and accuracy of risk identification. This makes the subsequent behavioral risk identification results determined based on the image risk identification results more accurate, avoiding missed detections and misjudgments of risky operations.

[0129] The above is an illustrative scheme of an operation behavior recognition system according to this embodiment. It should be noted that the technical solution of this operation behavior recognition system and the technical solution of the operation behavior recognition method described above belong to the same concept. For details not described in detail in the technical solution of the operation behavior recognition system, please refer to the description of the technical solution of the operation behavior recognition method described above.

[0130] Referring to Figure 6, Figure 6 shows a flowchart of a second operation behavior recognition method provided according to an embodiment of the present disclosure, which is applied to a target object server and specifically includes the following steps.

[0131] Step 602: In response to the current operation sent by the target object client, determine the first page image corresponding to the current operation and determine the second page image corresponding to the previous operation.

[0132] Step 604: Based on the first page image and the second page image, determine the difference page image between the first page image and the second page image, and send the difference page image to the risk identification server.

[0133] Step 606: Receive the behavior risk identification result of the current operation behavior sent by the risk identification server, wherein the behavior risk identification result is determined based on the image risk identification result, and the image risk identification result is obtained by the risk identification server based on the risk identification model to perform risk identification on the data of at least two modalities included in the difference page image.

[0134] In one embodiment of this disclosure, the target object server, in response to a current operation sent by the target object client, determines a first page image corresponding to the current operation and a second page image corresponding to the previous operation. Based on the first and second page images, a difference page image is determined between the first and second page images, thereby obtaining a difference page image generated based on the target object client's current operation that differs from the previous operation. This difference page image is then sent to a risk identification server. The risk identification server performs risk identification on the difference page image according to a risk identification model, eliminating the need to perform risk identification on the entire first page image, thus reducing the cost of using the risk identification model and saving resources on the risk identification server. Furthermore, the risk identification model can also perform risk identification on data from at least two modalities included in the difference page image, further ensuring the comprehensiveness and accuracy of risk identification. This makes the subsequent behavioral risk identification results determined based on the image risk identification results more accurate, avoiding missed detections and misjudgments of risky operations.

[0135] The above is an illustrative scheme of an operation behavior recognition method according to this embodiment. It should be noted that the technical solution of this operation behavior recognition method belongs to the same concept as the operation behavior recognition method described above. For details not described in detail in the technical solution of the operation behavior recognition method, please refer to the description of the technical solution of the operation behavior recognition method described above.

[0136] Corresponding to the above method embodiments, this disclosure also provides an embodiment of an operation behavior recognition device, applied to a target object server. Figure 7 shows a schematic diagram of the structure of a second operation behavior recognition device provided in an embodiment of this disclosure. As shown in Figure 7, the device includes the following modules.

[0137] The first determining module 702 is configured to, in response to the current operation behavior sent by the target object client, determine the first page image corresponding to the current operation behavior, and determine the second page image corresponding to the previous operation behavior of the current operation behavior.

[0138] The second determining module 704 is configured to determine the difference page image between the first page image and the second page image based on the first page image and the second page image, and send the difference page image to the risk identification server.

[0139] The receiving module 706 is configured to receive the behavior risk identification result of the current operation behavior sent by the risk identification server, wherein the behavior risk identification result is determined based on the image risk identification result, and the image risk identification result is obtained by the risk identification server through risk identification of at least two modalities included in the difference page image according to the risk identification model.

[0140] In one embodiment of this disclosure, the target object server, in response to a current operation sent by the target object client, determines a first page image corresponding to the current operation and a second page image corresponding to the previous operation. Based on the first and second page images, a difference page image is determined between the first and second page images, thereby obtaining a difference page image generated based on the target object client's current operation that differs from the previous operation. This difference page image is then sent to a risk identification server. The risk identification server performs risk identification on the difference page image according to a risk identification model, eliminating the need to perform risk identification on the entire first page image, thus reducing the cost of using the risk identification model and saving resources on the risk identification server. Furthermore, the risk identification model can also perform risk identification on data from at least two modalities included in the difference page image, further ensuring the comprehensiveness and accuracy of risk identification. This makes the subsequent behavioral risk identification results determined based on the image risk identification results more accurate, avoiding missed detections and misjudgments of risky operations.

[0141] The above is an illustrative scheme of an operation behavior recognition device according to this embodiment. It should be noted that the technical solution of this operation behavior recognition device and the technical solution of the operation behavior recognition method described above belong to the same concept. For details not described in detail in the technical solution of the operation behavior recognition device, please refer to the description of the technical solution of the operation behavior recognition method described above.

[0142] Referring to Figure 8, Figure 8 shows a flowchart of a third operational behavior recognition method provided according to an embodiment of the present disclosure, which is applied to a risk recognition server and specifically includes the following steps.

[0143] Step 802: Receive the current operation behavior and the difference page image sent by the target object server in response to the target object client, wherein the difference page image is determined by the target object server based on the first page image corresponding to the current operation behavior and the second page image corresponding to the previous operation behavior.

[0144] Step 804: Based on the risk identification model, perform risk identification on the data of at least two modalities included in the difference page image to obtain the image risk identification result of the difference page image.

[0145] Step 806: Determine the behavioral risk identification result of the current operation behavior based on the image risk identification result.

[0146] In one embodiment of this disclosure, the target object server, in response to a current operation sent by the target object client, determines a first page image corresponding to the current operation and a second page image corresponding to the previous operation. Based on the first and second page images, a difference page image is determined between the first and second page images, thereby obtaining a difference page image generated based on the target object client's current operation that differs from the previous operation. This difference page image is then sent to a risk identification server. The risk identification server performs risk identification on the difference page image according to a risk identification model, eliminating the need to perform risk identification on the entire first page image, thus reducing the cost of using the risk identification model and saving resources on the risk identification server. Furthermore, the risk identification model can also perform risk identification on data from at least two modalities included in the difference page image, further ensuring the comprehensiveness and accuracy of risk identification. This makes the subsequent behavioral risk identification results determined based on the image risk identification results more accurate, avoiding missed detections and misjudgments of risky operations.

[0147] The above is an illustrative scheme of an operation behavior recognition method according to this embodiment. It should be noted that the technical solution of this operation behavior recognition method belongs to the same concept as the operation behavior recognition method described above. For details not described in detail in the technical solution of the operation behavior recognition method, please refer to the description of the technical solution of the operation behavior recognition method described above.

[0148] Corresponding to the above method embodiments, this disclosure also provides an embodiment of an operation behavior recognition device, applied to a risk recognition server. Figure 9 shows a schematic diagram of the structure of a third operation behavior recognition device provided in an embodiment of this disclosure. As shown in Figure 9, the device includes the following modules.

[0149] The receiving module 902 is configured to receive the current operation behavior and the difference page image sent by the target object server in response to the target object client, wherein the difference page image is determined by the target object server based on the first page image corresponding to the current operation behavior and the second page image corresponding to the previous operation behavior.

[0150] The identification module 904 is configured to perform risk identification on the data of at least two modalities included in the difference page image according to the risk identification model, and obtain the image risk identification result of the difference page image.

[0151] The determination module 906 is configured to determine the behavior risk identification result of the current operation behavior based on the image risk identification result.

[0152] In one embodiment of this disclosure, the target object server, in response to a current operation sent by the target object client, determines a first page image corresponding to the current operation and a second page image corresponding to the previous operation. Based on the first and second page images, a difference page image is determined between the first and second page images, thereby obtaining a difference page image generated based on the target object client's current operation that differs from the previous operation. This difference page image is then sent to a risk identification server. The risk identification server performs risk identification on the difference page image according to a risk identification model, eliminating the need to perform risk identification on the entire first page image, thus reducing the cost of using the risk identification model and saving resources on the risk identification server. Furthermore, the risk identification model can also perform risk identification on data from at least two modalities included in the difference page image, further ensuring the comprehensiveness and accuracy of risk identification. This makes the subsequent behavioral risk identification results determined based on the image risk identification results more accurate, avoiding missed detections and misjudgments of risky operations.

[0153] The above is an illustrative scheme of an operation behavior recognition device according to this embodiment. It should be noted that the technical solution of this operation behavior recognition device and the technical solution of the operation behavior recognition method described above belong to the same concept. For details not described in detail in the technical solution of the operation behavior recognition device, please refer to the description of the technical solution of the operation behavior recognition method described above.

[0154] Figure 10 shows a structural block diagram of a computing device 1000 according to an embodiment of the present disclosure. The components of the computing device 1000 include, but are not limited to, a memory 1010 and a processor 1020. The processor 1020 is connected to the memory 1010 via a bus 1030, and a database 1050 is used to store data.

[0155] The computing device 1000 also includes an access device 1040, which enables the computing device 1000 to communicate via one or more networks 1060. Examples of these networks include Public Switched Telephone Network (PSTN), Local Area Network (LAN), Wide Area Network (WAN), Personal Area Network (PAN), or combinations of communication networks such as the Internet. The access device 1040 may include one or more of any type of wired or wireless network interface (e.g., a network interface controller (NIC)), such as an IEEE 802.11 Wireless Local Area Network (WLAN) wireless interface, a Wi-MAX (Worldwide Interoperability for Microwave Access) interface, an Ethernet interface, a Universal Serial Bus (USB) interface, a cellular network interface, a Bluetooth interface, or a Near Field Communication (NFC) interface.

[0156] In one embodiment of this disclosure, the aforementioned components of the computing device 1000, as well as other components not shown in FIG. 10, may be interconnected, for example, via a bus. It should be understood that the computing device block diagram shown in FIG. 10 is merely for illustrative purposes and is not intended to limit the scope of this disclosure. Those skilled in the art can add or replace other components as needed.

[0157] The computing device 1000 can be any type of stationary or mobile computing device, including mobile computers or mobile computing devices (e.g., tablet computers, personal digital assistants, laptop computers, notebook computers, netbooks, etc.), mobile phones (e.g., smartphones), wearable computing devices (e.g., smartwatches, smart glasses, etc.) or other types of mobile devices, or stationary computing devices such as desktop computers or personal computers (PCs). The computing device 1000 can also be a mobile or stationary server.

[0158] The processor 1020 is used to execute the following computer program / instruction, which, when executed by the processor, implements the steps of the above-mentioned operation behavior recognition method.

[0159] The various embodiments in this disclosure are described in a progressive manner. Similar or identical parts between embodiments can be referred to mutually. Each embodiment focuses on describing the differences from other embodiments. In particular, the computing device embodiments are basically similar to the operation behavior recognition method embodiments, so the description is relatively simple; relevant parts can be referred to in the description of the operation behavior recognition method embodiments.

[0160] An embodiment of this disclosure also provides a computer-readable storage medium storing a computer program / instructions that, when executed by a processor, implement the steps of the above-described operation behavior recognition method.

[0161] The various embodiments in this disclosure are described in a progressive manner. Similar or identical parts between embodiments can be referred to mutually. Each embodiment focuses on describing the differences from other embodiments. In particular, the computer-readable storage medium embodiments are basically similar to the operation behavior recognition method embodiments, so the description is relatively simple; relevant parts can be referred to in the description of the operation behavior recognition method embodiments.

[0162] An embodiment of this disclosure also provides a computer program product, including a computer program / instructions that, when executed by a processor, implement the steps of the above-described operation behavior recognition method.

[0163] The above is an illustrative scheme of a computer program product according to this embodiment. It should be noted that the technical solution of this computer program product and the technical solution of the above-described operation behavior recognition method belong to the same concept. For details not described in detail in the technical solution of the computer program product, please refer to the description of the technical solution of the above-described operation behavior recognition method.

[0164] The foregoing has described specific embodiments of this disclosure. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims may be performed in a different order than that shown in the embodiments and may still achieve the desired results. Furthermore, the processes depicted in the drawings do not necessarily require the specific or sequential order shown to achieve the desired results. In some embodiments, multitasking and parallel processing are also possible or may be advantageous.

[0165] The computer instructions include computer program code, which may be in the form of source code, object code, executable file, or certain intermediate forms. The computer-readable medium may include: any entity or device capable of carrying the computer program code, recording media, USB flash drive, portable hard drive, magnetic disk, optical disk, computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signals, telecommunication signals, and software distribution media, etc. It should be noted that the content included in the computer-readable medium may be appropriately added or removed according to the requirements of patent practice. For example, in some regions, according to patent practice, computer-readable media may not include electrical carrier signals and telecommunication signals.

[0166] It should be noted that, for the sake of simplicity, the foregoing method embodiments are all described as a series of actions. However, those skilled in the art should understand that the embodiments of this disclosure are not limited to the described order of actions, because according to the embodiments of this disclosure, some steps can be performed in other orders or simultaneously. Furthermore, those skilled in the art should also understand that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily essential to the embodiments of this disclosure.

[0167] In the above embodiments, the descriptions of each embodiment have different focuses. For parts not described in detail in a certain embodiment, please refer to the relevant descriptions of other embodiments.

[0168] The preferred embodiments disclosed above are merely illustrative of this disclosure. The optional embodiments do not exhaustively describe all details, nor do they limit the invention to the specific implementations described. Clearly, many modifications and variations can be made based on the embodiments of this disclosure. These embodiments are selected and specifically described in this disclosure to better explain the principles and practical applications of the embodiments of this disclosure, thereby enabling those skilled in the art to better understand and utilize this disclosure. This disclosure is limited only by the claims and their full scope and equivalents.

Claims

1. An operating behavior recognition method applied to an operating behavior recognition system, the operating behavior recognition system comprising a target object server and a risk identification server, wherein, The method comprises: The target object server determines a first page image corresponding to the current operation behavior and a second page image corresponding to a previous operation behavior of the current operation behavior in response to the current operation behavior sent by the target object client, determines a difference page image between the first page image and the second page image according to the first page image and the second page image, and sends the difference page image to the risk identification server; The risk identification server performs risk identification on data of at least two modalities included in the difference page image according to a risk identification model to obtain an image risk identification result of the difference page image, determines a behavior risk identification result of the current operation behavior according to the image risk identification result.

2. The operation behavior identification method of claim 1, after determining the first page image corresponding to the current operation behavior in response to the current operation behavior sent by the target object client, further comprising: The target object server sends the first page image to the risk identification server in a case where it is determined that there is no previous operation behavior of the current operation behavior; The risk identification server performs risk identification on data of at least two modalities included in the first page image according to a risk identification model to obtain a first image risk identification result of the first page image, determines a behavior risk identification result of the current operation behavior according to the first image risk identification result.

3. The operation behavior identification method of claim 1, wherein determining the behavior risk identification result of the current operation behavior according to the image risk identification result comprises: According to the image risk identification result, the image risk identification result and the reference behavior rule are matched to obtain a matching result, and the behavior risk identification result of the current operation behavior is determined according to the matching result; Wherein, the reference behavior rule is a pre-defined risk behavior rule.

4. The operation behavior identification method of claim 1, after determining the behavior risk identification result of the current operation behavior, further comprising: The risk identification server determines a control instruction for the current operation behavior according to the behavior risk identification result of the current operation behavior, and sends the control instruction to the target object server; The target object server processes the current operation behavior according to the control instruction.

5. The operation behavior identification method of claim 1, wherein determining the difference page image between the first page image and the second page image according to the first page image and the second page image comprises: Calculate the image similarity between the first page image and the second page image; According to the image similarity, the difference page image between the first page image and the second page image is determined.

6. The operation behavior identification method of claim 5, wherein calculating the image similarity between the first page image and the second page image comprises: determine a plurality of first pixel points in the first page image, and determine a plurality of second pixel points in the second page image, which are at the same positions as the plurality of first pixel points; calculate pixel similarity between the first pixel points and the second pixel points; the determining the difference page image between the first page image and the second page image according to the image similarity comprises: determining a second pixel point with a pixel similarity less than a preset similarity threshold as a target pixel point, and determining the difference page image according to the target pixel point.

7. The operation behavior recognition method according to claim 1, before the sending the difference page image to the risk recognition server, further comprising: determining data of at least two modalities included in the difference page image; the sending the difference page image to the risk recognition server comprises: packaging the difference page image and the data of the at least two modalities into a data frame and sending to the risk recognition server.

8. The operation behavior recognition method according to claim 1, the determining the first page image corresponding to the current operation behavior in response to the current operation behavior sent by the target object client, comprising: determining page data corresponding to the current operation behavior in response to the current operation behavior sent by the target object client; rendering the page data to obtain the first page image corresponding to the current operation behavior.

9. The operation behavior recognition method according to claim 8, after the obtaining the first page image corresponding to the current operation behavior, further comprising: sending the first page image to the target object client for display on a display interface of the target object client.

10. The operation behavior recognition method according to claim 1, after the determining the first page image corresponding to the current operation behavior, further comprising: storing the first page image to a database, so that in response to a next operation behavior of the current operation behavior, the first page image is obtained from the database to perform behavior risk identification on the next operation behavior.

11. The method of claim 1-10, wherein, the risk recognition model is a multi-modal machine learning model.

12. The method of claim 1-10, wherein, the target object server is a server of a cloud computer, and the target object client is a client of the cloud computer.

13. An operation behavior recognition method applied to a target object server, wherein, the method comprises: determining a first page image corresponding to a current operation behavior in response to the current operation behavior sent by a target object client, and determining a second page image corresponding to a previous operation behavior of the current operation behavior; determining a difference page image between the first page image and the second page image according to the first page image and the second page image, and sending the difference page image to a risk recognition server; receive a behavior risk identification result of the current operation behavior sent by the risk identification server, wherein the behavior risk identification result is determined according to an image risk identification result, and the image risk identification result is obtained by the risk identification server according to a risk identification model by performing risk identification on data of at least two modalities included in the difference page image.

14. An operation behavior recognition method applied to a risk identification server, wherein, The method comprises: receiving a difference page image sent by a target object server in response to a current operation behavior sent by a target object client, wherein the difference page image is determined by the target object server according to a first page image corresponding to the current operation behavior and a second page image corresponding to a previous operation behavior of the current operation behavior; performing risk identification on data of at least two modalities included in the difference page image according to a risk identification model to obtain an image risk identification result of the difference page image; determining a behavior risk identification result of the current operation behavior according to the image risk identification result.

15. An operation behavior identification system, comprising a target object server and a risk identification server, wherein the target object server is configured to determine a first page image corresponding to a current operation behavior and determine a second page image corresponding to a previous operation behavior of the current operation behavior in response to the current operation behavior sent by a target object client, determine a difference page image between the first page image and the second page image according to the first page image and the second page image, and send the difference page image to the risk identification server; the risk identification server is configured to perform risk identification on data of at least two modalities included in the difference page image according to a risk identification model to obtain an image risk identification result of the difference page image, determine a behavior risk identification result of the current operation behavior according to the image risk identification result.

16. A computing device, comprising: a memory and a processor; the memory is used to store computer programs / instructions, and the processor is used to execute the computer programs / instructions, which realize the steps of the method in any one of claims 1 to 14 when executed by the processor.

17. A computer readable storage medium storing computer programs / instructions, which realize the steps of the method in any one of claims 1 to 14 when executed by the processor.

18. A computer program product comprising computer programs / instructions, which realize the steps of the method in any one of claims 1 to 14 when executed by the processor.

Citation Information

Patent Citations

  • User behavior abnormity detection method and device, electronic equipment and readable storage medium

    CN109727058A

  • Risk operation behavior identification method and device, equipment and storage medium

    CN113849810A

  • Internet behavior data monitoring method and device, equipment and storage medium

    CN115914032A

  • Mouse operation event identification method and device, electronic equipment and readable medium

    CN116434121A

  • Method and device for detecting user interface

    CN117413256A