A data processing method, apparatus and device

By constructing target graph structured data and using large language models to process user operation behavior log data, the problems of risk detection efficiency and accuracy under the complexity of user data and large data volume are solved, and more efficient risk user identification is achieved.

CN119848834BActive Publication Date: 2025-12-19ZHEJIANG UNIV +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411997927.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-12-31
Publication Date
2025-12-19
Estimated Expiration
2044-12-31

AI Technical Summary

Technical Problem

In existing technologies, user data structures are complex and the data volume is large, resulting in low efficiency and accuracy of risk detection through manual analysis.

Method used

By receiving risk detection requests, obtaining the target user's operation behavior log data, constructing the target graph structure data, performing feature extraction and key information extraction, and using a large language model to determine whether the user is a risk user.

Benefits of technology

It improves the efficiency and accuracy of risk detection, enabling more effective identification of users' risky behaviors.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119848834B_ABST
    Figure CN119848834B_ABST
Patent Text Reader

Abstract

Embodiments of the present specification provide a data processing method, device and equipment, wherein the method comprises: receiving a risk detection request for a target user; in response to the risk detection request, obtaining log data containing operation behaviors of the target user; based on the operation behaviors of the target user contained in the log data, determining target graph structure data, the target graph structure data containing target nodes determined according to the operation behaviors, and edges between nodes determined according to logical relationships between the operation behaviors; performing feature extraction processing on the target graph structure data to obtain a target graph embedding vector corresponding to the target graph structure data, and performing key information extraction processing on the log data to obtain target information; and determining whether the target user is a risk user based on the target graph embedding vector and the target information according to a large language model.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present document relates to the technical field of computer technology, and particularly relates to a data processing method, device and equipment. BACKGROUND

[0002] With people paying more and more attention to their private data, in order to protect user privacy and ensure data security, it is necessary to detect whether a user is a risk user to guarantee data security in the process of business execution. For example, a security personnel can manually analyze user data to detect whether a user is a risk user.

[0003] However, as the data structure of user data becomes more and more complex and the data volume becomes larger and larger, the detection efficiency and detection accuracy of risk detection by manual analysis are low. Therefore, the present specification embodiment provides a more optimal technical solution for risk detection of a user. SUMMARY

[0004] The purpose of the present specification embodiment is to provide a more optimal technical solution for risk detection of a user.

[0005] In order to achieve the above technical solution, the present specification embodiment is implemented as follows:

[0006] The data processing method provided by the present specification embodiment comprises the following steps: receiving a risk detection request for a target user; in response to the risk detection request, acquiring log data containing operation behaviors of the target user; based on the operation behaviors of the target user contained in the log data, determining target graph structure data, the target graph structure data containing target nodes determined according to the operation behaviors and edges between nodes determined according to logical relationships between the operation behaviors; performing feature extraction processing on the target graph structure data to obtain a target graph embedding vector corresponding to the target graph structure data, and performing key information extraction processing on the log data to obtain target information; based on the target graph embedding vector and the target information, determining whether the target user is a risk user according to a large language model.

[0007] An embodiment of the present specification provides a data processing apparatus, the apparatus comprising: a request receiving module configured to receive a risk detection request for a target user; a first obtaining module configured to, in response to the risk detection request, obtain log data containing operation behaviors of the target user; a first determining module configured to determine target graph structure data based on the operation behaviors of the target user contained in the log data, the target graph structure data containing target nodes determined according to the operation behaviors and edges between nodes determined according to logical relationships between the operation behaviors; a first extracting module configured to perform feature extraction processing on the target graph structure data to obtain a target graph embedding vector corresponding to the target graph structure data, and perform key information extraction processing on the log data to obtain target information; and a risk detection module configured to determine whether the target user is a risk user based on the target graph embedding vector and the target information according to a large language model.

[0008] An embodiment of the present specification provides a data processing device, the data processing device comprising: a processor; and a memory arranged to store computer executable instructions that, when executed, cause the processor to: receive a risk detection request for a target user; in response to the risk detection request, obtain log data containing operation behaviors of the target user; determine target graph structure data based on the operation behaviors of the target user contained in the log data, the target graph structure data containing target nodes determined according to the operation behaviors and edges between nodes determined according to logical relationships between the operation behaviors; perform feature extraction processing on the target graph structure data to obtain a target graph embedding vector corresponding to the target graph structure data, and perform key information extraction processing on the log data to obtain target information; and determine whether the target user is a risk user based on the target graph embedding vector and the target information according to a large language model.

[0009] An embodiment of the present specification also provides a storage medium for storing computer executable instructions, the executable instructions, when executed by a processor, implement the following processes: receiving a risk detection request for a target user; in response to the risk detection request, obtaining log data containing operation behaviors of the target user; determining target graph structure data based on the operation behaviors of the target user contained in the log data, the target graph structure data containing target nodes determined according to the operation behaviors and edges between nodes determined according to logical relationships between the operation behaviors; performing feature extraction processing on the target graph structure data to obtain a target graph embedding vector corresponding to the target graph structure data, and performing key information extraction processing on the log data to obtain target information; and determining whether the target user is a risk user based on the target graph embedding vector and the target information according to a large language model.

[0010] The embodiment of the present specification further provides a computer program product comprising a computer program which, when executed by a processor, implements the following process: receiving a risk detection request for a target user; in response to the risk detection request, obtaining log data containing operation behaviors of the target user; determining target graph structure data based on the operation behaviors of the target user contained in the log data, the target graph structure data containing target nodes determined according to the operation behaviors, and edges between nodes determined according to logical relationships between the operation behaviors; performing feature extraction processing on the target graph structure data to obtain a target graph embedding vector corresponding to the target graph structure data, and performing key information extraction processing on the log data to obtain target information; determining whether the target user is a risk user based on the target graph embedding vector and the target information according to a large language model. BRIEF DESCRIPTION OF DRAWINGS

[0011] In order to more clearly illustrate the technical solutions in the embodiments of the present specification or the prior art, the following will briefly introduce the drawings needed to be used in the embodiments or prior art description. Obviously, the drawings in the following description are only some embodiments described in the present specification, and for those skilled in the art, other drawings can also be obtained from these drawings without creative labor;

[0012] Figure 1 for an embodiment of a data processing method of the present specification;

[0013] Figure 2 for a schematic diagram of a data processing process of the present specification;

[0014] Figure 3 for another embodiment of a data processing method of the present specification;

[0015] Figure 4 for another embodiment of a data processing method of the present specification;

[0016] Figure 5 for a schematic diagram of a data processing process of the present specification;

[0017] Figure 6 for an embodiment of a data processing device of the present specification;

[0018] Figure 7 for an embodiment of a data processing device of the present specification. DETAILED DESCRIPTION

[0019] The embodiment of the present specification provides a data processing method, device and equipment.

[0020] In order for those skilled in the art to better understand the technical solutions in the specification, the technical solutions in the specification will be clearly and completely described below in conjunction with the drawings in the specification. Obviously, the described embodiments are only part of the embodiments of the specification, not all. Based on the embodiments in the specification, all other embodiments obtained by those of ordinary skill in the art without creative labor should be within the scope of protection of the specification.

[0021] The embodiments of the specification provide a detection mechanism for whether a user is a risk user. As people pay more and more attention to their own privacy data, in order to protect user privacy and ensure data security, it is necessary to detect whether a user is a risk user to ensure data security during business execution. For example, a security personnel can manually analyze user data to detect whether a user is a risk user. However, as the data structure of user data becomes more and more complex and the data volume becomes larger and larger, the detection efficiency and detection accuracy of risk detection by manual analysis are low. Therefore, the embodiments of the specification provide a more optimal technical solution for risk detection of a user. In this solution, a risk detection request for a target user is received, log data containing operation behaviors of the target user is obtained in response to the risk detection request, target graph structure data is determined based on the operation behaviors of the target user contained in the log data, wherein the target graph structure data can contain target nodes determined according to the operation behaviors, and edges between nodes determined according to logical relationships between the operation behaviors, feature extraction processing is performed on the target graph structure data to obtain target graph embedding vectors corresponding to the target graph structure data, key information extraction processing is performed on the log data to obtain target information, and whether the target user is a risk user is determined based on the target graph embedding vectors and the target information according to a large language model. In this way, on the one hand, as the log data contains more content, the operation behaviors of the target user can be extracted from the log data to construct target graph structure data with small size and clear structure, so as to improve the detection efficiency and detection accuracy of subsequent risk detection based on the target graph structure data. On the other hand, as the target information containing key information is natural language that can be understood by the large language model, and the target graph embedding vectors determined according to the target graph structure data contain logical relationships between the operation behaviors, the risk detection accuracy and detection efficiency of the user can be improved based on the target information and the target graph embedding vectors by the large language model. Specific processing can be referred to the specific content in the following embodiments.

[0022] For example, the log data can be obtained by monitoring the operation behaviors of the target user, and the operation behaviors of the target user can be obtained by analyzing the log data. Figure 1As shown, the embodiment of the present specification provides a data processing method, the execution subject of the method can be a server, wherein the server can be an independent server, or a server cluster composed of multiple servers, etc. The server can be a background server of a financial service or a network shopping service, etc., or a background server of an application program, etc. In the embodiment, the execution subject is taken as the server as an example for detailed description. The method can specifically include the following steps:

[0023] In step S102, a risk detection request for a target user is received.

[0024] The target user can be any user to be detected, for example, the target user can be a user triggering the execution of a certain service, or the target user can also be a user related to the execution of a certain service, for example, the target user can be a user triggering the execution of a resource transfer service, or the target user can also be a resource transfer object in the execution of a resource transfer service.

[0025] In implementation, the server can determine a user corresponding to a certain service when receiving a triggering execution instruction of the service, and determine the user as a target user, that is, the server can receive a risk detection request for the target user.

[0026] For example, the server can determine user 1 and / or user 2 as the target user when receiving a triggering execution instruction of a resource transfer service of user 2 by user 1.

[0027] Alternatively, the server can also determine user 1 and / or user 2 as the target user when receiving a triggering execution instruction of an information sharing service of user 2 by user 1.

[0028] Alternatively, the server can determine user 1 as the target user when receiving a triggering execution instruction of a certain database access service by user 1.

[0029] The above determination method of the target user is an optional and implementable determination method. In addition, there can be many determination methods, and different determination methods can be selected according to different actual application scenarios, which are not limited in the embodiment of the present specification.

[0030] In step S104, log data containing the operation behavior of the target user is obtained in response to the risk detection request.

[0031] In implementation, the server can obtain log data containing the operation behavior of the target user in a preset detection period in response to the risk detection request, wherein the preset detection period can be any detection period such as one day, three days, etc.

[0032] Alternatively, the server can also obtain log data related to the target user triggering the execution of a certain service, for example, taking the target user as an example, the server can obtain log data of the target user triggering the execution of the resource transfer service.

[0033] In step S106, the target graph structure data is determined based on the operation behavior of the target user contained in the log data.

[0034] The target graph structure data can include target nodes determined according to the operation behavior, and edges between nodes determined according to the logical relationship between the operation behaviors. The target graph structure data can be knowledge graph data, i.e., the target graph structure data can be graph structure data used to represent and organize knowledge, which can include entities, attributes and relationships, and can be used to describe the association between things. The target graph structure data can be NEO4J (high-performance image database) graph data.

[0035] In implementation, the server can preprocess the log data, and determine the target graph structure data based on the operation behavior of the target user contained in the preprocessed log data. The preprocessing can include regularization processing, redundant information removal processing, etc. In this way, the subsequent data processing efficiency can be improved by preprocessing the log data.

[0036] Specifically, the server can unify the format of the log data based on a preset format, and then remove redundant sentences and redundant behavior information in the log data. Finally, the server can rearrange and divide the log data according to different processes, thread identifiers contained in each process, and logical relationships of operation behaviors, etc., to construct the target graph structure data.

[0037] The logical relationship of the operation behavior can include the trigger order relationship, data calling relationship, and behavior association relationship between the operation behaviors.

[0038] In step S108, the target graph structure data is subjected to feature extraction processing to obtain a target graph embedding vector corresponding to the target graph structure data, and the log data is subjected to key information extraction processing to obtain target information.

[0039] In implementation, the server can perform feature extraction processing on the target graph structure data through a pre-trained feature extraction model to obtain a target graph embedding vector corresponding to the target graph structure data. The feature extraction model can be a model constructed based on a preset deep learning algorithm.

[0040] For example, the server can build a feature extraction model through an encoder and a decoder, train the feature extraction model through historical graph structure data and corresponding historical graph embedding vectors, and obtain a trained feature extraction model. Then, the server can perform feature extraction processing on the target graph structure data according to the encoder in the trained feature extraction model, and obtain a target graph embedding vector corresponding to the target graph structure data.

[0041] In addition, the above is to obtain a target graph embedding vector by performing feature extraction processing on target graph structure data according to a feature extraction model constructed by a deep learning algorithm. In actual application scenarios, there can be many methods for determining a target graph embedding vector, and different methods can be selected according to different actual application scenarios. The embodiments of the present specification do not make specific limitations in this regard.

[0042] The server can perform key information extraction processing on the log data through a pre-trained information extraction model to obtain target information, where the information extraction model can be a model constructed based on a preset machine learning algorithm.

[0043] Alternatively, the server can also perform key information extraction processing on the log data through a large language model to obtain target information. In addition, there can be many different key information extraction processing methods, and different key information extraction methods can be selected according to different actual application scenarios. The embodiments of the present specification do not make specific limitations in this regard.

[0044] In step S110, it is determined whether the target user is a risk user based on the target graph embedding vector and the target information according to the large language model.

[0045] The large language model (LLM) can be an advanced natural language processing (NLP) model constructed based on deep learning technology. The large language model has a large number of parameters and a wide range of training data sets, has strong capabilities in language understanding and generation tasks, can be effectively applied to text generation, machine translation, content summarization, question and answer systems, and various complex dialogue scenarios, and can capture the richness and subtle differences of language expression and display the captured information through the powerful computing power of tens of billions to hundreds of billions of parameters.

[0046] In implementation, the server can integrate the target embedding vector and the target information, and then determine whether the target user is a risk user based on preset prompt information and the information obtained through integration according to the large language model.

[0047] In fact, the preset prompt information can be determined according to the detection requirement corresponding to the target user. For example, taking a target user triggered to perform a user of a database access service as an example, the target user's access behavior can be detected for risk, that is, the corresponding preset prompt information can be "please describe whether the target user's access behavior is an attack behavior, if it is an attack behavior, please describe the attack source, attack range and attack description".

[0048] The embodiment of the present specification provides a data processing method, by receiving a risk detection request for a target user, in response to the risk detection request, obtaining log data containing operation behavior of the target user, determining target graph structure data based on the operation behavior of the target user contained in the log data, wherein the target graph structure data can contain target nodes determined according to the operation behavior, and edges between nodes determined according to the logical relationship between the operation behaviors, performing feature extraction processing on the target graph structure data to obtain target graph embedding vectors corresponding to the target graph structure data, and performing key information extraction processing on the log data to obtain target information, determining whether the target user is a risk user based on the target graph embedding vectors and the target information according to a large language model. In this way, on the one hand, since the log data contains more content, the operation behavior of the target user can be extracted from the log data to construct target graph structure data with small size and clear structure, so as to improve the detection efficiency and detection accuracy of subsequent risk detection based on the target graph structure data. On the other hand, since the target information containing key information is natural language that can be understood by the large language model, and the target graph embedding vectors determined according to the target graph structure data contain the logical relationship between the operation behaviors, the risk detection accuracy and detection efficiency of the user can be improved based on the target information and the target graph embedding vectors through the large language model.

[0049] In actual application, the specific processing mode of determining the target graph structure data based on the operation behavior of the target user contained in the log data in step S106 can be various, and the following provides an optional processing mode, as shown in the following steps S1062~S1064. Figure 2 As shown, the specific processing can include the following steps S1062~S1064.

[0050] In step S1062, the operation behavior of the target user contained in the log data is clustered to obtain a first clustering result, and the events contained in the log data are determined according to the first clustering result.

[0051] In implementation, the server can determine each class contained in the first clustering result as an event.

[0052] In actual applications, the specific processing manner of clustering the operation behaviors of the target user included in the log data in step S1062 to obtain the first clustering result can be various. The following provides an optional processing manner, which can specifically include the following steps A1-A3.

[0053] In step A1, the first graph structure data is determined based on the thread information corresponding to the operation behaviors of the target user included in the log data.

[0054] The first graph structure data can include first nodes determined according to the thread information and edges between the nodes determined according to the logical relationship between threads.

[0055] In step A2, a second node in the first nodes is determined according to the out-degree information and / or the in-degree information of each first node in the first graph structure data.

[0056] In implementation, the server can determine the node association degree between the first node and other first nodes according to the out-degree information and / or the in-degree information of the first node, and then the server can filter out the nodes in the first node with a higher association degree with other first nodes as the second node according to a preset association degree threshold.

[0057] In step A3, the second nodes are clustered to obtain a second clustering result, and the first clustering result is determined according to the second clustering result.

[0058] In implementation, the server can cluster the second nodes according to a preset clustering algorithm (such as k-means algorithm) to obtain the second clustering result.

[0059] The server can determine the second clustering result as the first clustering result, or the server can further perform filtering processing on the second clustering result, and determine the first clustering result according to the filtered second clustering result.

[0060] For example, the server can perform filtering processing on the second clustering result according to the number of second nodes included in each class in the second clustering result, specifically, the server can filter out the class in the second clustering result that includes a number of second nodes greater than a preset number threshold, to determine the first clustering result.

[0061] In step S1064, the target nodes in the target graph structure data are constructed according to the events, and the edges between the nodes in the target graph structure data are determined according to the logical relationship between the events.

[0062] In actual application, the specific processing manner of the target graph structure data in step S108 can be various, and an optional processing manner is provided as follows, as shown in Figure 2 The specific processing manner can include the following step S1082.

[0063] In step S1082, the target graph structure data is subjected to feature extraction processing according to a pre-trained feature extraction model to obtain a target graph embedding vector corresponding to the target graph structure data, and the log data is subjected to key information extraction processing to obtain target information.

[0064] The feature extraction model can be a model constructed based on a heterogeneous graph attention network (HAN) algorithm. The HAN algorithm can fully capture rich semantic information in a heterogeneous graph by introducing node-level attention and semantic-level attention.

[0065] In implementation, the server can construct the feature extraction model based on the HAN algorithm and a graph neural network (GNN) algorithm. In this way, the feature extraction model constructed based on the HAN algorithm can be used for feature extraction processing, and the GNN network can fully understand the relationship between nodes in the case that the nodes have multiple types.

[0066] In actual application, before determining the risk type corresponding to the operation behavior based on the target graph embedding vector and the target information according to the large language model, the large language model can be subjected to fine-tuning processing, and the specific processing manner of the fine-tuning processing can be various, and an optional processing manner is provided as follows, as shown in Figure 3 The specific processing manner can include the following steps S302-S312.

[0067] In step S302, a pre-trained large language model is obtained.

[0068] In implementation, the server can pre-train the large language model based on sample data with a large amount of data to obtain the pre-trained large language model.

[0069] In step S304, historical log data and a first risk type corresponding to the operation behavior of a user included in the historical log data are obtained.

[0070] In step S306, second graph structure data is determined based on the operation behavior of the user included in the historical log data.

[0071] The second graph structure data can include nodes determined according to operation behaviors in the historical log data, and edges between nodes determined according to logical relationships between the operation behaviors in the historical log data.

[0072] In step S308, feature extraction processing is performed on the second graph structure data to obtain a first graph embedding vector corresponding to the second graph structure data, and key information extraction processing is performed on the historical log data to obtain first key information.

[0073] In implementation, the specific processing process of step S308 can refer to the processing process of step S108, which will not be described here.

[0074] In step S310, according to the pre-trained large language model, based on the preset prompt information, the first graph embedding vector and the first key information, a second risk type corresponding to the operation behavior of the user contained in the historical log data is determined.

[0075] In actual application, the specific processing mode of step S310 for determining the second risk type corresponding to the operation behavior of the user contained in the historical log data according to the pre-trained large language model based on the preset prompt information, the first graph embedding vector and the first key information can be various, and the following provides an optional processing mode, which can include the following steps B1-B2.

[0076] In step B1, according to the detection requirement corresponding to the target user, a plurality of sub-prompt information having a logical reasoning relationship corresponding to the preset prompt information is generated.

[0077] In implementation, the server can use the thought chain (Chain of Thought, CoT) idea to decompose the preset prompt information into a plurality of sub-prompt information, and gradually guide the large language model to solve these sub-prompt information, thereby improving the reasoning ability and adaptability of the large language model, realizing that the large language model can detect risks while answering questions such as attack source points and attack behavior characteristics.

[0078] In step B2, according to the pre-trained large language model, based on the plurality of sub-prompt information, the first graph embedding vector and the first key information, a second risk type corresponding to the operation behavior of the user contained in the historical log data is determined.

[0079] In step S312, according to the first risk type and the second risk type, the pre-trained large language model is fine-tuned to obtain a trained large language model.

[0080] In practical applications, the key information extraction and processing of log data in step S108 can be performed in various ways to obtain the target information. One optional processing method is provided below, such as... Figure 4 As shown, the specific process may include the following steps, S1084.

[0081] In step S1084, feature extraction processing is performed on the target graph structure data to obtain the target graph embedding vector corresponding to the target graph structure data. Then, according to the preset summary generation algorithm, key information extraction processing is performed on the log data to obtain the summary information of the log data, and the summary information of the log data is determined as the target information.

[0082] In practical applications, log data can include audit log data and / or operation log data. Audit log data can be log file data that records information such as system operating status, user operation behavior, and security events. It can include data such as log time, source, level, and specific operation behavior, and is a kind of system behavior description data.

[0083] In implementation, such as Figure 5 As shown, the server can obtain audit log data, convert it into a unified format, remove redundant sentences and behaviors, and finally rearrange and divide the audit log data according to different processes, thread identifiers, and the logical order of operation behaviors to construct NEO4J graph data (i.e., target graph structure data).

[0084] It solves the problem of messy log data content in information extraction, making it difficult to directly extract information and classify events. Moreover, it can build a small-volume, clearly structured graph data through auditing log data.

[0085] In addition, the server can train a feature extraction model built by GNN and HAN using the number of households in the second graph structure constructed from historical log data. Then, the trained feature extraction model can be used to perform feature extraction on the target graph structure data to obtain the target graph embedding vector.

[0086] Simultaneously, the server can generate summaries using natural language processing techniques. This involves extracting key information from log data using a pre-defined summary generation algorithm to obtain a summary of the log data, which can then be used to identify target information. The server can then integrate the target information with graph embedding to generate integrated detection data, which can be used for attack detection.

[0087] Wherein, when fine-tuning the large language model, the server can utilize the Chain-Of-Thought principle to fine-tune the large language model to efficiently infer and detect the log summary and the graph embedding, realize attack behavior detection with high accuracy and low false alarm rate, and solve the problem of poor universality of advanced persistent threat (APT) attack detection.

[0088] In addition, the server can also flexibly design a large model fine-tuning scheme according to actual needs, such as designing prompt information (attack source point, attack range, attack story description, etc.) according to needs to realize the question and answer function for the needs.

[0089] In actual application, the specific processing manner of determining whether the target user is a risk user according to the large language model based on the target graph embedding vector and the target information in the step S110 can be various, and an optional processing manner is provided as follows, such as Figure 4 As shown, the specific processing can include the following step S1102.

[0090] In the step S1102, the operation behavior corresponding risk type is determined according to the large language model based on the target graph embedding vector and the target information, and whether the target user is a risk user is determined according to the operation behavior corresponding risk type.

[0091] Wherein, the risk type can include high risk, medium risk, low risk, no risk and the like.

[0092] In implementation, the server can determine whether the target user is a risk user according to the behavior type and the risk type of each operation behavior.

[0093] For example, the server can determine the risk score corresponding to each operation behavior according to the behavior type and the risk type of each operation behavior, and then determine the risk score corresponding to the target user according to the risk score corresponding to each operation behavior, so as to determine whether the target user is a risk user according to the risk score corresponding to the target user.

[0094] In addition, in the case of determining that the target user is a risk user, the preset alarm information and the target operation behavior with a risk score greater than a preset score threshold can be sent to the preset processing party.

[0095] The embodiment of the present specification provides a data processing method, by receiving a risk detection request for a target user, in response to the risk detection request, obtaining log data containing the operation behavior of the target user, determining target graph structure data based on the operation behavior of the target user contained in the log data, wherein the target graph structure data can contain target nodes determined according to the operation behavior, and edges between nodes determined according to the logical relationship between the operation behaviors, performing feature extraction processing on the target graph structure data to obtain target graph embedding vectors corresponding to the target graph structure data, and performing key information extraction processing on the log data to obtain target information, determining whether the target user is a risk user based on the target graph embedding vectors and the target information according to a large language model. In this way, on the one hand, since the content contained in the log data is more, the operation behavior of the target user can be extracted from the log data, and the target graph structure data with small size and clear structure can be constructed, so as to improve the detection efficiency and detection accuracy of subsequent risk detection based on the target graph structure data. On the other hand, since the target information containing the key information is natural language that can be understood by the large language model, and the target graph embedding vectors determined according to the target graph structure data contain the logical relationship between the operation behaviors, the risk detection accuracy and detection efficiency of the user can be improved based on the target information and the target graph embedding vectors through the large language model.

[0096] The above is the data processing method provided by the embodiment of the present specification, based on the same idea, the embodiment of the present specification also provides a data processing device, as shown in Figure 6 .

[0097] The data processing device comprises a request receiving module 601, a first obtaining module 602, a first determining module 603, a first extracting module 604 and a risk detection module 605, wherein:

[0098] The request receiving module 601 is configured to receive a risk detection request for a target user;

[0099] The first obtaining module 602 is configured to obtain log data containing the operation behavior of the target user in response to the risk detection request;

[0100] The first determining module 603 is configured to determine target graph structure data based on the operation behavior of the target user contained in the log data, wherein the target graph structure data contains target nodes determined according to the operation behavior, and edges between nodes determined according to the logical relationship between the operation behaviors;

[0101] The first extracting module 604 is configured to perform feature extraction processing on the target graph structure data to obtain target graph embedding vectors corresponding to the target graph structure data, and perform key information extraction processing on the log data to obtain target information;

[0102] The risk detection module 605 is configured to determine whether the target user is a risk user based on the target graph embedding vector and the target information according to the large language model.

[0103] In an embodiment of the present specification, the first determination module 603 is configured to:

[0104] perform clustering processing on the operation behavior of the target user contained in the log data to obtain a first clustering result, and determine an event contained in the log data according to the first clustering result.

[0105] construct a target node in the target graph structure data according to the event, and determine an edge between nodes in the target graph structure data according to a logical relationship between the events.

[0106] In an embodiment of the present specification, the first determination module 603 is configured to:

[0107] determine first graph structure data based on thread information corresponding to the operation behavior of the target user contained in the log data, the first graph structure data including first nodes determined according to the thread information and edges between nodes determined according to the thread information;

[0108] determine a second node in the first nodes according to out-degree information and / or in-degree information of each of the first nodes in the first graph structure data.

[0109] perform clustering processing on the second node to obtain a second clustering result, and determine the first clustering result according to the second clustering result.

[0110] In an embodiment of the present specification, the first extraction module 604 is configured to:

[0111] perform feature extraction processing on the target graph structure data according to a pre-trained feature extraction model to obtain a target graph embedding vector corresponding to the target graph structure data, wherein the feature extraction model is a model constructed based on a heterogeneous graph attention network algorithm.

[0112] In an embodiment of the present specification, the apparatus further includes:

[0113] a model acquisition module configured to acquire a pre-trained large language model;

[0114] a second acquisition module configured to acquire historical log data and a first risk type corresponding to operation behavior of a user contained in the historical log data, the data quantity of the historical log data being less than the data quantity of the sample data;

[0115] The second determining module is configured to determine second graph structure data based on the operation behavior of the user contained in the historical log data.

[0116] The second extracting module is configured to perform feature extraction processing on the second graph structure data to obtain a first graph embedding vector corresponding to the second graph structure data, and perform key information extraction processing on the historical log data to obtain first key information.

[0117] The risk determining module is configured to determine, according to the pre-trained large language model, a second risk type corresponding to the operation behavior of the user contained in the historical log data based on preset prompt information, the first graph embedding vector and the first key information.

[0118] The model training module is configured to perform fine-tuning processing on the pre-trained large language model according to the first risk type and the second risk type to obtain a trained large language model.

[0119] In the embodiments of the present specification, the type determining module is configured to:

[0120] generate a plurality of sub-prompt information having a logical reasoning relationship corresponding to the preset prompt information according to the detection requirement corresponding to the target user;

[0121] determine, according to the pre-trained large language model, a second risk type corresponding to the operation behavior of the user contained in the historical log data based on the plurality of sub-prompt information, the first graph embedding vector and the first key information.

[0122] In the embodiments of the present specification, the first extracting module 604 is configured to:

[0123] perform key information extraction processing on the log data according to a preset abstract generation algorithm to obtain abstract information of the log data, and determine the abstract information of the log data as the target information.

[0124] In the embodiments of the present specification, the log data includes audit log data and / or operation log data, and the risk detection module 605 is configured to:

[0125] determine, according to the large language model, a risk type corresponding to the operation behavior based on the target graph embedding vector and the target information, and determine whether the target user is a risk user according to the risk type corresponding to the operation behavior.

[0126] The embodiment of the present specification provides a data processing apparatus, by receiving a risk detection request for a target user, in response to the risk detection request, obtaining log data containing the operation behavior of the target user, determining target graph structure data based on the operation behavior of the target user contained in the log data, wherein the target graph structure data can contain target nodes determined according to the operation behavior, and edges between nodes determined according to the logical relationship between the operation behaviors, performing feature extraction processing on the target graph structure data to obtain target graph embedding vectors corresponding to the target graph structure data, and performing key information extraction processing on the log data to obtain target information, determining whether the target user is a risk user based on the target graph embedding vectors and the target information according to a large language model. In this way, on the one hand, since the content contained in the log data is more, the operation behavior of the target user can be extracted from the log data, and the target graph structure data with small size and clear structure is constructed, so as to improve the detection efficiency and detection accuracy of subsequent risk detection based on the target graph structure data. On the other hand, since the target information containing the key information is natural language that can be understood by the large language model, and the target graph embedding vectors determined according to the target graph structure data contain the logical relationship between the operation behaviors, the risk detection accuracy and detection efficiency of the user can be improved based on the target information and the target graph embedding vectors through the large language model.

[0127] The data processing apparatus provided by the embodiment of the present specification is based on the same idea, and the data processing device provided by the embodiment of the present specification is shown in the following. Figure 7

[0128] The data processing device can be a terminal device or a server provided by the above embodiment.

[0129] The data processing device can have a large difference due to different configurations or performances, and can include one or more processors 701 and memories 702, and the memories 702 can store one or more storage application programs or data. The memory 702 can be temporary storage or persistent storage. The application programs stored in the memory 702 can include one or more modules (not shown in the figure), and each module can include a series of computer executable instructions in the data processing device. Further, the processor 701 can be configured to communicate with the memory 702 and execute a series of computer executable instructions in the memory 702 on the data processing device. The data processing device can also include one or more power supplies 703, one or more wired or wireless network interfaces 704, one or more input output interfaces 705, and one or more keyboards 706.

[0130] ​In particular embodiments, a data processing apparatus can include a memory and one or more programs, wherein one or more programs are stored in the memory and configured to be implemented by one or more processors of the data processing apparatus, the one or more programs including instructions for:

[0131] receiving a risk detection request for a target user;

[0132] in response to the risk detection request, obtaining log data containing operation behaviors of the target user;

[0133] determining target graph structure data based on the operation behaviors of the target user contained in the log data, the target graph structure data containing target nodes determined according to the operation behaviors, and edges between nodes determined according to logical relationships between the operation behaviors;

[0134] performing feature extraction processing on the target graph structure data to obtain a target graph embedding vector corresponding to the target graph structure data, and performing key information extraction processing on the log data to obtain target information;

[0135] determining whether the target user is a risk user based on the target graph embedding vector and the target information according to a large language model.

[0136] Each of the embodiments in the specification is described in a progressive manner, and the same or similar parts between the embodiments can be referred to each other. Each embodiment focuses on the difference from other embodiments. In particular, the data processing apparatus embodiment is described simply because it is basically similar to the method embodiment, and the relevant parts can be referred to the part of the method embodiment.

[0137] The embodiment of the present specification provides a data processing device, by receiving a risk detection request for a target user, in response to the risk detection request, obtaining log data containing the operation behavior of the target user, determining target graph structure data based on the operation behavior of the target user contained in the log data, wherein the target graph structure data can contain target nodes determined according to the operation behavior, and edges between nodes determined according to the logical relationship between the operation behaviors, performing feature extraction processing on the target graph structure data to obtain target graph embedding vectors corresponding to the target graph structure data, and performing key information extraction processing on the log data to obtain target information, determining whether the target user is a risk user based on the target graph embedding vectors and the target information according to a large language model. In this way, on the one hand, since the content contained in the log data is more, the operation behavior of the target user can be extracted from the log data, and the target graph structure data with small size and clear structure can be constructed, so as to improve the detection efficiency and detection accuracy of subsequent risk detection based on the target graph structure data. On the other hand, since the target information containing the key information is natural language that can be understood by the large language model, and the target graph embedding vectors determined according to the target graph structure data contain the logical relationship between the operation behaviors, the risk detection accuracy and detection efficiency of the user can be improved based on the target information and the target graph embedding vectors through the large language model.

[0138] Further, based on the above Figures 1 to 5 The one or more embodiments of the present specification also provide a storage medium for storing computer executable instruction information, in a specific embodiment, the storage medium can be a U disk, an optical disk, a hard disk, etc. The computer executable instruction information stored in the storage medium can realize the following processes when executed by a processor:

[0139] receiving a risk detection request for a target user;

[0140] in response to the risk detection request, obtaining log data containing the operation behavior of the target user;

[0141] determining target graph structure data based on the operation behavior of the target user contained in the log data, wherein the target graph structure data contains target nodes determined according to the operation behavior, and edges between nodes determined according to the logical relationship between the operation behaviors;

[0142] performing feature extraction processing on the target graph structure data to obtain target graph embedding vectors corresponding to the target graph structure data, and performing key information extraction processing on the log data to obtain target information;

[0143] determining whether the target user is a risk user based on the target graph embedding vectors and the target information according to a large language model.

[0144] Each of the embodiments in the specification is described in a progressive manner, and the same or similar parts between the embodiments can be referred to each other. Each of the embodiments focuses on the difference from other embodiments. In particular, for the above-mentioned storage medium embodiment, since it is basically similar to the method embodiment, the description is relatively simple, and the relevant parts can be referred to the part of the method embodiment.

[0145] The embodiment of the specification provides a storage medium, by receiving a risk detection request for a target user, in response to the risk detection request, obtaining log data containing the operation behavior of the target user, determining target graph structure data based on the operation behavior of the target user contained in the log data, wherein the target graph structure data can contain target nodes determined according to the operation behavior, and edges between nodes determined according to the logical relationship between the operation behaviors, performing feature extraction processing on the target graph structure data to obtain target graph embedding vectors corresponding to the target graph structure data, and performing key information extraction processing on the log data to obtain target information, determining whether the target user is a risk user based on the target graph embedding vectors and the target information according to a large language model. In this way, on the one hand, since the content contained in the log data is more, the operation behavior of the target user can be extracted from the log data, and the target graph structure data with small size and clear structure can be constructed, so as to improve the detection efficiency and detection accuracy of subsequent risk detection based on the target graph structure data. On the other hand, since the target information containing the key information is natural language that can be understood by the large language model, and the target graph embedding vectors determined according to the target graph structure data contain the logical relationship between the operation behaviors, the risk detection accuracy and detection efficiency of the user can be improved based on the target information and the target graph embedding vectors through the large language model.

[0146] Further, based on the above-mentioned Figures 1 to 5 The one or more embodiments of the specification also provide a computer program product including a computer program, the computer program in the computer program product can implement the following flow when executed by a processor.

[0147] Receiving a risk detection request for a target user;

[0148] In response to the risk detection request, obtaining log data containing the operation behavior of the target user;

[0149] Determining target graph structure data based on the operation behavior of the target user contained in the log data, wherein the target graph structure data contains target nodes determined according to the operation behavior, and edges between nodes determined according to the logical relationship between the operation behaviors;

[0150] perform feature extraction processing on the target graph structure data to obtain a target graph embedding vector corresponding to the target graph structure data, and perform key information extraction processing on the log data to obtain target information;

[0151] According to the large language model, it is determined whether the target user is a risk user based on the target graph embedding vector and the target information.

[0152] Each of the embodiments in the specification is described in a progressive manner, and the same or similar parts between the embodiments can be referred to each other. Each embodiment focuses on the difference from other embodiments. In particular, for the above-mentioned computer program product embodiment, since it is basically similar to the method embodiment, the description is relatively simple, and the relevant parts can be referred to the part of the method embodiment.

[0153] The embodiment of the specification provides a computer program product, which receives a risk detection request for a target user, acquires log data containing operation behaviors of the target user in response to the risk detection request, determines target graph structure data based on the operation behaviors of the target user contained in the log data, wherein the target graph structure data can contain target nodes determined according to the operation behaviors, and edges between nodes determined according to the logical relationship between the operation behaviors, performs feature extraction processing on the target graph structure data to obtain a target graph embedding vector corresponding to the target graph structure data, and performs key information extraction processing on the log data to obtain target information, according to the large language model, it is determined whether the target user is a risk user based on the target graph embedding vector and the target information. In this way, on the one hand, since the log data contains more content, the operation behaviors of the target user can be extracted from the log data, and the target graph structure data with small size and clear structure can be constructed, so as to improve the detection efficiency and detection accuracy of subsequent risk detection based on the target graph structure data. On the other hand, since the target information containing key information is natural language that can be understood by the large language model, and the target graph embedding vector determined according to the target graph structure data contains the logical relationship between the operation behaviors, the large language model can be used to improve the risk detection accuracy and detection efficiency of the user based on the target information and the target graph embedding vector.

[0154] The above describes specific embodiments of the specification. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims can be performed in an order different than the order in the embodiments and still achieve the desired result. In addition, the processes depicted in the figures do not necessarily require the particular order shown, or sequential order, to achieve the desired results. In certain implementations, multitasking and parallel processing can be advantageous or possible.

[0155] In the 1990s, it was possible to distinguish whether an improvement in a technology was a hardware improvement (e.g., an improvement in the circuit structure of a diode, transistor, switch, etc.) or a software improvement (an improvement in a method flow). However, as technology has advanced, many improvements in method flows today can be considered as direct improvements in hardware circuit structures. Designers almost always obtain the corresponding hardware circuit structure by programming the improved method flow into a hardware circuit. Therefore, it cannot be said that an improvement in a method flow cannot be implemented using a hardware entity module. For example, a programmable logic device (PLD) (e.g., a field programmable gate array (FPGA)) is an integrated circuit whose logic function is determined by user programming of the device. A digital system is "integrated" on a PLD by the designer programming it themselves, without having to ask a chip manufacturer to design and fabricate a custom integrated circuit chip. Moreover, instead of manually fabricating an integrated circuit chip, this programming is now mostly implemented using "logic compiler" software, which is similar to the software compiler used when developing a program, and the original code before compilation must also be written in a specific programming language, which is called a hardware description language (HDL), and there are many types of HDL, such as ABEL (Advanced Boolean Expression Language), AHDL (Altera Hardware Description Language), Confluence, CUPL (Cornell University Programming Language), HDCal, JHDL (Java Hardware Description Language), Lava, Lola, MyHDL, PALASM, RHDL (Ruby Hardware Description Language), etc., and the most commonly used are VHDL (Very-High-Speed Integrated Circuit Hardware Description Language) and Verilog. Those skilled in the art should also be aware that it is very easy to obtain a hardware circuit that implements a logical method flow by simply logically programming the method flow in one of the above-mentioned hardware description languages and programming it into an integrated circuit.

[0156] The controller can be implemented in any suitable way, for example, the controller can take the form of a microprocessor or processor and a computer readable medium storing computer readable program code, such as software or firmware, executable by the microprocessor, logic gates, switches, an application specific integrated circuit (ASIC), a programmable logic controller and an embedded microcontroller, examples of which include but are not limited to the following microcontrollers: ARC 625D, Atmel AT91SAM, Microchip PIC18F26K20 and Silicone Labs C8051F320, the memory controller can also be implemented as part of the control logic of the memory. Those skilled in the art will also know that, in addition to implementing the controller in pure computer readable program code, it is possible to implement the same functionality in the form of logic gates, switches, application specific integrated circuits, programmable logic controllers and embedded microcontrollers by logically programming the method steps. Such a controller can therefore be considered to be a hardware component, and the means included therein for implementing the various functions can also be considered to be structures within the hardware component. Alternatively, or even additionally, the means for implementing the various functions can be considered to be both a software module implementing the method and a structure within the hardware component.

[0157] The systems, apparatuses, modules or units illustrated by the above embodiments can be specifically implemented by computer chips or entities, or by products with certain functions. A typical implementation device is a computer. Specifically, the computer can be, for example, a personal computer, a laptop computer, a cellular phone, a camera phone, a smart phone, a personal digital assistant, a media player, a navigation device, an email device, a game console, a tablet computer, a wearable device, or a combination of any of these devices.

[0158] For the sake of description, the above apparatuses are described in functional division and are described respectively. Of course, the functions of each unit can be implemented in the same or multiple software and / or hardware when implementing one or more embodiments of the present specification.

[0159] Those skilled in the art will understand that the embodiments of the present specification can be provided as a method, a system, or a computer program product. Therefore, one or more embodiments of the present specification can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, one or more embodiments of the present specification can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to magnetic disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0160] The embodiments of the present specification are described with reference to flowcharts and / or block diagrams of the method, device (system) and computer program product according to the embodiments of the present specification. It should be understood that each flow and / or block in the flowcharts and / or block diagrams, and the combination of flows and / or blocks in the flowcharts and / or block diagrams can be implemented by computer program instructions. These computer program instructions can be provided to a general-purpose computer, a special-purpose computer, an embedded processor or other programmable electronic devices to produce a machine, so that the instructions executed by the computer or other programmable electronic devices generate a device that implements the functions specified in the flowcharts and / or block diagrams. Figure 1 The functions of one or more flows and / or blocks Figure 1 The functions of one or more flows and / or blocks

[0161] These computer program instructions can also be stored in a computer readable storage medium that can direct the computer or other programmable electronic devices to work in a specific manner, so that the instructions stored in the computer readable storage medium produce a product including instruction devices that implement the functions specified in the flowcharts and / or block diagrams. Figure 1 The functions of one or more flows and / or blocks Figure 1 The functions of one or more flows and / or blocks

[0162] These computer program instructions can also be loaded into a computer or other programmable electronic devices, so that a series of operation steps are performed on the computer or other programmable electronic devices to produce a computer implemented process, so that the instructions executed on the computer or other programmable electronic devices provide steps for implementing the functions specified in the flowcharts and / or block diagrams. Figure 1 The functions of one or more flows and / or blocks Figure 1 The functions of one or more flows and / or blocks

[0163] In a typical configuration, the computing device includes one or more processors (CPU), input / output interface, network interface and memory.

[0164] The memory can include non-persistent memory in the computer readable medium, random access memory (RAM) and / or non-volatile memory, such as read-only memory (ROM) or flash memory (flash RAM). The memory is an example of computer readable media.

[0165] Computer-readable media includes permanent and non-permanent, movable and non-movable media that can be implemented by any method or technology to store information. The information can be computer-readable instructions, data structures, program modules or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassette, magnetic tape, magnetic disk storage or other magnetic storage devices, or any other non-transmission medium that can be used to store information accessible to a computing device. According to the definition herein, computer-readable media does not include transitory media such as modulated data signals and carriers.

[0166] It should also be noted that the terms "comprising", "containing", or any other variant thereof are intended to cover non-exclusive inclusion, such that a process, method, article or apparatus that comprises a list of elements does not only include those elements, but can also include other elements not expressly listed or inherent to such process, method, article or apparatus. Without more limitations, the element defined by the statement "comprising a" does not exclude the presence of additional identical elements in the process, method, article or apparatus that includes the element.

[0167] Those skilled in the art will appreciate that embodiments of the present specification can be provided as methods, systems or computer program products. Therefore, one or more embodiments of the present specification can take the form of an entirely hardware embodiment, an entirely software embodiment or an embodiment combining software and hardware aspects. Moreover, one or more embodiments of the present specification can take the form of a computer program product implemented on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROMs, optical storage devices, etc.) containing computer-usable program code.

[0168] One or more embodiments of the present specification can be described in the general context of computer-executable instructions, such as program modules, being executed by a computer. Generally, program modules include routines, programs, objects, components, data structures, etc. that perform particular tasks or implement particular abstract data types. One or more embodiments of the present specification can also be practiced in distributed computing environments where tasks are performed by remote processing devices that are connected through a communication network. In a distributed computing environment, program modules can be located in both local and remote computer storage media including storage devices.

[0169] Each of the embodiments in the specification is described in a progressive manner, and the same or similar parts between the embodiments can be referred to each other. Each of the embodiments focuses on the difference from other embodiments. In particular, for the system embodiments, since they are basically similar to the method embodiments, the description is relatively simple, and the relevant parts can be referred to the part of the method embodiments.

[0170] The above only describes the embodiments of the specification and is not used to limit the file. For those skilled in the art, the specification can have various changes and changes. Any modification, equivalent replacement, improvement, etc. within the spirit and principle of the specification shall be included in the scope of claims of the specification.

Claims

1. A data processing method, comprising: Receive risk detection requests targeting specific users; In response to the risk detection request, log data containing the target user's operational behavior is obtained; Based on the target user's operational behavior contained in the log data, target graph structure data is determined. The target graph structure data includes target nodes determined according to the operational behavior, and edges between nodes determined according to the logical relationship between the operational behaviors. Based on a pre-trained feature extraction model, feature extraction processing is performed on the target graph structure data to obtain the target graph embedding vector corresponding to the target graph structure data, and key information extraction processing is performed on the log data to obtain target information; wherein, the feature extraction model is a model built based on the heterogeneous graph attention network algorithm; Based on the large language model, and using the target graph embedding vector and the target information, it is determined whether the target user is a risky user.

2. The method according to claim 1, wherein determining the target graph structure data based on the target user's operational behavior contained in the log data includes: The operation behaviors of the target user contained in the log data are clustered to obtain a first clustering result, and the events contained in the log data are determined based on the first clustering result. The target nodes in the target graph structure data are constructed based on the events, and the edges between the nodes in the target graph structure data are determined based on the logical relationships between the events.

3. The method according to claim 2, wherein clustering the target user's operational behavior contained in the log data to obtain a first clustering result includes: Based on the thread information corresponding to the target user's operation behavior contained in the log data, a first graph structure data is determined. The first graph structure data includes a first node determined according to the thread information, and edges between nodes determined according to the thread information. Based on the out-degree and / or in-degree information of each of the first nodes in the first graph structure data, determine the second node among the first nodes; The second node is clustered to obtain a second clustering result, and the first clustering result is determined based on the second clustering result.

4. The method according to claim 1, before determining the risk type corresponding to the operation behavior based on the target graph embedding vector and the target information according to the large language model, further comprising: Obtain a pre-trained large language model; Obtain historical log data, and the first risk type corresponding to the user's operation behavior contained in the historical log data; Based on the user's operational behavior contained in the historical log data, the second graph structure data is determined; Feature extraction processing is performed on the second graph structure data to obtain the first graph embedding vector corresponding to the second graph structure data, and key information extraction processing is performed on the historical log data to obtain the first key information; Based on the pre-trained large language model, and based on the preset prompt information, the first graph embedding vector, and the first key information, the second risk type corresponding to the user's operation behavior contained in the historical log data is determined; Based on the first risk type and the second risk type, the pre-trained large language model is fine-tuned to obtain the trained large language model.

5. The method according to claim 4, wherein determining the second risk type corresponding to the user's operation behavior contained in the historical log data based on the pre-trained large language model, preset prompt information, the first graph embedding vector, and the first key information includes: Based on the detection requirements corresponding to the target user, generate multiple sub-prompt messages with logical reasoning relationships corresponding to the preset prompt messages; Based on the pre-trained large language model, and using the multiple sub-cue information, the first graph embedding vector, and the first key information, the second risk type corresponding to the user's operation behavior contained in the historical log data is determined.

6. The method according to claim 1, wherein the key information extraction processing of the log data to obtain target information includes: According to a preset summary generation algorithm, key information is extracted from the log data to obtain summary information of the log data, and the summary information of the log data is determined as the target information.

7. The method according to claim 1, wherein the log data includes audit log data and / or operation log data, and the step of determining whether the target user is a risk user based on the target graph embedding vector and the target information according to the large language model includes: Based on the large language model, the risk type corresponding to the operation behavior is determined based on the target graph embedding vector and the target information, and the target user is determined to be a risk user based on the risk type corresponding to the operation behavior.

8. A data processing apparatus, comprising: The request receiving module is used to receive risk detection requests for target users; The first acquisition module is used to acquire log data containing the operation behavior of the target user in response to the risk detection request; The first determining module is used to determine target graph structure data based on the target user's operation behavior contained in the log data. The target graph structure data includes target nodes determined according to the operation behavior, and edges between nodes determined according to the logical relationship between the operation behaviors. The first extraction module is used to perform feature extraction processing on the target graph structure data according to a pre-trained feature extraction model to obtain the target graph embedding vector corresponding to the target graph structure data, and to perform key information extraction processing on the log data to obtain target information; wherein, the feature extraction model is a model built based on the heterogeneous graph attention network algorithm; The risk detection module is used to determine whether the target user is a risk user based on the target graph embedding vector and the target information, according to the large language model.

9. A data processing apparatus, the data processing apparatus comprising: processor; as well as A memory configured to store computer-executable instructions, which, when executed, cause the processor to: Receive risk detection requests targeting specific users; In response to the risk detection request, log data containing the target user's operational behavior is obtained; Based on the target user's operational behavior contained in the log data, target graph structure data is determined. The target graph structure data includes target nodes determined according to the operational behavior, and edges between nodes determined according to the logical relationship between the operational behaviors. Based on a pre-trained feature extraction model, feature extraction processing is performed on the target graph structure data to obtain the target graph embedding vector corresponding to the target graph structure data, and key information extraction processing is performed on the log data to obtain target information; wherein, the feature extraction model is a model built based on the heterogeneous graph attention network algorithm; Based on the large language model, and using the target graph embedding vector and the target information, it is determined whether the target user is a risky user.

Citation Information

Patent Citations

  • Construction behavior safety risk identification method based on knowledge graph and large language model

    CN119066151A

  • Abnormal user detection method and device for data security, equipment, storage medium and product

    CN119167356A