Abnormal behavior detection method and device, equipment, medium and program product

By constructing an abnormal behavior feature dataset and a user behavior model, and combining it with a large language model to generate an abnormal behavior baseline, the problem of low detection efficiency in UEBA technology is solved, and automated adaptation and efficient abnormal behavior detection are achieved.

CN121786337APending Publication Date: 2026-04-03CHINA MOBILE GROUP JIANGSU +1
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-12-25
Publication Date
2026-04-03

AI Technical Summary

Technical Problem

UEBA technology relies on known attack patterns, resulting in a lag in response to changes in business and personnel, requiring continuous manual updates to the rule base, and leading to low detection efficiency.

Method used

By constructing an abnormal behavior feature dataset, combining it with data lineage training set and user behavior attention model data, an initial abnormal behavior baseline is generated using a large language model, and then verified by a data security intelligent agent to achieve abnormal behavior detection.

Benefits of technology

It achieves automated adaptation of abnormal behavior baselines, covers risks throughout the entire lifecycle, reduces reliance on manual intervention, and improves detection efficiency and accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121786337A_ABST
    Figure CN121786337A_ABST
Patent Text Reader

Abstract

The invention provides an abnormal behavior detection method and device, equipment, a medium and a program product, and relates to the technical field of cloud computing, and the method comprises the steps: constructing an abnormal behavior feature data set based on a behavior template and sampling data, and obtaining a data consanguinity training set and user behavior attention model data based on the sampling data; based on the abnormal behavior feature data set, processing a data consanguinity training set and user behavior attention model data meeting a preset condition to construct an abnormal behavior information training set; performing feature processing on the abnormal behavior information training set and the baseline template knowledge base to obtain target feature data, inputting the target feature data into a large language model to generate an initial abnormal behavior baseline, and sending the initial abnormal behavior baseline to a data security agent for storage; and obtaining a target baseline passing through the data security intelligent experience certificate, and performing abnormal behavior detection based on the target baseline.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of data security technology, and in particular to an abnormal behavior detection method, apparatus, equipment, medium, and program product. Background Technology

[0002] User and Entity Behavior Analytics (UEBA) is increasingly being used in data security. Leveraging big data and artificial intelligence, it collects and analyzes massive amounts of data to establish baselines for user and entity behavior, thereby identifying anomalous activity and effectively compensating for the shortcomings of traditional rule-based protection solutions. However, currently, UEBA technology relies on known attack patterns, resulting in a lag in response to changes in business and personnel, and the rule base requires continuous manual updates, leading to low detection efficiency. Summary of the Invention

[0003] This application provides an abnormal behavior detection method and apparatus, which can solve the problems of UEBA technology relying on known attack patterns, lagging response to changes in business and personnel, the need for continuous manual updates to the rule base, and low detection efficiency.

[0004] To solve the above-mentioned technical problems, this application is implemented as follows: In a first aspect, embodiments of this application provide an abnormal behavior detection method, the method comprising: An abnormal behavior feature dataset is constructed based on behavior templates and sampled data, and data lineage training set and user behavior attention model data are obtained based on the sampled data. Based on the abnormal behavior feature dataset, the data lineage training set and user behavior attention model data that meet the preset conditions are processed to construct an abnormal behavior information training set. The abnormal behavior information training set and baseline template knowledge base are subjected to feature processing to obtain target feature data. The target feature data is input into a large language model to generate an initial abnormal behavior baseline. The initial abnormal behavior baseline is sent to the data security intelligent agent for storage. Obtain the target baseline verified by the data security intelligent agent, and perform abnormal behavior detection based on the target baseline.

[0005] Secondly, embodiments of this application provide an abnormal behavior detection device, the device comprising: The module is used to construct an abnormal behavior feature dataset based on behavior templates and sampled data, and to obtain a data lineage training set and user behavior attention model data based on the sampled data. The acquisition module is used to process the data lineage training set and user behavior attention model data that meet preset conditions based on the abnormal behavior feature dataset, so as to construct an abnormal behavior information training set. The processing module is used to perform feature processing on the abnormal behavior information training set and the baseline template knowledge base to obtain target feature data, input the target feature data into the large language model to generate an initial abnormal behavior baseline, and send the initial abnormal behavior baseline to the data security intelligent agent for storage. The generation module is used to obtain the target baseline verified by the data security intelligent agent and to perform abnormal behavior detection based on the target baseline.

[0006] Thirdly, embodiments of this application provide an electronic device, including a processor and a memory, wherein the memory stores a program or instructions that can run on the processor, and when the program or instructions are executed by the processor, they implement the steps of the abnormal behavior detection method as described in the first aspect.

[0007] Fourthly, embodiments of this application provide a readable storage medium storing a program or instructions that, when executed by a processor, implement the steps of the abnormal behavior detection method as described in the first aspect.

[0008] In this embodiment, cloud resource pool status data is acquired, and based on this data, an abnormal behavior feature dataset is constructed using behavior templates and sampled data. A data lineage training set and user behavior attention model data are then obtained based on the sampled data. Based on the abnormal behavior feature dataset, the data lineage training set and user behavior attention model data that meet preset conditions are processed to construct an abnormal behavior information training set. Feature processing is performed on the abnormal behavior information training set and baseline template knowledge base to obtain target feature data. This target feature data is then input into a large language model to generate an initial abnormal behavior baseline, which is then sent to a data security agent for storage. Finally, a target baseline verified by the data security agent is obtained, and abnormal behavior detection is performed based on this target baseline.

[0009] In this way, by integrating behavioral templates, data lineage, and user behavior model data, and combining feature processing, large model baseline generation, and intelligent agent verification closed loop, the automatic adaptation of abnormal behavior baselines is achieved. This can more accurately cover risks throughout the entire lifecycle, reduce reliance on manual intervention, and improve detection efficiency and accuracy to a certain extent. Attached Figure Description

[0010] Figure 1 A flowchart illustrating the abnormal behavior detection method provided in this application embodiment; Figure 2A flowchart of the AI ​​autonomous baseline analysis data behavior anomaly method provided in the embodiments of this application; Figure 3 A flowchart for constructing a behavior information recognition database provided in this application embodiment; Figure 4 A flowchart for creating an abnormal behavior information dataset provided in this application embodiment; Figure 5 A flowchart for generating an AI baseline is provided as an embodiment of this application; Figure 6 A flowchart for abnormal behavior analysis provided in the embodiments of this application; Figure 7 This is a schematic diagram of the abnormal behavior detection device provided in the embodiments of this application; Figure 8 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. Detailed Implementation

[0011] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0012] This application provides an abnormal behavior detection method, apparatus, device, medium, and program product. The embodiments of this application will be described in detail below with reference to the accompanying drawings and specific embodiments and application scenarios.

[0013] Please see Figure 1 , Figure 1 A flowchart illustrating an abnormal behavior detection method provided in this application embodiment is shown in the figure. The method includes: Step 110: Construct an abnormal behavior feature dataset based on behavior templates and sampled data, and obtain data lineage training set and user behavior attention model data based on the sampled data; In this step, the aforementioned behavioral templates can be understood as a standardized set of rules defining normal behavioral characteristics and patterns, and are a core component of the knowledge base. For example... Figure 2This can be user behavior paradigm templates (such as login modes, operation frequency), application behavior paradigm templates (such as business process specifications, data interaction protocols), security behavior paradigm templates (such as authorized access rules, authentication processes), or personalized rules adapted to different industries (such as telecom operators, government and enterprise customers) or business scenarios (such as data access, network configuration modification). The above-mentioned sampled data can be understood as raw data collected from business systems for behavioral analysis. This can be business logs (such as user access logs, server configuration change logs, base station operation and maintenance logs, etc.), security device logs, operation records at each stage of the data lifecycle (such as collection, storage, processing, and transmission logs, etc.), or time-series data, streaming data, or historical static data.

[0014] The aforementioned abnormal behavior feature dataset can be understood as a structured data collection that records abnormal behaviors and their corresponding features. It can include features extracted through static features, such as behavior frequency and distribution features, and features extracted through dynamic features, such as time domain, frequency domain, and deep learning. It can also cover feature samples of different anomaly types, such as unauthorized queries, unauthorized configuration modifications, and data leaks.

[0015] The aforementioned data lineage training set is training data that characterizes the entire lifecycle of data flow and its association with user access. It can be constructed using data matrices, access relationship matrices, and data access relationship views, and can also integrate supplementary data such as behavior timestamps, access frequency, and operation types to cover all stages of data collection, storage, use, and destruction.

[0016] User behavior focus model data refers to model-related data used to characterize normal or abnormal user behavior patterns. This data can include historical access datasets, target feature sets, model parameters such as Hofdinger tree and random forest ensemble model parameters, GMM probability distribution parameters, and can also cover node features and behavior-related weight data after graph network enhancement.

[0017] Step 120: Based on the abnormal behavior feature dataset, process the data lineage training set and user behavior attention model data that meet the preset conditions to construct an abnormal behavior information training set. In this step, the aforementioned preset conditions represent the criteria for judging the selection of data lineage training sets and user behavior attention model data. These can be feature integrity thresholds, data quality compliance conditions such as no missing values, uniform format, and abnormal sample ratio requirements, or business scenario adaptation rules, such as retaining only data related to operator 5G services.

[0018] The aforementioned abnormal behavior information training set can be understood as a comprehensive abnormal data set used for model training. It can be an integrated result of abnormal behavior feature datasets, data lineage training sets, and user behavior attention model data. It can also be supplemented with typical abnormal cases, such as internal employees accessing the site without authorization or external brute-force attacks, or it can be a normal behavior control sample.

[0019] Step 130: Perform feature processing on the abnormal behavior information training set and baseline template knowledge base to obtain target feature data, input the target feature data into the large language model to generate an initial abnormal behavior baseline, and send the initial abnormal behavior baseline to the data security intelligent agent for storage. In this step, the aforementioned baseline template knowledge base can be understood as a database that stores normal behavior rules and baseline construction templates. It can include industry compliance standards (such as data security regulations), business specifications (such as customer service operation procedures and operation and maintenance approval rules), historical baseline parameters, or structured templates categorized by scenario (such as data access and network configuration).

[0020] The feature processing described above can be understood as a series of operations to optimize the features of the data. This can include feature fusion, feature importance assessment, dimensionality reduction, temporal feature extraction, as well as preprocessing operations such as feature cleaning, normalization, and format conversion.

[0021] Target feature data can be understood as the core data used for baseline generation after feature processing. It can be low-dimensional feature vectors after dimensionality reduction, time-series behavioral features, high-value fusion features, or feature data combined with business scenario annotations (such as sensitive data operations, high-risk business behaviors, etc.).

[0022] The aforementioned large language model can be understood as an artificial intelligence model with natural language understanding and generation capabilities. It can be the DeepSeek large model, or other large language models that support prompt-driven generation, such as the GPT series and Tongyi Qianwen, adapting to scenarios such as baseline generation and anomaly detection rule derivation.

[0023] The aforementioned initial abnormal behavior baseline can be understood as the initial standard generated by the large language model for judging abnormal behavior. It can be a numerical threshold, such as querying more than 20 sensitive data entries in a single day, a rule-based judgment condition, or a probabilistic judgment standard. It can also be a specific baseline divided by business module, such as data access and network configuration, or user type, such as internal employees and external customers.

[0024] The aforementioned data security intelligent agent is an intelligent module running on a large model for data security management. It can store initial abnormal behavior baselines, core model parameters, and verification metrics, and can also have capabilities such as baseline verification, anomaly inference, and new threat identification, adapting to the security protection needs of different enterprises.

[0025] Step 140: Obtain the target baseline verified by the data security intelligent agent, and perform abnormal behavior detection based on the target baseline.

[0026] The aforementioned target baseline refers to a valid and fine-tuned standard for judging abnormal behavior. It can be the result of an initial baseline verified by experts and optimized through gray-scale testing, or it can be a baseline dynamically adjusted in response to business changes. The aforementioned abnormal behavior detection can be understood as the process of identifying abnormal behavior based on the target baseline. Real-time monitoring can be achieved through an abnormal behavior analysis engine, or abnormal risks can be analyzed through risk matrix construction, periodic indicator comparison, etc., covering both known and unknown novel anomalies.

[0027] In the abnormal behavior detection method implemented in this application, by integrating behavior templates, data lineage and user behavior model data, and combining feature processing, large model baseline generation and intelligent agent verification closed loop, the abnormal behavior baseline is automatically adapted, which can more accurately cover the risks throughout the entire life cycle, reduce reliance on manual methods, and improve detection efficiency and accuracy to a certain extent.

[0028] Optionally, the construction of the abnormal behavior feature dataset based on behavior templates and sampled data includes: A knowledge base is built based on behavior templates to store normal behavioral characteristics and judgment rules; Behavioral features are extracted from the sampled data using both static and dynamic feature extraction methods. Based on the knowledge base, abnormal behavior is identified from the behavioral features, and an abnormal behavior feature dataset is constructed based on the results of the abnormal behavior identification.

[0029] In this implementation, firstly, in "building a knowledge base based on behavior templates to store normal behavior features and judgment rules", "behavior templates" can be understood as a set of standardized rules that define the core features and boundaries of normal behavior, and are the core component of the knowledge base; "knowledge base" can be understood as a rule base that integrates normal behavior templates and abnormal judgment logic, and is used as the basis for subsequent abnormal behavior identification.

[0030] In some alternative implementations, building this knowledge base requires first establishing a clearly categorized template system based on different business scenarios and behavior types, specifically revolving around three types of paradigm templates: User behavior paradigm templates can cover features such as login mode, operation time and frequency, data access path, geographical location, terminal IMSI, and IP address. For example, for internal employees, the normal template for customer service to query user information is "associated work order number, operation during working hours, anonymized data display, ≤50 queries per day". The normal template for remote login by maintenance personnel is "VPN access, two-factor authentication, ≤3 logins per day, operation log". For external customers, the normal template for logging into the online business hall is "password + SMS verification, ≤10 high-risk business operations per session, high-risk operations outside of working hours triggering two-factor authentication". Application behavior paradigm templates can include features such as business process specifications, data interaction protocols, transaction processing modes, sensitive business, data collection, and data display. For example, the normal template for a billing system is "bill calculation error rate ≤ 0.1%, data backup once a day, and authentication required for external interface calls". The normal template for a base station management system is "parameter modifications require review, configuration change logs are uploaded in real time, and abnormal traffic alarm response time ≤ 1 minute". Security behavior paradigm templates: These involve features such as authorized access rules, authentication processes, and permission change modes. For example, a normal template for a firewall is "blocking unknown IPs from accessing the core database, blocking ≥100 attacks per day, and updating the malicious traffic feature database weekly." A normal template for data encryption is "encrypting user sensitive data storage, using the TLS 1.3 protocol for transmission, and changing the key every quarter."

[0031] Based on the three types of behavioral paradigm templates mentioned above, rules for judging abnormal behavior can be further defined, thus completing the construction of the knowledge base. For example, based on the user behavior paradigm template, rules such as "illegal query (querying multiple sensitive data without a work order or during non-working hours)" and "excessive operation (a customer's single session involves more than 10 high-risk services)" can be defined; based on the application behavior paradigm template, rules such as "error rate exceeding the standard (billing error exceeds 0.1%)" and "modification without review (single-person modification of base station parameters)" can be defined; based on the security behavior paradigm template, rules such as "insufficient interception rate (firewall blocks less than 50 attacks per day)" and "outdated encryption protocol (transmission using TLS 1.2 or below)" can be defined.

[0032] In the above statement "using static feature extraction method and dynamic feature extraction method to extract behavioral features from the sampled data respectively", "sampled data" can be understood as various log data generated by the business system, such as user data access logs, server configuration change logs, network operation logs, etc.

[0033] In some optional implementations, a combination of static and dynamic methods is used to extract behavioral features to ensure comprehensive feature coverage: The aforementioned static feature extraction refers to analyzing historical normal log data to statistically analyze the frequency and distribution characteristics of behaviors, and extracting key nodes and state transition rules in the business process. For example, it involves statistically analyzing the average number of times customer service personnel query user data per day and the distribution of the proportion of sensitive data queries. The aforementioned dynamic feature extraction refers to the capture of behavioral patterns, regularities, and anomalous signals that change over time in time-series or streaming data. Specifically, a dynamic feature library can be constructed using a combination of time-domain feature extraction, frequency-domain feature extraction, and deep learning extraction. Time-domain features include time window features, sequence pattern features, trend and abrupt change features, and periodic features; frequency-domain features are extracted based on Fourier transform, converting time-domain signals into frequency-domain components to capture periodic frequency features; deep learning extraction includes cluster analysis, time-series pattern mining, and sequence feature extraction, such as using LSTM to extract login IP time series or association rule mining.

[0034] The above-mentioned "identifying abnormal behavior based on the knowledge base and constructing an abnormal behavior feature dataset based on the results of abnormal behavior identification" is a process of matching the "behavioral features" obtained in the first two steps with the "knowledge base rules" to determine whether the behavior is abnormal and integrating abnormal behavior records to form a dataset.

[0035] For example, this step can be implemented as follows: Log data collection and preprocessing: Collect business system logs (such as user data access logs and server configuration change logs), clean and format the logs to ensure data availability; Abnormal Behavior Identification: Using the knowledge base constructed above, preprocessed log data is sampled and analyzed. Behavioral characteristics in the logs are matched against normal templates and abnormal rules in the knowledge base. For example, the behavior of "at 2 AM, customer service employee A, without associated work orders, queried 50 complete user ID numbers" is identified from the customer service system logs. Its characteristics of "outside working hours," "no work orders," and "high-frequency query of sensitive data" match the "illegal query" rule, thus determining it as abnormal behavior. The above-mentioned abnormal behavior feature dataset construction can organize and archive all identified abnormal behaviors and their corresponding features (such as behavior occurrence time, operating subject, data type, behavior frequency, etc.) to form an abnormal behavior feature dataset, providing data support for subsequent model training.

[0036] In the abnormal behavior detection method implemented in this application, by building a behavior template library, extracting dual-mode features and matching with the knowledge base, a more comprehensive coverage of behavioral features is achieved, the accuracy of abnormal identification is improved, and a high-quality abnormal behavior feature dataset can be constructed more efficiently.

[0037] Optionally, the dynamic feature extraction method includes time-domain feature extraction, frequency-domain feature extraction, and deep learning extraction; The time-domain features include at least one of the following: time window features, sequence pattern features, trend and mutation features, and periodic features; The frequency domain features are extracted based on Fourier transform; The deep learning extraction includes at least one of cluster analysis, temporal pattern mining, sequence feature extraction, and association rule mining.

[0038] The aforementioned "dynamic feature extraction methods include time-domain feature extraction, frequency-domain feature extraction, and deep learning extraction" clearly defines the core path of dynamic feature extraction. These three extraction methods can complement each other and cover dynamic behavior patterns in different dimensions. The aforementioned "temporal feature extraction" can be understood as directly analyzing the time dimension data of behavior, and is a method of extracting behavioral patterns, trends or sudden change signals that occur over time.

[0039] In some optional implementations, time window features can focus on the density of behavior within a specific time interval (e.g., the number of queries within one hour); sequence pattern features can capture the sequential logic of behavior (e.g., a fixed sequence of "login - query sensitive data - export"); trend and abrupt change features can identify sudden increases or decreases in behavior frequency (e.g., a sudden increase in daily query volume from 10 to 100); and periodicity features can uncover recurring behavioral patterns (e.g., regular access to a database every Monday morning). For example, when monitoring the time-series data of internal employee operations, if employee B continuously queries 20 pieces of sensitive user data during non-working hours (22:00-06:00), the characteristic of "non-working hours + high-frequency continuous operation," identified through trend and abrupt change features and time window features in time-domain feature extraction, is determined to be an abnormal time-domain feature.

[0040] "Frequency domain feature extraction" can be understood as a method of extracting hidden periodic frequency patterns by using Fourier transform to convert time-domain signals (such as the frequency of behavior that changes over time) into frequency-domain components. This "frequency domain feature extraction based on Fourier transform" can discover periodic anomaly patterns that are difficult to detect from a time-domain perspective through frequency dimension analysis. In some optional implementations, the Fast Fourier Transform can be used to optimize the computational model, improving the efficiency of frequency domain feature extraction and avoiding delays caused by excessive data volume. For example, when performing frequency domain analysis on time-series data of firewall-intercepted attacks, if a periodic peak is found in the number of attacks intercepted per hour at 2 AM, and this peak is unrelated to the normal business cycle, an anomaly pattern suspected of malicious scanning can be identified through frequency domain feature extraction. Similarly, performing frequency domain analysis on time-series data of high-risk user operations can capture the frequency characteristics corresponding to "intensive operations within a short period," which can be identified as abnormal behavior signals.

[0041] The aforementioned "deep learning extraction" can be understood as an extraction method that utilizes machine learning and deep learning algorithms to deeply mine log data and identify complex associations, clustering patterns, or time-series sequence patterns. In some optional implementations, cluster analysis can use machine learning algorithms to group log data according to behavioral characteristics and identify abnormal behaviors that deviate from normal clusters (e.g., most users' daily query volume is concentrated between 0-30 records, while a user with 100 daily queries is clustered into an abnormal cluster); time-series pattern mining can identify periodically recurring behavioral logic (e.g., batch data export behavior on fixed dates each month); sequence feature extraction can learn behavioral sequence patterns through models such as LSTM (e.g., learning the sequence pattern of "common IP → abnormal IP → login failure" to detect account theft risk); association rule mining can discover implicit associations between different behaviors (e.g., the combined association of "querying user ID number + exporting data + external IP login"). For example, when using LSTM to model the time-series data of user "login-operation-logout", if a sequence of "high-frequency querying of sensitive data directly after login without business operation intervals" is identified, which deviates significantly from the normal behavior pattern (login-business query-normal operation-logout), the deep learning model will determine that the sequence is an abnormal dynamic feature. As another example, by discovering the behavioral associations of "low-privilege employees, access to highly sensitive data and high-frequency communication with external IPs" through association rule mining, potential data leakage risks that are easily overlooked by humans can be captured.

[0042] In the abnormal behavior detection method of this application, by combining the above three types of dynamic feature extraction methods, dynamic abnormal signals in log data can be comprehensively captured from three dimensions: time regularity, frequency regularity, and complex correlation. This complements static feature extraction and ultimately improves the system's ability to identify unknown anomalies.

[0043] Optionally, the step of obtaining the data lineage training set based on the sampled data includes: Construct a data matrix, an access relationship matrix, and a data access relationship view based on the sampled data; The matrix elements of the data access relationship view are determined to be the product of the access relationship matrix elements and the data matrix elements; Based on the data access relationship view, the sampled data is integrated to construct a data lineage training set; The data matrix includes the relationship between the source and data type of the sampled data, and the access relationship matrix represents the user access behavior of the dataset involved in the sampled data. The dimensions of the data access relationship view include the user set, the sampled data source set, and the sampled data type set.

[0044] In this implementation, the aforementioned "data matrix" can be understood as a basic data structure that depicts the "source-type" relationship of the sampled data. Its core function is to clarify "what types of data were collected from different channels." In some optional implementations, the specific process of constructing the data matrix is ​​as follows: Data Collection and Preliminary Association: Data collection activities are subdivided according to specific scenarios into internet data collection, business hall data collection, partner data collection, sensor data collection, and third-party shared data. The corresponding data source set is denoted as S={s1,s2,⋯,sn} (where s1 represents internet data collection, s2 represents business hall data collection, and so on). Simultaneously, the collected data types are categorized, including user information, location, terminal information, access records, card information, service plans, consumption records, product information, partner manufacturers, business systems, API names, biometric data, voice data, video data, etc., and the data type set is denoted as T={t1,t2,...,tm} (e.g., t1 represents user information, t2 represents location information, etc.). For each data source s... i ∈S and data type t j For any ∈T, there exists a binary relation r. ij This indicates whether there is a correlation between the two, if the data source is s. i Data type t was collected j Then r ij =1; otherwise r ij =0, thus forming a preliminary data association matrix R=[r ij ]n×m.

[0045] Data Classification and Association Deepening: The initially constructed data matrix is ​​classified, such as into user basic information (including account opening information, etc.), business operation (including package changes, data usage, etc.), and base station operation and maintenance (including operation and maintenance records, etc.). Different types of data are associated through user ID and business ID to establish the data association relationship of "account opening information - package usage - data consumption - base station interaction (businesses involving base station services)", and to clarify the "data flow logic from account opening to package usage, data consumption and base station interaction" in the data lineage.

[0046] Dynamic updates to the data matrix: During the data lifecycle stages such as storage, use, processing, transmission, provision, publication, and destruction, new operational result data will be generated. Let the set of data lifecycle stages be L={l1,l2,...,l...} k Each stage of data manipulation alters the values ​​of elements in the data matrix R. For a given stage l... p Operations ∈L, if for data source s i and data type t j If the correlation has an impact (e.g., the data content is changed during data processing, causing previously unrelated data types to become related to their source), then r will be updated according to the operation rules. ij The updated data matrix is ​​still denoted as R, and its update process can be expressed as: R′=f(R,o lp ), where R′ is the updated data matrix, o lp Indicating in stage l p The specific operation is as follows: f is the update function, which adjusts the matrix elements accordingly based on different operation types (such as adding associations, deleting associations, modifying association data, etc.).

[0047] Meanwhile, the data lifecycle of the business system was analyzed. The initial data included user account opening information (name, ID number, network access time) and basic package information (package type, data allowance, tariff). The operation result data included package change records (change time, old and new packages), data usage records (usage time, data consumption, accessed content), and base station operation and maintenance records (operation and maintenance personnel, operation time, configuration change). These data were organized and supplemented into the data matrix by field (such as user ID, business ID, timestamp), with each row representing a business operation record and each column corresponding to data attributes, to ensure the integrity of the data matrix.

[0048] The aforementioned "access relationship matrix" can be understood as structured data recording the access relationships between "personnel and datasets," and its core function is to clarify "who accessed which data (including direct access and indirect inference access)." In some optional implementations, the specific process of constructing the access relationship matrix is ​​as follows: Personnel and Dataset Definition: Clearly define the set of personnel involved in each stage of the data lifecycle, P = {p1, p2, ..., p...} q This includes data collection personnel, maintenance personnel, marketing personnel, and management personnel during the data usage / processing phase; the dataset set D = {d1, d2, ..., d...} r In addition to the data that users directly access, it also includes datasets that users can infer from known data.

[0049] Establishing relationships: For person p i ∈P and dataset d j ∈D, there exists an association a ij Personnel p i Is it for dataset d? j If there is access behavior, then a ij =1; otherwise a ij =0, thus constructing the access relationship matrix A=[a ij ]q×r.

[0050] When determining associations, user reasoning behavior must be considered. Assume there exists a set of reasoning rules I = {i1, i2, ..., is}, for a given user p... i and known dataset D known Through reasoning rule i k ∈I can derive a new dataset D derived If p i Visited D known And satisfy inference rule i k Then p i With D derived The relationship a ij (d) j ∈D derived Let D be 1. The reasoning process can be represented as: D derived =g(D known i k ).

[0051] Where g is the inference function, which derives a new dataset from a known dataset based on different inference rules.

[0052] Supplementary information on reasoning behavior: When determining associations, user reasoning behavior must be considered. Assume there exists a set of reasoning rules I = {i1, i2, ..., i...}. s For a certain person p i and known dataset D known Through reasoning rule i k ∈I can derive a new dataset D derived If p i Visited D knownAnd satisfy inference rule i k Then p i With D derived The relationship a ij (d) j ∈D derived Let ) be 1, the reasoning process can be represented as: D derived =g(D known i k ), where g is the inference function, which derives a new dataset from a known dataset based on different inference rules.

[0053] Meanwhile, based on the data matrix and its relationships, the system tracks user and maintenance personnel access to different data (e.g., user A checks their remaining data allowance today and consumes 5GB of data; maintenance personnel B logs into the base station management system to modify configurations). The access frequency or operation type is then added to the access relationship matrix to enrich the behavioral description dimensions of the matrix.

[0054] The aforementioned "data access relationship view" can be understood as a three-dimensional relational model that integrates the "data matrix" and the "access relationship matrix." Its core function is to clearly present the complete lineage path of "personnel-data source-data type." The claims explicitly state that "the dimensions of the data access relationship view include the user set, the sampled data source set, and the sampled data type set," and that "the matrix elements of the data access relationship view are the product of the access relationship matrix elements and the data matrix elements." Specifically, in some optional embodiments, the data matrix R is combined with the access relationship matrix A to form a data access relationship view matrix V, whose dimensions are jointly determined by the personnel set, the data source set, and the data type set, expressed as V=[v ijk ]q×n×m. View matrix element v ijk The calculation method is as follows: v ijk =a ij ⋅r ik , where a ij For personnel p i With dataset d j Access relationship, r ik For data source s i With data type t k The relationship between v; ijk When =1, it indicates that there are personnel p. i Through data source s j Accessed data type t k Relevant data; when v ijkA value of 0 indicates that the access relationship does not exist. The data access relationship view clearly identifies a user's data lineage, i.e., how users acquire, use, and manipulate data, thus providing structured training data for subsequent anomaly behavior analysis. Furthermore, the access relationship matrix can be visualized to generate an intuitive access relationship view, clearly presenting the interaction between users / operations personnel and data, serving as the foundation for data lineage analysis.

[0055] The above-mentioned "integrating the sampled data to construct a data lineage training set based on the data access relationship view" is the final integration and improvement of the aforementioned structured related data, specifically as follows: In some optional implementations, the data access relationship view constructed above is combined with other relevant data (such as behavior timestamps, behavior frequencies, etc.) to supplement the behavioral context information of the data, forming a complete abnormal behavior information training set for subsequent model training and abnormal behavior detection tasks. For example, data lineage integration can be achieved by sorting out the data lifecycle lineage, such as user information from "collection in the business hall → storage in the business system → access by customer service / maintenance on demand → periodic archiving / destruction". Abnormal data flow in each link (such as "user information is directly transmitted to a third party without entering the business system after collection") can be extracted as abnormal data lineage samples to supplement the training set. At the same time, using the data lineage results (access relationship view, etc.), an initial user behavior model is constructed using a Gaussian Mixture Model (GMM). GMM clusters the patterns of user access data (such as the access frequency distribution during normal package query periods, the proportion of different traffic consumption intervals; the frequency of routine base station configuration modifications by maintenance personnel) to identify the routine behavior patterns of users and maintenance personnel (such as most users consume traffic between 07:00-23:00, and the daily traffic consumption is in the range of 0-50GB; maintenance personnel modify base station configuration ≤5 times per month). This serves as the normal behavior baseline of the initial model, further improving the abnormal judgment benchmark of the data lineage training set.

[0056] For example, the data matrix R can be combined with the access relationship matrix A to form a data access relationship view. Let the data access relationship view matrix be V, whose dimensions are jointly determined by the set of personnel, the set of data sources, and the set of data types, and expressed as V=[vijk]q×n×m. View matrix element v ijk The calculation method is as follows: v ijk =a ij ⋅r i; Where a ij For personnel p i With dataset d j Access relationship, r ik For data source s i With data type t kThe relationship between v. ijk When =1, it indicates that there are personnel p. i Through data source s j Accessed data type t k Relevant data; when v ijk When =0, it means that the access relationship does not exist.

[0057] In the abnormal behavior detection method of this application, by constructing a three-dimensional correlation matrix, performing matrix element multiplication operations and integrating sampled data, a structured model of the entire data lineage is achieved, which can provide accurate underlying data support for abnormal behavior detection.

[0058] Optionally, obtaining the user behavior attention model data includes: Obtain the historical access dataset from the sampled data, the historical access dataset containing access record feature vectors and abnormal behavior annotation labels; The information gain method is used to obtain the information gain value of each feature vector in the historical access dataset. Feature vectors with information gain values ​​greater than a preset threshold are selected to form a target feature set. Based on the target feature set, the Hofding tree model is used as the initial anomaly recognition model, and the random forest algorithm is used for ensemble learning to obtain the anomaly recognition ensemble model. The model based on the anomaly detection model generates initial user behavior model data based on the inference results of the historical access dataset. The initial user behavior model data is trained using expectation maximization and graph network augmentation to obtain user behavior attention model data.

[0059] In this implementation, the "historical access dataset" can be understood as a basic dataset extracted from the sampled data, containing past operation records and anomaly annotations of users and maintenance personnel, and is the core data source for model construction. In some optional implementations, the process of obtaining the historical access dataset is as follows: The scope of data collection can be understood as collecting historical access data of users and maintenance personnel of the operator's 5G business system over the past year, covering login time, accessed business modules (package management, traffic query, base station maintenance), operation content (such as package change type, traffic usage peak, base station configuration modification items), etc., to ensure that the data covers core business scenarios.

[0060] The above-mentioned dataset structure can be defined as follows: Let the historical access dataset be D={x} i ,y i} (i from 1 to N), where x i ∈RD is the feature vector of the i-th access record, y i∈{0,1} indicates whether it is abnormal behavior (1 is abnormal); the feature space can be represented as x i =[x i1 ,x i2 ,…,x iD ], where x id The d-th feature of the i-th record (such as access time, page dwell time, click frequency, etc.) is defined, specifying the feature dimension and tag attribute of each data point.

[0061] The aforementioned "information gain method" can be understood as a mathematical method that filters out high-value features by quantifying their contribution to "distinguishing between normal and abnormal behavior," thus avoiding redundant features from affecting model efficiency. The specific implementation process is as follows: Feature extraction: Extract candidate features from historical access datasets, such as login time period, operation frequency, traffic consumption, package change cycle, base station configuration modification content, etc., covering the core dimensions of operation behavior.

[0062] Information gain calculation: Information entropy of the original dataset: H(D) = −∑ 1 k=0 p(k)log2p(k), Where p(k) = |{yi=k}| / N is the probability of category k; Conditional entropy after feature d partitioning: H(D|x) d )=∑ v∈Vd |D v d | H (D v d ) / N, Where V d Let D be the set of values ​​for feature d. v d ={x i |x i d =v}; Information gain: IG(d) = H(D) − H(D|x) d ).

[0063] Feature selection: Select features where IG(d) > τ (τ is a preset threshold) to form the target feature set F = {f1, f2, ..., f M (M≤D).

[0064] The aforementioned "Hoeffding Tree model" can be understood as a decision tree model suitable for streaming data and capable of efficiently handling dynamic features. Its core advantage lies in rapidly determining node splitting rules based on statistical theory. In some optional implementations, the model construction process is as follows: Furthermore, the splitting condition can be constructed based on the Hoeffding inequality, namely P(μ^1−μ^2≥ϵ)≤2e −2nϵ2 , where μ^1, μ^2 are the means of the two classes of samples, n is the number of samples, and ϵ is the error bound; Node splitting rules: Find the feature f and the threshold θ so that the information gain after splitting meets the preset conditions; During model training, the model can be trained to learn the judgment rules of abnormal behavior based on the feature patterns of normal / abnormal behavior in historical data (such as the normal package change cycle is 3-12 months, while abnormal changes may occur multiple times within 1 month; normal base station configuration modifications are concentrated on working days, while abnormal changes occur frequently on non-working days), and initially have the ability to identify abnormal user and maintenance personnel behavior.

[0065] The aforementioned "random forest ensemble learning" can be understood as constructing multiple independent Hofting trees, fusing the prediction results of each tree, reducing the bias and variance of a single model, and improving recognition stability. In some optional implementations, the optimization process can be as follows: the above ensemble model construction uses random forest as the ensemble method, assuming the construction of T Hofting trees {ht} (t from 1 to T), and the prediction function of the ensemble model is y^=argmax. k ∑ T t=1 Ⅱ(h t (x)=k), where Ⅱ(⋅) is the indicator function and k is the category label; Loss function definition: The loss function for an ensemble model is: L ensemble =−1 / N∑ N i=1 ∑ 1 k=0 Ⅱ(y i =k)log(1 / T∑ T t=1 I(h t (x i The model parameters can be optimized by minimizing the loss function. Sample input optimization: Input historical data and labeled abnormal samples (such as known package theft (frequent changes in package within a short period of time, excessive data consumption with low rates), and unauthorized access to base stations (unauthorized personnel modifying base station configurations)) to train and optimize the random forest model, enabling it to more accurately identify abnormal behavior, such as accurately distinguishing between normal user traffic peaks (such as watching videos or downloading large files) and malicious traffic attacks (such as being implanted with malicious programs that continuously upload data).

[0066] The aforementioned "initial user behavior model data" can be understood as basic data characterizing the probability distribution of normal user behavior based on historical access data and ensemble model inference results. It is modeled using a Gaussian mixture model (GMM), and in some optional implementations, the generation process is: p(x|θ)=∑ M m=1 π m N(x∣μ m ,Σ m ); Where θ={π m ,μ m ,Σ m} M m=1 , π m For the mixing coefficient, (∑ M m=1 π m =1), N(x|μ) m ,Σ m The distribution is Gaussian. The above modeling method can transform the inference results of the ensemble model on the historical access dataset (such as the probability of abnormal behavior and the clustering results of normal behavior) into structured initial user behavior model data, which can lay the foundation for subsequent training enhancement.

[0067] The aforementioned "Expectation-Maximization (EM) algorithm" can be understood as a training method that optimizes the fitting effect of a Gaussian mixture model by iteratively calculating the posterior probabilities of latent variables and updating model parameters. In some optional implementations, the training process can be: E-step: Calculate the posterior probabilities of latent variables γi,m = π m N(x i |μ m ,Σ m ) / ∑ M j=1 π j N(x i |μ j ,Σ j ); M-step parameter update: π m new =1 / N∑ N i=1γi,m; μ m new =∑ i=1 N γi,m xi / ∑ i=1 N γi,m; Σ m new =∑ i=1 N γi,m(x i -μ m new (x) i -μ m new )T / ∑ i=1 N γi,m.

[0068] The aforementioned "graph network enhancement" can be understood as modeling user behavior as a heterogeneous graph. Through message passing and attention mechanisms, it can strengthen the process of extracting behavioral association features and ultimately optimize the model's ability to identify complex related behaviors.

[0069] For example, the process of supplementing and improving the user behavior attention model data can be as follows: Based on the model enhanced by the above training, collect abnormal behavior data: internal employees accessing the platform beyond their authorized permissions (such as interns accessing high-value user business information), external users brute-forcing logins (trying high-frequency incorrect passwords), etc. Integrate these data with abnormal data lineage samples to form an abnormal behavior information training set, further enrich the scenario coverage of the user behavior attention model data, and ensure that the model can accurately capture various abnormal behavior patterns.

[0070] In the abnormal behavior detection method of this application, by using information gain to filter features, Hofding trees, random forest ensemble modeling, EM training, and network enhancement, more accurate feature filtering and improved model robustness can be achieved, thereby improving the accuracy of characterizing user behavior patterns and strengthening the ability to identify abnormal behavior.

[0071] Optionally, the step of performing expectation-maximization training and graph network augmentation on the initial user behavior model data to obtain user behavior attention model data includes: Construct a heterogeneous graph, wherein the set of nodes in the heterogeneous graph includes users, pages, and operation types, and the set of edges in the heterogeneous graph represents the behavioral relationships between the nodes; The heterogeneous graph is aggregated based on the message passing mechanism of the graph neural network, and edge weights and node features are fused to output a node feature representation vector. The node feature representation vector is processed by multi-level graph attention to obtain behavior association weight data and optimized node feature representation vector; The user behavior attention model data is formed by integrating the model parameter data after expected value maximization training, the sampled weight data, the node feature representation vector, and the behavior-related weight data.

[0072] In this implementation, "Expectation-Maximization (EM) training" can be understood as a process of iteratively calculating the posterior probabilities of latent variables, updating model parameters, and then fusing and optimizing the outputs of the initial user behavior model and the anomaly detection model. In some optional implementations, the fusion process can refer to... Figure 4 The inference results of the initial user behavior model of S301 (characterizing the probability distribution of user behavior based on Gaussian mixture model GMM) and the anomaly recognition ensemble model of S302 (Hofding tree + random forest) are combined. The latent variable posterior probability is calculated by the E-step of the EM algorithm and the model parameters are updated by the M-step. This yields basic model data that better reflects the actual behavior patterns of users (including maintenance personnel), laying the foundation for subsequent graph network enhancement.

[0073] The aforementioned "heterogeneous graph" can be understood as a graph structure containing multiple types of nodes and edges, which can be used to characterize the behavioral relationships between different entities and is the core carrier for graph network enhancement. The specific construction process is as follows: The nodes and edges mentioned above can be defined as follows: The set of nodes V={user, page, operation type} of the heterogeneous graph G=(V,E) corresponds to the operation subject, access object and behavior action in the business scenario, respectively; the set of edges E represents the behavioral association between each node, such as the association of "user A - query operation - user information page" and "operation and maintenance personnel B - modification operation - base station configuration page".

[0074] Furthermore, to improve model processing efficiency, random walk sampling (RWR) can be used to define the sampling probability, with the formula: π(v ′ )=(1−α)π(v ′ )+α∑ v∈N(v) π(v) / d(v); Where α is the restart probability, N(v) are the neighbors of node v, and d(v) is the node degree. This sampling method focuses on nodes associated with core behaviors, reducing the impact of redundant data on the model. The above "graph aggregation processing" can be understood as using the message passing mechanism of a graph neural network (GNN) to fuse the features of a node itself with those of its neighbors, generating a more discriminative node feature representation vector. In some optional implementations, the aggregation process can be achieved using the following formula: h v (l+1) =σ(∑ u∈N(v) wv,u ⋅h u (l) +W (l) h v (l) ); Where hv(l) is the representation of node v at layer l, wv,u are the edge weights (characterizing the strength of the association between nodes), W(l) is the learnable matrix, and σ is the activation function (such as ReLU). This process, by fusing edge weights and node features, can expand the features of a single node into comprehensive features that include the associated context. For example, it can fuse the features of "user C" with the associated features of "high-frequency query operations" and "sensitive data pages" to highlight the association signals of abnormal behavior.

[0075] The aforementioned "multi-level graph attention processing" can be understood as the process of allocating node association weights through an attention mechanism, strengthening important association features and weakening irrelevant associations, thereby obtaining optimized node feature representation vectors and behavior association weight data. Specifically, it consists of two steps: a combination of single-head attention and multi-head attention. 1. Single-head attention mechanism: Calculate the attention coefficient between nodes: This can be done using the formula e. v,u =LeakyReLU(a T [Wh v ⊕Wh u ]) to achieve, where a T Here, W is the attention weight vector, W is the feature transformation matrix, and ⊕ represents concatenation, which can be used to quantify the importance of the relationship between nodes. Attention coefficient normalization: can be achieved using the formula α v,u =exp(e v,u ) / ∑ k∈N(v) exp(e v,k The coefficients are normalized to ensure that the sum of the weights is 1; Node representation update: Update node features based on normalized weights, using the formula h. v ′=σ(∑ u∈N(v) α v, u Wh u This allows node features to focus more on key relationships.

[0076] 2. Multi-headed attention combination: To further enhance feature representation capabilities, multi-head attention combination optimization can be employed, with the formula: hvmulti=∑k=1Kσ(∑u∈N(v)αv,ukWkhu), where K is the number of heads and α k v,u and W kLet be the parameters of the k-th head. By combining multiple attention mechanisms, node association features can be captured from different dimensions, improving the model's adaptability to complex behaviors.

[0077] Furthermore, combining anomaly detection and behavior representation learning, the joint loss L=L rec +λL anomaly Where: Reconstruction loss (based on graph autoencoder): ; It can be used to ensure the accuracy of graph structure reconstruction, p(e^ v,u ) represents the probability of the predicted edge existing; Anomaly detection loss:

[0078] It can be used to enhance the model's ability to distinguish abnormal behaviors.

[0079] Furthermore, the parameter θ is updated using stochastic gradient descent (SGD), as shown in the formula: θ t+1 =θ t −η∇θL(θ t ), where η is the learning rate, and the gradient ∇θL is calculated through backpropagation to continuously optimize the model parameters to minimize the loss.

[0080] It can integrate the expected value maximization training model parameter data (fusion of basic model parameters), sampling weight data (weights obtained by random walk sampling), node feature representation vectors (node ​​features optimized by graph aggregation and graph attention), and behavior association weight data (association weights calculated by the attention mechanism) to form structured user behavior attention model data.

[0081] For example, the model data can be adapted to the risk prevention and control scenarios of operators' 5G services: graph aggregation maps user behavior, operation and maintenance and data lineage to graph structure, and multi-level graph attention network focuses on the association weight between different users and different base station nodes, which strengthens the ability to identify complex related behaviors such as "group package theft, such as multiple accounts frequently changing packages and maliciously consuming traffic", and finally realizes accurate monitoring and early warning of abnormal user and equipment behavior in operators' 5G service systems; at the same time, it can also be adapted to other scenarios such as fixed network services and broadband operation and maintenance, and adapt to different service needs by adjusting data and feature dimensions.

[0082] Optionally, the step of processing the abnormal behavior information training set and baseline template knowledge base, and inputting the processed feature data into the large language model to generate an initial abnormal behavior baseline includes: The abnormal behavior information training set and the baseline template knowledge base are fused to obtain a fused feature vector, and the fused feature vector is then dimensionality-reduced to obtain dimensionality-reduced features. Combining the aforementioned dimensionality reduction features, typical abnormal samples, and normal behavior rules, a structured prompt word containing business scenarios, feature dimensions, and anomaly detection requirements is constructed; A time-series dataset is constructed using dimensionality reduction features as input and behavior type as label. The time-series dataset is then input into a long short-term memory network to obtain the time-series behavior features output by the long short-term memory network. The temporal behavior features and the structured cue words are input into a large language model to generate an initial abnormal behavior baseline.

[0083] In this implementation, the aforementioned "feature fusion" can be understood as the process of integrating the behavioral features of the abnormal behavior information training set with the rule features of the baseline template knowledge base to form a fused feature vector that comprehensively covers the dimensions of data security risks; "dimensionality reduction" is the operation of simplifying high-dimensional fused features through algorithms while retaining core risk information. In some optional implementations, reference can be made to... Figure 5 The specific process is as follows: Feature fusion and importance assessment: First, abnormal features (such as features of unauthorized queries and illegal configuration modifications) in the abnormal behavior information training set are fused with normal behavior rule features (such as operations during working hours and double-person review processes) in the baseline template knowledge base to form a fused feature vector; then, information gain, mutual information, or XGBoost algorithms are used to assess feature importance—for example, by using the information gain formula IG(x)=H(y)−H(y|x) (where H(y) is the information entropy of the category label, H(y|x) is the conditional entropy under feature x, and C is the number of categories), high-contribution features such as "accessing sensitive data outside of working hours" and "modifying core parameters without approval" are selected, and redundant information is eliminated.

[0084] in, ; Dimensionality Reduction: For high-dimensional fused feature vectors, Principal Component Analysis (PCA) or t-distributed random neighborhood embedding (t-SNE) is used for dimensionality reduction. Taking PCA as an example, the covariance matrix Σ of the feature vectors is calculated, and its eigenvalues ​​λi and eigenvectors vi are obtained through eigenvalue decomposition. The eigenvectors corresponding to the m largest eigenvalues ​​are selected to construct the projection matrix P=[v1,v2,...,vm]. Then, f′=P T f projects the original feature vector f into a low-dimensional space; if the t-SNE algorithm is used, the original 20-plus features can be compressed into 6-8 core dimensions, which simplifies the subsequent model calculation and highlights the core features of data security risks.

[0085] Furthermore, model co-training can be performed based on LSTM and cross-entropy to adapt to data security temporal behavior, specifically as follows: Operator data access and operations exhibit temporal patterns (e.g., internal personnel routinely access data during work hours according to procedures, and network configuration modifications occur at fixed intervals). Therefore, a Long Short-Term Memory (LSTM) network is used for modeling. Taking internal personnel accessing user data as an example, the LSTM learns historical access patterns (e.g., employee A queries user information in a specific area every Monday and Wednesday from 9-11 AM). Combined with a cross-entropy loss function, the difference between the predicted "compliant access probability distribution" and the actual "compliant / illegal access labels" is compared to optimize the LSTM parameters and improve the ability to identify data violation behaviors (e.g., employees frequently downloading user information outside of work hours, and unknown IP addresses tampering with base station configurations).

[0086] Furthermore, baselines can be generated and parameter tuning verified, thus ensuring the accuracy of the data security baseline: Model prediction: Using a trained LSTM model, predict the new data access and operation behaviors of operators, output a "data security risk score", and initially generate a data security behavior baseline (such as normal employees accessing ≤50 pieces of sensitive data per day, accessing sensitive data 0 times during non-working hours; base station configuration modifications need to match the approval process with a probability of more than 80%).

[0087] Validation dataset: Select one year of labeled data from the operator's history (including compliant operations and violations, such as employees illegally selling user data and external attacks tampering with configurations) as the validation set, and compare the model's prediction results with the labeled data. If the baseline is found to be insufficient in identifying "internal personnel abuse of power (high-privilege account violations)," adjust the model parameters (such as increasing the LSTM memory step size and strengthening the weights of high-risk features), and retrain the model until the baseline achieves a detection accuracy of ≥97% and a recall of ≥92% for data violations on the validation set, meeting data security control requirements.

[0088] The aforementioned "structured prompt words" can be understood as standardized text instructions that integrate dimensionality reduction features, abnormal samples, normal rules, and business scenarios. These instructions provide clear generation basis for large models, preventing them from deviating from the baseline of actual needs. In some optional implementations, the construction process involves combining the core features after dimensionality reduction (such as "excessive querying of sensitive data" and "parameter modification without review"), typical abnormal samples (such as "querying more than 10 pieces of data before anonymization without work order association"), and normal behavior rules in the baseline template knowledge base (such as "core parameter modification requires double review"). This process clearly includes core elements such as business scenarios (such as operator data security and government / enterprise customer data protection), feature dimensions (such as access time, operation type, and data volume), and abnormal judgment requirements (such as "identifying behaviors exceeding permissions, exceeding frequency, and non-compliant processes"), assembling them into structured prompt words.

[0089] The aforementioned "time-series dataset" refers to training data organized along the time dimension, using dimensionality-reduced features as input and behavioral type (normal / abnormal) as label; the "Long Short-Term Memory Network (LSTM)" is used to learn the time-series patterns of behavior, outputting "time-series behavioral features" that reflect temporal correlations. In some optional implementations, the specific process can be as follows: Time-series dataset construction: Arrange the dimensionality-reduced features in chronological order, label the behavior type corresponding to each feature (1 for abnormal, 0 for normal), and form a time-series dataset to adapt to the temporal characteristics of data access and operation behavior (such as internal personnel regularly accessing data during working hours and configuration modifications having a fixed cycle).

[0090] LSTM Model Training and Feature Extraction: The time-series dataset is input into the Long Short-Term Memory network. The cross-entropy loss function is used as the optimization objective (formula: y^(N-C) = y^(N-C), where N is the number of samples and C is the number of classes. Stochastic Gradient Descent (SGD) is employed to adjust the model parameters, minimizing the difference between the predicted output y^ and the true label y. The cross-entropy loss function is: ; LSTM learns historical time-series patterns (such as employees querying specific data at fixed times each week) to capture time-series correlation features such as "common IP - abnormal IP - login failure" and "collection - not stored - external transmission", and outputs time-series behavioral features that can characterize the time regularity of behavior.

[0091] Further, a baseline is generated and validated through parameter tuning until the preset target is achieved. The validated abnormal behavior baseline data is then sent to the data security agent running on the DeepSeek large model.

[0092] After training, the model is used to predict the training dataset, and the predicted output distribution of normal behavior samples is calculated. For numerical prediction results, the mean μ and standard deviation σ are calculated, and the baseline range of normal behavior is defined as [μ−kσ,μ+kσ], where k is a coefficient set according to actual needs; for probabilistic prediction results, a probability threshold τ is set, and behaviors with a predicted probability greater than τ are judged as normal behaviors.

[0093] The generated baseline is validated using a reserved validation dataset. The proportion of normal behavior samples falling within the baseline range (recall) and the proportion of abnormal behavior samples misclassified as normal behavior (false positive rate) are calculated. Baseline parameters (such as k and τ) are adjusted based on the validation results until satisfactory validation performance is achieved.

[0094] The finalized baseline data of abnormal behavior is organized into a standard format and encapsulated and sent according to the interface requirements of the DeepSeek large model. The data security agent receives the baseline definition parameters, feature descriptions, and relevant verification metrics, enabling it to utilize DeepSeek's inference capabilities for subsequent abnormal behavior detection and analysis tasks.

[0095] In this implementation, the extracted temporal behavioral features and the constructed structured prompts are input into a large language model. Leveraging the model's knowledge reasoning capabilities, an initial baseline of abnormal behavior conforming to the business scenario is generated. In some optional implementations, the large language model can be the DeepSeek model. The input content can include both temporal behavioral features (such as the temporal pattern of "more than 50 sensitive data queries per day") and structured prompts (such as "Based on the operator's data security scenario, identify whether the following behaviors are abnormal:..."). 1. Query more than 10 user records before data anonymization, excluding work order-related queries; 2. Modification of relevant parameters in the business system without review... Please generate an abnormal behavior judgment baseline). The model combines the two to output a clear initial abnormal behavior baseline, such as "more than 20 queries of sensitive user data by a single account in a single day are considered abnormal", "modification of core system parameters without review process is considered abnormal", "more than 3 high-risk operations outside of working hours trigger abnormal alarms", etc.

[0096] For example, this process can be adapted to operator data security scenarios: core features such as "downloading more than 10GB of user communication records in a single transaction" and "unauthorized modification of base station keys" are selected through feature fusion. After t-SNE dimensionality reduction, a time series dataset is constructed. LSTM learns the time series pattern of "base station configuration modification - no approval record - traffic surge". Then, combined with structured prompt words input into the DeepSeek large model, an initial baseline adapted to scenarios such as 5G services and network configuration security is generated, which can cover time series correlation anomalies and rule-based anomalies relatively accurately.

[0097] Furthermore, the verified "data security behavior baseline," "model core parameters (LSTM structure, feature weights)," and key data security prevention and control scenarios for operators (such as monitoring cross-border transmission of user information and protecting core network element configuration data) can be packaged into security detection tasks and sent to the DeepSeek data security agent. The agent can then rely on the knowledge reasoning and generalization capabilities of the large model to conduct in-depth analysis of real-time data from operators (such as real-time user data access streams and network configuration change logs). On the one hand, it can identify new data security threats not covered by the model (such as using AI to generate false identities to steal user data and configuration tampering attacks targeting new 5G protocols). On the other hand, it can combine operator data security policies (such as requirements for local storage of user data) to output more realistic security warnings (such as a third-party partner attempting to illegally export user data overseas) and handling solutions (such as automatically blocking illegal operations and triggering manual audit processes), helping operators build an intelligent and collaborative data security protection system. This embodiment focuses on the core data security needs of operators, and realizes the construction of data security baseline and intelligent agent collaboration through process implementation. It can be adapted to scenarios such as user data management, network configuration security, and business system data protection. It dynamically adjusts features and models to enhance the operator's data security governance capabilities and protect user privacy and network operation security.

[0098] Optionally, the abnormal behavior detection based on the target baseline includes: sending the target baseline configuration to the abnormal behavior analysis engine for abnormal behavior detection; The method further includes: The monitoring rules of the target baseline are loaded and distributed through the abnormal behavior analysis engine, and then integrated into the corresponding risk model. Based on the monitoring rules, the data security events and potential events monitored in real time by the filtering engine are filtered, the type information of the data security events and potential events is extracted, and then the life cycle stage information and security attribute information corresponding to the type information are obtained. A data risk matrix is ​​constructed based on the aforementioned lifecycle stage information and security attribute information; Based on the aforementioned data risk matrix, security risk indicators for each period are obtained. Identify and generate alerts for abnormal risk indicators that exceed the baseline threshold from the security risk indicators of each period.

[0099] In this implementation, the "target baseline" can be understood as a standardized judgment rule that, after verification by a data security intelligent agent and fine-tuning by an expert system, can be directly implemented for anomaly detection; the "abnormal behavior analysis engine" is the core execution module that carries the baseline monitoring rules and enables real-time detection. In some optional implementations, this step requires the verification, fine-tuning, and activation of the target baseline before configuration distribution, as detailed below: Build a remote expert system to manually verify, fine-tune, and activate the initial abnormal behavior baseline generated by AI: Baseline remote verification: This includes the distribution of baseline review tasks, system verification, and recording of verification results.

[0100] Baseline review task distribution: The system automatically pushes the AI-generated baseline draft to experts in the corresponding fields, intelligently assigns tasks based on the experts' areas of expertise and online status, and reminds experts to process the task through system notifications, emails, and other means.

[0101] System verification: Provides an intuitive visualization of baseline parameters, making it easy for experts to view data such as the behavioral sample size covered by the baseline, the accuracy rate of historical anomaly detection, and parameter distribution curves, to help determine the rationality of the baseline.

[0102] Verification result record: Experts can make three judgments on the baseline: "passed", "needs fine adjustment" or "rejected and regenerated", and fill in the specific reasons. The system automatically records the verification process.

[0103] Baseline fine-tuning module: Provides flexible parameter adjustment tools for baselines that "need fine-tuning", including manual fine-tuning, rule template calling, batch fine-tuning, etc.

[0104] Manual fine-tuning: Experts can directly modify key parameters such as behavioral feature thresholds and anomaly detection weights, and the system displays the baseline effect simulation data after adjustment in real time.

[0105] Rule template calling: Built-in fine-tuning rule templates for common scenarios, which experts can directly call and adapt.

[0106] Batch fine-tuning: Supports batch parameter adjustment of multiple baselines of the same type, improving fine-tuning efficiency.

[0107] Baseline Secondary Review and Activation: The fine-tuned baseline requires a secondary review (multi-level review processes can be set). After passing the review, it enters the activation pending confirmation state. Baselines in the activation pending confirmation state can further verify their effectiveness through gray-scale testing. That is, the baseline is first activated in some business scenarios or devices to monitor the operational effects before deciding whether to fully activate it. After the baseline is activated, the system automatically records information such as activation time, activation scope, and activation version, forming a complete version traceability chain.

[0108] Configuration distribution: The enabled target baseline configuration is distributed to the abnormal behavior analysis engine, specifically including: The enabled baseline parameters are automatically converted into a configuration format (such as XML, JSON, etc.) that the abnormal behavior analysis engine can recognize, ensuring parameter compatibility.

[0109] Configurations are distributed to different types of abnormal behavior analysis engines (such as network intrusion detection engines and user operation audit engines), and configuration synchronization is achieved through engine interfaces.

[0110] After the configuration is distributed, a verification request is automatically sent to the analysis engine to confirm whether the configuration has taken effect. If it has not taken effect, a retry mechanism is triggered and the operations and maintenance personnel are notified.

[0111] Engine loading rules: After receiving the baseline monitoring rules, the abnormal behavior analysis engine loads the monitoring rules of the target baseline and integrates the monitoring rules into the corresponding risk model, laying the foundation for subsequent real-time detection.

[0112] Based on monitoring rules incorporating a risk model, the abnormal behavior analysis engine monitors data security incidents and potential incidents in real time. After filtering out events that meet the correlation scope of the monitoring rules, it extracts key information, specifically as follows: "Data security incident" can be understood as an event that has occurred and has a real impact on data security; "data security potential incident" can be understood as an event that has not yet caused a real impact but has potential risks. In some optional implementation methods, refer to... Figure 5 The process of extracting information is as follows: Extracting data security event type information: Operators collect data security event information, covering network-side and user-side data-related events, including user data leakage events (such as a business hall employee illegally downloading 1,000 user ID numbers and call details and distributing them externally, the event type is marked as "internal personnel illegally exporting data"), network configuration data tampering events (external attacks tampering with 5G base station parameters (such as power and frequency band configuration), causing local network signal abnormalities, the event type is defined as "external malicious attack tampering"), and data transmission abnormality events (unreported cross-border transmission of user data, involving sensitive communication records, the event type is marked as "illegal cross-border data transmission"), etc.; through log auditing systems (such as user data access logs and network configuration change logs) and security monitoring platforms, the type information of these events is automatically identified and extracted to form an event type list; then, based on the lifecycle stage corresponding to the event type, the data lifecycle information corresponding to the data security event is obtained, and based on the attribute information corresponding to the event type, the security attribute information corresponding to the data security event is obtained.

[0113] Extracting data security vulnerability incident type information: Identifying vulnerabilities at each stage of the operator's data lifecycle, including: "Unauthorized data collection" in the collection stage (e.g., third-party partners illegally collecting user location data); "Storage media security vulnerabilities" in the storage stage (e.g., user data storage servers not being encrypted, posing a risk of physical theft); "Data misuse / unmasked use" in the usage stage (e.g., business systems accessing user data without anonymization (e.g., directly displaying complete ID numbers)); and "Data residue leakage" in the destruction stage (e.g., data not being thoroughly cleared from abandoned servers, leaving residual user communication records). Using regular security inspections and vulnerability scanning tools, identifying vulnerabilities at each stage, extracting risk types, and entering them into a risk information database; then, based on the lifecycle stage corresponding to the vulnerability type, obtaining data lifecycle information corresponding to the data security vulnerability; and based on the attribute information corresponding to the vulnerability type, obtaining security attribute information corresponding to the data security vulnerability.

[0114] Furthermore, a data risk matrix can be constructed based on the extracted lifecycle stage information and security attribute information to achieve a structured correlation between events and risks. Specifically, the "data risk matrix" can be understood as a structured analysis tool that integrates data security events and potential risks using data lifecycle stages and security attributes as dimensions. In some optional implementations, the construction process is as follows: Matrix dimension definition: Each stage of the lifecycle is used as a column of the data risk analysis matrix, and each element of the security attribute is used as a row of the data risk analysis matrix.

[0115] Matrix data filling: Based on the data lifecycle information and security attribute information corresponding to the data security incident information and data security risk information for each statistical period, the data security incident information and data security risk information are filled into the corresponding positions of the data risk analysis matrix.

[0116] Matrix value definition: The matrix value is the correlation (or frequency of occurrence, degree of impact) between the corresponding event and the risk. The correlation is determined by historical event statistics and expert evaluation (such as security team scoring). The higher the value, the closer the correlation between the event type and the risk type and the greater the impact.

[0117] For example, a two-dimensional matrix can be constructed with "Event Type" as the rows and "Risk Type" as the columns. The matrix values ​​represent the correlation (or frequency of occurrence, degree of impact) between the corresponding event and risk. For example: The correlation is determined through historical event statistics and expert evaluation (such as security team scoring). The higher the value, the closer the correlation between the event type and the risk type, and the greater the impact.

[0118] Based on the constructed data risk matrix, security risk indicators for each period are calculated, anomalies are identified through cross-period comparisons, and alarms are ultimately generated, as follows: Obtain security risk indicators for each period: Based on the data risk analysis matrix for each statistical period, obtain the hazard name, hazard type, hazard level, and number of hazards corresponding to each life cycle stage and each security attribute, as well as the security event name, security event type, security event level, and number of security events; based on this information, calculate the data security risk indicators for each statistical period.

[0119] Identify and generate alerts for abnormal risk indicators: Identify and generate alerts for abnormal risk indicators that exceed the baseline threshold from the data security risk indicators for each period.

[0120] For example, the cross-cycle comparison and alarm process is as follows: Compare the risk matrices of different stages of the data lifecycle (e.g., the beginning and end of the quarter), and compare the data collection stage: if the correlation between "unauthorized data collection" and "internal personnel's unauthorized data export" events increases from 0.2 at the beginning of the quarter to 0.6 at the end of the quarter, and the number of times third-party partners illegally collect user data increases by 3 times in the corresponding actual events, then "unauthorized data collection - internal unauthorized export" is determined to be an abnormal indicator, and it is necessary to strengthen the control of collection channels (e.g., auditing partner interfaces, restricting the scope of collection) and triggering an alarm; Usage phase comparison: The correlation between "data misuse / unmasked use" and the "external malicious attack tampering" event dropped sharply from 0.6 to 0.1. However, in actual business, complaints about unmasked user data display in business systems increased. Combined with matrix anomalies, it was found that an attack caused the masking function to fail. "Data misuse / unmasked use - attack impact" was used as an anomaly indicator, triggering system repair and security hardening and generating an alarm. The abnormal behavior analysis engine monitors operator business logs in real time based on baselines: when it detects that "employee B queries 30 user call details in a single day (exceeding the baseline limit of 20)," it is judged as abnormal and an alarm is triggered; when it finds that "maintenance personnel C has a work order without faults, but modifies base station frequency band parameters without review," it immediately issues an alarm and pushes it to the security operations center for handling by designated personnel to block unauthorized operations and trace data risks.

[0121] In the abnormal behavior detection method of this application, the effectiveness of the baseline is ensured by an expert system and the abnormal behavior analysis engine performs full-process detection to realize the abnormal detection of data access and operation behavior. It can be adapted to scenarios such as main business data security and government and enterprise customer data protection, dynamically optimize the baseline and detection logic, improve the accuracy of data security governance, and protect the security of operator data assets.

[0122] Please see Figure 7 , Figure 7This is a schematic diagram of an abnormal behavior detection device 700 provided in an embodiment of this application. As shown in the figure, the abnormal behavior detection device 700 includes: The construction module 710 is used to construct an abnormal behavior feature dataset based on behavior templates and sampled data, and to obtain a data lineage training set and user behavior attention model data based on the sampled data; The acquisition module 720 is used to process the data lineage training set and user behavior attention model data that meet the preset conditions based on the abnormal behavior feature dataset, so as to construct an abnormal behavior information training set. Processing module 730 is used to perform feature processing on the abnormal behavior information training set and baseline template knowledge base to obtain target feature data, input the target feature data into a large language model to generate an initial abnormal behavior baseline, and send the initial abnormal behavior baseline to the data security intelligent agent for storage. The generation module 740 is used to obtain the target baseline verified by the data security intelligent agent and to perform abnormal behavior detection based on the target baseline.

[0123] Optionally, the building module 710 may be specifically used for; A knowledge base is built based on behavior templates to store normal behavioral characteristics and judgment rules; Behavioral features are extracted from the sampled data using both static and dynamic feature extraction methods. Based on the knowledge base, abnormal behavior is identified from the behavioral features, and an abnormal behavior feature dataset is constructed based on the results of the abnormal behavior identification.

[0124] Optionally, the dynamic feature extraction method includes time-domain feature extraction, frequency-domain feature extraction, and deep learning extraction; The time-domain features include at least one of the following: time window features, sequence pattern features, trend and mutation features, and periodic features; The frequency domain features are extracted based on Fourier transform; The deep learning extraction includes at least one of cluster analysis, temporal pattern mining, sequence feature extraction, and association rule mining.

[0125] Optionally, the acquisition module 720 may be specifically used for: Construct a data matrix, an access relationship matrix, and a data access relationship view based on the sampled data; The matrix elements of the data access relationship view are determined to be the product of the access relationship matrix elements and the data matrix elements; Based on the data access relationship view, the sampled data is integrated to construct a data lineage training set; The data matrix includes the relationship between the source and data type of the sampled data, and the access relationship matrix represents the user access behavior of the dataset involved in the sampled data. The dimensions of the data access relationship view include the user set, the sampled data source set, and the sampled data type set.

[0126] Optionally, the acquisition module 720 may be specifically used for: Obtain the historical access dataset from the sampled data, the historical access dataset containing access record feature vectors and abnormal behavior annotation labels; The information gain method is used to obtain the information gain value of each feature vector in the historical access dataset. Feature vectors with information gain values ​​greater than a preset threshold are selected to form a target feature set. Based on the target feature set, the Hofding tree model is used as the initial anomaly recognition model, and the random forest algorithm is used for ensemble learning to obtain the anomaly recognition ensemble model. The model based on the anomaly detection model generates initial user behavior model data based on the inference results of the historical access dataset. The initial user behavior model data is trained using expectation maximization and graph network augmentation to obtain user behavior attention model data.

[0127] Optionally, the acquisition module 720 may be specifically used for: Construct a heterogeneous graph, wherein the set of nodes in the heterogeneous graph includes users, pages, and operation types, and the set of edges in the heterogeneous graph represents the behavioral relationships between the nodes; The heterogeneous graph is aggregated based on the message passing mechanism of the graph neural network, and edge weights and node features are fused to output a node feature representation vector. The node feature representation vector is processed by multi-level graph attention to obtain behavior association weight data and optimized node feature representation vector; The user behavior attention model data is formed by integrating the model parameter data after expected value maximization training, the sampled weight data, the node feature representation vector, and the behavior-related weight data.

[0128] Optionally, the processing module 730 may specifically be used for: The abnormal behavior information training set and the baseline template knowledge base are fused to obtain a fused feature vector, and the fused feature vector is then dimensionality-reduced to obtain dimensionality-reduced features. Combining the aforementioned dimensionality reduction features, typical abnormal samples, and normal behavior rules, a structured prompt word containing business scenarios, feature dimensions, and anomaly detection requirements is constructed; A time-series dataset is constructed using dimensionality reduction features as input and behavior type as label. The time-series dataset is then input into a long short-term memory network to obtain the time-series behavior features output by the long short-term memory network. The temporal behavior features and the structured cue words are input into a large language model to generate an initial abnormal behavior baseline.

[0129] Optionally, the generation module 740 may be specifically used to: send the target baseline configuration to the abnormal behavior analysis engine for abnormal behavior detection; The abnormal behavior detection device 700 can also be used for: The monitoring rules of the target baseline are loaded and distributed through the abnormal behavior analysis engine, and then integrated into the corresponding risk model. Based on the monitoring rules, the data security events and potential events monitored in real time by the filtering engine are filtered, the type information of the data security events and potential events is extracted, and then the life cycle stage information and security attribute information corresponding to the type information are obtained. A data risk matrix is ​​constructed based on the aforementioned lifecycle stage information and security attribute information; Based on the aforementioned data risk matrix, security risk indicators for each period are obtained. Identify and generate alerts for abnormal risk indicators that exceed the baseline threshold from the security risk indicators of each period.

[0130] The abnormal behavior detection device in this application embodiment can be an electronic device, such as an electronic device with an operating system, or a component in an electronic device, such as an integrated circuit or a chip.

[0131] The abnormal behavior detection device provided in this application embodiment can achieve... Figure 1 The various processes implemented in the method embodiments shown achieve the same technical effects, and will not be described again here to avoid repetition.

[0132] For details, see Figure 8 As shown in the figure, this application embodiment also provides an electronic device, including a bus 801, a transceiver 802, an antenna 803, a bus interface 804, a processor 805, and a memory 806.

[0133] Processor 805, used for: An abnormal behavior feature dataset is constructed based on behavior templates and sampled data, and data lineage training set and user behavior attention model data are obtained based on the sampled data. Based on the abnormal behavior feature dataset, the data lineage training set and user behavior attention model data that meet the preset conditions are processed to construct an abnormal behavior information training set. The abnormal behavior information training set and baseline template knowledge base are subjected to feature processing to obtain target feature data. The target feature data is input into a large language model to generate an initial abnormal behavior baseline. The initial abnormal behavior baseline is sent to the data security intelligent agent for storage. Obtain the target baseline verified by the data security intelligent agent, and perform abnormal behavior detection based on the target baseline.

[0134] exist Figure 8 In this document, a bus architecture (represented by bus 801) is used. Bus 801 can include any number of interconnected buses and bridges, linking various circuits including one or more processors represented by processor 805 and memory represented by memory 806. Bus 801 can also link various other circuits such as peripheral devices, voltage regulators, and power management circuits, which are well known in the art and therefore will not be described further herein. Bus interface 804 provides an interface between bus 801 and transceiver 802. Transceiver 802 can be a single element or multiple elements, such as multiple receivers and transmitters, providing a unit for communicating with various other devices over a transmission medium. Data processed by processor 805 is transmitted over a wireless medium via antenna 803, which further receives data and transmits data to processor 805.

[0135] The processor 805 manages the bus 801 and handles general processing, and also provides various functions, including timing, peripheral interfaces, voltage regulation, power management, and other control functions. The memory 806 can be used to store data used by the processor 805 during operation.

[0136] Alternatively, the processor 805 may be a central processing unit (CPU), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or a complex programmable logic device (CPLD).

[0137] Optionally, the processor 805 is specifically used for: A knowledge base is built based on behavior templates to store normal behavioral characteristics and judgment rules; Behavioral features are extracted from the sampled data using both static and dynamic feature extraction methods. Based on the knowledge base, abnormal behavior is identified from the behavioral features, and an abnormal behavior feature dataset is constructed based on the results of the abnormal behavior identification.

[0138] Optionally, the dynamic feature extraction method includes time-domain feature extraction, frequency-domain feature extraction, and deep learning extraction; The time-domain features include at least one of the following: time window features, sequence pattern features, trend and mutation features, and periodic features; The frequency domain features are extracted based on Fourier transform; The deep learning extraction includes at least one of cluster analysis, temporal pattern mining, sequence feature extraction, and association rule mining.

[0139] Optionally, the processor 805 is specifically used for: Construct a data matrix, an access relationship matrix, and a data access relationship view based on the sampled data; The matrix elements of the data access relationship view are determined to be the product of the access relationship matrix elements and the data matrix elements; Based on the data access relationship view, the sampled data is integrated to construct a data lineage training set; The data matrix includes the relationship between the source and data type of the sampled data, and the access relationship matrix represents the user access behavior of the dataset involved in the sampled data. The dimensions of the data access relationship view include the user set, the sampled data source set, and the sampled data type set.

[0140] Optionally, the processor 805 is specifically used for: Obtain the historical access dataset from the sampled data, the historical access dataset containing access record feature vectors and abnormal behavior annotation labels; The information gain method is used to obtain the information gain value of each feature vector in the historical access dataset. Feature vectors with information gain values ​​greater than a preset threshold are selected to form a target feature set. Based on the target feature set, the Hofding tree model is used as the initial anomaly recognition model, and the random forest algorithm is used for ensemble learning to obtain the anomaly recognition ensemble model. The model based on the anomaly detection model generates initial user behavior model data based on the inference results of the historical access dataset. The initial user behavior model data is trained using expectation maximization and graph network augmentation to obtain user behavior attention model data.

[0141] Optionally, the processor 805 is specifically used for: Construct a heterogeneous graph, wherein the set of nodes in the heterogeneous graph includes users, pages, and operation types, and the set of edges in the heterogeneous graph represents the behavioral relationships between the nodes; The heterogeneous graph is aggregated based on the message passing mechanism of the graph neural network, and edge weights and node features are fused to output a node feature representation vector. The node feature representation vector is processed by multi-level graph attention to obtain behavior association weight data and optimized node feature representation vector; The user behavior attention model data is formed by integrating the model parameter data after expected value maximization training, the sampled weight data, the node feature representation vector, and the behavior-related weight data.

[0142] Optionally, the processor 805 is specifically used for: The abnormal behavior information training set and the baseline template knowledge base are fused to obtain a fused feature vector, and the fused feature vector is then dimensionality-reduced to obtain dimensionality-reduced features. Combining the aforementioned dimensionality reduction features, typical abnormal samples, and normal behavior rules, a structured prompt word containing business scenarios, feature dimensions, and anomaly detection requirements is constructed; A time-series dataset is constructed using dimensionality reduction features as input and behavior type as label. The time-series dataset is then input into a long short-term memory network to obtain the time-series behavior features output by the long short-term memory network. Optionally, the processor 805 is specifically used to: send the target baseline configuration to the abnormal behavior analysis engine for abnormal behavior detection; The processor 805 can also be used for: The monitoring rules of the target baseline are loaded and distributed through the abnormal behavior analysis engine, and then integrated into the corresponding risk model. Based on the monitoring rules, the data security events and potential events monitored in real time by the filtering engine are filtered, the type information of the data security events and potential events is extracted, and then the life cycle stage information and security attribute information corresponding to the type information are obtained. A data risk matrix is ​​constructed based on the aforementioned lifecycle stage information and security attribute information; Based on the aforementioned data risk matrix, security risk indicators for each period are obtained. Identify and generate alerts for abnormal risk indicators that exceed the baseline threshold from the security risk indicators of each period.

[0143] It should be noted that the electronic device provided in this application embodiment is a device capable of executing the above-described abnormal behavior detection method. Therefore, all implementation methods in the above-described abnormal behavior detection method embodiments are applicable to this electronic device and can achieve the same or similar beneficial effects. To avoid repetition, this embodiment will not elaborate further.

[0144] This invention also provides an electronic device, including: a processor, a memory, and a program stored in the memory and executable on the processor. When the program is executed by the processor, it implements the various processes of the above-described abnormal behavior detection method embodiments and achieves the same technical effect. To avoid repetition, it will not be described again here.

[0145] This application also provides a computer-readable storage medium storing a computer program. When executed by a processor, this computer program implements the various processes of the above-described abnormal behavior detection method embodiments and achieves the same technical effects. To avoid repetition, it will not be described again here. The computer-readable storage medium may be a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk, etc.

[0146] This application also provides a computer program product, including computer instructions. When executed by a processor, the computer instructions implement the various processes of the above-described abnormal behavior detection method embodiments and achieve the same technical effect. To avoid repetition, they will not be described again here.

[0147] It should be noted that, in this document, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Unless otherwise specified, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes that element.

[0148] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods of the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions to cause a terminal (which may be a mobile phone, computer, server, air conditioner, or network device, etc.) to execute the methods described in the various embodiments of this application.

[0149] The embodiments of this application have been described above with reference to the accompanying drawings. However, this application is not limited to the specific embodiments described above. The specific embodiments described above are merely illustrative and not restrictive. Those skilled in the art can make many other forms under the guidance of this application without departing from the spirit and scope of the claims, and all of these forms are within the protection scope of this application.

Claims

1. A method for detecting abnormal behavior, characterized in that, The method includes: An abnormal behavior feature dataset is constructed based on behavior templates and sampled data, and data lineage training set and user behavior attention model data are obtained based on the sampled data. Based on the abnormal behavior feature dataset, the data lineage training set and user behavior attention model data that meet the preset conditions are processed to construct an abnormal behavior information training set. The abnormal behavior information training set and baseline template knowledge base are subjected to feature processing to obtain target feature data. The target feature data is input into a large language model to generate an initial abnormal behavior baseline. The initial abnormal behavior baseline is sent to the data security intelligent agent for storage. Obtain the target baseline verified by the data security intelligent agent, and perform abnormal behavior detection based on the target baseline.

2. The method according to claim 1, characterized in that, The abnormal behavior feature dataset constructed based on behavior templates and sampled data includes: A knowledge base is built based on behavior templates to store normal behavioral characteristics and judgment rules; Behavioral features are extracted from the sampled data using both static and dynamic feature extraction methods. Based on the knowledge base, abnormal behavior is identified from the behavioral features, and an abnormal behavior feature dataset is constructed based on the results of the abnormal behavior identification.

3. The method according to claim 1, characterized in that, The dynamic feature extraction methods include time-domain feature extraction, frequency-domain feature extraction, and deep learning extraction. The time-domain features include at least one of the following: time window features, sequence pattern features, trend and mutation features, and periodic features; The frequency domain features are extracted based on Fourier transform; The deep learning extraction includes at least one of cluster analysis, temporal pattern mining, sequence feature extraction, and association rule mining.

4. The method according to any one of claims 1 to 3, characterized in that, The data lineage training set obtained based on the sampled data includes: Construct a data matrix, an access relationship matrix, and a data access relationship view based on the sampled data; The matrix elements of the data access relationship view are determined to be the product of the access relationship matrix elements and the data matrix elements; Based on the data access relationship view, the sampled data is integrated to construct a data lineage training set; The data matrix includes the relationship between the source and data type of the sampled data, and the access relationship matrix represents the user access behavior of the dataset involved in the sampled data. The dimensions of the data access relationship view include the user set, the sampled data source set, and the sampled data type set.

5. The method according to any one of claims 1 to 3, characterized in that, Obtaining the user behavior attention model data includes: Obtain the historical access dataset from the sampled data, the historical access dataset containing access record feature vectors and abnormal behavior annotation labels; The information gain method is used to obtain the information gain value of each feature vector in the historical access dataset. Feature vectors with information gain values ​​greater than a preset threshold are selected to form a target feature set. Based on the target feature set, the Hofding tree model is used as the initial anomaly recognition model, and the random forest algorithm is used for ensemble learning to obtain the anomaly recognition ensemble model. The model based on the anomaly detection model generates initial user behavior model data based on the inference results of the historical access dataset. The initial user behavior model data is trained using expectation maximization and graph network augmentation to obtain user behavior attention model data.

6. The method according to claim 5, characterized in that, The step of performing expectation-maximization training and graph network augmentation on the initial user behavior model data to obtain user behavior attention model data includes: Construct a heterogeneous graph, wherein the set of nodes in the heterogeneous graph includes users, pages, and operation types, and the set of edges in the heterogeneous graph represents the behavioral relationships between the nodes; The heterogeneous graph is aggregated based on the message passing mechanism of the graph neural network, and edge weights and node features are fused to output a node feature representation vector. The node feature representation vector is processed by multi-level graph attention to obtain behavior association weight data and optimized node feature representation vector; The user behavior attention model data is formed by integrating the model parameter data after expected value maximization training, the sampled weight data, the node feature representation vector, and the behavior-related weight data.

7. The method according to any one of claims 1 to 3, characterized in that, The process of processing the abnormal behavior information training set and baseline template knowledge base, and then inputting the processed feature data into the large language model to generate the initial abnormal behavior baseline includes: The abnormal behavior information training set and the baseline template knowledge base are fused to obtain a fused feature vector, and the fused feature vector is then dimensionality-reduced to obtain dimensionality-reduced features. Combining the aforementioned dimensionality reduction features, typical abnormal samples, and normal behavior rules, a structured prompt word containing business scenarios, feature dimensions, and anomaly detection requirements is constructed; A time-series dataset is constructed using dimensionality reduction features as input and behavior type as label. The time-series dataset is then input into a long short-term memory network to obtain the time-series behavior features output by the long short-term memory network. The temporal behavior features and the structured cue words are input into a large language model to generate an initial abnormal behavior baseline.

8. The method according to any one of claims 1 to 3, characterized in that, The abnormal behavior detection based on the target baseline includes: sending the target baseline configuration to the abnormal behavior analysis engine for abnormal behavior detection; The method further includes: The monitoring rules of the target baseline are loaded and distributed through the abnormal behavior analysis engine, and then integrated into the corresponding risk model. Based on the monitoring rules, the data security events and potential events monitored in real time by the filtering engine are filtered, the type information of the data security events and potential events is extracted, and then the life cycle stage information and security attribute information corresponding to the type information are obtained. A data risk matrix is ​​constructed based on the aforementioned lifecycle stage information and security attribute information; Based on the aforementioned data risk matrix, security risk indicators for each period are obtained. Identify and generate alerts for abnormal risk indicators that exceed the baseline threshold from the security risk indicators of each period.

9. An abnormal behavior detection device, characterized in that, include: The module is used to construct an abnormal behavior feature dataset based on behavior templates and sampled data, and to obtain a data lineage training set and user behavior attention model data based on the sampled data. The acquisition module is used to process the data lineage training set and user behavior attention model data that meet preset conditions based on the abnormal behavior feature dataset, so as to construct an abnormal behavior information training set. The processing module is used to perform feature processing on the abnormal behavior information training set and the baseline template knowledge base to obtain target feature data, input the target feature data into the large language model to generate an initial abnormal behavior baseline, and send the initial abnormal behavior baseline to the data security intelligent agent for storage. The generation module is used to obtain the target baseline verified by the data security intelligent agent and to perform abnormal behavior detection based on the target baseline.

10. An electronic device, characterized in that, include: A processor, a memory, and a program stored in the memory and executable on the processor, wherein the program, when executed by the processor, implements the steps of the abnormal behavior detection method as described in any one of claims 1 to 8.

11. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program that, when executed by a processor, implements the steps of the abnormal behavior detection method as described in any one of claims 1 to 8.

12. A computer program product, characterized in that, It includes computer instructions that, when executed by a processor, implement the steps of the abnormal behavior detection method as described in any one of claims 1 to 8.