Abnormal behavior identification method and system based on multi-dimensional data analysis
Through the multi-dimensional data analysis method, structured and unstructured data are processed, multi-dimensional feature sets and entity relationship diagrams are constructed, and the correlation intensity and abnormal scores are calculated. The problems of insufficient processing capabilities and correlation relationship analysis in the existing technology are solved, and a flexible abnormal scoring mechanism and highly accurate abnormal behavior recognition are realized.
Patent Information
- Application Number
- CN202510489869.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-18
- Publication Date
- 2025-05-16
- Estimated Expiration
- 2045-04-18
AI Technical Summary
The prior art lacks the ability to process unstructured data in abnormal behavior recognition, lacks in-depth analysis of correlation relationships between entities, and the model scoring mechanism is rigid, making it difficult to flexibly adjust weights according to different scenarios.
An abnormal behavior recognition method based on multi-dimensional data analysis is adopted, and data cleaning and standardization are carried out by receiving structured and unstructured data, and entity information and relational information in document texts are extracted using natural language processing technology to generate structured feature vectors. Then, a multidimensional feature set is constructed, an entity relationship diagram is established, and the correlation strength between entities is calculated through low-rank tensors and p-AAA algorithm, the entity relationship diagram is optimized, a flexible feature weight system is set, a multidimensional anomaly score is calculated and an early warning list is output.
The processing ability of unstructured data is improved, the correlation relationship between entities is deeply analyzed, and the exception scoring mechanism is optimized, so that it can flexibly adjust weights according to different scenarios, which improves the accuracy and interpretability of abnormal behavior recognition.
Smart Images

Figure CN120012004A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of data analysis technology, and in particular to an abnormal behavior identification method and system based on multidimensional data analysis. Background Art
[0002] Abnormal behavior recognition technology has important application value in the fields of financial risk control, security supervision, etc. This technology mainly analyzes various types of data to identify behavioral characteristics that deviate from normal patterns and discover potential risks in a timely manner.
[0003] At present, the common abnormal behavior identification methods are mainly based on a single data source or simple rules. For example, judgment is made by setting a fixed numerical threshold, or only analyzing structured transaction data for pattern recognition. These methods are relatively simple and direct, but the analysis dimension is single and it is difficult to meet the needs of abnormal behavior identification in complex scenarios.
[0004] With the development of technology, more advanced abnormal behavior recognition technology uses machine learning methods to detect anomalies by establishing behavioral models. This technology first collects historical data to establish a baseline model, then compares the new behavior data with the model, and calculates the deviation to determine whether it is abnormal.
[0005] However, this technology still has the following problems: first, the processing capabilities of unstructured data are insufficient, making it difficult to fully utilize information such as text; second, there is a lack of in-depth analysis of the relationships between entities, and it is easy to miss complex abnormal patterns based on relationship networks; finally, the model scoring mechanism is relatively rigid, making it difficult to flexibly adjust weights according to different scenarios. Summary of the invention
[0006] In view of this, the present application provides an abnormal behavior identification method and system based on multidimensional data analysis, which solves the problems of insufficient unstructured data processing capabilities, lack of entity association relationship analysis and rigid scoring mechanism in the prior art.
[0007] The present application embodiment provides a method for identifying abnormal behavior based on multidimensional data analysis, including: Acquire structured data and unstructured data through a business system interface, perform data cleaning and standardization on the structured data and the unstructured data, and obtain a data set in a unified format; Receiving the document text in the unified format data set, processing the document text using natural language processing technology, extracting entity information, relationship information and behavior information, and generating a structured feature vector; Using the structured data in the unified format data set and the structured feature vector, constructing behavioral statistical features, temporal features and correlation features, using a sampling algorithm based on the generalized Golub-Kahan method to select feature combinations, and constructing a multidimensional feature set; According to the associated data in the multidimensional feature set, an entity relationship graph is established, the association strength between entities is calculated by low-rank tensor and p-AAA algorithm, and the entity relationship graph is optimized based on the association strength to establish a network structure model; According to the network structure model and the multidimensional feature set, a feature weight system is set, and a multidimensional anomaly score is calculated using a pruned tensor structure measurement method. The multidimensional anomaly score is comprehensively calculated using a low-rank tensor recovery technique to obtain an anomaly score, and a warning list is output based on the anomaly score.
[0008] Optionally, the structured data and the unstructured data are cleaned and standardized to obtain a data set in a unified format, including: For the structured data and the unstructured data, a statistical method is used to identify abnormal values of numerical data and make corrections, the coding of categorical data is standardized, and the format of time data is unified to obtain cleaned data; According to the cleaned data, the field names, data types, and value ranges are unified to obtain a data set in a unified format.
[0009] For the document text, a conditional random field model is used to perform word segmentation processing to obtain a word sequence; For the term sequence, named entity recognition technology is used to identify key entities, and relationship description words between entities are extracted to construct entity-relationship pairs; According to the entity-relationship pairs, dependency syntactic analysis is used to extract target grammatical components, identify behavior types and behavior features, and output structured feature vectors.
[0010] Optionally, the method of constructing the behavior statistical features by using the structured data in the data set in the unified format and the structured feature vector includes: Based on the structured data in the unified format data set and the structured feature vector, the frequency and time distribution of the behavior are counted, and the target statistics of the behavior frequency are calculated to form a frequency feature; According to the frequency characteristics, a sliding time window is set, and the frequency changes at different time scales are calculated to obtain the behavior statistical characteristics.
[0011] Optionally, the method of constructing time series features using the structured data in the data set in the unified format and the structured feature vector includes: Based on the structured data in the unified format data set and the structured feature vector, the periodic pattern of the behavior is detected by Fourier transform, the autocorrelation coefficient is calculated, and the periodicity index is constructed; For the periodic indicators, the moving average method is used to analyze the long-term trend, calculate the volatility and amplitude, and output the time series characteristics.
[0012] Optionally, the method of constructing the associated features using the structured data in the data set in the unified format and the structured feature vector includes: Based on the structured data in the unified format data set and the structured feature vector, the frequency and intensity of direct interactions between entities are counted, and the weighted association coefficient is calculated to obtain the direct association degree; Based on the direct correlation degree, a multi-hop relationship path is constructed, and the path importance weight is calculated to form a correlation feature.
[0013] Optionally, the association strength between entities is calculated using a low-rank tensor and a p-AAA algorithm, including: Using the correlation data in the multidimensional feature set, constructing a multidimensional correlation tensor, using Tucker decomposition to reduce the tensor dimension, and obtaining main feature information; According to the main feature information, a rational function approximation is constructed through the p-AAA algorithm, the approximation accuracy is iteratively optimized, and the association strength between entities is obtained.
[0014] Optionally, the optimizing the entity relationship diagram based on the association strength includes: According to the association strength, an association strength threshold is set, edges below the association strength threshold are deleted, and an optimized entity relationship graph is constructed; The optimized entity relationship diagram is used to identify nodes whose similarity exceeds a preset threshold, merge the nodes, update the association relationship, and form a network structure model.
[0015] Optionally, setting a feature weight system based on the network structure model and the multi-dimensional feature set includes: According to the network structure model and the multi-dimensional feature set, the weight value of the closeness of the associated party, the weight value of the abnormal behavior and the weight value of the abnormal timing are set to construct a weight configuration; Based on the weight configuration, the associated party closeness weight value, the behavior abnormality weight value and the time series abnormality weight value are refined to obtain a feature weight system.
[0016] The embodiment of the present application also provides an abnormal behavior identification device based on multi-dimensional data analysis, including: A data preprocessing module is used to obtain structured data and unstructured data through a business system interface, and to perform data cleaning and standardization on the structured data and the unstructured data to obtain a data set in a unified format; A text processing module, used to receive the document text in the unified format data set, process the document text using natural language processing technology, extract entity information, relationship information and behavior information, and generate a structured feature vector; A feature engineering module, for constructing behavioral statistical features, temporal features and correlation features by using the structured data in the unified format data set and the structured feature vector, and selecting feature combinations by using a sampling algorithm based on the generalized Golub-Kahan method to construct a multidimensional feature set; A network construction module, used to establish an entity relationship graph according to the associated data in the multidimensional feature set, calculate the association strength between entities through low-rank tensor and p-AAA algorithm, and optimize the entity relationship graph based on the association strength to establish a network structure model; The scoring calculation module is used to set a feature weight system according to the network structure model and the multidimensional feature set, calculate the multidimensional anomaly score using the pruned tensor structure measurement method, use the low-rank tensor recovery technology to comprehensively calculate the multidimensional anomaly score to obtain the anomaly score, and output a warning list according to the anomaly score.
[0017] This application has the following technical effects: The present application obtains structured data and unstructured data through a business system interface, performs data cleaning and standardization processing on the structured data and the unstructured data, and obtains a data set in a unified format; receives a document text in the data set in the unified format, processes the document text using natural language processing technology, extracts entity information, relationship information and behavior information, and generates a structured feature vector; uses the structured data in the data set in the unified format and the structured feature vector to construct behavioral statistical features, time series features and association features, uses a sampling algorithm based on the generalized Golub-Kahan method to select feature combinations, and constructs a multidimensional feature set; establishes an entity relationship graph based on the associated data in the multidimensional feature set, calculates the association strength between entities through a low-rank tensor and a p-AAA algorithm, optimizes the entity relationship graph based on the association strength, and establishes a network structure model; sets a feature weight system based on the network structure model and the multidimensional feature set, calculates a multidimensional anomaly score using a pruned tensor structure measurement method, uses a low-rank tensor recovery technology to comprehensively calculate the multidimensional anomaly score, obtains an anomaly score, and outputs an early warning list based on the anomaly score.
[0018] Among them, the generalized Golub-Kahan method is used for feature sampling to improve the processing efficiency of large-scale hierarchical Bayesian inverse problems; the low-rank tensor and p-AAA algorithm are combined to realize multivariate rational approximation, which improves the construction accuracy of entity relationship networks; the pruned tensor structure measurement and efficient low-rank tensor recovery technology are introduced to optimize the scoring calculation process of abnormal behavior; and, the innovative combination of text parsing and relational network analysis realizes the unified processing of structured and unstructured data; a flexible multi-dimensional weight scoring mechanism is designed to improve the accuracy and interpretability of abnormal behavior identification. BRIEF DESCRIPTION OF THE DRAWINGS
[0019] In order to more clearly illustrate the technical solutions of the embodiments of the present disclosure, the following is a brief introduction to the drawings required for use in the embodiments. The drawings herein are incorporated into the specification and constitute a part of the specification. These drawings illustrate embodiments consistent with the present disclosure and are used together with the specification to illustrate the technical solutions of the present disclosure. It should be understood that the following drawings only illustrate certain embodiments of the present disclosure and should not be regarded as limiting the scope. For ordinary technicians in this field, other relevant drawings can also be obtained based on these drawings without creative work.
[0020] Figure 1 A flowchart of an abnormal behavior identification method provided in an embodiment of the present application; Figure 2 Schematic diagram of the processing flow of the text processing module in the embodiment of the present application; Figure 3 This is a schematic diagram of the processing flow of the feature engineering module in an embodiment of the present application; Figure 4 A schematic diagram of the processing flow of the network construction module in an embodiment of the present application; Figure 5 Schematic diagram of the processing flow of the score calculation module in the embodiment of the present application.
[0021] Figure 6 It is a structural diagram of an abnormal behavior identification system based on multidimensional data analysis in an application embodiment. DETAILED DESCRIPTION
[0022] In order to make the purpose, technical scheme and advantages of the embodiments of the present disclosure clearer, the technical scheme in the embodiments of the present disclosure will be clearly and completely described below in conjunction with the drawings in the embodiments of the present disclosure. Obviously, the described embodiments are only part of the embodiments of the present disclosure, rather than all of the embodiments. The components of the embodiments of the present disclosure generally described and shown in the drawings here can be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of the present disclosure provided in the drawings is not intended to limit the scope of the present disclosure for protection, but merely represents the selected embodiments of the present disclosure. Based on the embodiments of the present disclosure, all other embodiments obtained by those skilled in the art without making creative work belong to the scope of protection of the present disclosure.
[0023] Figure 1 The following is a flow chart of the abnormal behavior identification method provided in the embodiment of the present application. Figure 1 As shown, the abnormal behavior identification method based on multidimensional data analysis provided by the present application may include the following steps S1 to S5: S1: Acquire structured data and unstructured data through a business system interface, perform data cleaning and standardization on the structured data and the unstructured data, and obtain a data set in a unified format; In the embodiments of the present application, the data obtained through the business system interface can be divided into two categories: structured data and unstructured data. Among them, structured data mainly includes numerical data (such as transaction amount, frequency statistics), category data (such as business type, risk level) and time data (such as transaction time, registration date), etc. Unstructured data mainly includes various types of document texts, such as business reports, transaction instructions, risk assessment reports, audit records, customer feedback, business communication records, contract agreement texts, regulatory filing documents, etc. These document texts contain rich entity information, relationship information and behavior information, and are an important data source for abnormal behavior identification.
[0024] During the data cleaning and standardization stage, the system uses a special processing flow for unstructured data such as document texts. First, text preprocessing is performed, including special character filtering, punctuation normalization, typo correction, and redundant information removal. Then, the system unifies the text format and converts documents of various formats (such as DOC, PDF, TXT, etc.) into a unified text format (such as the standard text format encoded in UTF-8) to ensure the consistency of subsequent processing. The system also performs preliminary structural tagging on the text, identifies key elements such as time expressions, monetary amounts, and institution names in the text, and converts them into standardized representations. For example, "December 15, 2023" is standardized to "2023-12-15", etc.
[0025] Through the above processing, the system integrates the document text into a unified format data set, retaining the semantic integrity of the original text and improving the processability of the data. For example, for a transaction description text "Company A conducted transactions with several affiliated companies under Group B in the fourth quarter of 2023", after standardized processing, the system not only retains the complete original text for deep semantic analysis, but also extracts preliminary structured information such as the transaction subject "Company A", the counterparty "affiliated companies under Group B", and the transaction time "fourth quarter of 2023". This processing method provides high-quality text input for the natural language processing module in the subsequent steps, enabling the system to more effectively identify abnormal behavior patterns from document texts.
[0026] In one embodiment, step S1 may specifically include steps S1.1 and S1.2: S1.1: For the structured data and the unstructured data, a statistical method is used to identify abnormal values of numerical data and make corrections, the coding of categorical data is standardized, and the format of time data is unified to obtain cleaned data; Specifically, in step S1.1, the system first performs targeted processing on different types of data: For numerical data, an outlier identification scheme based on statistical methods is adopted: first, the quartiles (Q1, Q2, Q3) and interquartile range (IQR) of the data are calculated, and the values outside the range of [Q1-1.5IQR, Q3+1.5IQR] are marked as potential outliers. Then, according to business rules and historical data distribution, these outliers are corrected, and methods such as median filling, linear interpolation or nearest neighbor interpolation can be selected.
[0027] For categorical data, the system uses a standardized coding scheme for processing. First, a unified category value dictionary is established to map category values that express the same meaning in different source systems to standard categories, such as uniformly mapping "male", "M", "1", etc. to "male". For categorical variables with an order relationship (such as risk level), ordinal coding is used; for unordered categorical variables, one-hot encoding is used to convert them into numerical form.
[0028] For time data, the system performs format unification. First, time strings in different formats (such as "2024-02-06", "2024 / 02 / 06", "20240206", etc.) are parsed into standard timestamp formats. Then, according to business needs, the timestamp is converted to a unified date and time format, and the time zone information is supplemented to ensure the consistency and comparability of time data. In addition, for missing time data, reasonable time interpolation or filling processing is performed according to business rules.
[0029] S1.2: According to the cleaned data, the field names, data types, and value ranges are unified to obtain a data set in a unified format; In step S1.2, the system further standardizes the cleaned data. First, according to the pre-defined data dictionary, the field names of each data source are unified, and the fields with the same meaning in different systems are mapped to the standard field names. For example, "user_id", "uid", "account_no", etc. are uniformly mapped to "user_identifier".
[0030] Secondly, standardize the data types. The system converts the same type of data from different data sources into a unified data type, such as unifying numeric data into double-precision floating point data (double), unifying character data into variable-length character data (varchar), etc. At the same time, set appropriate precision and length limits based on business needs.
[0031] Finally, the system normalizes the range of the data. For numerical features, methods such as minimum-maximum normalization or Z-score normalization can be used to map the data to a unified range. For categorical features, ensure that all possible values are within the predefined range, and handle values that are out of range appropriately. Through these processes, a standardized data set with a unified format and standardized structure is finally formed, laying the foundation for subsequent feature extraction and model building.
[0032] S2: receiving the document text in the unified format data set, processing the document text using natural language processing technology, extracting entity information, relationship information and behavior information, and generating a structured feature vector; like Figure 2 As shown, step S2 includes steps S2.1 to S2.3: S2.1: For the document text, a conditional random field model is used to perform word segmentation processing to obtain a word sequence; Specifically, step S2.1 uses an improved conditional random field (CRF) model to perform text segmentation. First, the system builds and maintains a special domain dictionary, including professional terms, institution names, and specific expressions in the fields of finance and risk control. On this basis, feature templates are designed, including character features (such as single words and double word combinations), position features (such as word beginning and word ending marks), and context features (such as context character combinations). Through these feature templates, the CRF model can learn the contextual dependencies of words, thereby improving the accuracy of word segmentation of professional texts. For example, for the text "an investment company conducts large-scale financial transactions with affiliated companies", the system can accurately identify professional terms such as "investment company", "affiliated companies", and "fund transactions". In addition, the system also introduces a dynamic programming optimization scheme based on the Viterbi algorithm to improve the efficiency of word segmentation.
[0033] S2.2: for the term sequence, using named entity recognition technology to identify key entities, extract relationship description words between entities, and construct entity-relationship pairs; In step S2.2, the system first uses the BiLSTM-CRF-based named entity recognition model to process the word sequence after word segmentation. This model captures contextual semantic information through a bidirectional LSTM network and combines the CRF layer for sequence labeling, which can accurately identify key entities in the text, such as institution names, names, time, and place. For example, it can label "a certain investment company" as an institutional entity and "the first quarter of 2024" as a time entity. Next, the system introduces an attention mechanism-enhanced relationship extraction module to analyze the dependencies between entities and extract key descriptors representing entity relationships. These descriptors usually include semantic information such as transaction categories (such as "transfer", "investment"), control categories (such as "holding", "holding"), etc. Finally, the system organizes the identified entity and relationship information into entity-relationship pairs to provide structured basic data for subsequent behavior analysis.
[0034] In step S2.2, the system first uses the BiLSTM-CRF-based named entity recognition model to process the word sequence after segmentation. Specifically, the training data sources of the model include two parts: one is the manually annotated financial field text data set (containing about 500,000 annotated sentences), and the other is the entity sample library collected and annotated from the business system (containing about 100,000 standard entities). In terms of model architecture, a two-layer BiLSTM structure is adopted, each layer contains 256 hidden units, and a Dropout layer (dropout rate is 0.5) is added in the middle to prevent overfitting. The input layer uses a 300-dimensional pre-trained word vector, and character features are extracted through character-level CNN (convolution kernel size is 3, 4, 5, 128 each). After BiLSTM, the CRF layer is connected for sequence labeling, and the loss function uses negative log-likelihood.
[0035] It should be noted that the Adam optimizer is used for model training, with an initial learning rate set to 0.001, a batch size of 64, and 50 training epochs. Early stopping is performed when the F1 value on the validation set has not improved for 5 consecutive epochs. To improve the generalization ability of the model, entity dictionary augmentation and adversarial training strategies are adopted during training. For example, "a certain investment company" can be labeled as an institutional entity (confidence 0.95), and "the first quarter of 2024" can be labeled as a time entity (confidence 0.98).
[0036] Next, the system introduces a relation extraction module based on the multi-head attention mechanism. This module contains 6 attention heads, each with a dimension of 64. The feed-forward network consists of two fully connected layers (with dimensions of 512 and 256 respectively). The scaled dot-product attention mechanism is used to calculate the attention scores, and the scaling factor is set to 8. The model uses relative position encoding to enhance the position perception ability, and the maximum distance of the position encoding is set to 100. The training data of this module contains approximately 300,000 pairs of manually annotated entity-relation pairs, covering 20 core relation types.
[0037] To improve the accuracy of relation extraction, the system also integrates rule constraints based on dependency syntax trees. For example, for a text like "Company A holds a controlling stake in Company B", through syntactic dependency relations, it can be determined that "holds a controlling stake" is the core predicate connecting the two company entities, thus extracting the triple structure <Company A, holds a controlling stake, Company B>. The extracted relations are normalized according to a predefined relation ontology. For example, similar expressions such as "holds a controlling stake", "owns shares", and "has a stake" are uniformly mapped to the standard relation type of "equity control".
[0038] Finally, the system organizes the identified entity and relation information into normalized entity-relation pairs. Each entity-relation pair contains fields such as the ID of the entity pair, entity type (with a confidence score), relation type (with a confidence score), relation direction, and timestamp. For example: { "entity1": {"id": "E001", "text": "Company A", "type": "Organization", "confidence": 0.95}, "entity2": {"id": "E002", "text": "Company B", "type": "Organization", "confidence": 0.93}, "relation": {"type": "equity control", "direction": "forward", "confidence": 0.89}, "timestamp": "2024-02-06 10:30:00" } This structured representation ensures the integrity and traceability of entity relationship information and provides standardized basic data for subsequent behavior analysis. The overall performance indicators of the model on the test set are: the entity recognition F1 value reaches 0.92, and the relationship extraction F1 value reaches 0.87.
[0039] S2.3: According to the entity-relationship pair, extract the target grammatical components by using dependency syntactic analysis, identify the behavior type and behavior features, and output a structured feature vector; In step S2.3, the system uses a dependency-based syntactic analysis method to deeply analyze the sentence structure. First, a syntactic analysis tree is constructed to identify the subject-predicate-object structure and various modifiers of the sentence. Then, based on predefined grammatical rules and semantic templates, the behavior-related target components are extracted from the syntactic tree. Specifically, the system focuses on verbal components (indicating the behavior type) and their related modifiers (indicating the behavior characteristics), while considering special components such as negative words and modal words that may affect the semantics of the behavior. For example, for the sentence "The company frequently conducts large-scale financial transactions with multiple related parties", the system can extract the behavior type as "fund transactions", and the behavior characteristics include "frequent" (frequency attribute) and "large amount" (scale attribute). Finally, the system converts the extracted semantic information into a structured feature vector according to the preset feature template, which contains information in multiple dimensions such as entity type, relationship type, behavior type, and behavior attribute. This structured representation method provides a standardized data foundation for subsequent feature engineering and anomaly identification.
[0040] S3: Using the structured data in the unified format data set and the structured feature vector, constructing behavioral statistical features, temporal features and correlation features, using a sampling algorithm based on the generalized Golub-Kahan method to select feature combinations, and constructing a multidimensional feature set; In step S3, the system innovatively uses a feature selection strategy based on the generalized Golub-Kahan method to construct a multidimensional feature set. First, the system performs a structured analysis on the original feature space and divides the features into three categories: behavioral statistical features, temporal features, and association features. For each type of feature, the system designs a special feature extraction method and quantitative indicators.
[0041] Specifically, for the behavior statistical characteristics, the system has built a multi-level indicator system. The basic layer includes behavior frequency statistics (such as daily transaction times, weekly transaction amounts, etc.), the middle layer includes behavior pattern characteristics (such as counterparty concentration, transaction time distribution, etc.), and the advanced layer includes behavior complexity indicators (such as behavior entropy, behavior diversity, etc.). For example, for a company's transfer behavior, the following characteristic values can be obtained: { "daily_transaction_count": 12.5, "weekly_amount_std": 0.85, "counterparty_concentration": 0.72, "behavior_entropy": 1.23 } In terms of time series features, the system uses multi-scale time windows for feature extraction. First, three basic time windows are set: short-term (3 days), medium-term (15 days), and long-term (90 days). The statistical features, trend features, and fluctuation features of the behavior sequence are calculated in each window. For example, for the transaction amount sequence, we can get: { "short_term_trend": 0.45, / / short-term trend slope "mid_term_volatility": 0.28, / / mid-term volatility "long_term_periodicity": 0.65 / / Long-term periodicity strength } The construction of association features is based on the entity relationship network, including direct association features (such as the number of first-degree connections, weights, etc.) and indirect association features (such as path diversity, group features, etc.). For example, the association features of a node can be expressed as: { "degree_centrality": 0.34, "clustering_coefficient": 0.56, "path_diversity": 0.78 } In the feature selection stage, the system innovatively applies the improved generalized Golub-Kahan algorithm for feature combination optimization. The core parameters of the algorithm are set as follows: projection matrix dimension: k = 100; number of iterations: max_iter = 1000; convergence threshold: tol = 1e-6; regularization parameter: lambda = 0.01; The algorithm execution process includes the following steps: constructing the feature matrix X (number of samples × number of features) and the label matrix Y; applying the bilateral Lanczos process to calculate the Krylov subspace basis; performing SVD decomposition to obtain the feature projection matrix; and calculating the feature importance score based on the projection result.
[0042] For example, for a dataset with 500 original features, the output of the algorithm might be as follows: { "selected_features": { "behavior_entropy": 0.92, / / Importance score "mid_term_volatility": 0.85, "path_diversity": 0.78, ... }, "feature_combinations": [ ["behavior_entropy", "path_diversity"], / / Optimal feature combination ["mid_term_volatility", "clustering_coefficient"], ... ] } Finally, the system selected about 150 core features to form a multidimensional feature set based on feature importance scores and combination effects. The performance of these features on the test data showed: feature redundancy: 0.15 (lower than 0.35 of the traditional method); feature coverage: 0.92 (higher than 0.78 of the baseline method); model prediction improvement: AUC value increased by 15% compared with the original features.
[0043] Through this feature engineering method, the system not only significantly reduces the dimension of the feature space, but also maintains the expressiveness and interpretability of the features. For example, in a certain financial risk control scenario, the feature set successfully captured more than 95% of the known abnormal patterns and discovered multiple previously undiscovered abnormal behavior types.
[0044] In one embodiment, Figure 3 As shown, step S3 includes steps S3.1 to S3.6: S3.1: Based on the structured data in the unified format data set and the structured feature vector, count the behavior frequency and time distribution, calculate the target statistics of the behavior frequency, and form a frequency feature; In step S3.1, the system constructs behavioral statistical features based on a unified format of data sets. First, basic statistics are performed on the behavior of each entity, including daily average frequency, cumulative number of times, maximum / minimum time intervals, etc. The system uses a hierarchical statistical method to calculate the statistics of behavioral frequency at different time granularities (such as hours, days, weeks, and months), including mean, standard deviation, kurtosis, skewness, etc. For different types of behaviors, the system will also consider their specific attributes. For example, in addition to the frequency of transaction behaviors, the distribution characteristics of transaction amounts will also be counted. These statistics can reflect the basic patterns and changing laws of entity behavior.
[0045] S3.2: According to the frequency characteristics, a sliding time window is set to calculate the frequency changes at different time scales to obtain the behavior statistical characteristics; In step S3.2, the system introduces a sliding time window mechanism to capture the dynamic changes of behavior patterns. Specifically, multiple time windows of different sizes are set (such as 7 days, 30 days, 90 days, etc.), and the changing trend of behavioral characteristics is calculated in each window. The system not only pays attention to the change in absolute frequency, but also calculates the rate of change relative to the historical baseline, as well as the difference compared with similar entities. By using the sliding window method, the sudden change or gradual change of the behavior pattern can be discovered in time, providing an important basis for anomaly identification.
[0046] S3.3: Based on the structured data in the unified format data set and the structured feature vector, detect the periodic pattern of the behavior through Fourier transform, calculate the autocorrelation coefficient, and construct a periodicity index; In step S3.3, the system uses Fourier transform technology to analyze the periodic characteristics of the behavior. First, the time series data is converted to the frequency domain space to identify the main frequency components to discover the potential periodic patterns. The system calculates the autocorrelation coefficient sequence to quantify the repeatability of the behavior pattern and analyzes the short-term, medium-term and long-term periodic characteristics by setting different time delays. This periodic analysis can help distinguish normal business cycles from abnormal behavior patterns.
[0047] In the behavior cycle pattern analysis, the system first preprocesses the time series data. For a given behavior sequence, the system uses interpolation and smoothing techniques to handle missing values and noise to ensure the continuity and reliability of the data. At the same time, the system normalizes the data to eliminate the impact of dimensions and make different types of behavior data comparable.
[0048] Next, the system applies a fast Fourier transform (FFT) to convert the time domain data into frequency domain space. In the frequency domain analysis, the system first calculates the power spectral density to identify significant frequency components. By setting an energy threshold, the main periodic components are screened out and these periods are sorted by energy size. The system pays special attention to those frequency components whose energy percentage exceeds a preset threshold (usually 10%), which often correspond to the main periodic patterns of behavior.
[0049] After obtaining preliminary periodic features, the system calculates the autocorrelation function (ACF) to verify and refine the periodic analysis results. Specifically, the system sets different time delays (lags) and calculates the correlation coefficient between the sequence and its delayed version. By analyzing the change pattern of the autocorrelation coefficient, the system can accurately identify the period length and period stability of the behavior sequence. For example, a significant peak in the autocorrelation coefficient at a certain delay time point often indicates the existence of a corresponding periodic pattern.
[0050] Based on the results of Fourier analysis and autocorrelation analysis, the system constructs a set of periodic indicators. These indicators include: the length and strength of the main cycle, the stability coefficient of the cycle, the proportion of harmonic components, the significance level of the periodicity, etc. The system also calculates the persistence index of the periodic pattern to evaluate the stability of the periodic behavior. For example, for a certain trading behavior sequence, the system may find that 7 days is the main cycle (strength 0.85), 30 days is the secondary cycle (strength 0.45), and calculates the period stability to be 0.72.
[0051] In order to improve the accuracy of periodic analysis, the system uses a sliding window mechanism to repeat periodic analysis at different time scales. This multi-scale analysis method can capture the changes in behavior patterns in different time periods. At the same time, the system will adjust analysis parameters such as window size and frequency resolution according to the characteristics of the business scenario to obtain the best detection effect.
[0052] Finally, the system integrates all periodic indicators into a standardized feature vector. This vector contains not only quantitative periodic features, but also qualitative descriptions of periodic patterns, such as the regularity of the period, the trend of periodic changes, etc. These periodic indicators provide important feature support for subsequent abnormal behavior identification, especially when discovering abnormal behaviors that deviate from normal periodic patterns.
[0053] S3.4: For the periodic indicators, use the moving average method to analyze the long-term trend, calculate the volatility and amplitude, and output the time series characteristics; In step S3.4, based on the calculated periodic indicators, the system uses an improved moving average method to analyze the long-term trend of the behavior. Specifically, it uses techniques such as exponential moving average (EMA) and weighted moving average (WMA) to give higher weights to recent data to better reflect changes in trends. At the same time, the system calculates volatility indicators (such as standard deviation, coefficient of variation, etc.) and amplitude characteristics (such as maximum fluctuation range, peak-to-valley ratio, etc.) to describe the stability and intensity of the behavior pattern.
[0054] After obtaining the periodic indicators, the system first applies a multi-level moving average method to analyze the long-term trend. Specifically, the system uses three methods: simple moving average (SMA), exponential moving average (EMA) and weighted moving average (WMA). For EMA, the system uses three levels of attenuation factors α=0.1, 0.2 and 0.3 to capture trend characteristics at different time scales. For WMA, the system designs a weight function based on time distance, so that recent data obtains a higher weight, while the influence of long-term data gradually decreases. This combination of multiple moving averages can more comprehensively characterize the changing characteristics of trends.
[0055] Based on trend analysis, the system calculates a series of volatility indicators. The first is the basic volatility, which is measured by the standard deviation of logarithmic returns. The system sets three calculation windows of 15 days, 30 days and 60 days to obtain volatility characteristics of different time scales. The second is the volatility indicator that considers directionality, calculating the upward volatility and downward volatility respectively, which is used to characterize the asymmetry of behavioral changes. In addition, the system also introduces a volatility estimation method based on extreme value theory, which is specifically used to capture abnormal volatility events.
[0056] The calculation of amplitude characteristics is more detailed. The system first identifies local extreme points in different time windows and calculates the maximum amplitude and average amplitude. Then, by comparing the time intervals and amplitude changes of adjacent extreme points, an oscillation strength indicator is constructed. The system pays special attention to the sudden change of amplitude. When the amplitude change exceeds 2 times the historical standard deviation, it is marked as a potential anomaly. At the same time, the system also calculates the attenuation characteristics of the amplitude to evaluate the persistence of behavioral fluctuations.
[0057] In order to improve the robustness of time series features, the system adopts an adaptive threshold mechanism. Specifically, the system dynamically adjusts the judgment criteria of volatility and amplitude based on the distribution characteristics of historical data. For example, during periods of volatile market fluctuations, the system will appropriately increase the threshold of volatility; during stable periods, the threshold will be lowered to increase sensitivity to small anomalies. This adaptive mechanism significantly improves the adaptability of features.
[0058] When integrating various indicators, the system constructs a hierarchical time series feature structure. The first layer is the basic features, including various moving averages and raw volatility indicators; the second layer is the derived features, including trend strength, volatility change rate, amplitude pattern, etc.; the third layer is the combined features, which are obtained through the cross-operation of multiple basic indicators. For example, the system may find that a certain behavior sequence has an upward trend (trend strength 0.75), accompanied by a gradual increase in volatility (volatility change rate 0.15) and asymmetric amplitude characteristics (upper and lower amplitude ratio 1.8).
[0059] Finally, the system standardizes all time series features to ensure that features of different dimensions are comparable. At the same time, the system also calculates the correlation between features, removes redundant features, and retains the most representative indicator combination. The time series features processed in this way can not only fully describe the dynamic change characteristics of the behavior, but also have high computational efficiency and interpretability.
[0060] Through this multi-level, multi-angle time series feature analysis, the system can accurately capture the long-term evolution trend, periodic change characteristics and abnormal fluctuation patterns of behavior patterns, providing reliable feature support for subsequent abnormal behavior identification. Especially when discovering gradual and sudden abnormalities, these fine time series features can often provide key early warning signals.
[0061] S3.5: Based on the structured data in the unified format data set and the structured feature vector, the frequency and intensity of direct interactions between entities are counted, and the weighted association coefficient is calculated to obtain the direct association degree; First, the frequency of interactions between entity pairs is counted, and the weighted coefficient is calculated based on the type, scale and other attributes of the interaction. The system uses an improved correlation coefficient calculation method to consider the timing characteristics, directionality and intensity of the interaction to generate a standardized correlation metric. These direct correlation features can reflect the closeness and dependency between entities.
[0062] First, in S3.5, the system conducts a multi-dimensional analysis of the direct interaction relationship between entities. The system first constructs an interaction matrix. Each element of the matrix contains not only the basic interaction frequency, but also attribute information such as the time distribution and amount distribution of the interaction. For each interaction between a pair of entities, the system calculates statistics in multiple dimensions: the mean and variance of the interaction frequency, the distribution characteristics of the single interaction amount (such as the mean, median, quantile, etc.), the interval characteristics of the interaction time, etc. These basic statistics provide data support for the subsequent calculation of the association strength.
[0063] When calculating the weighted correlation coefficient, the system uses a multi-level weight system. The first level is the interaction attribute weight, which assigns different basic weights according to the nature of the interaction (such as financial transactions, business cooperation, etc.). The second level is the time decay weight, which uses an exponential decay function to give recent interactions a higher weight. The third level is the abnormal intensity weight. When the attributes of an interaction (such as amount, frequency) deviate significantly from the historical pattern, the system will assign a higher weight. This multi-level weight design ensures that the correlation coefficient can accurately reflect the importance and abnormality of the relationship between entities.
[0064] The system constructs a scoring model for direct association by combining weighted association coefficients of multiple dimensions. This model not only considers the quantitative characteristics of the interaction, but also introduces qualitative factors, such as the continuity and stability of the relationship. For example, for two entities that interact frequently, the system will analyze whether their interaction patterns are regular and whether there are abnormal interaction periods or changes in amounts. This comprehensive analysis helps the system more accurately evaluate the strength of direct associations between entities.
[0065] S3.6: Based on the direct correlation degree, construct a multi-hop relationship path, calculate the path importance weight, and form a correlation feature; The indirect association paths between entities are explored through graph algorithms, and the importance weights of the paths are calculated based on factors such as path length and intermediate node features. The system also considers the diversity and stability of the paths and evaluates the complexity of the relationship network. Finally, these direct and indirect association features are integrated to form a complete set of association features, providing a basis for subsequent network structure analysis. Through this multi-level feature construction method, the system can comprehensively characterize the behavior patterns and association relationships of entities and improve the accuracy of anomaly identification.
[0066] In S3.6, the system starts to build and analyze multi-hop relationship paths based on the calculated direct association. First, the system uses an improved breadth-first search algorithm to find all possible multi-hop paths in the entity relationship network. In order to control the computational complexity, the system sets a maximum hop limit (usually 3 or 4 hops) and uses an association threshold to prune paths, retaining only paths with higher association strength.
[0067] In the process of path search, the system uses dynamic programming to calculate the importance weight of the path. Specifically, the calculation of path weight takes into account the following key factors: path length (number of hops), importance of each node on the path, attenuation effect of the strength of association between hops, uniqueness of the path, etc. The system designs a weight transfer model based on Markov chain so that the path weight can reasonably reflect the strength attenuation of indirect associations.
[0068] In order to improve the accuracy of path analysis, the system also introduces a path pattern recognition mechanism. First, the system defines a series of typical path patterns, such as loop pattern, star pattern, chain pattern, etc. Then, pattern matching is performed on each multi-hop path to identify the pattern type to which it belongs. Paths of different pattern types will receive different weight adjustment coefficients. For example, a loop pattern usually indicates a stronger association, so it will receive a higher weight bonus.
[0069] The system pays special attention to the temporal characteristics of the path. For each multi-hop path, the system analyzes the time series of each interaction on the path to identify whether there are obvious temporal patterns or abnormal time differences. For example, if the interactions on a path show obvious continuity or regularity, this may suggest some kind of organized behavior pattern, and the system will adjust the importance weight of the path accordingly.
[0070] Finally, the system integrates the direct association features and multi-hop path features to form a complete association feature set. This feature set contains multiple levels: node-level features (such as node degree centrality, betweenness centrality, etc.), direct association-level features (such as association strength, interaction mode, etc.), path-level features (such as path diversity, path strength, etc.), and network-level features (such as local clustering coefficient, community structure features, etc.). This multi-level feature system can comprehensively characterize the association relationship between entities and provide rich feature support for subsequent abnormal behavior identification.
[0071] Through this in-depth correlation analysis, the system can effectively identify complex correlation patterns and potential abnormal relationships. For example, the system may find that although some entities do not interact directly, they form close indirect connections through multiple intermediate nodes. This pattern may suggest some deliberate behavior to evade monitoring. At the same time, the time series analysis of correlation features can also help to discover dynamically evolving abnormal patterns, such as the gradually formed chain of fund transfers or the gradually expanding network of related transactions.
[0072] Assume that we analyze the fund transactions of Company A. The original data contains the transfer records for the past year: In step S3.1 (building behavioral statistical features): The system conducts multi-dimensional statistical analysis on the transfer behavior of enterprise A and obtains the following characteristic values: { "Basic statistical characteristics": { "Daily average transaction frequency": 8.5 times, "Monthly cumulative number of transactions": 255 times, "Maximum time interval": 36 hours, "Minimum time interval": 0.5 hours }, "Amount distribution characteristics": { "Daily average transaction amount": 1.8 million yuan, "Standard deviation of amount": 450,000 yuan, "Amount Kurtosis": 3.2, "Amount Skewness": 0.8 }, "Counterparty Characteristics": { "Monthly unique counterparties": 25, "Fixed counterparty ratio": 0.35, "New counterparty rate": 0.15 } } In step S3.2 (sliding time window analysis): The system sets up multiple time windows for trend analysis: { "7-day window": { "Transaction frequency change rate": +25%, "Amount change rate": +45%, "Counterparty change rate": +15% }, "30-day window": { "Transaction frequency change rate": +15%, "Amount change rate": +30%, "Counterparty change rate": +10% }, "90-day window": { "Transaction frequency change rate": +5%, "Amount change rate": +12%, "Counterparty change rate": +3% } } In step S3.3 (Periodic analysis): Through Fourier transform analysis, we get: { "Main cycle": { "Short-term cycle": "7 days", "Cycle Strength": 0.85, "Autocorrelation coefficient": 0.72 }, "Minor Cycle": { "Medium term": "30 days", "Cycle Strength": 0.45, "Autocorrelation coefficient": 0.38 } } In step S3.4 (Long-term trend analysis): The system calculates trend indicators on different time scales: { "EMA indicator": { "Short-term (7 days)": 1.85 million yuan, "Medium term (30 days)": 1.65 million yuan, "Long-term (90 days)": 1.5 million yuan }, "Volatility Indicator": { "Intraday Volatility": 0.25, "Weekly Volatility": 0.18, "Monthly Volatility": 0.12 }, "Trend Features": { "Uptrend Strength": 0.65, "Trend Stability": 0.78 } } In step S3.5 (direct association analysis): Systematic analysis of the direct relationship between Enterprise A and its counterparties: { "High Frequency Trading Counterparties": { "Company B": { "Interaction frequency": 45 times / month, "Average amount": 1.2 million yuan, "Strength of association": 0.82 }, "Company C": { "Interaction frequency": 35 times / month, "Average amount": 900,000 yuan, "Association Strength": 0.75 } }, "association mode": { "Two-way transaction ratio": 0.45, "Net capital flow": "Mainly outflow", "Time Concentration": 0.68 } } In step S3.6 (multi-hop relationship analysis): The system builds and analyzes indirect association networks: { "Second degree association": { "Number of nodes": 15, "Average Path Length": 2.3, "Key intermediate nodes": ["Enterprise D", "Enterprise E"], "Path Importance": { "A->B->F": 0.65, "A->C->G": 0.58 } }, "Third degree correlation": { "Number of nodes": 45, "Clustering coefficient": 0.42, "Number of critical paths": 8 }, "Network Features": { "Network Density": 0.35, "Centrality": 0.68, "Community structure strength": 0.72 } } Based on the above analysis, the system found that Enterprise A has the following abnormal characteristics: The transaction frequency and amount have increased suddenly recently (7-day window), exceeding the historical fluctuation range; The capital flow with enterprise B shows obvious one-way outflow characteristics; With enterprise D as the intermediate node, a close capital circulation path is formed.
[0073] These features together constitute the behavior profile of enterprise A, providing multi-dimensional feature support for subsequent anomaly identification. Based on these features, the system gave a risk score of 0.82 (high risk) and recommended manual verification. This example shows how to comprehensively characterize the behavior patterns and associations of enterprises through multi-level feature analysis, effectively supporting risk monitoring decisions.
[0074] S4: establishing an entity relationship graph according to the associated data in the multidimensional feature set, calculating the association strength between entities by using a low-rank tensor and a p-AAA algorithm, optimizing the entity relationship graph based on the association strength, and establishing a network structure model; When constructing the initial entity relationship graph based on the associated data in the multidimensional feature set, the system first structures the associated data. Each associated data contains information in multiple dimensions: entity pair identification, association type, association time, association strength, interaction features, etc. The system organizes these multidimensional data into a tensor structure, where each dimension of the tensor corresponds to a feature dimension. For example, a fourth-order tensor may contain: entity dimension, time dimension, relationship type dimension, and feature dimension. This tensor structure can fully preserve the multidimensional characteristics of the data.
[0075] In order to process high-dimensional tensor data, the system uses low-rank tensor decomposition technology. Specifically, the system first uses Tucker decomposition to decompose the original tensor into a combination of core tensors and factor matrices. During the decomposition process, the system determines the optimal rank parameter through cross-validation, and usually selects a rank value that can retain 85%-95% of the information. This decomposition can not only significantly reduce the data dimension, but also filter out the noise components in the data. For example, for a data set containing 1,000 entities, a time span of 12 months, and 10 types of relationships, it may be compressed to a rank of (50, 6, 5) through low-rank decomposition.
[0076] Based on the low-rank representation, the system applies the p-AAA (parallel Adaptive Anderson-Antoulas Algorithm) algorithm to construct an accurate representation of the strength of associations between entities. The core idea of the algorithm is to capture the nonlinear characteristics of associations through rational function approximation. The algorithm first selects the most representative interpolation points in the compressed feature space, and then iteratively optimizes the numerator and denominator polynomial coefficients of the rational function. The system adopts a parallel computing strategy to process multiple rational function approximation tasks at the same time, which significantly improves the computing efficiency.
[0077] An important feature of the p-AAA algorithm is its adaptability. The algorithm can automatically adjust the approximation strategy according to the local characteristics of the data, using more interpolation points in areas with complex relationships and fewer interpolation points in areas with simple relationships. The system sets an approximation accuracy threshold (usually 1e-6) and a maximum number of iterations (usually 1000) to ensure that the algorithm can converge to the desired accuracy within a reasonable time. This adaptive approximation strategy enables the system to accurately characterize different types of association patterns.
[0078] After obtaining the association strength estimate, the system starts to optimize the structure of the entity relationship graph. First, the system uses an adaptive threshold mechanism to filter weakly associated edges. The threshold is set based on the distribution characteristics of the association strength, and is usually selected so that the number of retained edges is between 20% and 40% of the total number of edges. In the process of deleting weakly associated edges, the system will also evaluate the connectivity of the network to ensure that the overall structure of the network is not destroyed due to excessive pruning.
[0079] Next, the system optimizes the network at the node level. The system designs a node similarity calculation method based on multi-dimensional features, taking into account the attribute characteristics, topological characteristics, and behavioral characteristics of the nodes. When the similarity between two nodes exceeds a preset threshold (usually 0.85), the system evaluates their possibility of merging. The merging process requires careful handling of the node's attribute inheritance and reconstruction of the association relationship to ensure that no false associations are introduced.
[0080] During the network optimization process, the system pays special attention to the identification of abnormal structures. For example, the system will detect abnormally dense subgraph structures, abnormal star structures or ring structures, etc. These special structures often imply potential abnormal association patterns. The system will give these structures special marks and pay more attention to them in subsequent analysis.
[0081] Finally, the system builds a complete network structure model. This model not only contains the optimized network topology, but also contains rich attribute information: node importance indicators (such as various centrality measures), edge weights and type information, subgraph structural characteristics, etc. The system also calculates a series of network-level statistical indicators, such as network density, average path length, clustering coefficient, etc. These indicators help to understand the overall characteristics of the network.
[0082] Through this sophisticated network construction and optimization process, the system finally obtains a network model that not only retains key structural information but also has good interpretability. This model can accurately reflect the complex relationships between entities and provide reliable network structure support for subsequent abnormal behavior identification. For example, the system may find that a group of companies has formed a closely related community structure through complex shareholding relationships. This discovery is of great value in identifying related-party transactions.
[0083] like Figure 4 As shown, step S4 also includes S4.1 to S4.4: S4.1: Use the correlation data in the multidimensional feature set to construct a multidimensional correlation tensor, use Tucker decomposition to reduce the tensor dimension, and obtain the main feature information; Specifically, a high-order tensor is constructed, whose dimensions include entities, relationship types, time and other aspects, and each tensor element represents the strength of association under the corresponding dimension combination. In order to process this high-dimensional data structure, the system uses Tucker decomposition technology to decompose the original tensor into a combination of core tensors and factor matrices. This decomposition can not only significantly reduce the data dimension, but also retain key feature information. For example, for data containing tens of thousands of entities and dozens of relationship types, Tucker decomposition can compress it to a suitable dimension while maintaining more than 90% of the information.
[0084] In S4.1, the system first constructs a high-dimensional association tensor. This tensor contains multiple key dimensions: entity dimension (I×J, where I and J represent the number of source entities and target entities, respectively), time dimension (T, representing the number of sampling points in the observation time period), relationship type dimension (R, containing different types of association relationships), and feature dimension (F, containing various interaction features). For example, for a data set containing 1,000 entities, 12 months of daily data, 10 relationship types, and 20 features, the dimension of the initial tensor is 1000×1000×365×10×20. This multidimensional structure can fully preserve the temporal evolution characteristics and multidimensional attribute information of entity associations.
[0085] When constructing tensors, the system uses a sparse storage format (COO format) to handle large-scale sparse data. For each non-zero element, the system records its complete index information and the corresponding eigenvalue. At the same time, the system normalizes the eigenvalues, using Min-Max normalization or Z-score normalization methods to ensure that features of different dimensions are comparable. In addition, the system also handles the problem of missing values, using appropriate interpolation methods to fill in missing tensor elements based on the temporal correlation and entity similarity of the data.
[0086] In order to reduce the tensor dimension, the system uses Tucker decomposition technology. Tucker decomposition decomposes the original tensor into the product of a core tensor and multiple factor matrices. Specifically, for a D-dimensional tensor X, its Tucker decomposition can be expressed as:
[0087] Where G is the core tensor and Ui is the factor matrix of the i-th dimension. The system solves this decomposition problem by alternating least squares (ALS) and uses a multi-start strategy to avoid local optimal solutions.
[0088] The system uses an adaptive strategy to determine the rank parameter of Tucker decomposition. First, the singular value spectrum of the tensor is calculated and the energy distribution is analyzed. Then, the minimum rank value combination that can retain more than 90% of the information is selected. For example, the original tensor may be compressed to a rank of (50,50,30,5,10), which significantly reduces the data dimension while maintaining key structural information. The system also verifies the compression quality through reconstruction error and cross-validation to ensure that the compressed representation can accurately reflect the characteristics of the original data.
[0089] S4.2: Based on the main feature information, construct a rational function approximation through the p-AAA algorithm, iteratively optimize the approximation accuracy, and obtain the association strength between entities; First, based on the low-dimensional features obtained by Tucker decomposition, a rational function is constructed to approximate the association relationship between entities. Through an iterative optimization process, the algorithm adaptively selects the optimal interpolation points and continuously adjusts the numerator and denominator polynomial coefficients of the rational function until the preset approximation accuracy is reached. Compared with traditional methods, the p-AAA algorithm has better numerical stability and convergence, and can more accurately characterize the nonlinear association relationship between entities. For example, when dealing with association patterns with strong periodicity or mutation characteristics, the algorithm can maintain a high fitting accuracy.
[0090] In S4.2, the system uses the p-AAA algorithm to construct an accurate representation of the strength of association between entities based on the compressed feature representation. The core idea of the p-AAA algorithm is to capture nonlinear association patterns through rational function approximation. The algorithm constructs rational function approximations in each dimension, and then combines these approximations in the form of tensor products to finally obtain a representation of multi-dimensional association strength.
[0091] Specifically, for each pair of entities (i, j), the system constructs a rational function of the following form:
[0092] Where z represents the characteristic variable, p(z) and q(z) are the numerator and denominator polynomials, respectively. The system determines the coefficients of the polynomials through iterative optimization, so that the rational function can accurately approximate the observed correlation pattern.
[0093] In the implementation of the p-AAA algorithm, the system adopts a parallel computing strategy. First, the entity pairs are divided into multiple batches, and each batch simultaneously calculates multiple rational function approximations. The system uses GPU to accelerate matrix operations, which significantly improves the execution efficiency of the algorithm. For each approximation task, the system sets the following key parameters: maximum degree of numerator polynomial: m = 5; maximum degree of denominator polynomial: n = 4; convergence threshold: ε = 1e-6; maximum number of iterations: maxIter = 1000; The iterative process of the algorithm includes the following steps: adaptively select interpolation points, give priority to data points with significant correlation features; construct the Loewner matrix and calculate its SVD decomposition; solve the linear equations and update the rational function coefficients; calculate the approximation error and determine whether the convergence conditions are met; in the approximation process, the system pays special attention to numerical stability. Regularization technology is used to control the size of the coefficients to prevent overfitting. At the same time, the system monitors the condition number, and automatically adjusts the regularization parameters or reselects the interpolation points when the numerical instability is found. Finally, the system obtains an estimate of the correlation strength between each pair of entities.
[0094] These estimates contain not only scalar strengths but also confidence interval information, reflecting the reliability of the estimates. The system organizes these association strengths into a matrix form as the basis for subsequent network optimization. For example, for two entities that interact frequently, the system may obtain the following results: { "Association strength": 0.85, "Confidence Interval": [0.82, 0.88], "Approximation error": 0.003, "Convergence rounds": 245 } Through this precise calculation of association strength, the system can accurately characterize the complex association patterns between entities, providing a reliable quantitative basis for subsequent network structure optimization. Especially when dealing with nonlinear association patterns, this rational function-based approximation method shows obvious advantages and can capture complex association features that traditional linear methods may ignore.
[0095] S4.3: according to the association strength, setting an association strength threshold, deleting edges below the association strength threshold, and constructing an optimized entity relationship graph; First, the system uses an adaptive threshold setting method to dynamically determine the cutoff threshold of the association strength based on the connection density of the overall network and the characteristics of the business scenario. For edges below the threshold, the system will delete them from the network to reduce the impact of noise. This process is progressive, and the system will re-evaluate the connectivity and structural characteristics of the network after each edge deletion to ensure that the optimized network still maintains a reasonable topological structure. For example, for a network containing tens of thousands of edges, this optimization can delete about 60% of the weakly associated edges while maintaining the main connectivity of the network.
[0096] In S4.3, the system first uses an adaptive method to set the association strength threshold based on the statistical distribution characteristics of the association strength. Specifically, the system calculates the distribution statistics of the association strength, including the mean (μ), standard deviation (σ), quantiles, etc. Based on these statistics, the system designs a three-level threshold mechanism: Strong correlation threshold (T 1 =μ+1.5σ); Medium correlation threshold (T 2 =μ+0.5σ); Weak correlation threshold (T 3 =μ-0.5σ).
[0097] This hierarchical threshold setting allows the system to adopt differentiated processing strategies for associations of different strengths.
[0098] In the process of edge filtering, the system adopts a progressive strategy. First, the edges below the weak correlation threshold (T 3 ), these edges usually represent noise or accidental associations. 3 and T 2 The system will further evaluate the time persistence and stability of the edges between them. If an edge has a weak association strength but shows significant time persistence (such as existing for several consecutive months), the system will consider retaining the edge. This filtering strategy that considers the time dimension can avoid accidentally deleting important long-term stable associations.
[0099] The system pays special attention to maintaining the connectivity of the network. During the process of deleting edges, the system monitors several key indicators of the network in real time: average degree, clustering coefficient, number and size distribution of connected components, etc. When it is found that an edge deletion may cause a significant change in the network structure (such as the generation of isolated large connected components), the system will re-evaluate the importance of the relevant edges. For example, if deleting an edge will cause an important community structure to be split, the system may choose to retain the edge even if its association strength is relatively low.
[0100] S4.4: using the optimized entity relationship diagram, identifying nodes whose similarity exceeds a preset threshold, merging the nodes, updating the association relationship, and forming a network structure model; In step S4.4, the system further optimizes the structure of the optimized entity relationship graph. First, a node similarity calculation method based on multidimensional features is designed, which comprehensively considers the attribute characteristics, topological characteristics and behavioral characteristics of the nodes. When the similarity of two nodes exceeds the preset threshold, the system will evaluate their possibility of merging. During the merging process, the system will carefully handle the inheritance and merging of association relationships to ensure that no false associations are introduced. At the same time, for the merged nodes, the system will recalculate the strength of its association with other nodes and update the network structure. This optimization can not only reduce the complexity of the network, but also discover potential entity association groups, providing a clearer network structure foundation for subsequent abnormal behavior identification. For example, through node merging, entities that are apparently independent but actually highly related can be identified as an association group, effectively preventing the behavior of evading monitoring through decentralized operations.
[0101] In S4.4, the system conducts node-level optimization. First, a node similarity calculation framework is constructed, which comprehensively considers the characteristics of multiple dimensions: attribute similarity (such as the basic characteristics of entities), structural similarity (such as the local network structure of nodes), and behavioral similarity (such as interaction patterns). For each dimension, the system designs a special similarity calculation method: Attribute similarity: Use the weighted Jaccard coefficient to calculate the similarity of categorical features, and use cosine similarity to calculate the similarity of numerical features; Structural similarity: Calculates structural similarity based on the node's neighbor set, taking into account the importance weights of neighbor nodes; Behavior similarity: The dynamic time warping (DTW) algorithm is used to calculate the similarity of behavior sequences; The system weights the similarities of these dimensions and combines them to get a comprehensive similarity score. The weights are set based on the importance and reliability of the features, and the optimal weight combination is determined through cross-validation. When the comprehensive similarity of two nodes exceeds a preset threshold (usually set to 0.85), the system marks them as potential merge candidates.
[0102] During the node merging process, the system adopts a cautious strategy. First, for each pair of candidate nodes, the system conducts an in-depth correlation analysis to check whether they have obvious complementarity or conflict. For example, if two nodes show obvious complementarity in time (one node is active when the other is inactive), this may suggest that they are actually the manifestations of the same entity in different periods, and merging is reasonable at this time.
[0103] The merge operation needs to deal with three key issues: attribute inheritance: determine how the merged node inherits the attribute value of the original node, which may adopt weighted average, maximum value or other appropriate aggregation methods; association reconstruction: handle the association relationship between the merged node and other nodes, which requires reasonable accumulation or average association strength; time series information preservation: ensure that the time series characteristics of the merged node can accurately reflect the time evolution pattern of the original node; For example, when merging two highly similar enterprise nodes, the system may generate the following merge record: { "Merge Node Pair": ["Enterprise A", "Enterprise B"], "Similarity score": 0.89, "attribute inheritance strategy": { "Registered Capital": "Maximum", "Business Scope": "Combination", "Risk Level": "Highest" }, "Relation Refactoring": { "Update Edges": 15, "Maximum Intensity Change": 0.12 } } After completing the node merger, the system updates and optimizes the entire network structure. This includes recalculating the network's topological features, updating the node's centrality index, adjusting the community division results, etc. The system also generates a detailed optimization report to record changes in the network structure, such as the reduction ratio of the number of nodes, changes in edge density, and the evolution of the community structure.
[0104] This optimized network structure model provides a reliable basis for subsequent abnormal behavior identification. For example, the system can more accurately identify abnormal capital flow paths or suspicious related transaction patterns based on this optimized network structure.
[0105] S5: according to the network structure model and the multidimensional feature set, a feature weight system is set, a multidimensional anomaly score is calculated using a pruned tensor structure measurement method, a low-rank tensor recovery technique is used to comprehensively calculate the multidimensional anomaly score to obtain an anomaly score, and an early warning list is output according to the anomaly score; like Figure 5 As shown, step S5 includes steps S5.1 and S5.2: S5.1: According to the network structure model and the multi-dimensional feature set, the weight value of the closeness of the associated party, the weight value of the abnormal behavior and the weight value of the abnormal timing are set to construct a weight configuration; In step S5.1, the system builds a multi-level weight configuration system based on the network structure model and the multidimensional feature set. First, for the setting of the weight value of the closeness of the related party, the system considers the position importance and connection pattern of the entity in the network. Specifically, the influence of the entity in the network is evaluated by calculating indicators such as the degree centrality, betweenness centrality and eigenvector centrality of the node. At the same time, the path diversity and clustering coefficient between entities are considered to quantify the stability of the association relationship. For example, for entities that are at the core of the network and have diverse connections, the system will assign a higher closeness weight to the related party, usually between 0.3-0.5.
[0106] For the weight value of abnormal behavior, the system mainly sets it based on the distribution of the entity's behavioral characteristics. First, the frequency, scale, and complexity of the behavior are analyzed to establish a baseline behavior pattern. Then, the weight value is dynamically adjusted according to the degree to which the behavior deviates from the baseline pattern. For example, for trading behavior, the system will focus on features such as sudden changes in trading frequency and abnormal changes in counterparties, and set weights based on the significance of these features, with typical values ranging from 0.2 to 0.4.
[0107] The setting of the time series anomaly weight value focuses on the time evolution characteristics of the behavior pattern. The system evaluates the degree of abnormality of the time series pattern by analyzing the periodicity, trend and mutation of the behavior sequence. For different time scales (such as day, week, month), the system will set different weight coefficients to capture short-term, medium-term and long-term abnormal patterns. Usually, the weight value of short-term anomalies will be slightly higher, set between 0.3-0.5.
[0108] S5.2: Based on the weight configuration, the weight value of the closeness of the associated party, the weight value of the abnormal behavior, and the weight value of the abnormal time series are refined to obtain a feature weight system; In step S5.2, the system makes fine adjustments to the initially set weight values. First, the system introduces a hierarchical adjustment mechanism for the closeness weights of related parties. For direct relationships, the system refines the weights based on factors such as interaction frequency, interaction scale, and relationship duration; for indirect relationships, the system attenuates the weights based on path length, intermediate node characteristics, etc. For example, the weight of a second-degree relationship will usually decay by 40%-60% based on the original weight.
[0109] For the refinement of the weight of abnormal behavior, the system adopts a multi-dimensional decomposition strategy. The abnormal behavior is decomposed into multiple sub-dimensions, such as transaction abnormality, operation abnormality, relationship abnormality, etc., and the weight ratio of each sub-dimension is set according to the specific business scenario. The system also considers the severity and scope of impact of the behavior, and assigns higher weights to high-risk behavior types.
[0110] In the process of refining the weight of time series anomaly, the system introduces an adaptive adjustment mechanism. First, the benchmark interval of the time series pattern is established based on the fluctuation characteristics of historical data. Then, the weight value is dynamically adjusted according to the degree of deviation between the real-time data and the benchmark interval. The system also considers the impact of seasonal factors and special periods and adjusts the weight sensitivity appropriately. For example, during business peak periods, the system will appropriately lower the weight threshold of time series anomalies to reduce the false alarm rate.
[0111] Through this multi-level, refined weight system, the system can more accurately characterize the various dimensions of abnormal behavior and improve the accuracy and explainability of identification. At the same time, this weight system has strong adaptability and can be dynamically adjusted according to changes in business scenarios and the evolution of abnormal patterns.
[0112] Suppose that a financial transaction network consisting of 2,000 companies is monitored, and the goal is to identify possible abnormal transaction behaviors. In this case, the multidimensional feature set includes the following dimensions: transaction behavior characteristics (such as transaction frequency, amount, time distribution), basic enterprise characteristics (such as registered capital, years of operation, industry category), network structure characteristics (such as node centrality, community affiliation) and time series evolution characteristics (such as the changing trend of transaction patterns).
[0113] First, the system sets a weight system for these features. Taking the transaction behavior dimension as an example, the system may give the following weight distribution: Abnormal transaction amount: 0.35 (high weight, directly reflects the degree of risk); Abnormal transaction frequency: 0.25 (medium-high weight, reflecting changes in behavior patterns); Counterparty anomaly: 0.25 (medium-high weight, reflecting correlation risk); Abnormal trading time: 0.15 (medium-low weight, auxiliary indicator); Taking a specific company A as an example, its abnormal characteristics during a certain observation period are as follows: The transaction amount suddenly increased: the average daily transaction amount in the past month is 5 times the historical average; The trading frequency has increased significantly: the number of daily transactions has increased from the historical average of 3 times to 15 times; Added multiple counterparties that had never been traded before: 10 new counterparties were added in one week; Abnormal trading hours: A large number of transactions occur during non-working hours; Using the pruned tensor structure measurement method, the system organizes these abnormal features into a fourth-order tensor (company × feature × time × associated entity). For company A, the system calculates the abnormal score of each dimension: Abnormal score of the amount dimension: 0.85 (serious abnormality); Frequency dimension anomaly score: 0.78 (significant anomaly); Adversary dimension anomaly score: 0.92 (severe anomaly); Time dimension anomaly score: 0.65 (moderate anomaly); The system then applies low-rank tensor recovery technology to comprehensively calculate these multi-dimensional anomaly scores. Considering the correlation between different dimensions, the system uses Tucker decomposition for low-rank approximation to obtain a more robust anomaly score. For company A, the final anomaly score calculation process is as follows: Initial comprehensive score: Abnormality score = 0.85 × 0.35 + 0.78 × 0.25 + 0.92 × 0.25 + 0.65 × 0.15 = 0.82; Adjustments to consider timing dependencies: Abnormal duration: 15 consecutive days; Timing adjustment factor: 1.2 (abnormal duration exceeds 10 days); Adjusted score: 0.82 × 1.2 = 0.984; Consider adjusting the network structure: Abnormal proportion of associated entities: 30%; Network adjustment factor: 1.1 (high-risk community); Final anomaly score: 0.984 × 1.1 = 0.982; Based on this anomaly score, the system generates the following warning information: Warning level: Level 1 (highest level) warning targets: Company A Anomaly score: 0.982; Main abnormal features: The transaction amount fluctuates dramatically; Rapid expansion of trading counterparties; The transaction frequency increases abnormally; Associated risks: Three of the 10 new trading counterparties also showed abnormalities; The risk level of the community you belong to increases; Recommended actions: Immediately initiate on-site verification; Temporary restrictions on large transactions; Closely monitor affiliated companies; At the same time, the system also identified a risk transmission chain closely related to Enterprise A: Enterprise A → Enterprise B (anomaly score 0.875) → Enterprise C (anomaly score 0.823) → Enterprise D (anomaly score 0.792); This indicates the possible existence of an organized risk network, and the system has accordingly expanded the monitoring scope to include these related companies in the key monitoring list.
[0114] On a larger scale, the system generates a tiered alert list based on anomaly scores: Level 1 warning (score>0.9): 5 companies; Level 2 warning (score 0.8-0.9): 12 companies; Level 3 warning (score 0.7-0.8): 25 companies; For each warning object, the system generates a detailed abnormal feature analysis report, including: A time-series evolution diagram of abnormal behavior; Contribution analysis of key abnormal features; Visual display of risk transmission paths; Comparative analysis of historical abnormal behavior; Targeted regulatory recommendations; This multi-dimensional abnormal scoring and early warning mechanism can help regulators quickly locate high-risk entities and understand the specific characteristics and transmission paths of abnormal behaviors, so as to take targeted regulatory measures. For example, in the case of Enterprise A, the regulatory authorities may immediately initiate an on-site inspection and implement joint supervision of its affiliated companies to effectively prevent the spread and evolution of risks.
[0115] In addition, if Figure 6 As shown, the present application also provides an abnormal behavior identification system 600 based on multi-dimensional data analysis, which may specifically include: The data preprocessing module 601 is used to obtain structured data and unstructured data through the business system interface, and perform data cleaning and standardization on the structured data and the unstructured data to obtain a data set in a unified format; A text processing module 602 is used to receive the document text in the unified format data set, process the document text using natural language processing technology, extract entity information, relationship information and behavior information, and generate a structured feature vector; A feature engineering module 603 is used to construct behavioral statistical features, temporal features and correlation features by using the structured data in the unified format data set and the structured feature vector, and to select feature combinations by using a sampling algorithm based on the generalized Golub-Kahan method to construct a multidimensional feature set; A network construction module 604 is used to establish an entity relationship diagram according to the associated data in the multi-dimensional feature set, calculate the association strength between entities through low-rank tensor and p-AAA algorithm, and optimize the entity relationship diagram based on the association strength to establish a network structure model; The scoring calculation module 605 is used to set a feature weight system based on the network structure model and the multidimensional feature set, calculate the multidimensional anomaly score using the pruned tensor structure measurement method, use the low-rank tensor recovery technology to comprehensively calculate the multidimensional anomaly score to obtain the anomaly score, and output a warning list based on the anomaly score.
[0116] The system of the embodiments of the present disclosure can execute the method provided by the embodiments of the present disclosure, and the implementation principles are similar. The actions performed by each module in the device of each embodiment of the present disclosure correspond to the steps in the method of each embodiment of the present disclosure. For the detailed functional description of each module of the device, please refer to the description in the corresponding method shown in the previous text, which will not be repeated here.
[0117] Finally, it should be noted that the above-described embodiments are only specific implementation methods of the present disclosure, which are used to illustrate the technical solutions of the present disclosure, rather than to limit them. The protection scope of the present disclosure is not limited thereto. Although the present disclosure is described in detail with reference to the aforementioned embodiments, ordinary technicians in the field should understand that any technician familiar with the technical field can still modify the technical solutions recorded in the aforementioned embodiments within the technical scope disclosed in the present disclosure, or can easily think of changes, or make equivalent replacements for some of the technical features therein; and these modifications, changes or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present disclosure, and should be included in the protection scope of the present disclosure. Therefore, the protection scope of the present disclosure shall be based on the protection scope of the claims.
Claims
1. A method for identifying abnormal behavior based on multidimensional data analysis, characterized in that: include: Acquire structured data and unstructured data through a business system interface, perform data cleaning and standardization on the structured data and the unstructured data, and obtain a data set in a unified format; Receiving the document text in the unified format data set, processing the document text using natural language processing technology, extracting entity information, relationship information and behavior information, and generating a structured feature vector; Using the structured data in the unified format data set and the structured feature vector, constructing behavioral statistical features, temporal features and correlation features, using a sampling algorithm based on the generalized Golub-Kahan method to select feature combinations, and constructing a multidimensional feature set; According to the associated data in the multidimensional feature set, an entity relationship graph is established, the association strength between entities is calculated by low-rank tensor and p-AAA algorithm, and the entity relationship graph is optimized based on the association strength to establish a network structure model; According to the network structure model and the multidimensional feature set, a feature weight system is set, and a multidimensional anomaly score is calculated using a pruned tensor structure measurement method. The multidimensional anomaly score is comprehensively calculated using a low-rank tensor recovery technique to obtain an anomaly score, and a warning list is output based on the anomaly score.
2. The method according to claim 1, characterized in that The structured data and the unstructured data are cleaned and standardized to obtain a data set in a unified format, including: For the structured data and the unstructured data, a statistical method is used to identify abnormal values of numerical data and make corrections, the coding of categorical data is standardized, and the format of time data is unified to obtain cleaned data; According to the cleaned data, the field names, data types, and value ranges are unified to obtain a data set in a unified format.
3. The method according to claim 1, characterized in that The document text is processed using natural language processing technology, including: For the document text, a conditional random field model is used to perform word segmentation processing to obtain a word sequence; For the term sequence, named entity recognition technology is used to identify key entities, and relationship description words between entities are extracted to construct entity-relationship pairs; According to the entity-relationship pairs, dependency syntactic analysis is used to extract target grammatical components, identify behavior types and behavior features, and output structured feature vectors.
4. The method according to claim 1, characterized in that The method of constructing the behavior statistical features by using the structured data in the unified format data set and the structured feature vector includes: Based on the structured data in the unified format data set and the structured feature vector, the frequency and time distribution of the behavior are counted, and the target statistics of the behavior frequency are calculated to form a frequency feature; According to the frequency characteristics, a sliding time window is set, and the frequency changes at different time scales are calculated to obtain the behavior statistical characteristics.
5. The method according to claim 1, characterized in that The method of constructing time series features by using the structured data in the unified format data set and the structured feature vector includes: Based on the structured data in the unified format data set and the structured feature vector, the periodic pattern of the behavior is detected by Fourier transform, the autocorrelation coefficient is calculated, and the periodicity index is constructed; For the periodic indicators, the moving average method is used to analyze the long-term trend, calculate the volatility and amplitude, and output the time series characteristics.
6. The method according to claim 1, characterized in that The method of constructing the associated features by using the structured data in the unified format data set and the structured feature vector includes: Based on the structured data in the unified format data set and the structured feature vector, the frequency and intensity of direct interactions between entities are counted, and the weighted association coefficient is calculated to obtain the direct association degree; Based on the direct correlation degree, a multi-hop relationship path is constructed, and the path importance weight is calculated to form a correlation feature.
7. The method according to claim 1, characterized in that The association strength between entities is calculated using low-rank tensors and the p-AAA algorithm, including: Using the correlation data in the multidimensional feature set, constructing a multidimensional correlation tensor, using Tucker decomposition to reduce the tensor dimension, and obtaining main feature information; According to the main feature information, a rational function approximation is constructed through the p-AAA algorithm, the approximation accuracy is iteratively optimized, and the association strength between entities is obtained.
8. The method according to claim 1, characterized in that The optimizing the entity relationship diagram based on the association strength includes: According to the association strength, an association strength threshold is set, edges below the association strength threshold are deleted, and an optimized entity relationship graph is constructed; The optimized entity relationship diagram is used to identify nodes whose similarity exceeds a preset threshold, merge the nodes, update the association relationship, and form a network structure model.
9. The method according to claim 1, characterized in that: The step of setting a feature weight system based on the network structure model and the multi-dimensional feature set includes: According to the network structure model and the multi-dimensional feature set, the weight value of the closeness of the associated party, the weight value of the abnormal behavior and the weight value of the abnormal timing are set to construct a weight configuration; Based on the weight configuration, the associated party closeness weight value, the behavior abnormality weight value and the time series abnormality weight value are refined to obtain a feature weight system.
10. An abnormal behavior identification system based on multidimensional data analysis, characterized in that: include: A data preprocessing module is used to obtain structured data and unstructured data through a business system interface, and to perform data cleaning and standardization on the structured data and the unstructured data to obtain a data set in a unified format; A text processing module, used to receive the document text in the unified format data set, process the document text using natural language processing technology, extract entity information, relationship information and behavior information, and generate a structured feature vector; A feature engineering module, for constructing behavioral statistical features, temporal features and correlation features by using the structured data in the unified format data set and the structured feature vector, and selecting feature combinations by using a sampling algorithm based on the generalized Golub-Kahan method to construct a multidimensional feature set; A network construction module, used to establish an entity relationship graph according to the associated data in the multidimensional feature set, calculate the association strength between entities through low-rank tensor and p-AAA algorithm, and optimize the entity relationship graph based on the association strength to establish a network structure model; The scoring calculation module is used to set a feature weight system according to the network structure model and the multidimensional feature set, calculate the multidimensional anomaly score using the pruned tensor structure measurement method, use the low-rank tensor recovery technology to comprehensively calculate the multidimensional anomaly score to obtain the anomaly score, and output a warning list according to the anomaly score.
Citation Information
Patent Citations
Customer portrait key data mining method and system based on space-time big data
CN118797542A
Intelligent risk control method, device and equipment for multi-source data fusion and storage medium
CN119046647A
Financial transaction anomaly detection and risk assessment method and device based on artificial intelligence
CN119693111A
Data anomaly diagnosis method and system based on knowledge graph and large model
CN119807960A
Methods and systems for anomaly and pattern detection of unstructured big data
US20230186120A1
Cited By
File co-processing method and system based on cloud computing
CN120561626A
Intelligent walking stick early warning method and system based on multi-parameter physiological monitoring
CN120938373A
Personalized error analysis and practice generation system based on AI algorithm
CN120996172A
AI-based personalized error analysis and practice generation system
CN120996172B
Software improvement suggestion generation method based on fault feature matching and role guidance
CN121233477A