Risk abnormal behavior event identification method based on multi-source risk abnormal behavior data fusion analysis

By constructing a multi-source risk and abnormal behavior data fusion analysis method, and utilizing knowledge graphs and large language models, the problem of low accuracy in identification by a single data model is solved, and efficient identification and early warning of complex risk and abnormal behaviors are achieved.

CN120930001APending Publication Date: 2025-11-11HENAN XINDA WANGYU TECH CO LTD +1
View PDF 0 Cites 7 Cited by

Patent Information

Application Number
CN202510987917.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-17
Publication Date
2025-11-11

AI Technical Summary

Technical Problem

In existing technologies, the identification of risky and abnormal behaviors relies on data models from a single industry or field, resulting in low accuracy and recall rates, and making it impossible to effectively identify complex cross-industry risky and abnormal behaviors.

Method used

By constructing a knowledge graph and integrating multi-source data such as communication, the Internet, funds, and cloud video conferencing, we use graph attention networks and large language models to encode and align multimodal features, and design prompt templates to identify risky behaviors.

Benefits of technology

It achieves highly accurate identification of complex risk and abnormal behaviors, improves the recall rate and identification capability of detection, can identify cross-industry risk and abnormal behavior chains, and provides detailed risk warning information.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120930001A_ABST
    Figure CN120930001A_ABST
Patent Text Reader

Abstract

The invention provides a risk abnormal behavior event identification method based on multi-source risk abnormal behavior data fusion analysis. The method comprises the following steps: S1, obtaining abnormal behavior label data; s2, constructing a knowledge graph ontology structure; s3, extracting entities, attributes and relationships involved in the structured data of the abnormal behavior label data into the constructed knowledge graph ontology structure; s4, for the constructed knowledge graph ontology structure, encoding graph data to obtain corresponding modal features; aiming at the structured data of the knowledge graph ontology structure, coding each source by adopting a corresponding feature coding method to obtain a corresponding modal feature; s5, the obtained modal features are input and mapped to the same vector space for alignment fusion; s6, performing fine tuning training to obtain an abnormal risk behavior recognition model LLM; and S7, superposing the fused multi-modal features, inputting the superposed multi-modal features to the LLM, and guiding the LLM to generate a corresponding output or decision according to the prompt of the specified input.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of data mining technology, and in particular relates to a method for identifying risky and abnormal behavior events based on the fusion analysis of multi-source risky and abnormal behavior data. Background Technology

[0002] In recent years, incidents of abnormal risk behavior have occurred frequently, and the forms of such behavior have become increasingly diversified, organized, and covert, becoming prominent incidents that undermine people's sense of happiness, security, and ability to obtain benefits.

[0003] With the development of network technology and the rapid growth of internet users in China, the forms and media of abnormal risk behaviors are constantly evolving. From the earliest abnormal risk behaviors via telephone and SMS, they have gradually developed into a combination of multiple methods, including telephone, SMS, apps, websites, and cloud video conferencing. The tools used to analyze abnormal risk behaviors have also diversified beyond traditional telephone and SMS methods. This has rendered traditional models based on data from single industry sectors insufficient to support the analysis of increasingly complex abnormal risk behavior events.

[0004] Figure 1 This paper describes the general process for identifying and issuing early warnings of risky and abnormal behavior events using data in various industries. The main shortcomings of the current technical solutions are as follows: (1) Currently, the monitoring and identification of abnormal behaviors in telecommunications networks are all based on data models from a single industry or field. This results in a single dimension of data, leading to low accuracy and recall when identifying abnormal behaviors. For example, only telephone call records or bank transaction records may be used, without integrating other relevant data, which will inevitably lead to missed detections or false positives.

[0005] Identifying anomalous risk behavior is a complex process involving multiple stages, encompassing various data types such as communication data (e.g., phone calls, text messages), network data (e.g., website visits, video conferencing), financial data (e.g., fund flows, account information), and personal information (e.g., user identity information). Current monitoring methods primarily rely on single-domain data models. This approach can only identify localized features of anomalous risk behavior events within a specific vertical domain, limiting analysis to a single stage and resulting in numerous missed and false detections, thus severely restricting the ability to identify anomalous risk behavior events.

[0006] To improve the accuracy and comprehensiveness of identifying risky and abnormal behaviors, it is essential to overcome the limitations of single data models and emphasize the integration of multi-source data and the construction of multimodal cognition. Only through comprehensive analysis of data from various aspects can we more accurately identify and predict risky and abnormal behaviors, thereby improving the accuracy and recall of identification and reducing missed and false detections. Summary of the Invention

[0007] To address the aforementioned issues, it is necessary to provide a method for identifying risky and abnormal behavior events based on the fusion analysis of multi-source risky and abnormal behavior data.

[0008] The first aspect of this invention proposes a method for identifying risky and abnormal behavior events based on the fusion analysis of multi-source risky and abnormal behavior data, comprising: S1: Obtain abnormal behavior tag data; the abnormal behavior tag data includes abnormal behavior tag data from the communication side, abnormal behavior tag data from the Internet side, abnormal behavior tag data from the funding side, abnormal behavior tag data from cloud video conferencing, and case and police data. S2: Abstract the entities and relationships involved in the abnormal behavior tag data, define the entities, attributes and relationships of the knowledge graph, and construct the knowledge graph ontology structure; S3: Based on the defined entities, attributes, and relationships, the acquired abnormal behavior label data is extracted into the constructed knowledge graph ontology structure through ETL. All entities have a globally unique identifier (vid). S4: For the constructed knowledge graph ontology structure, use the Graph Attention Network (GAT) to encode the graph data and obtain the corresponding modal features; For structured data of knowledge graph ontology, corresponding feature encoding methods are used for each source to obtain corresponding modal features; S5: Map the modal features obtained after encoding by different encoding methods in S4 to the same vector space, and align the modal features to ensure the matching of multimodal features in size and semantics, thereby achieving the fusion of multimodal features; S6: Based on the large language model, and using the fused features, fine-tuning training yields an abnormal risk behavior recognition model LLM based on the large language model; S7: The fused multimodal features are superimposed on the specified input and fed to the LLM, thereby guiding the LLM to generate corresponding outputs or decisions based on the prompts of the specified input; wherein, the specified input is a prompt template designed using the prompt learning method.

[0009] Based on the above, in step S1: Abnormal behavior tag data from the communication side includes: abnormal calling frequency, abnormal called number dispersion, abnormal called number attribution dispersion, abnormal calling time pattern, abnormal calling number empty number rate, abnormal called hang-up rate, and abnormal call success rate. Abnormal behavior tag data from the Internet side includes: abnormal behavior URLs, abnormal channel app downloads, apps that illegally obtain personal information, impersonation scripts, joining group chats, order brushing, and gaining followers; Abnormal behavior tags from the funding side include: small transfers, password changes, dormant accounts, malicious overdrafts, and large cash withdrawals; Abnormal behavior tags from cloud video conferencing include: abnormal voiceprint features, abnormal sharing, abnormal login accounts, and voice alteration.

[0010] Based on the above, for the structured data of the knowledge graph ontology, the corresponding feature encoding methods are used for encoding each source to obtain the corresponding modal features. For structured data in the abnormal behavior label data on the communication side, feature extraction is performed using the following method: Abnormal calling frequency: Calculate the number of calls per unit time and standardize these values ​​or convert them to relative values ​​using quantiles; Abnormal dispersion of called numbers: This is measured by calculating the diversity of called numbers; Abnormal dispersion of called number attribution: Measured by calculating the diversity of the regions to which the called number belongs; Abnormal call timing patterns: Analyze the distribution of call times throughout the day to identify call patterns during atypical time periods; Abnormal caller ID non-existent rate: Calculate the proportion of invalid numbers dialed; Abnormal call hang-up rate: Statistics on the probability that the called party hangs up the phone; Call success rate anomaly: Calculate the proportion of successfully connected calls out of all attempted calls; The numerical features obtained above are encoded using Z-score or quantile discretization + Embedding layer to obtain the modal features on the communication side. For the structured data in the abnormal behavior tag data on the Internet side, natural language processing technology is used to convert it into vector representation, that is, to obtain the modal features on the Internet side; For the structured data in the abnormal behavior tag data on the funding side, the amount and operation type are directly used as numerical features, and at the same time, the features are constructed and encoded based on the user's behavior pattern. Dormant accounts are marked as accounts that have not been active for a long time, and their trading behavior is analyzed after they are reactivated. Malicious overdrafts: Track changes in the ratio between account balance and credit limit, paying particular attention to the occurrence of overdraft behavior; Large cash withdrawals: Records large cash withdrawal events; Collect features including but not limited to the above, perform binning encoding and behavior sequence encoding, and then fuse them; The structured data in the abnormal behavior tagging data of cloud video conferencing includes video conferencing data and conference voiceprint data; Video conferencing data encoding: The ViViT model is used to encode the video conferencing data to obtain the modal features of the video conferencing; Conference voiceprint data encoding: The conference voiceprint data is encoded using the ECAPA-TDNN model to obtain the modal features of the conference voiceprint.

[0011] Based on the above, the method for achieving feature alignment across modalities is as follows: Based on the Multi-Head Self-Attention implementation in the Transformer architecture, in order to match the feature dimensions of different modal features, the different modal features extracted in S4 are processed as follows: One-dimensional convolutional layers are used to reduce the feature dimension of modal features respectively; Then, a linear layer is used to adjust the feature dimensions of each modality feature to be consistent with the text feature dimensions output by the LLM; Finally, Cross Attention is used to align the visual and audio features from each modal feature after dimension adjustment as the query for Cross Attention, and the text features as the key and value. The aligned modal features are concatenated with the text features to complete the fusion of multimodal features.

[0012] Based on the above, the method for fine-tuning training to obtain an abnormal risk behavior recognition model LLM based on a large language model is as follows: Based on the multimodal coding method in step S4, collect and construct sufficient training data, and perform fine-tuning training on the large language model; The fine-tuning strategy is to use a pre-trained large language model as a base, combine the fused features as additional input with text prompts, and input them into the LLM; Based on a specific risk identification task, a fine-tuned objective function is designed; large-scale computing resources are used to iteratively update the model parameters, enabling the model to learn the feature representations and decision-making patterns of multi-source risk abnormal behaviors. The fine-tuning process includes: Initialization: Load the pre-trained LLM parameters and initialize the fusion feature input layer and task-specific output layer; Data loading: Multimodal training data is loaded in batches, with each batch containing a certain number of positive and negative samples; Forward propagation: Input the fused features and text prompts into the model and calculate the output; Loss calculation: Calculate the loss value according to the task type and evaluate the difference between the model prediction and the true label; Backpropagation: Calculates gradients, updates model parameters, and optimizes the direction to minimize the loss function; Iterative loop: Repeat the above steps until the model achieves satisfactory performance or converges on the validation set.

[0013] Based on the above, the prompt template designed using the prompting learning method is as follows: "Please determine whether any abnormal behavior of the [risk behavior type] exists based on the following information. Relevant information includes: [text description], image content [key visual feature description], and audio recording [speech content summary]. Please output your judgment and provide the basis for your assessment." The fused multimodal features are transformed into natural language descriptions or feature vectors and superimposed on a specified position in the prompt template to form a complete input prompt, which is then input into the fine-tuned LLM. LLM generates corresponding outputs or decisions based on the prompts.

[0014] A second aspect of the present invention provides a risk abnormal behavior event identification system based on data fusion analysis, comprising: The abnormal behavior tag data acquisition module is used to acquire abnormal behavior tag data; the acquired abnormal behavior tag data includes abnormal behavior tag data from the communication side, abnormal behavior tag data from the Internet side, abnormal behavior tag data from the funding side, abnormal behavior tag data from cloud video conferencing, as well as case and police data. The knowledge graph ontology structure construction module is used to abstract the entities and relationships involved in the abnormal behavior tag data, define the entities, attributes and relationships of the knowledge graph, and construct the knowledge graph ontology structure. The entity relationship extraction module is used to extract the entities, attributes, and relationships involved in the structured data of the acquired abnormal behavior label data into the constructed knowledge graph ontology structure based on the defined entities, attributes, and relationships; where all entities have a globally unique identifier (vid); The multimodal encoding module is used to encode the graph data using the graph attention network GAT for the constructed knowledge graph ontology structure to obtain the corresponding modal features; It is also used for structured data of knowledge graph ontology structure, and corresponding feature encoding methods are used to encode each source to obtain corresponding modal features; The alignment module is used to map the modal features obtained after encoding by different encoding methods in the multimodal encoding module to the same vector space, and to align the modal features to ensure the matching of multimodal features in terms of size and semantics, thereby realizing the fusion of multimodal features; The cognitive module is used to fine-tune the training of an abnormal risk behavior recognition model (LLM) based on a large language model and the fused features, using the large language model as the foundation. The output module is used to overlay the fused multimodal features onto a specified input and feed it to the LLM, thereby guiding the LLM to generate corresponding outputs or decisions based on the prompts of the specified input; wherein, the specified input is a prompt template designed using a prompt learning method.

[0015] A third aspect of the present invention provides an electronic device, comprising: One or more processors; Storage device for storing one or more programs. When the one or more programs are executed by the one or more processors, the one or more processors implement the risk abnormal behavior event identification method based on multi-source risk abnormal behavior data fusion analysis as described above.

[0016] A fourth aspect of the present invention provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the risk abnormal behavior event identification method based on multi-source risk abnormal behavior data fusion analysis as described above.

[0017] The fifth aspect of the present invention provides a computer program product, including a computer program / instruction, which, when executed by a processor, implements the risk abnormal behavior event identification method based on multi-source risk abnormal behavior data fusion analysis as described above.

[0018] This invention has outstanding substantive features and significant progress compared to the prior art, specifically: This invention provides a method for constructing a knowledge graph by integrating risky and abnormal behavior data from various industries. The method associates and maps the implementation elements (mobile phone number, WeChat ID, ID card, fund account, website, transaction, cloud video conferencing, etc.) and operational rules involved in risky and abnormal behavior events with entities and relationships in the knowledge graph.

[0019] This invention, based on multimodal coding and feature alignment, also provides a fusion analysis method based on "multi-source risk abnormal behavior label" data. This method overcomes the limitations of local features in risk abnormal behavior events by fusing abnormal label features from multiple stages. It comprehensively considers the temporal correlation, logical coupling, and co-evolutionary relationships between behavioral features across various dimensions, thus characterizing the user's abnormal behavior trajectory from a holistic perspective. Through joint modeling of abnormal label features from multiple stages, it can effectively identify potential risk signals hidden within complex behavioral patterns.

[0020] The feature alignment method of this invention differs from traditional feature extraction and data statistical mining methods, fully considering the heterogeneity and semantic consistency of features across different modalities. Through multimodal feature mapping and alignment techniques, the encoded features of each modality are mapped to a unified semantic space. MACAW-LLM is used to adjust the dimensions of each modal feature, and an attention mechanism is employed to calculate the correlation weights between features of different modalities, achieving dynamic feature alignment and fusion. This combination of encoding method and alignment can maximize the fusion of multiple encoded information and fully explore the potential correlations and complementary information in the data from different modalities. Attached Figure Description

[0021] The above and / or additional aspects and advantages of the present invention will become apparent and readily understood from the description of the embodiments taken in conjunction with the following drawings, in which: Figure 1 The process for identifying abnormal data risk events in a single industry sector is illustrated.

[0022] Figure 2 A flowchart of the method of the present invention is shown.

[0023] Figure 3 A structural diagram of the method of the present invention is shown.

[0024] Figure 4 The knowledge graph ontology structure constructed in this invention is shown.

[0025] Figure 5 A schematic diagram of video encoding in the method of the present invention is shown.

[0026] Figure 6 A schematic diagram of voiceprint encoding in the method of the present invention is shown. Detailed Implementation

[0027] To better understand the above-mentioned objectives, features, and advantages of the present invention, the present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments. It should be noted that, unless otherwise specified, the embodiments and features described in these embodiments can be combined with each other.

[0028] Many specific details are set forth in the following description in order to provide a full understanding of the invention. However, the invention may also be practiced in other ways different from those described herein, and therefore the scope of protection of the invention is not limited to the specific embodiments disclosed below.

[0029] Example 1 like Figure 2 and Figure 3 As shown in the figure, this embodiment proposes a method for identifying risky abnormal behavior events based on the fusion analysis of multi-source risky abnormal behavior data, including: S1: Set up the software and database environment.

[0030] The model and data processing provided by this invention rely on the Nebulagraph graph database and big data components such as Hadoop, Hive, Spark, and Flink. Among them, big data components such as Hadoop, Hive, Spark, and Flink are mainly used for data preprocessing to connect the processed entities and relationships into the graph database.

[0031] S2: Obtain abnormal behavior label data.

[0032] The abnormal behavior labeling data primarily originates from abnormal behavior data with accompanying labels identified across various industry sectors. This includes: Communication-side data includes voice, SMS, and call signaling data; abnormal behavior tag data from the communication side mainly includes: abnormal calling frequency, abnormal called number dispersion, abnormal called number attribution dispersion, abnormal call time pattern, abnormal calling number empty number rate, abnormal called hang-up rate, and abnormal call success rate. The internet side includes apps, websites, IPs, QQ, WeChat, etc.; abnormal behavior tag data from the internet side mainly includes: abnormal behavior websites, apps downloaded through abnormal channels, apps that illegally obtain personal information, impersonation tactics, joining group chats, order brushing, and abnormal follower growth. Funding-side data includes user fund account data, transaction data, and account information data; abnormal behavior tag data from the funding side mainly includes: small-amount transfers, password changes, silent accounts, malicious overdrafts, large-amount cash withdrawals, and other abnormal behavior tag data. Cloud video conferencing includes Zoom Meeting, Tencent Meeting, and Zoom Meeting; abnormal behavior tag data from cloud video conferencing mainly includes: abnormal voiceprint features, abnormal sharing, abnormal login accounts, voice changing, and other abnormal behavior tag data. In addition, case and police data can be used to determine whether the data that has been warned can be linked to cases or police incidents, and whether abnormal behavior events have been reported to the police or have become cases.

[0033] S3: Construct the ontology structure of the knowledge graph.

[0034] The entities and relationships involved in the abnormal behavior tag data are abstracted, and the entities, attributes, and relationships of the knowledge graph are defined to construct the knowledge graph ontology structure. The constructed knowledge graph ontology structure is as follows: Figure 4 As shown.

[0035] Define the following entities: Mobile Number Entity Table, Name Entity Table, Meeting Entity Table, Bank Account Entity Table, Account Transaction Entity Table, Case Entity Table, IP Entity Table, APP Entity Table, IMEI Entity Table, MAC Entity Table, Voice Entity Table; Define attributes: describe information about the entity, such as meeting attributes (duration, number of participants, whether the meeting shares the screen, etc.), and app attributes (app package name, app type, etc.). Define relationships: Define the relationships between entities, such as the relationship between a bank account and a transaction being a related account.

[0036] S4: Graph entity relation extraction; The abnormal behavior label data obtained in S2 is extracted into the constructed knowledge graph ontology structure based on the entities, attributes, and relationships defined in S3 using ETL. Each entity has a globally unique identifier (vid), which is generated as follows: The strings vertexType and name are concatenated, and their hexadecimal hash value is calculated using the MD5 algorithm. The result is then assigned to the variable vid. vid=MD5Hex(vertexType⊕_⊕name) Where: ⊕ represents a string concatenation operation.

[0037] S5: Multimodal coding construction; Graph Encoding: For the knowledge graph ontology structure constructed by S4, the graph attention network (GAT) is used to encode the graph data to obtain graph modal features.

[0038] The GAT network introduces an attention mechanism, enabling each node to dynamically adjust its weights based on the importance of its neighbors. The GAT network is well-suited for inductive learning and is well-suited for risk-anomalous behavior graph structures. It can fully mine and represent the graph information, including communication, application, financial, and meeting information, within the knowledge graph ontology structure.

[0039] The method for encoding graph data using a graph attention network (GAT) to obtain graph modal features is as follows: Step 1: Calculate the similarity coefficient between each entity in the knowledge graph ontology structure and itself: e ij = a ([ Wh i ||Wh j ]), j ∈ N i h i For the i-th entity, h j Let be the j-th entity, a be a single-layer feedforward neural network, and N be...i Let be the set of neighbors of node i. The linear mapping with shared parameter W is used to increase the dimensionality of the features of the knowledge graph ontology structure. || is used to concatenate the transformed features of graph vertices i and j. Finally, the concatenated high-dimensional features are mapped to a real number. Step 2: Normalize to obtain the attention coefficient;

[0040] a ij Let be the normalized attention weight (importance score) of node j to node i, with a value range of [0, 1], and all neighbors. ; e ij The raw attention score (unnormalized) between nodes i and j; N i Let i be the set of neighboring nodes of node i. In the knowledge graph, there may be neighbors with multiple relationship types. The exp(·) exponential function is used to amplify differences and strengthen the weight of important neighbors; LeakyReLU is a modified linear unit with leakage (the slope is usually taken as 0.2), and its formula is: LeakyReLU(x) = max(0.2x, x). It retains negative value information to avoid gradient vanishing. Step 3: Through a multi-head attention mechanism, the features of each node that incorporate neighborhood information are obtained:

[0041] For nodes The output feature (a new representation after fusing neighbor information) has a dimension of K×d' (d' is the single-head output dimension). K represents the number of heads in the multi-head attention mechanism, where each head independently calculates attention and captures features from different subspaces. || is the vector concatenation operation, which concatenates the outputs of K heads into the final feature; W k Let be the trainable weight matrix for the k-th head, with each head having an independent weight matrix. To achieve feature space projection; The attention weight of node j to i is calculated for the k-th head. Different heads may pay attention to different neighbors (e.g., head 1 pays attention to semantic similarity, head 2 pays attention to structural similarity). h j is the input feature vector of node j, the initial feature, or the output of the previous layer.

[0042] Structured data encoding: Data from the communication side, the internet side, and the funding side are all structured data. For structured data, the corresponding modal features can be obtained by using feature encoding methods that are suitable for the characteristics of the data.

[0043] For structured data in the abnormal behavior label data on the communication side, feature extraction is performed using the following method: Abnormal calling frequency: Calculate the number of calls per unit time (e.g., per day or per hour) and standardize these values ​​or convert them to relative values ​​using quantiles; Abnormal dispersion of called numbers: This is measured by calculating the diversity of called numbers, for example, by using indicators such as entropy or Gini impurity. Abnormal dispersion of called number attribution: Measured by calculating the diversity of the regions to which the called number belongs; Abnormal call timing patterns: Analyze the distribution of call times throughout the day to identify call patterns during atypical time periods; time series analysis methods can be used. Abnormal caller ID non-existent rate: Calculate the proportion of invalid numbers dialed; Abnormal call hang-up rate: Statistics on the probability that the called party hangs up the phone; Call success rate anomaly: Calculate the proportion of successfully connected calls out of all attempted calls; The numerical features obtained above are encoded using Z-score or quantile discretization + Embedding layer to obtain the modal features on the communication side.

[0044] For structured data in abnormal behavior tag data on the Internet side, text data such as abnormal behavior URLs, abnormal channel app downloads, URLs, wording, joining group chats, order brushing, and follower growth are converted into vector representations using natural language processing technology. The semantic features of the text are then encoded in a high-order manner to obtain the modal features on the Internet side.

[0045] For structured data in the abnormal behavior tag data on the funding side, the amount and operation type are directly used as numerical features, and at the same time, features (such as transfer frequency, amount distribution, etc.) are constructed based on the user's behavior pattern for encoding. Dormant accounts are marked as accounts that have not been active for a long time, and their trading behavior is analyzed after they are reactivated. Malicious overdrafts: Track changes in the ratio between account balance and credit limit, paying particular attention to the occurrence of overdraft behavior; Large cash withdrawals: Records large cash withdrawal events; Small transfers, record small transfer events; Change the password and record the password change event; The data collection includes, but is not limited to, the above features. Binning encoding is performed (continuous amounts can use equal-frequency / equal-width binning + business rule binning, employing one-hot encoding / label encoding; for example, transfer amount binning: [0,100), [100,1k), [1k,10k), ≥10k; frequency features can use dynamic threshold binning (based on standard deviation), employing ordinal encoding, for example, transfer frequency: low frequency (<μ-σ), medium frequency, high frequency (>μ+σ); proportion features can use business key point binning, employing binary encoding, for example, overdraft proportion: safe (<80%), risky (≥80%)); behavior sequence encoding (transaction operation type sequences can use Skip-gram / W2V embedding, employing fixed-length vectors (e.g., 64-dimensional); password modification time sequences can use time difference encoding (Δt1,Δt2,...), employing variable-length sequences → Padding), and then fusion (directly concatenating binning encoding + sequence mean vector).

[0046] The structured data in the abnormal behavior tagging data of cloud video conferencing includes video conferencing data and conference voiceprint data.

[0047] Video conferencing data encoding: The ViViT model is used to encode the video conferencing data to obtain the modal features of the video conferencing, such as... Figure 5 As shown; The ViViT model is an extension of the Vision Transformer (ViT) to the video domain. It treats video sequences as a series of "spatiotemporal blocks" and learns feature representations of the video through spatial and temporal self-attention mechanisms. ViViT can simultaneously capture spatial information within video frames and temporal information between frames, thereby better understanding video content with potentially risky or abnormal behavior.

[0048] By analyzing the motion density and probability of important events in the video stream, the number of frames captured per second is dynamically adjusted. For example, the sampling frequency is increased when high-risk behavioral patterns (such as abnormal eye contact or gestures) are detected, while unnecessary computational burden is reduced in low-risk scenarios. This transformer-based motion recognition module pre-assesses the importance of each frame. Based on the assessment results, the frame sequence input to the ViViT model is adjusted in real time to ensure that important information is fully expressed without wasting resources on non-critical parts. An attention mechanism is combined to optimize the feature extraction process, further enhancing the focus on key details in videos with potentially risky or abnormal behavior.

[0049] Conference audioprint data encoding: The conference audioprint data is encoded using the ECAPA-TDNN model to obtain the modal features of the conference audioprint, such as... Figure 6 As shown.

[0050] Based on the ECAPA-TDNN model, robustness to different environmental noise and channel variations is improved by enhancing the channel and perturbation attention mechanism. This ensures accurate user identification even in complex environments (such as noisy conferences, background phone calls, or communication between different devices), thereby improving the reliability and accuracy of identifying abnormal risk behaviors.

[0051] The integration of voiceprint coding into a multimodal big data model allows for the identification of risky and abnormal behavior that goes beyond static voiceprint matching. It can also be combined with a user's daily calling habits, frequently used vocabulary, tone changes, and other behavioral patterns for a comprehensive assessment. For example, if an account suddenly exhibits abnormal calling patterns (such as making a large number of calls to unknown numbers in a short period of time), it will trigger further attention to risky and abnormal behavior.

[0052] S6: Feature alignment; Modal features obtained by encoding different encoding methods in S5 are mapped to the same vector space, and the modal features are aligned to ensure the matching of multimodal features in terms of size and semantics, thereby achieving the fusion of multimodal features.

[0053] Encoders for each modality are typically trained separately, which can lead to potential differences in the representations generated by different encoders. Therefore, these independent representations are aligned in a joint space. By mapping inputs from different modalities to the same vector space and aligning the modal vectors, the matching of multimodal vectors in terms of size and semantics is ensured, thereby achieving the fusion of multimodal information. The core implementation is based on Multi-Head Self-Attention (MHSA) in the Transformer architecture. To match the feature dimensions of various modalities, the extracted feature sequences are processed as follows: Reduce the feature sequence using a one-dimensional convolutional layer, for example: ; Among them, h i For text modality, h v For the visual modality, h a For audio modality; Then, a linear layer is used to adjust the dimensions of each modality feature to match the dimensions of the text features output by the LLM, that is:

[0054] Make ; Finally, CrossAttention alignment is used, with the dimension-adjusted visual, audio, and other modal features as the query for CrossAttention, and the text features as the key and value, for feature alignment. The alignment formula is:

[0055] in, It can be and Any one of them, It is a feature after alignment.

[0056] Feature fusion and input LLM concatenates aligned visual, audio, and other modal features with textual features as input to the LLM. For example:

[0057] Its Embed(x) t ) is an embedded representation of text input.

[0058] Cross-attention is used to treat one modality as the query and the other as the key and value, calculating their similarity and generating new feature representations. This approach dynamically adjusts the importance ratio between different modalities based on task requirements, thus better reflecting the actual risk situation. The feature vectors processed by cross-attention are concatenated to form the final fused features. Attention weights are applied to the fused features to ensure that features crucial for identifying risky and abnormal behavior play a greater role in subsequent classification or prediction.

[0059] In some exemplary embodiments, a target function encompassing multiple tasks is constructed through multi-task learning, such as simultaneously performing text classification, image recognition, and joint judgment of multimodal data to determine the existence of danger. This allows the Transformer model to fully utilize the characteristics and interrelationships of different modal features during the learning process, improving its understanding and comprehensive judgment capabilities regarding multimodal data. Then, comparative learning is used to train the model by constructing positive and negative sample pairs. For example, text descriptions and corresponding images of the same dangerous scene form positive sample pairs, while unrelated text and images form negative sample pairs. The Transformer model learns to make positive sample pairs closer in semantic space and negative sample pairs further apart, thereby enhancing its semantic understanding and association judgment capabilities regarding different modal data. Finally, self-supervised learning is used to perform predefined masking operations on the multimodal data, such as masking certain keywords in text or obscuring certain areas in images, allowing the Transformer model to predict the masked content based on the unmasked portions. This approach enables the Transformer model to better learn the intrinsic connections and semantic information between different modal data, and gradually master how to align and fuse these multimodal data to complete the task during the training process.

[0060] S7: Construct the recognition model; Based on a large language model, and using the fused features, a large language model-based abnormal risk behavior recognition model (LLM) is obtained through fine-tuning and training.

[0061] Based on the aforementioned multimodal coding method, sufficient training data is collected and constructed, and fine-tuned training is performed on large-scale computing resources to obtain a large-scale model suitable for multi-source risk and abnormal behavior recognition scenarios. The obtained large-scale model can understand and generate natural language, while integrating information from other modalities (such as vision and audio), and then combining it with custom instructions to enable it to perform specific tasks effectively. In the application phase, a cue learning method is used to design specific cue templates, and the fused multimodal data is superimposed on specified inputs to the LLM, thereby guiding the LLM to generate corresponding outputs or decisions based on the cue.

[0062] The fine-tuning strategy is based on a pre-trained large language model, using fused features as additional input, combined with text prompts, and then input into the LLM. For specific risk identification tasks, fine-tuning objective functions are designed; for example, cross-entropy loss is used for classification tasks, and language model loss is used for generation tasks. Large-scale computing resources are utilized to iteratively update the model parameters, enabling the model to learn feature representations and decision patterns of multi-source risky and abnormal behaviors.

[0063] The fine-tuning process includes: Initialization: Load the pre-trained LLM parameters and initialize the fusion feature input layer and task-specific output layer; Data loading: Multimodal training data is loaded in batches, with each batch containing a certain number of positive and negative samples; Forward propagation: Input the fused features and text prompts into the model and calculate the output; Loss calculation: Calculate the loss value according to the task type and evaluate the difference between the model prediction and the true label; Backpropagation: Calculates gradients, updates model parameters, and optimizes the direction to minimize the loss function; Iterative loop: Repeat the above steps until the model achieves satisfactory performance or converges on the validation set.

[0064] S8: Identification of Risky and Abnormal Behavior Events; The fused multimodal features are superimposed on a specified input and fed into the LLM, thereby guiding the LLM to generate corresponding outputs or decisions based on the specified input prompts; where the specified input is a prompt template designed using a prompt learning method.

[0065] During the application phase, carefully designed prompt templates are used to guide the model to output results that meet expectations. For example, the following prompt template can be used for risk behavior detection tasks: "Based on the following information, determine whether any abnormal behavior of the [risk behavior type] exists. Relevant information includes: [text description], image content [key visual feature description], and audio recording [speech content summary]. Please output your judgment and provide the basis for your assessment." The fused multimodal features are transformed into natural language descriptions or feature vectors, which are then superimposed onto designated positions in the prompt template to form complete input prompts, which are then fed into the fine-tuned LLM. The model generates corresponding outputs or decisions based on the prompt content, such as the probability of risky behavior and category labels.

[0066] The method of the present invention has the following beneficial effects: 1. Data fusion improves detection accuracy Multi-source data integration: Traditional methods rely on data from a single industry, leading to incomplete detection. This method integrates data from multiple sources, including communications, finance, and social media, to construct a knowledge graph that comprehensively reflects the multi-stage characteristics of risky and abnormal behaviors, significantly improving detection accuracy and recall.

[0067] Cross-industry analysis: Risky and abnormal activities often occur across industries, and a single data source cannot capture all clues. Multi-source data fusion enables the identification of complex cross-industry chains of risky and abnormal activities, improving overall detection capabilities.

[0068] 2. Large model technology enhances recognition capabilities Risk Pattern Recognition: A multimodal coding module integrating multi-source data provides a unified representation of heterogeneous risk label data from multiple channels, including communication, internet, and funding sources, extracting high-order feature vectors with semantic consistency and discriminative capabilities. Subsequently, feature alignment is used to perform cross-modal mapping and alignment of feature spaces from different sources, eliminating semantic biases caused by differences in data distribution and improving the fusionability of multi-source information. Large-scale multimodal technology enables cross-modal feature extraction to deeply mine the intrinsic characteristics of the data, and knowledge graph association analysis effectively integrates knowledge and reasoning capabilities. This allows for the extraction of subtle signs of abnormal risk behavior from massive amounts of data, prediction of the next steps in abnormal risk activities, and provision of detailed chain information for each stage of abnormal risk events, enabling earlier warnings and interventions.

[0069] 3. Knowledge graph construction deepens understanding Entity and Relationship Mapping: This method maps the elements and operational patterns of risky and abnormal behaviors into entities and relationships in a knowledge graph, forming a deeper understanding of risky and abnormal behavior activities. It not only identifies known risky and abnormal behaviors but also predicts unknown patterns. At the same time, it constrains the reasoning process of large models through the knowledge graph.

[0070] Risk Abnormal Behavior Chain: Provides detailed information on risk abnormal behavior chains to help relevant organizations more effectively track and handle groups with risk abnormal behavior.

[0071] Example 2 This embodiment provides a risk abnormal behavior event identification system based on data fusion analysis, including: The abnormal behavior tag data acquisition module is used to acquire abnormal behavior tag data; the acquired abnormal behavior tag data includes abnormal behavior tag data from the communication side, abnormal behavior tag data from the Internet side, abnormal behavior tag data from the funding side, abnormal behavior tag data from cloud video conferencing, as well as case and police data. The knowledge graph ontology structure construction module is used to abstract the entities and relationships involved in the abnormal behavior tag data, define the entities, attributes and relationships of the knowledge graph, and construct the knowledge graph ontology structure. The entity relationship extraction module is used to extract the entities, attributes, and relationships involved in the structured data of the acquired abnormal behavior label data into the constructed knowledge graph ontology structure based on the defined entities, attributes, and relationships; where all entities have a globally unique identifier (vid); The multimodal encoding module is used to encode the graph data using the graph attention network GAT for the constructed knowledge graph ontology structure to obtain the corresponding modal features; It is also used for structured data of knowledge graph ontology structure, and corresponding feature encoding methods are used to encode each source to obtain corresponding modal features; The alignment module is used to map the modal features obtained after encoding by different encoding methods in the multimodal encoding module to the same vector space, and to align the modal features to ensure the matching of multimodal features in terms of size and semantics, thereby realizing the fusion of multimodal features; The cognitive module is used to fine-tune the training of an abnormal risk behavior recognition model (LLM) based on a large language model and the fused features, using the large language model as the foundation. The output module is used to overlay the fused multimodal features onto a specified input and feed it to the LLM, thereby guiding the LLM to generate corresponding outputs or decisions based on the prompts of the specified input; wherein, the specified input is a prompt template designed using a prompt learning method.

[0072] For a detailed implementation method of this embodiment, please refer to "A Method for Identifying Risk Abnormal Behavior Events Based on Multi-Source Risk Abnormal Behavior Data Fusion Analysis", which will not be elaborated here.

[0073] Example 3 This embodiment provides an electronic device, including: One or more processors; Storage device for storing one or more programs. When the one or more programs are executed by the one or more processors, the one or more processors implement the risk abnormal behavior event identification method based on multi-source risk abnormal behavior data fusion analysis as described in Embodiment 1.

[0074] Example 4 This embodiment provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the risk abnormal behavior event identification method based on multi-source risk abnormal behavior data fusion analysis as described in Embodiment 1.

[0075] Example 5 This embodiment provides a computer program product, including a computer program / instruction, characterized in that, when the computer program / instruction is executed by a processor, it implements the risk abnormal behavior event identification method based on multi-source risk abnormal behavior data fusion analysis as described in Embodiment 1.

[0076] The various embodiments in this specification are described in a progressive manner, with each embodiment focusing on the differences from other embodiments. The same or similar parts between the various embodiments can be referred to each other.

[0077] Those skilled in the art will understand that embodiments of this application can be provided as methods, systems, media, or program products. Therefore, embodiments of this application can take the form of entirely hardware embodiments, entirely software embodiments, or embodiments combining software and hardware aspects. Furthermore, embodiments of this application can take the form of program products implemented on one or more usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing usable program code.

[0078] This application describes embodiments of methods, systems, devices, storage media, and program products according to embodiments of this application with reference to flowchart illustrations and / or block diagrams. It should be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by program instructions. These program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing terminal device to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing terminal device, generate instructions for implementing the flowchart illustrations. Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.

[0079] The methods, systems, devices, and media provided in this application have been described in detail above. Specific examples have been used to illustrate the principles and implementation methods of this application. The descriptions of the above embodiments are only for the purpose of helping to understand the methods and core ideas of this application. At the same time, for those skilled in the art, there will be changes in the specific implementation methods and application scope based on the ideas of this application. Therefore, the content of this specification should not be construed as a limitation of this application.

Claims

1. A method for identifying risky abnormal behavior events based on the fusion analysis of multi-source risky abnormal behavior data, characterized in that, include: S1: Obtain abnormal behavior tag data; the abnormal behavior tag data includes abnormal behavior tag data from the communication side, abnormal behavior tag data from the Internet side, abnormal behavior tag data from the funding side, abnormal behavior tag data from cloud video conferencing, and case and police data. S2: Abstract the entities and relationships involved in the abnormal behavior tag data, define the entities, attributes and relationships of the knowledge graph, and construct the knowledge graph ontology structure; S3: Based on the defined entities, attributes, and relationships, the acquired abnormal behavior label data is extracted into the constructed knowledge graph ontology structure through ETL. All entities have a globally unique identifier (vid). S4: For the constructed knowledge graph ontology structure, use the Graph Attention Network (GAT) to encode the graph data and obtain the corresponding modal features; For structured data of knowledge graph ontology, corresponding feature encoding methods are used for each source to obtain corresponding modal features; S5: Map the modal features obtained after encoding by different encoding methods in S4 to the same vector space, and align the modal features to ensure the matching of multimodal features in size and semantics, thereby achieving the fusion of multimodal features; S6: Based on the large language model, and using the fused features, fine-tuning training yields an abnormal risk behavior recognition model LLM based on the large language model; S7: The fused multimodal features are superimposed on the specified input and fed to the LLM, thereby guiding the LLM to generate corresponding outputs or decisions based on the prompts of the specified input; wherein, the specified input is a prompt template designed using the prompt learning method.

2. The method for identifying risky abnormal behavior events based on multi-source risky abnormal behavior data fusion analysis according to claim 1, characterized in that, In step S1: Abnormal behavior tag data from the communication side includes: abnormal calling frequency, abnormal called number dispersion, abnormal called number attribution dispersion, abnormal calling time pattern, abnormal calling number empty number rate, abnormal called hang-up rate, and abnormal call success rate. Abnormal behavior tag data from the Internet side includes: abnormal behavior URLs, abnormal channel app downloads, apps that illegally obtain personal information, impersonation scripts, joining group chats, order brushing, and gaining followers; Abnormal behavior tags from the funding side include: small transfers, password changes, dormant accounts, malicious overdrafts, and large cash withdrawals; Abnormal behavior tags from cloud video conferencing include: abnormal voiceprint features, abnormal sharing, abnormal login accounts, and voice alteration.

3. The method for identifying risky abnormal behavior events based on multi-source risky abnormal behavior data fusion analysis according to claim 2, characterized in that, For structured data of knowledge graph ontology, corresponding feature encoding methods are used for each source to obtain the corresponding modal features. For structured data in the abnormal behavior label data on the communication side, feature extraction is performed using the following method: Abnormal calling frequency: Calculate the number of calls per unit time and standardize these values ​​or convert them to relative values ​​using quantiles; Abnormal dispersion of called numbers: This is measured by calculating the diversity of called numbers; Abnormal dispersion of called number attribution: Measured by calculating the diversity of the regions to which the called number belongs; Abnormal call timing patterns: Analyze the distribution of call times throughout the day to identify call patterns during atypical time periods; Abnormal caller ID non-existent rate: Calculate the proportion of invalid numbers dialed; Abnormal call hang-up rate: Statistics on the probability that the called party hangs up the phone; Call success rate anomaly: Calculate the proportion of successfully connected calls out of all attempted calls; The numerical features obtained above are encoded using Z-score or quantile discretization + Embedding layer to obtain the modal features on the communication side. For the structured data in the abnormal behavior tag data on the Internet side, natural language processing technology is used to convert it into vector representation, that is, to obtain the modal features on the Internet side; For the structured data in the abnormal behavior tag data on the funding side, the amount and operation type are directly used as numerical features, and at the same time, the features are constructed and encoded based on the user's behavior pattern. Dormant accounts are marked as accounts that have not been active for a long time, and their trading behavior is analyzed after they are reactivated. Malicious overdrafts: Track changes in the ratio between account balance and credit limit, paying particular attention to the occurrence of overdraft behavior; Large cash withdrawals: Records large cash withdrawal events; Collect features including but not limited to the above, perform binning encoding and behavior sequence encoding, and then fuse them; The structured data in the abnormal behavior tagging data of cloud video conferencing includes video conferencing data and conference voiceprint data; Video conferencing data encoding: The ViViT model is used to encode the video conferencing data to obtain the modal features of the video conferencing; Conference voiceprint data encoding: The conference voiceprint data is encoded using the ECAPA-TDNN model to obtain the modal features of the conference voiceprint.

4. The method for identifying risky abnormal behavior events based on multi-source risky abnormal behavior data fusion analysis according to claim 1, characterized in that, The method to achieve feature alignment across modalities is as follows: Based on the Multi-Head Self-Attention implementation in the Transformer architecture, in order to match the feature dimensions of different modal features, the different modal features extracted in S4 are processed as follows: One-dimensional convolutional layers are used to reduce the feature dimension of modal features respectively; Then, a linear layer is used to adjust the feature dimensions of each modality feature to be consistent with the text feature dimensions output by the LLM; Finally, Cross Attention is used to align the visual and audio features from each modal feature after dimension adjustment as the query for Cross Attention, and the text features as the key and value. The aligned modal features are concatenated with the text features to complete the fusion of multimodal features.

5. The method for identifying risky abnormal behavior events based on multi-source risky abnormal behavior data fusion analysis according to claim 1, characterized in that, The method for fine-tuning training to obtain an abnormal risk behavior identification model (LLM) based on a large language model is as follows: Based on the multimodal coding method in step S4, collect and construct sufficient training data, and perform fine-tuning training on the large language model; The fine-tuning strategy is to use a pre-trained large language model as a base, combine the fused features as additional input with text prompts, and input them into the LLM; Based on a specific risk identification task, a fine-tuned objective function is designed; large-scale computing resources are used to iteratively update the model parameters, enabling the model to learn the feature representations and decision-making patterns of multi-source risk abnormal behaviors. The fine-tuning process includes: Initialization: Load the pre-trained LLM parameters and initialize the fusion feature input layer and task-specific output layer; Data loading: Multimodal training data is loaded in batches, with each batch containing a certain number of positive and negative samples; Forward propagation: Input the fused features and text prompts into the model and calculate the output; Loss calculation: Calculate the loss value according to the task type and evaluate the difference between the model prediction and the true label; Backpropagation: Calculates gradients, updates model parameters, and optimizes the direction to minimize the loss function; Iterative loop: Repeat the above steps until the model achieves satisfactory performance or converges on the validation set.

6. The method for identifying risky abnormal behavior events based on multi-source risky abnormal behavior data fusion analysis according to claim 1, characterized in that, The prompt template designed using the prompting learning method is as follows: The fused multimodal features are transformed into natural language descriptions or feature vectors and superimposed on a specified position in the prompt template to form a complete input prompt, which is then input into the fine-tuned LLM. LLM generates corresponding outputs or decisions based on the prompts.

7. A risk abnormal behavior event identification system based on data fusion analysis, characterized in that, include: The abnormal behavior tag data acquisition module is used to acquire abnormal behavior tag data; the acquired abnormal behavior tag data includes abnormal behavior tag data from the communication side, abnormal behavior tag data from the Internet side, abnormal behavior tag data from the funding side, abnormal behavior tag data from cloud video conferencing, as well as case and police data. The knowledge graph ontology structure construction module is used to abstract the entities and relationships involved in the abnormal behavior tag data, define the entities, attributes and relationships of the knowledge graph, and construct the knowledge graph ontology structure. The entity relationship extraction module is used to extract the entities, attributes, and relationships involved in the structured data of the acquired abnormal behavior label data into the constructed knowledge graph ontology structure based on the defined entities, attributes, and relationships; where all entities have a globally unique identifier (vid); The multimodal encoding module is used to encode the graph data using the graph attention network GAT for the constructed knowledge graph ontology structure to obtain the corresponding modal features; It is also used for structured data of knowledge graph ontology structure, and corresponding feature encoding methods are used to encode each source to obtain corresponding modal features; The alignment module is used to map the modal features obtained after encoding by different encoding methods in the multimodal encoding module to the same vector space, and to align the modal features to ensure the matching of multimodal features in terms of size and semantics, thereby realizing the fusion of multimodal features; The cognitive module is used to fine-tune the training of an abnormal risk behavior recognition model (LLM) based on a large language model and the fused features, using the large language model as the foundation. The output module is used to overlay the fused multimodal features onto a specified input and feed it to the LLM, thereby guiding the LLM to generate corresponding outputs or decisions based on the prompts of the specified input; wherein, the specified input is a prompt template designed using a prompt learning method.

8. An electronic device, characterized in that, include: One or more processors; Storage device for storing one or more programs. When the one or more programs are executed by the one or more processors, the one or more processors implement the risk abnormal behavior event identification method based on multi-source risk abnormal behavior data fusion analysis as described in any one of claims 1-6.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that, When executed by the processor, the program implements the risk abnormal behavior event identification method based on multi-source risk abnormal behavior data fusion analysis as described in any one of claims 1-6.

10. A computer program product comprising a computer program / instructions, characterized in that, When the computer program / instruction is executed by the processor, it implements the risk abnormal behavior event identification method based on multi-source risk abnormal behavior data fusion analysis as described in any one of claims 1-6.

Citation Information

Cited By

  • Method and system for monitoring installation risk of partition board based on video

    CN121147857A

  • A method and system for installation risk based on video monitoring partition

    CN121147857B

  • Intelligent decision generation method and device for water supply project

    CN121365322A

  • Intelligent decision generation method and device for water supply projects

    CN121365322B

  • Method, device and equipment for analyzing SIP behavior based on real-time flow and medium

    CN121530751A