Campus fraud identification method based on multi-source data fusion and time sequence knowledge graph

The campus fraud identification method based on multi-source data fusion and temporal knowledge graph solves the problems of insufficient data integration and temporal reasoning in existing campus fraud identification methods, and achieves accurate identification and early warning of fraudulent behavior, thereby improving defense capabilities.

CN121579979APending Publication Date: 2026-02-27GUIZHOU UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511744532.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-11-25
Publication Date
2026-02-27

AI Technical Summary

Technical Problem

Existing methods for identifying campus fraud rely on monitoring single-source data, making it difficult to build a panoramic risk view and effectively integrate multi-source heterogeneous data. Static knowledge graphs cannot characterize the dynamic evolution of entity relationships in the campus environment and lack temporal reasoning capabilities, resulting in lagging defense.

Method used

By employing multi-source data fusion and temporal knowledge graph methods, this study collects heterogeneous data from multiple sources within the campus information ecosystem, performs data cleaning and preprocessing, constructs a dynamic knowledge graph with temporal awareness capabilities, designs a temporally enhanced knowledge graph pattern, and utilizes a temporal graph neural network model for risk identification, thereby enabling dynamic analysis and early warning of fraudulent activities.

Benefits of technology

It enables accurate identification and early detection of campus fraud, improves the timeliness of early warning and proactive defense, solves the problem of data silos, and can capture the gradual characteristics of fraud and the risk transmission path.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121579979A_ABST
    Figure CN121579979A_ABST
Patent Text Reader

Abstract

The invention provides a campus fraud identification method based on multi-source data fusion and a time sequence knowledge graph, and relates to the field of information security and artificial intelligence. The campus fraud identification method based on the multi-source data fusion and the time sequence knowledge graph comprises the following steps: S1, collecting and preprocessing multi-source heterogeneous data; s2, constructing a time sequence knowledge graph oriented to dynamic risk reasoning; s3, constructing a campus fraud recognition model based on multi-source data fusion and a time sequence knowledge graph; and S4, inputting real-time behavior data to be recognized into the campus fraud recognition model. Through fusion and entity alignment of multi-source heterogeneous data, a unified campus behavior data view is constructed, the problem of data islands is effectively solved, through dynamic risk scoring and a multi-level early warning mechanism, conversion from passive rule matching to active risk intervention is realized, and early warning timeliness and defense initiative of campus anti-fraud work are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the fields of information security and artificial intelligence technology, specifically to a campus fraud identification method based on multi-source data fusion and time-series knowledge graphs. Background Technology

[0003] In recent years, targeted fraud against specific student groups has become a key area of ​​focus for crackdown and prevention. In terms of fraud types, fake job offers, impersonating customer service, and fake shopping scams are particularly prevalent in university environments, posing a major threat to campus safety.

[0004] Traditional campus fraud identification methods largely rely on monitoring single-point risk signals, such as intercepting fraudulent text messages based on keyword matching, blocking suspicious calls based on blacklists, or alerting on abnormal transactions through preset rules. These methods are limited to shallow, static feature matching, making it difficult to capture cross-data source correlations to reconstruct the complete attack chain of "online inducement - social infiltration - financial fraud." Furthermore, they are ill-equipped to address the challenges posed by the rapid evolution of fraud methods, as they are essentially defenses based on historically known patterns, lacking the necessary dynamic perception and reasoning capabilities for new variants and combined attacks.

[0005] Knowledge graph technology has been widely used in the field of risk control due to its powerful ability to express relationships. By constructing a topological network of entities (such as students, devices, and accounts) and relationships (such as friend relationships, transaction relationships, and login relationships), static knowledge graphs can support a certain degree of risk transmission analysis, such as discovering potential victims associated with known fraudulent accounts.

[0006] However, directly applying static knowledge graphs to highly dynamic campus fraud scenarios faces significant challenges. The core issue lies in the inherent temporal dynamism and continuous evolution of risk within campus behavioral data. Fraudulent behavior is not a momentary, isolated event, but a dynamic process: a social group may gradually evolve from normal communication into a breeding ground for fraudulent information; a device may gradually exhibit abnormal behavior patterns after being infected with malware. The "snapshot" representation of static knowledge graphs cannot capture this continuous evolutionary trajectory of entity states and relational risks, resulting in the system only being able to identify the consequences after a risk erupts, but unable to warn of the risk accumulation process. Even more problematic is that fraud strategies themselves possess adversarial evolutionary characteristics, deliberately evading detection based on static patterns.

[0007] Given the highly dynamic evolution of campus behavioral data, static graph analysis methods are unable to characterize the transmission path of risks or assess their cumulative effects over time. Therefore, those skilled in the art have provided a campus fraud identification method based on multi-source data fusion and temporal knowledge graphs to address the problems mentioned in the background. Summary of the Invention

[0008] (a) Technical problems to be solved To address the shortcomings of existing technologies, this invention provides a campus fraud identification method based on multi-source data fusion and temporal knowledge graphs. This method solves the problems of existing campus fraud identification methods, which rely on single-source data monitoring and thus cannot construct a panoramic risk view, cannot effectively integrate multi-source heterogeneous data, cannot characterize the dynamic evolution characteristics of entity relationships in the campus environment using static knowledge graph analysis methods, cannot accurately capture risk transmission paths, and suffer from defensive lag due to the lack of temporal reasoning capabilities in traditional early warning mechanisms.

[0009] (II) Technical Solution To achieve the above objectives, this invention provides the following technical solution: a campus fraud identification method based on multi-source data fusion and temporal knowledge graphs, comprising the following steps: S1. Multi-source heterogeneous data collection and preprocessing: The multi-source heterogeneous data is mainly collected from the campus information ecosystem. Through preprocessing operations such as data cleaning, normalization, entity alignment and association, high-quality standardized data is formed to provide data support for the construction of dynamic knowledge graphs. S1.1 Multi-source data collection, with data acquisition methods mainly including: Network behavior data: collected from campus network authentication systems, firewalls, DNS servers, etc., including network access logs of user devices, abnormal connection events, etc.; Consumer transaction data: collected from campus card systems, online payment platforms, etc., including transaction records such as catering consumption, supermarket shopping, online recharge, etc., covering information such as transaction time, location, amount, merchant type, etc.; Social interaction data: collected from campus forums, course groups, official social platforms, etc., after de-identification processing under the premise of complying with privacy protection policies, including post / reply content, user follow relationships, group chat frequency and time patterns, etc.; External threat intelligence data: access to external threat intelligence sources such as anti-fraud blacklist databases, known fraudulent number databases, and malicious website databases provided by public security, financial regulatory agencies, etc. S1.2. Data preprocessing: Standardize and preprocess the collected multi-source heterogeneous raw data to eliminate noise and solve data inconsistency problems for subsequent entity association and fusion. S2. Construction of Temporal Knowledge Graph for Dynamic Risk Reasoning: In order to transform fraud identification into the detection of abnormal evolution patterns of individual behavior in the global network context, this invention first constructs a global temporal knowledge graph that evolves over time; then, through temporal slicing and model attention mechanisms, it realizes dynamic analysis of the behavioral trajectory of specific subjects, thereby more accurately capturing the progressive characteristics of fraud inducement. Based on the standardized "entity-behavior-time" triple sequence generated in step S1, a knowledge graph with time-series awareness is constructed. This graph not only serves as a structured storage carrier for entities and relationships, but also achieves a complete record of the time evolution process by embedding a time dimension, providing a structured input rich in time-series context for subsequent dynamic risk reasoning. S2.1. Temporally Enhanced Knowledge Graph Pattern Definition: To effectively characterize the dynamic behavioral features in campus fraud scenarios, a knowledge graph ontology pattern supporting temporal representation is designed, specifically including the following elements: Core Entity and Relationship Definition: Clearly define the entity types in the graph, including but not limited to students, devices, bank accounts, URLs, and social media accounts, as well as their relationship types, to form a semantically consistent ontology structure; Extended event attributes for relation edges: Each relation edge is assigned time information, supporting the expression of time intervals or timestamp sequences, such as... The relationship can record a sequence of specific time points for multiple visits. , The relationship can be marked with the time when the friendship was established and when it was terminated, so as to reflect the dynamic changes of the relationship; Dynamic attribute modeling of nodes: Introduce attribute vectors that change over time to entity nodes; for example, for student entities, indicators such as total transaction amount and social activity can be recorded daily to form a traceable time-series attribute sequence, thereby characterizing the evolution process of entity state. S2.2. Graph construction and persistence based on time-series triples: The processed triple sequences are imported into the graph database in batches to construct a graph with time-series information. Each quadruple record is processed sequentially. The specific process is as follows: Dynamic updates of nodes and relationships: If the head or tail entity node does not yet exist in the graph, the corresponding node is automatically created; then, relationship edges are created or updated between the corresponding nodes. If the relationship edge already exists, the new timestamp is added to the timestamp sequence attribute of the edge; if the relationship edge does not exist, a new relationship edge is created and its time attribute is initialized with the current timestamp. To support efficient time-range queries, a time-tree index is built on the time attributes of all relation edges. With this index, time-aware query tasks such as "querying all entities that have interacted with high-risk URLs within a certain time period" can be realized, which significantly improves the retrieval efficiency in complex reasoning scenarios. S2.3 Generation of Temporal Subgraph Snapshot Sequences: To adapt to the subsequent model's requirements for processing temporal structure data, the continuously evolving knowledge graph needs to be converted into a discretized temporal subgraph snapshot sequence. The specific process is as follows: Subgraph extraction based on sliding window: Set sliding event window parameters for each time step. The system automatically extracts the time range. Construct a snapshot of a local subgraph containing the temporal attributes of all active entities and relationships within the graph. ; Time-series subgraph sequence construction: As the window slides forward at a set step size, a new subgraph snapshot is generated every hour, and so on. ..., forming a sequence of subgraphs arranged in chronological order. ; Snapshot of each subgraph All of them fully preserved the dynamic interaction between campus entities within the corresponding time period, providing structured input for subsequent time series models; Vectorization of snapshot images: To facilitate subsequent model processing, preliminary feature extraction needs to be performed on each time series sub-image snapshot, converting it into a numerical representation; the specific process is as follows: Initial feature construction for each node: for each subgraph snapshot For nodes in the data, extract static and dynamic temporal features, and the nodes within the current time window. The internal behavioral dynamics, including but not limited to indicators such as transaction frequency and number of visits, constitute its initial feature vector; Structured graph data object generation: Each subgraph snapshot is encapsulated into a structured data object, which specifically includes: a set of nodes and their initial feature vectors, a set of edges and their corresponding time attributes, such as timestamp sequences, and the graph's topology represented by an adjacency matrix or edge index; At this point, the flowing, unstructured raw data has been transformed into a structured knowledge graph sequence with time-series labels; this sequence fully preserves the dynamic evolution information of entity relationships, providing a highly adaptable and directly input data foundation for subsequent models; S3. Construct a campus fraud identification model based on multi-source data fusion and temporal knowledge graph. A temporal graph neural network model that integrates spatial and temporal dependencies is constructed. The model consists of a spatiotemporal feature encoding module, a spatiotemporal fusion attention module, and a multi-granularity risk identification module. S3.1. Spatiotemporal Feature Encoding Module Extracts and Fuses Temporal and Spatial Features: To obtain rich feature representations of the spatial and temporal information of each temporal subgraph snapshot and improve the inference ability of subsequent models, this invention designs a spatiotemporal feature encoding module. Firstly, it uses a graph attention network to encode time slices... subgraph The topological structure is encoded, and the nodes are obtained by aggregating their neighbor information. Spatial perception embedding Meanwhile, by extracting nodes In time slice The dynamic behavioral characteristics within constitute temporal behavioral embedding Finally, embed the space. With temporal embedding The nodes are then stitched together, fused, and reduced in dimensionality using a fully connected layer. exist Unified spatiotemporal feature vector at time step ; S3.2. Spatiotemporal Fusion Attention Reasoning Module Further Enhances Feature Representation: In order to solve the problem that static models are difficult to capture the dynamic transmission path of risks, a spatiotemporal fusion attention reasoning module is designed. By modeling the long-term spatiotemporal dependencies of node behaviors, it realizes continuous perception and reasoning of the risk accumulation process. This module first arranges the unified spatiotemporal feature vectors of each node across all historical time slices in chronological order to construct the spatiotemporal evolution sequence of the nodes. Each feature block in the sequence All have been integrated with the corresponding time. The spatial topology and self-behavioral information lay the foundation for subsequent joint spatiotemporal analysis. Then, a multi-layer Transformer encoder is used as the core component, and its self-attention mechanism is used to achieve deep fusion of spatiotemporal information, and key time nodes and spatial pattern changes are identified in the node behavior history. For any two moments in the sequence and Through a learnable weight matrix , , Generate query vector and key vector respectively; since both query and key vectors are composed of fused spatial and temporal information. and As a derivative, the calculated attention weight can simultaneously reflect the similarity of behavioral patterns and network connections between different time points, thereby enabling quantitative analysis of risk transmission paths; By encoding the time series, node state representations containing historical evolution information are generated. The final encoder layer is then used at the last moment. The output vector is used as a node. Dynamic risk state embedding ; ; The embedding vector By integrating the spatiotemporal characteristics of nodes at the current and historical moments, a continuous representation of the risk status of nodes is formed, providing high-order features of temporal context for risk identification; S3.3. Campus Fraud Identification and Classification, using vector... Input a classifier consisting of a fully connected layer and a Softmax function, and map the extracted features to the feature space of the real labels. Output the final probability distribution of fraud risk levels, which are divided into normal, suspicious and high risk. S4. Input the real-time behavioral data to be identified into the campus fraud identification model to identify fraud behavior characteristics, calculate dynamic risk scores and levels; based on the dynamic risk levels, trigger and output multi-level early warning information from low to high to complete the proactive intervention in campus fraud.

[0010] Preferably, the collected data in step S1.2 is processed as follows: Data cleaning and noise reduction: handling missing values, outliers, and duplicate records in raw data; for example, smoothing or removing transaction amounts that are clearly outside the reasonable range, and removing meaningless records such as heartbeat packets from network logs; Data normalization and structuring: converting data from different sources and in different formats into a unified spatiotemporal standard and data pattern; for example, extracting key entities and sentiment from unstructured text data using natural language processing techniques and transforming it into structured or semi-structured data. Entity recognition and alignment: Based on a unified identification system, such as student ID and device ID, identify and associate the same entity from different data sources; for example, accurately map IP addresses in network logs, campus card numbers in consumption records, and anonymous user IDs in social platforms to the same student entity through technologies such as session association and behavior pattern matching, thus solving the "data silo" problem. Feature engineering and vectorization: Extracting temporal, statistical, and contextual features related to fraud risk for different data modalities; for example, calculating the variance of users' short-term transaction frequency, constructing social network centrality indicators, and encoding device behavior sequences into feature vectors to provide effective input representations for subsequent model learning; After the above processing, the original multi-source data is transformed into a triple sequence with "entity-behavior-time" as the core, providing a standardized and time-series data foundation for the construction of time-series knowledge graphs.

[0011] Preferably, the network access log in step S1.1 includes, but is not limited to, URL records, access duration, and traffic volume.

[0012] (III) Beneficial Effects This invention provides a campus fraud identification method based on multi-source data fusion and time-series knowledge graph. It has the following beneficial effects: 1. In this invention, a unified campus behavior data view is constructed by fusing and aligning multi-source heterogeneous data with entities, effectively solving the problem of data silos.

[0013] 2. In this invention, by utilizing temporal knowledge graphs and dynamic graph neural networks, the continuous evolution process of entity behavior and relationships is depicted, thereby achieving accurate identification and early detection of gradual and evolutionary fraud methods.

[0014] 3. In this invention, through dynamic risk scoring and a multi-level early warning mechanism, the transformation from passive rule matching to proactive risk intervention is realized, thereby improving the timeliness of early warning and the initiative of defense in campus anti-fraud work. Attached Figure Description

[0015] Figure 1 This is a schematic diagram of the overall system flow of the present invention; Figure 2 This is a schematic diagram of the overall structure of the campus fraud identification model based on multi-source data fusion and temporal knowledge graph in this invention. Detailed Implementation

[0016] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0017] Example 1: like Figure 1-2 As shown, this embodiment of the invention provides a campus fraud identification method based on multi-source data fusion and temporal knowledge graph, including the following steps: S1. Multi-source heterogeneous data collection and preprocessing: Multi-source heterogeneous data is mainly collected from the campus information ecosystem. Through preprocessing operations such as data cleaning, normalization, entity alignment and association, high-quality standardized data is formed to provide data support for the construction of dynamic knowledge graphs. S1.1 Multi-source data collection, with data acquisition methods mainly including: Network behavior data: collected from campus network authentication systems, firewalls, DNS servers, etc., including network access logs of user devices, abnormal connection events, etc.; Consumer transaction data: collected from campus card systems, online payment platforms, etc., including transaction records such as catering consumption, supermarket shopping, online recharge, etc., covering information such as transaction time, location, amount, merchant type, etc.; Social interaction data: collected from campus forums, course groups, official social platforms, etc., after de-identification processing under the premise of complying with privacy protection policies, including post / reply content, user follow relationships, group chat frequency and time patterns, etc.; External threat intelligence data: access to external threat intelligence sources such as anti-fraud blacklist databases, known fraudulent number databases, and malicious website databases provided by public security, financial regulatory agencies, etc. S1.2. Data preprocessing: Standardize and preprocess the collected multi-source heterogeneous raw data to eliminate noise and solve data inconsistency problems for subsequent entity association and fusion. S2. Construction of Temporal Knowledge Graph for Dynamic Risk Reasoning: In order to transform fraud identification into the detection of abnormal evolution patterns of individual behavior in the global network context, this invention first constructs a global temporal knowledge graph that evolves over time; then, through temporal slicing and model attention mechanisms, it realizes dynamic analysis of the behavioral trajectory of specific subjects, thereby more accurately capturing the progressive characteristics of fraud inducement. Based on the standardized "entity-behavior-time" triple sequence generated in step S1, a knowledge graph with time-series awareness is constructed. This graph not only serves as a structured storage carrier for entities and relationships, but also achieves a complete record of the time evolution process by embedding a time dimension, providing a structured input rich in time-series context for subsequent dynamic risk reasoning. S2.1. Temporally Enhanced Knowledge Graph Pattern Definition: To effectively characterize the dynamic behavioral features in campus fraud scenarios, a knowledge graph ontology pattern supporting temporal representation is designed, specifically including the following elements: Core Entity and Relationship Definition: Clearly define the entity types in the graph, including but not limited to students, devices, bank accounts, URLs, and social media accounts, as well as their relationship types, to form a semantically consistent ontology structure; Extended event attributes for relation edges: Each relation edge is assigned time information, supporting the expression of time intervals or timestamp sequences, such as... The relationship can record a sequence of specific time points for multiple visits. , The relationship can be marked with the time when the friendship was established and when it was terminated, so as to reflect the dynamic changes of the relationship; Dynamic attribute modeling of nodes: Introduce attribute vectors that change over time to entity nodes; for example, for student entities, indicators such as total transaction amount and social activity can be recorded daily to form a traceable time-series attribute sequence, thereby characterizing the evolution process of entity state. S2.2. Graph construction and persistence based on time-series triples: The processed triple sequences are imported into the graph database in batches to construct a graph with time-series information. Each quadruple record is processed sequentially. The specific process is as follows: Dynamic updates of nodes and relationships: If the head or tail entity node does not yet exist in the graph, the corresponding node is automatically created; then, relationship edges are created or updated between the corresponding nodes. If the relationship edge already exists, the new timestamp is added to the timestamp sequence attribute of the edge; if the relationship edge does not exist, a new relationship edge is created and its time attribute is initialized with the current timestamp. To support efficient time-range queries, a time-tree index is built on the time attributes of all relation edges. With this index, time-aware query tasks such as "querying all entities that have interacted with high-risk URLs within a certain time period" can be realized, which significantly improves the retrieval efficiency in complex reasoning scenarios. S2.3 Generation of Temporal Subgraph Snapshot Sequences: To adapt to the subsequent model's requirements for processing temporal structure data, the continuously evolving knowledge graph needs to be converted into a discretized temporal subgraph snapshot sequence. The specific process is as follows: Subgraph extraction based on sliding window: Set sliding event window parameters for each time step. The system automatically extracts the time range. Construct a snapshot of a local subgraph containing the temporal attributes of all active entities and relationships within the graph. ; Time-series subgraph sequence construction: As the window slides forward at a set step size, a new subgraph snapshot is generated every hour, and so on. ..., forming a sequence of subgraphs arranged in chronological order. ; Snapshot of each subgraph All of them fully preserved the dynamic interaction between campus entities within the corresponding time period, providing structured input for subsequent time series models; Vectorization of snapshot images: To facilitate subsequent model processing, preliminary feature extraction needs to be performed on each time series sub-image snapshot, converting it into a numerical representation; the specific process is as follows: Initial feature construction for each node: for each subgraph snapshot For nodes in the data, extract static and dynamic temporal features, and the nodes within the current time window. The internal behavioral dynamics, including but not limited to indicators such as transaction frequency and number of visits, constitute its initial feature vector; Structured graph data object generation: Each subgraph snapshot is encapsulated into a structured data object, which specifically includes: a set of nodes and their initial feature vectors, a set of edges and their corresponding time attributes, such as timestamp sequences, and the graph's topology represented by an adjacency matrix or edge index; At this point, the flowing, unstructured raw data has been transformed into a structured knowledge graph sequence with time-series labels; this sequence fully preserves the dynamic evolution information of entity relationships, providing a highly adaptable and directly input data foundation for subsequent models; S3. Construct a campus fraud identification model based on multi-source data fusion and temporal knowledge graph. A temporal graph neural network model that integrates spatial and temporal dependencies is constructed. The model consists of a spatiotemporal feature encoding module, a spatiotemporal fusion attention module, and a multi-granularity risk identification module. S3.1. Spatiotemporal Feature Encoding Module Extracts and Fuses Temporal and Spatial Features: To obtain rich feature representations of the spatial and temporal information of each temporal subgraph snapshot and improve the inference ability of subsequent models, this invention designs a spatiotemporal feature encoding module. Firstly, it uses a graph attention network to encode time slices... subgraph The topological structure is encoded, and the nodes are obtained by aggregating their neighbor information. Spatial perception embedding Meanwhile, by extracting nodes In time slice The dynamic behavioral characteristics within constitute temporal behavioral embedding Finally, embed the space. With temporal embedding The nodes are then stitched together, fused, and reduced in dimensionality using a fully connected layer. exist Unified spatiotemporal feature vector at time step ; S3.2. Spatiotemporal Fusion Attention Reasoning Module Further Enhances Feature Representation: In order to solve the problem that static models are difficult to capture the dynamic transmission path of risks, a spatiotemporal fusion attention reasoning module is designed. By modeling the long-term spatiotemporal dependencies of node behaviors, it realizes continuous perception and reasoning of the risk accumulation process. This module first arranges the unified spatiotemporal feature vectors of each node across all historical time slices in chronological order to construct the spatiotemporal evolution sequence of the nodes. Each feature block in the sequence All have been integrated with the corresponding time. The spatial topology and self-behavioral information lay the foundation for subsequent joint spatiotemporal analysis. Then, a multi-layer Transformer encoder is used as the core component, and its self-attention mechanism is used to achieve deep fusion of spatiotemporal information, and key time nodes and spatial pattern changes are identified in the node behavior history. For any two moments in the sequence and Through a learnable weight matrix , , Generate query vector and key vector respectively; since both query and key vectors are composed of fused spatial and temporal information. and As a derivative, the calculated attention weight can simultaneously reflect the similarity of behavioral patterns and network connections between different time points, thereby enabling quantitative analysis of risk transmission paths; By encoding the time series, node state representations containing historical evolution information are generated. The final encoder layer is then used at the last moment. The output vector is used as a node. Dynamic risk state embedding ; ; The embedding vector By integrating the spatiotemporal characteristics of nodes at the current and historical moments, a continuous representation of the risk status of nodes is formed, providing high-order features of temporal context for risk identification; S3.3. Campus Fraud Identification and Classification, using vector... Input a classifier consisting of a fully connected layer and a Softmax function, and map the extracted features to the feature space of the real labels. Output the final probability distribution of fraud risk levels, which are divided into normal, suspicious and high risk. S4. Input the real-time behavioral data to be identified into the campus fraud identification model to identify fraud behavior characteristics, calculate dynamic risk scores and levels; based on the dynamic risk levels, trigger and output multi-level early warning information from low to high to complete the proactive intervention in campus fraud.

[0018] In step S1.2, the collected data is processed as follows: Data cleaning and noise reduction: handling missing values, outliers, and duplicate records in raw data; for example, smoothing or removing transaction amounts that are clearly outside the reasonable range, and removing meaningless records such as heartbeat packets from network logs; Data normalization and structuring: Converting data from different sources and in different formats into a unified spatiotemporal standard and data pattern; for example, converting unstructured text data, such as forum posts, into structured or semi-structured data by extracting key entities and sentiment tendencies through natural language processing techniques. Entity recognition and alignment: Based on a unified identification system, such as student ID and device ID, identify and associate the same entity from different data sources; for example, accurately map IP addresses in network logs, campus card numbers in consumption records, and anonymous user IDs in social platforms to the same student entity through technologies such as session association and behavior pattern matching, thus solving the "data silo" problem. Feature engineering and vectorization: Extracting temporal, statistical, and contextual features related to fraud risk for different data modalities; for example, calculating the variance of users' short-term transaction frequency, constructing social network centrality indicators, and encoding device behavior sequences into feature vectors to provide effective input representations for subsequent model learning; After the above processing, the original multi-source data is transformed into a triple sequence with "entity-behavior-time" as the core, providing a standardized and time-series data foundation for the construction of time-series knowledge graphs.

[0019] In step S1.1, network access logs include, but are not limited to, URL records, access duration, and traffic volume.

[0020] Although embodiments of the invention have been shown and described, it will be understood by those skilled in the art that various changes, modifications, substitutions and alterations can be made to these embodiments without departing from the principles and spirit of the invention, the scope of which is defined by the appended claims and their equivalents.

Claims

1. A campus fraud identification method based on multi-source data fusion and time-series knowledge graph, characterized in that: Comprise the following steps: S1. Multi-source heterogeneous data collection and preprocessing, the multi-source heterogeneous data is mainly collected from the campus information ecology, through data cleaning, normalization, entity alignment and correlation, etc. Preprocessing operation, form high quality standardized data, provide data support for the construction of dynamic knowledge graph; S1.1 Multi-source data collection, the data acquisition source method mainly includes: network behavior data: collected from campus network authentication system, firewall, DNS server, etc. Including user equipment network access log, abnormal connection event, etc.; Consumption transaction data: collected from campus card system, online payment platform, etc. Including dining consumption, supermarket shopping, online recharge, etc. Transaction records, covering transaction time, place, amount, merchant type, etc. Information; Social interaction data: under the premise of complying with the privacy protection policy, after de-identification processing, collected from campus forum, course group, official social platform, etc. Including post / post content, user attention relationship, group chat frequency and time mode, etc.; External threat intelligence data: access to anti-fraud black list library, known fraud number library, malicious website library and other external threat intelligence sources provided by public security, financial supervision agencies, etc.; S1.

2. Data preprocessing, the collected multi-source heterogeneous raw data is standardized and pretreated to eliminate noise and solve the problem of data inconsistency, for subsequent entity correlation and fusion; S2. Time sequence knowledge graph construction for dynamic risk reasoning, in order to transform fraud identification into abnormal evolution pattern detection of individual behavior in the global network context, the invention first constructs a global, time-evolving time sequence knowledge graph; Further, through time sequence slicing and model attention mechanism, the dynamic analysis of the behavior track of a specific subject (such as a student) is realized, so as to more accurately capture the progressive characteristics induced by fraud; On the basis of generating standardized "entity-behavior-time" triple sequence in S1 step, a knowledge graph with time sequence perception ability is constructed, which not only serves as a structured storage carrier for entities and relationships, but also realizes complete recording of the time evolution process by embedding the time dimension, providing structured input rich in time sequence context for subsequent dynamic risk reasoning; S2.

1. Knowledge graph mode definition enhanced by time sequence, in order to effectively depict the dynamic behavior characteristics in the campus fraud scene, a knowledge graph ontology mode supporting time sequence expression is designed, which specifically includes the following elements: Core entity and relationship definition: clearly define the entity types in the graph, including but not limited to students, devices, bank accounts, websites and social accounts, and their relationship types, forming a semantically consistent ontology structure; Event attribute extension of relationship edge: give each relationship edge time information, support the expression of time interval or timestamp sequence, such as , the specific time point sequence of multiple visits can be recorded in the relationship , , the establishment time and the release time of the friend relationship can be marked in the relationship to reflect the dynamic change of the relationship; Dynamic attribute modeling of nodes: introduce time-varying attribute vectors for entity nodes; For example, for student entities, daily transaction amount, social activity level, etc. Indexes can be recorded to form traceable time sequence attribute sequences, thereby depicting the entity state evolution process; S2.

2. Graph construction and persistence based on time sequence triplets, the processed triple sequence is batch imported into the graph database, realizing the construction of graph with time sequence information, processing each four-tuple record in sequence, the specific process is as follows: Dynamic update of nodes and relations: if the head entity or tail entity node does not exist in the graph, the corresponding node is automatically created; then, a relation edge is created or updated between the corresponding nodes (if the relation edge already exists, a new timestamp is added to the timestamp sequence attribute of the edge; if the relation edge does not exist, a new relation edge is created, and its time attribute is initialized with the current timestamp); Time index construction: to support efficient time range queries, a time tree index is established for the time attribute on all relation edges. With the help of the index, time-aware query tasks such as "query all entities that interact with high-risk websites within a certain time period" can be implemented, significantly improving retrieval efficiency in complex reasoning scenarios; S2.3 Generation of time sequence subgraph snapshot sequence: to adapt to the processing needs of subsequent models for time sequence structure data, it is necessary to convert the continuously evolving knowledge graph into a discrete time sequence subgraph snapshot sequence. The specific process is as follows: Subgraph extraction based on sliding window: set sliding event window parameters, for each time , the system automatically extracts all active entities and relationships within the time interval , construct local subgraph snapshots containing their temporal attributes ; Temporal subgraph sequence construction: As the window slides forward by the set step size, a new subgraph snapshot is generated every hour, and in turn ,..., forming a sequence of subgraphs arranged in chronological order ; Each subgraph snapshot completely retains the dynamic interaction between campus entities within the corresponding period, providing structured input for subsequent temporal models; Vectorization representation of snapshot graph: to facilitate subsequent model processing, preliminary feature extraction is required for each time sequence subgraph snapshot, which is converted into a numerical representation. The specific process is as follows: Node initial feature construction: for each node in the subgraph snapshot , extract static features and dynamic timing features, the behavior of the node in the current time window , including but not limited to transaction frequency, access frequency, etc. The indicators constitute its initial feature vector; Structured graph data object generation: each subgraph snapshot is encapsulated into a structured data object, which specifically includes: node set and its initial feature vector, edge set and corresponding time attribute (such as timestamp sequence), and graph topology (represented by adjacency matrix or edge index); At this point, the flowing unstructured raw data has been converted into a structured knowledge graph sequence with time sequence labels; this sequence completely retains the dynamic evolution information of entity relations, providing a high-adaptation and directly-input data foundation for subsequent models; S3. Constructing a campus fraud identification model based on multi-source data fusion and time sequence knowledge graph, a time sequence graph neural network model that integrates spatial and temporal dependencies is constructed. The model consists of a spatio-temporal feature encoding module, a spatio-temporal fusion attention module, and a multi-granularity risk identification module. S3.

1. Spatiotemporal Feature Encoding Module Extracts and Fuses Temporal and Spatial Features: To obtain rich feature representations of the spatial and temporal information of each temporal subgraph snapshot and improve the inference ability of subsequent models, this invention designs a spatiotemporal feature encoding module. Firstly, it uses a graph attention network to encode time slices... subgraph The topological structure is encoded, and the nodes are obtained by aggregating their neighbor information. Spatial perception embedding Meanwhile, by extracting nodes In time slice The dynamic behavioral characteristics within constitute temporal behavioral embedding Finally, embed the space. With temporal embedding The nodes are then stitched together, fused, and reduced in dimensionality using a fully connected layer. exist Unified spatiotemporal feature vector at time step ; S3.

2. Spatio-temporal fusion attention reasoning module further enhances feature representation: to solve the problem that static models are difficult to capture risk dynamic transmission paths, a spatio-temporal fusion attention reasoning module is designed to model the long-term spatio-temporal dependency of node behavior, enabling continuous perception and reasoning of risk accumulation process; The module first arranges the unified spatio-temporal feature vector of each node at all historical time slices in chronological order to construct a spatio-temporal evolution sequence of the node Each feature block in the sequence has fused the spatial topology and behavior information of the corresponding moment , laying a foundation for subsequent joint spatio-temporal analysis, and then adopts a multi-layer Transformer encoder as the core component to realize deep fusion of spatio-temporal information through its self-attention mechanism, and identify key time nodes and spatial mode changes in the node behavior history; For any two time points in the sequence and , the query vector and the key vector are respectively generated through a learnable weight matrix , , Since the query and key vectors are derived from the space-time information and that have been fused, the attention weight calculated can simultaneously reflect the similarity in behavior patterns and network associations between different time points, thereby realizing quantitative analysis of the risk transmission path. By encoding the time series, a node state representation containing historical evolution information is generated, and the output vector of the last layer encoder at the final time is taken as the dynamic risk state embedding of the node ; ; ; ; The embedding vector The spatio-temporal features of the nodes at the current and historical time points are integrated to form a continuous representation of the risk state of the nodes, thereby providing high-order features of the timing context for risk identification. S3.

3. Campus fraud recognition and classification, the vector Input a classifier composed of a fully connected layer Softmax function, and map the extracted features to the feature space of the true label, output the final fraud risk level probability distribution, the risk level is divided into normal, suspicious and high risk; S4. Input the real-time behavior data to be identified into the campus fraud identification model to identify fraud behavior characteristics and calculate dynamic risk scores and levels. According to the dynamic risk level, multi-level early warning information from low to high is triggered and output to complete the proactive intervention of campus fraud. 2.The campus fraud identification method of multi-source data fusion and time-series knowledge graph according to claim 1, characterized in that: The following processing is performed on the collected data in S1.2: Data cleaning and denoising: handle missing values, outliers, and duplicate records in the original data; for example, smooth or remove obviously excessive transaction amounts, and remove meaningless records such as heartbeat packets in network logs; Data normalization and structuring: convert data of different sources and formats into unified spatio-temporal standards and data patterns; for example, extract key entities and sentiment orientations from unstructured text data (such as forum posts) through natural language processing techniques, and convert them into structured or semi-structured data; Entity recognition and alignment: Based on unified identification system (such as student ID, device ID, etc.), identify and associate the same entity from different data sources; For example, the IP address in the network log, the card number in the consumption record, and the anonymous user ID in the social platform are accurately mapped to the same student entity through session association, behavior pattern matching and other technologies, solving the "data island" problem; Feature engineering and vectorization: For different data modalities, extract time series features, statistical features and context features related to fraud risk; For example, calculate the transaction frequency variance of the user in the short term, construct the social network centrality index, and encode the device behavior sequence into a feature vector, etc. Provide effective input representation for subsequent model learning; After the above processing, the original multi-source data is transformed into a triple sequence with "entity-behavior-time" as the core, providing a standardized and time-series data foundation for the construction of a time-series knowledge graph. 3.The campus fraud identification method of multi-source data fusion and time-series knowledge graph according to claim 1, characterized in that: The network access log in the S1.1 step includes but is not limited to URL record, access duration and traffic size.