A deep spoofing detection and tracing system for network phishing attacks

By constructing a unified attack association graph and using PathSim similarity calculation, the problem of unified association between deepfake content and phishing propagation characteristics in phishing attacks is solved, improving the accuracy of identification and the reliability of tracing the source, and addressing the lack of stability in the identification of fake content and propagation behavior in existing technologies.

CN122394861APending Publication Date: 2026-07-14NINGXIA PUSHI INFORMATION TECH SERVICE CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
NINGXIA PUSHI INFORMATION TECH SERVICE CO LTD
Filing Date
2026-04-16
Publication Date
2026-07-14

AI Technical Summary

Technical Problem

Existing technologies lack a unified framework for linking deepfake content with phishing propagation characteristics in phishing attack detection, making it difficult to effectively identify the intrinsic connection between fake content and propagation behavior. Furthermore, existing solutions are susceptible to spoofing drift in similarity calculations, affecting the stability and accuracy of the identification results.

Method used

A unified attack association graph is constructed for the forged content association layer and the phishing propagation association layer. By combining PathSim similarity calculation and verification of forged evidence and propagation evidence, and through multi-source data collection, feature extraction, association graph generation, path filtering and iterative calculation and updating, the accuracy of deepfake content identification and the stability of attack association analysis are improved.

Benefits of technology

It enables the complete and accurate identification of deepfake content in phishing attack scenarios, enhances the ability to express the correlation of heterogeneous attack objects, and improves the stability of attack correlation analysis and the reliability of source tracing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122394861A_ABST
    Figure CN122394861A_ABST
Patent Text Reader

Abstract

The application discloses a kind of deep fake detection and tracing systems for network phishing attack, including the following steps: multi-source data acquisition module is used to form pre-processing multi-source data;Feature extraction module is used to form associated feature data;Association graph generation module is used to build fake content association layer and phishing propagation association layer, form unified attack association graph;Path screening construction module is used to identify effective relationship section, and build calculation path set;Similarity calculation backwrite module is used to perform PathSim similarity calculation, form the unified attack association graph after state update;Iterative calculation update module is used to re-execute PathSim similarity calculation, form candidate association result;Back-checking analysis output module is used to verify based on fake evidence and propagation evidence, and convert the candidate association result after verification into attack analysis result.The application improves the accuracy of deep fake content identification in network phishing attack scenario.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of network security technology, and in particular to a deepfake detection and tracing system for phishing attacks. Background Technology

[0002] With the development of generative artificial intelligence, deep synthesis, and multimedia forgery technologies, the methods of phishing attacks have evolved from traditional malicious links, fake pages, and misleading texts to composite attack forms that integrate image forgery, voice forgery, video forgery, webpage spoofing, and propagation link spoofing. Existing detection solutions for phishing attacks mostly focus on domain name identification, page similarity judgment, email content filtering, blacklist matching, and log auditing, while detection solutions for deep forgery content mainly concentrate on image authenticity recognition, audio synthesis detection, video tampering analysis, and text anomaly recognition.

[0003] Existing technologies have several shortcomings. Firstly, current solutions typically handle deepfake detection and phishing propagation analysis separately, lacking a unified approach to linking deepfake and phishing features. This makes it difficult to identify the inherent connection between forged content and propagation behavior holistically. Secondly, in terms of association modeling, most existing technologies only reach the level of ordinary graph connections or simple link splicing, failing to effectively fold and integrate the cross-layer correspondences between the forged content association layer and the phishing propagation association layer, resulting in insufficient ability to express heterogeneous associations between attack targets. Thirdly, existing path filtering and similarity analysis methods mostly process the original graph structure directly, lacking technical mechanisms to peel away surface-level change segments, retain effective relationship segments, and reorganize computational paths by incorporating attack camouflage variation features. This makes similarity calculations susceptible to camouflage drift interference, affecting the stability of the same-source identification results.

[0004] Therefore, how to provide a deepfake detection and tracing system for phishing attacks is a problem that urgently needs to be solved by those skilled in the art. Summary of the Invention

[0005] One objective of this invention is to propose a deepfake detection and tracing system for phishing attacks. This invention constructs a unified attack association graph by using a fake content association layer and a phishing propagation association layer, and combines it with a calculation path set, PathSim similarity calculation, contribution write-back, and verification of fake evidence and propagation evidence to improve the accuracy of deepfake content identification, the stability of attack association analysis, and the reliability of source tracing in phishing attack scenarios.

[0006] A deepfake detection and tracing system for phishing attacks according to an embodiment of the present invention includes: The multi-source data acquisition module is used to acquire multi-source data corresponding to the object to be detected and to perform data preprocessing to form preprocessed multi-source data. The feature extraction module is used to extract deepfake features and phishing propagation features from preprocessed multi-source data to form associated feature data; The association graph generation module is used to construct a forged content association layer and a phishing propagation association layer based on association feature data, and connect and fold them according to cross-layer anchoring rules to form a unified attack association graph. The path filtering construction module is used to identify effective relationship segments based on a unified attack association graph and the attack camouflage change characteristics of the current target object, and to construct a set of computational paths based on the effective relationship segments. The similarity calculation write-back module is used to perform PathSim similarity calculation on the set of calculation paths to form initial candidate association results, and to perform contribution analysis on the association calculation elements involved in the calculation. The contribution analysis results are then written back to the unified attack association graph to form the unified attack association graph after state update. The iterative computation update module is used to reconstruct the computation path set based on the unified attack association graph after the state update, and re-execute PathSim similarity calculation on computation objects that meet the state preservation condition to form candidate association results; The verification analysis output module is used to perform verification confirmation on candidate association results, verify based on forged evidence and propagation evidence, and convert the verified candidate association results into attack analysis results.

[0007] Optionally, the multi-source data acquisition module includes: Collect raw multi-source data corresponding to the object to be detected, and classify and aggregate the data according to the data source. Record the corresponding data source identifier, object to be detected identifier and collection time identifier for each type of raw multi-source data to form a labeled raw multi-source dataset. Based on the original labeled multi-source dataset, unified labeling and time alignment are performed to form a time-consistent associated multi-source dataset. Standardization is performed on time-consistent, correlated multi-source datasets to obtain preprocessed multi-source data.

[0008] Optionally, the feature extraction module includes: Preprocessed multi-source data is aggregated according to the identifier of the object to be detected, and the aggregated preprocessed multi-source data is arranged in chronological order according to the collection time identifier. The chronologically arranged preprocessed multi-source data is divided into content processing sequence and propagation processing sequence according to the data source. The content processing sequence is subjected to target localization, fragment segmentation, correspondence alignment and difference comparison to extract forgery traces, tampering traces and inconsistency traces to form deep forgery features. The propagation processing sequence is subjected to relationship expansion, link matching, propagation tracking and anomaly screening to form phishing propagation features. Deepfake features and phishing propagation features are associated, organized, and uniformly packaged according to the identifier of the object to be detected and the time of collection to form associated feature data.

[0009] Optionally, the association graph generation module includes: Based on the identifier of the object to be detected and the identifier of the collection time, the deep forgery features in the associated feature data are compressed into forgery content evolution units that characterize the forgery change process, and a forgery content association layer is constructed based on the evolutionary association relationship between the forgery content evolution units. The phishing propagation features in the associated feature data are merged according to the object identifier to be detected, sorted according to the collection time identifier, and linked together according to the sequential connection relationship after sorting to form a propagation relationship evolution unit. Based on the propagation association between the propagation relationship evolution units, a phishing propagation association layer is formed. According to the cross-layer anchoring rules, the corresponding retrieval of the forged content evolution unit in the forged content association layer and the propagation relationship evolution unit in the phishing propagation association layer is performed. At the same time, according to the risk gating rules, the forged content evolution unit and propagation relationship evolution unit that have completed the corresponding retrieval are gating and filtered. For the forged content evolution unit and propagation relationship evolution unit that meet the cross-layer anchoring and risk gating, cross-layer connection, semantic compression and unified mapping are performed to form attack semantic units and establish the cross-layer connection relationship corresponding to the attack semantic units. The content associations in the forged content association layer, the propagation associations in the phishing propagation association layer, and the cross-layer connection relationships are folded and integrated. The folded and integrated content associations, propagation associations, and cross-layer connection relationships are then organized into connection relationships between attack nodes to form a unified attack association graph.

[0010] Optionally, the path filtering construction module includes: Extract content associations, propagation associations, and cross-layer connections corresponding to the same object to be detected from the unified attack association graph. Expand the connection segments according to the connection order of the attack edges in the unified attack association graph. Extract the content change position and cross-layer change position from the deep forgery feature corresponding to the current object to be detected. Extract the propagation change position from the phishing propagation feature corresponding to the current object to be detected to form attack camouflage change features. Perform position matching on the expanded connection segments to form a candidate relation segment set. A stripping process is performed on each candidate relation segment in the candidate relation segment set. The stripping process involves identifying connection segments that only cover the content change location and the propagation change location but do not cover the cross-layer change location, deleting the identified connection segments, and re-establishing the connection order of the attack nodes before and after the deletion of the connection segments to form the stripped candidate relation segment set. The stripped candidate relation segment set is subjected to retention processing. The retention processing identifies connection segments that still maintain continuous correspondence after cross-layer connection. The identified connection segments are retained and bundled and organized according to the cross-layer connection position, the first and last connection position of the attack node, and the connection position before and after the collection time mark to form a valid relation segment set. Around the continuous cross-layer connection positions in the set of effective relationship segments, the terminating attack node of the previous effective relationship segment and the starting attack node of the next effective relationship segment are connected and spliced ​​together. The terminating attack node of the previous effective relationship segment and the starting attack node of the next effective relationship segment are established to establish a connection relationship. The connected effective relationship segments are then organized into paths according to the continuous order of content association, the continuous order of propagation association, and the continuous order of cross-layer connection to form a set of computational paths.

[0011] Optionally, the similarity calculation write-back module includes: Each computation path in the computation path set is expanded into a path instance. For each attack node to be computed, a retained path instance is extracted that maintains continuous content association, continuous propagation association, and continuous cross-layer connection. The continuity of cross-layer connection positions and the corresponding situation before and after the attack node in each retained path instance are statistically analyzed to form the cross-layer anchoring preservation degree. The position matching between each retained path instance and the attack camouflage change feature is statistically analyzed to form the camouflage drift suppression value. The cross-layer anchoring preservation degree and the camouflage drift suppression value are written into the corresponding retained path instance to form the path contribution weight. Based on the path contribution weight, a weighted PathSim similarity calculation is performed. The weighted path matching sequence corresponding to any attack node to be computed is accumulated and statistically analyzed. The accumulated statistical result is normalized with the total amount of weighted path matching sequence corresponding to each attack node to be computed to obtain the similarity result of any attack node pair to be computed, forming the initial candidate association result. Contribution analysis is performed on the initial candidate association results. Each calculation path in the calculation path set is compared according to the path contribution weight to form the path contribution result. The attack nodes to be calculated are compared according to the cross-layer anchoring retention degree to form the node contribution result. The content association, propagation association and cross-layer connection are compared according to the camouflage drift suppression value to form the connection contribution result. The contribution analysis result is formed by the path contribution result, node contribution result and connection contribution result. Drift suppression processing is performed on the contribution analysis results, and the contribution analysis results are matched with the attack camouflage change features. Suppression processing is performed on the contribution part of the corresponding stripped candidate relation segment set, and retention processing is performed on the contribution part of the corresponding valid relation segment set that maintains the cross-layer anchorage retention degree, thus forming the suppressed contribution analysis results. The contribution rewrite is performed on the suppressed contribution analysis results, and the association retention state of the computation path set and the unified attack association graph is adjusted according to the rewrite results to form the unified attack association graph after state update.

[0012] Optionally, the iterative calculation and update module includes: Based on the unified attack association graph after state update, the retention state retrieval is performed on the computation paths in the computation path set to identify the computation paths, attack nodes, content associations, propagation associations and cross-layer connections that maintain the corresponding relationship between retention and removal, and then reorganize them into an iterative computation path set. For computational objects in the iterative computation path set that meet the condition of preserving the state, perform path instance reconstruction, re-establish the path matching sequence according to the continuous order of content association, continuous order of propagation association, and continuous order of cross-layer connection, and re-execute PathSim similarity calculation based on the reconstructed path matching sequence to form iterative similarity results.

[0013] The iterative similarity results are associated and merged. According to the identifier of the object to be detected and the collection time identifier, the attack nodes corresponding to the similarity results are combined into candidate homologous objects. The content association, propagation association and cross-layer connection corresponding to the candidate homologous objects are combined into candidate association links. The candidate homologous objects and candidate association links are consistent and organized to form candidate association results. A closed-loop organization is performed on the candidate homologous objects and candidate association links in the candidate association results, and the deepfake features corresponding to the candidate homologous objects are bound to the phishing propagation features corresponding to the candidate association links to form candidate association results.

[0014] Optionally, the proof analysis output module includes: The candidate association results are split to extract candidate homologous objects and candidate association links. The deepfake features corresponding to the candidate homologous objects are merged and organized to form a forgery evidence set. The phishing propagation features corresponding to the candidate association links are merged and organized to form a propagation evidence set. Based on the candidate homologous objects, the forgery traces, tampering traces and inconsistency traces in the forgery evidence set are compared to form forgery verification results. Based on the candidate association links, the propagation nodes, propagation directions and propagation connection relationships in the propagation evidence set are verified to form propagation verification results. The system binds the forged verification results and the propagation verification results accordingly. It confirms the candidate association results that simultaneously maintain the correspondence between the forged content and the propagation link, and converts the confirmed candidate association results into attack analysis results.

[0015] The beneficial effects of this invention are: This invention uses a multi-source data acquisition module and a feature extraction module to perform unified preprocessing and associated feature extraction on data in phishing attack scenarios. It can simultaneously acquire deepfake features and phishing propagation features in the same technical chain, thereby improving the completeness and accuracy of deepfake content recognition in phishing attack scenarios.

[0016] This invention constructs a forged content association layer and a phishing propagation association layer through an association graph generation module, and combines cross-layer anchoring rules, risk gating rules, cross-layer connections, semantic compression, and unified mapping to form a unified attack association graph. This can organize the originally scattered forged content information and propagation behavior information into the same attack association space, thereby enhancing the ability to express the association between heterogeneous attack objects and the ability to characterize attack relationships.

[0017] This invention uses a path filtering construction module to peel off and retain candidate relationship segments based on attack camouflage change features. In the similarity calculation and write-back module, it combines PathSim similarity calculation, path contribution weight, contribution parsing results, and contribution write-back to update the state of the unified attack association graph. This can effectively suppress the interference caused by surface camouflage changes and improve the stability of attack association analysis and the credibility of candidate association results.

[0018] This invention re-executes PathSim similarity calculation on computational objects that meet the condition of retaining the state through an iterative calculation update module, and verifies the candidate association results by combining forged evidence and propagation evidence through a back-evidence analysis output module. This enables gradual convergence from the initial candidate association to the candidate association results and then to the attack analysis results, thereby improving the reliability of same-source object identification, the accuracy of propagation link analysis, and the credibility of attack tracing conclusions in phishing attack scenarios. Attached Figure Description

[0019] The accompanying drawings are provided to further illustrate the invention and form part of the specification. They are used in conjunction with embodiments of the invention to explain the invention and do not constitute a limitation thereof. In the drawings: Figure 1 The flowchart shows a deepfake detection and tracing system for phishing attacks proposed in this invention. Figure 2 This is a schematic diagram of the meta-path evolution and back-verification of a deepfake detection and tracing system for phishing attacks proposed in this invention, based on PathSim. Detailed Implementation

[0020] The present invention will now be described in further detail with reference to the accompanying drawings. These drawings are simplified schematic diagrams, illustrating only the basic structure of the invention, and therefore only show the components relevant to the invention.

[0021] refer to Figures 1-2 A deepfake detection and attribution system for phishing attacks, comprising: The multi-source data acquisition module is used to acquire multi-source data corresponding to the object to be detected and to perform data preprocessing to form preprocessed multi-source data. The feature extraction module is used to extract deepfake features and phishing propagation features from preprocessed multi-source data to form associated feature data; The association graph generation module is used to construct a forged content association layer and a phishing propagation association layer based on association feature data, and connect and fold them according to cross-layer anchoring rules to form a unified attack association graph. The path filtering construction module is used to identify effective relationship segments based on a unified attack association graph and the attack camouflage change characteristics of the current target object, and to construct a set of computational paths based on the effective relationship segments. The similarity calculation write-back module is used to perform PathSim similarity calculation on the set of calculation paths to form initial candidate association results, and to perform contribution analysis on the association calculation elements involved in the calculation. The contribution analysis results are then written back to the unified attack association graph to form the unified attack association graph after state update. The iterative computation update module is used to reconstruct the computation path set based on the unified attack association graph after the state update, and re-execute PathSim similarity calculation on computation objects that meet the state preservation condition to form candidate association results; The verification analysis output module is used to perform verification confirmation on candidate association results, verify based on forged evidence and propagation evidence, and convert the verified candidate association results into attack analysis results.

[0022] In this embodiment, the multi-source data acquisition module includes: Collect raw multi-source data corresponding to the object to be detected, and classify and aggregate the data according to the data source. The raw multi-source data includes email data, web page data, image data, audio data, video data, text data, log data, and infrastructure data. Record the corresponding data source identifier, object to be detected identifier, and collection time identifier for each type of raw multi-source data to form a labeled raw multi-source dataset. Based on the labeled original multi-source dataset, unified labeling and time alignment are performed. Unified labeling associates different source data corresponding to the same object to be detected to the same object label. Time alignment arranges the different source data in order and time correspondence according to the collection time label corresponding to each original multi-source data, forming a time-consistent associated multi-source dataset. Standardization is performed on time-consistent multi-source datasets. Standardization includes unifying field structure, data type, and storage format. After standardization, preprocessed multi-source data is obtained.

[0023] In this embodiment, the feature extraction module includes: Preprocessed multi-source data is aggregated according to the identifier of the object to be detected, and the aggregated preprocessed multi-source data is arranged in chronological order according to the collection time identifier. The chronologically arranged preprocessed multi-source data is divided into content processing sequence and propagation processing sequence according to the data source. The system performs target localization, segmentation, correspondence alignment, and difference comparison on the content processing sequence to extract forgery traces, tampering traces, and inconsistency traces, forming deepfake features. It also performs relation expansion, link matching, propagation tracing, and anomaly screening on the propagation processing sequence to form phishing propagation features. Target localization extracts corresponding image regions, audio segments, video segments, web page regions, and text segments from the content processing sequence to form content units to be analyzed. Segmentation segments the content units to be analyzed to form comparable content analysis units. Correspondence alignment establishes time and positional correspondences for the content analysis units. Difference comparison performs consistency comparison and difference extraction on the aligned content analysis units. Relationship expansion sequentially expands the propagation records in the propagation processing sequence and establishes preceding and following relationships. Link matching connects the propagation records to form a propagation link. Propagation tracing identifies propagation nodes, propagation directions, and propagation connection relationships along the propagation link. Anomaly screening identifies anomalies in the propagation link. Deepfake features and phishing propagation features are associated, organized, and uniformly packaged according to the identifier of the object to be detected and the time of collection to form associated feature data.

[0024] In this embodiment, the association graph generation module includes: Based on the identifier of the object to be detected and the identifier of the collection time, the deep forgery features in the associated feature data are compressed into forgery content evolution units that characterize the forgery change process, and a forgery content association layer is constructed based on the evolutionary association relationship between the forgery content evolution units. The phishing propagation features in the associated feature data are merged according to the object identifier to be detected, sorted according to the collection time identifier, and linked together according to the sequential connection relationship after sorting to form a propagation relationship evolution unit. Based on the propagation association between the propagation relationship evolution units, a phishing propagation association layer is formed. According to the cross-layer anchoring rule, corresponding retrieval is performed on the forged content evolution units in the forged content association layer and the propagation relationship evolution units in the phishing propagation association layer. Simultaneously, according to the risk gating rule, the forged content evolution units and propagation relationship evolution units that have completed the corresponding retrieval are gated and filtered. Forged content evolution units and propagation relationship evolution units that satisfy cross-layer anchoring and risk gating, cross-layer connection, semantic compression, and unified mapping are performed to form attack semantic units, and the corresponding cross-layer connection relationships are established. The cross-layer anchoring rule is a matching rule that establishes a corresponding relationship between the forged content association layer and the phishing propagation association layer. Based on the consistency of the target object identifier and the correspondence of the collection time identifier, cross-layer matching is performed on deepfake features and phishing propagation features, and the deepfake features that have completed cross-layer matching are... The forgery features and phishing propagation features are established as cross-layer connections. The risk gating rules are based on the degree of forgery change in the forgery content evolution unit, the degree of propagation change in the propagation relationship evolution unit, and the degree of correspondence preservation in the cross-layer matching results. Forgery content evolution units, propagation relationship evolution units, and cross-layer matching results that do not meet the risk gating are blocked, while forgery content evolution units, propagation relationship evolution units, and cross-layer matching results that meet the risk gating are preserved. Cross-layer connection is to establish corresponding connection relationships for forgery content evolution units and propagation relationship evolution units that have completed cross-layer anchoring. Semantic compression is to merge and converge forgery content evolution units and propagation relationship evolution units that have completed corresponding connections. Unified mapping is to map the merged and converged units as attack nodes in the unified attack association graph. The content associations in the forged content association layer, the propagation associations in the phishing propagation association layer, and the cross-layer connection relationships are folded and integrated. The folded and integrated content associations, propagation associations, and cross-layer connection relationships are then organized into connection relationships between attack nodes to form a unified attack association graph.

[0025] This invention constructs a forged content association layer and a phishing propagation association layer, and combines cross-layer anchoring rules, risk gating rules, cross-layer connections, semantic compression, and unified mapping to form a unified attack association graph. This enables the association and folding integration of deepfake features and phishing propagation features within the same attack semantic space, improving the ability to organize associations between heterogeneous attack objects, characterize attack relationships, and unify modeling capabilities. As a result, it enhances the accuracy of deepfake content identification and the stability of attack association analysis in phishing attack scenarios.

[0026] In this embodiment, the path filtering construction module includes: Extract content associations, propagation associations, and cross-layer connections corresponding to the same object to be detected from the unified attack association graph. Expand the connection segments according to the connection order of the attack edges in the unified attack association graph. Extract the content change position and cross-layer change position from the deep forgery feature corresponding to the current object to be detected. Extract the propagation change position from the phishing propagation feature corresponding to the current object to be detected to form attack camouflage change features. Perform position matching on the expanded connection segments to form a candidate relation segment set. A stripping process is performed on each candidate relation segment in the candidate relation segment set. The stripping process involves identifying connection segments that only cover the content change location and the propagation change location but do not cover the cross-layer change location, deleting the identified connection segments, and re-establishing the connection order of the attack nodes before and after the deletion of the connection segments to form the stripped candidate relation segment set. The candidate relation segment set after stripping is subjected to retention processing. The retention processing identifies connection segments that still maintain continuous correspondence after cross-layer connection. The identified connection segments are retained and bundled according to the cross-layer connection position, the first and last connection position of the attack node, and the connection position before and after the collection time mark to form a valid relation segment set. Around the continuous cross-layer connection positions in the set of effective relationship segments, the terminating attack node of the previous effective relationship segment and the starting attack node of the next effective relationship segment are connected and spliced ​​together. The terminating attack node of the previous effective relationship segment and the starting attack node of the next effective relationship segment are established to establish a connection relationship. The connected effective relationship segments are then organized into paths according to the continuous order of content association, the continuous order of propagation association, and the continuous order of cross-layer connection to form a set of computational paths.

[0027] This invention effectively eliminates interfering connection segments that only reflect surface-level camouflage changes by performing position matching, stripping, retention, and path organization on content associations, propagation associations, and cross-layer connections in a unified attack association graph. It retains stable connection segments that maintain continuous correspondence after cross-layer connections, enhancing the accuracy of effective relationship segment identification and the targeting of path set construction. This improves the ability to identify attack camouflage changes, the stability of path filtering, and the accuracy of source association analysis in phishing attack scenarios.

[0028] In this embodiment, the similarity calculation write-back module includes: Each computation path in the computation path set is expanded into a path instance. For each attack node to be computed, a retained path instance is extracted that maintains continuous content association, continuous propagation association, and continuous cross-layer connection. The continuity of cross-layer connection positions and the corresponding situation before and after the attack node in each retained path instance are statistically analyzed to form the cross-layer anchoring preservation degree. The position matching between each retained path instance and the attack camouflage change feature is statistically analyzed to form the camouflage drift suppression value. The cross-layer anchoring preservation degree and the camouflage drift suppression value are written into the corresponding retained path instance to form the path contribution weight. Based on the path contribution weight, a weighted PathSim similarity calculation is performed. The weighted path matching sequence corresponding to any attack node to be computed is accumulated and statistically analyzed. The accumulated statistical result is normalized with the total amount of weighted path matching sequence corresponding to each attack node to be computed to obtain the similarity result of any attack node pair to be computed, forming the initial candidate association result. Contribution analysis is performed on the initial candidate association results. Each calculation path in the calculation path set is compared according to the path contribution weight to form the path contribution result. The attack nodes to be calculated are compared according to the cross-layer anchoring retention degree to form the node contribution result. The content association, propagation association and cross-layer connection are compared according to the camouflage drift suppression value to form the connection contribution result. The contribution analysis result is formed by the path contribution result, node contribution result and connection contribution result. Drift suppression processing is performed on the contribution analysis results, and the contribution analysis results are matched with the attack camouflage change features. Suppression processing is performed on the contribution part of the corresponding stripped candidate relation segment set, and retention processing is performed on the contribution part of the corresponding valid relation segment set that maintains the cross-layer anchorage retention degree, thus forming the suppressed contribution analysis results. The suppressed contribution analysis results are written back to the corresponding computation path in the computation path set. The suppressed node contribution results are written back to the corresponding attack node in the unified attack association graph. The suppressed connection contribution results are written back to the corresponding content association, propagation association and cross-layer connection in the unified attack association graph. The association retention status of the computation path set and the unified attack association graph is adjusted according to the write-back results to form the unified attack association graph with updated status. The association retention status is the correspondence between retention and removal in the computation path set and the unified attack association graph.

[0029] This invention introduces cross-layer anchoring preservation, camouflage drift suppression, path contribution weight, contribution parsing results, and contribution write-back mechanism to link and update the calculation path set and unified attack association graph. This effectively suppresses the interference of attack camouflage changes on similarity calculation, enhances the ability of weighted PathSim similarity calculation to identify stable association paths, improves the credibility of initial candidate association results, the adaptive update capability of the unified attack association graph, and the accuracy of deepfake homology analysis and attack attribution in phishing attack scenarios.

[0030] In this embodiment, the iterative calculation and update module includes: Based on the unified attack association graph after state update, the retention state retrieval is performed on the computation paths in the computation path set to identify the computation paths, attack nodes, content associations, propagation associations and cross-layer connections that maintain the corresponding relationship between retention and removal, and then reorganize them into an iterative computation path set. For computational objects in the iterative computational path set that meet the state retention condition, path instance reconstruction is performed. The path matching sequence is re-established according to the continuous order of content association, continuous order of propagation association, and continuous order of cross-layer connection. Based on the reconstructed path matching sequence, PathSim similarity calculation is re-executed to form iterative similarity results. The state retention condition is the judgment condition that the computational object still maintains the corresponding relationship between retention and removal in the unified attack association graph after the state update.

[0031] The iterative similarity results are associated and merged. According to the identifier of the object to be detected and the collection time identifier, the attack nodes corresponding to the similarity results are combined into candidate homologous objects. The content association, propagation association and cross-layer connection corresponding to the candidate homologous objects are combined into candidate association links. The candidate homologous objects and candidate association links are consistent and organized to form candidate association results. A closed-loop organization is performed on the candidate homologous objects and candidate association links in the candidate association results, and the deepfake features corresponding to the candidate homologous objects are bound to the phishing propagation features corresponding to the candidate association links to form candidate association results.

[0032] In this embodiment, the proof analysis output module includes: The candidate association results are split to extract candidate homologous objects and candidate association links. The deepfake features corresponding to the candidate homologous objects are merged and organized to form a forgery evidence set. The phishing propagation features corresponding to the candidate association links are merged and organized to form a propagation evidence set. Based on the candidate homologous objects, the forgery traces, tampering traces and inconsistency traces in the forgery evidence set are compared to form forgery verification results. Based on the candidate association links, the propagation nodes, propagation directions and propagation connection relationships in the propagation evidence set are verified to form propagation verification results. The forgery verification results and propagation verification results are bound to each other. Candidate association results that simultaneously maintain the correspondence between forged content and propagation links are confirmed. The confirmed candidate association results are converted into attack analysis results. The attack analysis results are the output results of the candidate association results after forgery evidence verification and propagation evidence verification. They represent the deep forgery association status, propagation association status and source tracing association status of the object to be detected in the phishing attack scenario, including confirmed candidate homologous objects, confirmed candidate association links and corresponding attack association conclusions.

[0033] Example 1: To verify the feasibility of this invention in practice, it was applied to a network office and financial collaboration security scenario of an enterprise. This enterprise uses email systems, online office websites, audio and video conferencing tools, instant messaging tools, and a unified identity authentication platform simultaneously. Its business processes involve high-frequency interactive behaviors such as approval notifications, financial reviews, meeting invitations, password confirmation, login verification, and file transfers. Attackers often use methods such as impersonating login pages, forging approval emails, synthesizing manager voices, splicing short video notifications, and forging chat text to carry out phishing attacks in this type of environment. Traditional solutions often only allow independent analysis of email bodies, web page links, or terminal logs. When attack samples simultaneously contain deepfake content and propagation chain disguises, problems arise such as front-end detection hitting the target but being unable to trace the source, suspicious links but being unable to confirm the source of the forgery, and scattered and difficult-to-verify identification results. To address this issue, this invention integrates deepfake features and phishing propagation features into a single attack analysis process through a multi-source data acquisition module, a feature extraction module, a correlation graph generation module, a path filtering and construction module, a similarity calculation and write-back module, an iterative calculation and update module, and a verification analysis and output module. This improves the accuracy of deepfake content identification, the stability of attack correlation analysis, and the reliability of source tracing.

[0034] In this scenario, the multi-source data acquisition module first continuously collects email data, webpage data, image data, audio data, video data, text data, log data, and infrastructure data corresponding to the object to be detected. It then performs unified identification processing, time alignment processing, and standardization processing on the collected results to form preprocessed multi-source data. After preprocessing, the feature extraction module establishes content processing sequences and propagation processing sequences around the same object to be detected. The content processing sequences are then subjected to target localization, segmentation, correspondence alignment, and difference comparison to extract forgery traces, tampering traces, and inconsistency traces, forming deepfake features. The propagation processing sequences are subjected to relationship expansion, link matching, propagation tracking, and anomaly screening to form phishing propagation features. Next, the association graph generation module compresses the deepfake features into forged content evolution units to form a forged content association layer, and compresses the phishing propagation features into propagation relationship evolution units to form a phishing propagation association layer. Finally, based on cross-layer anchoring rules and risk gating rules, the forged content evolution units and propagation relationship evolution units are cross-layer connected, semantically compressed, and uniformly mapped to form a unified attack association graph.

[0035] After the unified attack association graph is formed, the path filtering module does not directly send all connection segments into similarity calculation. Instead, it filters based on attack camouflage change features, extracting content change positions and cross-layer change positions from the deep forgery features corresponding to the current target object, and extracting propagation change positions from the phishing propagation features, forming attack camouflage change features. Then, it expands the content association, propagation association, and cross-layer connection in the unified attack association graph and performs position matching to form a candidate relationship segment set. For the candidate relationship segment set, the system identifies connection segments that only cover content change positions and propagation change positions but not cross-layer change positions and deletes them, thereby removing interference segments that only reflect surface camouflage changes. For connection segments that remain continuously corresponding after cross-layer connection, the system performs retention processing and consolidates them according to cross-layer connection positions, attack node start and end connection positions, and collection time identifier connection positions, forming a valid relationship segment set. Finally, around the continuous cross-layer connection positions, the valid relationship segments are connected and organized to form a set of calculation paths.

[0036] The similarity calculation and write-back module is the part with the most obvious verification effect in this embodiment. The system first expands each calculation path in the calculation path set into path instances, and extracts retained path instances around each attack node to be calculated, which simultaneously maintain the continuity of content association, propagation association, and cross-layer connection. Then, it statistically analyzes the continuity of cross-layer connection positions and the corresponding situation before and after the attack node in each retained path instance to form the cross-layer anchoring preservation degree. It also statistically analyzes the position matching between each retained path instance and the attack camouflage change features to form the camouflage drift suppression value. Then, it writes the cross-layer anchoring preservation degree and the camouflage drift suppression value into the corresponding retained path instance to form the path contribution weight. Finally, it performs weighted PathSim similarity calculation based on the path contribution weight to obtain the initial candidate association results. After the initial candidate association results are formed, the system continues to perform contribution parsing, separating the path contribution results, node contribution results, and connection contribution results. Then, the contribution parsing results are matched with the attack camouflage change characteristics. The contribution part corresponding to the stripped candidate relation segment set is suppressed, and the contribution part corresponding to the valid relation segment set and maintaining the cross-layer anchorage retention degree is retained. The suppressed contribution parsing results are formed, and the suppressed results are written back to the calculation path set and the unified attack association graph to form the unified attack association graph after state update.

[0037] After the unified attack association graph with updated state enters the iterative calculation and update module, the system performs a state-preserving retrieval of the calculation path. Only for calculation objects that meet the state-preserving conditions, a new path matching sequence is established and PathSim similarity calculation is performed to form candidate association results. The evidence analysis output module then splits the candidate association results into a forgery evidence set and a propagation evidence set. Forgery verification is performed around the candidate homologous objects, and propagation verification is performed around the candidate association links. The forgery verification results and propagation verification results are then bound to each other. Candidate association results that simultaneously maintain the correspondence between forged content and propagation links are confirmed to form attack analysis results. Finally, the system output is no longer just "suspicious or not", but confirmed candidate homologous objects, confirmed candidate association links, and corresponding attack association conclusions, thus meeting the requirements of integrated detection and tracing.

[0038] To ensure the authenticity of the comparison results, the same batch of input samples was used to continuously test the traditional webpage and email detection scheme, the ordinary graph association analysis scheme, and the scheme of this invention. The traditional webpage and email detection scheme is a detection scheme that performs spoofing page identification, link detection, email content filtering, and rule matching on webpage data and email data in phishing attacks, respectively. The ordinary graph association analysis scheme is an association analysis scheme that performs graph structure organization, relationship connection, and link analysis on the relationships between attack objects. The specific comparative experimental data is shown in Table 1: Table 1 Performance Evaluation of Different Solutions Against Phishing Attacks

[0039] As shown in Table 1, the solution of this invention significantly outperforms traditional webpage and email detection solutions and ordinary graph association analysis solutions in terms of accuracy in deepfake content identification, phishing attack identification, same-origin object identification, propagation link reconstruction, and dual evidence verification pass rate for forgery and propagation. Specifically, the accuracy in deepfake content identification reaches 94.8%, the accuracy in phishing attack identification reaches 95.6%, the accuracy in same-origin object identification reaches 92.3%, the propagation link reconstruction pass rate reaches 91.5%, and the dual evidence verification pass rate for forgery and propagation reaches 89.7%. At the same time, the mismatch rate of candidate association results is reduced to 4.1%, indicating that this invention can more effectively improve the deepfake identification capability, attack association analysis capability, and same-origin tracing reliability in phishing attack scenarios.

[0040] The above are merely preferred embodiments of the present invention, but the scope of protection of the present invention is not limited thereto. Any equivalent substitutions or modifications made by those skilled in the art within the scope of the technology disclosed in the present invention, based on the technical solution and inventive concept of the present invention, should be covered within the scope of protection of the present invention.

Claims

1. A deepfake detection and tracing system for phishing attacks, characterized in that, include: The multi-source data acquisition module is used to acquire multi-source data corresponding to the object to be detected and to perform data preprocessing to form preprocessed multi-source data. The feature extraction module is used to extract deepfake features and phishing propagation features from preprocessed multi-source data to form associated feature data; The association graph generation module is used to construct a forged content association layer and a phishing propagation association layer based on association feature data, and connect and fold them according to cross-layer anchoring rules to form a unified attack association graph. The path filtering construction module is used to identify effective relationship segments based on a unified attack association graph and the attack camouflage change characteristics of the current target object, and to construct a set of computational paths based on the effective relationship segments. The similarity calculation write-back module is used to perform PathSim similarity calculation on the set of calculation paths to form initial candidate association results, and to perform contribution analysis on the association calculation elements involved in the calculation. The contribution analysis results are then written back to the unified attack association graph to form the unified attack association graph after state update. The iterative computation update module is used to reconstruct the computation path set based on the unified attack association graph after the state update, and re-execute PathSim similarity calculation on computation objects that meet the state preservation condition to form candidate association results; The verification analysis output module is used to perform verification confirmation on candidate association results, verify based on forged evidence and propagation evidence, and convert the verified candidate association results into attack analysis results.

2. The deepfake detection and tracing system for phishing attacks according to claim 1, characterized in that, The multi-source data acquisition module includes: Collect raw multi-source data corresponding to the object to be detected, and classify and aggregate the data according to the data source. Record the corresponding data source identifier, object to be detected identifier and collection time identifier for each type of raw multi-source data to form a labeled raw multi-source dataset. Based on the original labeled multi-source dataset, unified labeling and time alignment are performed to form a time-consistent associated multi-source dataset. Standardization is performed on time-consistent, correlated multi-source datasets to obtain preprocessed multi-source data.

3. The deepfake detection and tracing system for phishing attacks according to claim 1, characterized in that, The feature extraction module includes: Preprocessed multi-source data is aggregated according to the identifier of the object to be detected, and the aggregated preprocessed multi-source data is arranged in chronological order according to the collection time identifier. The chronologically arranged preprocessed multi-source data is divided into content processing sequence and propagation processing sequence according to the data source. The content processing sequence is subjected to target localization, fragment segmentation, correspondence alignment and difference comparison to extract forgery traces, tampering traces and inconsistency traces to form deep forgery features. The propagation processing sequence is subjected to relationship expansion, link matching, propagation tracking and anomaly screening to form phishing propagation features. Deepfake features and phishing propagation features are associated, organized, and uniformly packaged according to the identifier of the object to be detected and the time of collection to form associated feature data.

4. The deepfake detection and tracing system for phishing attacks according to claim 1, characterized in that, The association graph generation module includes: Based on the identifier of the object to be detected and the identifier of the collection time, the deep forgery features in the associated feature data are compressed into forgery content evolution units that characterize the forgery change process, and a forgery content association layer is constructed based on the evolutionary association relationship between the forgery content evolution units. The phishing propagation features in the associated feature data are merged according to the object identifier to be detected, sorted according to the collection time identifier, and linked together according to the sequential connection relationship after sorting to form a propagation relationship evolution unit. Based on the propagation association between the propagation relationship evolution units, a phishing propagation association layer is formed. According to the cross-layer anchoring rules, the corresponding retrieval of the forged content evolution unit in the forged content association layer and the propagation relationship evolution unit in the phishing propagation association layer is performed. At the same time, according to the risk gating rules, the forged content evolution unit and propagation relationship evolution unit that have completed the corresponding retrieval are gating and filtered. For the forged content evolution unit and propagation relationship evolution unit that meet the cross-layer anchoring and risk gating, cross-layer connection, semantic compression and unified mapping are performed to form attack semantic units and establish the cross-layer connection relationship corresponding to the attack semantic units. The content associations in the forged content association layer, the propagation associations in the phishing propagation association layer, and the cross-layer connection relationships are folded and integrated. The folded and integrated content associations, propagation associations, and cross-layer connection relationships are then organized into connection relationships between attack nodes to form a unified attack association graph.

5. A deepfake detection and tracing system for phishing attacks according to claim 1, characterized in that, The path filtering construction module includes: Extract content associations, propagation associations, and cross-layer connections corresponding to the same object to be detected from the unified attack association graph. Expand the connection segments according to the connection order of the attack edges in the unified attack association graph. Extract the content change position and cross-layer change position from the deep forgery feature corresponding to the current object to be detected. Extract the propagation change position from the phishing propagation feature corresponding to the current object to be detected to form attack camouflage change features. Perform position matching on the expanded connection segments to form a candidate relation segment set. A stripping process is performed on each candidate relation segment in the candidate relation segment set. The stripping process involves identifying connection segments that only cover the content change location and the propagation change location but do not cover the cross-layer change location, deleting the identified connection segments, and re-establishing the connection order of the attack nodes before and after the deletion of the connection segments to form the stripped candidate relation segment set. The stripped candidate relation segment set is subjected to retention processing. The retention processing identifies connection segments that still maintain continuous correspondence after cross-layer connection. The identified connection segments are retained and bundled and organized according to the cross-layer connection position, the first and last connection position of the attack node, and the connection position before and after the collection time mark to form a valid relation segment set. Around the continuous cross-layer connection positions in the set of effective relationship segments, the terminating attack node of the previous effective relationship segment and the starting attack node of the next effective relationship segment are connected and spliced ​​together. The terminating attack node of the previous effective relationship segment and the starting attack node of the next effective relationship segment are established to establish a connection relationship. The connected effective relationship segments are then organized into paths according to the continuous order of content association, the continuous order of propagation association, and the continuous order of cross-layer connection to form a set of computational paths.

6. The deepfake detection and tracing system for phishing attacks according to claim 1, characterized in that, The similarity calculation write-back module includes: Each computation path in the computation path set is expanded into a path instance. For each attack node to be computed, a retained path instance is extracted that maintains continuous content association, continuous propagation association, and continuous cross-layer connection. The continuity of cross-layer connection positions and the corresponding situation before and after the attack node in each retained path instance are statistically analyzed to form the cross-layer anchoring preservation degree. The position matching between each retained path instance and the attack camouflage change feature is statistically analyzed to form the camouflage drift suppression value. The cross-layer anchoring preservation degree and the camouflage drift suppression value are written into the corresponding retained path instance to form the path contribution weight. Based on the path contribution weight, a weighted PathSim similarity calculation is performed. The weighted path matching sequence corresponding to any attack node to be computed is accumulated and statistically analyzed. The accumulated statistical result is normalized with the total amount of weighted path matching sequence corresponding to each attack node to be computed to obtain the similarity result of any attack node pair to be computed, forming the initial candidate association result. Contribution analysis is performed on the initial candidate association results. Each calculation path in the calculation path set is compared according to the path contribution weight to form the path contribution result. The attack nodes to be calculated are compared according to the cross-layer anchoring retention degree to form the node contribution result. The content association, propagation association and cross-layer connection are compared according to the camouflage drift suppression value to form the connection contribution result. The contribution analysis result is formed by the path contribution result, node contribution result and connection contribution result. Drift suppression processing is performed on the contribution analysis results, and the contribution analysis results are matched with the attack camouflage change features. Suppression processing is performed on the contribution part of the corresponding stripped candidate relation segment set, and retention processing is performed on the contribution part of the corresponding valid relation segment set that maintains the cross-layer anchorage retention degree, thus forming the suppressed contribution analysis results. The contribution rewrite is performed on the suppressed contribution analysis results, and the association retention state of the computation path set and the unified attack association graph is adjusted according to the rewrite results to form the unified attack association graph after state update.

7. A deepfake detection and tracing system for phishing attacks according to claim 1, characterized in that, The iterative calculation and update module includes: Based on the unified attack association graph after state update, the retention state retrieval is performed on the computation paths in the computation path set to identify the computation paths, attack nodes, content associations, propagation associations and cross-layer connections that maintain the corresponding relationship between retention and removal, and then reorganize them into an iterative computation path set. For computational objects in the iterative computation path set that meet the condition of preserving the state, perform path instance reconstruction, re-establish the path matching sequence according to the continuous order of content association, continuous order of propagation association, and continuous order of cross-layer connection, and re-execute PathSim similarity calculation based on the reconstructed path matching sequence to form iterative similarity results.

8. The iterative similarity results are associated and merged. According to the object to be detected and the collection time, the attack nodes corresponding to the similarity results are combined into candidate homologous objects. The content association, propagation association and cross-layer connection corresponding to the candidate homologous objects are combined into candidate association links. The candidate homologous objects and candidate association links are consistent and organized to form candidate association results. A closed-loop organization is performed on the candidate homologous objects and candidate association links in the candidate association results, and the deepfake features corresponding to the candidate homologous objects are bound to the phishing propagation features corresponding to the candidate association links to form candidate association results.

9. A deepfake detection and tracing system for phishing attacks according to claim 1, characterized in that, The proof analysis output module includes: The candidate association results are split to extract candidate homologous objects and candidate association links. The deepfake features corresponding to the candidate homologous objects are merged and organized to form a forgery evidence set. The phishing propagation features corresponding to the candidate association links are merged and organized to form a propagation evidence set. Based on the candidate homologous objects, the forgery traces, tampering traces and inconsistency traces in the forgery evidence set are compared to form forgery verification results. Based on the candidate association links, the propagation nodes, propagation directions and propagation connection relationships in the propagation evidence set are verified to form propagation verification results. The system binds the forged verification results and the propagation verification results accordingly. It confirms the candidate association results that simultaneously maintain the correspondence between the forged content and the propagation link, and converts the confirmed candidate association results into attack analysis results.