Dynamic anti-fraud decision-making method based on adversarial evolution
By constructing an attacker model to generate adversarial samples with multi-dimensional perturbation features and dynamically training them with the defender model, the problem of decreased recognition accuracy of existing anti-fraud models when facing complex fraud methods is solved, a dynamic anti-fraud decision-making method based on adversarial evolution is realized, and the recognition ability and robustness of the model are improved.
Patent Information
- Application Number
- CN202510880331.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-27
- Publication Date
- 2025-10-10
AI Technical Summary
The existing anti-fraud models have reduced recognition accuracy when faced with complex and covert online fraud methods. In addition, the existing adversarial training methods lack a dynamic game mechanism, resulting in the generated adversarial samples being unable to effectively improve the model's recognition capabilities, and ignoring the semantic masking and word order perturbation characteristics at the language level.
A dynamic anti-fraud decision-making method based on adversarial evolution is constructed. The attacker model is used to generate adversarial samples with sensitive word disguise, word order perturbation, and semantic masking features. The attacker model is combined with the defender model for supervised training. A game mechanism is introduced to optimize the attacker model. The defender model with a Transformer architecture is used for identification. The model robustness is improved through dynamic purification and semantic compensation mechanisms.
It achieves highly robust recognition of the disguised features of real fraudulent language. Through attack-defense co-evolution and semantic-layer perturbation adversarial training, it significantly improves the stability and robustness of the model when facing highly disguised fraudulent texts, and achieves efficient recognition and precise interception of new types of online fraud behaviors.
Smart Images

Figure CN120764597A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of intelligent anti-fraud technology, and in particular to a dynamic anti-fraud decision-making method based on adversarial evolution. Background Art
[0002] Online fraud tactics are becoming increasingly sophisticated and covert. Fraudsters are gradually abandoning traditional overt tactics and instead employing covert tactics such as obfuscated language, disguised identities, and emotional manipulation. This has led to a decline in the recognition accuracy of existing anti-fraud models in practical applications. Currently, mainstream anti-fraud methods rely on manual rule-based keyword matching, shallow semantic model classification, or historical corpus retrospective analysis. However, these methods often struggle to achieve stable and robust recognition of fraudulent texts that are highly disguised and contain volatile language.
[0003] On the other hand, existing text generation strategies capable of adversarial perturbation are used to simulate variations in fraudulent language, thereby testing the anti-fraud system's ability to withstand attacks. However, in existing technologies, attack sample construction and defense model training are often separated, lacking a dynamic game mechanism. As a result, the generated adversarial samples cannot effectively guide the model's recognition capabilities. Furthermore, existing adversarial training methods focus on improving adversarial robustness at the model structure level, ignoring common variational features in real fraud scenarios, such as semantic masking, word order perturbations, and stylistic disguises at the linguistic level. Summary of the Invention
[0004] To solve the above problems, the present invention provides a dynamic anti-fraud decision-making method based on adversarial evolution.
[0005] To achieve the above object, the technical solution adopted by the present invention is:
[0006] A dynamic anti-fraud decision-making method based on adversarial evolution, comprising:
[0007] S1. Extract fraud conversation texts from fraud case data and construct an original fraud sample dataset.
[0008] S2. Construct an attacker model based on the original fraud sample dataset. The attacker model uses a guiding prompt word generation strategy to call a language generation model to output a set of adversarial samples with features such as sensitive word disguise, word order perturbation, and semantic masking.
[0009] S3. Construct a hybrid training dataset based on the adversarial sample set and the original fraud sample dataset, input the dataset into the supervised training of the defender model for fraud intent identification, and output an identification label;
[0010] S4, comparing the identification tag with the true label of the adversarial sample, performing optimization on the attacker model prompt according to the identification accuracy, updating the adversarial sample set and re-executing S3 until the round reaches a preset number or the change rate of the attack success rate in a preset window is less than a preset change rate threshold, and then stopping executing S3 to obtain a trained defender model;
[0011] S5, performing anti-fraud identification and matching decision through the trained defender model.
[0012] Further, the S1 comprises the following steps:
[0013] Periodically acquire the original data file containing the fraud event record through the anti-fraud database interface;
[0014] Preprocess the text content in the original data file, including sentence division, stop word removal, uniform format marking and role dialogue extraction, to obtain structured dialogue corpus;
[0015] Perform semantic segmentation and intent annotation operations on the structured dialogue corpus, extract language segments highly associated with fraud behavior, and construct a labeled fraud sample corpus set;
[0016] Cluster the labeled fraud sample corpus set according to time window and fraud technique category to generate an original fraud sample data set for downstream adversarial sample generation.
[0017] Further, the attacker model is used to perform the following steps:
[0018] Based on the content and label information of each fraud text in the original fraud sample data set, a guiding prompt is constructed, which contains semantic transformation requirements, fraud behavior disguise strategies and current typical technique characteristics;
[0019] Splice the guiding prompt template and the corresponding fraud sample to form an input sequence, and call a pre-trained language model to perform text generation on the input sequence to output a preliminary generated adversarial text candidate set;
[0020] Based on each text in the adversarial text candidate set, sequentially perform word-level perturbation processing, syntax-level perturbation processing and semantic-level perturbation processing to construct adversarial samples with multi-dimensional perturbation features;
[0021] Perform rule-driven compliance screening on the adversarial samples with multi-dimensional perturbation features, eliminate samples that are not in compliance with the format, have incoherent semantics or do not achieve the attack purpose, and generate an adversarial sample set.
[0022] Furthermore, the step of sequentially performing lexical-level perturbation processing, grammatical-level perturbation processing, and semantic-level perturbation processing on each text in the adversarial text candidate set includes:
[0023] Based on the lexical content of each adversarial text candidate, high-risk keywords are identified, and synonym replacement, character insertion, character deletion, and case character replacement operations are performed on the high-risk keywords to obtain an intermediate text containing lexical-level perturbation features;
[0024] Based on the syntactic structure of the intermediate text, the subject, predicate, and object components of the sentence and the modifier structure are identified, and word order transformation, sentence reconstruction, and grammatical fuzzification operations are performed to obtain a second-stage text containing grammatical-level perturbation features;
[0025] Based on the semantic expression of the text in the second stage, non-critical background descriptions, positive evaluation content or logical disguised sentences are injected to enhance the non-fraud representation and semantic induction effect of the text, and obtain adversarial samples with multi-dimensional perturbation features.
[0026] Furthermore, the step S3 includes the following steps:
[0027] Based on the adversarial sample set and the original fraud sample dataset, a mixed training dataset is constructed according to a preset ratio, and the training dataset is batch-divided and vector-encoded;
[0028] Input the encoded training data into the defender model, which is built based on the language model of the Transformer architecture, and sets its preset layer parameters to a frozen state for supervised training, and outputs the recognition label;
[0029] The cross entropy loss calculation is performed based on the output recognition label to obtain the loss result, and the model parameters are back-propagated based on the loss result to obtain the defender model.
[0030] Furthermore, the defender model includes:
[0031] The feature extraction module is used to encode the input training data based on the transformer architecture to obtain high-dimensional feature representation;
[0032] A dynamic purification module, configured to introduce a small perturbation to the high-dimensional feature representation, calculate the reverse gradient after the perturbation to obtain a feature sensitivity map, filter highly sensitive areas according to the feature sensitivity map, and perform a zeroing operation;
[0033] A semantic compensation module, configured to perform weighted superposition on the compensation values calculated based on the attention weights of the highly sensitive areas, restore semantic integrity, and output a purified feature representation;
[0034] A classifier is used to generate a fraud intent identification label based on the purified feature representation.
[0035] Furthermore, the dynamic purification module is used to perform the following steps:
[0036] generating a random noise tensor having the same shape as the high-dimensional feature representation based on the high-dimensional feature representation, and adding the random noise tensor to the high-dimensional feature representation to obtain a perturbation feature representation;
[0037] Performing cross entropy loss calculation based on the perturbation feature representation and the corresponding error label, and performing backpropagation according to the loss value to obtain the perturbation gradient tensor;
[0038] Calculating the L2 norm of each feature dimension based on the perturbation gradient tensor to obtain a feature sensitivity map representing the sensitivity of each feature position;
[0039] Feature positions with sensitivities higher than a preset threshold are identified according to the feature sensitivity map, and a zeroing operation is performed on feature dimensions corresponding to the feature positions to obtain a partially suppressed feature representation.
[0040] Furthermore, the semantic compensation module is used to perform the following steps:
[0041] Based on the partially suppressed feature representation, calling the attention mechanism module to extract the feature dimensions that are not set to zero in the context window and their corresponding attention weights, and constructing a candidate compensation area;
[0042] Performing weighted summation on the feature representations of the candidate compensation regions according to the attention weights to generate a compensation vector for semantic restoration;
[0043] The compensation vector is mapped to the corresponding zeroed feature position and fused with the original zeroed representation to output a purified feature representation.
[0044] Furthermore, the optimization of the attacker model prompt according to the recognition accuracy includes:
[0045] The recognition label results of the adversarial sample set based on the defender model are compared with the true labels of the adversarial samples to determine whether each adversarial sample is correctly identified and construct a feedback sequence containing attack success and failure marks;
[0046] Match the feedback sequence with the guiding prompt words, perturbation strategy, and generation semantic information during the generation process of each adversarial sample, and construct an adversarial trajectory sample set containing generation status, guiding prompts, and reward signals;
[0047] Based on the adversarial trajectory sample set, a policy gradient algorithm is used to optimize the guiding prompts in the attacker model.
[0048] Furthermore, the preset window is 5 rounds, and the preset change rate threshold is 5%.
[0049] The beneficial effects of the present invention are as follows: the present invention constructs a dynamic anti-fraud decision-making method with anti-evolution capability, introduces a game mechanism between the attacker model and the defender model, and realizes highly robust recognition of the disguised features of real fraud language. First, an original fraud sample data set is constructed through historical fraud cases to provide a corpus basis for subsequent adversarial sample generation and model training. Then, an attacker model is constructed based on the language generation model, and combined with a guiding prompt strategy, a multi-dimensional adversarial sample with sensitive word disguise, word order perturbation and semantic masking features is generated to restore the hidden variants of fraudulent speech in real scenarios. Further, a training data set is constructed by mixing adversarial samples with original samples, and input into the defender model for supervised training, and the model's recognition accuracy of adversarial samples is used as feedback to optimize the attacker model prompt, thereby realizing dynamic evolution between attack and defense. This mechanism overcomes the problem of separation between attack sample construction and defense model training in the existing technology, so that the generated adversarial samples continuously reveal the vulnerability of the defender model in each round of evolution, and promote the gradual enhancement of the model's recognition capability. Furthermore, the Defender model incorporates dynamic purification and semantic compensation mechanisms. By suppressing and restoring the gradient-sensitive regions of perturbed features, the model significantly improves its stability and robustness against highly disguised fraudulent texts. Through attack-defense co-evolution and semantic-layer perturbation adversarial training, the defense model's recognition capabilities are effectively enhanced, enabling efficient identification and precise interception of new types of online fraud. BRIEF DESCRIPTION OF THE DRAWINGS
[0050] Figure 1 This is a flowchart of the steps of a dynamic anti-fraud decision-making method based on adversarial evolution in the present invention.
[0051] Figure 2 It is a flow chart of the execution steps of the dynamic purification module in the present invention. DETAILED DESCRIPTION
[0052] See also Figure 1-Figure 2 As shown, the present invention relates to a dynamic anti-fraud decision-making method based on adversarial evolution, comprising:
[0053] S1. Extract fraud conversation texts from fraud case data and construct an original fraud sample dataset.
[0054] S2. Construct an attacker model based on the original fraud sample dataset. The attacker model uses a guiding prompt word generation strategy to call a language generation model to output a set of adversarial samples with features such as sensitive word disguise, word order perturbation, and semantic masking.
[0055] S3, constructing a mixed training data set based on the set of adversarial samples and the original fraud sample data set, inputting into a defender model supervised training for fraud intent recognition, and outputting an identification label;
[0056] S4, comparing the identification label with the true label of the adversarial sample, performing optimization on the attacker model prompt according to the identification accuracy, updating the set of adversarial samples and re-executing S3 until the round reaches a preset number or the change rate of the attack success rate within a preset window is less than a preset change rate threshold, and then stopping executing S3, to obtain a trained defender model;
[0057] S5, performing anti-fraud identification and matching decision through the trained defender model.
[0058] In some embodiments, first, based on the data resources of the anti-fraud data center interface, the dialogue text in real fraud cases over the years is periodically extracted, including common fraud scenarios such as "impersonating customer service", "fake investment", "brushing rebates", etc. The above text is structured and the speech fragments of the fraud initiator are separated, the role identity is uniformly formatted, and the original fraud sample dataset is constructed combined with the intent label, providing corpus support for subsequent adversarial sample generation and recognition model training. Based on the above original fraud sample dataset, an attacker model is constructed. The attacker model is not simply static perturbation, but uses a guided prompt word generation strategy to reconstruct the text combined with a language generation model (such as the GPT series). Unlike the existing technology of directly replacing keywords or splicing templated sentences, the present scheme generates disguised rhetoric with stronger semantic deception through customized prompt guidance model, such as expressing "high yield guarantee" as "project profit stable but need to keep secret", and at the same time performing syntax transformation to make the fraud text not only formally deviate from the existing template, but also achieve concealment and misdirection in semantic expression. This guided generation mechanism reflects the high flexibility and simulation of the present scheme in the attack sample construction stage, which is significantly different from traditional perturbation methods. Subsequently, the adversarial samples generated by the attacker model are mixed with the original fraud samples in proportion to construct a training set for training the defender model. During the training process, the defender model uses the Transformer architecture, keeps the pre-training ability by freezing part of the layer parameters, and strengthens the recognition ability of semantic variants after introducing adversarial data. During the model training process, the recognition accuracy of the model on the adversarial samples is used as feedback to optimize the attacker model prompt, thereby realizing the dynamic evolution between attack and defense. The attacker model improves the text disguise ability in each round of optimization, and the defender model optimizes the recognition ability of complex rhetoric in continuous learning, forming an attack-defense co-evolution path. Finally, when the attacker model cannot significantly improve the attack success rate (i.e. the probability of the attack sample being recognized tends to be stable) in several rounds of confrontation, the training process is terminated, and the defender model with strong robustness and generalization ability is output. The defense model can be deployed and applied in real-time scenarios to identify fraud intent in user input dialogue content and output matching response strategies or alarm results, realizing dynamic and accurate anti-fraud decision support. Different from the existing technology: 1. In the attack sample construction process, the language generation model is guided instead of simple text perturbation, achieving higher simulation and language deception; 2. The attack and defense process forms a closed-loop linkage, and the model co-evolution is driven by the game mechanism, which is different from the static mode of the existing attack-defense decoupling in adversarial training; 3. The dynamic termination mechanism is introduced, which makes the training process self-adaptively converge after meeting the robustness improvement target, improving the model deployment practicability.
[0059] Further, the S1 comprises the following steps:
[0060] Periodically obtain original data files containing fraud event records through the anti-fraud database interface;
[0061] Preprocessing the text content in the original data file, including sentence segmentation, stop word removal, unified formatting and character dialogue extraction, to obtain structured dialogue data;
[0062] Performing semantic segmentation and intent labeling operations on the structured dialogue corpus, extracting language segments that are highly correlated with fraudulent behavior, and constructing a labeled fraud sample corpus;
[0063] The labeled fraud sample corpus is clustered according to time windows and fraud speech categories to generate an original fraud sample dataset for downstream adversarial sample generation.
[0064] In some embodiments, raw data files containing fraud incidents are periodically pulled. The files are usually in a structured or semi-structured format (such as JSON, CSV, XML), and the content includes fields such as case number, conversation content, role labeling, and timestamps. For the "call log" or "text conversation" field in the raw data, the basic cleaning process in natural language processing is used to pre-process the text content, specifically including: regularizing the sentence segmentation (based on punctuation and grammatical structure), performing stop word list filtering (to exclude high-frequency words with weak semantic load such as "的", "是", "啊", etc.), unifying the conversation format (standardizing role labels, such as "fraud party" and "victim"), and extracting role conversations through regular expressions and semantic templates, and finally obtaining structured conversation data segmented by role. Subsequently, for the structured conversation data, a sequence labeling model based on the BiLSTM-CRF structure is introduced to perform semantic segmentation on each sentence, identifying the semantic units therein, such as "excuse opening", "emotional manipulation", "fund inducement", "transfer request" and other fraudulent intent components. Based on this, combined with existing anti-fraud intent labeling systems (such as intent classifiers trained with BERT or ERN IE models), intent labeling is performed on each semantic unit, resulting in fraudulent sentence fragments labeled with intent categories, forming a labeled fraud sample corpus. To enhance the coverage and semantic diversity of this sample corpus, the system grouped these samples using a dual clustering strategy. First, a sliding time window was constructed based on the time of the fraud incident to extract time series distribution characteristics and avoid corpus bias caused by sample concentration in specific historical periods. Second, a sentence vector representation based on BERT encoding was used, and cosine similarity was introduced as a similarity metric to semantically cluster the sentences using the K-means clustering algorithm. Each category in the clustering results corresponds to a common fraudulent phrasing variant, such as "impersonating customer service to notify a refund" and "part-time job fraud promises cash back." The resulting raw fraud sample dataset, with good semantic discrimination and label integrity, is stored in a database with a structure of "sample text - intent label - cluster category - time label," serving as a foundation for the attacker model to generate guiding cues and attack text.
[0065] Furthermore, the attacker model is used to perform the following steps:
[0066] Based on the content and label information of each fraudulent text in the original fraud sample dataset, construct guiding prompts, which include semantic transformation requirements, fraudulent behavior disguise strategies and current typical speech characteristics;
[0067] The guiding prompt template and the corresponding fraud sample are spliced together to form an input sequence, and a pre-trained language model is called to perform text generation on the input sequence, outputting a preliminary generated adversarial text candidate set;
[0068] Based on each text in the adversarial text candidate set, lexical-level perturbation processing, grammatical-level perturbation processing, and semantic-level perturbation processing are performed in sequence to construct adversarial samples with multi-dimensional perturbation features;
[0069] A rule-driven compliance screening is performed on the adversarial samples of the multi-dimensional perturbation features to eliminate samples that are format-incompliant, semantically incoherent, or fail to achieve the attack purpose, and generate an adversarial sample set.
[0070] In some embodiments, first, based on each labeled fraud text in the original fraud sample data set, a guiding prompt is constructed, including: semantic transformation requirements (such as "please keep the original meaning, but disguise the keywords"), fraud disguise strategies (such as "covering the transfer request with a customer service tone"), and typical fraud speech features (such as "brushing rebates" and "unfreezing accounts" and other expressions). The template is designed with parameters so that it can automatically be spliced with the fraud sample content to generate an input sequence. For example: Input sequence construction: Prompt = "Disguise the following fraud text with XX to make it difficult to be identified as a fraudulent behavior: {original fraud text}" Input sequence construction: Prompt = "Disguise the following fraud text with XX to make it difficult to be identified as a fraudulent behavior: {original fraud text}" Then, the input sequence is input into a pre-trained language model (such as ChatGLM, GPT or Qwen), and the text generation task is performed to obtain a preliminary adversarial text candidate set. The language model parses the semantic instructions in the Prompt and outputs disguised text with word order changes and diversified expressions. After generating the preliminary adversarial text, the candidate samples are further subjected to multi-level perturbation processing, including three stages: lexical level, grammatical level, and semantic level. To ensure the availability and offensiveness of adversarial samples, the system designs a rule-driven compliance screening mechanism. This mechanism includes three types of checking rules: format legitimacy: regular rules are used to check whether the text meets the basic grammatical structure, and samples that are obviously incoherent or generated by the model are eliminated; semantic consistency: a semantic similarity matching model (such as SBERT) is introduced to calculate the semantic overlap between the adversarial sample and the original sample, ensuring that the adversarial text achieves surface disguise while maintaining the semantic core. Through the above steps, the attacker model not only has the ability to control language diversity, but also can generate high-quality adversarial samples that are closer to real scenarios through semantic perturbation strategies. This is significantly different from traditional static samples generated by simple word replacement or word order disruption. It reflects the creativity of this solution in the multi-dimensional perturbation modeling capability and dynamic guided generation mechanism in the construction dimension of adversarial samples, providing more challenging input corpus for subsequent defense model training.
[0071] Furthermore, the step of sequentially performing lexical-level perturbation processing, grammatical-level perturbation processing, and semantic-level perturbation processing on each text in the adversarial text candidate set includes:
[0072] Based on the lexical content of each adversarial text candidate, high-risk keywords are identified, and synonym replacement, character insertion, character deletion, and case character replacement operations are performed on the high-risk keywords to obtain an intermediate text containing lexical-level perturbation features;
[0073] Based on the syntactic structure of the intermediate text, the subject, predicate, and object components of the sentence and the modifier structure are identified, and word order transformation, sentence reconstruction, and grammatical fuzzification operations are performed to obtain a second-stage text containing grammatical-level perturbation features;
[0074] Based on the semantic expression of the text in the second stage, non-critical background descriptions, positive evaluation content or logical disguised sentences are injected to enhance the non-fraud representation and semantic induction effect of the text, and obtain adversarial samples with multi-dimensional perturbation features.
[0075] In some embodiments, first, during the lexical perturbation processing phase, the system uses a high-risk word identification algorithm (e.g., based on TF-IDF weighting or a predefined sensitive word dictionary) to extract keywords from the adversarial text candidates, screening out highly sensitive words with fraudulent potential, such as "transfer," "verification code," and "account freeze." For these keywords, the following operations are applied: Synonym replacement: A semantic similarity retrieval model based on Word2Vec or BERT embedding vector space is used to replace words with similar semantics but different expressions, such as replacing "transfer" with "fund movement"; Character perturbation: Key characters are slightly damaged by character insertion (e.g., "转*账"), deletion (e.g., "转账"), and replacement (e.g., "账") to create morphological disturbances; Character deformation: Full-width and half-width switching, Unicode variants, phonetic combinations, or similar characters (e.g., "zhuǎnzhàng") are used to further circumvent the keyword identification algorithm. The above processing produces a lexically perturbed intermediate text. While maintaining a certain level of readability, it significantly reduces the recognition capabilities of static dictionary-based fraud prevention models. Next, the grammatical perturbation processing phase begins. This phase uses dependency syntactic analysis of the intermediate text (e.g., constructing a dependency tree using tools like spaCy or Stanza) to identify the sentence's basic subject, predicate, and object components, as well as their modifier structure. Restructuring the sentence structure is then performed: Word order is shifted: the subject, predicate, and object components are rationally rearranged. For example, "Please transfer the money to the designated account as soon as possible" can be changed to "Please complete the transfer to the designated account as soon as possible." Sentence structure is rewritten: declarative sentences are rewritten into imperative, interrogative, or parallel variations. For example, "I'll help you with your account problem" becomes "Do you need my assistance with your account?" Grammatical fuzzification: Using the subjunctive mood, adverbial modifiers, and the passive voice to weaken the statement, for example, by adding words like "maybe," "suggest," and "maybe" to construct ambiguous statements. These manipulations alter the text's structure while preserving the core fraudulent intent. This increases the attack strength of the generated samples against structure-sensitive models, and outputs a second-stage text containing grammatical-level perturbation features. Finally, semantic-level perturbation processing is performed. This stage no longer directly rewrites the original text content. Instead, it employs a semantic injection mechanism to enhance the "non-fraudulent" appearance of the utterance, thereby misleading defense models: Injecting background descriptions: Neutral background information unrelated to the fraudulent intent is introduced through template expansion or generative models, such as "I am an employee of the XX platform, and the system has recently been upgraded." Positive comments: Including commendatory statements to reduce alertness, such as "Thank you very much for your patience" and "This is standard procedure." Logical disguise statements: Inserting logical loops or misleading explanations, such as "The funds were frozen because you did not complete the authentication process."
[0076] Furthermore, the step S3 includes the following steps:
[0077] Based on the adversarial sample set and the original fraud sample dataset, a mixed training dataset is constructed according to a preset ratio, and the training dataset is batch-divided and vector-encoded;
[0078] Input the encoded training data into the defender model, which is built based on the language model of the Transformer architecture, and sets its preset layer parameters to a frozen state for supervised training, and outputs the recognition label;
[0079] The cross entropy loss calculation is performed based on the output recognition label to obtain the loss result, and the model parameters are back-propagated based on the loss result to obtain the defender model.
[0080] In some embodiments, the system first proportionally fuses the previously generated adversarial sample set with the original fraud sample dataset to construct a hybrid training dataset. This ratio is generally set to 30%-50% adversarial samples to ensure the model maintains its ability to recognize both real fraudulent speech and potential language variants. After data fusion, the training set undergoes batch partitioning and vector encoding. The batch partitioning utilizes a dynamic padding strategy to ensure consistent sequence length within each batch, thereby improving GPU acceleration efficiency. The encoding phase uses pre-trained embedders such as BERT or RoBERTa to convert the text into a fixed-dimensional high-order vector representation, providing the input format for downstream model processing. Next, the encoded training data is fed into the defender model. This model is built based on the Transformer architecture, comprising an embedding layer, a multi-head attention mechanism layer, a feedforward neural network layer, and a classification output layer. To ensure training stability and preserve semantic representation, the first N layers of the model employ a frozen parameter strategy, locking their pre-trained weights and updating only the top classifier. This strategy effectively avoids overfitting of the underlying language representation and ensures robust feature extraction despite the semantic perturbations of adversarial examples. During model inference, the output layer classifies the input sample for fraudulent intent and generates a label. This label is compared with the sample's true label to calculate the cross entropy loss. Based on the error calculation results, the system executes a backpropagation algorithm and, combined with gradient descent or the Adam optimizer, updates the weights of the model's unfrozen layers, thereby optimizing the model's recognition capabilities.
[0081] Furthermore, the defender model includes:
[0082] The feature extraction module is used to encode the input training data based on the transformer architecture to obtain high-dimensional feature representation;
[0083] A dynamic purification module, configured to introduce a small perturbation to the high-dimensional feature representation, calculate the reverse gradient after the perturbation to obtain a feature sensitivity map, filter highly sensitive areas according to the feature sensitivity map, and perform a zeroing operation;
[0084] A semantic compensation module, configured to perform weighted superposition on the compensation values calculated based on the attention weights of the highly sensitive areas, restore semantic integrity, and output a purified feature representation;
[0085] A classifier is used to generate a fraud intent identification label based on the purified feature representation.
[0086] It should be noted that, first, the feature extraction module is implemented based on the Transformer architecture. It processes the input text data (including the original fraudulent samples and adversarial samples) through word embedding encoding and a multi-layer self-attention mechanism, and outputs context-related feature representations with consistent dimensions. The output of this module is in the form of a high-dimensional tensor, which preserves the structural dependencies and contextual association information of the text at the semantic level. The dynamic purification module introduces controllable perturbations to the feature extraction results. Specifically, it uses a noise tensor with the same shape as the input to form a perturbation representation. The reverse gradient of this perturbation representation under the target error label is calculated to obtain the sensitivity of each dimension of the feature to the output change. Then, the sensitivity map is calculated through L2 norm calculation, and a threshold is set to filter out highly sensitive areas. The feature channels corresponding to these areas are zeroed to achieve targeted suppression of semantic interference. This process strengthens the model's robust discrimination ability for key features and avoids being misled by deceptive local disguised perturbations in adversarial samples. Next, the semantic compensation module is responsible for restoring the semantic information lost due to perturbation suppression. This module relies on the attention weight distribution in the original Transformer structure, selects unmasked key features in the context of the zeroed area, and generates a compensation vector through attention weighting. The compensation vector is then fused with the zeroed result, so that the purified feature representation not only weakens the impact of the attack disturbance, but also retains enough semantic clues to maintain the discrimination ability. Finally, the classifier receives the above-mentioned purified feature representation and outputs the predicted label of fraudulent intent through the feedforward network. As the decision-making end of the overall model, this part uses the high-quality representation after purification and compensation for supervised training, so as to maintain a high recognition accuracy when facing adversarial speech with highly concealed strategies (such as word order deformation, sensitive word masking, etc.).
[0087] Furthermore, the dynamic purification module is used to perform the following steps:
[0088] generating a random noise tensor having the same shape as the high-dimensional feature representation based on the high-dimensional feature representation, and adding the random noise tensor to the high-dimensional feature representation to obtain a perturbation feature representation;
[0089] perform cross-entropy loss calculation based on the perturbed feature representation and the corresponding wrong label, and perform back propagation according to the loss value to obtain a perturbation gradient tensor;
[0090] calculate L2 norm of each feature dimension according to the perturbation gradient tensor to obtain a feature sensitivity map representing sensitivity of each feature position;
[0091] identify feature positions with sensitivity higher than a preset threshold according to the feature sensitivity map, and perform a zeroing operation on feature dimensions corresponding to the feature positions to obtain a partially inhibited feature representation.
[0092] It should be noted that, first, based on the high-dimensional feature representation output by the Transformer model, a random noise tensor consistent with the shape thereof is constructed. The noise tensor generally obeys a normal distribution and is used to simulate the input perturbation environment that the model may encounter in real application. The noise tensor is added element by element to the original feature representation to generate a perturbed feature representation to introduce response fluctuations of potential deceptive variants. Subsequently, in order to measure the degree of deviation of the model under the perturbation condition, the dynamic purification module targets a specific misleading label, performs a forward propagation on the perturbed feature representation and calculates a cross-entropy loss value. The loss value reflects the strength of the model being misled after the perturbation. On this basis, the perturbation gradient tensor is obtained through back propagation technology, thereby revealing which feature dimensions play a leading role in promoting the model misjudgment process. Further, the system calculates the L2 norm of each feature dimension according to the perturbation gradient tensor as a sensitivity index thereof under the current perturbation situation. The sensitivity values of all feature dimensions constitute a feature sensitivity map consistent with the shape of the original feature. The map is essentially a positioning heat map that can accurately depict which feature positions pose the greatest challenge to the model robustness. Finally, the dynamic purification module filters out feature dimensions higher than the threshold from the map based on a set sensitivity threshold, and performs a zeroing operation on the values of these high-sensitivity regions in the feature space to generate a partially inhibited feature representation. The inhibition strategy effectively weakens the interference of deceptive expressions in fraudulent texts on the model judgment result, thereby improving the stability and reliability of the overall recognition. Through the above perturbation purification mechanism based on gradient sensitivity feedback, the vulnerable feature regions in the recognition process of the deep neural model can be structurally remodeled, and the dynamic adaptive ability against adversarial perturbations can be realized at the language representation level. This mechanism breaks through the traditional idea of relying on static feature elimination or surface rule filtering, and is a processing method of establishing a game feedback closed loop at the feature space level.
[0093] Further, the semantic compensation module is configured to perform the following steps:
[0094] Based on the partially inhibited feature representation, a attention mechanism module is called to extract the feature dimensions within the context window that are not zeroed and their corresponding attention weights, and to construct a candidate compensation region;
[0095] The feature representation of the candidate compensation region is weighted and summed according to the attention weights to generate a compensation vector for semantic repair;
[0096] The compensation vector is mapped to the corresponding zeroed feature position and fused with the original zeroed representation to output the purified feature representation.
[0097] Specifically, first, the partially inhibited feature representation output by the dynamic purification module is received, in which some dimensions are zeroed due to high sensitivity, which may cause the breakage of the original semantic chain. To avoid semantic sparsity causing model understanding imbalance, the semantic compensation module takes the feature representation as input, calls the self-attention mechanism in the Transformer architecture, scores the attention of each feature dimension within the context window that is not zeroed, extracts its weight value and original representation, and constructs a candidate compensation region. In the candidate compensation region, the system performs attention weighting operation on all un-inhibited feature representations, that is, the corresponding features are weighted and summed according to the attention weights, thereby generating a compensation vector that is structurally matched with the zeroed region. The vector carries representative semantic information in the context and aims to compensate for the representation loss caused by the elimination of sensitive areas through the comprehensive representation of the surrounding stable areas. Then, the compensation vector is mapped to the corresponding position of the zeroed high-dimensional feature space in the original, and a fusion operation is performed with the inhibited representation, such as residual connection or weighted superposition, to complete compensation injection. The fusion operation not only retains the information of the non-sensitive areas of the original features, but also injects context-driven semantic compensation signals, finally outputting a purified feature representation that is structurally complete and expressively coherent.
[0098] Further, the optimization of the attacker model prompt according to the recognition accuracy comprises:
[0099] Based on the identification label results of the defender model on the set of adversarial samples, the true labels of the adversarial samples are compared to determine whether each adversarial sample is correctly identified, and a feedback sequence containing attack success and failure marks is constructed;
[0100] The feedback sequence is corresponded to the guiding prompt words, perturbation strategies and generated semantic information in the generation process of each adversarial sample to construct an adversarial trajectory sample set containing generation state, guiding prompt and reward signal;
[0101] Based on the adversarial trajectory sample set, the guiding prompt in the attacker model is optimized using a policy gradient algorithm.
[0102] In some embodiments, the optimization process of the attacker model is designed to improve its ability to generate adversarial samples to deceive the defender model, and the core lies in the strategic iterative optimization of guiding prompts. First, the system compares the identification labels output by the defender model with the true labels of the adversarial samples one by one to determine whether the model accurately identifies the corresponding fraudulent intentions. For samples successfully identified by the defender model, they are marked as attack failures; samples that cannot be identified are marked as attack successes, thereby generating a sequence of attack success labels containing 0 / 1 binary feedback. This sequence is input into the optimization process as a reward feedback signal. Subsequently, the system aligns the feedback sequence with the guiding prompts, perturbation strategy selection path, generated semantic feature vectors, etc. used in the generation process of each adversarial sample to construct a complete set of adversarial trajectory samples. Each trajectory sample can be regarded as a state-action-reward triple. After the construction is completed, the policy gradient algorithm is introduced to perform gradient updates in order to maximize the long-term reward. The key advantage of this method is that it regards the generation of adversarial samples as a strategy sampling problem in reinforcement learning, guiding the attacker model to gradually explore more confusing prompt word combinations in different samples and contexts, thereby forming an adaptive attack evolution mechanism for the game environment, effectively improving the realism and deceptiveness of adversarial text, and breaking through the limitations of traditional static data enhancement strategies in adversarial scenarios.
[0103] Furthermore, the preset window is 5 rounds, and the preset change rate threshold is 5%.
[0104] Specifically, a preset window of five rounds is set. After completing five rounds of attack-defense iterations, the relative change in the success rate of the current round is compared with the previous round. If the change in attack success rate for five consecutive rounds is less than 5% (i.e., the preset change rate threshold), the attacker's model strategy is considered to have converged and the defender's model's recognition capability has reached a stable state. The system then terminates the training phase and outputs the final defense model. This mechanism ensures both the attacker and defender have flexible response space during the evolution process while ensuring reasonable convergence in terms of actual training time and computing resources. This distinguishes itself from traditional static training models that rely solely on single-round optimization error judgment, enhancing the system's stability and scalability in real-world deployments.
[0105] The above embodiments are merely descriptions of preferred embodiments of the present invention and are not intended to limit the scope of the present invention. Without departing from the design spirit of the present invention, various modifications and improvements made to the technical solutions of the present invention by ordinary engineering technicians in this field should fall within the scope of protection determined by the claims of the present invention.
Claims
1. A dynamic anti-fraud decision-making method based on adversarial evolution, characterized in that: include: S1. Extract fraud conversation texts from fraud case data and construct an original fraud sample dataset. S2. Construct an attacker model based on the original fraud sample dataset. The attacker model uses a guiding prompt word generation strategy to call a language generation model to output a set of adversarial samples with features such as sensitive word disguise, word order perturbation, and semantic masking. S3. Construct a hybrid training dataset based on the adversarial sample set and the original fraud sample dataset, input the dataset into the supervised training of the defender model for fraud intent identification, and output an identification label; S4. Compare the identified label with the true label of the adversarial sample, optimize the attacker model prompt based on the recognition accuracy, update the adversarial sample set, and re-execute S3 until the number of rounds reaches a preset number or the change rate of the attack success rate within a preset window is less than a preset change rate threshold. S3 is stopped and the trained defender model is obtained. S5. Perform anti-fraud identification and matching decisions through the trained defender model.
2. A dynamic anti-fraud decision-making method based on adversarial evolution according to claim 1, characterized in that: Said S1 comprises the following steps: Periodically obtain original data files containing fraud event records through the anti-fraud database interface; Preprocessing the text content in the original data file, including sentence segmentation, stop word removal, unified formatting and character dialogue extraction, to obtain structured dialogue data; Performing semantic segmentation and intent labeling operations on the structured dialogue corpus, extracting language segments that are highly correlated with fraudulent behavior, and constructing a labeled fraud sample corpus; The labeled fraud sample corpus is clustered according to time windows and fraud speech categories to generate an original fraud sample dataset for downstream adversarial sample generation.
3. The dynamic anti-fraud decision-making method based on adversarial evolution according to claim 1 is characterized in that: The attacker model is used to perform the following steps: Based on the content and label information of each fraudulent text in the original fraud sample dataset, construct guiding prompts, which include semantic transformation requirements, fraudulent behavior disguise strategies and current typical speech characteristics; The guiding prompt template and the corresponding fraud sample are spliced together to form an input sequence, and a pre-trained language model is called to perform text generation on the input sequence, outputting a preliminary generated adversarial text candidate set; Based on each text in the adversarial text candidate set, lexical-level perturbation processing, grammatical-level perturbation processing, and semantic-level perturbation processing are performed in sequence to construct adversarial samples with multi-dimensional perturbation features; A rule-driven compliance screening is performed on the adversarial samples of the multi-dimensional perturbation features to eliminate samples that are format-incompliant, semantically incoherent, or fail to achieve the attack purpose, and generate an adversarial sample set.
4. The dynamic anti-fraud decision-making method based on adversarial evolution according to claim 3 is characterized in that: The step of sequentially performing lexical-level perturbation processing, grammatical-level perturbation processing, and semantic-level perturbation processing on each text in the adversarial text candidate set includes: Based on the lexical content of each adversarial text candidate, high-risk keywords are identified, and synonym replacement, character insertion, character deletion, and case character replacement operations are performed on the high-risk keywords to obtain an intermediate text containing lexical-level perturbation features; Based on the syntactic structure of the intermediate text, the subject, predicate, and object components of the sentence and the modifier structure are identified, and word order transformation, sentence reconstruction, and grammatical fuzzification operations are performed to obtain a second-stage text containing grammatical-level perturbation features; Based on the semantic expression of the text in the second stage, non-critical background descriptions, positive evaluation content or logical disguised sentences are injected to enhance the non-fraud representation and semantic induction effect of the text, and obtain adversarial samples with multi-dimensional perturbation features.
5. The dynamic anti-fraud decision-making method based on adversarial evolution according to claim 3 is characterized in that: The S3 includes the following steps: Based on the adversarial sample set and the original fraud sample dataset, a mixed training dataset is constructed according to a preset ratio, and the training dataset is batch-divided and vector-encoded; Input the encoded training data into the defender model, which is built based on the language model of the Transformer architecture, and sets its preset layer parameters to a frozen state for supervised training, and outputs the recognition label; The cross entropy loss calculation is performed based on the output recognition label to obtain the loss result, and the model parameters are back-propagated based on the loss result to obtain the defender model.
6. The dynamic anti-fraud decision-making method based on adversarial evolution according to claim 5 is characterized in that: The defender model includes: The feature extraction module is used to encode the input training data based on the transformer architecture to obtain high-dimensional feature representation; A dynamic purification module, configured to introduce a small perturbation to the high-dimensional feature representation, calculate the reverse gradient after the perturbation to obtain a feature sensitivity map, filter highly sensitive areas according to the feature sensitivity map, and perform a zeroing operation; A semantic compensation module, configured to perform weighted superposition on the compensation values calculated based on the attention weights of the highly sensitive areas, restore semantic integrity, and output a purified feature representation; A classifier is used to generate a fraud intent identification label based on the purified feature representation.
7. The dynamic anti-fraud decision-making method based on adversarial evolution according to claim 6 is characterized in that: The dynamic purification module is used to perform the following steps: generating a random noise tensor having the same shape as the high-dimensional feature representation based on the high-dimensional feature representation, and adding the random noise tensor to the high-dimensional feature representation to obtain a perturbation feature representation; Performing cross entropy loss calculation based on the perturbation feature representation and the corresponding error label, and performing backpropagation according to the loss value to obtain the perturbation gradient tensor; Calculating the L2 norm of each feature dimension based on the perturbation gradient tensor to obtain a feature sensitivity map representing the sensitivity of each feature position; Feature positions with sensitivities higher than a preset threshold are identified according to the feature sensitivity map, and a zeroing operation is performed on feature dimensions corresponding to the feature positions to obtain a partially suppressed feature representation.
8. The dynamic anti-fraud decision-making method based on adversarial evolution according to claim 7 is characterized in that: The semantic compensation module is used to perform the following steps: Based on the partially suppressed feature representation, calling the attention mechanism module to extract the feature dimensions that are not set to zero in the context window and their corresponding attention weights, and constructing a candidate compensation area; Performing weighted summation on the feature representations of the candidate compensation regions according to the attention weights to generate a compensation vector for semantic restoration; The compensation vector is mapped to the corresponding zeroed feature position and fused with the original zeroed representation to output a purified feature representation.
9. The dynamic anti-fraud decision-making method based on adversarial evolution according to claim 3 is characterized in that: The optimization of the attacker model prompt according to the recognition accuracy includes: The recognition label results of the adversarial sample set based on the defender model are compared with the true labels of the adversarial samples to determine whether each adversarial sample is correctly identified and construct a feedback sequence containing attack success and failure marks; Match the feedback sequence with the guiding prompt words, perturbation strategy, and generation semantic information during the generation process of each adversarial sample, and construct an adversarial trajectory sample set containing generation status, guiding prompts, and reward signals; Based on the adversarial trajectory sample set, a policy gradient algorithm is used to optimize the guiding prompts in the attacker model.
10. The dynamic anti-fraud decision-making method based on adversarial evolution according to claim 1 is characterized in that: The preset window is 5 rounds, and the preset change rate threshold is 5%.
Citation Information
Cited By
Large-scale correction model-oriented camouflage disturbance method and system
CN121071943A
A camouflage perturbation method and system for large-scale grading models
CN121071943B
Active defense method and system based on large model
CN121396685A