Fraud identification method and device, electronic equipment and readable storage medium

CN122824422APending Publication Date: 2026-09-25CHINA MOBILE GROUP ZHEJIANG +2
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610675200.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-05-15
Publication Date
2026-09-25

AI Technical Summary

Technical Problem

[0003]在相关技术中,涉诈用户识别通常依赖静态规则匹配和人工审查,在处理大规模、复杂数据时,对涉诈用户识别的准确率和效率较低

Benefits of technology

本申请实施例提供了一种涉诈用户识别方法,通过获取目标用户在预设时间范围内的电信网络行为序列和用户特征,基于该用户特征和电信网络行为序列,确定全量特征向量,该全量特征向量表征该用户特征、以及该目标用户的电信网络行为的内在风险状态演化规律,然后通过将该全量特征向量分别输入第一预测模型和第二预测模型,得到第一预测模型输出的与目标用户对应的第一涉诈结果、以及第二预测模型输出的与目标用户对应的第二涉诈结果,该第一预测模型为基于多维度特征能够实现分类功能的非量子模型,该第二预测模型为基于多维度特征能够实现分类功能的量子模型,基于该第一涉诈结果和第二涉诈结果,确定与目标用户对应的决策结果,该决策结果用于指示该目标用户是否为涉诈用户。本申请的方案,通过使用能够较为全面的对目标用户进行刻画的全量特征向量,以及利用量子模型(即上述第二预测模型)与经典机器学习模型(即上述第一预测模型)的互补性,能够提高涉诈用户识别的效率和准确率。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122824422A_ABST
    Figure CN122824422A_ABST
Patent Text Reader

Abstract

The application discloses a fraud-involved user identification method and device, electronic equipment and a readable storage medium. The method comprises the following steps: acquiring a telecommunication network behavior sequence and user characteristics of a target user within a preset time range; determining a full-amount feature vector based on the user characteristics and the telecommunication network behavior sequence, wherein the full-amount feature vector represents the user characteristics and the internal risk state evolution law of the telecommunication network behavior of the target user; inputting the full-amount feature vector into a first prediction model and a second prediction model respectively to obtain a first fraud-involved result corresponding to the target user output by the first prediction model and a second fraud-involved result corresponding to the target user output by the second prediction model, wherein the first prediction model is a non-quantum model capable of realizing classification function based on multi-dimensional characteristics, and the second prediction model is a quantum model capable of realizing classification function based on multi-dimensional characteristics; and determining a decision result corresponding to the target user based on the first fraud-involved result and the second fraud-involved result.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of cybersecurity, and in particular to a method, apparatus, electronic device, and readable storage medium for identifying users suspected of fraud. Background Technology

[0002] In recent years, telecommunications fraudsters have adopted various forms and technologies, greatly increasing the difficulty of combating fraud and making it difficult for victims to defend themselves.

[0003] In related technologies, the identification of users suspected of fraud usually relies on static rule matching and manual review. When dealing with large-scale and complex data, the accuracy and efficiency of identifying users suspected of fraud are relatively low. Summary of the Invention

[0004] This application discloses a method, apparatus, electronic device, and readable storage medium for identifying users suspected of fraud, which can improve the accuracy and efficiency of identifying users suspected of fraud.

[0005] To solve the above problems, this application adopts the following technical solution: In a first aspect, embodiments of this application disclose a method for identifying users suspected of fraud, comprising: acquiring a telecommunications network behavior sequence and user characteristics of a target user within a preset time range; determining a full feature vector based on the user characteristics and the telecommunications network behavior sequence, wherein the full feature vector characterizes the user characteristics and the inherent risk state evolution law of the target user's telecommunications network behavior; obtaining a first fraud-related result corresponding to the target user output by the first prediction model and a second fraud-related result corresponding to the target user output by the second prediction model by respectively inputting the full feature vector into a first prediction model and a second prediction model, wherein the first prediction model is a non-quantum model that can achieve classification function based on multi-dimensional features, and the second prediction model is a quantum model that can achieve classification function based on multi-dimensional features; determining a decision result corresponding to the target user based on the first fraud-related result and the second fraud-related result, wherein the decision result is used to indicate whether the target user is a user suspected of fraud.

[0006] Secondly, this application discloses a device for identifying users suspected of fraud, comprising: an acquisition module for acquiring a sequence of telecommunications network behaviors and user characteristics of a target user within a preset time range; a determination module for determining a full feature vector based on the user characteristics and the sequence of telecommunications network behaviors, wherein the full feature vector characterizes the user characteristics and the inherent risk state evolution law of the target user's telecommunications network behaviors; a obtaining module for obtaining a first fraud-related result corresponding to the target user output by the first prediction model and a second fraud-related result corresponding to the target user output by the second prediction model by inputting the full feature vector into a first prediction model and a second prediction model respectively, wherein the first prediction model is a non-quantum model that can achieve classification function based on multi-dimensional features, and the second prediction model is a quantum model that can achieve classification function based on multi-dimensional features; the determination module is further configured to determine a decision result corresponding to the target user based on the first fraud-related result and the second fraud-related result, wherein the decision result is used to indicate whether the target user is a user suspected of fraud.

[0007] Thirdly, embodiments of this application provide an electronic device including a processor and a memory, the memory storing programs or instructions executable on the processor, the programs or instructions, when executed by the processor, implementing the steps of the method described in the first aspect.

[0008] Fourthly, embodiments of this application provide a readable storage medium on which a program or instructions are stored, which, when executed by a processor, implement the steps of the method described in the first aspect.

[0009] Fifthly, embodiments of this application provide a computer program product, the computer program product including a computer program stored on a non-transitory computer-readable storage medium, the computer program including program instructions, which, when executed by a computer, cause the computer to perform the steps of the method described in the first aspect.

[0010] The technical solution adopted in this application can achieve the following beneficial effects: This application provides a method for identifying users suspected of fraud. It acquires a sequence of telecommunications network behaviors and user characteristics of a target user within a preset time range. Based on these user characteristics and the telecommunications network behavior sequence, a full feature vector is determined. This full feature vector characterizes the user characteristics and the inherent risk state evolution law of the target user's telecommunications network behavior. Then, by inputting this full feature vector into a first prediction model and a second prediction model, a first fraud-related result corresponding to the target user is obtained from the output of the first prediction model, and a second fraud-related result corresponding to the target user is obtained from the output of the second prediction model. The first prediction model is a non-quantum model that can achieve classification based on multi-dimensional features, while the second prediction model is a quantum model that can achieve classification based on multi-dimensional features. Based on the first and second fraud-related results, a decision result corresponding to the target user is determined. This decision result indicates whether the target user is a fraud-related user. This application's solution, by using a full feature vector that can comprehensively characterize the target user and utilizing the complementarity between the quantum model (i.e., the aforementioned second prediction model) and the classical machine learning model (i.e., the aforementioned first prediction model), can improve the efficiency and accuracy of identifying users suspected of fraud. Attached Figure Description

[0011] Figure 1 This is a flowchart illustrating a method for identifying users suspected of fraud, as disclosed in an embodiment of this application. Figure 2 This is a flowchart illustrating a quantum algorithm feature combination selection method disclosed in an embodiment of this application; Figure 3 This is a flowchart of a method for identifying fraudulent users disclosed in an embodiment of this application; Figure 4 This is a schematic diagram of the structure of a fraud-related user identification device disclosed in an embodiment of this application; Figure 5 This is a schematic diagram of the structure of an electronic device disclosed in an embodiment of this application. Detailed Implementation

[0012] The technical solutions of the embodiments of this application will be clearly described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this application. All other embodiments obtained by those skilled in the art based on the embodiments of this application are within the scope of protection of this application.

[0013] The terms "first," "second," etc., used in the specification and claims of this application are used to distinguish similar objects and not to describe a specific order or sequence. It should be understood that such terms can be used interchangeably where appropriate so that embodiments of this application can be implemented in orders other than those illustrated or described herein, and the objects distinguished by "first," "second," etc., are generally of the same class and the number of objects is not limited; for example, a first object can be one or more. Furthermore, in the specification and claims, "and / or" indicates at least one of the electrically connected objects, and the character " / " generally indicates that the preceding and following objects are in an "or" relationship.

[0014] The following description, in conjunction with the accompanying drawings, details the fraud-related user identification method, apparatus, electronic device, and readable storage medium disclosed in this application through specific embodiments and application scenarios.

[0015] This application discloses a method for identifying users suspected of fraud. Figure 1 This is a flowchart illustrating a method for identifying fraudulent users disclosed in an embodiment of this application. Figure 1 As shown, the method includes the following steps: S120. Obtain the telecommunications network behavior sequence and user characteristics of the target user within a preset time range.

[0016] The attribute information of each telecommunications network behavior in the telecommunications network behavior sequence may include behavior type, identification information, behavior duration, behavior occurrence time, and base station location at the time of behavior occurrence.

[0017] In this application, the preset time range can be a preset observation period. For example, the telecommunications network behavior sequence of target user u within one observation period is as follows: ,in, The total number of actions of the target user within the observation period, and each telecommunications network action in the telecommunications network action sequence. It is a tuple containing K attributes. ,in, It can be categorized by behavior type such as visiting websites or answering phone calls. It can provide identifying information such as URLs (Uniform Resource Locators) and mobile number categories. It can measure the duration of actions such as website visit duration and call duration. It can be based on the time when the action occurs. It can be the location of the base station when the action occurs.

[0018] In this application, the user characteristics of the target user u This can include age, gender, number of calls, types of websites visited, and statistical measures such as mean, variance, and standard deviation.

[0019] S140. Based on the user characteristics and the telecommunications network behavior sequence, determine the full feature vector, wherein the full feature vector represents the user characteristics and the inherent risk state evolution law of the target user's telecommunications network behavior.

[0020] This application is based on the full feature vector determined from the user characteristics of the target user and the telecommunications network behavior sequence. It can provide a relatively comprehensive profile of the target user u.

[0021] S160. By inputting the full feature vector into the first prediction model and the second prediction model respectively, a first fraud-related result corresponding to the target user is obtained from the output of the first prediction model, and a second fraud-related result corresponding to the target user is obtained from the output of the second prediction model. The first prediction model is a non-quantum model that can achieve classification function based on multi-dimensional features, and the second prediction model is a quantum model that can achieve classification function based on multi-dimensional features.

[0022] In one implementation, the first fraud result may include a first probability and a first prediction label, wherein the first probability is the probability that the first prediction model believes the target user is involved in fraud, and the first prediction label is used to indicate whether the first prediction model believes the target user is a fraudulent user. The second fraud result may include a second prediction label and a second probability, wherein the second prediction label is used to indicate whether the second prediction model believes the target user is a fraudulent user, and the second probability is the probability that the second prediction model believes the target user is involved in fraud.

[0023] For example, the first prediction model This can be the LightGBM model, XBboost model, random forest model, etc. Taking the LightGBM model as the first prediction model as an example, let's denote this first prediction model as... Where d represents the full feature vector The feature dimensions. The first prediction model receives this full feature vector. Output the first probability , ,in, These are the model parameters of the LightGBM model, and the corresponding first predicted label. , This is the classification threshold. It should be noted that the first predicted label here is a hard label.

[0024] For example, the second prediction model This can be a quantum support vector machine (QSVM), which uses the full feature vector. The second prediction model is first input through a quantum feature mapping. The full feature vector Mapped to a higher-dimensional Hilbert space quantum states in The decision function that QSVM solves in this space is: ,in, It is a set of support vectors. These are the corresponding Lagrange multipliers. These are the labels for support vectors (+1 corresponds to fraud). These are samples of support vectors. It is a bias term. It is the quantum kernel function, and the final second prediction label output by the second prediction model. Furthermore, the decision value is transformed into a second probability using a scaled sigmoid function. Where A is the preset calibration parameter. It should be noted that the second prediction label here is a hard label.

[0025] The high parallelism and processing speed of quantum computing enable the rapid analysis of large amounts of data, thereby more effectively identifying potential fraud patterns and risky behaviors, and improving the accuracy and real-time performance of fraud detection.

[0026] S180. Based on the first fraud-related result and the second fraud-related result, determine a decision result corresponding to the target user, wherein the decision result is used to indicate whether the target user is a fraud-related user.

[0027] This application determines the decision result corresponding to the target user based on the first fraud-related result output by the first prediction model and the second fraud-related result output by the second prediction model. By utilizing the complementarity of the quantum model (i.e., the second prediction model) and the classical machine learning model (i.e., the first prediction model), the former uses quantum kernel functions to process high-dimensional dynamic features to accurately match risk patterns, while the latter efficiently integrates all features to detect generalized anomalies, forming a complementary dual-channel discrimination mechanism. This improves the efficiency and accuracy of identifying fraud-related users, thereby reducing economic losses caused by fraudulent activities. Furthermore, this application's solution, through a hybrid intelligent system that integrates the high parallelism of quantum computing with the stability of classical computing, enables deeper, more dynamic, and more efficient probabilistic modeling and real-time identification of fraudulent user behavior on telecommunications networks, thereby improving the accuracy, real-time performance, adaptability, and interpretability of fraud detection.

[0028] This application provides a method for identifying users suspected of fraud. It acquires a sequence of telecommunications network behaviors and user characteristics of a target user within a preset time range. Based on these user characteristics and the telecommunications network behavior sequence, a full feature vector is determined. This full feature vector characterizes the user characteristics and the inherent risk state evolution law of the target user's telecommunications network behavior. Then, by inputting this full feature vector into a first prediction model and a second prediction model, a first fraud-related result corresponding to the target user is obtained from the output of the first prediction model, and a second fraud-related result corresponding to the target user is obtained from the output of the second prediction model. The first prediction model is a non-quantum model that can achieve classification based on multi-dimensional features, while the second prediction model is a quantum model that can achieve classification based on multi-dimensional features. Based on the first and second fraud-related results, a decision result corresponding to the target user is determined. This decision result indicates whether the target user is a fraud-related user. This application's solution, by using a full feature vector that can comprehensively characterize the target user and utilizing the complementarity between the quantum model (i.e., the aforementioned second prediction model) and the classical machine learning model (i.e., the aforementioned first prediction model), can improve the efficiency and accuracy of identifying users suspected of fraud.

[0029] In this embodiment of the application, determining the full feature vector based on the user characteristics and the telecommunications network behavior sequence in step S140 may include the following steps S141 to S143.

[0030] S141. By performing quantum state mapping on each telecommunications network behavior in the telecommunications network behavior sequence, a quantum state sequence corresponding to the telecommunications network behavior sequence is obtained, wherein the quantum state sequence carries the correlation information between two adjacent telecommunications network behaviors in the telecommunications network behavior sequence, and the correlation information includes semantic correlation and temporal correlation.

[0031] In this application, the telecommunications network behavior sequence of the target user u within a preset time range is obtained. Next, data preprocessing is performed, and the preprocessed telecommunications network behavior is denoted as... ,in, These are the preprocessed feature dimensions. It should be noted that preprocessing here may include cleaning missing values, handling outliers, encoding categorical variables (such as website categories), and normalizing numerical variables (such as duration).

[0032] This application defines a quantum system consisting of n qubits, whose Hilbert space... The dimension is This application computes each ground state of this space. (in Assigning a specific, fine-grained semantic label to a behavior is a method that predefines the behavior through clustering and semantic analysis of massive amounts of historical behavior data. For example: - "Briefly visit dating websites on rest days from 12:00 to 18:00"; - "Spending extended periods of time accessing overseas gambling websites between 8:00 PM and midnight"; - "Frequently receiving text calls from unknown numbers during weekdays from 8:00 AM to 6:00 PM"; - etc.

[0033] For preprocessed user telecommunications network behavior This application does not rigidly assign it to a specific ground state, but rather maps it to a quantum state composed of the superposition of all relevant ground states to characterize the fuzziness and multifaceted nature of the behavior. The mapping is performed through a classical trainable neural network. accomplish, ,in, It is A complex vector of dimension 1. The vector is ensured to satisfy the normalization condition through a softmax function and a phase generation network. Then telecommunications network behavior corresponding quantum state Defined as: , where the coefficient The model This indicates that the behavior is related to the semantic ground state. The strength of the association, for example, the act of "visiting a vague financial website on a weekday evening" may be associated with multiple ground states such as "nighttime activities", "financial behavior", and "information gathering" with different probability amplitudes.

[0034] To capture common cross-timestep correlation patterns in fraudulent activities at the data level (e.g., receiving a suspicious phone call followed immediately by visiting a phishing website link), this application explicitly introduces quantum entanglement in the coding and designs a parameterized two-body quantum gate circuit. Quantum state pairs acting on adjacent time steps Above, among which, These are trainable parameters. Specifically, this application uses a series of controlled rotating doors to construct... For each pair of ground states that may have semantic association (The correlation is discovered through prior knowledge), a controlled Y-turn gate is applied. Its function is: if the first quantum state of the previous quantum state... One component is excited (i.e., in the state of...). (state), then for the second quantum state, the first... Each component is applied an angle. Y-axis rotation. Adjusted through training. The system can learn the strong positive correlation between specific combinations of behaviors (such as "behavior A" and "behavior B"). This means that after behavior A occurs, the probability of behavior B occurring will increase, which can be used to encode the sequence of key steps in a fraud behavior pattern.

[0035] After pairwise entanglement operations on the entire sequence, a representation of a many-body quantum state sequence with a specific correlation pattern in the time dimension is finally obtained. For the target user u, its complete encoded output is a quantum state sequence. , , of which each yes A normalized vector. On a classical computer, it stores the coefficient vector for each state. It should be noted that, due to the introduction of quantum entanglement, the resulting quantum state sequence It also carries information on the correlation and strength between two adjacent telecommunications network behaviors in the telecommunications network behavior sequence. This association information includes semantic association and temporal association.

[0036] This application embeds key domain knowledge (such as fraud script steps) into the underlying data representation as trainable parameters by semantic and temporal correlations between the coding behaviors of quantum entangled circuits. This provides inputs rich in intrinsic correlations for subsequent quantum algorithms and enhances the model's ability to learn complex patterns from the source.

[0037] S142. Based on the quantum state sequence, determine the dynamic feature vector corresponding to the target user, wherein the dynamic feature vector captures the inherent risk state evolution law of the target user's telecommunications network behavior.

[0038] In one implementation, determining the dynamic feature vector corresponding to the target user based on the quantum state sequence may include: determining the dynamic feature vector corresponding to the target user based on the quantum state sequence, the model parameters of the normal model, and the model parameters of the fraud model, wherein the normal model is a model trained on a quantum hidden Markov model through a normal user behavior sequence, and the fraud model is a model trained on a quantum hidden Markov model through a fraudulent user behavior sequence.

[0039] This application uses the Quantum Hidden Markov Model (QHMM) as the generative model. Through a pre-set quantum probability model that simulates the dynamic evolution of behavior, it analyzes the hidden risk state change patterns in the user's telecommunications network behavior and extracts a set of quantitative and discriminable high-dimensional dynamic features as dynamic feature vectors corresponding to the target user.

[0040] This application describes user behavior using a QHMM, which consists of the following five parts: (1) Set of hidden states This is the model's assumption of the user's inherent risk state, of which there are N possibilities. Each state is represented by a quantum state. express All states span an N-dimensional Hilbert space. For example, in an anti-fraud scenario, this application defines: : Normal state - Represents the baseline of risk-free behaviors such as daily browsing and social interaction.

[0041] : Exploratory state - Represents initial behaviors that may be guided, such as traffic generation and information search.

[0042] Interactive state - Represents a state of continuous interaction with high-risk objects (such as fraudulent phone calls or fraudulent links).

[0043] : Execution state - Represents performing sensitive operations, such as transferring funds, disclosing verification codes, and other critical actions.

[0044] : Hidden state - represents a sudden period of inactivity after a high-risk operation.

[0045] These states are internal to the model, abstractions that cannot be directly observed, but correspond to interpretable risk stages.

[0046] (2) Set of observed states This refers to the quantum state sequence obtained above, which is the quantum state encoded by each specific telecommunications network behavior of the user. All possible observation states constitute the observation Hilbert space. Its dimensions are (n is the number of encoded qubits).

[0047] (3) Initial state distribution This is an N-dimensional probability vector. ,in, This indicates the state at the beginning of the sequence. The probability satisfies .

[0048] (4) State transition matrix This is A matrix whose elements Represents the state from time t Transition to state at time t+1 The probability. In a classic Hidden Markov Model (HMM), this is a random matrix; in a QHMM, By a quantum operation (quantum channel) To describe this, the operation acts on the density matrix of the hidden state. superior, ,in, It is the density matrix representation of the classical probability distribution of the hidden state at time t in quantum form. It is a set of conditions Kraus operator. It can be obtained through calculation. This allows for both classical randomness and the possibility of coherent superposition in state transitions.

[0049] (5) Emission probability matrix B: This is a key component that correlates the hidden state with the observed state, and its elements This indicates that the system is in a hidden state. When a specific quantum state is observed The probability density. Within the quantum framework, this is determined by a set of POVM (Positive Operator-Valued Measure) operators. To define, ,in, It is a state of observation The relevant positive definite operators are usually designed as follows: ,in, It is a hybrid parameter used to control the accuracy and noise tolerance of the measurement.

[0050] In this application, the training process of the normal model obtained by training QHMM with normal user behavior sequences and the fraudulent model obtained by training QHMM with fraudulent user behavior sequences are as follows: Given a set of training sequences The training objective is to maximize the likelihood probability of the observed sequence, along with its labels (normal or fraudulent). Since direct optimization is difficult, a variational posterior distribution is introduced. To approximate the true hidden state sequence posterior distribution And iteratively optimize the following lower bounds of evidence:

[0051] By alternating between the expected step and the maximization step, convergence is achieved.

[0052] Expected step: Fixed model parameters Updating variational distribution using quantum forward-backward algorithm This makes it approximate the true posterior under the current model. This algorithm can efficiently calculate the "marginal posterior probability" at all time points. and "transfer expectation count" .

[0053] Maximize step: fixed Update model parameters To maximize It is usually transformed into utilization and The parameters are re-estimated using the statistics. For example, the new initial distribution estimate is... The new transition probability related parameters are from Updated in the statistics.

[0054] After training and obtaining the above-mentioned normal model and fraud model, determine the model parameters of the normal model. and the model parameters of the fraud model. Then, based on the above quantum state sequence Model parameters of a normal model and the model parameters of the fraud model. Determine the dynamic feature vector corresponding to the target user. This dynamic feature vector captures the risk evolution patterns, certainty, and anomalies of users' telecommunications network behavior over time.

[0055] It should be noted that the dynamic feature vector corresponding to the target user Based on log-likelihood ratio (LLR) and quantum state sequence Average risk level of T quantum states The percentage of total time spent in high-risk states Temporal average entropy of posterior distribution From the sequence, from the "exploratory state / interactive state" ( ) is transferred to the "execution state" ( ) number of times and the abnormal score of the transition probability Confirmed. Model parameters for the normal model. Used to determine dynamic feature vectors The log-likelihood ratio (LLR) in the model is related to the model parameters of the fraud model. Used to determine dynamic feature vectors The log-likelihood ratio (LLR) and the average risk level of T quantum states in a quantum state sequence are shown in the figure. The percentage of total time spent in high-risk states Temporal average entropy of posterior distribution Dynamic feature vectors The log-likelihood ratio (LLR) in LLR is based on the sequence of quantum states. Probabilities generated by fraud models and quantum state sequences The quantum state sequence is determined by the probability generated by the normal model. The probability generated by the fraud model is based on a sequence of quantum states. Model parameters of fraud models Confirmed, quantum state sequence The probability generated by the normal model is based on the quantum state sequence. Model parameters of the normal model Determine the dynamic feature vector. Quantum state sequence Average risk level of T quantum states Based on posterior probability Determine the posterior probability. Based on quantum state sequences Model parameters of fraud models Determine the dynamic feature vector. The percentage of total time spent in high-risk states Based on posterior probability Determine the dynamic feature vector. Temporal average entropy of the posterior distribution in Based on posterior probability Sure.

[0056] Among them, determining the dynamic feature vector corresponding to the target user. The process is as follows: (1) Determine the log-likelihood ratio (LLR).

[0057] ),in, Represents the observation sequence (i.e., the above quantum state sequence) The probability (likelihood) generated by the fraud model. Represents the observation sequence (i.e., the above quantum state sequence) LLR is the probability (likelihood) generated by a normal model. LLR > 0 indicates that the sequence is more consistent with the fraud model, and LLR much greater than 0 is a strong fraud signal.

[0058] (2) Determine the occupancy statistic of risk status.

[0059] ① Based on posterior probability Calculate the average risk level of T quantum states in a quantum state sequence. ,in, It is to assign a state Predefined risk weights, for example, .

[0060] ② High-risk state (such as execution state) (Total duration percentage) It should be noted that the high-risk status here is predefined.

[0061] (3) Determine the time-averaged entropy of the posterior distribution .

[0062] ,in, It is the probability of being in state i at time t, and the entropy. To measure the uncertainty of the state at that moment. Low entropy indicates that user behavior patterns are clear and predictable. Fraudulent behavior may exhibit low entropy (strong purposefulness) during the interaction and execution phases, but may have high entropy (chaotic behavior) during the investigation phase.

[0063] (4) Determine the anomalies of state transitions.

[0064] Model parameters based on the normal model using the quantum Viterbi algorithm Decode the most likely hidden state sequence .

[0065] ① Determine the number of critical transition triggers. Calculate the sequence from "exploratory state / interaction state" ( ) is transferred to the "execution state" ( ) number of times Normal users rarely make such critical transfers.

[0066] ② Transition probability anomaly score ,in, It is the transition probability from state i to j in the normal model. This score measures the degree of abnormality of the user's actual state transition path in the view of the normal model. The higher the score, the more abnormal the transition pattern.

[0067] Based on the above calculations, the scalar characteristic log-likelihood ratio (LLR) and quantum state sequence are obtained. Average risk level of T quantum states The percentage of total time spent in high-risk states Temporal average entropy of posterior distribution From the sequence, from the "exploratory state / interactive state" ( ) is transferred to the "execution state" ( ) number of times and the abnormal score of the transition probability Determine the dynamic feature vector corresponding to the target user. .

[0068] This application's solution considers that fraudulent behavior is a low-probability, high-harm anomaly pattern within the probability distribution of normal behavior. Furthermore, its patterns often manifest as atypical correlations and state transitions of specific behavioral units over time. Therefore, more efficiently and profoundly modeling and extracting these atypical correlations and state transition patterns from massive user behavior data can effectively improve the accuracy of identifying users suspected of fraud. Moreover, this application utilizes QHMM, a generative model with explicit probabilistic interpretation, to determine dynamic feature vectors with clear physical meaning. The features in these dynamic feature vectors directly correlate with risk evolution stages, exhibiting high information density and strong interpretability.

[0069] In addition, this application introduces a quantum hidden Markov model, which utilizes the superposition property of quantum states to evaluate the probability of multiple behavioral evolution paths simultaneously with exponential parallelism in principle. This enables more refined and efficient modeling of the micro-dynamics of user risk states and has the potential to accelerate exponentially when dealing with long-sequence complex patterns.

[0070] S143. Based on the user characteristics and the dynamic feature vector, determine the full feature vector.

[0071] User characteristics based on target user u and dynamic feature vectors Determine the full feature vector ,in, , This represents the feature dimension of the full feature vector. Feature dimensions representing user characteristics This represents the feature dimension of a dynamic feature vector.

[0072] In this embodiment of the application, determining the decision result corresponding to the target user based on the first fraud-related result and the second fraud-related result may include: when the first predicted label and the second predicted label are the same, determining a target predicted label corresponding to the target user based on the first predicted label and the second predicted label, wherein the target predicted label is used to indicate whether the target user is a fraud-related user; and determining a target confidence level corresponding to the target predicted label based on the first probability and the second probability.

[0073] If the first prediction model outputs the first predicted label The second predicted label output by the second prediction model Same, that is If the first prediction model and the second prediction model reach a consensus, then the target predicted label corresponding to the target user u is considered to be... Target confidence level corresponding to the target prediction label , The first probability is the output of the first prediction model. This is the second probability output by the second prediction model.

[0074] Then predict the label of the target. and target confidence The data is transmitted to downstream fraud early warning systems, triggering corresponding handling procedures (such as monitoring, alerting, and interception). Simultaneously, the processed samples and their final confirmed tags are collected for periodic incremental updates to the base model. and This allows the system to adapt to the ever-evolving methods of fraud.

[0075] In this embodiment of the application, determining the decision result corresponding to the target user based on the first fraud-related result and the second fraud-related result may include: when the first predicted label and the second predicted label are different, determining a meta-feature vector based on the first probability, the second probability, the classification accuracy of the first prediction model, the classification accuracy of the second prediction model, and the dynamic feature vector corresponding to the target user, wherein the dynamic feature vector is determined based on the telecommunications network behavior sequence; obtaining the target probability output by the meta-learner by inputting the meta-feature vector, wherein the target probability is the probability that the meta-learner considers the target user to be involved in fraud; and determining a target predicted label corresponding to the target user based on the target probability, wherein the target predicted label is used to indicate whether the target user is a fraud-related user.

[0076] If the first prediction model outputs the first predicted label The second predicted label output by the second prediction model Difference, that is This indicates that the current sample is located in the fuzzy region between the decision boundaries of the first and second prediction models, and may be a new type of fraud or a special normal behavior. The system will then activate the meta-learner. In order to make a final ruling.

[0077] The first predicted label output by the first prediction model The second predicted label output by the second prediction model Under different circumstances, the first probability output by the first prediction model The second probability output by the second prediction model Classification accuracy of the first prediction model The classification accuracy of the second prediction model and dynamic feature vectors corresponding to the target user. Determine the meta-feature vector , The feature dimension of the meta-feature vector is used to provide richer contextual information for arbitration.

[0078] For example, the confidence level of the first prediction model Confidence of the second prediction model The degree of discrepancy between the first and second prediction models In the feature space In the middle, find Find the k nearest neighbors (based on Euclidean distance), and calculate the classification accuracy of the first and second prediction models on the historical data of these k nearest neighbors. and , respectively as meta-features and .from Select the preset features that are highly correlated with global anomalies and determine them. ,For example etc. Based on , , , , and Determine the meta-feature vector .

[0079] Meta-learner It is a Given a simple logistic regression model as input, output a arbitrated target probability. , Among them, the target probability This represents the probability that the meta-learner considers the target user to be involved in fraud. The model parameters of the meta-learner are the target prediction labels for the final decision. in, The classification threshold is used. The meta-learner is trained on historically inconsistent samples and their final human annotations, learning in what situations which model should be trusted.

[0080] Then predict the label of the target. and target probability The confidence level is then transmitted to the downstream fraud early warning system, triggering corresponding handling procedures (such as monitoring, alarms, and interception). Simultaneously, all samples processed by the meta-learner and their final confirmed labels (obtainable through a business feedback loop) are collected for periodic retraining of the meta-learner. and incremental update base model and This allows the system to adapt to the ever-evolving methods of fraud.

[0081] By adopting the solution of this application, only the first predicted label output by the first prediction model is used. The second predicted label output by the second prediction model By initiating high-cost meta-learning arbitration under different circumstances, computing resources are allocated on demand, significantly improving the average processing efficiency of the system while ensuring overall accuracy.

[0082] Because iterative computation on a quantum computer is extremely expensive, this application employs a feedforward feature selection method to choose the optimal features. This quantum algorithm is integrated into the overall framework. The final number of selected features requires a trade-off between the maximum number of qubits and the model evaluation metrics obtained in the iterations. The flowchart for the quantum algorithm's feature combination selection is shown below. Figure 2 As shown, AUC represents the probability that the model will place a positive sample before a negative sample when a positive sample and a negative sample are randomly selected. AUC is an effective metric for evaluating the performance of QSVM. The feature corresponding to the end of the flowchart step is the best selected feature.

[0083] like Figure 3 As shown, the fraud-related user identification method of this application includes the following steps: acquiring the telecommunications network behavior sequence and user characteristics of the target user within a preset time range, and performing data preprocessing; determining the full feature vector based on the user characteristics and the telecommunications network behavior sequence. This full feature vector characterizes the user's characteristics and the inherent risk state evolution law of the target user's telecommunications network behavior; by using the full feature vector Input the first prediction model respectively Second prediction model The system obtains the first fraud-related result corresponding to the target user output by the first prediction model, and the second fraud-related result corresponding to the target user output by the second prediction model; if the first prediction label in the first fraud-related result... Second prediction label in the second fraud outcome If they are the same, then based on the first and second fraud-related results, the final prediction result (including the target prediction tag corresponding to the target user) is determined. and the target confidence level corresponding to the target prediction label. (), and transmit it to the downstream fraud early warning system, triggering the corresponding handling process, if the first prediction tag in the first fraud result Second prediction label in the second fraud outcome If they are different, then determine the meta-feature vector according to the method described above. By using meta-feature vectors Input meta-learner Determine the final prediction results (including target prediction tags corresponding to the target users). and target probability This information is then transmitted to the downstream fraud warning system, triggering the corresponding handling process.

[0084] By adopting the scheme of this application, a hybrid architecture is constructed in which quantum model (i.e., the second prediction model mentioned above) and non-quantum model (i.e., the first prediction model mentioned above) work together. Combined with the meta-learning arbitration mechanism, it can not only use quantum computing to accurately match known high-risk patterns, but also use the powerful generalization ability of classical models to capture unknown anomalies, and resolve model discrepancies through meta-learning intelligence. The system as a whole is more robust when facing new and mutated fraud.

[0085] The fraud-related user identification method provided in this application can be executed by a fraud-related user identification device. This application uses the example of a fraud-related user identification device executing the fraud-related user identification method to illustrate the fraud-related user identification device provided in this application.

[0086] Figure 4 This is a schematic diagram of the structure of a fraud-related user identification device disclosed in an embodiment of this application. Figure 4 As shown, the fraud-related user identification device 400 includes: an acquisition module 410, a determination module 420, and a obtaining module 430.

[0087] In this application, the acquisition module 410 is used to acquire the telecommunications network behavior sequence and user characteristics of the target user within a preset time range; the determination module 420 is used to determine a full feature vector based on the user characteristics and the telecommunications network behavior sequence, wherein the full feature vector represents the user characteristics and the inherent risk state evolution law of the target user's telecommunications network behavior; the obtaining module 430 is used to obtain a first fraud-related result corresponding to the target user by inputting the full feature vector into a first prediction model and a second prediction model, respectively, wherein the first prediction model is a non-quantum model that can achieve classification function based on multi-dimensional features, and the second prediction model is a quantum model that can achieve classification function based on multi-dimensional features; the determination module 420 is also used to determine a decision result corresponding to the target user based on the first fraud-related result and the second fraud-related result, wherein the decision result is used to indicate whether the target user is a fraud-related user.

[0088] In one implementation, the determining module 420 determines a full feature vector based on the user characteristics and the telecommunications network behavior sequence, including: performing quantum state mapping on each telecommunications network behavior in the telecommunications network behavior sequence to obtain a quantum state sequence corresponding to the telecommunications network behavior sequence, wherein the quantum state sequence carries correlation information between two adjacent telecommunications network behaviors in the telecommunications network behavior sequence, the correlation information including semantic correlation and temporal correlation; determining a dynamic feature vector corresponding to the target user based on the quantum state sequence, wherein the dynamic feature vector captures the inherent risk state evolution law of the target user's telecommunications network behavior; and determining a full feature vector based on the user characteristics and the dynamic feature vector.

[0089] In one implementation, the determining module 420 determines the dynamic feature vector corresponding to the target user based on the quantum state sequence, including: determining the dynamic feature vector corresponding to the target user based on the quantum state sequence, the model parameters of the normal model, and the model parameters of the fraud model, wherein the normal model is a model trained on a quantum hidden Markov model through a normal user behavior sequence, and the fraud model is a model trained on a quantum hidden Markov model through a fraudulent user behavior sequence.

[0090] In one implementation, the first fraud result includes a first probability and a first prediction label. The first probability is the probability that the first prediction model believes the target user is involved in fraud, and the first prediction label is used to indicate whether the first prediction model believes the target user is a fraudulent user. The second fraud result includes a second prediction label and a second probability. The second prediction label is used to indicate whether the second prediction model believes the target user is a fraudulent user, and the second probability is the probability that the second prediction model believes the target user is involved in fraud.

[0091] In one implementation, the determining module 420 determines a decision result corresponding to the target user based on the first fraud-related result and the second fraud-related result, including: when the first predicted label and the second predicted label are the same, determining a target predicted label corresponding to the target user based on the first predicted label and the second predicted label, wherein the target predicted label is used to indicate whether the target user is a fraud-related user; and determining a target confidence level corresponding to the target predicted label based on the first probability and the second probability.

[0092] In one implementation, the determining module 420 determines a decision result corresponding to the target user based on the first fraud-related result and the second fraud-related result, including: when the first predicted label and the second predicted label are different, determining a meta-feature vector based on the first probability, the second probability, the classification accuracy of the first prediction model, the classification accuracy of the second prediction model, and the dynamic feature vector corresponding to the target user, wherein the dynamic feature vector is determined based on the telecommunications network behavior sequence; obtaining a target probability output by the meta-learner by inputting the meta-feature vector, wherein the target probability is the probability that the meta-learner considers the target user to be involved in fraud; and determining a target predicted label corresponding to the target user based on the target probability, wherein the target predicted label is used to indicate whether the target user is a fraud-related user.

[0093] The fraud-related user identification device provided in this application embodiment can implement all the processes implemented in the fraud-related user identification method embodiment, and will not be described again here to avoid repetition.

[0094] Optionally, such as Figure 5 As shown, this application embodiment also provides an electronic device 500, including a processor 501 and a memory 502. The memory 502 stores a program or instructions that can run on the processor 501. When the program or instructions are executed by the processor 501, they implement the various steps of the above-described fraudulent user identification method embodiment and can achieve the same technical effect. To avoid repetition, they will not be described again here.

[0095] It should be noted that the electronic devices in the embodiments of this application include mobile electronic devices and non-mobile electronic devices.

[0096] This application also provides a readable storage medium storing a program or instructions. When the program or instructions are executed by a processor, they implement the various processes of the above-described fraud-related user identification method embodiments and achieve the same technical effect. To avoid repetition, they will not be described again here.

[0097] The processor is the processor in the electronic device described in the above embodiments. The readable storage medium includes computer-readable storage media, such as computer read-only memory (ROM), random access memory (RAM), magnetic disk, or optical disk.

[0098] This application also provides a computer program product, which includes a computer program stored on a non-transitory computer-readable storage medium. The computer program includes program instructions that, when executed by a computer, cause the computer to perform the steps of the fraudulent user identification method described above.

[0099] It should be understood that the training and prediction processes of the AI ​​models involved in the various embodiments of this specification all adhere to multiple legal and compliant principles, including legal data sources, compliant data content, compliant data governance, compliant training objectives and schemes, compliant training processes, compliant training environments and tools, and compliant ethical verification of training results, and comply with the requirements of Article 5 of the Patent Law. Among them: Data source legitimacy: All datasets used for AI model training were obtained through legal means, covering three categories: publicly authorized data, data authorized by partners, and self-collected compliant data. Publicly authorized data comes from compliant data sources following open-source licenses such as Apache 2.0, with complete copyright attribution and authorization scope clearly marked, and no unauthorized open-source code or data reuse. Data authorized by partners has been subject to formal data usage agreements, clearly defining the scope, duration, and confidentiality obligations, and possessing a complete authorization chain. For self-collected data involving personal information, strict informed consent procedures have been followed, and anonymization processes (including but not limited to field masking, feature anonymization, and differential privacy technology applications) have been implemented to remove personally identifiable information, fully complying with the requirements of relevant laws and regulations such as the "Interim Measures for the Administration of Generative Artificial Intelligence Services" and the "Personal Information Protection Law."

[0100] Data governance norms: A complete data traceability system is established during the AI ​​model training process to automatically record the source, collection time, annotation process, cleaning rules, and permission allocation of training data, generating traceable compliance reports to ensure that the data is verifiable throughout its entire lifecycle. The dataset annotation process for AI models is completed by a professional human R&D team, clearly defining the proportion of human creative contributions and avoiding reliance on AI-generated data that has not undergone substantial human modification, thus meeting the examination requirements for "human main contributions" in AI patent applications.

[0101] Training process compliance: A closed-loop training framework is adopted to ensure compliance and controllability of the training process. The specific process is as follows: First, training samples are obtained through compliant data sources. After the aforementioned data cleaning and desensitization, they are input into the neural network model to generate preliminary training results. Second, an expert system is introduced to verify the preliminary results. Based on preset rules and human expert experience, the feasibility of the results is evaluated, and outputs that may pose ethical risks or compliance hazards are corrected (such as removing decision-making logic that violates public order and good morals, and adjusting model parameters that do not comply with safety regulations). Finally, the loss function weights are dynamically optimized based on expert system feedback to strengthen the model's learning of compliant results, avoid overfitting errors or non-compliant labels, and form a closed-loop control of "data input - model training - expert verification - parameter optimization - result feedback" to ensure that the entire training process complies with A5 ethical review requirements.

[0102] Training environment and tool compliance: AI model training is implemented using nationally licensed chips and a compliant training platform. All open-source frameworks and components used in the training process have obtained their corresponding licenses, and copyright statements and patent citation information are fully retained, with no instances of infringement or reuse. The training environment is built using virtual devices (containers / virtual machines) with fixed random seeds and initial parameter configurations to ensure the reproducibility of the training process. Furthermore, through access control and operation log recording, risks such as data leakage and parameter tampering during training are prevented, ensuring the security and compliance of the training process.

[0103] Training results ethical verification compliance: After the model is trained, it undergoes additional third-party ethical compliance assessment and algorithm filing review to verify that the model output does not violate social morality or harm public interests. For potentially sensitive scenarios (such as public services and intelligent decision-making), a special result verification mechanism is established to ensure that the model always complies with Article 5 of the Patent Law and relevant laws and regulations in practical applications.

[0104] In summary, the data and training process used in the AI ​​model of this specification strictly comply with the relevant provisions of Article 5 of the Patent Law and the Patent Examination Guidelines (2023 Edition), and there are no violations of laws, social ethics, public interests, or illegal use of genetic resources. It fully meets the compliance requirements for patent authorization.

[0105] The above embodiments of this application focus on describing the differences between the various embodiments. As long as the different optimization features between the various embodiments are not contradictory, they can be combined to form a better embodiment. For the sake of brevity, they will not be described in detail here.

[0106] The above description is merely an embodiment of this application and is not intended to limit this application. Various modifications and variations can be made to this application by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principle of this application should be included within the scope of the claims of this application.

Claims

1. A method for identifying users suspected of fraud, characterized in that, include: Obtain the telecommunications network behavior sequence and user characteristics of the target user within a preset time range; Based on the user characteristics and the telecommunications network behavior sequence, a full feature vector is determined, wherein the full feature vector characterizes the user characteristics and the inherent risk state evolution law of the target user's telecommunications network behavior; By inputting the full feature vector into the first prediction model and the second prediction model respectively, the first fraud result corresponding to the target user output by the first prediction model and the second fraud result corresponding to the target user output by the second prediction model are obtained. The first prediction model is a non-quantum model that can achieve classification function based on multi-dimensional features, and the second prediction model is a quantum model that can achieve classification function based on multi-dimensional features. Based on the first fraud-related result and the second fraud-related result, a decision result corresponding to the target user is determined, wherein the decision result is used to indicate whether the target user is a fraud-related user.

2. The method according to claim 1, characterized in that, The step of determining the full feature vector based on the user characteristics and the telecommunications network behavior sequence includes: By performing quantum state mapping on each telecommunications network behavior in the telecommunications network behavior sequence, a quantum state sequence corresponding to the telecommunications network behavior sequence is obtained. The quantum state sequence carries the correlation information between two adjacent telecommunications network behaviors in the telecommunications network behavior sequence, and the correlation information includes semantic correlation and temporal correlation. Based on the quantum state sequence, a dynamic feature vector corresponding to the target user is determined, wherein the dynamic feature vector captures the inherent risk state evolution law of the target user's telecommunications network behavior; Based on the user characteristics and the dynamic feature vector, the full feature vector is determined.

3. The method according to claim 2, characterized in that, The step of determining the dynamic feature vector corresponding to the target user based on the quantum state sequence includes: Based on the quantum state sequence, the model parameters of the normal model, and the model parameters of the fraud model, a dynamic feature vector corresponding to the target user is determined. The normal model is a model trained on a quantum hidden Markov model using a normal user behavior sequence, and the fraud model is a model trained on a quantum hidden Markov model using a fraudulent user behavior sequence.

4. The method according to claim 1, characterized in that, The first fraud result includes a first probability and a first prediction label. The first probability is the probability that the first prediction model believes the target user is involved in fraud. The first prediction label is used to indicate whether the first prediction model believes the target user is involved in fraud. The second fraud result includes a second prediction label and a second probability. The second prediction label is used to indicate whether the second prediction model believes the target user is involved in fraud. The second probability is the probability that the second prediction model believes the target user is involved in fraud.

5. The method according to claim 4, characterized in that, The step of determining the decision result corresponding to the target user based on the first fraud-related result and the second fraud-related result includes: If the first prediction tag and the second prediction tag are the same, a target prediction tag corresponding to the target user is determined based on the first prediction tag and the second prediction tag, wherein the target prediction tag is used to indicate whether the target user is a fraudulent user; Based on the first probability and the second probability, the target confidence level corresponding to the target prediction label is determined.

6. The method according to claim 4, characterized in that, The step of determining the decision result corresponding to the target user based on the first fraud-related result and the second fraud-related result includes: When the first predicted label and the second predicted label are different, a meta-feature vector is determined based on the first probability, the second probability, the classification accuracy of the first prediction model, the classification accuracy of the second prediction model, and the dynamic feature vector corresponding to the target user, wherein the dynamic feature vector is determined based on the telecommunications network behavior sequence. By inputting the meta-feature vector into the meta-learner, the target probability output by the meta-learner is obtained, wherein the target probability is the probability that the meta-learner believes the target user is involved in fraud. Based on the target probability, a target prediction label is determined corresponding to the target user, wherein the target prediction label is used to indicate whether the target user is a fraudulent user.

7. A device for identifying users suspected of fraud, characterized in that, include: The acquisition module is used to acquire the telecommunications network behavior sequence and user characteristics of the target user within a preset time range; The determination module is used to determine a full feature vector based on the user characteristics and the telecommunications network behavior sequence, wherein the full feature vector represents the inherent risk state evolution law of the user characteristics and the telecommunications network behavior of the target user; The module is used to obtain a first fraud-related result corresponding to the target user by inputting the full feature vector into a first prediction model and a second prediction model, respectively. The first prediction model is a non-quantum model that can achieve classification based on multi-dimensional features, and the second prediction model is a quantum model that can achieve classification based on multi-dimensional features. The determining module is further configured to determine a decision result corresponding to the target user based on the first fraud-related result and the second fraud-related result, wherein the decision result is used to indicate whether the target user is a fraud-related user.

8. An electronic device, characterized in that, It includes a processor and a memory, the memory storing a program or instructions that can run on the processor, the program or instructions being executed by the processor to implement the steps of the fraud-related user identification method as described in any one of claims 1-6.

9. A readable storage medium, characterized in that, The readable storage medium stores a program or instructions that, when executed by a processor, implement the steps of the fraud-related user identification method as described in any one of claims 1-6.

10. A computer program product, characterized in that, The computer program product includes a computer program stored on a non-transitory computer-readable storage medium, the computer program including program instructions that, when executed by a computer, cause the computer to perform the steps of the fraud-related user identification method as described in any one of claims 1-6.