Financial fraud early warning method and device, computer equipment and storage medium
Through preliminary evaluation of user-side lightweight models and combined with cloud-based full-scale model review, the problem of insufficient sentiment analysis in financial risk control is solved, efficient identification and defense of telecommunications fraud is achieved, and the accuracy and efficiency of early warnings are improved.
Patent Information
- Application Number
- CN202510654024.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-20
- Publication Date
- 2025-08-22
AI Technical Summary
The existing financial risk control technology has failed to effectively integrate the emotional dimension analysis in voice interaction scenarios, resulting in the inability to identify the abnormal emotional characteristics of users during telecommunications fraud, and it is difficult to deal with the dynamic fraudulent tactics generated by AI, resulting in a lag in defense.
By deploying a lightweight model on the user side for multi-scale acoustic feature extraction and semantic analysis, financial vocabulary features and vocabulary risk scores are obtained, and the features are uploaded to the cloud full model for review when the preset threshold is reached, and timing risk scores are output to trigger fraud warnings.
It improves the accuracy and reliability of financial fraud warnings, reduces cloud resource usage, improves warning efficiency, and can more comprehensively capture abnormal information in user calls.
Smart Images

Figure CN120526802A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of financial risk control technology, and in particular to a financial fraud early warning method, device, computer equipment and storage medium. Background Art
[0002] In the current field of financial risk control technology, risk control models based on user transaction behavior data (such as voice chats) have formed a basic application framework, but they suffer from a significant flaw: insufficient multi-dimensional data integration. While existing technologies can identify potential risks through behavioral characteristics such as transaction frequency and fluctuations in transaction amounts, they fail to integrate emotional analysis in voice interaction scenarios. This results in an inability to effectively identify abnormal emotional characteristics such as nervousness and coercion displayed by perpetrators of telecommunications fraud when they trick users into transferring money.
[0003] In terms of emotion recognition technology, existing related technologies in the financial field mostly use general-purpose sentiment analysis models. Their emotion classification only remains at the coarse-grained division of "positive / negative", and they fail to build a fine-grained emotion classification system that conforms to the characteristics of financial fraud. As a result, it is impossible to accurately identify key risk signals such as the "pseudo-enthusiasm" language in fake customer service scenarios or the panic emotions of victims when they encounter fraud.
[0004] In addition, existing voice content detection technology relies too much on a predefined keyword rule library. When faced with dynamic induced speech generated by AI, it lacks adaptive semantic understanding capabilities and is difficult to effectively identify new fraud speech that has been semantically reconstructed. This leads to the risk of defense lag in the financial anti-fraud system when dealing with rapidly iterating intelligent criminal methods. Summary of the Invention
[0005] The present invention provides a financial fraud early warning method, apparatus, computer equipment and storage medium to solve the problem of how to improve the accuracy of financial fraud early warning.
[0006] In a first aspect, a financial fraud early warning method is provided, comprising:
[0007] The lightweight model deployed on the user side is used to obtain the current voice segment of the user during a call in real time;
[0008] Perform multi-scale acoustic feature extraction on the current speech segment to obtain multi-dimensional acoustic features;
[0009] Perform semantic analysis on the current speech segment to obtain financial vocabulary features and corresponding vocabulary risk scores;
[0010] When the lexical risk score reaches a first preset threshold, the multi-dimensional acoustic features and financial lexical features of multiple consecutive speech segments before and after the current speech segment are input and uploaded to the full model deployed in the cloud for review, and multiple temporal risk scores corresponding to the multiple speech segments are output;
[0011] When multiple temporal risk scores reach a second preset threshold, a fraud warning is triggered and corresponding actions are performed.
[0012] In a second aspect, a financial fraud early warning device is provided, comprising:
[0013] An acquisition unit is used to acquire the current voice segment of the user during a call in real time through a lightweight model deployed on the user side;
[0014] The acoustic feature extraction unit is used to extract multi-scale acoustic features of the current speech segment to obtain multi-dimensional acoustic features;
[0015] Semantic analysis unit, used to perform semantic analysis on the current speech segment to obtain financial vocabulary features and corresponding vocabulary risk scores;
[0016] A risk scoring unit is configured to, when the vocabulary risk score reaches a first preset threshold, input the multidimensional acoustic features and financial vocabulary features of multiple speech segments preceding and following the current speech segment into a full-scale model deployed in the cloud for review, and output multiple temporal risk scores corresponding to the multiple speech segments;
[0017] The fraud warning unit is used to trigger a fraud warning and perform corresponding operations when multiple time series risk scores reach a second preset threshold.
[0018] In a third aspect, a computer device is provided, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the steps of the above-mentioned financial fraud early warning method when executing the computer program.
[0019] In a fourth aspect, a computer-readable storage medium is provided, which stores a computer program. When the computer program is executed by a processor, the steps of the above-mentioned financial fraud early warning method are implemented.
[0020] In the scheme implemented by the above-mentioned financial fraud early warning method, device, computer equipment and storage medium, the current voice segment of the user during the call can be obtained in real time through the lightweight model deployed on the user side; multi-scale acoustic feature extraction is performed on the current voice segment to obtain multi-dimensional acoustic features; semantic analysis is performed on the current voice segment to obtain financial vocabulary features and corresponding vocabulary risk scores; when the vocabulary risk score reaches a first preset threshold, the multi-dimensional acoustic features and financial vocabulary features of multiple consecutive voice segments before and after the current voice segment are input and uploaded to the full model deployed on the cloud for review, and multiple time series risk scores corresponding to the multiple voice segments are output; when the multiple time series risk scores reach a second preset threshold, a fraud warning is triggered and corresponding operations are performed. In the present invention, after the risk is assessed lightweightly by the user side, the full cloud model is called for review, which not only ensures the accuracy of the warning, but also reduces the occupation of cloud resources and improves the efficiency of the warning. At the same time, the combination of multi-scale acoustic feature extraction and semantic analysis can more comprehensively capture abnormal information in the user's call, thereby improving the accuracy and reliability of the financial fraud warning. BRIEF DESCRIPTION OF THE DRAWINGS
[0021] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments of the present invention. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.
[0022] Figure 1 2 is a schematic diagram of an application environment of a financial fraud early warning method in an embodiment of the present invention.
[0023] Figure 2 It is a flow chart of a financial fraud early warning method in one embodiment of the present invention.
[0024] Figure 3 This is a flowchart of a specific implementation of step S202 in one embodiment of the present invention.
[0025] Figure 4 This is a flowchart of a specific implementation of step S203 in one embodiment of the present invention.
[0026] Figure 5 It is a structural diagram of a financial fraud early warning device in one embodiment of the present invention.
[0027] Figure 6 It is a structural diagram of a computer device in one embodiment of the present invention.
[0028] Figure 7 FIG. 2 is another structural diagram of a computer device according to an embodiment of the present invention. DETAILED DESCRIPTION
[0029] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of them. All other embodiments obtained by ordinary technicians in this field based on the embodiments of the present invention without making any creative efforts shall fall within the scope of protection of the present invention.
[0030] The financial fraud early warning method provided by the embodiment of the present invention can be applied in the following situations: Figure 1 In an application environment, a user terminal communicates with the cloud via a network. The user terminal provides a data node to receive the user's current voice segment during a call in real time; multi-scale acoustic feature extraction is performed on the current voice segment to obtain multidimensional acoustic features; semantic analysis is performed on the current voice segment to obtain financial vocabulary features and corresponding vocabulary risk scores; when the vocabulary risk score reaches a first preset threshold, the multi-dimensional acoustic features and financial vocabulary features of multiple consecutive voice segments before and after the current voice segment are input and uploaded to a full-scale model deployed in the cloud for review, outputting multiple temporal risk scores corresponding to the multiple voice segments; when the multiple temporal risk scores reach a second preset threshold, a fraud warning is triggered and corresponding actions are executed. In this invention, a lightweight risk assessment is performed on the user terminal, followed by a full-scale cloud-based review, ensuring the accuracy of the warning while reducing cloud resource usage and improving warning efficiency. Furthermore, the combination of multi-scale acoustic feature extraction and semantic analysis can more comprehensively capture abnormal information in user calls, improving the accuracy and reliability of financial fraud warnings. The user terminal can include, but is not limited to, various personal computers, laptops, smartphones, tablets, and portable wearable devices. The cloud can be implemented with an independent server or a server cluster consisting of multiple servers. The present invention is described in detail below through specific embodiments.
[0031] See also Figure 2 As shown, Figure 2 A flowchart of a financial fraud early warning method provided by an embodiment of the present invention includes the following steps S201-S205.
[0032] S201, obtaining the current voice segment of the user during the call in real time through the lightweight model deployed on the user end;
[0033] In this step, when the user makes a call through the user end, the current voice segment of the user during the call is obtained in real time through the lightweight model deployed on the user end. The lightweight model here is used to preliminarily assess the possible financial fraud risks during the user's call, and has the characteristics of faster assessment and higher timeliness.
[0034] S202, extracting multi-scale acoustic features from the current speech segment to obtain multi-dimensional acoustic features;
[0035] In this step, multidimensional acoustic features are extracted from different scales and angles. These features reflect the characteristics of the user's voice during a call, providing a crucial data foundation for subsequent semantic analysis and financial fraud warnings. Specifically, by extracting multi-scale acoustic features from the current voice segment, we can more comprehensively and accurately capture abnormalities in user calls, improving the accuracy and reliability of financial fraud warnings.
[0036] For example, in a financial fraud call between a scammer and an investment victim, extracting the fundamental frequency characteristics of the scammer's current speech segment reveals that coercion or tension can significantly increase fundamental frequency fluctuations (measured data shows that the jitter index of fraudulent speech is 37% higher than normal speech on average), reflecting the abnormal emotional state of the fraudulent behavior. Furthermore, multidimensional acoustic features such as sound energy, formant frequency, and voice quality can be extracted. These features also exhibit different patterns in fraudulent scenarios than in normal calls. For example, the scammer's speech energy may decrease and the formant frequency may shift due to tension or deliberate disguise. By comprehensively analyzing these multidimensional acoustic features, we can uncover fraud clues hidden in the speech.
[0037] S203: Perform semantic analysis on the current speech segment to obtain financial vocabulary features and corresponding vocabulary risk scores;
[0038] In this step, the current voice segment is converted into text information, and all financial-related financial vocabulary features in the text information are extracted. Each financial vocabulary feature has a corresponding vocabulary risk score. The vocabulary risk score can be calculated based on factors such as the frequency of financial vocabulary in the context of financial fraud, its use in historical fraud cases, and the sensitivity of the vocabulary itself. By analyzing these financial vocabulary features and their risk scores, potential financial fraud behaviors in the voice can be further identified. For example, certain sensitive words such as "transfer", "password", "verification code", etc. often appear frequently in fraudulent calls and are accompanied by higher vocabulary risk scores, prompting the system to focus on these calls and conduct further analysis.
[0039] S204. When the vocabulary risk score reaches a first preset threshold, the multi-dimensional acoustic features and financial vocabulary features of multiple consecutive speech segments before and after the current speech segment are input and uploaded to the full model deployed in the cloud for review, and multiple temporal risk scores corresponding to the multiple speech segments are output.
[0040] In step S204, the full model uses a bidirectional LSTM+Attention model, which can integrate multi-dimensional acoustic features and financial vocabulary features to output a temporal risk score. Specifically, the temporal risk score of each speech segment can be calculated according to the following formula:
[0041] h t =BiLSTM(f t ⊕e t ),s t =Softmax(W a h t );
[0042] Among them, f t represents the multidimensional acoustic features extracted from the speech segment, e t represents the financial vocabulary features extracted from the speech segment, ⊕ represents the vector concatenation operation, and W a Represents a trainable attention weight matrix.
[0043] The Attention mechanism in the full-scale model enables the model to focus on key time segments (such as paragraphs where high-frequency keywords appear suddenly in the sales pitch). Semantic analysis in the full-scale model includes both explicit and implicit semantic analysis. Explicit semantic analysis can directly match fraudulent sales pitches using 368 keywords, while implicit semantic analysis analyzes contextual relevance based on word embeddings (e.g., the co-occurrence pattern between "system upgrade" and "account anomaly").
[0044] In step S204, when the vocabulary risk score reaches the first preset threshold, it indicates that the user-side lightweight model has preliminarily assessed the possibility of financial fraud. To more accurately judge and reduce the false alarm rate, the system uploads the multidimensional acoustic features and financial vocabulary features of the current speech segment and multiple adjacent speech segments to the full model deployed in the cloud for review. The full cloud model has stronger processing power and richer data resources, capable of in-depth fusion analysis and comprehensive evaluation of these features, ultimately outputting multiple time series risk scores corresponding to multiple speech segments. These time series risk scores can reflect the changing trends of fraudulent behavior over time, providing a more comprehensive and accurate basis for subsequent early warning decisions.
[0045] S205: When multiple time series risk scores reach a second preset threshold, a fraud warning is triggered and corresponding operations are performed.
[0046] In step S205, among the temporal risk scores of multiple voice segments, when the weighted scores of the temporal risk scores of a preset number of consecutive voice segments reach a second preset threshold, a fraud warning is triggered and a corresponding real-time interception operation is performed. For example, assuming that 10 voice segments are uploaded, 10 corresponding temporal risk scores can be obtained. When the weighted scores of the temporal risk scores of a preset number (for example, 3) of consecutive voice segments reach the second preset threshold, a fraud warning is triggered and a corresponding real-time interception operation is performed.
[0047] In step S205, the review of the full model confirms that multiple time series risk scores reach the second preset threshold, which means that the full model in the cloud also determines that there is a high risk of financial fraud in the current voice segment and its adjacent voice segments. At this time, the system will immediately trigger the fraud warning mechanism and automatically execute the preset countermeasures. For countermeasures, for example, the system may immediately interrupt the transaction process conducted by the user during the current call, freeze the user's related accounts, or send a security verification request to the user to ensure the safety of the user's funds. At the same time, the system will also record the fraud warning event and report the relevant information to the financial regulatory authorities for further tracking and handling of potential fraudulent behaviors. In addition, the system will also conduct an in-depth analysis of the voice segment that triggers the financial fraud warning and its adjacent voice segments according to the preset rules, extract the characteristics of the fraudulent behavior, and use it to optimize and update the model to improve the accuracy and robustness of the model, thereby better protecting the user's property safety.
[0048] As can be seen, in the solution of steps S201-S205 above, risk is assessed lightly on the user side, and full cloud verification is only performed when the user side detects suspicion. This ensures the accuracy of the warning, reduces cloud resource usage, and improves warning efficiency. Furthermore, the combination of multi-scale acoustic feature extraction and semantic analysis can more comprehensively capture abnormal information in user calls, improving the accuracy and reliability of financial fraud warnings.
[0049] In one embodiment, if Figure 3 As shown, step S202 includes:
[0050] S301, extracting the fundamental frequency feature, short-time energy feature and MFCC feature of the current speech segment to obtain basic features;
[0051] S302, extracting the tremor index and harmonic noise ratio of the current speech segment to obtain nonlinear features;
[0052] S303: Extract the speech rate change rate and the percentage of silence intervals of the current speech segment to obtain context features;
[0053] S304: Collect basic features, nonlinear features, and context features to obtain multi-dimensional acoustic features.
[0054] In step S302, the tremor index is extracted according to the following formula:
[0055]
[0056] Among them, F0 i represents the fundamental frequency feature of the i-th frame, N represents the total number of frames of the speech segment, F0 i -F0 i+1 It represents the absolute difference between the fundamental frequencies of adjacent frames. It can reflect the trembling characteristics of the speaker's voice by quantifying the short-term fluctuation of the fundamental frequency (F0) in the current speech segment signal.
[0057] In step S302, the harmonic-to-noise ratio is extracted according to the following formula:
[0058]
[0059] Here, H(k) represents the harmonic spectrum amplitude, N(k) represents the noise spectrum amplitude, and K represents the number of frequencies within the spectrum analysis bandwidth. Regarding the harmonic-to-noise ratio (HNR), for example, the HNR of fraudulent speech (such as impersonating customer service) is typically lower than that of normal speech because the harmonic energy is reduced by deliberately controlling vocal cord vibration.
[0060] In step S303, the percentage of silent intervals refers to the percentage of silent time (speech energy below -50dB and lasting >200ms) during the call. Experimental data shows that the percentage of silent intervals can effectively help identify financial fraud. For example, the percentage of silent intervals in calls impersonating public security, procuratorial, or legal authorities (waiting for the victim's response) is 12-15%, while the percentage of silent intervals in normal customer service calls is only 3-5%.
[0061] In this embodiment, steps S301-S304 extract 128 features across three categories. These features more closely reflect the user's emotional fluctuations and unusual behavior during a call, providing richer data support for subsequent fraud risk assessment. These multidimensional acoustic features encompass not only the basic acoustic properties of speech but also the user's emotional reactions and speech habits during the call, enabling a more comprehensive assessment of fraud risk during the call. Furthermore, by integrating multiple features, their complementarity can be enhanced, further improving the accuracy of financial fraud warnings.
[0062] For example, in the financial field, when a user is talking to customer service, if the user attempts to impersonate someone else to commit fraud, their voice features will often show a pattern different from that of a normal call. Through steps S301 to S304 described in the embodiment, the system can accurately capture these subtle changes in voice features. For example, if the user deliberately controls the vibration of the vocal cords during a call, resulting in a decrease in harmonic energy, the system's harmonic-to-noise ratio feature can accurately identify this anomaly. At the same time, if the user frequently remains silent during a call, and the proportion of silence intervals exceeds the normal range, the system can also capture this abnormal behavior through contextual features. These multi-dimensional acoustic features can not only help the system more accurately determine whether the user has committed fraud, but also provide more detailed data support for subsequent fraud risk assessments. Therefore, the financial fraud early warning method in this embodiment has broad application prospects and important practical value in the financial field.
[0063] In one embodiment, if Figure 4 As shown, step S203 includes:
[0064] S401. Pre-build a domain word risk analysis model containing multiple financial fraud-related words;
[0065] S402: The text information of the current speech segment is input into the domain dictionary analysis model for semantic matching analysis, and the financial vocabulary features of the current speech segment and the corresponding vocabulary risk score are output.
[0066] In this example, a domain dictionary containing 368 financial fraud-related words (such as "security account" and "verification code") is constructed. Through word vector weighted enhanced semantic analysis, the text information of the current speech segment is input into the following domain dictionary analysis model for calculation to obtain the vocabulary risk score in the current speech segment:
[0067] α w =TF-IDF(w)·RiskScore(w)
[0068] Where D represents the domain dictionary of 368 financial fraud keywords (such as "security account"), v w The word vector "w" represents a pre-trained word (typically 300 dimensions). TF-IDF(w) represents the importance weight of a word in the corpus. RiskScore(w) represents the word's risk score (0-1) based on historical fraud corpus statistics. RiskScore(w) can be derived from historical fraud corpus statistics. Semantic enhancement is achieved through weighted aggregation. For example, the RiskScore for "verification code" is 0.92, significantly higher than that of common words (such as "transfer" at 0.45).
[0069] As can be seen, in the above solution, for financial fraud early warning risk control, the lightweight model deployed on the user side first assesses the risk, and then calls the full model deployed in the cloud for verification. This not only ensures the accuracy of the early warning, but also reduces cloud resource usage and improves early warning efficiency. Furthermore, the combination of multi-scale acoustic feature extraction and semantic analysis can more comprehensively capture abnormal information in user calls, improving the accuracy and reliability of financial fraud early warnings.
[0070] Furthermore, the lightweight model deployed on the user side is a trimmed version of the full model deployed on the cloud. The lightweight model removes the Attention module from the full model and reduces the number of LSTM layers from three to one. A performance comparison shows that the lightweight model has less than 5MB of parameters and less than 90ms of latency, while the full model has 120MB of parameters and a latency of 800ms.
[0071] Fraud detection accuracy tests in several financial scenarios are as follows: In real-world banking data, the system achieved a 94.5% recall rate for telecom fraud, with a false positive rate (FPR) of 1.2% (compared to 78% and 8% for traditional models). The lightweight model on the client side can trigger an alert within 5 seconds of the start of a fraudulent attack, a 15-fold increase compared to cloud-based solutions, and an 82% success rate. Preliminary evaluation of the lightweight model on the client side has reduced invalid data uploads by 90%, saving over 100,000 yuan in daily bandwidth costs (based on an estimated user base of millions).
[0072] It should be understood that the size of the serial numbers of the steps in the above embodiments does not mean the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present invention.
[0073] The embodiment of the present invention further provides a financial fraud early warning device, which is used to execute any embodiment of the above-mentioned financial fraud early warning method. Figure 5 , Figure 5 It is a schematic block diagram of a financial fraud early warning device provided by an embodiment of the present invention.
[0074] like Figure 5 As shown, the financial fraud early warning device 500 includes: an acquisition unit 501, an acoustic feature extraction unit 502, a semantic analysis unit 503, a risk scoring unit 504 and a fraud early warning unit 505.
[0075] An acquisition unit 501 is configured to acquire, in real time, a current voice segment of a user during a call using a lightweight model deployed on a user terminal;
[0076] The acoustic feature extraction unit 502 is used to extract multi-scale acoustic features from the current speech segment to obtain multi-dimensional acoustic features;
[0077] Semantic analysis unit 503, used to perform semantic analysis on the current speech segment to obtain financial vocabulary features and corresponding vocabulary risk scores;
[0078] Risk scoring unit 504 is configured to, when the vocabulary risk score reaches a first preset threshold, input the multi-dimensional acoustic features and financial vocabulary features of multiple speech segments preceding and following the current speech segment into the full-scale model deployed in the cloud for review, and output multiple temporal risk scores corresponding to the multiple speech segments;
[0079] The fraud warning unit 505 is configured to trigger a fraud warning and perform corresponding operations when multiple temporal risk scores reach a second preset threshold.
[0080] In one embodiment, the acoustic feature extraction unit 502 specifically includes:
[0081] The first extraction unit is used to extract the fundamental frequency feature, short-time energy feature and MFCC feature of the current speech segment to obtain basic features;
[0082] The second extraction unit is used to extract the tremor index of the current speech segment and the harmonic noise ratio to obtain nonlinear features;
[0083] The third extraction unit is used to extract the speech rate change rate and the proportion of silence intervals of the current speech segment to obtain context features;
[0084] The aggregation unit is used to aggregate basic features, nonlinear features and context features to obtain multi-dimensional acoustic features.
[0085] In one embodiment, the second extraction unit is specifically configured to:
[0086] The tremor index is extracted according to the following formula:
[0087]
[0088] Among them, F0 i represents the fundamental frequency feature of the i-th frame, N represents the total number of frames of the speech segment, F0 i -F0 i+1 Indicates the absolute difference between the base frequencies of adjacent frames;
[0089] The harmonic-to-noise ratio is extracted using the following formula:
[0090]
[0091] Where H(k) represents the harmonic spectrum amplitude, N(k) represents the noise spectrum amplitude, and K represents the number of frequency points within the spectrum analysis bandwidth.
[0092] In one embodiment, the semantic analysis unit 503 is specifically configured to:
[0093] Pre-build a domain word risk analysis model containing multiple financial fraud-related words;
[0094] The text information of the current speech segment is input into the domain dictionary analysis model for semantic matching analysis, and the financial vocabulary features of the current speech segment and the corresponding vocabulary risk score are output.
[0095] In one embodiment, the risk scoring unit 504 is specifically configured to:
[0096] The temporal risk score of each speech segment is calculated according to the following formula:
[0097] h t =BiLSTM(f t ⊕e t ),s t =Softmax(W a h t );
[0098] Among them, f t represents the multidimensional acoustic features extracted from the speech segment, e t represents the financial vocabulary features extracted from the speech segment, ⊕ represents the vector concatenation operation, and W a Represents a trainable attention weight matrix.
[0099] In one embodiment, the fraud warning unit 505 is specifically configured to:
[0100] In the temporal risk scores of multiple voice segments, when the weighted scores of the temporal risk scores of a preset number of consecutive voice segments reach a second preset threshold, a fraud warning is triggered and a corresponding real-time interception operation is performed.
[0101] This invention provides a financial fraud early warning device. This device, specifically designed for risk control, performs a lightweight risk assessment on the user side before fully verifying the risk in the cloud. This ensures the accuracy of early warnings while reducing cloud resource usage and improving early warning efficiency. Furthermore, by combining multi-scale acoustic feature extraction with semantic analysis, it can more comprehensively capture abnormal information in user calls, enhancing the accuracy and reliability of financial fraud early warnings.
[0102] The specific definition of the financial fraud early warning device can be found in the definition of the financial fraud early warning method above and will not be repeated here. Each module in the aforementioned financial fraud early warning device can be implemented in whole or in part through software, hardware, or a combination thereof. Each of the aforementioned modules can be embedded in or independent of a processor in a computer device in hardware form, or can be stored in a computer device memory in software form, so that the processor can call and execute the corresponding operations of each module.
[0103] In one embodiment, a computer device is provided. The computer device may be a server, and its internal structure diagram may be as follows: Figure 6 As shown. The computer device includes a processor, memory, network interface and database connected via a system bus. The processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes non-volatile and / or volatile storage media and internal memory. The non-volatile storage medium stores an operating system, a computer program and a database. The internal memory provides an environment for the operation of the operating system and computer program in the non-volatile storage medium. The network interface of the computer device is used to communicate with an external client via a network connection. When the computer program is executed by the processor, it implements the functions or steps on the service side of a financial fraud early warning method.
[0104] In one embodiment, a computer device is provided. The computer device may be a client, and its internal structure diagram may be as follows: Figure 7 As shown. The computer device includes a processor, memory, a network interface, a display screen, and an input device connected via a system bus. The processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and computer program in the non-volatile storage medium. The network interface of the computer device is used to communicate with an external server via a network connection. When executed by the processor, the computer program implements the functions or steps on the client side of a financial fraud early warning method.
[0105] In one embodiment, a computer device is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the following steps are performed:
[0106] The lightweight model deployed on the user side is used to obtain the current voice segment of the user during a call in real time;
[0107] Perform multi-scale acoustic feature extraction on the current speech segment to obtain multi-dimensional acoustic features;
[0108] Perform semantic analysis on the current speech segment to obtain financial vocabulary features and corresponding vocabulary risk scores;
[0109] When the vocabulary risk score reaches a first preset threshold, the multi-dimensional acoustic features and financial vocabulary features of multiple consecutive speech segments before and after the current speech segment are uploaded to the full-scale model deployed in the cloud for review, and multiple temporal risk scores corresponding to the multiple speech segments are output;
[0110] When multiple temporal risk scores reach a second preset threshold, a fraud warning is triggered and corresponding actions are performed.
[0111] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the following steps are implemented:
[0112] The lightweight model deployed on the user side is used to obtain the current voice segment of the user during a call in real time;
[0113] Perform multi-scale acoustic feature extraction on the current speech segment to obtain multi-dimensional acoustic features;
[0114] Perform semantic analysis on the current speech segment to obtain financial vocabulary features and corresponding vocabulary risk scores;
[0115] When the vocabulary risk score reaches a first preset threshold, the multi-dimensional acoustic features and financial vocabulary features of multiple consecutive speech segments before and after the current speech segment are uploaded to the full-scale model deployed in the cloud for review, and multiple temporal risk scores corresponding to the multiple speech segments are output;
[0116] When multiple temporal risk scores reach a second preset threshold, a fraud warning is triggered and corresponding actions are performed.
[0117] It should be noted that the above functions or steps that can be implemented by the computer-readable storage medium or computer device can be found in the relevant descriptions of the server side and the client side in the aforementioned method embodiment. To avoid repetition, they will not be described one by one here.
[0118] Those skilled in the art will appreciate that all or part of the processes in the above-mentioned embodiments can be implemented by instructing the relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to memory, storage, database or other media used in the embodiments provided in this application can include non-volatile and / or volatile memory. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM) or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link (Synchlink) DRAM (SLDRAM), memory bus (Rambus) direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM) and memory bus dynamic RAM (RDRAM).
[0119] Those skilled in the art will clearly understand that for the sake of convenience and brevity of description, only the division of the above-mentioned functional units and modules is used as an example. In actual applications, the above-mentioned functions can be distributed and completed by different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above.
[0120] The above embodiments are only used to illustrate the technical solutions of the present invention, not to limit them. Although the present invention has been described in detail with reference to the above embodiments, it should be understood by those skilled in the art that the technical solutions described in the above embodiments can still be modified, or some of the technical features thereof can be replaced by equivalents. However, these modifications or replacements do not deviate from the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present invention, and should be included in the scope of protection of the present invention. The non-Company software tools or components that appear in the embodiments of this application are merely examples and do not represent actual use.
Claims
1. A financial fraud early warning method, characterized in that: include: The lightweight model deployed on the user side is used to obtain the current voice segment of the user during a call in real time; Perform multi-scale acoustic feature extraction on the current speech segment to obtain multi-dimensional acoustic features; Perform semantic analysis on the current speech segment to obtain financial vocabulary features and corresponding vocabulary risk scores; When the lexical risk score reaches a first preset threshold, the multi-dimensional acoustic features and financial lexical features of multiple consecutive speech segments before and after the current speech segment are input and uploaded to the full model deployed in the cloud for review, and multiple temporal risk scores corresponding to the multiple speech segments are output; When multiple temporal risk scores reach a second preset threshold, a fraud warning is triggered and corresponding actions are performed.
2. The financial fraud early warning method according to claim 1, characterized in that: The multi-scale acoustic feature extraction of the current speech segment to obtain multi-dimensional acoustic features includes: Extract the fundamental frequency feature, short-time energy feature and MFCC feature of the current speech segment to obtain the basic features; Extract the tremor index and harmonic noise ratio of the current speech segment to obtain nonlinear features; Extract the speech rate change rate and silence interval ratio of the current speech segment to obtain context features; The basic features, nonlinear features and context features are combined to obtain multi-dimensional acoustic features.
3. The financial fraud early warning method according to claim 2, characterized in that: The step of extracting the tremor index of the current speech segment includes: The tremor index is extracted according to the following formula: Among them, F0 i represents the fundamental frequency feature of the i-th frame, N represents the total number of frames of the speech segment, F0 i -F0 i+1 Indicates the absolute difference in fundamental frequencies between adjacent frames.
4. The financial fraud early warning method according to claim 2, characterized in that: The extracting the harmonic-to-noise ratio of the current speech segment includes: The harmonic-to-noise ratio is extracted as follows: Where H(k) represents the harmonic spectrum amplitude, N(k) represents the noise spectrum amplitude, and K represents the number of frequency points within the spectrum analysis bandwidth.
5. The financial fraud early warning method according to claim 1, characterized in that: The semantic analysis of the current speech segment to obtain financial vocabulary features and corresponding vocabulary risk scores includes: Pre-build a domain word risk analysis model containing multiple financial fraud-related words; The text information of the current speech segment is input into the domain dictionary analysis model for semantic matching analysis, and the financial vocabulary features of the current speech segment and the corresponding vocabulary risk score are output.
6. The financial fraud early warning method according to claim 1, characterized in that: The multi-dimensional acoustic features and financial vocabulary features of multiple consecutive speech segments before and after the current speech segment are input and uploaded to the full model deployed in the cloud for review, and multiple temporal risk scores corresponding to the multiple speech segments are output, including: The temporal risk score of each speech segment is calculated according to the following formula: Among them, f t represents the multidimensional acoustic features extracted from the speech segment, e t represents the financial vocabulary features extracted from the speech segment, Represents vector concatenation operation, W a Represents a trainable attention weight matrix.
7. The financial fraud early warning method according to claim 1, characterized in that: When the multiple time series risk scores reach a second preset threshold, a fraud warning is triggered and corresponding operations are performed, including: In the temporal risk scores of multiple voice segments, when the weighted scores of the temporal risk scores of a preset number of consecutive voice segments reach a second preset threshold, a fraud warning is triggered and a corresponding real-time interception operation is performed.
8. A financial fraud early warning device, characterized in that: include: An acquisition unit is used to acquire the current voice segment of the user during a call in real time through a lightweight model deployed on the user side; The acoustic feature extraction unit is used to extract multi-scale acoustic features of the current speech segment to obtain multi-dimensional acoustic features; Semantic analysis unit, used to perform semantic analysis on the current speech segment to obtain financial vocabulary features and corresponding vocabulary risk scores; A risk scoring unit is configured to, when the vocabulary risk score reaches a first preset threshold, input the multidimensional acoustic features and financial vocabulary features of multiple speech segments preceding and following the current speech segment into a full-scale model deployed in the cloud for review, and output multiple temporal risk scores corresponding to the multiple speech segments; The fraud warning unit is used to trigger a fraud warning and perform corresponding operations when multiple time series risk scores reach a second preset threshold.
9. A computer device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the computer program, the financial fraud early warning method according to any one of claims 1 to 7 is implemented.
10. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a computer program, which, when executed by a processor, causes the processor to execute the financial fraud early warning method according to any one of claims 1 to 7.
Citation Information
Patent Citations
Telcom phone phishing-resistant method and system based on discrimination and identification content analysis
CN103179122A
Temporal semantic fusion association determining sub-system based on multimodal emotion recognition system
CN108805087A
Telecommunication fraud detection method and device
CN110349586A
Anti-communication network fraud studying, judging, early warning and intercepting integrated platform
CN115102789A
System for identifying financial risk website based on fingerprint penetration technology
CN115879110A
Cited By
Detecting statistical anomalies in voice interactions
US12738274B1