Voice fraud analysis method, device and equipment and storage medium

By using an independent microphone array and voice activity detection technology, combined with a pre-trained fraud analysis model, the accuracy and real-time performance issues of existing voice fraud analysis technologies in complex scenarios are solved, achieving efficient voice fraud detection.

CN119181380BActive Publication Date: 2025-11-28PING AN TECH (SHENZHEN) CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411243438.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-09-05
Publication Date
2025-11-28
Estimated Expiration
2044-09-05

AI Technical Summary

Technical Problem

Existing voice fraud analysis technologies lack accuracy and real-time performance when processing complex dialogues and multi-channel data. They also lack intelligent analysis capabilities, are difficult to adapt to diverse fraud patterns, cannot meet the needs of specific industries, and rely on manual judgment, resulting in low efficiency and accuracy.

Method used

By acquiring speech signals, two independently configured microphone arrays are used to collect speech data from different sources and record them into two independent and isolated audio channels. Speech activity detection technology is used to identify and mark speech activity time periods, and effective speech segments are obtained through segmentation to generate dialogue text. Finally, the fraud probability is determined based on a pre-trained fraud analysis model.

Benefits of technology

It achieves precise processing and analysis of voice signals, significantly improving the quality of voice data and the accuracy of fraud detection, and is suitable for complex voice interaction scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119181380B_ABST
    Figure CN119181380B_ABST
Patent Text Reader

Abstract

The application relates to a voice fraud analysis method, which comprises the following steps: acquiring a voice signal, collecting voice data of different sources through two independently configured microphone arrays, and recording the voice data into two independent and isolated audio channels; detecting the voice data in the two audio channels through a voice activity detection technology, identifying and marking a voice activity time period in each audio channel, performing a fragmentation operation on the voice data based on the voice activity time period, and extracting a plurality of effective voice segments; generating a dialogue text according to the extracted effective voice segments, and screening a text part to be analyzed in the generated text; selecting a pre-trained fraud analysis model according to the data characteristics of the text to be analyzed, and determining a fraud probability of the voice data through the analysis model. The application can realize accurate processing and analysis of voice signals, significantly improve the quality of voice data and the accuracy of fraud detection, and is suitable for complex voice interaction scenes.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of artificial intelligence, and in particular to a method, apparatus, device, and storage medium for voice fraud analysis. Background Technology

[0002] Current technologies in voice fraud analysis primarily rely on traditional speech recognition and rule-based detection systems. While these methods may be effective in handling simple, straightforward fraudulent activities, their accuracy and real-time performance are often limited when faced with complex dialogues and multi-channel data. In particular, existing systems are significantly inadequate in distinguishing between different speakers in two channels, understanding implicit intentions, and adapting to evolving fraud strategies. Furthermore, most existing technologies lack intelligent analytical capabilities, struggle to adapt to diverse fraud patterns, and perform poorly in terms of user customization and flexibility, failing to meet the specific needs of particular industries or businesses. These shortcomings limit the application potential of existing technologies in high-risk business scenarios, especially in the banking and insurance industries, where a more efficient and intelligent voice fraud analysis method is urgently needed.

[0003] Furthermore, black market activities and fraud are rampant in various business scenarios in the banking and insurance industries. For example, there are professional black market teams that engage in counter-collection, insurance fraud, and agency cancellation of policies, causing huge losses to banks and insurance companies. Currently, the identification of such situations mainly relies on manual judgment, which is not only inefficient but also difficult to guarantee accuracy. There is an urgent need for a more intelligent method to detect and warn of these fraudulent activities in a timely manner. Summary of the Invention

[0004] The main objective of this invention is to provide a voice fraud analysis method, apparatus, device, and storage medium, aiming to solve the technical problems of low efficiency and accuracy in identifying fraud scenarios in existing voice fraud analysis technologies.

[0005] To achieve the above objectives, the present invention provides a voice fraud analysis method, which includes the following steps:

[0006] Acquire a voice signal, and collect voice data from different sources from the voice signal through two independently configured microphone arrays, and record the voice data into two independent and isolated audio channels respectively;

[0007] Speech activity detection technology is used to detect speech data in two audio channels, identify and mark the speech activity time period in each audio channel, and perform a segmentation operation on the speech data based on the identified speech activity time period, dividing the continuous speech data in each audio channel into multiple independent speech segments to obtain multiple valid speech segments.

[0008] Generate dialogue text based on the effective speech segments, and determine the text to be analyzed in the dialogue text;

[0009] A pre-trained fraud analysis model is selected based on the data features of the text to be analyzed, and the fraud probability of the voice data is determined based on the text to be analyzed and the pre-trained fraud analysis model.

[0010] Optionally, the step of generating dialogue text based on the valid speech fragment includes:

[0011] Identify the start and end times of the valid speech segments in the speech data, and identify the source of the valid speech segments;

[0012] The valid audio segments are sorted according to the start time and the source to obtain time-series audio segments;

[0013] The dialogue text is generated based on the time-series speech segments.

[0014] Optionally, the step of generating dialogue text based on the temporal speech segment includes:

[0015] The time-series speech segments are converted into speech text using speech recognition technology, and the speech text is generated by tagging the speech text according to the source.

[0016] Optionally, the step of acquiring the voice signal includes:

[0017] Select a device that supports dual-channel recording, including a recording device with independent left and right channel recording capabilities;

[0018] Set the recording parameters of the recording device, including sampling rate and quantization accuracy;

[0019] Audio signals from two different directions or sources are recorded using the left and right channels of the recording device to obtain a stereo audio signal.

[0020] The dual-channel audio signal is preprocessed to identify the speech signal that belongs to human voice in the preprocessed dual-channel audio signal.

[0021] Optionally, the steps of identifying and marking the speech activity time periods in each audio channel, segmenting the speech data based on the identified speech activity time periods, and dividing the continuous speech data in each audio channel into multiple independent speech segments to obtain multiple valid speech segments include:

[0022] Construct a speech recognition model, detect speech activity indicators of the speech data through the speech recognition model, identify and mark time periods containing speech activity, and generate start and end time markers for each speech segment;

[0023] Segmentation is performed based on labeled speech activity time periods. Continuous speech data in each audio channel is segmented according to labeled time periods, and each segmented speech segment corresponds to an independent speech activity time period.

[0024] During the segmentation process, meaningless speech segments are excluded, including silent segments, background noise, and irrelevant interjections.

[0025] The effective speech segment is obtained, which is a filtered speech data segment that contains actual dialogue content.

[0026] Optionally, the step of determining the fraud probability of the voice data based on the text to be analyzed and the pre-trained fraud analysis model includes:

[0027] The text to be analyzed is input into the pre-trained fraud analysis model to determine the number of fraudulent clauses in the text to be analyzed.

[0028] The probability of fraud in the voice data is determined based on the number of fraud clauses and a preset set of fraud strategies.

[0029] Optionally, before the step of determining the fraud probability of the voice data based on the number of fraud terms and a preset fraud strategy set, the method further includes:

[0030] Obtain historical fraud data, and define the preset fraud strategy set based on business needs and the historical fraud data.

[0031] Furthermore, to achieve the above objectives, the present invention also provides a voice fraud analysis device, the voice fraud analysis device comprising:

[0032] The speech signal processing module acquires a speech signal, collects speech data from different sources from the speech signal through two independently configured microphone arrays, and records the speech data into two independent and isolated audio channels.

[0033] The effective segment extraction module uses speech activity detection technology to detect speech data in two audio channels, identify and mark the speech activity time period in each audio channel, and perform segmentation operation on the speech data based on the identified speech activity time period, dividing the continuous speech data in each audio channel into multiple independent speech segments to obtain multiple effective speech segments.

[0034] The text conversion module generates dialogue text based on the effective speech segments and determines the text to be analyzed in the dialogue text.

[0035] The fraud analysis module selects a pre-trained fraud analysis model based on the data features of the text to be analyzed, and determines the fraud probability of the voice data based on the text to be analyzed and the pre-trained fraud analysis model.

[0036] Furthermore, to achieve the above objectives, the present invention also provides a voice fraud analysis device, the voice fraud analysis device including a memory, a processor, and a voice fraud analysis program stored in the memory and executable on the processor, wherein the voice fraud analysis program, when executed by the processor, implements the steps of the voice fraud analysis method as described above.

[0037] Furthermore, to achieve the above objectives, the present invention also provides a storage medium storing a voice fraud analysis program, which, when executed by a processor, implements the steps of the voice fraud analysis method as described above.

[0038] This invention relates to a voice fraud analysis method. It involves acquiring voice signals by using two independently configured microphone arrays to collect voice data from different sources and recording this data into two independent and isolated audio channels. Voice activity detection technology is used to detect the voice data in both audio channels, identifying and labeling the voice activity time periods in each channel. The voice data is then segmented based on these activity time periods, and multiple valid voice segments are extracted. Dialogue text is generated from the extracted valid voice segments, and the text to be analyzed is selected from the generated text. A pre-trained fraud analysis model is selected based on the data characteristics of the text to be analyzed, and the fraud probability of the voice data is determined through the analysis model. This invention enables accurate processing and analysis of voice signals, significantly improving the quality of voice data and the accuracy of fraud detection, and is suitable for complex voice interaction scenarios. Attached Figure Description

[0039] Figure 1 This is a schematic diagram of the hardware operating environment of the device involved in the embodiment of the voice fraud analysis device of the present invention;

[0040] Figure 2 This is a flowchart illustrating the first embodiment of the voice fraud analysis method of the present invention;

[0041] Figure 3 This is a schematic diagram of the functional modules of a preferred embodiment of the voice fraud analysis device of the present invention.

[0042] The realization of the objective, functional features and advantages of the present invention will be further explained in conjunction with the embodiments and with reference to the accompanying drawings. Detailed Implementation

[0043] It should be understood that the specific embodiments described herein are for illustrative purposes only and are not intended to limit the scope of the invention.

[0044] It should be noted that existing technologies in the field of voice fraud analysis primarily rely on traditional speech recognition and rule-based detection systems. While these methods may be effective in handling simple, straightforward fraudulent activities, their accuracy and real-time performance are often limited when faced with complex dialogues and multi-channel data. In particular, existing systems are significantly inadequate in distinguishing between different speakers in two channels, understanding implicit intentions, and adapting to constantly evolving fraud strategies. Furthermore, most existing technologies lack intelligent analytical capabilities, struggle to adapt to diverse fraud patterns, and perform poorly in terms of user customization and flexibility, failing to meet the specific needs of particular industries or businesses. These shortcomings limit the application potential of existing technologies in high-risk business scenarios, especially in the banking and insurance industries, where a more efficient and intelligent voice fraud analysis method is urgently needed.

[0045] Furthermore, black market activities and fraud are rampant in various business scenarios in the banking and insurance industries. For example, there are professional black market teams that engage in counter-collection, insurance fraud, and agency cancellation of policies, causing huge losses to banks and insurance companies. Currently, the identification of such situations mainly relies on manual judgment, which is not only inefficient but also difficult to guarantee accuracy. There is an urgent need for a more intelligent method to detect and warn of these fraudulent activities in a timely manner.

[0046] To address the aforementioned shortcomings, this invention provides a voice fraud analysis device, referring to... Figure 1 , Figure 1 This is a schematic diagram of the hardware operating environment of the device involved in the embodiment of the voice fraud analysis device of the present invention.

[0047] like Figure 1 As shown, the voice fraud analysis device may include: a processor 1001, such as a CPU; a communication bus 1002; a user interface 1003; a network interface 1004; and a memory 1005. The communication bus 1002 is used to establish communication between these components. The user interface 1003 may include a display screen and an input unit such as a keyboard; optionally, the user interface 1003 may also include a standard wired interface or a wireless interface. The network interface 1004 may optionally include a standard wired interface or a wireless interface (such as a Wi-Fi interface). The memory 1005 may be high-speed RAM or non-volatile memory, such as a disk drive. Optionally, the memory 1005 may also be a storage device independent of the aforementioned processor 1001.

[0048] Those skilled in the art will understand that Figure 1 The hardware structure of the voice fraud analysis device shown in the figure does not constitute a limitation on the voice fraud analysis device. It may include more or fewer components than shown, or combine certain components, or have different component arrangements.

[0049] like Figure 1 As shown, the memory 1005, which serves as a storage medium, may include an operating system, a network communication module, a user interface module, and a voice fraud analysis program. The operating system is a program that manages and controls the voice fraud analysis device and software resources, supporting the operation of the network communication module, the user interface module, the voice fraud analysis program, and other programs or software. The network communication module manages and controls the network interface 1004; the user interface module manages and controls the user interface 1003.

[0050] exist Figure 1 In the hardware structure of the voice fraud analysis device shown, the network interface 1004 is mainly used to connect to the backend server and communicate with it; the user interface 1003 is mainly used to connect to the client and communicate with it; the processor 1001 can call the voice fraud analysis program stored in the memory 1005 and perform the following operations:

[0051] Acquire a voice signal, and collect voice data from different sources from the voice signal through two independently configured microphone arrays, and record the voice data into two independent and isolated audio channels respectively;

[0052] Speech activity detection technology is used to detect speech data in two audio channels, identify and mark the speech activity time period in each audio channel, and perform a segmentation operation on the speech data based on the identified speech activity time period, dividing the continuous speech data in each audio channel into multiple independent speech segments to obtain multiple valid speech segments.

[0053] Generate dialogue text based on the effective speech segments, and determine the text to be analyzed in the dialogue text;

[0054] A pre-trained fraud analysis model is selected based on the data features of the text to be analyzed, and the fraud probability of the voice data is determined based on the text to be analyzed and the pre-trained fraud analysis model.

[0055] Furthermore, the step of generating dialogue text based on the effective speech fragments includes:

[0056] Identify the start and end times of the valid speech segments in the speech data, and identify the source of the valid speech segments;

[0057] The valid audio segments are sorted according to the start time and the source to obtain time-series audio segments;

[0058] The dialogue text is generated based on the time-series speech segments.

[0059] Furthermore, the step of generating dialogue text based on the temporal speech segments includes:

[0060] The time-series speech segments are converted into speech text using speech recognition technology, and the speech text is generated by tagging the speech text according to the source.

[0061] Furthermore, the step of acquiring the voice signal includes:

[0062] Select a device that supports dual-channel recording, including a recording device with independent left and right channel recording capabilities;

[0063] Set the recording parameters of the recording device, including sampling rate and quantization accuracy;

[0064] Audio signals from two different directions or sources are recorded using the left and right channels of the recording device to obtain a stereo audio signal.

[0065] The dual-channel audio signal is preprocessed to identify the speech signal that belongs to human voice in the preprocessed dual-channel audio signal.

[0066] Furthermore, the steps of identifying and marking the speech activity time periods in each audio channel, segmenting the speech data based on the identified speech activity time periods, and dividing the continuous speech data in each audio channel into multiple independent speech segments to obtain multiple valid speech segments include:

[0067] Construct a speech recognition model, detect speech activity indicators of the speech data through the speech recognition model, identify and mark time periods containing speech activity, and generate start and end time markers for each speech segment;

[0068] Segmentation is performed based on labeled speech activity time periods. Continuous speech data in each audio channel is segmented according to labeled time periods, and each segmented speech segment corresponds to an independent speech activity time period.

[0069] During the segmentation process, meaningless speech segments are excluded, including silent segments, background noise, and irrelevant interjections.

[0070] The effective speech segment is obtained, which is a filtered speech data segment that contains actual dialogue content.

[0071] Furthermore, the step of determining the fraud probability of the voice data based on the text to be analyzed and the pre-trained fraud analysis model includes:

[0072] The text to be analyzed is input into the pre-trained fraud analysis model to determine the number of fraudulent clauses in the text to be analyzed.

[0073] The probability of fraud in the voice data is determined based on the number of fraud clauses and a preset set of fraud strategies.

[0074] Furthermore, before the step of determining the fraud probability of the voice data based on the number of fraud terms and a preset fraud strategy set, the method further includes:

[0075] Obtain historical fraud data, and define the preset fraud strategy set based on business needs and the historical fraud data.

[0076] The specific implementation of the voice fraud analysis device of the present invention is basically the same as the embodiments of the voice fraud analysis method described below, and will not be repeated here.

[0077] The present invention also provides a voice fraud analysis method based on the above-mentioned voice fraud analysis device.

[0078] Reference Figure 2 , Figure 2 This is a flowchart illustrating the first embodiment of the voice fraud analysis method of the present invention.

[0079] This invention provides an embodiment of a voice fraud analysis method. It should be noted that although the logical order is shown in the flowchart, in some cases, the steps shown or described may be performed in a different order than that shown here.

[0080] In various embodiments of the voice fraud analysis method, the executing entity is the voice fraud analysis device.

[0081] The voice fraud analysis method is applied to a voice fraud analysis device, and the voice fraud analysis method includes:

[0082] Step S10: Acquire the voice signal. Collect voice data from different sources from the voice signal using two independently configured microphone arrays, and record the voice data into two independent and isolated audio channels.

[0083] In this embodiment, during a call, voice signals are acquired through two independently configured microphone arrays. These microphone arrays are capable of capturing voice data from different directions or sources and recording this data into two independent and isolated audio channels. This step ensures that the voice signals recorded in each channel are clear and do not interfere with each other.

[0084] Two-channel audio signals captured by a microphone array naturally possess spatial separation characteristics, meaning the left channel corresponds to a speech signal from one source (such as the user's voice), and the right channel corresponds to a speech signal from another source (such as the other party's voice). This spatial separation is achieved through the independent configuration of the microphone array, ensuring that the speech data from the two channels can be processed and analyzed independently, without the need for complex signal separation processing later.

[0085] In the separated left and right channel audio data, speech activity detection (VAD) technology is used to identify the time periods of speech activity. VAD technology can accurately locate the actual time periods containing speech content in each channel, eliminate invalid segments (such as silence and noise), and extract valid speech segments, providing high-quality input for subsequent processing.

[0086] Step S20: Use speech activity detection technology to detect the speech data in the two audio channels respectively, identify and mark the speech activity time period in each audio channel, and perform a segmentation operation on the speech data based on the identified speech activity time period, dividing the continuous speech data in each audio channel into multiple independent speech segments to obtain multiple valid speech segments.

[0087] In this embodiment, a voice activity segment refers to the portion of an audio signal that actually contains speech. These segments are identified using Voice Activity Detection (VAD) technology, typically based on audio signal characteristics such as energy level, spectral properties, and zero-crossing rate. VAD is a technique used to identify and distinguish speech signals from non-speech signals (such as silence, background noise, etc.). By analyzing the energy, spectrum, and zero-crossing rate of the audio signal, VAD technology can detect the portion of the audio signal that actually contains speech activity and mark these segments. VAD is commonly used in real-time communication, voice recording, and voice processing systems to help reduce interference from non-speech data, thereby improving the system's processing efficiency and accuracy.

[0088] In a two-channel recording, each channel (left and right) may record speech signals from different sources (such as two different speakers). To ensure the validity and accuracy of the speech data in each channel, VAD technology is applied separately to the two audio channels.

[0089] Within each channel, the VAD operates independently, identifying the time segments that actually contain speech and marking these activity fragments. This allows the speech data from the two channels to be processed independently, avoiding mutual interference.

[0090] Segmentation refers to dividing a continuous audio signal into multiple independent speech segments based on the identified speech activity time periods. Each segment corresponds to a speech activity time period, ensuring that the segmented segments are logically independent and complete speech data. Based on the identified speech activity, corresponding time stamps are generated, recording the start and end times of the speech activity. According to these time stamps, the audio signal is segmented into several independent speech segments. The segmentation process is performed using signal processing algorithms, achieving millisecond-level accuracy and ensuring the accuracy of segment boundaries.

[0091] Effective speech segments refer to the segments that actually contain speech content after segmentation and filtering. These segments remove meaningless parts (such as silence, background noise, and irrelevant interjections) and retain only valuable data related to speech activity. After segmentation, the system further filters out segments that do not contain actual dialogue content. Filtering can be based on the intensity of speech activity, segment length, speech features, etc. The retained effective speech segments are high-quality speech data that can be used for subsequent text generation and speech analysis.

[0092] The process of identifying valid speech segments ensures that useful dialogue content is extracted from the recorded data while eliminating background noise and meaningless filler words. Valid speech segments in each channel are detected and labeled in real time, while meaningless vocalizations (such as "um," "ah," etc.) are excluded. Edge computing is used to achieve local real-time speech activity detection, reducing latency and improving efficiency. The steps may include:

[0093] First, feature extraction is performed, such as Mel frequency cepstral coefficients (MFCCs), spectra, energy, and zero-crossing rate. These features can help the model distinguish speech from meaningless noise.

[0094] Mel-frequency cepstral coefficients (MFCCs): MFCCs are features extracted from speech signals and can effectively represent the spectral characteristics of speech.

[0095] Spectrogram: A spectrogram represents the energy distribution of a speech signal at different frequencies.

[0096] Energy: The energy characteristics of a speech signal, which can be used to distinguish speech from silence.

[0097] Zero Crossing Rate: This indicates the number of times a signal crosses zero and can be used to detect speech activity.

[0098] Then, valid speech segments can be identified from the speech data through speech activity detection. Speech activity detection is used to distinguish speech signals from silence and identify valid speech segments. The steps may include:

[0099] Choose a suitable VAD algorithm, such as a simple energy-based VAD, a statistical model-based VAD, or a deep learning model (such as CNN or RNN) for detection. Deep learning models can be trained to recognize speech activities more accurately.

[0100] Extracted speech features (such as MFCCs, spectrograms, etc.) are input into the VAD model. The model outputs the probability or binary result (speech / non-speech) of whether each time frame contains speech activity. The detection results are smoothed to reduce false positives caused by noise. Methods such as sliding windows or low-pass filters can be used. Meaningless filler words and background noise are filtered out, retaining valid dialogue content.

[0101] The extracted and separated valid speech segments are converted into text to generate call logs. Then, the text portions that require further analysis are filtered and marked for semantic analysis or fraud detection.

[0102] Step S30: Generate dialogue text based on the effective speech segments, and determine the text to be analyzed in the dialogue text;

[0103] In this embodiment, after identifying valid speech segments, these segments need to be converted into dialogue text, and the text to be analyzed for fraud analysis needs to be further determined. The steps may include:

[0104] First, determine the speech recognition model. One or more speech recognition models can be used, such as deep neural networks (DNN), convolutional neural networks (CNN), and Transformer models. Ensemble learning can improve the transcription accuracy. Multi-model ensemble can effectively reduce the error of a single model and improve the overall recognition performance.

[0105] Furthermore, the speech recognition model can be trained according to user needs, adapting it to the specialized terminology and expressions of different fields. For example, the terminology used in the financial industry differs from that in the insurance industry, requiring separate training and adjustment of the model.

[0106] Furthermore, speech recognition models can be trained to support the recognition of multiple languages ​​and dialects, improving their applicability in multilingual environments. For example, by adding training data on common dialects, the application performance in different regions can be improved.

[0107] During the process of transcribing valid speech segments into text using a speech recognition model, the timestamp information of each speech segment is maintained to ensure the synchronization of the generated text with the original speech. Timestamp synchronization helps to quickly locate specific speech segments during subsequent analysis.

[0108] The transcribed text fragments are then integrated chronologically to generate a complete dialogue text. This ensures the dialogue is natural, fluent, and logical, facilitating subsequent analysis.

[0109] Label each sentence in the dialogue text with the speaker's identity (e.g., customer service or customer) to ensure that the analysis can distinguish the speech of different roles. Speaker labeling can be done using voiceprint recognition or binaural recording data.

[0110] Clean up the dialogue text by removing noise words and meaningless filler words such as "um," "ah," and "oh." Pre-trained language models (such as GPT-4 and BERT) can be used for semantic analysis to identify and remove meaningless content.

[0111] Finally, identify the text to be analyzed from the generated dialogue text that may be susceptible to fraud detection. The steps may include:

[0112] One implementation involves selecting text from a specific source as the text to be analyzed from the dialogue text. For example, in a dialogue between a customer service representative and a customer, the customer's dialogue text is selected as the text to be analyzed.

[0113] Another approach is to extract key phrases and descriptions of potential fraudulent activities from the generated dialogue text. Natural Language Processing (NLP) techniques, such as keyword extraction and topic modeling, can be used to identify important information in the dialogue. Alternatively, the dialogue content can be understood based on context to identify key paragraphs and sentences, which can then be used as the text to be analyzed. For example, through semantic analysis and contextual correlation, important information such as customer inquiries and customer service responses can be identified.

[0114] Step 40: Select a pre-trained fraud analysis model based on the data features of the text to be analyzed, and determine the fraud probability of the voice data based on the text to be analyzed and the pre-trained fraud analysis model.

[0115] In this embodiment, data features refer to information extracted from the text to be analyzed, including but not limited to keywords, semantic patterns, grammatical structures, contextual relationships, and sentiment analysis results. These features reflect the nature of the text content and provide a basis for model selection and analysis.

[0116] Use Natural Language Processing (NLP) techniques to extract valuable features from text. This may include bag-of-words models, word embeddings, TF-IDF values, sentence structure analysis, and named entity recognition.

[0117] Pre-trained models refer to machine learning or deep learning models that have undergone preliminary training on a large corpus. These models already possess strong language understanding capabilities and can be fine-tuned to adapt to specific tasks.

[0118] Based on the extracted text features, the most suitable pre-trained model for analyzing the text is selected. Selection criteria may include the model's architecture (e.g., BERT, GPT-3, RoBERTa), its performance on similar tasks, and its ability to handle domain-specific terminology. The optimal model is determined by comparing the text features with the model's capabilities. For example, text containing a large amount of financial terminology might require selecting a model pre-trained in the financial domain.

[0119] Fraud probability refers to the likelihood that text content contains fraudulent activity, calculated through model analysis. This probability value is typically represented as a decimal between 0 and 1, with values ​​closer to 1 indicating a higher probability of fraud. The text to be analyzed is input into a selected pre-trained model. The model analyzes the text features using its inherent parameters and structure, and outputs a fraud probability. This process may involve calculations using multi-layered neural networks, feature vector matching, and the application of risk assessment algorithms.

[0120] This invention relates to a voice fraud analysis method. It involves acquiring voice signals by using two independently configured microphone arrays to collect voice data from different sources and recording this data into two independent and isolated audio channels. Voice activity detection technology is used to detect the voice data in both audio channels, identifying and labeling the voice activity time periods in each channel. The voice data is then segmented based on these activity time periods, and multiple valid voice segments are extracted. Dialogue text is generated from the extracted valid voice segments, and the text to be analyzed is selected from the generated text. A pre-trained fraud analysis model is selected based on the data characteristics of the text to be analyzed, and the fraud probability of the voice data is determined through the analysis model. This invention enables accurate processing and analysis of voice signals, significantly improving the quality of voice data and the accuracy of fraud detection, and is suitable for complex voice interaction scenarios.

[0121] Furthermore, a second embodiment of the voice fraud analysis method of the present invention is proposed, wherein the step S20 above, the step of generating dialogue text based on the effective voice segment, includes:

[0122] Step a1: Identify the start and end times of the valid speech segment in the speech data, and identify the source of the valid speech segment;

[0123] Step a2: Sort the valid speech segments according to the start time and the source to obtain time-series speech segments;

[0124] Step a3: Generate dialogue text based on the time-series speech segments.

[0125] In this embodiment, the start and end times of each valid speech segment in the speech data are identified. Timestamp information is crucial for ensuring the timing and synchronization of the dialogue. The source of each valid speech segment (e.g., customer service or customer) is determined. The voices of different speakers can be distinguished using dual-channel recording data, voiceprint recognition technology, or other identification methods.

[0126] Based on the start time and source of the identified audio, all valid audio segments are sorted to obtain chronological audio segments. This ensures the natural flow and logical consistency of the dialogue.

[0127] The sorted temporal speech segments are converted into text to generate a complete dialogue text. The temporal sequence and coherence of the dialogue text are ensured to facilitate subsequent analysis.

[0128] Label each sentence in the dialogue text with the speaker's identity (such as customer service or customer) to ensure that the analysis can distinguish the speech of different roles.

[0129] For example:

[0130] Segment 1: Start time - 00:00:05, End time - 00:00:10;

[0131] Segment 2: Start time - 00:00:12, End time - 00:00:15;

[0132] Segment 3: Start time - 00:00:08, End time - 00:00:11.

[0133] Identify the source of a speech segment.

[0134] Segment 1: Source - Customer Service;

[0135] Segment 2: Source - Customer;

[0136] Segment 3: Source - Client.

[0137] Based on the start time and source of the identified audio, all valid audio segments are sorted to obtain chronological audio segments. This ensures the natural flow and logical consistency of the dialogue.

[0138] Segment 1: Start time - 00:00:05, End time - 00:00:10, Source - Customer service;

[0139] Segment 3: Start time - 00:00:08, End time - 00:00:11, Source - Client;

[0140] Segment 2: Start time - 00:00:12, End time - 00:00:15, Source - Client.

[0141] The sorted time-series speech segments are converted into text to generate complete dialogue text.

[0142] Segment 1 (Customer Service): Hello, this is the bank's customer service center. How can I help you?

[0143] Segment 3 (Customer): I have some questions about my credit card bill.

[0144] Segment 2 (Customer): There's an expense on the bill that I don't remember.

[0145] Clean the dialogue text, removing noise words and meaningless filler words such as "um," "ah," and "oh." Also, label the speaker's identity (e.g., customer service or customer) for each sentence in the dialogue text to ensure that the analysis can distinguish the speech of different roles.

[0146] Customer Service: Hello, this is the bank's customer service center. How can I help you?

[0147] Customer: I have some questions about my credit card bill.

[0148] Customer: There's an expense on the bill that I don't remember.

[0149] This embodiment, through the above steps, can improve the naturalness, fluency, and logic of the dialogue, providing high-quality input data for subsequent fraud analysis.

[0150] Furthermore, a third embodiment of the voice fraud analysis method of the present invention is proposed, wherein step a3 above includes:

[0151] Step b1: Use speech recognition technology to convert the time-series speech segment into speech text, and generate the dialogue text by marking the speech text according to the source.

[0152] In this embodiment, after identifying valid speech segments, these segments need to be converted into dialogue text, and the text to be analyzed for fraud analysis needs to be further determined. The step of generating dialogue text from temporal speech segments may include:

[0153] Speech recognition technology can be used to convert sequential speech segments into speech text. Various speech recognition models can be employed, such as deep neural networks (DNNs), convolutional neural networks (CNNs), and Transformer models. Ensemble learning can improve the accuracy of transcription. Multi-model ensembles can effectively reduce the error of a single model and improve the overall recognition performance. Annotating the speaker's identity (e.g., customer service representative or customer) for each sentence in the dialogue text ensures that the analysis can distinguish the speech of different roles.

[0154] For example:

[0155] Segment 1: Start time - 00:00:05, End time - 00:00:10, Source - Customer service;

[0156] Segment 3: Start time - 00:00:08, End time - 00:00:11, Source - Client;

[0157] Segment 2: Start time - 00:00:12, End time - 00:00:15, Source - Client.

[0158] The sorted time-series speech segments are converted into text to generate complete dialogue text.

[0159] Segment 1 (Customer Service): Hello, this is the bank's customer service center. How can I help you?

[0160] Segment 3 (Customer): I have some questions about my credit card bill.

[0161] Segment 2 (Customer): There's an expense on the bill that I don't remember.

[0162] Clean the dialogue text, removing noise words and meaningless filler words such as "um," "ah," and "oh." Also, label the speaker's identity (e.g., customer service or customer) for each sentence in the dialogue text to ensure that the analysis can distinguish the speech of different roles.

[0163] Customer Service: Hello, this is the bank's customer service center. How can I help you?

[0164] Customer: I have some questions about my credit card bill.

[0165] Customer: There's an expense on the bill that I don't remember.

[0166] This embodiment uses the above method to accurately convert effective speech segments into dialogue text and to mark the source of the dialogue content, so as to filter the text to be analyzed based on the dialogue subject.

[0167] Furthermore, a fourth embodiment of the voice fraud analysis method of the present invention is proposed, wherein the step of acquiring the voice signal in step S10 above includes:

[0168] Step c1: Select a device that supports dual-channel recording, including a recording device with independent left and right channel recording functions;

[0169] Step c2: Set the recording parameters of the recording device, including sampling rate and quantization accuracy;

[0170] Step c3: Record audio signals from two different directions or sources through the left and right channels of the recording device to obtain a stereo audio signal;

[0171] Step c4: Preprocess the dual-channel audio signal to identify the speech signal that belongs to human voice in the preprocessed dual-channel audio signal.

[0172] In this embodiment, a device capable of recording two-channel audio signals is selected. A two-channel recording device has independent left and right channels, capable of simultaneously capturing sound signals from two different directions or sources. This ensures that the recorded audio signal has spatial separation characteristics, allowing subsequent processing to better identify and distinguish speech signals from different sources. Key parameters of the recording device are set, such as sampling rate and quantization precision. The sampling rate determines the number of audio samples recorded per second; a high sampling rate helps retain more audio details. Quantization precision determines the dynamic range and signal-to-noise ratio of the audio signal; higher quantization precision improves the clarity and sound quality of the recording. Microphone arrays for the left and right channels record audio signals from two different directions or sources, respectively. This step ensures that the recorded audio signal accurately reflects the sounds from different sources and maintains the independence between channels, which is helpful for subsequent signal processing and analysis. After recording, the two-channel audio signal is preprocessed.

[0173] Generally, call content includes not only spoken dialogue but also various environmental noises and other audio. It's necessary to extract the spoken dialogue data from this cluttered audio. High-fidelity stereo recording equipment can be used to acquire two-channel audio data, capturing the dialogue between customer service representatives and customers. This audio data can come from call center recordings, telephone customer service recordings, etc.

[0174] The collected audio data is preprocessed, including: using noise suppression algorithms to reduce background noise interference; ensuring consistent volume across different audio sources to improve speech recognition accuracy; and reducing the impact of echoes on the speech data to ensure audio quality.

[0175] Identify human voices from preprocessed two-channel audio data. Distinguish between human voices and background noise using feature extraction and voice activity detection (VAD) techniques.

[0176] Feature extraction methods can include:

[0177] Mel frequency cepstral coefficients (MFCCs) were extracted, and 13 MFCCs were selected as spectral features of the speech signal. The specific steps are as follows: The audio signal was pre-emphasized to enhance high-frequency components. The signal was framed, with each frame typically 20-40ms in length. Each frame was windowed, such as using a Hamming window. The Fast Fourier Transform (FFT) of each frame was calculated to obtain the power spectrum. A Mel filter bank was applied to perform a weighted summation of the power spectrum. The logarithm of the Mel power spectrum was taken, and the Discrete Cosine Transform (DCT) was calculated to obtain the MFCCs.

[0178] Extract the spectrogram, use Short Time Fourier Transform (STFT) to generate the spectrum of the audio signal, decompose the audio signal into multiple time frames, perform Fourier transform on each frame to obtain the frequency distribution.

[0179] Energy is extracted by calculating the energy of each frame of the audio signal, which serves as a feature to distinguish speech from silence. Frames with higher energy typically contain speech signals.

[0180] Zero-crossing rate is extracted by calculating the zero-crossing rate of each frame of the signal, which is the number of times the signal crosses zero. Zero-crossing rate can be used to detect speech activity because speech signals have a relatively high zero-crossing rate.

[0181] Then, a suitable VAD algorithm is selected, which can be an energy-based method, a statistical model-based method, or a deep learning model (such as CNN or RNN). The extracted MFCCs, spectrograms, energy, and zero-crossing rate features are input into the VAD model. The VAD model is trained on a large amount of speech data and can accurately distinguish between speech and non-speech components.

[0182] This embodiment, through the steps described above, accurately identifies human voice data from the preprocessed two-channel audio data. This ensures high accuracy and reliability for subsequent speech-to-text and fraud analysis.

[0183] Furthermore, a fifth embodiment of the voice fraud analysis method of the present invention is proposed. In step S10 above, the steps of identifying and marking the voice activity time period in each audio channel, performing a segmentation operation on the voice data based on the identified voice activity time period, and dividing the continuous voice data in each audio channel into multiple independent voice segments to obtain multiple valid voice segments include:

[0184] Step d1: Construct a speech recognition model, detect speech activity indicators of the speech data through the speech recognition model, identify and mark time periods containing speech activity, and generate start and end time markers for each speech segment;

[0185] Step d2: Segment the speech data based on the marked speech activity time periods. The continuous speech data in each audio channel is segmented according to the marked time periods, and each segmented speech segment corresponds to an independent speech activity time period.

[0186] Step d3: During the segmentation process, meaningless speech segments are excluded, including silent segments, background noise, and irrelevant interjections.

[0187] Step d4: Obtain the effective speech segment, which is a filtered speech data segment containing actual dialogue content.

[0188] In this embodiment, a model specifically designed for speech activity detection is constructed. This model identifies which time periods contain actual speech activity by analyzing speech activity indicators in speech data. Speech activity indicators may include the energy level of the audio signal, spectral characteristics, zero-crossing rate, phoneme identification, etc.

[0189] This model determines the presence of speech activity by analyzing various indicators in the speech signal. For example, when the energy of the audio signal exceeds a certain threshold, speech activity can be considered to exist within that time period. The model generates start and end time markers for each speech segment based on these indicators.

[0190] Based on detected speech activity metrics, the model identifies time periods containing speech activity. It then generates corresponding time stamps for these time periods, recording the start and end times of the speech activity. These time stamps provide precise location information for the speech activity periods within the audio signal, indicating the start and end points of each speech segment. These stamps guide subsequent segmentation operations.

[0191] After identifying and labeling the speech activity time periods, the continuous speech data in each audio channel is segmented according to these labeled time periods. Each segmented speech fragment corresponds to an independent speech activity time period. The segmentation operation ensures that each fragment is logically independent, meaning that each fragment represents a complete speech activity time period. This independence makes subsequent processing and analysis more convenient and accurate.

[0192] The model can be implemented using architectures such as Convolutional Neural Networks (CNNs), Recurrent Neural Networks (RNNs), or Transformer-based architectures. Choosing a suitable speech recognition model involves energy-based methods, statistical model-based methods, or deep learning models (such as CNNs and RNNs). Deep learning models are typically trained on large amounts of data and are able to recognize speech activities more accurately. The main function of the model is to analyze the speech data in each audio channel and identify the time periods containing speech activities. These time periods may include the moment the speaker utters words, while excluding silence and noise interference.

[0193] Definition of meaningless speech segments: Meaningless speech segments include silent segments, background noise, and interjections that do not involve actual dialogue (such as "um," "ah," etc.). These segments usually do not contain valuable information and may interfere with subsequent speech recognition and analysis.

[0194] Through preprocessing steps and a speech recognition model, the system can automatically detect and eliminate these meaningless segments. Silent segments and background noise can be identified and filtered using energy detection or spectral analysis, while interjections can be eliminated through semantic analysis or specific keyword filtering mechanisms.

[0195] One implementation method that can exclude meaningless human voices is:

[0196] To filter meaningless human voices based on rules, firstly, define meaningless human voice words. For example, define a set of common meaningless words (such as "um", "ah", "oh", "uh-huh", etc.).

[0197] After converting the speech data to text, filtering is performed by matching these predefined meaningless words. Regular expressions or string matching methods can be used to remove these meaningless words from the text.

[0198] For example, suppose a dialogue, when transcribed into text, contains meaningless words:

[0199] Original text: "Um, I'd like to ask about something, uh-huh, there's a problem with my bill."

[0200] The filtered text reads: "I'd like to inquire about a problem with my bill."

[0201] Another implementation method could be:

[0202] Semantic analysis based on NLP models utilizes pre-trained language models (such as BERT and GPT) for semantic analysis. These pre-trained models can understand context, thus identifying meaningless content in dialogue. Semantic analysis is performed on each sentence using the pre-trained model to identify filler words, modal particles, and other meaningless content in the dialogue. This meaningless content is then filtered out, retaining only meaningful speech segments.

[0203] For example, using a pre-trained model to perform semantic analysis on text, identifying and filtering meaningless words:

[0204] Original text: "Oh, I'd like to know how much this service costs?"

[0205] The model identifies and filters the text: "I'd like to know how much this service costs?".

[0206] Definition of a valid speech segment: A valid speech segment refers to a segment of speech data that, after processing by a speech recognition model, retains actual speech content containing meaningful dialogue. These segments have had silence, noise, and irrelevant content removed, ensuring high-quality speech data.

[0207] After identifying and eliminating meaningless segments, the remaining speech data are the valid speech segments. These segments will be used for subsequent speech-to-text processing, text analysis, and fraud detection.

[0208] This embodiment, through the steps described above, can effectively eliminate meaningless human voices from speech data, retaining meaningful speech segments. This ensures high accuracy and reliability for subsequent speech-to-text and fraud analysis.

[0209] Furthermore, a sixth embodiment of the voice fraud analysis method of the present invention is proposed. In step S30 above, the step of determining the fraud probability of the voice data based on the text to be analyzed and the pre-trained fraud analysis model includes:

[0210] Step e1: Input the text to be analyzed into the pre-trained fraud analysis model to determine the number of fraudulent clauses in the text to be analyzed;

[0211] Step e2: Determine the fraud probability of the voice data based on the number of fraud clauses and the preset fraud strategy set.

[0212] In this embodiment, to determine the probability of fraud in voice data, the text to be analyzed needs to be input into the analysis model to identify the number of fraudulent clauses, and the fraud probability needs to be determined based on a preset set of fraud strategies. Specific methods may include:

[0213] Basic preprocessing is performed on the text to be analyzed, such as removing stop words and punctuation. The preprocessed text is then input into a pre-trained analysis model. This model can be a deep learning-based language model (such as BERT or GPT) that has been fine-tuned using a large amount of fraud-related data.

[0214] The analysis model performs semantic analysis on the input text to identify fraudulent clauses. Fraudulent clauses are key phrases or sentences in the text related to fraudulent activities. The number of identified fraudulent clauses is counted. Each identified fraudulent clause represents a potential fraudulent activity.

[0215] A pre-defined set of fraud strategies is provided, which includes a range of fraud probabilities corresponding to different numbers of fraudulent terms. These strategies can be set through statistical analysis of historical data and expert experience.

[0216] Based on the number of identified fraudulent terms, the system searches for the corresponding fraud probability range within a pre-defined set of fraud strategies. Probability calculations can be performed using a rule engine or a logical reasoning system.

[0217] For example, suppose that after the valid speech segments extracted from the speech data are converted into text, the determined text to be analyzed is:

[0218] "Hello, my credit card has been stolen. Can you help me cancel this transaction?"

[0219] "Of course, we will process your request immediately. Please provide your card number and ID information."

[0220] Preprocess the text to be analyzed by removing irrelevant words and punctuation:

[0221] "Hello, my credit card has been stolen. Can you help me cancel this transaction?"

[0222] "Of course, we will process your request immediately. Please provide your card number and ID information."

[0223] The pre-processed text is input into the pre-trained fraud analysis model.

[0224] The analysis model identified fraudulent terms in the text:

[0225] Fraud Clause 1: "Credit card stolen";

[0226] Fraudulent Clause 2: "Cancel this transaction";

[0227] Fraud Clause 3: "Provide your card number and ID information".

[0228] Number of fraudulent clauses identified: 3.

[0229] Pre-set a set of fraud strategies, for example:

[0230] 0-1 fraudulent terms: Low probability of fraud (0-20%);

[0231] 2-3 fraudulent terms: Medium probability of fraud (20-60%);

[0232] Four or more fraudulent clauses: High probability of fraud (60-100%).

[0233] Based on the three identified fraudulent terms, the fraud probability range corresponding to the preset fraud strategy set was found and determined to be medium fraud probability (20-60%).

[0234] This embodiment, through the steps described above, can effectively identify and count fraudulent clauses in the text to be analyzed, and determine the fraud probability of the voice data based on a preset set of fraud strategies. This ensures the accuracy and reliability of fraud analysis.

[0235] Furthermore, a seventh embodiment of the voice fraud analysis method of the present invention is proposed, which includes the following steps before step e2:

[0236] Step f1: Obtain historical fraud data and define the preset fraud strategy set based on business needs and the historical fraud data.

[0237] In this embodiment, in order to define a preset set of fraud strategies, it is necessary to obtain historical fraud data and set specific strategies based on business needs and data analysis.

[0238] First, historical fraud data can be obtained. Data sources can include internal databases, such as internal transaction records, customer service conversations, and customer complaint records; industry reports, such as fraud analysis reports and research papers in the financial industry; or public cases, such as fraud cases and news reports obtained through the internet and public resources.

[0239] Data preprocessing includes cleaning the data, removing invalid records, duplicate data, and missing values ​​to ensure accuracy and completeness; format standardization, unifying the formats of data from different sources to facilitate subsequent processing and analysis; and data labeling, which identifies legitimate transactions from fraudulent ones. This can be done manually or using automated classification tools.

[0240] The cleaned and standardized data is stored in a database or data warehouse to facilitate subsequent analysis and model training.

[0241] Then, data analysis is performed, which may include the following steps:

[0242] Fraudulent clause extraction uses Natural Language Processing (NLP) techniques to extract common fraudulent clauses from historical fraud data. For example, it identifies key phrases and sentences in the text through keyword extraction and entity recognition.

[0243] Fraud pattern recognition: This involves using machine learning algorithms to identify fraud patterns in historical data. Classification models (such as decision trees, random forests, and support vector machines) can be used to identify common features of fraudulent behavior. The process of using classification models is as follows: extract features from the raw data, such as transaction amount, transaction time, transaction location, and customer behavior; train the classification model using labeled historical data to learn fraud patterns; and evaluate the model's accuracy, recall, F1 score, etc., to ensure the model's reliability.

[0244] Statistical analysis is performed on historical fraud data to calculate the frequency and combination of different fraudulent terms. For example, the number of times each fraudulent term appears in all fraudulent transactions, and the probability of multiple fraudulent terms appearing simultaneously, are calculated.

[0245] Based on data analysis results and business needs, define specific rules for a pre-set set of fraud strategies. The strategy set should include the probability range of fraud corresponding to different numbers of fraudulent clauses. Specific steps may include: setting risk weights for different fraudulent clauses based on statistical analysis results. For example, clauses that appear frequently can be assigned higher risk weights. Define the probability of fraud when multiple fraudulent clauses appear simultaneously. For example, certain clauses appearing together may indicate a higher risk of fraud.

[0246] Furthermore, identification and analysis can be conducted based on business needs, and the specific parameters of the strategy set can be adjusted. For example, for high-risk transactions (such as international transactions and large transfers), the threshold for triggering fraud detection can be lowered; for low-risk transactions (such as local small-amount consumption), the threshold can be raised.

[0247] Finally, the strategy set can be validated using historical data to ensure its rationality and effectiveness. Validation can be performed using methods such as cross-validation or A / B testing, and the strategy set parameters can be adjusted based on the validation results.

[0248] One embodiment is as follows:

[0249] We collect credit card transaction records from the past five years, including both fraudulent and legitimate transactions, from the bank's internal database. We also gather common fraud cases and patterns from financial industry research reports. We clean the data, removing invalid and duplicate records to ensure accuracy. Finally, we standardize the data format across different sources for easier subsequent processing.

[0250] Based on the bank's internal audit findings and customer complaint records, fraud and legitimate labels are added to transaction data. NLP techniques are used to extract common fraudulent terms, such as "credit card stolen" and "unauthorized transaction," from transaction notes and customer conversations. Decision tree models are used to identify fraud patterns in historical data, extracting features such as transaction amount, transaction location, time, and device information.

[0251] Train the model and evaluate its accuracy, recall, and F1 score to ensure it can accurately identify fraudulent activities. Analyze the frequency of each fraudulent clause in historical fraudulent transactions, as well as the combinations of different clauses. For example, calculate the probability of "unauthorized transaction" and "credit card theft" occurring simultaneously.

[0252] Based on the statistical analysis results, a preset set of fraud strategies is set:

[0253] 0-1 fraudulent terms: Low probability of fraud (0-20%);

[0254] 2-3 fraudulent terms: Medium probability of fraud (20-60%);

[0255] Four or more fraudulent clauses: High probability of fraud (60-100%).

[0256] For high-risk international transactions, set stricter fraud detection thresholds. For example:

[0257] 1 fraud clause: Medium probability of fraud (20-60%);

[0258] Two or more fraudulent clauses: High probability of fraud (60-100%).

[0259] The strategy set is validated using trading data from the past year. The strategy set parameters are adjusted through cross-validation to ensure the accuracy and reliability of the strategy set.

[0260] Through the above steps, this embodiment can define a reasonable set of preset fraud strategies based on historical fraud data and business needs. Based on historical data and statistical analysis, it ensures the scientific nature and accuracy of the strategy set, and ensures that fraudulent behavior can be accurately identified and prevented in practical applications, thereby improving business security.

[0261] The present invention also provides a voice fraud analysis device.

[0262] Reference Figure 3 , Figure 3 This is a functional module diagram of a first embodiment of the voice fraud analysis device of the present invention. The voice fraud analysis device includes:

[0263] The speech signal processing module acquires a speech signal, collects speech data from different sources from the speech signal through two independently configured microphone arrays, and records the speech data into two independent and isolated audio channels.

[0264] The effective segment extraction module uses speech activity detection technology to detect speech data in two audio channels, identify and mark the speech activity time period in each audio channel, and perform segmentation operation on the speech data based on the identified speech activity time period, dividing the continuous speech data in each audio channel into multiple independent speech segments to obtain multiple effective speech segments.

[0265] The text conversion module generates dialogue text based on the effective speech segments and determines the text to be analyzed in the dialogue text.

[0266] The fraud analysis module selects a pre-trained fraud analysis model based on the data features of the text to be analyzed, and determines the fraud probability of the voice data based on the text to be analyzed and the pre-trained fraud analysis model.

[0267] Furthermore, embodiments of the present invention also propose a storage medium.

[0268] The storage medium stores a voice fraud analysis program, which, when executed by the processor, implements the steps of the voice fraud analysis method described above.

[0269] The specific implementation of the storage medium of the present invention is basically the same as the embodiments of the above-described voice fraud analysis method, and will not be repeated here.

Claims

1. A voice fraud analysis method, characterized by, The voice fraud analysis method comprises the following steps: Acquire voice signals, collect voice data of different sources from the voice signals through two independently configured microphone arrays, and record the voice data into two independent and isolated audio channels respectively; Detect the voice data in the two audio channels using voice activity detection technology, identify and mark the voice activity time period in each audio channel, perform fragmentation on the voice data based on the identified voice activity time period, cut the continuous voice data in each audio channel into multiple independent voice segments, and acquire multiple valid voice segments; Generate a dialogue text according to the valid voice segments, determine the text to be analyzed in the dialogue text; Select a pre-trained fraud analysis model according to the data characteristics of the text to be analyzed, and determine the fraud probability of the voice data according to the text to be analyzed and the pre-trained fraud analysis model.

2. The voice fraud analysis method of claim 1, wherein, The step of generating a dialogue text according to the valid voice segments comprises: Identify the start time and end time of the valid voice segment in the voice data, and identify the source person of the valid voice segment; Sort the valid voice segments according to the start time and the source person to obtain time-sequential voice segments; Generate a dialogue text according to the time-sequential voice segments.

3. The voice fraud analysis method as claimed in claim 2, characterized in that, The step of generating a dialogue text according to the time-sequential voice segments comprises: Convert the time-sequential voice segments into voice text using voice recognition technology, mark the voice text according to the source person, and generate the dialogue text.

4. The voice fraud analysis method of claim 1, wherein, The step of acquiring voice signals comprises: Select a device supporting dual-channel recording, wherein the device comprises a recording device with independent left-channel and right-channel recording functions; Set the recording parameters of the recording device, including the sampling rate and quantization accuracy; Record audio signals from two different sources through the left channel and the right channel of the recording device to obtain dual-channel audio signals; Preprocess the dual-channel audio signals, and identify the voice signals belonging to human voice in the preprocessed dual-channel audio signals.

5. The voice fraud analysis method of claim 1, wherein, The step of identifying and marking the voice activity time period in each audio channel, performing fragmentation on the voice data based on the identified voice activity time period, cutting the continuous voice data in each audio channel into multiple independent voice segments, and acquiring multiple valid voice segments comprises: Construct a voice recognition model, detect the voice activity indicators of the voice data through the voice recognition model, identify and mark the time period containing voice activity, and generate the start and end time markers of each voice segment; Perform fragmentation based on the marked voice activity time period, cut the continuous voice data in each audio channel according to the marked time period, and each voice segment after cutting corresponds to an independent voice activity time period; Exclude meaningless voice segments, including silent segments, background noise, and irrelevant emotional words, during the fragmentation process; Acquire the valid voice segments, which are voice data segments containing actual dialogue content after filtering.

6. The voice fraud analysis method of claim 1, wherein, The step of determining the fraud probability of the voice data according to the text to be analyzed and the pre-trained fraud analysis model comprises: inputting the text to be analyzed into the pre-trained fraud analysis model to determine the number of fraud clauses in the text to be analyzed; determining the fraud probability of the voice data according to the number of fraud clauses and a preset fraud strategy set.

7. The voice fraud analysis method of claim 6, wherein, Before the step of determining the fraud probability of the voice data according to the number of fraud clauses and a preset fraud strategy set, the method further comprises: obtaining historical fraud data, and defining the preset fraud strategy set according to business requirements and the historical fraud data.

8. A voice fraud analysis apparatus, characterized by, The voice fraud analysis device comprises: a voice signal processing module configured to obtain a voice signal, collect voice data of different sources from the voice signal through two independently configured microphone arrays, and record the voice data into two independent and isolated audio channels, respectively; an effective segment extraction module configured to detect the voice data in the two audio channels using a voice activity detection technology, identify and mark a voice activity time period in each audio channel, perform a segmentation operation on the voice data based on the identified voice activity time period, cut continuous voice data in each audio channel into a plurality of independent voice segments, and obtain a plurality of effective voice segments; a text conversion module configured to generate a dialogue text according to the effective voice segments, and determine a text to be analyzed in the dialogue text; a fraud analysis module configured to select a pre-trained fraud analysis model according to a data feature of the text to be analyzed, and determine the fraud probability of the voice data according to the text to be analyzed and the pre-trained fraud analysis model.

9. A voice fraud analysis device, characterized by, The voice fraud analysis device comprises a memory, a processor, and a voice fraud analysis program stored on the memory and executable on the processor, and the voice fraud analysis program, when executed by the processor, implements the steps of the voice fraud analysis method according to any one of claims 1-7.

10. A storage medium, characterized by The storage medium stores a voice fraud analysis program, and the voice fraud analysis program, when executed by the processor, implements the steps of the voice fraud analysis method according to any one of claims 1-7.

Citation Information

Patent Citations

  • Device and method for preventing telephone fraud

    CN117278670A

  • Speech recognition appliance and a control method of the same

    KR102163927B1