Account anomaly determination method and device, storage medium and electronic device
By analyzing connection data and voice feature matching in voice calls, and combining this with account tag information, and dynamically setting thresholds, the problem of inaccurate account anomaly detection in financial services is solved, achieving more efficient anomaly detection.
Patent Information
- Application Number
- CN202411955050.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-27
- Publication Date
- 2026-02-24
- Estimated Expiration
- 2044-12-27
AI Technical Summary
In existing technologies, customer identification methods in the financial services industry lack generalization capabilities and cannot effectively determine whether an account is abnormal. In particular, when anti-collection companies intervene, traditional biometric technologies are easily fraudulent, leading to inaccurate judgments of account abnormalities.
By acquiring connection data and voice data during voice calls, analyzing voice feature matching degree and connection method, combining account tag information, dynamically setting matching degree threshold, and integrating the analysis of connection data and voice data, abnormal accounts can be identified.
It improves the accuracy and generalization ability of account anomaly detection, avoids the limitations of judging from a single data source, and can adapt to different call conditions and speaker characteristics, effectively identifying account anomalies.
Smart Images

Figure CN119832934B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of computer technology, and more specifically, to a method and apparatus for determining account anomalies, a storage medium, and an electronic device. Background Technology
[0002] In financial services, especially the lending industry, verifying customer identity is a crucial step in ensuring business compliance and operational security. When customers fail to repay on time, financial institutions typically contact them by phone to remind them and negotiate repayment. However, in recent years, with the rise of anti-collection companies, the lending industry faces a new challenge: customers may entrust agents within these companies to answer collection calls. These agents often use sophisticated strategies and persuasive language to evade repayment obligations or demand more favorable repayment terms. In such cases, traditional methods of strong identity verification usually involve collecting and verifying specific biometric information of the customer, such as voiceprints, fingerprints, or facial recognition, to determine whether an account is abnormal by confirming a specific personal identity. This process often requires customers to actively register before using the service. With the development of technologies such as AI and speech synthesis, fraud and impersonation techniques are also constantly evolving. Strong identity verification methods lack generalization capabilities and may fail to effectively determine whether an account is abnormal. Summary of the Invention
[0003] This application provides a method, apparatus, storage medium, and electronic device for determining account abnormalities, addressing the problem that at least in related technologies, it is impossible to effectively determine whether an account is abnormal.
[0004] According to one embodiment of this application, a method for determining account abnormality is provided, comprising: during a voice call with a voice account of a target account, acquiring connection data and voice data, wherein the connection data is data generated during the connection of the voice account, and the voice data is sound data emitted by a first object through the voice account when the voice account is connected; if the account tag of the target account is a first tag, acquiring target historical voice data of the target account, wherein the target historical voice data is sound data emitted by a historical object through the voice account when the voice account is connected in a historical time period; and determining the target account as an abnormal account if the voice feature matching degree between the target historical voice data and the voice data is less than a preset threshold, and the connection data is abnormal.
[0005] In one exemplary embodiment, during a voice call with a target account's voice account, acquiring connection data and voice data includes: during the voice call with the voice account, detecting the metadata and voice quality of the voice call, wherein the metadata includes the connection method used during the voice call; if the connection method is determined from the metadata to be a voice transfer method, determining the metadata as the connection data, wherein the voice transfer method includes transferring the voice call to another recipient; and during the voice call with the voice account, extracting voice data whose voice quality is greater than a preset quality threshold from the generated voice data.
[0006] In one exemplary embodiment, before obtaining the historical voice data of the target account when the account tag of the target account is a first tag, the method further includes: searching for the account tag of the target account from the tag database according to the account information of the target account; determining that the target account is an account in a normal state when the account tag of the target account is a first tag; and determining that the target account is an account in an abnormal state when the account tag of the target account is a second tag; wherein the target account is an account that conducts financial transactions with other accounts.
[0007] In an exemplary embodiment, when the account tag of the target account is a first tag, obtaining the historical voice data of the target account includes: obtaining N historical voice data of the target account from the voiceprint database according to the account information of the target account, wherein N is a natural number greater than or equal to 1; selecting historical voice data with a voice quality greater than a preset quality threshold from the N historical voice data to obtain the target historical voice data.
[0008] In an exemplary embodiment, determining the target account as an abnormal account when the voice feature matching degree between the target historical voice data and the voice data is less than a preset threshold, and the connection data is abnormal, includes: extracting voice features from the target historical voice data and the voice features from the voice data to obtain a first target voice feature and a second target voice feature; calculating the matching degree between the first target voice feature and the second target voice feature to obtain the voice feature matching degree; and determining the target account as an abnormal account when abnormal data is parsed from the connection data and the voice feature matching degree is less than a preset threshold, wherein the connection method used during the voice call included in the abnormal data is a voice transfer method.
[0009] In one exemplary embodiment, extracting speech features from the target historical speech data and the speech data to obtain a first target speech feature and a second target speech feature includes: performing semantic filtering operations on the target historical speech data and the speech data respectively to remove background speech data from the target historical speech data and the speech data respectively, to obtain first speech data and second speech data; performing standardization operations on the first speech data and the second speech data respectively to unify the audio reference of the first speech data and the second speech data respectively, to obtain third speech data and fourth speech data; extracting speech features from the third speech data and the fourth speech data in segments respectively, to obtain multiple first speech features corresponding to the third speech data and multiple second speech features corresponding to the fourth speech data; fusing the multiple first speech features to obtain the first target speech feature; and fusing the multiple second speech features to obtain the second target speech feature.
[0010] In an exemplary embodiment, after determining that the voice feature matching degree between the target historical voice data and the voice data is less than a preset threshold, and that the connection data is abnormal, the method further includes: setting the account tag of the target account as a second tag in the tag database according to the account information of the target account, so as to mark the target account as an abnormal account; and triggering the processing operation of the abnormal account based on the second tag.
[0011] According to one embodiment of this application, an account anomaly determination device includes: a first acquisition module, configured to acquire connection data and voice data during a voice call with a target account's voice account, wherein the connection data is data generated during the connection of the voice account, and the voice data is sound data emitted by a first object through the voice account when the voice account is connected; a second acquisition module, configured to acquire target historical voice data of the target account when the target account's account tag is a first tag, wherein the target historical voice data is sound data emitted by a historical object through the voice account when the voice account is connected during a historical time period; and a first determination module, configured to determine the target account as an abnormal account when the voice feature matching degree between the target historical voice data and the voice data is less than a preset threshold, and the connection data is abnormal.
[0012] In an exemplary embodiment, the first acquisition module includes: a first detection submodule, configured to detect metadata and voice quality of the voice call during the voice call with the voice account, wherein the metadata includes the connection method used during the voice call; a first determination submodule, configured to determine the metadata as the connection data when the connection method is determined to be a voice transfer method from the metadata, wherein the voice transfer method includes transferring the voice call to another object for answering; and a first interception submodule, configured to intercept the voice with a voice quality greater than a preset quality threshold from the generated voice during the voice call with the voice account, thereby obtaining the voice data.
[0013] In one exemplary embodiment, the apparatus further includes: a first search module, configured to search for the account tag of the target account from a tag database according to the account information of the target account before acquiring the historical voice data of the target account if the account tag of the target account is a first tag; a second determination module, configured to determine that the target account is an account in a normal state if the account tag of the target account is a first tag; and a third determination module, configured to determine that the target account is an account in an abnormal state if the account tag of the target account is a second tag; wherein the target account is an account that conducts financial transactions with other accounts.
[0014] In an exemplary embodiment, the first acquisition module includes: a first acquisition submodule, configured to acquire N historical voice data of the target account from the voiceprint database according to the account information of the target account, wherein N is a natural number greater than or equal to 1; and a first selection submodule, configured to select historical voice data with voice quality greater than a preset quality threshold from the N historical voice data to obtain the target historical voice data.
[0015] In an exemplary embodiment, the first determining module includes: a first extraction submodule, configured to extract the speech features of the target historical speech data and the speech features of the speech data respectively, to obtain a first target speech feature and a second target speech feature; a first calculation submodule, configured to calculate the matching degree between the first target speech feature and the second target speech feature, to obtain the speech feature matching degree; and a second determining submodule, configured to determine the target account as the abnormal account when abnormal data is parsed from the connection data and the speech feature matching degree is less than a preset threshold, wherein the connection method used during the voice call included in the abnormal data is a voice transfer method.
[0016] In an exemplary embodiment, the first extraction submodule includes: a first removal unit, configured to perform semantic filtering operations on the target historical speech data and the speech data respectively, to remove background speech data from the target historical speech data and the speech data respectively, to obtain first speech data and second speech data; a first unification unit, configured to perform standardization operations on the first speech data and the second speech data respectively, to unify the audio reference of the first speech data and the second speech data respectively, to obtain third speech data and fourth speech data; a first extraction unit, configured to extract speech features from the third speech data and the fourth speech data in segments respectively, to obtain multiple first speech features corresponding to the third speech data and multiple second speech features corresponding to the fourth speech data; a first fusion unit, configured to fuse multiple first speech features to obtain the first target speech feature; and a second fusion unit, configured to fuse multiple second speech features to obtain the second target speech feature.
[0017] In an exemplary embodiment, the apparatus further includes: a first tagging module, configured to, after determining that the target account is an abnormal account when the voice feature matching degree between the target historical voice data and the voice data is less than a preset threshold and the connection data is abnormal, set the account tag of the target account in the tag database according to the account information of the target account as a second tag to mark the target account as an abnormal account; and trigger a processing operation for the abnormal account based on the second tag.
[0018] According to another embodiment of this application, a computer program product is also provided, including a computer program configured to have a processor perform the steps in any of the above method embodiments.
[0019] According to yet another embodiment of this application, a computer-readable storage medium is also provided, wherein a computer program is stored in the computer-readable storage medium, and the computer program is configured to execute the steps in any of the above method embodiments by a processor.
[0020] According to yet another embodiment of this application, an electronic device is also provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor is configured to execute the computer program to perform the steps in any of the above method embodiments.
[0021] This application, during a voice call with a target account's voice account, acquires connection data and voice data sent and generated by the first entity through the voice account. Using the target account's historical voice data, if the voice feature matching degree between the historical voice data and the voice data is less than a preset threshold, and the connection data shows anomalies, the target account is determined to be an abnormal account. Because this application, through the analysis of fused connection data and voice data, dynamically sets the matching degree threshold, compares historical voice data, and utilizes account tag information, it can not only effectively detect account anomalies, avoiding the limitations of judging from a single data source, but also adapt to different call conditions and speaker characteristics, improving the accuracy and generalization ability of detection. Therefore, it solves the problem in related technologies of being unable to effectively determine whether an account is abnormal, achieving the effect of effectively determining account anomalies. Attached Figure Description
[0022] Figure 1 This is a hardware structure block diagram of a mobile terminal for a method of determining account abnormalities according to an embodiment of this application.
[0023] Figure 2 This is a flowchart of a method for determining account anomalies according to an embodiment of this application;
[0024] Figure 3 This is a schematic diagram of a method for determining account abnormality according to a specific embodiment of this application;
[0025] Figure 4 This is a flowchart illustrating a method for determining account anomalies according to a specific embodiment of this application;
[0026] Figure 5 This is a structural block diagram of an account anomaly determination device according to an embodiment of this application. Detailed Implementation
[0027] The embodiments of this application will be described in detail below with reference to the accompanying drawings and examples.
[0028] It should be noted that the terms "first," "second," etc., in the specification, claims, and drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence.
[0029] The methods and embodiments provided in this application can be executed on a mobile terminal, computer terminal, or similar computing device. Taking running on a mobile terminal as an example, Figure 1 This is a hardware structure block diagram of a mobile terminal for a method of determining account anomalies according to an embodiment of this application. Figure 1 As shown, a mobile terminal may include one or more ( Figure 1Only one is shown in the diagram. A processor 102 (which may include, but is not limited to, a microprocessor MCU or a programmable logic device FPGA, etc.) and a memory 104 for storing data are also shown. The mobile terminal may further include a transmission device 106 for communication functions and an input / output device 108. Those skilled in the art will understand that... Figure 1 The structure shown is for illustrative purposes only and does not limit the structure of the mobile terminal described above. For example, the mobile terminal may also include components that are more... Figure 1 The more or fewer components shown, or having the same Figure 1 The different configurations shown.
[0030] The memory 104 can be used to store computer programs, such as application software programs and modules, like the computer program corresponding to a method for determining account anomalies in this embodiment. The processor 102 executes various functional applications and data processing by running the computer program stored in the memory 104, thus implementing the aforementioned method. The memory 104 may include high-speed random access memory and non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some instances, the memory 104 may further include memory remotely located relative to the processor 102, and these remote memories can be connected to the mobile terminal via a network. Examples of such networks include, but are not limited to, the Internet, corporate intranets, local area networks, mobile communication networks, and combinations thereof.
[0031] The transmission device 106 is used to receive or send data via a network. Specific examples of the network described above may include a wireless network provided by the mobile terminal's communication provider. In one example, the transmission device 106 includes a Network Interface Controller (NIC), which can connect to other network devices via a base station to communicate with the Internet. In another example, the transmission device 106 may be a Radio Frequency (RF) module, used for wireless communication with the Internet.
[0032] This embodiment provides a method for determining account anomalies. Figure 2 This is a flowchart of a method for determining account anomalies according to an embodiment of this application, such as... Figure 2 As shown, the process includes the following steps:
[0033] Step S202: During a voice call with the target account's voice account, connection data and voice data are obtained. Connection data is data generated during the process of connecting to the voice account, and voice data is the sound data emitted by the first object through the voice account when the voice account is connected.
[0034] Optionally, the voice call to the target account's voice account may be made to human customer service representatives or intelligent customer service representatives from financial institutions, including but not limited to those of financial institutions.
[0035] Optionally, connection data is used to detect whether a call transfer has occurred.
[0036] Optionally, the target account is an account at the opening bank. The voice account is the phone number linked to the target account. The first object can be either an object linked to the target account or an object not linked to the target account.
[0037] Optionally, other voice data of other objects are collected to generate an anti-collection database, wherein other objects are anti-collection personnel; if the voice feature matching degree between voice data and other voice data is less than a preset threshold and there are abnormalities in the connection data, the target account is identified as an abnormal account.
[0038] Step S204: If the account tag of the target account is the first tag, obtain the target historical voice data of the target account. The target historical voice data is the voice data emitted by the historical object through the voice account when the voice account is connected in a historical time period.
[0039] Optionally, the first label is used to indicate that the target account is not an abnormal account.
[0040] Step S206: If the voice feature matching degree between the target historical voice data and the voice data is less than a preset threshold, and there is an anomaly in the connection data, the target account is determined to be an abnormal account.
[0041] Optionally, an abnormal account indicates that multiple objects have made voice calls with staff through the voice account bound to the target account.
[0042] In this embodiment, the entity executing the above steps can be a terminal, a server, a specific processor set in the terminal or server, or a processor or processing device set relatively independently of the terminal or server, but is not limited thereto.
[0043] Optionally, such as Figure 3 As shown, Figure 3This is a schematic diagram illustrating a method for determining account anomalies according to a specific embodiment of this application. The communication platform module provides basic call services and collects voice data. The user center module, based on voice-related interfaces, obtains connection data and voice data from the communication platform module through the business application module. It then calls the abnormal voice recognition and basic voice services in the voice service module to determine the voice feature matching degree between the target's historical voice data and the current voice data, and to detect connection data, thereby determining the anomaly of the target account and tagging it. The business application module, based on voice-related interfaces, displays the identified target account tags to customer service for viewing from the user center module.
[0044] Through the above steps, during a voice call with the target account's voice account, the connection data and voice data sent and generated by the first object through the voice account are obtained. Using the target account's historical voice data, if the voice feature matching degree between the historical voice data and the voice data is less than a preset threshold, and the connection data shows anomalies, the target account is determined to be an abnormal account. Because this application, through the analysis of fused connection data and voice data, dynamically sets the matching degree threshold, compares historical voice data, and utilizes account tag information, it can not only effectively detect account anomalies and avoid the limitations of judging from a single data source, but also adapt to different call conditions and speaker characteristics, improving the accuracy and generalization ability of detection. Therefore, it solves the problem in related technologies of being unable to effectively determine whether an account is abnormal, achieving the effect of effectively determining account anomalies.
[0045] In one exemplary embodiment, during a voice call with a target account's voice account, acquiring connection data and voice data includes: during the voice call with the voice account, detecting the metadata and voice quality of the voice call, wherein the metadata includes the connection method used during the voice call; if the connection method is determined from the metadata to be a voice transfer method, determining the metadata as the connection data, wherein the voice transfer method includes transferring the voice call to another recipient; and during the voice call with the voice account, extracting voice data whose voice quality is greater than a preset quality threshold from the generated voice data.
[0046] Optionally, metadata includes, but is not limited to, information such as voice call time, duration, connection method, and device type. The connection method is used to determine whether the voice call was transferred to a different recipient than the original caller; this may include situations where the voice account is transferred to employees of an anti-collection company. Metrics for evaluating voice quality include, but are not limited to, signal strength, clarity, and volume level. Only when the voice quality exceeds a preset quality threshold will the voice segment be captured and saved, ensuring sufficient clarity and integrity of the voice data. The preset quality threshold can be dynamically adjusted based on application or voice scenarios; for example, if the environment is noisy or the voice is unclear, the preset quality threshold can be lowered.
[0047] Optionally, when a consumer finance company's collection department initiates a voice call with the target account's voice account, call metadata is automatically collected the moment the call is established, including call time, duration, connection method, and device type. For example, a call might begin at 2:30 PM on March 15, 2023, last for 10 minutes, be directly connected, and be made using a smartphone. During the call, the connection method is monitored in real-time to confirm whether it was a voice transfer. If the call is transferred to a different recipient than the original caller, such as an employee of an anti-collection company, the system will mark this transfer and identify the relevant metadata (including transfer time, device information before and after the transfer, etc.) as connection data. Voice quality is continuously evaluated during the call, checking indicators such as signal strength, clarity, and volume level. For example, if the preset voice quality threshold is a signal-to-noise ratio (SNR) greater than 70 dB, and the detected SNR is 75 dB after the call starts, and the SNR stabilizes above 75 dB 5 seconds after the call starts, then the voice for the next 10 seconds is extracted, i.e. from the 5th to the 15th second. After extracting 10 seconds of high-quality voice segments, they are identified as voice data.
[0048] This embodiment detects metadata and voice quality during voice calls, enabling it to quickly identify abnormal situations where an account might be answered by someone other than the account holder, thus achieving accurate and reliable detection. Simultaneously, it extracts data from voices with quality exceeding a preset threshold, avoiding the collection of invalid data in noisy environments or when device audio quality is poor at the beginning of a call. This ensures high-quality voice data and reduces recognition errors caused by environmental noise or device malfunctions.
[0049] In one exemplary embodiment, before obtaining the historical voice data of the target account when the account tag of the target account is a first tag, the method further includes: searching for the account tag of the target account from the tag database according to the account information of the target account; determining that the target account is an account in a normal state when the account tag of the target account is a first tag; and determining that the target account is an account in an abnormal state when the account tag of the target account is a second tag; wherein the target account is an account that conducts financial transactions with other accounts.
[0050] Optionally, the target account's information includes, but is not limited to, account identifier and voice account. If the target account is determined to be in a normal state, its historical voice data will be acquired and saved during subsequent voice calls for use in voiceprint recognition and matching analysis. For example, all call records of the target account within the past three months will be acquired, and high-clarity, high-signal-strength voice segments from each call will be extracted and saved to obtain historical voice data. Conversely, if the target account's account tag is a "second tag," it will be immediately determined that the target account is in an abnormal state.
[0051] This embodiment can quickly determine whether the target account is in a normal state by querying the account tag of the target account in advance, thereby deciding whether to obtain historical voice data for voiceprint comparison. For normal accounts marked as "first tag", the historical voice data retrieval process can be automatically entered, realizing the goal of quickly completing the feature matching verification of voice features and improving the efficiency of identifying the target account as an abnormal account.
[0052] In an exemplary embodiment, when the account tag of the target account is a first tag, obtaining the target historical voice data of the target account includes: obtaining N historical voice data of the target account from the voiceprint database according to the account information of the target account, wherein N is a natural number greater than or equal to 1; selecting historical voice data with a voice quality greater than a preset quality threshold from the N historical voice data to obtain the target historical voice data.
[0053] Optionally, a classification operation is performed on N historical speech data to obtain M historical classified speech data, where M is a natural number greater than 1; from the M historical classified speech data, historical classified speech data with speech quality greater than a preset quality threshold are selected to obtain the above-mentioned target historical speech data.
[0054] Optionally, before conducting a voice call with a target account, if the target account's account tag is marked as "first tag" (i.e., the account is currently in normal status), the collection system of a consumer finance company will perform the following steps to obtain and filter the target account's historical voice data for subsequent feature matching: First, the collection system retrieves all historical voice data of the target account from the voiceprint database within a certain period of time, based on the target account's account information (such as account identifier). For example, the system sets the retrieval time range to the past 6 months, and the goal of each retrieval is to obtain N historical voice data related to the account, where N is a natural number greater than or equal to 1, such as N=10. Then, the collection system uses a built-in voice quality assessment module to perform quality checks on the retrieved N historical voice data, evaluating key indicators such as signal strength, clarity, and volume level. From the N historical voice data, all voice data with a signal-to-noise ratio greater than 60dB are selected. These data are considered as target historical voice data, i.e., high-quality voice segments used for subsequent voiceprint comparison. For example, when N=10, if the signal-to-noise ratio of 7 historical speech data points meets the condition after screening, these 7 speech data points will be identified as high-quality target historical speech data.
[0055] This embodiment ensures that the data used for feature matching is clear, complete, and free of significant noise by selecting only historical speech data with a quality greater than a preset quality threshold, thereby reducing recognition errors caused by poor speech quality. Simultaneously, it avoids indiscriminate processing of all historical data, reducing unnecessary computational resource consumption and improving data processing efficiency. This is particularly effective in scenarios involving large amounts of historical speech data, significantly saving computation time and storage space.
[0056] In an exemplary embodiment, determining the target account as an abnormal account when the voice feature matching degree between the target historical voice data and the voice data is less than a preset threshold, and the connection data is abnormal, includes: extracting voice features from the target historical voice data and the voice features from the voice data to obtain a first target voice feature and a second target voice feature; calculating the matching degree between the first target voice feature and the second target voice feature to obtain the voice feature matching degree; and determining the target account as an abnormal account when abnormal data is parsed from the connection data and the voice feature matching degree is less than a preset threshold, wherein the connection method used during the voice call included in the abnormal data is a voice transfer method.
[0057] Optionally, when determining whether a target account is an abnormal account, the collection department of a consumer finance company will analyze the matching degree of voice features and the abnormality of the call connection data. The specific process is as follows: First, extract the voice features of the target's historical voice data and the voice features of the current voice call data. Voice features include, but are not limited to, voiceprint features, speech rate features, tone features, and the frequency of use of specific words. Then, use artificial intelligence algorithms, such as deep neural networks or recurrent neural networks, to analyze the voice data and extract the first target voice features and the second target voice features. Next, use a feature matching algorithm to calculate the matching degree between the first and second target voice features to obtain the voice feature matching degree. The voice feature matching degree can be the Euclidean distance, cosine similarity, or other applicable metrics between two feature vectors. The system will set a preset threshold, such as 0.7 or 80%, to distinguish the voiceprint matching degree. If the voice feature matching degree is less than the preset threshold, the system initially determines that the voiceprint of the current call is inconsistent with the voiceprint features in the historical voice data.
[0058] This embodiment effectively identifies whether a target account has multiple users by comparing the voice feature matching degree between the target's historical voice data and the current call voice data, combined with the metadata anomaly detection of the connection data.
[0059] In one exemplary embodiment, extracting speech features from the target historical speech data and the speech data to obtain a first target speech feature and a second target speech feature includes: performing semantic filtering operations on the target historical speech data and the speech data respectively to remove background speech data from the target historical speech data and the speech data respectively, to obtain first speech data and second speech data; performing standardization operations on the first speech data and the second speech data respectively to unify the audio reference of the first speech data and the second speech data respectively, to obtain third speech data and fourth speech data; extracting speech features from the third speech data and the fourth speech data in segments respectively, to obtain multiple first speech features corresponding to the third speech data and multiple second speech features corresponding to the fourth speech data; fusing the multiple first speech features to obtain the first target speech feature; and fusing the multiple second speech features to obtain the second target speech feature.
[0060] Optionally, to improve the accuracy of voiceprint recognition, the collection system of a consumer finance company employs the following steps to extract voice features from target historical voice data and real-time call voice data, respectively, to obtain first target voice features and second target voice features: First, semantic filtering is performed on the target historical voice data, using audio processing algorithms (such as noise thresholding and spectrum analysis) to identify and remove background noise, such as ambient sounds and other people's voices, ensuring that the extracted voice features originate only from the first target. Similarly, semantic filtering is also performed on the voice data in the real-time call to remove background noise, resulting in first voice data (processed target historical voice data) and second voice data (processed real-time call voice data). For example, if a call record in the target historical voice data contains noisy background noise from a coffee shop, the system uses a spectrum analysis algorithm to identify and filter out the background noise, retaining only the clear human voice. Then, standardization is performed on the first and second voice data, preprocessing the audio data, such as adjusting volume and frequency response, to unify the audio benchmark and ensure that voiceprint feature comparisons between different call records are conducted under the same conditions. After processing, third and fourth voice data (standardized target historical voice data) are obtained. For example, the volume in the target historical voice data may be lower than the volume in the real-time call. Through standardization, the volumes of the two are adjusted to the same level to avoid matching errors caused by volume differences. Then, the third and fourth voice data are segmented, with each segment divided into multiple fixed-length segments, such as 10-second segments. A deep learning model is then used to extract voice features from each segment, resulting in multiple first voice features corresponding to the third voice data and multiple second voice features corresponding to the fourth voice data. For example, five 10-second segments are extracted from the target historical voice data. After processing by the voiceprint recognition model, the corresponding first voice features are obtained for each segment. Finally, the multiple first voice features are fused using averaging, weighted averaging, or other feature fusion algorithms to obtain the first target voice feature. Similarly, the multiple second voice features are fused to obtain the second target voice feature, ensuring that the features extracted from different voice segments can comprehensively reflect the voiceprint features of the first object, thus improving the stability and accuracy of recognition.
[0061] In an exemplary embodiment, after determining that the voice feature matching degree between the target historical voice data and the voice data is less than a preset threshold, and that the connection data is abnormal, the method further includes: setting the account tag of the target account as a second tag in the tag database according to the account information of the target account, so as to mark the target account as an abnormal account; and triggering the processing operation of the abnormal account based on the second tag.
[0062] Optionally, once a target account is marked as an abnormal account, risk can be managed through abnormal account handling procedures. For example, for human agents, the system will display the abnormal account marker in real time, allowing agents to adjust their call strategies, increase vigilance, and implement additional verification steps. For intelligent customer service systems, the system will automatically switch to a preset abnormal account handling process, which may include multiple verifications via voice recognition, requiring customers to provide static identity verification (such as passwords, ID photos, etc.), or transferring the call to a human agent for manual review to ensure the security of transactions or communication. Optionally, detailed information about this voice call, including voice characteristics, connection method, and abnormal account marker, can be recorded in the target account's history. This information can be used for subsequent risk analysis and early warning, helping financial institutions build more comprehensive risk management models and prevent future fraudulent activities.
[0063] The present invention will now be described in conjunction with specific embodiments:
[0064] This embodiment illustrates the confirmation of a target account in a scenario involving human or intelligent customer service. Figure 4 This is a flowchart of a method for determining account anomalies according to a specific embodiment of this application, including the following steps:
[0065] S402, Start: The first target calls in or the financial institution's human or intelligent customer service calls into the customer service workbench page. The first target conducts a voice call with human or intelligent customer service through the voice account bound to the target account.
[0066] S404, The user center module feeds back account information and account tags to the business application module, including but not limited to: account identifier, voice account, and primary tag;
[0067] S406, The customer service console displays the target account as a normal account to the human customer service representative;
[0068] S408: After a voice call is connected, the first person uses their voice account to make a voice call with a human or intelligent customer service representative.
[0069] S410, the communication platform module calls the call extraction interface. If a voice call is transferred to a different recipient than the original caller, such as an employee of an anti-collection company, this transfer is marked, and the relevant metadata (including transfer time, device information before and after the transfer, etc.) is identified as connection data. During the voice call, voice quality is continuously evaluated, checking indicators such as signal strength, clarity, and volume level. For example, if the preset voice quality threshold is a signal-to-noise ratio (SNR) greater than 70 dB, and the detected SNR is 75 dB after the call starts, and remains above 75 dB for 5 seconds, then the next 10 seconds of voice are extracted (from the 5th to the 15th second). This 10-second high-quality voice segment is identified as voice data. The connection data and voice data are then saved to a recording file in a specified path in the storage module.
[0070] S412, the business application module retrieves the connection data and voice data from the storage module based on the path and filename of the recording file returned by the communication platform module.
[0071] S414, The business application module sends the target account's voice account, connection data, and voice data to the user center module.
[0072] S416, The user center module retrieves the first object identifier from the database based on the voice account;
[0073] S418, The user center module retrieves the voiceprint registration information table of the target account from the voiceprint database based on the voice account and the first object identifier;
[0074] S420: Based on the voiceprint registration information table, determine whether there is historical voice data of the target account in the voiceprint database. If yes, proceed to S422; otherwise, proceed to S434 to register this voice call.
[0075] S422: Retrieve a list of registered voiceprint data from the voiceprint database. This list includes N historical voice data points for the target account. First, based on the target account's account information (such as account identifier), retrieve all historical voice data for the target account from the voiceprint database within the past 6 months. Then, using the built-in voice quality assessment module, perform quality checks on the retrieved 10 historical voice data points, evaluating key indicators such as signal strength, clarity, and volume level. Select all voice data points with a signal-to-noise ratio (SNR) greater than 60 dB from these 10 historical voice data points. These are considered the target historical voice data, i.e., high-quality voice segments used for subsequent voiceprint comparison. For example, when N=10, if the SNR of 7 historical voice data points meets the requirement after filtering, these 7 voice data points will be identified as high-quality target historical voice data.
[0076] S424, determine the voice feature matching degree between the target historical voice data and the voice data, and whether there are any anomalies in the call connection data. First, perform semantic filtering on the target historical voice data, using audio processing algorithms (such as noise thresholding, spectrum analysis) to identify and remove background noise, such as ambient sounds and other people's voices, ensuring that the extracted voice features come only from the first object. Similarly, perform semantic filtering on the voice data in the real-time call to remove background noise, obtaining the first voice data (processed target historical voice data) and the second voice data (processed real-time call voice data). For example, if a call record in the target historical voice data contains noisy background sounds from a coffee shop, use a spectrum analysis algorithm to identify and filter out the background noise, retaining only the clear human voice. Then, perform standardization on the first and second voice data, preprocessing the audio data, such as adjusting volume and frequency response, to unify the audio benchmark and ensure that the voiceprint feature comparison between different call records is performed under the same conditions. After processing, obtain the third voice data (standardized target historical voice data) and the fourth voice data (standardized real-time call voice data). For example, the volume in the target historical speech data might be lower than the volume of the real-time call. Standardization is used to adjust the volumes of both to the same level to avoid matching errors caused by volume differences. Then, the third and fourth speech data are segmented, with each segment divided into multiple fixed-length fragments, such as 10-second segments. A deep learning model is then used to extract speech features from each fragment, resulting in multiple first speech features corresponding to the third speech data and multiple second speech features corresponding to the fourth speech data. For example, five 10-second segments are extracted from the target historical speech data; each segment is processed by a voiceprint recognition model to obtain corresponding first speech features. Finally, the multiple first speech features are fused using averaging, weighted averaging, or other feature fusion algorithms to obtain the first target speech feature. Similarly, the multiple second speech features are fused to obtain the second target speech feature. A feature matching algorithm is used to calculate the similarity between the first and second target speech features to obtain the speech feature matching degree. The speech feature matching degree can be the Euclidean distance, cosine similarity, or other applicable metrics between two feature vectors. If the voice feature matching degree is less than the preset threshold, it is initially determined that the voiceprint of the current call is inconsistent with the voiceprint features in the historical voice data.
[0077] S426, Save the calculated speech feature matching degree;
[0078] S428: If the voice feature matching degree between the target historical voice data and the voice data is less than a preset threshold, and there are abnormalities in the connection data, the target account is determined to be an abnormal account, indicating that multiple people share the same account. The abnormality of the target account is reported to the business application module.
[0079] S430: The customer service console displays the abnormal situation of the target account to the human customer service representative, who can then adjust the call strategy, increase vigilance, and perform additional verification steps accordingly.
[0080] S432, save the abnormal situation of the target account to the record of this voice call. The fields in the record include, but are not limited to: the identification information of the target account, the voice account, the voice call record number, the voice feature matching degree, the abnormal label of the target account, and the validity period of the abnormal label.
[0081] S434, the user center module determines that there is no historical voice data for the target account in the voiceprint database based on the voiceprint registration information table, and then sends a voiceprint not registered prompt to the human customer service through the business application module, and then processes the registration process asynchronously;
[0082] S436, The user center module calls the voiceprint service interface to trigger the voiceprint registration of the voice service module;
[0083] S438, the voice service module uses the user center voiceprint identification interface to obtain the target account's voice account, connection data, and voice data for voiceprint registration;
[0084] S440, Registration successful. The user center module saves the voiceprint registration information to the voiceprint registration information table of the target account. The fields in the voiceprint registration table include, but are not limited to: the target account's identification information, voice account, voice call record number, the file name of the recording file, the path of the recording file, and the first object identifier.
[0085] S442, End.
[0086] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods according to the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions to cause a terminal device (which may be a mobile phone, computer, server, or network device, etc.) to execute the methods of the various embodiments of this application.
[0087] This embodiment also provides an account anomaly determination device, which is used to implement the above embodiments and preferred embodiments, and will not be repeated as already described. As used below, the term "module" can be a combination of software and / or hardware that implements a predetermined function. Although the device described in the following embodiments is preferably implemented in software, hardware implementation, or a combination of software and hardware, is also possible and contemplated.
[0088] Figure 5 This is a structural block diagram of an account anomaly determination device according to an embodiment of this application, such as... Figure 5 As shown, the device includes:
[0089] The first acquisition module 502 is used to acquire connection data and voice data during a voice call with the voice account of the target account. The connection data is data generated during the connection of the voice account, and the voice data is the sound data emitted by the first object through the voice account when the voice account is connected.
[0090] The second acquisition module 504 is used to acquire the target historical voice data of the target account when the account tag of the target account is the first tag. The target historical voice data is the voice data emitted by the historical object through the voice account when the voice account was connected in a historical time period.
[0091] The first determining module 506 is used to determine the target account as an abnormal account when the voice feature matching degree between the target historical voice data and the voice data is less than a preset threshold and the connection data is abnormal.
[0092] In an exemplary embodiment, the first acquisition module includes: a first detection submodule, configured to detect metadata and voice quality of the voice call during the voice call with the voice account, wherein the metadata includes the connection method used during the voice call; a first determination submodule, configured to determine the metadata as the connection data when the connection method is determined to be a voice transfer method from the metadata, wherein the voice transfer method includes transferring the voice call to another object for answering; and a first interception submodule, configured to intercept the voice with a voice quality greater than a preset quality threshold from the generated voice during the voice call with the voice account, thereby obtaining the voice data.
[0093] In one exemplary embodiment, the apparatus further includes: a first search module, configured to search for the account tag of the target account from a tag database according to the account information of the target account before acquiring the historical voice data of the target account if the account tag of the target account is a first tag; a second determination module, configured to determine that the target account is an account in a normal state if the account tag of the target account is a first tag; and a third determination module, configured to determine that the target account is an account in an abnormal state if the account tag of the target account is a second tag; wherein the target account is an account that conducts financial transactions with other accounts.
[0094] In an exemplary embodiment, the first acquisition module includes: a first acquisition submodule, configured to acquire N historical voice data of the target account from the voiceprint database according to the account information of the target account, wherein N is a natural number greater than or equal to 1; and a first selection submodule, configured to select historical voice data with voice quality greater than a preset quality threshold from the N historical voice data to obtain the target historical voice data.
[0095] In an exemplary embodiment, the first determining module includes: a first extraction submodule, configured to extract the speech features of the target historical speech data and the speech features of the speech data respectively, to obtain a first target speech feature and a second target speech feature; a first calculation submodule, configured to calculate the matching degree between the first target speech feature and the second target speech feature, to obtain the speech feature matching degree; and a second determining submodule, configured to determine the target account as the abnormal account when abnormal data is parsed from the connection data and the speech feature matching degree is less than a preset threshold, wherein the connection method used during the voice call included in the abnormal data is a voice transfer method.
[0096] In an exemplary embodiment, the first extraction submodule includes: a first removal unit, configured to perform semantic filtering operations on the target historical speech data and the speech data respectively, to remove background speech data from the target historical speech data and the speech data respectively, to obtain first speech data and second speech data; a first unification unit, configured to perform standardization operations on the first speech data and the second speech data respectively, to unify the audio reference of the first speech data and the second speech data respectively, to obtain third speech data and fourth speech data; a first extraction unit, configured to extract speech features from the third speech data and the fourth speech data in segments respectively, to obtain multiple first speech features corresponding to the third speech data and multiple second speech features corresponding to the fourth speech data; a first fusion unit, configured to fuse multiple first speech features to obtain the first target speech feature; and a second fusion unit, configured to fuse multiple second speech features to obtain the second target speech feature.
[0097] In an exemplary embodiment, the apparatus further includes: a first tagging module, configured to, after determining that the target account is an abnormal account when the voice feature matching degree between the target historical voice data and the voice data is less than a preset threshold and the connection data is abnormal, set the account tag of the target account in the tag database according to the account information of the target account as a second tag to mark the target account as an abnormal account; and trigger a processing operation for the abnormal account based on the second tag.
[0098] Embodiments of this application provide a computer program product, including a computer program that, when executed by a processor, implements the steps in any of the above method embodiments.
[0099] Embodiments of this application also provide a computer-readable storage medium storing a computer program, wherein the computer program is configured to execute the steps in any of the above method embodiments when run.
[0100] In one exemplary embodiment, the aforementioned computer-readable storage medium may include, but is not limited to, various media capable of storing computer programs, such as a USB flash drive, read-only memory (ROM), random access memory (RAM), portable hard drive, magnetic disk, or optical disk.
[0101] Embodiments of this application also provide an electronic device, including a memory and a processor, wherein the memory stores a computer program and the processor is configured to run the computer program to perform the steps in any of the above method embodiments.
[0102] In one exemplary embodiment, the electronic device may further include a transmission device and an input / output device, wherein the transmission device is connected to the processor and the input / output device is connected to the processor.
[0103] Specific examples in this embodiment can be found in the examples described in the above embodiments and exemplary implementations, and will not be repeated here.
[0104] Obviously, those skilled in the art should understand that the modules or steps of this application described above can be implemented using general-purpose computing devices. They can be centralized on a single computing device or distributed across a network of multiple computing devices. They can be implemented using computer-executable program code, and thus can be stored in a storage device for execution by a computing device. In some cases, the steps shown or described can be performed in a different order than those presented here, or they can be fabricated as separate integrated circuit modules, or multiple modules or steps can be fabricated as a single integrated circuit module. Thus, this application is not limited to any particular combination of hardware and software.
[0105] The above are merely preferred embodiments of this application and are not intended to limit this application. Various modifications and variations can be made to this application by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the principles of this application should be included within the scope of protection of this application.
Claims
1. A method for determining account irregularities, characterized in that, include: During a voice call with the target account's voice account, connection data and voice data are acquired. The connection data is data generated during the process of connecting to the voice account, and the voice data is the sound data emitted by the first object through the voice account when connecting to the voice account. The connection data is used to detect whether the voice call has been transferred. When the account tag of the target account is the first tag, the target historical voice data of the target account is obtained. The target historical voice data is the voice data emitted by the target object through the voice account when the voice account was connected in a historical time period. The first tag is used to indicate that the target account is not an abnormal account. If the voice feature matching degree between the target historical voice data and the voice data is less than a preset threshold, and the connection data is abnormal, the target account is determined to be an abnormal account. The abnormal connection data indicates that the voice call has undergone the transfer operation.
2. The method according to claim 1, characterized in that, During a voice call with the target account's voice account, acquire connection data and voice data, including: During the voice call with the voice account, the metadata and voice quality of the voice call are detected, wherein the metadata includes the connection method used during the voice call; If the connection method is determined to be a voice transfer method from the metadata, the metadata is determined to be the connection data, wherein the voice transfer method includes transferring the voice call to another person for answering; During the voice call with the voice account, the voice data is obtained by extracting the voice data whose quality is greater than a preset quality threshold from the generated voice.
3. The method according to claim 1, characterized in that, If the account tag of the target account is the first tag, before obtaining the target historical voice data of the target account, the method further includes: Search the tag database for the account tag of the target account based on the account information of the target account; If the target account's account tag is the first tag, then the target account is determined to be an account in a normal state; If the target account's account tag is the second tag, then the target account is determined to be an account in an abnormal state. The target account is the account that conducts financial transactions with other accounts.
4. The method according to claim 1, characterized in that, If the account tag of the target account is the first tag, obtain the target historical voice data of the target account, including: According to the account information of the target account, N historical voice data of the target account are obtained from the voiceprint database, where N is a natural number greater than or equal to 1; The target historical voice data is obtained by selecting historical voice data with a voice quality greater than a preset quality threshold from N historical voice data.
5. The method according to claim 1, characterized in that, If the voice feature matching degree between the target historical voice data and the voice data is less than a preset threshold, and the connection data is abnormal, the target account is determined to be an abnormal account, including: The speech features of the target historical speech data and the speech features of the speech data are extracted respectively to obtain the first target speech features and the second target speech features; Calculate the matching degree between the first target speech feature and the second target speech feature to obtain the speech feature matching degree; If abnormal data is parsed from the connection data and the voice feature matching degree is less than a preset threshold, the target account is determined to be the abnormal account, wherein the connection method used during the voice call included in the abnormal data is the voice transfer method.
6. The method according to claim 5, characterized in that, The speech features of the target historical speech data and the speech features of the speech data are extracted respectively to obtain the first target speech feature and the second target speech feature, including: Semantic filtering operations are performed on the target historical speech data and the speech data respectively to remove background speech data from the target historical speech data and the speech data, respectively, to obtain the first speech data and the second speech data; Standardization operations are performed on the first speech data and the second speech data to unify the audio reference of the first speech data and the second speech data respectively, so as to obtain the third speech data and the fourth speech data. Speech features are extracted from the third speech data and the fourth speech data in segments respectively to obtain multiple first speech features corresponding to the third speech data and multiple second speech features corresponding to the fourth speech data; The first target speech feature is obtained by fusing multiple first speech features; By fusing multiple second speech features, the second target speech feature is obtained.
7. The method according to claim 1, characterized in that, After determining that the voice feature matching degree between the target historical voice data and the voice data is less than a preset threshold, and that the connection data is abnormal, the method further includes: Based on the account information of the target account, the account tag of the target account is set as the second tag in the tag database to mark the target account as an abnormal account; Based on the second tag, the processing operation for the abnormal account is triggered.
8. A device for determining account irregularities, characterized in that, include: The first acquisition module is used to acquire connection data and voice data during a voice call with the voice account of the target account. The connection data is the data generated during the process of connecting to the voice account, and the voice data is the sound data emitted by the first object through the voice account when connecting to the voice account. The connection data is used to detect whether the voice call has been transferred. The second acquisition module is used to acquire the historical voice data of the target account when the account tag of the target account is the first tag. The historical voice data is the voice data emitted by the target account through the voice account when the target account was connected in a historical time period. The first tag is used to indicate that the target account is not an abnormal account. The first determining module is used to determine the target account as an abnormal account when the voice feature matching degree between the historical voice data and the voice data is less than a preset threshold and the connection data is abnormal. The abnormal connection data indicates that the voice call has undergone the transfer operation.
9. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program, wherein the computer program, when executed by a processor, implements the steps of the method described in any one of claims 1 to 6.
10. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the steps of the method described in any one of claims 1 to 6.
Citation Information
Patent Citations
Abnormal identification method and device for user feedback information and computer equipment
CN117292712A