Target user rechecking method and device based on voiceprint portrait, equipment and medium
By combining a voiceprint recognition model with gender and age recognition models to filter similar recordings, the problems of voiceprint recognition error and high cost of manual verification are solved, achieving efficient and accurate identification of fake users.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- PING AN TECH (SHENZHEN) CO LTD
- Filing Date
- 2026-01-09
- Publication Date
- 2026-04-24
Smart Images

Figure CN121921027A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of big data technology, applied to the financial sector, and particularly to a method, apparatus, device, and medium for verifying target users based on voiceprint profiling. Background Technology
[0002] In the financial sector, fraudulent activities involving impersonation of customers are becoming increasingly rampant, posing a serious threat to user asset security and institutional risk control systems. To address this challenge, voiceprint recognition technology, due to its unique biometric characteristics, is widely used in identity verification scenarios. However, in practical applications, the accuracy of voiceprint recognition cannot reach 100%, and it is susceptible to factors such as environmental noise, channel quality, and changes in the speaker's state, resulting in a certain risk of false recognition.
[0003] Therefore, after the automated identification stage, a secondary verification stage involving manual listening is usually required to ensure the accuracy of the final judgment. This process is highly dependent on manual operation, resulting in high time and labor costs. To improve the efficiency of manual verification, existing technical solutions attempt to provide reviewers with specific labeled reference information (e.g., "male," "30 years old") through auxiliary voice attribute recognition models, such as gender recognition models and age recognition models. However, this method based on fixed and specific labels has inherent defects: the recognition error of the model itself can be directly transmitted to the reviewer, leading to misleading results; moreover, human acoustic characteristics are complex and variable, and simply classifying them into a few discrete labels lacks descriptions of more subtle features such as timbre and intonation, making it easy to misjudge due to slight deviations between the labels and the real person's characteristics, thus increasing the uncertainty of the decision. Summary of the Invention
[0004] This invention provides a method, apparatus, computer equipment, and storage medium for verifying target users based on voiceprint profiling, so as to effectively improve the accuracy of counterfeit identification and reduce the workload of manual verification.
[0005] In a first aspect, embodiments of the present invention provide a target user verification method based on voiceprint profiling, comprising: extracting voiceprint features of a target recording using a preset voiceprint recognition model; extracting multiple initial recordings similar to the target recording from a preset database based on the voiceprint features using the voiceprint recognition model, wherein the preset database stores recordings of the target user in advance; filtering the multiple initial recordings using a preset gender recognition model and an age recognition model to obtain similar recordings similar to the target recording; and sending the target recording and the similar recordings to a human verification terminal for verification.
[0006] Secondly, embodiments of the present invention also provide a target user verification device based on voiceprint profiling, which includes a unit for performing the above-described method.
[0007] Thirdly, embodiments of the present invention also provide a computer device, which includes a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the above-described method.
[0008] Fourthly, embodiments of the present invention also provide a computer-readable storage medium storing a computer program, the computer program including program instructions that, when executed by a processor, can implement the above-described method.
[0009] This application provides a target user verification method, apparatus, device, and medium based on voiceprint profiling. First, the voiceprint features of the target recording are analyzed using the voiceprint recognition model to obtain multiple initial recordings. Then, a preset gender recognition model and age recognition model are used to analyze the gender and age features of the multiple initial recordings to evaluate the similarity of the initial recordings, thereby filtering out interfering recordings among the multiple initial recordings. Finally, the most similar recording to the target recording is obtained. The number of similar recordings is small and the accuracy is high. Finally, the target recording and the similar recordings are manually verified, thereby effectively improving the accuracy of counterfeit identification and reducing the workload of manual verification. Attached Figure Description
[0010] To more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings used in the following description of the embodiments will be briefly introduced. Obviously, the drawings described below are some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0011] Figure 1 A schematic flowchart illustrating the target user verification method based on voiceprint profiling provided in an embodiment of the present invention; Figure 2 A schematic diagram of a sub-process of the target user verification method based on voiceprint profiling provided in an embodiment of the present invention; Figure 3 A schematic diagram of a sub-process of the target user verification method based on voiceprint profiling provided in an embodiment of the present invention; Figure 4 A schematic diagram of a sub-process of the target user verification method based on voiceprint profiling provided in an embodiment of the present invention; Figure 5 A schematic diagram of a sub-process of the target user verification method based on voiceprint profiling provided in an embodiment of the present invention; Figure 6 A schematic diagram of a sub-process of the target user verification method based on voiceprint profiling provided in an embodiment of the present invention; Figure 7A schematic block diagram of a target user verification device based on voiceprint profiling provided in an embodiment of the present invention; Figure 8 A schematic block diagram of a computer device provided for an embodiment of the present invention. Detailed Implementation
[0012] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0013] It should be understood that, when used in this specification and the appended claims, the terms "comprising" and "including" indicate the presence of the described features, integrals, steps, operations, elements and / or components, but do not exclude the presence or addition of one or more other features, integrals, steps, operations, elements, components and / or collections thereof.
[0014] It should also be understood that the terminology used in this specification is for the purpose of describing particular embodiments only and is not intended to limit the invention. As used in this specification and the appended claims, the singular forms “a,” “an,” and “the” are intended to include the plural forms unless the context clearly indicates otherwise.
[0015] It should also be further understood that the term "and / or" as used in this specification and the appended claims refers to any combination of one or more of the associated listed items and all possible combinations, and includes such combinations.
[0016] Please see Figure 1 This is a schematic flowchart illustrating the target user verification method based on voiceprint profiling provided in this embodiment of the invention. In this application, the target user verification method based on voiceprint profiling is applied in the field of big data, particularly in the financial or internet sectors. It is applicable to verification scenarios involving high-risk users, dedicated complaint users, and users engaging in fraudulent activities by impersonating customers, thereby improving the accuracy of target user identification, reducing the workload of manual verification, and thus improving verification efficiency and reducing costs.
[0017] This application provides a target user verification method, apparatus, computer device, and storage medium based on voiceprint profiling. The target user verification method based on voiceprint profiling includes: extracting voiceprint features of a target recording using a preset voiceprint recognition model; extracting multiple initial recordings similar to the target recording from a preset database based on the voiceprint features using the voiceprint recognition model, wherein the preset database stores recordings of the target user in advance; filtering the multiple initial recordings using a preset gender recognition model and an age recognition model to obtain similar recordings similar to the target recording; and sending the target recording and the similar recordings to a human verification terminal for verification.
[0018] This application first analyzes the voiceprint features of the target recording using the voiceprint recognition model to obtain multiple initial recordings. Then, it combines a preset gender recognition model and an age recognition model to analyze the gender and age features of the multiple initial recordings and performs a similarity evaluation on the initial recordings to filter out interfering recordings among the multiple initial recordings. Finally, it obtains the most similar recording to the target recording. The number of similar recordings is small and the accuracy is high. Finally, the target recording and the similar recordings are manually verified, thereby effectively improving the accuracy of counterfeit identification and reducing the workload of manual verification.
[0019] Figure 1 This is a flowchart illustrating the target user verification method based on voiceprint profiling provided in an embodiment of the present invention. Figure 1 As shown, the method includes the following steps S10-S30.
[0020] S10. Extract the voiceprint features of the target recording using a preset voiceprint recognition model; Specifically, the target user refers to the designated user requiring review, which may include high-risk users, professional complainants, or users engaging in fraudulent activities by impersonating customers. This applies to scenarios requiring review of designated users. The voiceprint recognition model is a pre-set audio feature extraction model trained based on a deep learning algorithm. The target recording is the recording data of the designated target user, typically a digitized audio file containing human voice, originating from telephone channel recordings, network voice communications, or on-site audio capture. The voiceprint feature refers to the dense floating-point vector of the target recording, extracted by the voiceprint recognition model.
[0021] After receiving a target recording from a target user, the system inputs the recording into the voiceprint recognition model. Upon receiving the recording, the voiceprint recognition model first preprocesses it, including but not limited to: noise reduction, removing environmental noise, electronic interference, and other irrelevant signals using filters; voice activity detection (VAD), automatically segmenting valid speech segments from silent segments and removing silent portions to reduce invalid data interference; and normalizing the valid speech segments to standardize the amplitude range and duration of the audio, ensuring consistency of the input data.
[0022] After preprocessing, the voiceprint recognition model extracts the voiceprint features of the target recording. These voiceprint features are fixed-length, high-dimensional numerical vectors that fully preserve the uniqueness of the target user's voice. The feature vectors extracted from different recordings of the same person will be very close in high-dimensional space, while the feature vectors from different people will be much farther apart. The finally extracted voiceprint features will serve as the digital identity credential for the target recording, providing the most direct data basis for subsequent voiceprint comparison, cluster analysis, and identity verification.
[0023] In one embodiment, step S10 includes step S11.
[0024] S11. Extract the voiceprint features using TDNN and ECAPA technologies.
[0025] Specifically, the voiceprint recognition model integrates two key technologies, TDNN and ECAPA, in its network architecture to extract voiceprint features. TDNN, or Delayed Neural Network, is a deep learning architecture specifically designed for processing temporal signals (such as speech). By introducing flexible context awareness in the time dimension, it can effectively capture the acoustic feature dependencies across different time periods in speech signals, such as the dynamic changes in formants, thereby providing rich underlying acoustic information for voiceprint recognition.
[0026] Building on this, the model further adopts the advanced architectural ideas of ECAPA. ECAPA is not an independent technology, but a comprehensive enhancement of the traditional TDNN architecture. Its name comes from the many innovations it integrates: channel attention mechanism, aggregation techniques, and multi-layer feature fusion.
[0027] More specifically, after preprocessing, the frame-level Mel-spectrum of the target recording is input into the voiceprint recognition model. The target recording data propagates forward through the network, sequentially passing through an enhanced TDNN layer for local temporal modeling, then undergoing feature recalibration via a channel attention module, and finally being converged by a global context-aware pooling layer into a highly condensed, fixed-length voiceprint feature vector. This vector integrates multiple characteristics of the speaker, exhibiting stronger robustness to speech content, channel differences, and transient environmental noise, providing the most crucial and reliable data foundation for subsequent high-precision voiceprint comparison and verification.
[0028] S20. Using a voiceprint recognition model, extract multiple initial recordings similar to the target recording from a preset database based on the voiceprint features, wherein the preset database contains recordings of the target user.
[0029] Specifically, the preset database is a specially optimized vector database. The core content of this database is a pre-stored set of voiceprint feature vectors corresponding to a massive number of recordings from target users. Each recording in the preset database is a recording sample, and each vector data is associated with the original recording's metadata, such as user identifier, recording time, case number, etc. The database supports rapid retrieval and comparison of feature vectors and possesses efficient data reading and processing capabilities. The initial recordings refer to recordings similar to the target recordings screened from the preset database. Multiple initial recordings with similar voiceprint features to the target recordings and originating from multiple different cases are combined into a single recording.
[0030] During the extraction process, the voiceprint recognition model combines vector database technology and vector indexing technology to compare the voiceprint similarity between the target recording and the recordings in the preset database. Specifically, the system uses the voiceprint feature vector of the target recording as the query input, and the preset database utilizes its internally constructed efficient index structure (e.g., HNSW graph or IVF index) to quickly find the K vectors that are closest to the query vector in the entire high-dimensional vector space. Here, "distance" is typically measured using cosine similarity or Euclidean distance; the closer the distance, the more similar the voiceprint features.
[0031] Finally, the preset database returns a candidate list sorted from highest to lowest similarity score. Based on this list, the voiceprint recognition model extracts metadata sets associated with the top-ranked initial recordings most similar to the target recording from the preset database according to preset extraction rules. These retrieved initial recordings constitute the basic candidate set for subsequent in-depth analysis and manual review.
[0032] The entire retrieval process fully leverages the advantages of vectorized computation and dedicated indexes, enabling precise filtering from hundreds of millions of voiceprint feature databases within milliseconds. This demonstrates the feasibility and efficiency of voiceprint verification technology in large-scale application scenarios.
[0033] S30. Filter multiple initial recordings using a preset gender recognition model and age recognition model to obtain similar recordings that are similar to the target recording.
[0034] In this embodiment, a specific implementation process is described, which uses a gender recognition model and an age recognition model to filter the initial recording to obtain a final candidate set that is highly similar to the target recording. This process is a refined filtering based on the initial screening of voiceprint similarity. It aims to utilize the speaker's auxiliary biometric attributes to further exclude recordings with similar acoustic features but significantly inconsistent identity attributes, thereby improving the accuracy and reliability of the verification results.
[0035] In practice, the system uses the initial recordings retrieved in the previous step as the processing objects for this stage. These recordings have already undergone preliminary screening based on voiceprint feature vector similarity. The system calls two independent, pre-trained dedicated models in parallel: a preset gender recognition model and a preset age recognition model.
[0036] The preset gender recognition model is a deep learning-based classifier, typically employing a convolutional neural network or recurrent neural network structure. It receives the spectral features of the audio as input and, through internally learned acoustic patterns (such as fundamental frequency distribution and formant structure), determines the speaker's gender, obtaining a "male" or "female" classification label and corresponding confidence level for each initial recording. It then removes recordings whose classification label and confidence level differ from all other initial recordings.
[0037] The preset age recognition model is also a complex regression or classification model designed to estimate the speaker's age range. This model analyzes age-related acoustic cues in the audio, such as pronunciation stability, speech rate, vocal roughness, and the energy of specific frequency components, to obtain an estimated age or an age range classification (e.g., "youth," "middle-aged," "elderly") for each initial recording. It also removes recordings where the age of a particular initial recording differs from the ages of all other initial recordings.
[0038] This embodiment uses a preset gender recognition model and an age recognition model to filter multiple initial recordings to obtain similar recordings that are similar to the target recording. Specifically, the system first uses the two models to identify the gender and age attributes of each initial recording, obtaining their respective gender labels (and confidence levels) and age estimates. Then, the gender recognition model and the age recognition model compare the attributes of each initial recording with each other according to preset comparison and filtering rules, filtering out initial recordings that are different from other initial recordings. The remaining initial recordings with high similarity are then output as similar recordings to the target recording. These similar recordings refer to recording samples that, after being filtered for both gender and age, are consistent with the target recording in terms of core user identity attributes. They must not only meet the voiceprint feature similarity requirement but also match the target recording in terms of gender and age among multiple similar recordings.
[0039] Finally, through this attribute filtering mechanism, the system obtains a more refined final candidate list of similar recordings from the initial coarse selection set (i.e., multiple initial recordings), which are highly similar in terms of voiceprint features, gender attributes, and age attributes. This process significantly reduces false alarms caused by the inherent limitations of the voiceprint model, providing higher-quality and more relevant data support for subsequent final decisions.
[0040] For example, in the financial field, it is necessary to verify whether "Zhang*fu" is a user who has committed fraud by impersonating a customer. First, the voiceprint features of the target recording of "Zhang*fu" are extracted. Then, based on the voiceprint features, multiple initial recordings are extracted from a pre-set database. For example, four recording samples similar to the target recording are extracted: "Chen*huan," "Zhao*xia," "Zhang*lin," and "Xu*yong," indicating that these four recording samples all come from the same person as the target recording. Then, the four recordings of "Chen*huan," "Zhao*xia," "Zhang*lin," and "Xu*yong" that are similar to the target recording are analyzed. Similar recording samples are filtered by age and gender. For example, if the age recognition model finds that the age of "Zhao*xia" is too different from the age of all other recording samples, then the recording sample of "Zhao*xia" is removed. Or, if the gender recognition model finds that the gender confidence of "Chen*huan" is too different from the gender confidence of all other recording samples, then the recording sample of "Chen*huan" is removed. Finally, a set of data that is highly similar to "Zhang*fu" in the three dimensions of voiceprint, gender and age is obtained. At this time, the similar recordings include the recording samples of "Zhang*lin" and "Xu*yong".
[0041] In one embodiment, such as Figure 2 As shown, step S30 may include steps S31-S34.
[0042] S31. Simultaneously input multiple initial recordings into the gender recognition model and the age recognition model; S32. Filter the multiple initial recordings using the gender recognition model to obtain a first similar recording that is similar to each other; S33. Filter the multiple initial recordings using the age recognition model to obtain mutually similar second similar recordings; S34. Merge the first similar recording and the second similar recording to obtain the similar recording.
[0043] Specifically, the system synchronously inputs multiple initial recordings (i.e., the preliminary candidate set) obtained after voiceprint similarity retrieval into the gender recognition model and the age recognition model. Here, "synchronous input" means that the system simultaneously sends the complete initial recording set to two independent models for processing through parallel computing or asynchronous task scheduling mechanisms, rather than executing them sequentially. This significantly improves the overall filtering efficiency.
[0044] The gender recognition model, as an efficient audio classifier, first determines the gender attribute of all initial recordings in the candidate set of initial recordings. Its subsequent filtering logic does not compare each initial recording with an external target, but rather compares each initial recording within the candidate set. The gender recognition model analyzes the gender recognition results of all initial recordings to obtain the first similar recordings, meaning the system clusters one or more subsets with highly consistent gender attributes. For example, the system might identify all recordings in the candidate set that are judged as "male" with a confidence level higher than a threshold, forming an internally gender-consistent group.
[0045] In parallel, the age recognition model performs age feature analysis on all initial recordings in the candidate set of initial recordings. It also performs intra-group filtering by comparing the age estimates of all initial recordings to obtain second-similar recordings that are similar to each other. This is typically achieved through clustering algorithms (such as K-means) or by setting age range intervals, thereby filtering out a compact subset of recordings whose age estimates are close to each other.
[0046] After obtaining two subsets (a first similar recording set and a second similar recording set) based on gender consistency and age similarity respectively, the system performs a merging operation. This involves merging the datasets of the first and second similar recordings to form a single similar recording dataset. If any recording samples are identical in the first and second similar recordings, only one of them is retained, resulting in a final group of similar recordings with highly cohesive features and strong correlations. This result indicates that these recordings may originate from the same person, meaning that the similar recordings are all impersonated by the target user corresponding to the target recording, thus identifying the target user as someone engaging in fraudulent activities by impersonating a customer.
[0047] Therefore, this embodiment describes a specific implementation process for performing group similarity filtering on multiple initial recordings using the gender recognition model and the age recognition model, and obtaining a final set of similar recordings through a merging strategy. The core feature of this process is that it focuses on the consistency of attributes within the candidate recording group, rather than comparing it only with a single target recording. This method can effectively identify potentially related groups with common characteristic attributes.
[0048] In one embodiment, such as Figure 3 As shown, step S32 may include steps S321-S323.
[0049] S321. Extract gender features corresponding to multiple initial recordings using the gender recognition model; S322. Determine whether there is a first interfering recording based on the gender characteristics corresponding to all the initial recordings; S323. If so, filter the first interfering recording, and the remaining initial recording is the first similar recording.
[0050] Specifically, the system inputs multiple initial recordings into a preset gender recognition model. The model first performs feature extraction to extract gender features corresponding to the multiple initial recordings. These gender features are not simple "male" or "female" classification labels, but rather high-dimensional numerical vectors output by the model's intermediate layers that can meticulously characterize the gender attributes of the voice. These vectors contain richer acoustic information than single labels and can more accurately reflect subtle differences in gender attributes.
[0051] Subsequently, the system enters the core outlier identification stage. Based on the gender characteristics corresponding to all the initial recordings, it determines whether there is a first interfering recording. A group consistency analysis algorithm can be used, which compares and analyzes all initial recordings to identify outliers. This algorithm calculates the distribution of all gender feature vectors in the feature space, and determines whether there is a first interfering recording by measuring the distance between feature vectors or calculating their deviation from the distribution center. The first interfering recording specifically refers to those individuals whose gender feature vectors show significant statistical differences from the gender feature vectors of all recordings within the group, i.e., recordings that exhibit "outlier" gender attributes.
[0052] After judging the first interfering recording, if the judgment result shows that the first interfering recording exists, that is, the system detects one or more such first interfering recordings through the algorithm, the system will perform a filtering operation to filter the first interfering recordings and remove these abnormal individuals with inconsistent gender attributes from the candidate set of the initial recordings. After this filtering operation, the remaining initial recordings are all first similar recordings, and all the first similar recordings constitute a group with high internal consistency in gender characteristics, that is, the first similar recording set.
[0053] In this embodiment, the gender recognition model performs group consistency filtering on multiple initial recordings through feature extraction and outlier detection to obtain the first similar recordings, thereby forming the first similar recording set. This process ensures that subsequent analysis can focus on the recording group that is consistent with each other in terms of the key biometric feature of gender, laying the foundation for more accurate association analysis.
[0054] In one embodiment, such as Figure 4 As shown, step S322 may include steps S3221-S3224.
[0055] S3221. Calculate the gender probability of the initial recording based on the gender characteristics; S3222. Calculate the absolute difference in gender probability between the gender probabilities of each of the initial recordings; S3223. Determine whether the absolute difference in gender probability between each initial recording and all other initial recordings is greater than a preset probability difference; S3224. If the absolute difference in the gender probability is greater than the preset probability difference, then the corresponding initial recording is the first interference recording.
[0056] Specifically, the gender recognition model first calculates the gender probability of the initial recording based on the gender features. This gender probability is a continuous scalar value derived from the output of the gender recognition model (a binary classification model). The model maps voice features to a continuous discrimination interval, such as [-100, 100]. Within this interval, negative values indicate that the model judges the voice to be male, and positive values indicate a female bias, while the absolute value reflects the confidence level of the judgment. To facilitate subsequent calculations, the system normalizes this discrimination score to an approximate [0, 1] probability space using the Sigmoid function, or directly uses the original discrimination score for subsequent calculations. It should be noted that the decision boundary (zero-point threshold) of the model's judgment can be calibrated and offset according to specific hardware pickup characteristics and business requirements to optimize the accuracy of identifying specific gender biases.
[0057] After obtaining the gender probabilities (or discrimination scores) of all initial recordings, the system calculates the absolute difference in gender probabilities between each of the initial recordings. This process constructs a complete difference matrix, quantifying the degree of difference in gender determination between any two recorded individuals within the group.
[0058] Subsequently, the gender recognition model determines whether the absolute difference in gender probability between each initial recording and all other initial recordings is greater than a preset probability difference. This preset probability difference is a key threshold value, set considering the distribution characteristics of the model output and the business's tolerance for attribute consistency. In this embodiment, the preset probability difference is twice the gender standard deviation, which is a fixed value calculated through multiple experiments. The system iterates through each initial recording in the group, checking whether it differs significantly from all other initial recordings in the group, exceeding the preset probability difference.
[0059] For any of the initial recordings that are traversed, if, after judgment, the absolute difference in gender probability is greater than the preset probability difference, that is, there is an insurmountable gap between this recording and all other members in the group in terms of gender attributes, then the system determines that the corresponding initial recording is the first interference recording.
[0060] If there is at least one gender probability absolute difference less than or equal to the preset probability difference, then the initial recording does not belong to the first interfering recording, and its eligibility in the candidate set is retained, providing basic data for age dimension compliance for subsequent merging of similar recordings.
[0061] This method, based on global consistency testing, can rigorously screen out anomalous individuals whose gender attributes are fundamentally opposed to the entire group, ensuring that the final first set of similar recordings is highly purified in terms of gender.
[0062] In one embodiment, such as Figure 5 As shown, step S33 may include steps S331-S333.
[0063] S331. Extract age features corresponding to multiple initial recordings using the age recognition model; S332. Determine whether there is a second interfering recording based on the age characteristics corresponding to all the initial recordings; S333. If so, filter the second interfering recording, and the remaining initial recording is the second similar recording.
[0064] Specifically, the system inputs multiple initial recordings into a preset age recognition model. This model first performs feature extraction, extracting age features corresponding to the initial recordings. These age features are not simple age estimates, but rather high-dimensional numerical vectors output by the model's intermediate layers that more precisely characterize the age attributes of the voice. This vector contains richer acoustic pattern information related to the aging of the vocal organs and changes in the vocal tract, providing a data foundation for subsequent detailed comparisons.
[0065] Subsequently, the system enters the core outlier detection stage. The age recognition model determines whether there is a second interfering recording based on the age characteristics corresponding to all the initial recordings, which can be achieved using an outlier detection algorithm for continuous variables. This algorithm analyzes the overall distribution structure of all age feature vectors in high-dimensional space, and determines whether there is a second interfering recording by calculating the distance of each feature vector to its K nearest neighbors, or by assessing its deviation from the center of the entire data distribution (such as using an isolated forest or a Z-score-based method). The second interfering recording specifically refers to those individuals whose age feature vectors are statistically significantly different from the age feature vectors of most recordings in the group, that is, recordings that are obviously "outliers" in terms of age attributes.
[0066] After judging the second interfering recording, if the judgment result shows that the second interfering recording exists, that is, the system detects one or more such second interfering recordings through the algorithm, the system will perform a filtering operation to filter the second interfering recordings and remove these abnormal individuals with inconsistent age attributes from the candidate set of the initial recordings. After this filtering operation, the remaining initial recordings are all second similar recordings, and all second similar recordings constitute a group with high internal consistency in age characteristics, that is, the set of second similar recordings.
[0067] In this embodiment, the age recognition model performs group consistency filtering on multiple initial recordings through feature extraction and outlier detection to obtain the second similar recordings, thereby forming a second similar recording set. This process ensures that subsequent analysis can focus on the recording group that is consistent with each other in terms of the key biometric feature of age, laying the foundation for more accurate association analysis.
[0068] In one embodiment, such as Figure 6 As shown, step S332 may include steps S3321-S3324.
[0069] S3321. Calculate the age range of the initial recording based on the age characteristics; S3322. Calculate the age difference between each of the initial recordings based on the age range; S3323. Determine whether the age difference between each initial recording and all other initial recordings is greater than a preset age difference; S3324. If all the age differences are greater than the preset age difference, then the corresponding initial recording is the first interference recording.
[0070] Specifically, when removing divorce points using an age recognition model, the model first calculates the age range for each initial recording based on the age characteristics. This age range refers to the age spectrum corresponding to the initial recording, determined by the age characteristics and obtained through quantifying the age characteristic parameters, possessing clear numerical boundaries. That is, based on the preset age recognition model, the model extracts the age characteristics of each initial recording, generating a standardized age feature vector. The model then analyzes the numerical distribution and correlation of each parameter in the feature vector, combined with the age-feature mapping rules learned during training, to calculate the age range corresponding to each initial recording, ensuring an accurate correspondence between the age range and the age characteristics.
[0071] Subsequently, based on the age ranges of all initial recordings, a pairwise comparison calculation matrix is constructed. The age difference between each initial recording and all other initial recordings is calculated one by one. This age difference refers to the numerical difference between the age ranges corresponding to any two initial recordings, used to measure the degree of fit between different recordings in the age dimension. In this embodiment, the median of two age ranges is used as the calculation benchmark, and the absolute value is taken to obtain the age difference for a single comparison combination, ensuring the objectivity and consistency of the difference calculation. For example, if the age range of the speaker in recording A is [30-35] and the age range of the speaker in recording B is [34-40], then the age difference between A and B is 4.
[0072] After calculating all age differences, for each initial recording, it is determined whether its age difference with all other initial recordings is greater than a preset age difference. The preset age difference is a fixed numerical threshold calibrated through experiments, used to determine whether the age difference exceeds a reasonable range, balancing the accuracy and tolerance of age screening. In this embodiment, the preset age difference is twice the age standard deviation, which is a preset fixed value obtained through multiple experiments.
[0073] If all age differences in an initial recording exceed the preset age difference, it indicates that its age attribute has no reasonable fit with other samples and significantly deviates from the age distribution pattern of most initial recordings. In this case, the initial recording is determined to be the second interference recording. If at least one age difference is less than or equal to the preset age difference, the initial recording does not belong to the second interference recording and retains its eligibility in the candidate set, providing compliant basic data for the merging of similar recordings in the future.
[0074] This method, based on global consistency testing, can rigorously screen out anomalous individuals whose age attributes are fundamentally opposed to the entire group, ensuring a high degree of purification in the age dimension of the final second set of similar recordings.
[0075] S40. Send the target recording and the similar recording to the manual review terminal for review.
[0076] Specifically, after the system completes automated screening based on voiceprint recognition, gender recognition, and age recognition models, it generates a candidate set for review. This set contains two core parts: the target recording requiring identity verification, and a group of similar recordings that the system considers highly similar to the target recording across multiple dimensions. The system encapsulates these audio data and their associated metadata into a structured data packet. This data packet contains not only original or securely processed audio file segments but also machine-readable decision-making assistance information generated by the system, such as voiceprint feature similarity scores, gender and age recognition results and confidence levels, case source identifiers for each recording, and other relevant contextual information.
[0077] Subsequently, the system sends this structured data packet to the human reviewer via its internal secure data transmission interface. The human reviewer is a hardware and software integrated platform designed specifically for reviewers, typically manifested as a high-privilege web application interface or dedicated client software. This platform is deployed within the organization's secure intranet environment and equipped with high-quality audio playback devices to ensure reviewers achieve optimal auditory identification results.
[0078] After receiving the task on the manual review platform, the reviewer's interface will clearly display the target recording and detailed information on all similar recordings side-by-side. The reviewer will then perform the core review process: first, listening to each recording one by one, using their professional experience to perceive and compare subjective characteristics such as the speaker's timbre, tone, accent, and language habits; simultaneously, referring to various quantitative data and tags provided by the system as supplementary criteria for judgment. Finally, based on their professional judgment, the reviewer makes a final decision on the review platform interface, such as confirming relevance, excluding relevance, or marking that further investigation is needed.
[0079] In this embodiment, all data interactions during the entire sending and review process are encrypted and recorded to ensure information security and operational traceability, thereby forming a complete and reliable technical closed loop from machine intelligence pre-screening to final decision-making by human experts.
[0080] In this embodiment, the manual review stage is carried out. By performing in-depth screening of similar recordings that are similar to the target recording in three dimensions—voiceprint, gender, and age—the number of similar recordings that finally reach the manual review stage is small and the accuracy is high, thereby reducing the workload of manual review, improving review efficiency, and improving the accuracy of review.
[0081] For example, in the financial technology field, in a scenario involving the verification of users who commit fraud by impersonating customers, this verification method is triggered when the risk control system identifies a suspected fraudulent loan application account (marked as the target user) through transaction behavior monitoring. The system first retrieves the voice verification recording left by the user during the application process as the target recording and extracts its voiceprint features using a voiceprint recognition model. Subsequently, these features are sent to a vector database storing the voiceprint features of historical suspicious application recordings for similarity retrieval, quickly identifying multiple initial recordings with highly similar voiceprints. These initial recordings are associated with dozens of other loan application accounts registered with different ID numbers.
[0082] The system further uses a gender and age recognition model to filter the initial batch of recordings, removing recordings with obviously inconsistent attributes, and finally obtains a set of similar recordings that are highly consistent in the three dimensions of voiceprint, gender, and age.
[0083] The system packages the target recordings and similar recordings, along with associated account information, application time, device fingerprints, and other data, into a structured review task order, which is then sent to the security operations team's manual review end via an encrypted channel. Review experts listen to all recordings using professional headphones and conduct a comprehensive analysis based on the correlation graph data provided by the system. They ultimately confirm that these accounts belong to different aliases controlled by the same person. Based on this, they execute batch bans on all related accounts and initiate further investigations, effectively preventing organized fraud.
[0084] This application first analyzes the voiceprint features of the target recording using the voiceprint recognition model to obtain multiple initial recordings. Then, it combines a preset gender recognition model and an age recognition model to analyze the gender and age features of the multiple initial recordings and performs a similarity evaluation on the initial recordings to filter out interfering recordings among the multiple initial recordings. Finally, it obtains the most similar recording to the target recording. The number of similar recordings is small and the accuracy is high. Finally, the target recording and the similar recordings are manually verified, thereby effectively improving the accuracy of counterfeit identification and reducing the workload of manual verification.
[0085] Figure 7 This is a schematic block diagram of a target user verification device 300 based on voiceprint profiling provided in an embodiment of the present invention. Figure 7 As shown, corresponding to the above-described target user verification method based on voiceprint profiling, the present invention also provides a target user verification device 300 based on voiceprint profiling. This target user verification device 300 includes a unit for executing the above-described target user verification method based on voiceprint profiling, and the device can be configured in a computer device. Specifically, please refer to... Figure 7 The target user verification device 300 based on voiceprint profile includes an extraction unit 301, a filtering unit 302, and a sending unit 303.
[0086] Extraction unit 301 is used to extract voiceprint features of target recordings using a preset voiceprint recognition model; extract multiple initial recordings similar to the target recordings from a preset database based on the voiceprint features using the voiceprint recognition model, wherein the preset database stores recordings of the target user in advance; and extract the voiceprint features using TDNN and ECAPA technologies. The filtering unit 302 is used to filter multiple initial recordings using a preset gender recognition model and an age recognition model to obtain similar recordings that are similar to the target recording; The sending unit 303 is used to send the target recording and the similar recording to the manual review terminal for review.
[0087] In one embodiment, the filtering unit 302 includes an input unit, a first filtering unit, a second filtering unit, and a merging unit.
[0088] An input unit is used to synchronously input multiple initial recordings into the gender recognition model and the age recognition model; The first filtering unit is configured to: filter multiple initial recordings using the gender recognition model to obtain first similar recordings that are similar to each other; extract gender features corresponding to multiple initial recordings using the gender recognition model; determine whether there is a first interfering recording based on the gender features corresponding to all initial recordings; if so, filter the first interfering recording, and the remaining initial recordings are the first similar recordings; calculate the gender probability of the corresponding initial recording based on the gender features; calculate the absolute difference in gender probability between the gender probabilities of each initial recording; determine whether the absolute difference in gender probability between each initial recording and all other initial recordings is greater than a preset probability difference; if the absolute difference in gender probability is greater than the preset probability difference, the corresponding initial recording is the first interfering recording. The second filtering unit is configured to: filter multiple initial recordings using the age recognition model to obtain mutually similar second similar recordings; extract age features corresponding to multiple initial recordings using the age recognition model; determine whether there is a second interfering recording based on the age features corresponding to all initial recordings; if so, filter the second interfering recording, and the remaining initial recordings are the second similar recordings; calculate the age range of the corresponding initial recording based on the age features; calculate the age difference between each initial recording based on the age range; determine whether the age difference between each initial recording and all other initial recordings is greater than a preset age difference; if the age difference is greater than the preset age difference, the corresponding initial recording is the first interfering recording. A merging unit is used to merge the first similar recording and the second similar recording to obtain the similar recording.
[0089] It should be noted that those skilled in the art can clearly understand that the specific implementation process of the above-mentioned target user verification device and its various units based on voiceprint profiling can be referred to the corresponding descriptions in the foregoing method embodiments. For the sake of convenience and brevity, these details will not be repeated here.
[0090] The aforementioned target user verification device 300 based on voiceprint profiling can be implemented as a computer program, which can, for example... Figure 8 It runs on the computer device shown.
[0091] Please see Figure 8 , Figure 8 This is a schematic block diagram of a computer device provided in an embodiment of this application. The computer device 500 can be a terminal or a server. The terminal can be an electronic device with communication functions, such as a smartphone, tablet, laptop, desktop computer, personal digital assistant, or wearable device. The server can be a standalone server or a server cluster composed of multiple servers.
[0092] See Figure 8 The computer device 500 includes a processor 502, a memory, and a network interface 505 connected via a system bus 501. The memory may include a non-volatile storage medium 503 and internal memory 504.
[0093] The non-volatile storage medium 503 may store an operating system 5031 and a computer program 5032. The computer program 5032 includes program instructions that, when executed, cause the processor 502 to perform a target user verification method based on voiceprint profiling.
[0094] The processor 502 provides computing and control capabilities to support the operation of the entire computer device 500.
[0095] The internal memory 504 provides an environment for the execution of the computer program 5032 in the non-volatile storage medium 503. When the computer program 5032 is executed by the processor 502, the processor 502 can execute a target user verification method based on voiceprint profile.
[0096] This network interface 505 is used for network communication with other devices. Those skilled in the art will understand that... Figure 8 The structure shown is merely a block diagram of a portion of the structure related to the present application and does not constitute a limitation on the computer device 500 to which the present application is applied. The specific computer device 500 may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements.
[0097] The processor 502 is used to run a computer program 5032 stored in the memory to implement the steps of the target user verification method based on voiceprint profiling described above.
[0098] It should be understood that in the embodiments of this application, the processor 502 may be a central processing unit (CPU), or it may be other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or any conventional processor.
[0099] It will be understood by those skilled in the art that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program includes program instructions and can be stored in a storage medium, which is a computer-readable storage medium. The program instructions are executed by at least one processor in the computer system to implement the process steps of the embodiments of the above methods.
[0100] Therefore, the present invention also provides a storage medium. This storage medium can be a computer-readable storage medium. The storage medium stores a computer program, wherein the computer program includes program instructions. When executed by a processor, the program instructions cause the processor to perform the steps of the above-described target user verification method based on voiceprint profiling.
[0101] The storage medium can be any computer-readable storage medium capable of storing program code, such as a USB flash drive, portable hard drive, read-only memory (ROM), magnetic disk, or optical disk.
[0102] Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the components and steps of the various examples have been generally described in terms of functionality in the foregoing description. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementations should not be considered beyond the scope of this invention.
[0103] In the several embodiments provided by this invention, it should be understood that the disclosed apparatus and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative. For example, the division of each unit is merely a logical functional division, and there may be other division methods in actual implementation. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed.
[0104] The steps in the method of this invention can be adjusted, merged, or reduced in order according to actual needs. The units in the device of this invention can be merged, divided, or reduced according to actual needs. Furthermore, the functional units in the various embodiments of this invention can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit.
[0105] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, a terminal, or a network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention.
[0106] The above description is merely a specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any person skilled in the art can easily conceive of various equivalent modifications or substitutions within the technical scope disclosed in the present invention, and these modifications or substitutions should all be covered within the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be determined by the scope of the claims.
Claims
1. A method for verifying target users based on voiceprint profiling, characterized in that, The method includes: The voiceprint features of the target recording are extracted using a pre-defined voiceprint recognition model. The voiceprint recognition model extracts multiple initial recordings similar to the target recording from a preset database based on the voiceprint features, wherein the preset database contains recordings of the target user. By using a preset gender recognition model and age recognition model, multiple initial recordings are filtered to obtain similar recordings that are similar to the target recording; The target recording and the similar recordings are sent to a human reviewer for verification.
2. The method according to claim 1, characterized in that, The step of filtering multiple initial recordings using a preset gender recognition model and age recognition model to obtain similar recordings to the target recording includes: Multiple initial recordings are simultaneously input into the gender recognition model and the age recognition model; The gender recognition model is used to filter multiple initial recordings to obtain first similar recordings that are similar to each other. The age recognition model is used to filter multiple initial recordings to obtain second similar recordings that are similar to each other. The first similar recording and the second similar recording are merged to obtain the similar recording.
3. The method according to claim 2, characterized in that, The step of filtering multiple initial recordings using the gender recognition model to obtain mutually similar first similar recordings includes: The gender recognition model is used to extract gender features corresponding to multiple initial recordings. Determine whether there is a first interfering recording based on the gender characteristics corresponding to all the initial recordings; If so, the first interfering recording is filtered out, and the remaining initial recording is the first similar recording.
4. The method according to claim 3, characterized in that, The step of determining whether there is a first interfering recording based on the gender characteristics corresponding to all the initial recordings includes: Calculate the gender probability of the initial recording based on the gender characteristics; Calculate the absolute difference in gender probability between the gender probabilities of each of the initial recordings; Determine whether there is an initial recording whose absolute difference in gender probability with respect to all other initial recordings is greater than a preset probability difference; If so, then one of the initial recordings is the first interference recording.
5. The method according to claim 2, characterized in that, The step of filtering multiple initial recordings using the age recognition model to obtain mutually similar second similar recordings includes: The age features corresponding to multiple initial recordings are extracted using the age recognition model. Determine whether there is a second interfering recording based on the age characteristics corresponding to all the initial recordings; If present, the second interfering recording is filtered out, and the remaining initial recording constitutes the second similar recording.
6. The method according to claim 5, characterized in that, The step of determining whether there is a second interfering recording based on the age characteristics corresponding to all the initial recordings includes: Calculate the age range corresponding to the initial recording based on the age characteristics; Calculate the age difference between each of the initial recordings based on the age range; Determine whether there is an initial recording whose age difference is greater than a preset difference from all other initial recordings; If so, then one of the initial recordings is a second interference recording.
7. The method according to claim 1, characterized in that, The step of extracting the voiceprint features of the target recording using a preset voiceprint recognition model includes: The voiceprint features were extracted using TDNN and ECAPA techniques.
8. A target user verification device based on voiceprint profiling, characterized in that, Includes a unit for performing the method as described in any one of claims 1-7.
9. A computer device, characterized in that, The computer device includes a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the method as described in any one of claims 1-7.
10. A storage medium, characterized in that, The storage medium stores a computer program, which includes program instructions that, when executed by a processor, can implement the method as described in any one of claims 1-7.