Risk control method and device, equipment, storage medium and computer program product
By applying voiceprint recognition risk control methods in telecommunications customer service scenarios, work orders and user risks are dynamically assessed, and voice preprocessing and confidence matching are performed. This solves the problem of low efficiency and difficulty in balancing security experience in existing technologies, and improves recognition accuracy and customer service experience.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- CHINA MOBILE ONLINE SERVICES CO LTD
- Filing Date
- 2026-01-21
- Publication Date
- 2026-05-12
AI Technical Summary
When existing voiceprint verification solutions are directly applied to telecommunications customer service scenarios, they suffer from low system efficiency and difficulty in balancing security and customer experience. This is mainly due to the lack of dynamic assessment of business risks and customer behavior, the inability to quantify the certainty of identity matching, and the lack of optimization for interference in telecommunications customer service scenarios.
A risk control method based on voiceprint recognition is adopted. The risk level of the work order is determined by a predefined business risk rule base, the customer's risk label is determined by obtaining the user's historical behavior data, preprocessing and voiceprint matching are performed before voiceprint verification, and the confidence score is output to execute a differentiated risk response process.
It improved the accuracy of voiceprint matching, achieved quantitative identity matching certainty, optimized customer service experience and security protection level, and reduced false alarms to compliant users.
Smart Images

Figure CN122024740A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of wireless communication technology, and in particular to a risk control method, apparatus, device, storage medium, and computer program product based on voiceprint recognition. Background Technology
[0002] Voiceprint recognition, as a biometric identification technology, has been widely used in the field of identity verification in recent years due to its convenience and contactless nature. In existing technologies, a typical voiceprint verification scheme usually includes the following steps: receiving a voice command, triggering full voiceprint verification, and outputting a binary judgment result.
[0003] However, directly transplanting the aforementioned general voiceprint verification solution to the telecommunications customer service field may present the following technical problems due to the high-risk business, complex call environment, and stringent user experience requirements of this specific scenario: First, the lack of dynamic assessment of business risks and customer behavior, coupled with the indiscriminate application of full-process voiceprint verification to all call tickets, leads to excessive system resource consumption by a large number of low-risk transactions, significantly reducing verification efficiency. Second, the mere binary output of "pass" or "reject" fails to quantify the certainty of identity matching, potentially creating security vulnerabilities and resulting in insufficient accuracy in risk identification. Finally, the existing general voiceprint model and processing flow used for voiceprint verification are not optimized for the special interference in telecommunications customer service scenarios, such as telephone channel compression, cross-talk, and emotional pronunciation, resulting in poor recognition accuracy in this scenario and failing to meet the reliability requirements of actual risk control operations.
[0004] Therefore, there is an urgent need for a voiceprint recognition risk control method that can be adapted to the telecommunications customer service scenario. Summary of the Invention
[0005] This application provides a risk control method based on voiceprint recognition to solve the problem that in the prior art, the traditional voiceprint verification scheme is directly applied to the telecommunications customer service scenario. Due to the use of full triggering and binary judgment mechanism and the lack of optimization for interference in the telecommunications customer service scenario, the system is inefficient and it is difficult to balance security and customer experience.
[0006] This application also provides a risk control device based on voiceprint recognition to solve the problem that in the prior art, when traditional voiceprint verification schemes are directly applied to telecommunications customer service scenarios, the system efficiency is low and it is difficult to balance security and customer experience due to the use of full-volume triggering, binary judgment mechanism and lack of optimization for interference in telecommunications customer service scenarios.
[0007] This application also provides a risk control device based on voiceprint recognition to solve the problem that in the prior art, the traditional voiceprint verification scheme is directly applied to the telecommunications customer service scenario. Due to the use of full-volume triggering and binary judgment mechanism and the lack of optimization for interference in the telecommunications customer service scenario, the system is inefficient and it is difficult to balance security and customer experience.
[0008] This application also provides a computer-readable storage medium to address the problem that in the prior art, when traditional voiceprint verification schemes are directly applied to telecommunications customer service scenarios, the system suffers from low efficiency and difficulty in balancing security and customer experience due to the use of full-volume triggering, binary judgment mechanisms, and lack of optimization for interference in telecommunications customer service scenarios.
[0009] A computer program product is provided to address the problem that in existing technologies, when traditional voiceprint verification schemes are directly applied to telecommunications customer service scenarios, the system suffers from low efficiency and an inability to balance security and customer experience due to the use of full-volume triggering, binary judgment mechanisms, and lack of optimization for interference in telecommunications customer service scenarios.
[0010] The embodiments of this application adopt the following technical solutions: A risk control method based on voiceprint recognition includes: determining the risk level of a customer service work order to be processed according to a predefined business risk rule base; acquiring historical behavior data of the user corresponding to the customer service work order to be processed, and determining the user's customer risk label based on the historical behavior data; determining whether to perform voiceprint verification on the customer service work order to be processed based on the work order risk level and the customer risk label; when the determination result is yes, preprocessing the original call voice corresponding to the customer service work order to be processed to obtain the call voice to be analyzed; performing voiceprint matching on the call voice to be analyzed according to a pre-trained voiceprint matching model to obtain the confidence level of the call voice to be analyzed, wherein the confidence level is used to represent the matching degree between the call voice to be analyzed and the user's voiceprint; and executing a risk response process corresponding to the risk control strategy according to the risk control strategy matched with the confidence level.
[0011] A risk control device based on voiceprint recognition includes: a work order risk level determination unit, used to determine the work order risk level corresponding to a customer service work order to be processed according to a predefined business risk rule base; a customer risk tag determination unit, used to acquire historical behavior data of the user corresponding to the customer service work order to be processed, and determine the customer risk tag of the user according to the historical behavior data; a judgment unit, used to judge whether to perform voiceprint verification on the customer service work order to be processed according to the work order risk level and the customer risk tag; a voice optimization unit, used to preprocess the original call voice corresponding to the customer service work order to be processed to obtain the call voice to be analyzed when the judgment result is yes; a voiceprint matching unit, used to perform voiceprint matching on the call voice to be analyzed according to a pre-trained voiceprint matching model to obtain the confidence level corresponding to the call voice to be analyzed, wherein the confidence level is used to represent the voiceprint matching degree between the call voice to be analyzed and the user; and a risk control unit, used to execute a risk response process corresponding to the risk control strategy according to the risk control strategy matched with the confidence level.
[0012] A risk control device based on voiceprint recognition includes: The system includes a processor and a memory configured to store computer-executable instructions, which, when executed, cause the processor to perform the following operations: determine the risk level of a customer service work order to be processed based on a predefined business risk rule base; acquire historical behavior data of the user corresponding to the customer service work order to be processed, and determine the user's customer risk label based on the historical behavior data; determine whether to perform voiceprint verification on the customer service work order to be processed based on the work order risk level and the customer risk label; when the determination result is yes, preprocess the original call voice corresponding to the customer service work order to be processed to obtain the call voice to be analyzed; perform voiceprint matching on the call voice to be analyzed based on a pre-trained voiceprint matching model to obtain the confidence level corresponding to the call voice to be analyzed, wherein the confidence level is used to represent the matching degree between the call voice to be analyzed and the user's voiceprint; and execute a risk response process corresponding to the risk control strategy based on the risk control strategy matched with the confidence level.
[0013] A computer-readable storage medium stores one or more programs that, when executed by an electronic device including multiple applications, cause the electronic device to perform the following operations: determining the risk level of a customer service work order to be processed based on a predefined business risk rule base; acquiring historical behavior data of the user corresponding to the customer service work order to be processed, and determining the user's customer risk label based on the historical behavior data; determining whether to perform voiceprint verification on the customer service work order to be processed based on the work order risk level and the customer risk label; when the determination result is yes, preprocessing the original call voice corresponding to the customer service work order to be processed to obtain call voice to be analyzed; performing voiceprint matching on the call voice to be analyzed based on a pre-trained voiceprint matching model to obtain a confidence level corresponding to the call voice to be analyzed, wherein the confidence level is used to represent the voiceprint matching degree between the call voice to be analyzed and the user; and executing a risk response process corresponding to a risk control strategy based on a risk control strategy matched with the confidence level.
[0014] A computer program product includes a computer program that, when executed by a processor, performs the following: determining the risk level of a customer service work order to be processed based on a predefined business risk rule base; acquiring historical behavior data of the user corresponding to the customer service work order to be processed, and determining the user's customer risk label based on the historical behavior data; determining whether to perform voiceprint verification on the customer service work order to be processed based on the work order risk level and the customer risk label; when the determination result is yes, preprocessing the original call voice corresponding to the customer service work order to be processed to obtain call voice to be analyzed; performing voiceprint matching on the call voice to be analyzed based on a pre-trained voiceprint matching model to obtain the confidence level corresponding to the call voice to be analyzed, wherein the confidence level is used to represent the matching degree between the call voice to be analyzed and the user's voiceprint; and executing a risk response process corresponding to a risk control strategy based on a risk control strategy matched with the confidence level.
[0015] The above-described technical solutions adopted in the embodiments of this application can achieve the following beneficial effects: The risk control method based on voiceprint recognition provided in this application provides the following steps for customer service risk control management: For pending work orders reported by users via telephone voice, the risk level of the pending customer service work order is first determined based on a predefined business risk rule base. Historical behavior data of the user corresponding to the pending customer service work order is then obtained, and the user's customer risk label is determined based on this data. Furthermore, based on the work order risk level and the customer risk label, it is determined whether to perform voiceprint verification on the pending customer service work order. If the determination is yes, the original call voice corresponding to the pending customer service work order is preprocessed to obtain the call voice to be analyzed. Voiceprint matching is performed on the call voice to be analyzed using a pre-trained voiceprint matching model to obtain the confidence level corresponding to the call voice. Finally, based on the risk control strategy matched with the confidence level, the risk response process corresponding to the risk control strategy is executed. The risk control method based on voiceprint recognition provided in this application has two advantages. First, by performing targeted preprocessing on the original call audio before voiceprint matching, it effectively suppresses the contamination of the original voiceprint features by various interference factors, thereby providing more accurate voice input for subsequent voiceprint matching and greatly improving the accuracy of subsequent voiceprint matching. Second, compared with the binary judgment result given by traditional speech recognition schemes, the risk control method based on voiceprint recognition provided in this application can accurately characterize the certainty of identity matching by outputting a quantitative confidence level. Subsequently, based on this confidence level, the system can trigger a differentiated and graded risk response process, which can effectively intercept identity fraud while minimizing false alarms to compliant users, thereby improving the level of security protection and optimizing the customer service experience. Attached Figure Description
[0016] The accompanying drawings, which are included to provide a further understanding of this application and form part of this application, illustrate exemplary embodiments and are used to explain this application, but do not constitute an undue limitation of this application. In the drawings: Figure 1 A schematic diagram illustrating the specific process of a risk control method based on voiceprint recognition provided in this application embodiment; Figure 2 A schematic diagram of the specific structure of a risk control device based on voiceprint recognition provided in this application embodiment; Figure 3 This is a schematic diagram of the specific structure of a risk control device based on voiceprint recognition, provided as an embodiment of this application. Detailed Implementation
[0017] To make the objectives, technical solutions, and advantages of this application clearer, the technical solutions of this application will be clearly and completely described below in conjunction with specific embodiments and corresponding drawings. Obviously, the described embodiments are only a part of the embodiments of this application, and not all of them. Based on the embodiments in this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0018] This application provides a risk control method based on voiceprint recognition to address the problem in the prior art where traditional voiceprint verification schemes are directly applied to telecommunications customer service scenarios. This results in low system efficiency and an inability to balance security and customer experience due to the use of full-volume triggering, binary judgment mechanisms, and a lack of optimization for interference in telecommunications customer service scenarios.
[0019] The execution subject of the risk control method based on voiceprint recognition provided in this application embodiment may be, but is not limited to, at least one of a customer management server, a risk control server, and a security server; in addition, the execution subject of the method may also be the system or application (APP) itself running on these servers.
[0020] For ease of description, the following description uses a risk control system as the implementing entity of this method as an example. It should be understood that using a risk control system as the implementing entity is merely an illustrative example and should not be construed as a limitation of the method.
[0021] The schematic diagram of the specific implementation process of the risk control method based on voiceprint recognition provided in this application is as follows: Figure 1 As shown, the main steps include the following: Step 11: Determine the risk level of the customer service work order to be processed based on the predefined business risk rule base. When a customer dials the customer service hotline and enters the voice customer service and generates a pending work order, the risk control system can analyze the pending work order to determine the business type or customer intent of the pending work order.
[0022] Meanwhile, the risk control system stores a pre-set business risk rule base based on business needs. The risk control system can determine the risk level of the customer service work order to be processed by matching keywords, business codes or acceptance items in the work order to be processed.
[0023] For example, in one implementation, suppose the business risk rule base has the following three types of risk levels: 1. High-risk work orders: In this application embodiment, a high-risk work order can refer to a work order involving operations that are directly related to account security or cause significant changes in tariffs, such as "account cancellation", "number portability", "downgrading of service plan", or "account unbinding".
[0024] 2. Medium-risk level work orders: In this application embodiment, a medium-risk level work order can refer to a work order involving operations such as "package change (monthly fee reduction <50%)" or "activation or cancellation of value-added services".
[0025] 3. Low-risk work orders: In this application embodiment, a low-risk work order can refer to a work order that involves only information exchange and does not involve changes to account permissions or tariffs, such as "business consultation", "bill inquiry" or "complaint feedback".
[0026] Based on the aforementioned pre-set business risk rule base, assuming the system receives a pending customer service order, and through semantic recognition of the pending customer service order, determines that the user corresponding to the pending customer service order is handling the "account cancellation" business, the risk control system can determine the risk level of the pending customer service order as "high risk" in the business risk rule base based on keyword matching.
[0027] Step 12: Obtain the historical behavior data of the user corresponding to the pending customer service work order, and determine the user's customer risk label based on the historical behavior data; In this embodiment of the application, the risk control system can determine the user corresponding to the pending customer service work order based on the incoming mobile phone number of the user who submitted the work order, and then obtain the user's historical behavior data within a preset period from the database.
[0028] It should be noted that since certain user behaviors are strongly correlated with user account security, the risk level of a user can be judged based on these user behaviors. Therefore, in this embodiment, multiple behavioral factors for determining user risk labels can be pre-set according to business needs. In this embodiment, the specific implementation of step 12 may include: determining the user's behavioral score on each behavioral factor based on the historical behavioral data and at least two preset behavioral factors; performing a weighted summation based on the behavioral scores to obtain the user's total risk score; comparing the total risk score with a preset risk threshold; if the total risk score is greater than the risk threshold, the customer's risk label is determined to be a high-risk label; if the total risk score does not exceed the risk threshold, the customer's risk label is determined to be a non-high-risk label.
[0029] In one implementation, the risk control system can use the following three core factors to determine a user's customer risk label, including: Behavioral Factor 1: Number of business changes; Specifically, the risk control system can determine the number of business change orders submitted by users within a specified period based on historical behavioral data. These business change orders may include, but are not limited to, modifying packages, unbinding devices, and subscribing to / unsubscribing to value-added services.
[0030] Behavioral Factor 2: Frequently Used Devices / Location Matching Degree; Specifically, the risk control system can compare the IMEI / IP address of the current incoming call device with its historically frequently used devices or location.
[0031] Behavioral factor 3: frequency of inbound calls from other networks; Specifically, the risk control system can count the number of times a user calls customer service using a number from a number other than the network operator's within a specified period.
[0032] Assuming the risk control system pre-sets the weights of behavioral factors 1, 2, and 3 to be 40%, 30%, and 30%, respectively, then: if the risk control system determines that a user submits ≥3 business change orders within a specified period, the user's behavioral score for behavioral factor 1 is 40 points; if the number of business change orders is <3, the user's behavioral score for behavioral factor 1 is 0 points. Similarly, if the risk control system determines that the IMEI / IP address of the current incoming call device is inconsistent with its historically used device or location, the user's behavioral score for behavioral factor 2 is 30 points; conversely, if they are different, the user's behavioral score for behavioral factor 2 is 0 points. Furthermore, if the risk control system determines that there are ≥2 inter-network calls within a specified period, the user's behavioral score for behavioral factor 3 is 30 points; if the risk control system determines that there are <2 inter-network calls within a specified period, the user's behavioral score for behavioral factor 3 is 0 points.
[0033] In this embodiment of the application, the risk control system can determine the user's total risk score according to the following formula [1]: T = S1 + S2 + S3 [1] Here, S1~S3 represent the user's behavioral scores in the dimensions of behavioral factor 1 to behavioral factor 3, respectively.
[0034] Furthermore, the risk control system can compare the calculated total risk score T with a preset risk threshold (e.g., 50 points): If T≥50, then the customer's risk label can be determined as "high-risk label".
[0035] If T < 50, then the customer's risk label can be determined as "non-high-risk label".
[0036] For example, continuing with the above example, suppose customer A changes services 4 times within 30 days (S1=40), but the device and region match (S2=0), and there are no calls from other networks (S3=0). Then the total score T=40, and it can be determined that customer A's customer risk label is "non-high-risk label".
[0037] Step 13: Based on the work order risk level obtained by executing Step 11 and the customer risk label obtained by executing Step 12, determine whether to perform voiceprint verification on the customer service work order to be processed. Specifically, in this embodiment of the application, the risk control system can execute binary classification decision logic based on the work order risk level and customer risk label determined in the aforementioned steps to achieve dynamic triggering of voiceprint verification, including: Rule 1, Mandatory Trigger: If the work order risk level is high risk, then regardless of whether the customer risk label is high risk, it is determined that voiceprint verification is required.
[0038] Rule 2, conditional trigger: If the work order risk level is medium risk and the customer risk label is high risk, then it is determined that voiceprint verification is required.
[0039] Rule 3, Exemption Trigger: If the work order risk level is low risk, or the work order risk level is medium risk but the customer risk label is not high risk, then it is determined that voiceprint verification is not required, and the pending customer service work order will directly enter the regular business processing flow.
[0040] Step 14: When the judgment result obtained by executing step 13 is yes, the original call voice corresponding to the customer service work order to be processed is preprocessed to obtain the call voice to be analyzed. When step 13 determines that voiceprint verification is required, the risk control system can extract the original call audio containing the customer's voice from the call recording that triggered the pending customer service work order. It's important to note that in telecommunications customer service scenarios, the original call audio is affected by various factors such as telephone channel compression, environmental noise, customer service interruptions, and fluctuations in the customer's emotions, leading to distortion of voice features. Without processing, this will severely impact the accuracy of subsequent voiceprint matching results. Therefore, preprocessing of the original call audio is necessary before score generation to eliminate interference.
[0041] In this embodiment of the application, the risk control system can preprocess the original call voice in the following manner, including: Preprocessing Step 1: Channel Compensation Since the compression coding of telecommunications channels commonly used by telecommunications customer service loses 3-4kHz high-frequency features, the reverse coding reconstruction algorithm can be used in this application embodiment. By analyzing the coding compression rules, such as the linear predictive coding (LPC) parameters of G.729, the compressed high-frequency spectrum information is reversed and partially recovered during the decoding process to obtain the first call voice, thereby reducing voiceprint feature distortion.
[0042] Preprocessing Step 2: Speaker Separation It should be noted that, due to the alternating "customer service interruption - customer response" speech patterns in customer service calls, and the potential presence of background noise such as office ambient noise or customer-side ambient noise, direct use for recognition would introduce non-customer voiceprint features, leading to score deviations.
[0043] Therefore, in this embodiment of the application, the risk control system can use a pre-trained deep learning-based speech separation model, such as Conv-TasNet, and process the first call speech according to the following sub-steps: Sub-step 1401 converts the first call speech, which is a mixture of customer service voice, customer voice and noise, into a time-domain-frequency domain feature map, such as STFT short-time Fourier transform features. Sub-step 1402: The encoder based on the speech separation model maps the feature map into a latent representation, and then the separator network distinguishes the speech features of customers and customer service staff based on the spectrum and rhythm differences of the speaker's speech. Sub-step 1403: The decoder based on the speech separation model restores the implicit representation of the customer's speech to the time-domain speech signal, and finally realizes the extraction of the customer's pure speech segment to obtain the second call speech.
[0044] Preprocessing Step 3: Emotion Normalization Processing It should be noted that users may experience anger, tension, or other emotions when consulting about high-risk services, which can cause variations in their voice tone, speaking speed, and volume. For example, when angry, their tone may rise and their speaking speed may increase. If these variations are not addressed, they can lead to false differences in the voiceprint characteristics of the same customer, thus affecting the accuracy of subsequent voiceprint recognition.
[0045] Therefore, in this embodiment of the application, the risk control system can perform emotion normalization processing on the second call voice according to the following sub-steps, including: In sub-step 14a, the risk control system extracts the dynamic features of the Mel frequency cepstral coefficients from the second call speech and determines the user's current emotional state, such as anger, calmness, or tension, based on these dynamic features.
[0046] Sub-step 14b, based on the preset "emotion-feature mapping table", normalizes the variability of voiceprint features, and adjusts the increased fundamental frequency when angry back to the statistical average range of the customer in a calm state through linear interpolation, thereby obtaining the voice of the call to be analyzed and improving the consistency of voiceprint features of the same customer under different emotions.
[0047] Through the above three preprocessing steps, the high-frequency features of the call speech to be analyzed can be recovered, the pure voice of the customer can be filtered out, redundant interference can be eliminated, and emotional fluctuations can be eliminated to ensure feature consistency. This provides more accurate voice input for subsequent voiceprint matching and greatly improves the accuracy of subsequent voiceprint matching.
[0048] Step 15: Based on the pre-trained voiceprint matching model, perform voiceprint matching on the speech to be analyzed to obtain the confidence level corresponding to the speech to be analyzed. In this embodiment of the application, the risk control system can determine the confidence level corresponding to the voice call to be analyzed by following the following sub-steps: Sub-step 1501: Based on the voice of the call to be analyzed, obtain the first voiceprint feature corresponding to the voice of the call to be analyzed; It should be noted that, since Mel frequency cepstral coefficients simulate the nonlinear perception of frequency by the human auditory system, they can effectively capture the spectral envelope characteristics of speech and are highly adaptable to telephone channel environments; while linear predictive cepstral coefficients extract the vocal tract resonance characteristics of speech, such as the differences in formant peaks caused by differences in the shape of the oral cavity and nasal cavity, which can serve as key features to distinguish different speakers; and the pitch period reflects the individual differences in the vibration frequency of the customer's vocal cords, such as the male pitch period of about 85-180Hz and the female pitch period of about 165-255Hz.
[0049] Therefore, in this embodiment, combined features such as Mel-frequency cepstral coefficients, linear predictive cepstral coefficients, and pitch period can be extracted from the speech to be analyzed to form the first voiceprint feature vector of the speech. It should be noted that extracting features such as Mel-frequency cepstral coefficients, linear predictive cepstral coefficients, and pitch period from speech is a common technique in the art, so the specific feature extraction method will not be described in detail here.
[0050] Additionally, it should be noted that, in order to ensure the accuracy of subsequent matching results, the extracted original features need to be preprocessed in this embodiment. Specific preprocessing schemes may include, but are not limited to, the following: 1. Mean removal: Eliminate the influence of the overall shift in speech energy on features; 2. Normalization: Map the original feature values to the [0, 1] interval to avoid feature magnitude deviation caused by differences in speech energy among different customers; 3. Differential processing: Extract the first and second differences of the features to capture the dynamic changes of the speech features, further improve the distinguishability of the features, and finally form a combined feature vector of 128-dimensional Meyer frequency cepstral coefficients, 64-dimensional linear prediction cepstral coefficients and pitch period, which constitutes the first voiceprint feature of the call.
[0051] Sub-step 1502: Based on the pre-trained voiceprint matching model, determine the feature matching degree between the first voiceprint feature and the user's pre-stored voiceprint features, and obtain the first score. It should be noted that, to ensure the accuracy of voiceprint feature matching, in this embodiment of the application, the risk control system can train a voiceprint matching model specifically for the telecommunications customer service scenario based on a real voice dataset. The specific training method of the voiceprint matching model is as follows: Step 1: Construct the training dataset; Specifically, by collecting real historical customer service call voice recordings from over 100,000 telecom customers, the collected historical customer service call voice recordings are ensured to cover the following key scenarios to guarantee model adaptability: 1. Channel type: Voice including mainstream telephone channel codes such as G.711 and G.729; 2. Voice type: Includes non-cooperative voice (such as cross-talk, emotionally variable voice) and normal cooperative voice; 3. User coverage: Covering users of different ages, genders, and dialect backgrounds to ensure the generalization ability of the model trained later.
[0052] Step 2: Construct the initial model using a deep learning matching model; Specifically, in the embodiments of this application, a hybrid architecture of convolutional neural network + recurrent neural network + attention mechanism can be selected for the construction of the initial model.
[0053] Among them, the CNN layer extracts the local spectral patterns of speech features through convolutional kernels to capture the static differences in voiceprints; An LSTM (Long Short-Term Memory) network is used to construct the RNN layer, and the RNN layer is used to process the temporal dependencies of speech features, such as the temporal changes in speech rate and intonation, to capture the dynamic differences in voiceprints. By using the Attention layer to identify the feature regions that contribute most to identity differentiation, such as the unique syllable features in a customer's pronunciation, the model's attention to key features is increased, thereby enhancing its discriminative ability.
[0054] Step 3: Construct the loss function; Specifically, in the embodiments of this application, a triplet loss function is used instead of the traditional cross-entropy loss function.
[0055] Specifically, each time the initial model is fed with: Anchor (i.e., the voiceprint features of a certain customer), Positive (other voiceprint features of the same customer), and Negative (voiceprint features of other customers), the distance between Anchor and Positive is minimized and the distance between Anchor and Negative is maximized through the loss function.
[0056] After training the voiceprint matching model, the risk control system can use the voiceprint matching model to calculate the cosine similarity between the first voiceprint feature and multiple historical voiceprint feature samples registered by the customer in the system. The similarity value is then taken as the average value and multiplied by 100, which is mapped to a score of 0-100, thus obtaining the first score reflecting the voiceprint matching degree.
[0057] Sub-step 1503: Determine the second score based on the voice quality of the call to be analyzed; In this embodiment of the application, the risk control system can determine the voice quality of the call to be analyzed based on the following three dimensions of quality indicators: Metric 1, Signal-to-Noise Ratio (SNR); In this embodiment of the application, when SNR≥30dB, the score of the voice to be analyzed in this dimension can be determined to be 100 points, and when SNR≤10dB, the score of the voice to be analyzed in this dimension can be determined to be 0 points.
[0058] Indicator 2, Speech Proficiency (PESQ); In this embodiment of the application, when PESQ≥4.0, the score of the voice to be analyzed in this dimension can be determined to be 100 points, and when PESQ≤2.0, the score of the voice to be analyzed in this dimension can be determined to be 0 points.
[0059] Indicator 3, Speech Integrity Score; In this application embodiment, when the duration without severe interruptions or clippings accounts for ≥95%, the score of the voice to be analyzed in this dimension can be determined to be 100 points, while when the duration without severe interruptions or clippings accounts for ≤80%, the score of the voice to be analyzed in this dimension can be determined to be 0 points.
[0060] The second score is determined by averaging the scores of the three dimensions mentioned above.
[0061] Sub-step 1504: Determine the confidence level based on the first score obtained by executing the above sub-step 1502 and the second score obtained by executing the above sub-step 1503. In this embodiment of the application, a weighted summation method can be used to calculate the final confidence score according to the following formula [2]: Score = A×70% + B×30% [2] Where A represents the first score and B represents the second score.
[0062] Step 16: Based on the risk control strategy that matches the confidence level determined by executing Step 15, execute the risk response process corresponding to the risk control strategy; In this embodiment, the risk control system presets multiple confidence intervals, each interval being associated with a different risk control strategy. Therefore, in this embodiment, risk control can be performed in the following manner: 1. When the confidence level falls within the first confidence interval (e.g., Score ≥ 85): If the voice message to be analyzed is determined to have a high confidence level, the risk control system can execute the processing procedure to allow the business to proceed, automatically pass the verification, and the customer service agent interface will display "Voiceprint verification passed", and the business can continue to be processed.
[0063] 2. When the confidence level falls within the second confidence interval (e.g., 60 ≤ Score < 85): At this point, the voice message to be analyzed is determined to have a medium confidence level. The risk control system can then execute the response procedure of sending a risk alert to the customer service agent. A pop-up message will appear on the agent's interface, displaying "Voiceprint matching is moderate; detailed information verification is recommended," along with a scripted question generated based on the customer's history, such as "What was the address where you installed your broadband last month?" The agent will then manually assess the customer's response and decide whether to continue.
[0064] 3. When the confidence level falls within the third confidence interval (e.g., Score < 60 points): At this point, the call to be analyzed is determined to have low confidence. The risk control system can then initiate a process to transfer the call to a second, manual verification. The work order will be automatically transferred to a complaint handling specialist or risk control specialist with higher authority. Simultaneously, the system generates a "secondary verification checklist," containing non-public information such as the last four digits of the ID card and recent spending details, for the specialist to rigorously verify. If verification fails, the service is rejected and a risk event report is generated.
[0065] Additionally, it's worth noting that after completing the above operations, the risk control system can further record risk control logs for the entire process, including the work order ID, customer ID, work order risk level, customer risk label, scores for each factor, confidence score, final response process, and results. These logs are then stored in a distributed database.
[0066] Subsequently, analysis and iteration are performed based on this log data according to the preset iteration cycle, specifically including: Model iteration: For samples with "medium confidence misjudgment", retrain or adjust the parameters of the voiceprint matching model.
[0067] Rule optimization: Analysis revealed that the actual risk of a certain type of "medium-risk work order" after triggering verification is extremely low. The triggering conditions of S103 or the weight of factors in S102 can be adjusted.
[0068] Threshold optimization: Based on the balance point between "acceptance rate of the person" and "blocking rate of the person not" the confidence interval threshold in S106 is fine-tuned. For example, the high confidence threshold is adjusted from 85 points to 82 points.
[0069] Through the above iterative operations, the risk control system can continuously adapt to business changes and maintain long-term effective risk control capabilities.
[0070] The risk control method based on voiceprint recognition provided in this application provides the following steps for customer service risk control management: For pending work orders reported by users via telephone voice, the risk level of the pending customer service work order is first determined based on a predefined business risk rule base. Historical behavior data of the user corresponding to the pending customer service work order is then obtained, and the user's customer risk label is determined based on this data. Furthermore, based on the work order risk level and the customer risk label, it is determined whether to perform voiceprint verification on the pending customer service work order. If the determination is yes, the original call voice corresponding to the pending customer service work order is preprocessed to obtain the call voice to be analyzed. Voiceprint matching is performed on the call voice to be analyzed using a pre-trained voiceprint matching model to obtain the confidence level corresponding to the call voice. Finally, based on the risk control strategy matched with the confidence level, the risk response process corresponding to the risk control strategy is executed. The risk control method based on voiceprint recognition provided in this application has two advantages. First, by performing targeted preprocessing on the original call audio before voiceprint matching, it effectively suppresses the contamination of the original voiceprint features by various interference factors, thereby providing more accurate voice input for subsequent voiceprint matching and greatly improving the accuracy of subsequent voiceprint matching. Second, compared with the binary judgment result given by traditional speech recognition schemes, the risk control method based on voiceprint recognition provided in this application can accurately characterize the certainty of identity matching by outputting a quantitative confidence level. Subsequently, based on this confidence level, the system can trigger a differentiated and graded risk response process, which can effectively intercept identity fraud while minimizing false alarms to compliant users, thereby improving the level of security protection and optimizing the customer service experience.
[0071] In one embodiment, this application also provides a risk control device based on voiceprint recognition to address the problem in the prior art where traditional voiceprint verification schemes are directly applied to telecommunications customer service scenarios. This is because the schemes employ full-scale triggering and a binary judgment mechanism without optimization for interference in telecommunications customer service scenarios, resulting in low system efficiency and an inability to balance security and customer experience. A schematic diagram of the specific structure of this risk control device based on voiceprint recognition is shown below. Figure 2 As shown, it includes: work order risk level determination unit 21, customer risk label determination unit 22, judgment unit 23, voice optimization unit 24, voiceprint matching unit 25, and risk control unit 26.
[0072] Among them, the work order risk level determination unit 21 is used to determine the work order risk level corresponding to the customer service work order to be processed according to the predefined business risk rule library. The customer risk label determination unit 22 is used to obtain the historical behavior data of the user corresponding to the customer service work order to be processed, and determine the customer risk label of the user based on the historical behavior data. The judgment unit 23 is used to determine whether to perform voiceprint verification on the pending customer service work order based on the work order risk level and the customer risk label. The voice optimization unit 24 is used to preprocess the original call voice corresponding to the customer service work order to be processed when the judgment result is yes, so as to obtain the call voice to be analyzed. Voiceprint matching unit 25 is used to perform voiceprint matching on the speech to be analyzed according to a pre-trained voiceprint matching model to obtain the confidence level corresponding to the speech to be analyzed, wherein the confidence level is used to represent the voiceprint matching degree between the speech to be analyzed and the user. The risk control unit 26 is used to execute a risk response process corresponding to the risk control strategy that matches the confidence level.
[0073] In one implementation, the customer risk label determination unit 22 is specifically configured to: determine the user's behavioral score on each behavioral factor based on the historical behavioral data and at least two preset behavioral factors; perform a weighted summation based on the behavioral scores to obtain the user's total risk score; compare the total risk score with a preset risk threshold, and determine the customer risk label as a high-risk label when the total risk score is greater than the risk threshold; and determine the customer risk label as a non-high-risk label when the total risk score does not exceed the risk threshold.
[0074] In one implementation, the work order risk level includes high risk, medium risk, and low risk. The judgment unit 23 is specifically configured to: if the work order risk level is high risk, determine that voiceprint verification is required for the pending customer service work order; if the work order risk level is medium risk and the customer risk label is a high-risk label, determine that voiceprint verification is required for the pending customer service work order; if the work order risk level is low risk, or the work order risk level is medium risk and the customer risk label is not a high-risk label, determine that voiceprint verification is not required for the pending customer service work order.
[0075] In one implementation, the voice optimization unit 24 is specifically used for: performing channel compensation processing on the original call voice according to the inverse coding reconstruction algorithm to restore the high-frequency voiceprint features of the call voice and obtain a first call voice; performing human voice filtering on the first call voice according to a pre-trained voice separation model to extract the call voice segment corresponding to the user and obtain a second call voice; determining the user's emotion parameters based on the second call voice; and performing normalization processing on the two call voices based on the emotion parameters to obtain the call voice to be analyzed.
[0076] In one embodiment, the voiceprint matching unit 25 is specifically configured to: obtain a first voiceprint feature corresponding to the voiceprint to be analyzed based on the voiceprint to be analyzed; determine the feature matching degree between the first voiceprint feature and the user's pre-stored voiceprint feature to obtain a first score; determine a second score based on the voice quality of the voiceprint to be analyzed; and determine the confidence level based on the first score and the second score.
[0077] In one implementation, the risk control unit 26 is specifically used to: execute a process to allow business processing when the confidence level is in the first confidence level range; execute a process to send a risk warning to customer service agents when the confidence level is in the second confidence level range; and execute a process to transfer the call to a second manual verification when the confidence level is in the third confidence level range.
[0078] Using the voiceprint recognition-based risk control device provided in this application embodiment, during the customer service risk control management process, for a pending work order reported by a user via telephone voice, the device first determines the work order risk level corresponding to the pending customer service work order based on a predefined business risk rule base; then it obtains the historical behavior data of the user corresponding to the pending customer service work order, determines the user's customer risk label based on the historical behavior data, and then determines whether to perform voiceprint verification on the pending customer service work order based on the work order risk level and the customer risk label; when the determination result is yes, it preprocesses the original call voice corresponding to the pending customer service work order to obtain the call voice to be analyzed; it performs voiceprint matching on the call voice to be analyzed based on a pre-trained voiceprint matching model to obtain the confidence level corresponding to the call voice to be analyzed; finally, it executes the risk response process corresponding to the risk control strategy based on the risk control strategy matched with the confidence level. The risk control device based on voiceprint recognition provided in this application has two advantages. First, by performing targeted preprocessing on the original call speech before voiceprint matching, it effectively suppresses the contamination of the original voiceprint features by various interference factors, thereby providing more accurate voice input for subsequent voiceprint matching and greatly improving the accuracy of subsequent voiceprint matching. Second, compared with the binary judgment result given by traditional speech recognition schemes, the risk control method based on voiceprint recognition provided in this application can accurately characterize the certainty of identity matching by outputting a quantitative confidence level. Subsequently, based on this confidence level, the system can trigger differentiated hierarchical risk response processes, which can effectively intercept identity fraud while minimizing false alarms to compliant users, thereby improving the level of security protection and optimizing the customer service experience.
[0079] Figure 3 This is a schematic diagram of the structure of an electronic device according to an embodiment of this application. Please refer to it. Figure 3 At the hardware level, the electronic device includes a processor, and optionally also includes an internal bus, a network interface, and memory. The memory may include main memory, such as high-speed random-access memory (RAM), or non-volatile memory, such as at least one disk drive. Of course, the electronic device may also include other hardware required for other business operations.
[0080] The processor, network interface, and memory can be interconnected via an internal bus, which can be an ISA (Industry Standard Architecture) bus, a PCI (Peripheral Component Interconnect) bus, or an EISA (Extended Industry Standard Architecture) bus, etc. This bus can be divided into address bus, data bus, control bus, etc. For ease of representation, Figure 3 The symbol is represented by a single double-headed arrow, but this does not mean that there is only one bus or one type of bus.
[0081] Memory is used to store programs. Specifically, programs may include program code, which includes computer operation instructions. Memory may include main memory and non-volatile memory, and provides instructions and data to the processor.
[0082] The processor reads the corresponding computer program from non-volatile memory into main memory and then runs it, forming a risk control device based on voiceprint recognition at the logical level. The processor executes the program stored in memory and specifically performs the following operations: Based on a predefined business risk rule base, the risk level of the customer service work order to be processed is determined; historical behavior data of the user corresponding to the customer service work order to be processed is obtained, and the customer risk label of the user is determined based on the historical behavior data; based on the work order risk level and the customer risk label, it is determined whether to perform voiceprint verification on the customer service work order to be processed; when the determination result is yes, the original call voice corresponding to the customer service work order to be processed is preprocessed to obtain the call voice to be analyzed; based on a pre-trained voiceprint matching model, voiceprint matching is performed on the call voice to be analyzed to obtain the confidence level corresponding to the call voice to be analyzed, wherein the confidence level is used to represent the matching degree between the call voice to be analyzed and the user's voiceprint; based on the risk control strategy matched with the confidence level, the risk response process corresponding to the risk control strategy is executed.
[0083] The above is as stated in this application. Figure 3The method for risk control electronic devices based on voiceprint recognition disclosed in the illustrated embodiments can be applied to a processor or implemented by a processor. The processor may be an integrated circuit chip with signal processing capabilities. During implementation, each step of the above method can be completed by integrated logic circuits in the processor's hardware or by instructions in software form. The processor can be a general-purpose processor, including a Central Processing Unit (CPU), a Network Processor (NP), etc.; it can also be a Digital Signal Processor (DSP), an Application Specific Integrated Circuit (ASIC), a Field-Programmable Gate Array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components. It can implement or execute the methods, steps, and logic block diagrams disclosed in the embodiments of this application. The general-purpose processor can be a microprocessor or any conventional processor. The steps of the method disclosed in the embodiments of this application can be directly embodied in the execution of a hardware decoding processor, or executed by a combination of hardware and software modules in the decoding processor. The software module can reside in a mature storage medium in the field, such as random access memory, flash memory, read-only memory, programmable read-only memory, electrically erasable programmable memory, or registers. This storage medium is located in memory, and the processor reads information from the memory and, in conjunction with its hardware, completes the steps of the above method.
[0084] Of course, in addition to software implementation, the electronic device of this application does not exclude other implementation methods, such as logic devices or a combination of hardware and software, etc. In other words, the execution subject of the following processing flow is not limited to each logic unit, but can also be hardware or logic devices.
[0085] This application also proposes a computer-readable storage medium that stores one or more programs, the programs including instructions that, when executed by a portable electronic device including multiple applications, enable the portable electronic device to perform... Figure 1 The risk control method based on voiceprint recognition in the illustrated embodiment is specifically used to perform the following operations: Based on a predefined business risk rule base, the risk level of the customer service work order to be processed is determined; historical behavior data of the user corresponding to the customer service work order to be processed is obtained, and the customer risk label of the user is determined based on the historical behavior data; based on the work order risk level and the customer risk label, it is determined whether to perform voiceprint verification on the customer service work order to be processed; when the determination result is yes, the original call voice corresponding to the customer service work order to be processed is preprocessed to obtain the call voice to be analyzed; based on a pre-trained voiceprint matching model, voiceprint matching is performed on the call voice to be analyzed to obtain the confidence level corresponding to the call voice to be analyzed, wherein the confidence level is used to represent the matching degree between the call voice to be analyzed and the user's voiceprint; based on the risk control strategy matched with the confidence level, the risk response process corresponding to the risk control strategy is executed.
[0086] It should be understood that the training and prediction processes of the AI models involved in the various embodiments of this specification all adhere to multiple legal and compliant principles, including legal data sources, compliant data content, compliant data governance, compliant training objectives and schemes, compliant training processes, compliant training environments and tools, and compliant ethical verification of training results, and comply with the requirements of Article 5 of the Patent Law. Among them: Data source legitimacy: All datasets used for AI model training were obtained through legal means, covering three categories: publicly authorized data, data authorized by partners, and self-collected compliant data. Publicly authorized data comes from compliant data sources following open-source licenses such as Apache 2.0, with complete copyright attribution and authorization scope clearly marked, and no unauthorized open-source code or data reuse. Data authorized by partners has been subject to formal data usage agreements, clearly defining the scope, duration, and confidentiality obligations, and possessing a complete authorization chain. For self-collected data involving personal information, strict informed consent procedures have been followed, and anonymization processes (including but not limited to field masking, feature anonymization, and differential privacy technology applications) have been implemented to remove personally identifiable information, fully complying with the requirements of relevant laws and regulations such as the "Interim Measures for the Administration of Generative Artificial Intelligence Services" and the "Personal Information Protection Law."
[0087] Data content compliance: The AI model's dataset undergoes multiple screenings and cleaning processes to remove all content that may violate social morality or harm public interests. It contains no obscene, pornographic, violent, discriminatory, or information that endangers national or public safety, nor does it involve the illegal acquisition or use of genetic resources. For data in sensitive fields (such as healthcare and finance), an additional privacy-preserving computation module (including federated learning and secure multi-party computation technologies) ensures that the data is "usable but not visible," avoiding compliance risks during the original data transmission process and ensuring that the data application scenarios and uses comply with public order and good morals and industry regulatory requirements.
[0088] Data governance norms: A complete data traceability system is established during the AI model training process to automatically record the source, collection time, annotation process, cleaning rules, and permission allocation of training data, generating traceable compliance reports to ensure that the data is verifiable throughout its entire lifecycle. The dataset annotation process for AI models is completed by a professional human R&D team, clearly defining the proportion of human creative contributions and avoiding reliance on AI-generated data that has not undergone substantial human modification, thus meeting the examination requirements for "human main contributions" in AI patent applications.
[0089] Training objectives and plans are compliant: The AI model training objective focuses on voiceprint matching and authentication. The training scheme and final output results do not violate any mandatory provisions of laws and administrative regulations, do not harm the public interest or the legitimate rights and interests of others, and do not pose any potential risks of being used for illegal activities, infringing on privacy, or disrupting public safety. It strictly adheres to the ethical principle of "intelligent for good".
[0090] Training process compliance: A closed-loop training framework is adopted to ensure compliance and controllability of the training process. The specific process is as follows: First, training samples are obtained through compliant data sources. After the aforementioned data cleaning and desensitization, they are input into the neural network model to generate preliminary training results. Second, an expert system is introduced to verify the preliminary results. Based on preset rules and human expert experience, the feasibility of the results is evaluated, and outputs that may pose ethical risks or compliance hazards are corrected (such as removing decision-making logic that violates public order and good morals, and adjusting model parameters that do not comply with safety regulations). Finally, the loss function weights are dynamically optimized based on expert system feedback to strengthen the model's learning of compliant results, avoid overfitting errors or non-compliant labels, and form a closed-loop control of "data input - model training - expert verification - parameter optimization - result feedback" to ensure that the entire training process complies with A5 ethical review requirements.
[0091] Training environment and tool compliance: AI model training is implemented using nationally licensed chips and a compliant training platform. All open-source frameworks and components used in the training process have obtained their corresponding licenses, and copyright statements and patent citation information are fully retained, with no instances of infringement or reuse. The training environment is built using virtual devices (containers / virtual machines) with fixed random seeds and initial parameter configurations to ensure the reproducibility of the training process. Furthermore, through access control and operation log recording, risks such as data leakage and parameter tampering during training are prevented, ensuring the security and compliance of the training process.
[0092] Training results ethical verification compliance: After the model is trained, it undergoes additional third-party ethical compliance assessment and algorithm filing review to verify that the model output does not violate social morality or harm public interests. For potentially sensitive scenarios (such as public services and intelligent decision-making), a special result verification mechanism is established to ensure that the model always complies with Article 5 of the Patent Law and relevant laws and regulations in practical applications.
[0093] In summary, the data and training process used in the AI model of this specification strictly comply with the relevant provisions of Article 5 of the Patent Law and the Patent Examination Guidelines (2023 Edition), and there are no violations of laws, social ethics, public interests, or illegal use of genetic resources. It fully meets the compliance requirements for patent authorization.
[0094] Those skilled in the art will understand that embodiments of the present invention can be provided as methods, systems, or computer program products. Therefore, the present invention can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, the present invention can take the form of a computer program product embodied on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0095] This invention is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart illustrations and / or block diagrams. Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.
[0096] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.
[0097] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.
[0098] In a typical configuration, a computing device includes one or more processors (CPU), input / output interfaces, network interfaces, and memory.
[0099] Memory may include non-persistent storage in computer-readable media, such as random access memory (RAM) and / or non-volatile memory, such as read-only memory (ROM) or flash RAM. Memory is an example of computer-readable media.
[0100] Computer-readable media includes both permanent and non-permanent, removable and non-removable media that can store information using any method or technology. Information can be computer-readable instructions, data structures, modules of programs, or other data. Examples of computer storage media include, but are not limited to, phase-change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, CD-ROM, digital versatile optical disc (DVD) or other optical storage, magnetic tape, magnetic magnetic disk storage or other magnetic storage devices, or any other non-transferable medium that can be used to store information accessible by a computing device. As defined herein, computer-readable media does not include transient computer-readable media, such as modulated data signals and carrier waves.
[0101] It should also be noted that the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitation, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.
[0102] Those skilled in the art will understand that embodiments of this application can be provided as methods, systems, or computer program products. Therefore, this application can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, this application can take the form of a computer program product embodied on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0103] The above description is merely an embodiment of this application and is not intended to limit this application. Various modifications and variations can be made to this application by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principle of this application should be included within the scope of the claims of this application.
Claims
1. A risk control method based on voiceprint recognition, characterized in that, include: Based on the predefined business risk rule base, determine the risk level of the customer service work order to be processed; Obtain the historical behavior data of the user corresponding to the pending customer service work order, and determine the customer risk label of the user based on the historical behavior data; Based on the work order risk level and the customer risk label, determine whether to perform voiceprint verification on the pending customer service work order; When the judgment result is yes, the original call voice corresponding to the customer service work order to be processed is preprocessed to obtain the call voice to be analyzed. Based on a pre-trained voiceprint matching model, voiceprint matching is performed on the speech to be analyzed to obtain the confidence level corresponding to the speech to be analyzed, wherein the confidence level is used to represent the matching degree between the speech to be analyzed and the user's voiceprint. Based on the risk control strategy that matches the confidence level, execute the risk response process corresponding to the risk control strategy.
2. The method according to claim 1, characterized in that, The step of determining the user's customer risk label based on the historical behavior data specifically includes: Based on the historical behavior data, and using at least two preset behavior factors, determine the user's corresponding behavior score on each behavior factor; The user's total risk score is obtained by weighted summation based on the behavioral scores. The total risk score is compared with a preset risk threshold. If the total risk score is greater than the risk threshold, the customer's risk label is determined to be a high-risk label. When the total risk score does not exceed the risk threshold, the customer risk label is determined to be a non-high-risk label.
3. The method according to claim 1, characterized in that, The work order risk levels include high risk, medium risk, and low risk; The step of determining whether to perform voiceprint verification on the pending customer service work order based on the work order risk level and the customer risk label specifically includes: If the risk level of the work order is high, it is determined that voiceprint verification is required for the pending customer service work order. If the work order risk level is medium risk and the customer risk label is high risk, then it is determined that voiceprint verification is required for the customer service work order to be processed. If the work order risk level is low risk, or the work order risk level is medium risk and the customer risk label is not a high risk label, then it is determined that voiceprint verification is not required for the pending customer service work order.
4. The method according to claim 1, characterized in that, The preprocessing of the original call audio corresponding to the customer service work order to be processed specifically includes: According to the inverse coding reconstruction algorithm, the original call voice is subjected to channel compensation processing to restore the high-frequency voiceprint features of the call voice and obtain the first call voice. Based on a pre-trained speech separation model, human voice filtering is performed on the first call speech, and the call speech segment corresponding to the user is extracted to obtain the second call speech. Based on the second voice recording, the user's emotional parameters are determined. Based on the emotional parameters, the two voice recordings are normalized to obtain the voice recording to be analyzed.
5. The method according to claim 1, characterized in that, The step of performing voiceprint matching on the speech to be analyzed based on a pre-trained voiceprint matching model to obtain the confidence level corresponding to the speech to be analyzed specifically includes: Based on the voice of the call to be analyzed, obtain the first voiceprint feature corresponding to the voice of the call to be analyzed; Determine the feature matching degree between the first voiceprint feature and the user's pre-stored voiceprint feature to obtain a first score; A second score is determined based on the voice quality of the call to be analyzed. The confidence level is determined based on the first score and the second score.
6. The method according to claim 1, characterized in that, The risk control strategy includes a tiered strategy corresponding to different confidence levels; The step of executing a risk response process corresponding to a risk control strategy that matches the confidence level includes: When the confidence level falls within the first confidence level range, the procedure for processing the business is executed. If the confidence level falls within the second confidence level range, the response process of sending a risk warning to the customer service agent will be executed. If the confidence level falls within the third confidence level range, a process to transfer the case to a second manual verification will be executed.
7. A risk control device based on voiceprint recognition, characterized in that, include: The work order risk level determination unit is used to determine the work order risk level corresponding to the customer service work order to be processed based on the predefined business risk rule base. The customer risk label determination unit is used to obtain the historical behavior data of the user corresponding to the customer service work order to be processed, and determine the customer risk label of the user based on the historical behavior data. The judgment unit is used to determine whether to perform voiceprint verification on the pending customer service work order based on the work order risk level and the customer risk label. The voice optimization unit is used to preprocess the original call voice corresponding to the customer service work order to be processed when the judgment result is yes, so as to obtain the call voice to be analyzed. The voiceprint matching unit is used to perform voiceprint matching on the speech to be analyzed based on a pre-trained voiceprint matching model to obtain the confidence level of the speech to be analyzed, wherein the confidence level is used to represent the voiceprint matching degree between the speech to be analyzed and the user. The risk control unit is used to execute risk response procedures corresponding to the risk control strategy that matches the confidence level.
8. A risk control device based on voiceprint recognition, comprising: processor; as well as A memory configured to store computer-executable instructions, which, when executed, cause the processor to perform the following operations: Based on the predefined business risk rule base, determine the risk level of the customer service work order to be processed; Obtain the historical behavior data of the user corresponding to the pending customer service work order, and determine the customer risk label of the user based on the historical behavior data; Based on the work order risk level and the customer risk label, determine whether to perform voiceprint verification on the pending customer service work order; When the judgment result is yes, the original call voice corresponding to the customer service work order to be processed is preprocessed to obtain the call voice to be analyzed. Based on a pre-trained voiceprint matching model, voiceprint matching is performed on the speech to be analyzed to obtain the confidence level corresponding to the speech to be analyzed, wherein the confidence level is used to represent the matching degree between the speech to be analyzed and the user's voiceprint. Based on the risk control strategy that matches the confidence level, execute the risk response process corresponding to the risk control strategy.
9. A computer-readable storage medium storing one or more programs that, when executed by an electronic device including a plurality of applications, cause the electronic device to perform the risk control method based on voiceprint recognition as described in any one of claims 1-6.
10. A computer program product, characterized in that, It includes a computer program that, when executed by a processor, implements the risk control generation method based on voiceprint recognition as described in any one of claims 1-6.