Risk identification method and device, computer equipment and storage medium
By analyzing voiceprints and semantics from recorded conversations, a comprehensive risk score is generated, which solves the problems of insufficient identification of black market intermediaries and inconsistencies in statements in existing technologies, thereby improving the accuracy of risk identification and timely interception of fraudulent activities.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- PING AN TECH (SHENZHEN) CO LTD
- Filing Date
- 2026-01-08
- Publication Date
- 2026-05-01
AI Technical Summary
Existing technologies are insufficient in identifying complex fraud patterns involving the involvement of black market intermediaries and contradictory statements, and cannot stop fraudulent activities in a timely manner, especially lacking effective means for real-time identity verification and judgment of the consistency of dialogue logic.
By acquiring multiple time-segment dialogue recordings of the same risk identification case, voiceprint recognition and semantic analysis are performed to generate voiceprint risk signals and semantic risk signals, which are then fused and calculated to determine the risk level.
It enables the systematic identification of complex fraudulent activities involving the intervention of black market intermediaries and contradictory statements, improving the accuracy of risk identification and allowing for timely interception of fraud risks during the process.
Smart Images

Figure CN121963777A_ABST
Abstract
Description
Risk identification methods, devices, computer equipment and storage media Technical Field
[0001] This application relates to the field of artificial intelligence technology and is applied to online processing business scenarios in financial technology, particularly to a risk identification method, device, computer equipment, and storage medium. Background Technology
[0003] Most existing risk control models focus on post-event analysis, exhibiting significant shortcomings in real-time intervention and deep semantic contradiction mining. For example, when malicious individuals impersonate customers to file malicious complaints over the phone, traditional methods are insufficient in real-time identity verification and consistency judgment of dialogue logic, failing to promptly prevent fraudulent activities.
[0004] While there are existing cases in the industry where voiceprint recognition technology has been applied to assist in identity verification and fraud prevention, this technology is typically used only as a single verification point and is not deeply integrated with the semantic analysis of complaint content. This isolated application makes it impossible for the system to systematically identify complex fraud patterns such as the involvement of black market intermediaries and contradictory statements, making it difficult to meet the growing risk control needs of the insurance industry. Therefore, a more effective technical solution is urgently needed to address these issues. Summary of the Invention
[0005] The purpose of this application is to propose a risk identification method, device, computer equipment, and storage medium, aiming to solve the problem of how to identify complex fraud patterns involving the intervention of black market intermediaries and contradictory statements, so as to improve the accuracy of risk identification.
[0006] Firstly, a risk identification method is provided, employing the following technical solution: acquiring dialogue recordings from multiple time periods corresponding to the same risk identification case, the dialogue recordings including current dialogue recordings and historical dialogue recordings; identifying the current dialogue recording to obtain a character-level confidence sequence with timestamps and the current recording text; semantically segmenting the current recording text to obtain several complete sentences and start and end time points; extracting audio intervals corresponding to each complete sentence based on the start and end time points; obtaining several sentence voiceprint features based on the audio intervals and the character-level confidence sequence; clustering the sentence voiceprint features to obtain voiceprint clusters; determining whether different speakers exist based on the voiceprint clusters and generating voiceprint risk signals; converting the historical dialogue recordings into historical recording text; using a language model to perform contradiction analysis on the historical recording text and the current recording text to obtain contradiction confidence scores and generate semantic risk signals; fusing the voiceprint risk signals and semantic risk signals to calculate a comprehensive risk score; and determining the risk level of the risk identification case based on the comprehensive risk score.
[0007] Secondly, a risk identification device is provided, employing the following technical solution: a recording acquisition module, used to acquire dialogue recordings from multiple time periods corresponding to the same risk identification case, the dialogue recordings including current dialogue recordings and historical dialogue recordings; a recording identification module, used to identify the current dialogue recording, obtaining a character-level confidence sequence with a timestamp and the current recording text, and performing semantic segmentation on the current recording text to obtain several complete sentences and start and end time points; a feature extraction module, used to extract the audio interval corresponding to each complete sentence based on the start and end time points, and extract the feature based on the audio interval and the character-level confidence sequence. The system obtains several sentence-by-sentence voiceprint features; a voiceprint clustering module is used to cluster the sentence-by-sentence voiceprint features to obtain voiceprint clusters, and determines whether different speakers exist based on the voiceprint clusters and generates a voiceprint risk signal; a contradiction analysis module is used to convert the historical dialogue recordings into historical recording text, and uses a language model to perform contradiction analysis on the historical recording text and the current recording text to obtain a contradiction confidence score and generate a semantic risk signal; a risk determination module is used to calculate a comprehensive risk score by fusing the voiceprint risk signal and the semantic risk signal, and determines the risk level of the risk identification case based on the comprehensive risk score.
[0008] Thirdly, a computer device is provided, including a memory and a processor, wherein the memory stores computer-readable instructions, and the processor executes the computer-readable instructions to implement the steps of the risk identification method described above.
[0009] Fourthly, a computer-readable storage medium is provided, on which computer-readable instructions are stored, which, when executed by a processor, implement the steps of the risk identification method described above.
[0010] Compared with existing technologies, the embodiments of this application have the following main advantages: This application acquires current and historical dialogue recordings from multiple time periods corresponding to the same risk identification case, converts the current dialogue recording into a character-level confidence sequence with timestamps and the current recording text, performs semantic segmentation on the current recording text to obtain several complete sentences and start and end time points, accurately locating each semantically complete sentence and its start and end time points; based on the start and end time points, it extracts the audio interval corresponding to each complete sentence, and obtains several sentence voiceprint features based on the audio intervals and character-level confidence sequences, effectively improving the quality and comparison accuracy of the sentence voiceprint features; by clustering the sentence voiceprint features, it obtains voiceprint clusters... This application employs a clustering method to determine the presence of different speakers based on voiceprint clusters and generates voiceprint risk signals. This allows for the identification of potential intermediary litigation activities through voiceprint clustering analysis, avoiding the risk of "one person, multiple people" (a fraudulent practice). By converting historical dialogue recordings into historical audio text, a language model is used to analyze the contradictions between the historical and current audio texts, obtaining a contradiction confidence score and generating a semantic risk signal, achieving consistency comparison across conversational information. A comprehensive risk score is calculated by fusing the voiceprint and semantic risk signals, determining the risk level of the identified case based on the comprehensive risk score. This allows for a comprehensive assessment of the overall risk value of "one person, multiple people" voiceprint alerts and semantic contradiction alerts, improving the accuracy of risk identification. This application identifies abnormal "one person, multiple people" patterns by analyzing voiceprint features in recordings and uses semantic analysis to detect logical contradictions in customer statements. It can systematically identify complex fraudulent behaviors such as black market intermediaries intervening in litigation midway and contradictory statements, improving the accuracy of risk identification and facilitating timely interception of fraud risks during the process. Attached Figure Description
[0011] To more clearly illustrate the solutions in this application, the accompanying drawings used in the description of the embodiments of this application will be briefly introduced below. Obviously, the accompanying drawings described below are some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0012] Figure 1 is an exemplary system architecture diagram in which this application can be applied; Figure 2 is a flowchart of an embodiment of the risk identification method of this application; Figure 3 is a flowchart of an embodiment of step S202 of this application; Figure 4 is a flowchart of another embodiment of step S202 of this application; Figure 5 is a structural diagram of an embodiment of the risk identification device of this application; Figure 6 is a structural diagram of an embodiment of the computer device according to this application. Detailed Implementation
[0013] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application pertains; the terminology used herein in the specification of the application is for the purpose of describing particular embodiments only and is not intended to be limiting of the application; the terms "comprising" and "having," and any variations thereof, in the specification, claims, and foregoing drawings of this application, are intended to cover non-exclusive inclusion. The terms "first," "second," etc., in the specification, claims, or foregoing drawings of this application are used to distinguish different objects, not to describe a particular order.
[0014] In this document, the term "embodiment" means that a particular feature, structure, or characteristic described in connection with an embodiment may be included in at least one embodiment of this application. The appearance of this phrase in various places throughout the specification does not necessarily refer to the same embodiment, nor is it a separate or alternative embodiment mutually exclusive with other embodiments. It will be explicitly and implicitly understood by those skilled in the art that the embodiments described herein can be combined with other embodiments.
[0015] To enable those skilled in the art to better understand the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings.
[0016] As shown in Figure 1, the system architecture 100 may include a terminal device 101, a network 102, and a server 103. The terminal device 101 may be a laptop computer 1011, a tablet computer 1012, or a mobile phone 1013. The network 102 is used as a medium to provide a communication link between the terminal device 101 and the server 103. The network 102 may include various connection types, such as wired or wireless communication links or fiber optic cables.
[0017] Customers can use terminal device 101 to interact with server 103 via network 102 to receive or send messages, etc. Various communication client applications can be installed on terminal device 101, such as web browser applications, shopping applications, search applications, instant messaging tools, email clients, social media platform software, etc.
[0018] Terminal device 101 can be various electronic devices with a display screen and support web browsing. In addition to laptops 1011, tablets 1012, or mobile phones 1013, terminal device 101 can also be an e-book reader, an MP3 player (Moving Picture Experts Group Audio Layer III), an MP4 player (Moving Picture Experts Group Audio Layer IV), a laptop computer, and a desktop computer, etc.
[0019] Server 103 can be a server that provides various services, such as a backend server that provides support for the pages displayed on terminal device 101.
[0020] It should be noted that the risk identification method provided in this application embodiment is generally executed by a server / terminal device, and correspondingly, the risk identification device is generally set in the server / terminal device.
[0021] It should be understood that the number of terminal devices, networks, and servers shown in Figure 1 is merely illustrative. Depending on implementation needs, any number of terminal devices, networks, and servers can be included.
[0022] Referring again to Figure 2, a flowchart of an embodiment of the risk identification method according to this application is shown. The risk identification method includes the following steps: Step S201, acquiring multiple time-segment dialogue recordings corresponding to the same risk identification case, the dialogue recordings including current dialogue recordings and historical dialogue recordings; In this embodiment, the electronic device on which the risk identification method runs (e.g., the server / terminal device shown in Figure 1) can acquire the dialogue recordings corresponding to the risk identification case via a wired connection or a wireless connection. It should be noted that the aforementioned wireless connection methods may include, but are not limited to, 3G / 4G / 5G connections, WiFi connections, Bluetooth connections, WiMAX connections, Zigbee connections, UWB (ultra-wideband) connections, and other currently known or future known wireless connection methods.
[0023] In this embodiment, a risk identification case refers to a specific business scenario instance that requires intervention in risk assessment, which may include multiple rounds of conversation records and case-related data, such as policy numbers, complaint times, basic information of the policyholder, mobile phone numbers, etc. For example, in the insurance industry, a surrender case may involve multiple rounds of conversation records between the policyholder and the customer service, and it is necessary to use risk identification methods to determine whether there are fraud behaviors such as black production intermediary agents in the conversation records. The conversation recording is audio data that records the conversations of the other party in the risk identification case, which is divided into the current conversation recording (i.e., the conversation in the latest period) and the historical conversation recording (i.e., the conversation in the past period).
[0024] Specifically, obtain the conversation recordings of multiple periods corresponding to the same risk identification case, and the conversation recordings may include the current conversation recording and the historical conversation recording. For example, in the insurance surrender scenario, collect multiple rounds of conversation recordings of the policyholder, suspected intermediary personnel, and customer service, covering the key periods before and after the occurrence of the case, and form a complete audio evidence chain with a complete time span to provide raw materials for subsequent voiceprint recognition and semantic analysis.
[0025] Step S202, identify the current conversation recording to obtain a character-level confidence sequence with timestamps and the current recording text, and perform semantic clause segmentation on the current recording text to obtain several complete clauses and their start and end time points; in this embodiment, use voiceprint recognition technology to perform frame-by-frame analysis on the current conversation recording to generate a character-level confidence sequence with timestamps, where the character-level confidence sequence with timestamps refers to a sequence composed of characters (or speech frames) corresponding to recognition confidence and time marks. For example, output the recognition probability of each character through a machine learning model (such as the confidence of the character "退" is 0.95), and mark its start / end time in the audio (such as 00:02:01 / 00:02:05). Based on the character-level confidence sequence with timestamps, select multiple characters from the sequence to generate the current recording text to convert the current conversation recording into text form. Then, use semantic recognition technology to perform semantic recognition on the current recording text, and based on the semantic recognition content, segment the continuous recording text into several logically complete sentence units, that is, complete clauses, and mark the start and end time points of each complete clause. For example, combine punctuation prediction, syntactic analysis, or endpoint detection algorithms in natural speech processing to determine the clause boundary and mark the start and end time points of the clause.
[0026] Step S203: Based on the start and end time points, extract the audio interval corresponding to each complete sentence, and obtain several sentence voiceprint features based on the audio interval and the character-level confidence sequence. In this embodiment, based on the start and end time points of each complete sentence, extract the corresponding audio interval from the original current dialogue recording. Determine the corresponding character (or speech frame) from the character-level confidence sequence based on the audio interval, and combine the recognition confidence and time stamp contained in the corresponding character (or speech frame) with the voiceprint features of the audio interval to obtain several sentence voiceprint features. Here, voiceprint features refer to biometric vectors extracted from the audio interval, which may include parameters such as fundamental frequency, spectrum, and Mel-frequency cepstral coefficients (MFCC).
[0027] Step S204: Cluster the sentence-by-sentence voiceprint features to obtain voiceprint clusters. Determine whether different speakers exist based on the voiceprint clusters and generate a voiceprint risk signal. In this embodiment, an unsupervised clustering algorithm is used to cluster the above-mentioned sentence-by-sentence voiceprint features, such as using K-means or DBSCAN algorithms, to calculate the Euclidean distance or cosine similarity between each sentence-by-sentence voiceprint feature, and generate at least one voiceprint cluster based on the calculation results. Determine whether different speakers exist based on the voiceprint clusters. If multiple voiceprint clusters are generated, it is determined that different speakers exist in the current dialogue recording; if only one voiceprint cluster is generated, it is determined that only one speaker exists in the current dialogue recording. Generate a voiceprint risk signal based on the determination result. For example, if different speakers exist, a high-risk warning signal is generated; if only one speaker exists, a low-risk warning signal is generated.
[0028] Step S205: Convert the historical dialogue recordings into historical recording text. Use a language model to perform contradiction analysis on the historical recording text and the current recording text, obtain a contradiction confidence score, and generate a semantic risk signal. In this embodiment, the historical dialogue recordings are converted into text form using speech recognition technology to form historical recording text. A language model is used to perform deep logical contradiction detection on the historical recording text and the current recording text. For example, using the semantic understanding capabilities of LLM, the system analyzes whether there are logical conflicts in the timeline, event descriptions, and identity information of the two recording texts, and calculates the corresponding contradiction confidence scores. Semantic risk information is generated based on the contradiction confidence scores, where a higher score indicates a greater semantic contradiction and a higher probability of fraud.
[0029] Step S206: Calculate a comprehensive risk score by fusing the voiceprint risk signal and the semantic risk signal, and determine the risk level of the risk identification case based on the comprehensive risk score.
[0030] In this embodiment, a weighted fusion algorithm is used to fuse voiceprint risk signals and semantic risk signals to obtain a comprehensive risk score. For example, a linear weighting or machine learning model can be used to dynamically fuse voiceprint risk signals and semantic risk signals. Then, the risk level of the risk identification case is determined based on the comprehensive risk score. For example, a mapping relationship between risk score ranges and risk levels is pre-constructed, and the comprehensive risk score is normalized to ensure a clear correspondence between the comprehensive risk score and the score range. Through a preset risk level classification standard, such as low, medium, and high levels, the comprehensive risk score is mapped to a specific risk level.
[0031] In this embodiment, by acquiring current and historical dialogue recordings from multiple time periods corresponding to the same risk identification case, the current dialogue recording is converted into a character-level confidence sequence with timestamps and the current recording text. Semantic segmentation is performed on the current recording text to obtain several complete sentences and their start and end times, enabling precise location of each semantically complete sentence and its start and end times. Based on the start and end times, audio intervals corresponding to each complete sentence are extracted. Several sentence voiceprint features are obtained based on the audio intervals and the character-level confidence sequence, effectively improving the quality and comparison accuracy of the sentence voiceprint features. By clustering the sentence voiceprint features, voiceprint clusters are obtained, and based on the voiceprint clusters... It can determine whether there are different speakers and generate voiceprint risk signals, identify potential intermediary litigation behavior through voiceprint clustering analysis, and avoid the risk of "one person, multiple people"; by converting historical dialogue recordings into historical recording texts, it uses language models to perform contradiction analysis on historical and current recording texts, obtains contradiction confidence scores and generates semantic risk signals, and achieves consistency comparison of cross-conversation information; by fusing voiceprint risk signals and semantic risk signals to calculate a comprehensive risk score, it determines the risk level of risk identification cases based on the comprehensive risk score, and can comprehensively assess the overall risk value of "one person, multiple people" voiceprint alarms and semantic contradiction alarms, thereby improving the accuracy of risk identification.
[0032] This application identifies the abnormal pattern of "one account used by multiple people" by analyzing the voiceprint features in the recordings, and uses semantic analysis to detect logical contradictions in customer statements. It can systematically identify complex fraudulent behaviors such as black market intermediaries intervening in the litigation process and contradictory statements, thereby improving the accuracy of risk identification and facilitating timely interception of fraud risks during the process.
[0033] In some optional implementations of this embodiment, referring to Figure 3, step S202, which involves recognizing the current dialogue recording to obtain a character-level confidence sequence with a timestamp and the current recording text, may include the following steps: Step S2021, inputting the current dialogue recording into a trained speech recognition model to segment the current dialogue recording into several audio frames; in this embodiment, inputting the current dialogue recording into a trained speech recognition model to segment the current dialogue recording, for example, dividing it into several audio frames of fixed duration. Each audio frame contains a continuous segment of speech signal and retains the original timestamp and other temporal information.
[0034] Optionally, before inputting the current dialogue recording into the model, a spectral subtraction-based algorithm or a deep learning-based model (such as RNNoise) can be used to denoise the current dialogue recording to improve the accuracy of subsequent voiceprint recognition and speech transcription.
[0035] Step S2022: Feature extraction is performed on several audio frames to obtain a spectrum-time sequence corresponding to each audio frame. The time sequence information includes the timestamp of each audio frame. In this embodiment, Fourier transform or Mel spectrum analysis is used to convert several audio frames into a spectrum-time sequence. The spectrum features reflect the frequency distribution characteristics of the sound, and the time sequence information includes the start and end timestamps of each audio frame, forming the basis for time-dimensional positioning.
[0036] Step S2023 involves using an encoder to extract semantic features from the spectrum-time sequence based on a self-attention mechanism, and then using a decoder to generate a character-level confidence sequence based on these semantic features. In this embodiment, the speech recognition model is constructed using an encoder-decoder architecture. The encoder extracts semantic features from the spectrum-time sequence based on a self-attention mechanism, such as capturing contextual features of speech, like the semantic coherence of continuous speech or changes in acoustic patterns. Then, the decoder predicts character probabilities frame by frame based on the semantic features extracted by the encoder, generating a character-level confidence sequence where each character corresponds to a confidence score (e.g., 0.0-1.0), representing the reliability of the recognition result.
[0037] Step S2024: Align the character-level confidence sequence with the timestamp of each audio frame one by one to obtain a character-level confidence sequence with timestamps; In this embodiment, the character-level confidence sequence output by the decoder is aligned with the timestamp of each audio frame one by one to locate the recognition result of each character to a specific time period, thereby obtaining a character-level confidence sequence with timestamps.
[0038] Step S2025: Select high-confidence characters from the timestamped character-level confidence sequence to form an optimal text sequence, and integrate the optimal text sequence and the corresponding timestamps to obtain the timestamped current recorded text.
[0039] In this embodiment, high-confidence characters are selected from the timestamped character-level confidence sequence. For example, high-confidence characters are extracted through threshold screening (such as confidence > 0.8) to exclude low-confidence noises, forming an optimal text sequence, and the timestamps of the corresponding characters are retained. The optimal text sequence and the timestamps corresponding to each character included in the sequence are integrated to obtain the timestamped current recorded text.
[0040] Optionally, after generating the timestamped current recorded text, it may further include: comparing the original current conversation recording with the generated current recorded text to check for timestamp misalignment or semantic discontinuity. For example, if the timestamp span of a certain character in the text exceeds its normal pronunciation duration, or if there is a logical break in the text sequence, the model parameters or feature extraction strategy need to be traced back and adjusted to ensure that the finally output recorded text conforms to the acoustic characteristics and has semantic integrity.
[0041] The solution of this embodiment realizes high-precision conversion from the original audio to the structured text through the segmented processing, feature extraction, semantic modeling, time alignment, and sequence optimization of the speech recognition model, provides a reliable data basis for subsequent voiceprint analysis and semantic contradiction detection, and at the same time ensures the credibility and traceability of the text through confidence screening and timestamp binding.
[0042] In some optional implementation manners of this embodiment, referring to FIG. 4, the above step S202, that is, performing semantic clause segmentation on the current recorded text to obtain a plurality of complete clauses and start and end time points, may include the following steps: Step S2026: Input the current recorded text into the trained semantic clause segmentation model, and analyze the correlation weights between each character in the current recorded text based on the self-attention mechanism to obtain the semantic dependency relationship between words; In this embodiment, the current recorded text is input into the trained semantic clause segmentation model, such as a variant model of the BERT architecture. This model analyzes the semantic correlation weights between each character in the current recorded text through the self-attention mechanism. For example, there is a weak correlation between "Hello" and "I", and a strong correlation between "I" and "want to surrender the insurance", to obtain the semantic dependency relationship between words.
[0043] Step S2027: Determine the punctuation probability distribution between each adjacent character based on the semantic dependency relationship, and determine the punctuation insertion position and punctuation type according to a preset probability threshold and the punctuation probability distribution. In this embodiment, calculating the punctuation probability distribution of adjacent characters based on the semantic dependency relationship may be to form character pairs of two adjacent characters, and output the probability values of various types of punctuation symbols for each character pair, such as a period, a comma, a question mark, etc., to obtain the punctuation probability distribution. The punctuation insertion position is determined through a preset threshold (such as probability ≥ 0.8) and this punctuation probability distribution, and the type of punctuation to be inserted is determined. For example, in "您好我要退保", the probability of inserting a comma between "您好" and "我" is 0.9, and the probability of inserting a period after the character "保" is 0.85. Then it is determined that the position between "您好" and "我" is the position for inserting a comma, and the position after the character "保" is the position for inserting a period, forming an independent sentence "您好,我要退保。".
[0044] Step S2028: Divide the current recorded text into sentences according to the punctuation insertion position and punctuation type, and generate several complete clauses. In this embodiment, divide the current recorded text into sentences according to the punctuation insertion position and the inserted punctuation type, and generate several complete clauses with punctuation marks, forming a clause list. Each clause satisfies semantic integrity. For example, "因为急用钱" as an adverbial clause of reason is separated from the main clause "您好,我要退保" into two independent clauses to ensure clear logical units.
[0045] Step S2029: Obtain the start and end characters of the complete clause, obtain the timestamps corresponding to the start and end characters, and generate the start and end time points corresponding to the complete clause according to the timestamps corresponding to the start and end characters.
[0046] In this embodiment, since the current recorded text is already bound with timestamps, each character has a corresponding start and end time. After dividing the text into several complete clauses, determine the start and end characters of each complete clause, and extract the start timestamps of the first characters and the end timestamps of the last characters of each complete clause to form the start and end time points corresponding to each complete clause, forming the time boundaries of each clause. For example, if the first character "我" of "我要退保" corresponds to 00:00:03 and the last character "保" corresponds to 00:00:06, then the start and end time of this clause is from 00:00:03 to 00:00:06, ensuring that the timestamps are strictly corresponding to the clause content.
[0047] The solution of this embodiment realizes the efficient conversion from text to clauses with timestamps through the self-attention mechanism of the semantic clause model, punctuation probability determination, sentence segmentation, and timestamp alignment, provides an accurate time positioning basis for subsequent voiceprint feature extraction and semantic contradiction analysis, and at the same time ensures the double accuracy of clauses in the semantic and time dimensions.
[0048] In some optional implementations of this embodiment, step S203, i.e., obtaining several sentence voiceprint features based on the audio interval and the character-level confidence sequence, may include the following steps: calculating the speech clarity score of each complete sentence based on the character-level confidence sequence; in this embodiment, a pre-constructed clarity scoring formula is used to calculate the speech clarity score of each complete sentence based on the character-level confidence sequence. The clarity scoring formula is constructed based on the entropy or confidence of the character prediction probability.
[0049] Specifically, for each complete sentence, a clarity score is calculated based on the character-level confidence sequence corresponding to that sentence. The clarity scoring formula can be the following formula (1): w_i=(1 / N)*Σ_{i=1}^{N} P(c_i), where w_i is the speech clarity score corresponding to the i-th complete sentence, N is the number of characters in the complete sentence, and P(c_i) is the confidence of the trained speech recognition model in predicting the i-th character. Ultimately, a higher speech clarity score indicates that the audio frame corresponding to the complete sentence is clearer.
[0050] The voiceprint feature vector for each complete sentence corresponding to an audio interval is extracted. In this embodiment, after extracting the corresponding audio interval from the original audio based on the start and end times of the complete sentence, a pre-trained voiceprint recognition model (such as ECAPA-TDNN) is used for feature extraction. The model outputs a fixed-length embedding vector (such as 256-dimensional) as the voiceprint feature vector for that audio interval. This vector contains the speaker's biometric information, such as fundamental frequency and spectral features.
[0051] The speech clarity score and the voiceprint feature vector are weighted and fused to obtain several sentence-by-sentence voiceprint features.
[0052] In this embodiment, a preset weighted algorithm is used to weight and fuse the speech clarity score with the voiceprint feature vector to obtain several sentence voiceprint features.
[0053] Specifically, after extracting the voiceprint feature vector of the audio interval corresponding to each complete sentence, a weighted fusion is performed based on the speech intelligibility score of the complete sentence to generate the final sentence voiceprint feature representing the audio frame. The weighting algorithm used can be expressed by the following formula (2): V_final = (Σ_{i=1}^{K} w_i * V_i) / (Σ_{i=1}^{K} w_i), where V_final is the final sentence voiceprint feature of the i-th complete sentence, K is the number of sentences, V_i is the voiceprint feature vector of the i-th complete sentence, and w_i is the speech intelligibility score corresponding to the i-th complete sentence.
[0054] This embodiment improves the robustness of the overall voiceprint features by weighting high-definition sentences, especially in scenarios with uneven audio quality, and can effectively suppress the interference of noisy sentences.
[0055] In some optional implementations of this embodiment, step S204, namely determining whether different speakers exist based on the voiceprint cluster and generating a voiceprint risk signal, may include the following steps: obtaining the target customer identifier of the risk identification case; in this embodiment, the corresponding target customer identifier, such as the customer's phone number or customer ID, is extracted from the risk identification case as the core object for voiceprint feature comparison. For example, in a "policy cancellation by agent" case, if the target customer is the policyholder, the target customer identifier may be the phone number used by the customer or a registered customer ID.
[0056] Based on the target customer identifier, the target registered voiceprint features are extracted from a pre-built historical voiceprint database. In this embodiment, the registered voiceprint features of the target customer identifier are extracted from the historical voiceprint database. For example, by querying the historical voiceprint records corresponding to the target customer identifier through a pre-built historical voiceprint database, such as registered baseline voiceprints or voiceprint features from past calls. If the target customer identifier has no associated historical records, initial voiceprint features are generated from the current conversation recording and stored in the database.
[0057] Based on a preset similarity algorithm, the voiceprint similarity between each voiceprint cluster and the target registered voiceprint feature is determined, wherein the number of voiceprint clusters is greater than or equal to one. In this embodiment, based on a preset similarity algorithm, such as cosine similarity or PLDA algorithm, the feature vector of each voiceprint cluster is compared with the target registered voiceprint one by one to determine the voiceprint similarity between each voiceprint cluster and the target registered voiceprint feature. The number of voiceprint clusters is greater than or equal to one. For example, if there are 3 voiceprint clusters under the same phone number, the voiceprint similarity between each cluster and the target registered voiceprint feature needs to be calculated separately, and the results recorded.
[0058] If each of the aforementioned voiceprint similarities meets a preset similarity threshold, then the number of speakers in the current dialogue recording is determined to be one, and a voiceprint risk signal for indicating a low risk level is generated. In this embodiment, the number of speakers is determined according to a preset similarity threshold. If each voiceprint similarity meets a preset similarity threshold, then the number of speakers in the current dialogue recording is determined to be one, that is, the current speaker is legal, and a voiceprint risk signal for indicating a low risk level is generated.
[0059] If one or more voiceprint similarities do not meet the preset similarity threshold, then it is determined that the number of speakers in the current dialogue recording is greater than one, and a voiceprint risk signal is generated to indicate a high risk level.
[0060] In this embodiment, if one or more voiceprint similarities do not meet the preset similarity threshold, it is determined that the number of speakers in the current dialogue recording is greater than one, that is, the current speaker is illegitimate, and a voiceprint risk signal is generated to indicate a high risk level.
[0061] This embodiment of the solution realizes a complete logical chain from voiceprint features to risk signals through historical voiceprint comparison, similarity calculation and threshold determination, effectively identifying the "one account, multiple people" abnormal pattern and supporting accurate risk level assessment.
[0062] In some optional implementations of this embodiment, step S205, namely, using a language model to perform contradiction analysis on the historical and current audio recordings to obtain a contradiction confidence score and generate a semantic risk signal, may include the following steps: organizing the historical and current audio recordings in chronological order, inputting them into a trained language model, and obtaining preset prompt word templates; in this embodiment, a language model is pre-trained, and this language model can be a large language model with strong semantic understanding and reasoning capabilities as its base. The dataset used for training the language model contains a large number of simulated dialogues between insurance customer service representatives and customers or anonymized real dialogues. Various types of logical contradictions (such as contradictions in basic information, contradictions in factual statements, contradictions in claims, etc.) are set in the dialogues, and the existence of logical contradictions, the points of contradiction, the contradiction issues, and the contradiction confidence scores are marked.
[0063] Specifically, historical and current audio recordings are organized chronologically to build the foundation for analysis. For example, historical and current dialogue recordings from the same risk case are sorted chronologically to form a complete dialogue timeline, which is then input into the trained language model. For instance, the recording of the policyholder's initial inquiry about canceling the policy is input first, followed by subsequent dialogue records, ensuring the model can capture the evolution of the customer's demands and potential points of conflict, providing a coherent context for subsequent semantic analysis. The language model then obtains pre-defined prompt word templates. These templates are designed based on the needs of conflict analysis in insurance scenarios and include several pre-defined key questions, such as "customer's zodiac sign," "first insurance purchase date," and "reason for cancellation."
[0064] Preset key questions are extracted from the prompt word template. Based on these key questions, key information is extracted from the historical and current audio recordings to form key question-and-answer pairs. In this embodiment, preset key questions are extracted from the prompt word template. Based on these key questions, key information is sequentially extracted from the historical and current audio recordings to form structured key question-and-answer pairs. For example, "customer's zodiac sign is Horse" is extracted from the historical text, and "customer's zodiac sign is Sheep" is extracted from the current text, forming a contradictory question-and-answer pair, providing clear comparison items for subsequent logical comparison.
[0065] A semantic-level logical consistency comparison is performed on the key question-answer pairs to obtain a contradiction confidence score. In this embodiment, a trained language model is used to perform a semantic-level logical consistency comparison on the key question-answer pairs. The model not only detects literal contradictions, such as the direct conflict between "born in the Year of the Horse" and "born in the Year of the Sheep," but also identifies implicit logical contradictions, such as the causal contradiction between "stated no prior medical history at the time of insurance application" and "stated currently has a chronic disease." A contradiction confidence score is generated by calculating the semantic similarity and contradiction probability between the key question-answer pairs. The score ranges from 0 to 1, with a higher score indicating a more significant contradiction.
[0066] The contradiction confidence score is compared with a preset contradiction threshold. If the contradiction confidence score does not meet the contradiction threshold, a semantic risk signal is generated based on the key question-and-answer pair corresponding to the contradiction confidence score.
[0067] In this embodiment, the contradiction confidence score is compared with a preset threshold, and a semantic risk signal is generated based on the comparison result. If the contradiction confidence score does not meet the contradiction threshold, or exceeds the contradiction threshold, the key question-answer pair corresponding to the contradiction confidence score is determined to have a high-risk semantic contradiction, and a "high-risk semantic signal" is generated.
[0068] If the contradiction confidence score is below the contradiction threshold, a "low-risk semantic signal" is generated. For example, if the contradiction confidence score of the key question-and-answer pair extracted for "customer's zodiac sign" reaches 0.8, a high-risk signal is triggered, indicating that there may be identity fraud or intermediary instigation; if the score is 0.4, it is considered low-risk, which may just be a customer's slip of the tongue or information update.
[0069] For example, the prompt template could be set as: "You are an insurance risk control expert. Please analyze the following conversation records of the same customer at different time periods to find out if there are any logical contradictions in their answers to the same or similar questions."
[0070] The conversation logs are arranged in chronological order: [Conversation 1 Date and Time]: {Text 1} [Conversation 2 Date and Time]: {Text 2}... Please pay special attention to the following aspects: 1. Basic information about the customer (such as family members, work experience, health status).
[0071] 2. Policy details and claims process.
[0072] 3. The client's demands and reasons.
[0073] Please output the analysis results in JSON format, including the following fields: - "has_contradiction": true / false - "contradiction_type": a string describing the type of contradiction - "contradiction_detail": a string detailing the point of contradiction, such as "Regarding the client's zodiac sign, the initial answer was 'horse,' and the later answer was 'sheep'" - "confidence_score": a floating-point number between 0 and 1, representing the confidence level of the contradiction.
[0074] This embodiment of the solution utilizes a deep semantic contradiction analysis framework based on a large language model, combined with a carefully designed prompt word engineering and contradiction quantification algorithm, to perform cross-conversation semantic analysis on key question-answer pairs. This effectively addresses subtle differences and contextual changes in customer responses instigated by black market operators, thereby improving the accuracy of semantic risk assessment in black and gray market identification.
[0075] In some optional implementations of this embodiment, step S206, namely, calculating a comprehensive risk score by fusing the voiceprint risk signal and the semantic risk signal, and determining the risk level of the risk identification case based on the comprehensive risk score, may include the following steps: obtaining a first weight coefficient corresponding to the voiceprint risk signal and obtaining a second weight coefficient corresponding to the semantic risk signal, wherein the first weight coefficient and the second weight coefficient are set based on feedback data from historical cases; in this embodiment, initial weights are pre-assigned to the voiceprint risk signal ("one number, multiple people") and the semantic risk signal respectively: a first weight coefficient W_v and a second weight coefficient W_t. These weight coefficients can be initialized based on feedback data from historical cases (such as the accuracy rate of each risk type in confirmed black market cases) and updated online or offline using machine learning models (such as logistic regression) based on continuously incoming feedback data.
[0076] Specifically, after generating voiceprint risk signals and semantic risk signals for the same risk identification case, the first weighting coefficient corresponding to the voiceprint risk signal and the second weighting coefficient corresponding to the semantic risk signal are obtained.
[0077] A first risk score for the voiceprint risk signal is determined according to preset voiceprint risk assessment rules, and a second risk score for the semantic risk signal is determined according to preset semantic risk assessment rules. In this embodiment, the voiceprint risk assessment rules refer to a series of standards and algorithms used to quantify the severity of voiceprint risk signals. It determines the risk score corresponding to the voiceprint risk signal by performing multi-dimensional analysis on the speech acoustic features (such as fundamental frequency, formants, energy distribution, etc.) and voiceprint uniqueness (such as voiceprint similarity matching, voiceprint stability analysis) contained in the voiceprint risk signal. For example, if the voiceprint describing a speech segment in the voiceprint risk signal does not highly match the voiceprint database of registered insured persons, or if there are abnormal speech rates or abrupt changes in tone in the speech, a corresponding risk score will be assigned according to the preset voiceprint risk assessment rules, thereby converting the voiceprint risk signal into a calculable first risk score. The higher the score, the higher the degree of voiceprint risk. The semantic risk assessment rules are a set of criteria used to measure the potential risk of semantic risk signals. It determines the risk score corresponding to a semantic risk signal by identifying logical contradictions (such as inconsistencies in statements, unreasonable demands, and sensitive words) contained within the semantic risk signal. For example, if a semantic risk signal describes a dialogue with inconsistent semantic logic or contains typical fraudulent keywords such as "transfer to a safe account," it will be assigned a corresponding risk score according to the semantic risk assessment rules. The score reflects the degree of semantic risk, thus converting the semantic risk signal into a calculable second risk score.
[0078] Specifically, the voiceprint risk signal is evaluated according to the preset voiceprint risk assessment rules (such as voiceprint feature matching degree, abnormal frequency and other indicators) to determine the first risk score of the voiceprint risk signal; at the same time, the semantic risk signal is evaluated according to the preset semantic risk assessment rules (such as keyword matching, semantic logic anomaly and other standards) to determine the second risk score of the semantic risk signal.
[0079] The comprehensive risk score is obtained by weighting and summing the first weighting coefficient, the second weighting coefficient, the first risk score, and the second risk score. In this embodiment, the comprehensive risk score is obtained by weighting and summing the first weighting coefficient, the second weighting coefficient, the first risk score, and the second risk score using a preset weighted summation formula.
[0080] For example, the dynamic weighting model formula Risk_Score = (W_v * S_v) + (W_t * S_t) can be used to calculate the comprehensive risk score. This involves multiplying the first weighting coefficient W_v by the first risk score S_v, and the second weighting coefficient W_t by the second risk score S_t, then adding the two products to obtain the comprehensive risk score. This formula dynamically adjusts the contribution of voiceprint and semantic cues through weighting coefficients. For instance, the system might assign higher weights to strong biometric alerts such as "one person, multiple users," while assigning lower weights to single low-confidence semantic contradictions. This ensures that high-confidence alerts (such as the first risk score of voiceprint risk signals) dominate the comprehensive score, avoiding decision-making bias caused by false alarms from a single dimension.
[0081] When the overall risk score is higher than or equal to a preset risk threshold, the risk-identified case is determined as a high-risk case; when the overall risk score is lower than the preset risk threshold, the risk-identified case is determined as a low-risk case.
[0082] In this embodiment, a risk threshold is preset to distinguish between high-risk and low-risk cases. When the overall risk score is higher than or equal to the threshold, the case is determined to be a high-risk case; when the overall risk score is lower than the threshold, it is determined to be a low-risk case.
[0083] For example, a high-risk threshold of Threshold_high = 0.75 and a low-risk threshold of Threshold_low = 0.45 are preset. When the overall risk score > 0.75, the case is marked as "high-risk suspected black market activity" and triggers a manual review queue; when the overall risk score is 0.45 < overall risk score ≤ 0.75, it is marked as "medium-risk suspected" and triggers an enhanced verification process; when the overall risk score is ≤ 0.45, it is marked as "low-risk" and enters a continuous monitoring process. This threshold setting is based on a balance between the false positive rate and the false negative rate of historical cases, ensuring that high-risk cases are prioritized and medium- and low-risk cases are efficiently processed.
[0084] This embodiment constructs a dynamic weighted decision model to comprehensively evaluate the combined risk score of "one person, multiple users" voiceprint alerts and semantic contradiction alerts. This weighted decision model is not simply based on the binary triggering of the two alert types, but rather assigns an adjustable weight coefficient to each alert type and sets a dynamic threshold. For example, the system might assign a higher weight to strong biometric alerts like "one person, multiple users," while assigning a lower weight to single, low-confidence semantic contradiction alerts. Finally, a weighted total risk score is calculated and compared with a risk threshold to output the risk level of the case (e.g., high risk, medium risk, low risk). This dynamic fusion strategy significantly reduces the false alarm rate and enhances the system's adaptability to changes in black market strategies.
[0085] Referring further to FIG5, as an implementation of the method shown in FIG2 above, this application provides an embodiment of a risk identification device, which corresponds to the method embodiment shown in FIG2, and the device can be specifically applied to various electronic devices.
[0086] As shown in Figure 5, the risk identification device 400 described in this embodiment includes: a recording acquisition module 401, a recording identification module 402, a feature extraction module 403, a voiceprint clustering module 404, a contradiction analysis module 405, and a risk determination module 406. Specifically: the recording acquisition module 401 is used to acquire dialogue recordings from multiple time periods corresponding to the same risk identification case, including current dialogue recordings and historical dialogue recordings; the recording identification module 402 is used to identify the current dialogue recording, obtaining a character-level confidence sequence with a timestamp and the current recording text, and performing semantic segmentation on the current recording text to obtain several complete sentences and start and end time points; the feature extraction module 403 is used to extract the audio interval corresponding to each complete sentence based on the start and end time points, and obtain several sentence voiceprint features based on the audio interval and the character-level confidence sequence; The voiceprint clustering module 404 is used to cluster the sentence-by-sentence voiceprint features to obtain voiceprint clusters, determine whether different speakers exist based on the voiceprint clusters, and generate a voiceprint risk signal; the contradiction analysis module 405 is used to convert the historical dialogue recording into historical recording text, use a language model to perform contradiction analysis on the historical recording text and the current recording text, obtain a contradiction confidence score, and generate a semantic risk signal; the risk determination module 406 is used to calculate a comprehensive risk score by fusing the voiceprint risk signal and the semantic risk signal, and determine the risk level of the risk identification case based on the comprehensive risk score.
[0087] In some optional implementations of this embodiment, the above-mentioned recording recognition module 402 includes a recording segmentation submodule, a feature extraction submodule, an encoding / decoding submodule, a time alignment submodule, and a text generation submodule, wherein: the recording segmentation submodule is used to input the current dialogue recording into a trained speech recognition model, and segment the current dialogue recording to obtain several audio frames; the feature extraction submodule is used to extract features from the several audio frames to obtain a spectrum-time sequence corresponding to the audio frames, wherein the time sequence information includes the timestamp of each audio frame; the encoding / decoding submodule is used to extract semantic features from the spectrum-time sequence using an encoder based on a self-attention mechanism, and to generate a character-level confidence sequence using a decoder based on the semantic features; the time alignment submodule is used to align the character-level confidence sequence with the timestamp of each audio frame one by one to obtain a character-level confidence sequence with timestamps; the text generation submodule is used to select high-confidence characters from the character-level confidence sequence with timestamps to form an optimal text sequence, and to integrate the optimal text sequence and the corresponding timestamps to obtain the current recording text with timestamps.
[0088] In some optional implementations of this embodiment, the above-mentioned recording recognition module 402 further includes an association analysis submodule, a punctuation determination submodule, a sentence segmentation submodule, and a time annotation submodule, wherein: the association analysis submodule is used to input the current recorded text into the trained semantic sentence segmentation model, analyze the association weights between each character in the current recorded text based on the self-attention mechanism, and obtain the semantic dependency relationship between words; the punctuation determination submodule is used to determine the punctuation probability distribution between each adjacent character based on the semantic dependency relationship, and determine the punctuation insertion position and punctuation type according to a preset probability threshold and the punctuation probability distribution; the sentence segmentation submodule is used to segment the current recorded text into sentences according to the punctuation insertion position and punctuation type, and generate several complete sentences; the time annotation submodule is used to obtain the first and last characters of the complete sentence, obtain the timestamps corresponding to the first and last characters, and generate the start and end time points corresponding to the complete sentence according to the timestamps corresponding to the first and last characters.
[0089] In some optional implementations of this embodiment, the feature extraction module 403 includes a score calculation submodule, a vector extraction submodule, and a weighted fusion submodule, wherein: the score calculation submodule is used to calculate the speech clarity score of each complete sentence based on the character-level confidence sequence; the vector extraction submodule is used to extract the voiceprint feature vector of the audio interval corresponding to each complete sentence; and the weighted fusion submodule is used to perform weighted fusion of the speech clarity score and the voiceprint feature vector to obtain several sentence voiceprint features.
[0090] In some optional implementations of this embodiment, the voiceprint clustering module 404 includes an identifier acquisition submodule, a voiceprint extraction submodule, a similarity determination submodule, a first signal generation submodule, and a second signal generation submodule, wherein: the identifier acquisition submodule is used to acquire the target customer identifier of the risk identification case; the voiceprint extraction submodule is used to extract the target registered voiceprint features from a pre-built historical voiceprint database based on the target customer identifier; the similarity determination submodule is used to determine the voiceprint similarity between each voiceprint cluster and the target registered voiceprint features based on a preset similarity algorithm, wherein the number of voiceprint clusters is greater than or equal to one; the first signal generation submodule is used to determine that the number of speakers in the current dialogue recording is one if each voiceprint similarity meets a preset similarity threshold, and generate a voiceprint risk signal to indicate a low-risk level; the second signal generation submodule is used to determine that the number of speakers in the current dialogue recording is greater than one if one or more voiceprint similarities do not meet the preset similarity threshold, and generate a voiceprint risk signal to indicate a high-risk level.
[0091] In some optional implementations of this embodiment, the aforementioned contradiction analysis module 405 includes a template acquisition submodule, a question-and-answer pair extraction submodule, a semantic comparison submodule, and a third information generation submodule, wherein: the template acquisition submodule is used to organize the historical recording text and the current recording text in chronological order, input them into the trained language model, and obtain a preset prompt word template; the question-and-answer pair extraction submodule is used to extract preset key questions from the prompt word template, and extract key information from the historical recording text and the current recording text based on the key questions to form key question-and-answer pairs; the semantic comparison submodule is used to perform a semantic-level logical consistency comparison based on the key question-and-answer pairs to obtain a contradiction confidence score; the third information generation submodule is used to compare the contradiction confidence score with a preset contradiction threshold, and if the contradiction confidence score does not meet the contradiction threshold, then a semantic risk signal is generated based on the key question-and-answer pair corresponding to the contradiction confidence score.
[0092] In some optional implementations of this embodiment, the risk determination module 406 includes a weight acquisition submodule, a score determination submodule, a score synthesis submodule, a first level determination submodule, and a second level determination submodule, wherein: the weight acquisition submodule is used to acquire a first weight coefficient corresponding to the voiceprint risk signal and a second weight coefficient corresponding to the semantic risk signal, wherein the first weight coefficient and the second weight coefficient are set based on feedback data of historical cases; the score determination submodule is used to determine a first risk score of the voiceprint risk signal according to a preset voiceprint risk assessment rule and a second risk score of the semantic risk signal according to a preset semantic risk assessment rule; the score synthesis submodule is used to perform a weighted summation of the first weight coefficient, the second weight coefficient, the first risk score, and the second risk score to obtain a comprehensive risk score; the first level determination submodule is used to determine the risk identification case as a high-risk case when the comprehensive risk score is higher than or equal to a preset risk threshold; the second level determination submodule is used to determine the risk identification case as a low-risk case when the comprehensive risk score is lower than the preset risk threshold.
[0093] The risk identification device 400 of this application acquires current and historical dialogue recordings from multiple time periods corresponding to the same risk identification case. It converts the current dialogue recording into a character-level confidence sequence with a timestamp and the current recording text. The current recording text is then semantically segmented to obtain several complete sentences and their start and end times, enabling precise location of each semantically complete sentence and its start and end times. Based on the start and end times, the device extracts the audio interval corresponding to each complete sentence. Several sentence voiceprint features are obtained based on the audio interval and the character-level confidence sequence, effectively improving the quality and comparison accuracy of the sentence voiceprint features. By clustering the sentence voiceprint features, voiceprint clusters are obtained, and based on the voiceprint... Clustering determines the existence of different speakers and generates voiceprint risk signals, enabling the identification of potential intermediary litigation behavior through voiceprint clustering analysis and avoiding the risk of "one person, multiple people" (a term used to describe individuals with the same identity but different identities). By converting historical dialogue recordings into historical audio texts, a language model is used to perform contradiction analysis on the historical and current audio texts, obtaining a contradiction confidence score and generating a semantic risk signal, achieving consistency comparison of cross-conversation information. By fusing the voiceprint risk signal and the semantic risk signal, a comprehensive risk score is calculated, and the risk level of the risk identification case is determined based on the comprehensive risk score. This allows for a comprehensive assessment of the overall risk value of "one person, multiple people" voiceprint alarms and semantic contradiction alarms, improving the accuracy of risk identification.
[0094] This application identifies the abnormal pattern of "one account used by multiple people" by analyzing the voiceprint features in the recordings, and uses semantic analysis to detect logical contradictions in customer statements. It can systematically identify complex fraudulent behaviors such as black market intermediaries intervening in the litigation process and contradictory statements, thereby improving the accuracy of risk identification and facilitating timely interception of fraud risks during the process.
[0095] To address the aforementioned technical problems, this application also provides a computer device. Please refer to Figure 6 for details; Figure 6 is a basic structural block diagram of the computer device according to this embodiment.
[0096] The computer device 6 includes a memory 61, a processor 62, and a network interface 63 that are interconnected via a system bus. It should be noted that only the computer device 6 with memory 61, processor 62, and network interface 63 is shown in the figure; however, it should be understood that it is not required to implement all the components shown, and more or fewer components can be implemented alternatively. Those skilled in the art will understand that the computer device described here is a device capable of automatically performing numerical calculations and / or information processing according to pre-set or stored instructions, and its hardware includes, but is not limited to, microprocessors, application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), digital signal processors (DSPs), embedded devices, etc.
[0097] The computer device can be a desktop computer, laptop, handheld computer, or cloud server, etc. The computer device can interact with the customer via a keyboard, mouse, remote control, touchpad, or voice control.
[0098] The memory 61 includes at least one type of readable storage medium, including flash memory, hard disk, multimedia card, card-type memory (e.g., SD or DX memory), random access memory (RAM), static random access memory (SRAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), programmable read-only memory (PROM), magnetic memory, disk, optical disk, etc. In some embodiments, the memory 61 may be an internal storage unit of the computer device 6, such as the hard disk or memory of the computer device 6. In other embodiments, the memory 61 may also be an external storage device of the computer device 6, such as a plug-in hard disk, smart media card (SMC), secure digital (SD) card, flash card, etc., equipped on the computer device 6. Of course, the memory 61 may also include both the internal storage unit and its external storage device of the computer device 6. In this embodiment, the memory 61 is typically used to store the operating system and various application software installed on the computer device 6, such as computer-readable instructions for risk identification methods. In addition, the memory 61 can also be used to temporarily store various types of data that have been output or will be output.
[0099] In some embodiments, the processor 62 may be a central processing unit (CPU), controller, microcontroller, microprocessor, or other data processing chip. The processor 62 is typically used to control the overall operation of the computer device 6. In this embodiment, the processor 62 is used to execute computer-readable instructions stored in the memory 61 or to process data, for example, to execute computer-readable instructions for the risk identification method.
[0100] The network interface 63 may include a wireless network interface or a wired network interface, which is typically used to establish communication connections between the computer device 6 and other electronic devices.
[0101] This application also provides another embodiment, namely, providing a computer-readable storage medium storing computer-readable instructions that can be executed by at least one processor to cause the at least one processor to perform the steps of the risk identification method described above.
[0102] This application provides a computer device and a computer-readable storage medium that, through a processor, acquires current and historical dialogue recordings from multiple time periods corresponding to the same risk identification case. The current dialogue recording is converted into a character-level confidence sequence with a timestamp and the current recording text. Semantic segmentation is performed on the current recording text to obtain several complete sentences and their start and end times, enabling precise location of each semantically complete sentence and its start and end times. Based on the start and end times, audio intervals corresponding to each complete sentence are extracted. Several sentence voiceprint features are obtained based on the audio intervals and the character-level confidence sequence, effectively improving the quality and comparison accuracy of the sentence voiceprint features. Finally, by clustering the sentence voiceprint features, voiceprints are obtained. The system uses voiceprint clustering to determine the existence of different speakers and generates voiceprint risk signals. It can identify potential intermediary litigation behavior through voiceprint clustering analysis, avoiding the risk of "one person, multiple people" (speakers with the same name but different identities). By converting historical dialogue recordings into historical audio text, a language model is used to perform contradiction analysis on the historical and current audio texts, obtaining a contradiction confidence score and generating a semantic risk signal, achieving consistency comparison of cross-conversation information. By fusing the voiceprint risk signal and the semantic risk signal, a comprehensive risk score is calculated. Based on the comprehensive risk score, the risk level of the risk identification case is determined, which can comprehensively assess the overall risk value of "one person, multiple people" voiceprint alarms and semantic contradiction alarms, improving the accuracy of risk identification.
[0103] This application identifies the abnormal pattern of "one account used by multiple people" by analyzing the voiceprint features in the recordings, and uses semantic analysis to detect logical contradictions in customer statements. It can systematically identify complex fraudulent behaviors such as black market intermediaries intervening in the litigation process and contradictory statements, thereby improving the accuracy of risk identification and facilitating timely interception of fraud risks during the process.
[0104] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods of the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions to cause a terminal device (which may be a mobile phone, computer, server, air conditioner, or network device, etc.) to execute the methods described in the various embodiments of this application.
[0105] Obviously, the embodiments described above are only some embodiments of this application, not all embodiments. The accompanying drawings show preferred embodiments of this application, but do not limit the patent scope of this application. This application can be implemented in many different forms; rather, these embodiments are provided to provide a more thorough and comprehensive understanding of the disclosure of this application. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art can still modify the technical solutions described in the foregoing specific embodiments, or make equivalent substitutions for some of the technical features. Any equivalent structures made using the content of this application's specification and drawings, directly or indirectly applied to other related technical fields, are similarly within the scope of patent protection of this application.
[0106] The software tools or components not belonging to our company that appear in the embodiments of this application are merely examples and do not represent actual use.
Claims
1. A risk identification method, characterized in that, The method includes the following steps: acquiring dialogue recordings from multiple time periods corresponding to the same risk identification case, wherein the dialogue recordings include current dialogue recordings and historical dialogue recordings; The current dialogue recording is identified to obtain a character-level confidence sequence with timestamps and the current recording text. The current recording text is semantically segmented to obtain several complete sentences and start and end time points. Based on the start and end time points, the audio interval corresponding to each complete sentence is extracted. Several sentence voiceprint features are obtained according to the audio intervals and the character-level confidence sequence. The sentence voiceprint features are clustered to obtain voiceprint clusters. Based on the voiceprint clusters, it is determined whether there are different speakers and a voiceprint risk signal is generated. The historical dialogue recording is converted into historical recording text. A language model is used to perform contradiction analysis on the historical recording text and the current recording text to obtain a contradiction confidence score and generate a semantic risk signal. A comprehensive risk score is obtained by fusing the voiceprint risk signal and the semantic risk signal, and the risk level of the risk identification case is determined based on the comprehensive risk score.
2. The risk identification method according to claim 1, characterized in that, The step of recognizing the current dialogue recording to obtain a timestamped character-level confidence sequence and the current recording text includes: inputting the current dialogue recording into a trained speech recognition model; segmenting the current dialogue recording to obtain several audio frames; extracting features from the several audio frames to obtain a spectrum-time sequence corresponding to each audio frame, wherein the time sequence information includes the timestamp of each audio frame; extracting semantic features from the spectrum-time sequence using an encoder based on a self-attention mechanism; generating a character-level confidence sequence using a decoder based on the semantic features; aligning the character-level confidence sequence with the timestamp of each audio frame to obtain a timestamped character-level confidence sequence; selecting high-confidence characters from the timestamped character-level confidence sequence to form an optimal text sequence; and integrating the optimal text sequence and its corresponding timestamp to obtain the timestamped current recording text.
3. The risk identification method according to claim 1, characterized in that, The step of semantically segmenting the current recorded text to obtain several complete sentences and start and end time points includes: inputting the current recorded text into a trained semantic segmentation model, analyzing the association weights between characters in the current recorded text based on a self-attention mechanism to obtain semantic dependencies between words; determining the punctuation probability distribution between each adjacent character based on the semantic dependencies, and determining the punctuation insertion position and punctuation type according to a preset probability threshold and the punctuation probability distribution; dividing the current recorded text into sentences according to the punctuation insertion position and punctuation type to generate several complete sentences; obtaining the first and last characters of the complete sentences, obtaining the timestamps corresponding to the first and last characters, and generating the start and end time points corresponding to the complete sentences according to the timestamps corresponding to the first and last characters.
4. The risk identification method according to claim 1, characterized in that, The step of obtaining several sentence voiceprint features based on the audio interval and the character-level confidence sequence includes: calculating the speech clarity score of each complete sentence based on the character-level confidence sequence; extracting the voiceprint feature vector of the audio interval corresponding to each complete sentence; and performing weighted fusion of the speech clarity score and the voiceprint feature vector to obtain several sentence voiceprint features.
5. The risk identification method according to claim 1, characterized in that, The step of determining whether different speakers exist based on the voiceprint clusters and generating a voiceprint risk signal includes: obtaining the target customer identifier of the risk identification case; extracting target registered voiceprint features from a pre-built historical voiceprint database based on the target customer identifier; determining the voiceprint similarity between each voiceprint cluster and the target registered voiceprint features based on a preset similarity algorithm, wherein the number of voiceprint clusters is greater than or equal to one; if each voiceprint similarity meets a preset similarity threshold, then the number of speakers in the current dialogue recording is determined to be one, and a voiceprint risk signal for indicating a low risk level is generated; if one or more voiceprint similarities do not meet the preset similarity threshold, then the number of speakers in the current dialogue recording is determined to be greater than one, and a voiceprint risk signal for indicating a high risk level is generated.
6. The risk identification method according to claim 1, characterized in that, The step of using a language model to perform contradiction analysis on the historical and current audio recordings, obtain a contradiction confidence score, and generate a semantic risk signal includes: organizing the historical and current audio recordings in chronological order, inputting them into a trained language model, and obtaining a preset prompt word template; extracting preset key questions from the prompt word template, and extracting key information from the historical and current audio recordings based on the key questions to form key question-answer pairs; performing a semantic-level logical consistency comparison based on the key question-answer pairs to obtain a contradiction confidence score; comparing the contradiction confidence score with a preset contradiction threshold, and if the contradiction confidence score does not meet the contradiction threshold, generating a semantic risk signal based on the key question-answer pairs corresponding to the contradiction confidence score.
7. The risk identification method according to any one of claims 1-6, characterized in that, The step of calculating a comprehensive risk score by fusing the voiceprint risk signal and the semantic risk signal, and determining the risk level of the risk identification case based on the comprehensive risk score, includes: obtaining a first weighting coefficient corresponding to the voiceprint risk signal and a second weighting coefficient corresponding to the semantic risk signal, wherein the first weighting coefficient and the second weighting coefficient are set based on feedback data from historical cases; determining a first risk score of the voiceprint risk signal according to a preset voiceprint risk assessment rule and a second risk score of the semantic risk signal according to a preset semantic risk assessment rule; performing a weighted summation of the first weighting coefficient, the second weighting coefficient, the first risk score, and the second risk score to obtain a comprehensive risk score; when the comprehensive risk score is higher than or equal to a preset risk threshold, the risk identification case is determined as a high-risk case; when the comprehensive risk score is lower than the preset risk threshold, the risk identification case is determined as a low-risk case.
8. A risk identification device, characterized in that, include: The recording acquisition module is used to acquire dialogue recordings from multiple time periods corresponding to the same risk identification case, the dialogue recordings including current dialogue recordings and historical dialogue recordings; The recording recognition module is used to recognize the current dialogue recording, obtain a character-level confidence sequence with timestamps and the current recording text, perform semantic segmentation on the current recording text, and obtain several complete sentences and start and end time points; the feature extraction module is used to extract the audio interval corresponding to each complete sentence based on the start and end time points, and obtain several sentence voiceprint features based on the audio intervals and the character-level confidence sequence. The voiceprint clustering module is used to cluster the sentence voiceprint features to obtain voiceprint clusters, determine whether there are different speakers based on the voiceprint clusters, and generate voiceprint risk signals; the contradiction analysis module is used to convert the historical dialogue recordings into historical recording text, use a language model to perform contradiction analysis on the historical recording text and the current recording text, obtain contradiction confidence scores, and generate semantic risk signals. The risk determination module is used to calculate a comprehensive risk score by fusing the voiceprint risk signal and the semantic risk signal, and to determine the risk level of the risk identification case based on the comprehensive risk score.
9. A computer device, characterized in that, The method includes a memory and a processor, wherein the memory stores computer-readable instructions, and the processor executes the computer-readable instructions to implement the steps of the risk identification method as described in any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer-readable instructions, which, when executed by a processor, implement the steps of the risk identification method as described in any one of claims 1 to 7.