Dialogue recording processing method and device, equipment and medium
By processing dialogue recording data streams in real time, generating recording segments, calculating negative emotions and semantic risks, and combining knowledge graphs to automatically handle risks, the problems of insufficient recording data coverage and delayed response are solved, achieving efficient risk identification and handling.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- PING AN TECH (SHENZHEN) CO LTD
- Filing Date
- 2026-01-09
- Publication Date
- 2026-04-24
AI Technical Summary
In regulatory and compliance-sensitive industries such as fintech, insurance, and healthcare, existing technologies lack comprehensive coverage for the quality auditing and risk management of recorded data. They struggle to accurately identify inappropriate language and emotionally charged risks in complex contexts, and risk response is delayed, making timely intervention impossible.
By receiving real-time dialogue recording data streams, continuously adding new recording segments, calculating negative sentiment index and target semantic hit frequency, constructing a long-term risk profile, and combining it with a regulatory compliance semantic knowledge graph to calculate semantic risk scores, high-risk scenarios are automatically identified and handled.
It achieves full coverage detection of recorded conversations, accurately identifies the risk of emotional manipulation, improves the accuracy of risk identification, significantly increases the speed of risk response, and reduces the regulatory compliance pressure on enterprises.
Smart Images

Figure CN121922158A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of artificial intelligence technology and is applied to online processing business scenarios such as fintech, insurance, and healthcare. In particular, it relates to a method, apparatus, device, and medium for processing recorded conversations. Background Technology
[0002] In regulatory and compliance-sensitive industries such as fintech, insurance, and healthcare, customer service center recordings are a core component of service quality assessment and compliance oversight. They not only need to meet regulatory requirements such as dual recording and anti-misleading sales, but also serve as crucial evidence for preventing business risks. However, current technologies for quality auditing and risk management of this type of recording data still face numerous technical bottlenecks that urgently need to be overcome, making it difficult to meet the needs of comprehensive, accurate, and real-time compliance control.
[0003] Current technologies largely rely on manual sampling, which is limited by labor costs and efficiency, and has insufficient recording coverage. This results in a large number of potential violations and high-risk transaction scenarios going undetected, leading to significant missed detection risks and substantial compliance pressure on enterprises facing regulatory penalties. Furthermore, while speech-to-text technology is widely used, existing systems only recognize regulatory semantics at the literal matching level. They struggle to accurately capture customer service representatives' attempts to circumvent regulations through vague expressions and synonyms in complex contexts, and cannot correlate customer request scenarios to assess the emotional manipulation risks of customer service language. In addition, existing solutions lack the ability to deeply correlate speaker identity with risk characteristics, resulting in ambiguous risk level assessments. Moreover, risk management relies on manual intervention, leading to delayed responses and an inability to provide timely intervention in high-risk scenarios. Summary of the Invention
[0004] The purpose of this application is to provide a method, apparatus, computer device, and storage medium for processing recorded conversations, so as to solve the problems of inaccurate risk identification, insufficient full coverage, and delayed risk handling response in the prior art.
[0005] Firstly, a method for processing recorded conversations is provided, which adopts the following technical solution: The system receives real-time dialogue recording data streams from the target service and generates continuous new recording segments according to preset rules. For each new recording segment and its corresponding new time period, it calculates the negative sentiment index and target semantic hit frequency of the corresponding server within the new time period based on the server-side text data in the new recording segment. Based on the server's historical risk data, the negative sentiment index of the current new time period, and the target semantic hit frequency, it continuously constructs a long-term risk profile of the server. Based on the current new recording segment, a preset regulatory compliance semantic knowledge graph, and the demand-side text data in the current new recording segment, it calculates the semantic risk score of the current new recording segment. When the semantic risk score is greater than a preset compliance threshold, it determines the comprehensive risk level of the current new recording segment based on the semantic risk score, the negative sentiment index of the current new time period, and the cumulative risk score of the long-term risk profile. Based on the comprehensive risk level, it determines the corresponding target agent to process the current new recording segment.
[0006] Secondly, a device for processing recorded conversations is provided, which adopts the following technical solution: The receiving module is used to receive the dialogue recording data stream of the target service in real time and generate continuous new recording segments according to preset rules. The first calculation module is used to calculate the negative sentiment index and target semantic hit frequency of the corresponding server in the newly added time period for each newly added recording segment, based on the server text data in the newly added recording segment. The module is used to continuously build a long-term risk profile for the server based on the server's historical risk data, the negative sentiment index of the current new time period, and the frequency of target semantic hits. The second calculation module is used to calculate the semantic risk score of the newly added recording segment based on the newly added recording segment, the preset regulatory compliance semantic knowledge graph, and the demand-side text data in the newly added recording segment. The determination module is used to determine the comprehensive risk level of the newly added recording segment based on the semantic risk score, the negative sentiment index in the current newly added time period, and the cumulative risk score of the long-term risk profile when the semantic risk score is greater than the preset compliance threshold. The processing module is used to determine the corresponding target agent based on the comprehensive risk level, so that the target agent can process the newly added audio segment.
[0007] Thirdly, a computer device is provided, which adopts the following technical solution: The system receives real-time dialogue recording data streams from the target service and generates continuous new recording segments according to preset rules. For each new recording segment and its corresponding new time period, it calculates the negative sentiment index and target semantic hit frequency of the corresponding server within the new time period based on the server-side text data in the new recording segment. Based on the server's historical risk data, the negative sentiment index of the current new time period, and the target semantic hit frequency, it continuously constructs a long-term risk profile of the server. Based on the current new recording segment, a preset regulatory compliance semantic knowledge graph, and the demand-side text data in the current new recording segment, it calculates the semantic risk score of the current new recording segment. When the semantic risk score is greater than a preset compliance threshold, it determines the comprehensive risk level of the current new recording segment based on the semantic risk score, the negative sentiment index of the current new time period, and the cumulative risk score of the long-term risk profile. Based on the comprehensive risk level, it determines the corresponding target agent to process the current new recording segment.
[0008] Fourthly, a computer-readable storage medium is provided, which adopts the following technical solution: The system receives real-time dialogue recording data streams from the target service and generates continuous new recording segments according to preset rules. For each new recording segment and its corresponding new time period, it calculates the negative sentiment index and target semantic hit frequency of the corresponding server within the new time period based on the server-side text data in the new recording segment. Based on the server's historical risk data, the negative sentiment index of the current new time period, and the target semantic hit frequency, it continuously constructs a long-term risk profile of the server. Based on the current new recording segment, a preset regulatory compliance semantic knowledge graph, and the demand-side text data in the current new recording segment, it calculates the semantic risk score of the current new recording segment. When the semantic risk score is greater than a preset compliance threshold, it determines the comprehensive risk level of the current new recording segment based on the semantic risk score, the negative sentiment index of the current new time period, and the cumulative risk score of the long-term risk profile. Based on the comprehensive risk level, it determines the corresponding target agent to process the current new recording segment.
[0009] Compared with existing technologies, the embodiments of this application have the following main advantages: By receiving the dialogue recording data stream in real time and splitting it into continuously newly added recording segments, the streaming processing design abandons the traditional manual sampling mode, performs full capture and automated processing of all recording data, overcomes the limitations of human cost and efficiency, achieves full coverage detection of dialogue recordings, and solves the problem of missed detection caused by insufficient coverage of existing technologies. Relying on the regulatory compliance semantic knowledge graph, combined with the correlation analysis of demand-side text data and server-side text data, it not only achieves accurate matching of regulatory terms, but also captures evasion behaviors such as ambiguous expressions and synonym substitutions in complex contexts. At the same time, by calculating the server-side negative sentiment index and the degree of conflict between demands and responses, it accurately identifies the risk of emotion inducement, significantly improving the accuracy of risk identification. By accumulating historical risk data on the server side, a dynamically updated long-term risk profile is constructed. Combined with real-time semantic risk scores and negative sentiment indices, the comprehensive risk level is quantified, and then the corresponding target intelligent agent is matched to complete the automated handling without human intervention, significantly improving the risk handling response speed, achieving timely intervention in high-risk scenarios, effectively reducing the regulatory compliance pressure on enterprises, and strengthening the overall effect of service quality and compliance control. Attached Figure Description
[0010] To more clearly illustrate the solutions in this application, the accompanying drawings used in the description of the embodiments of this application will be briefly introduced below. Obviously, the accompanying drawings described below are some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0011] Figure 1 This is an exemplary system architecture diagram to which this application can be applied; Figure 2 A flowchart of an embodiment of the method for processing recorded conversations according to this application; Figure 3 This is a schematic diagram of the structure of one embodiment of the dialogue recording processing apparatus according to this application; Figure 4 This is a schematic diagram of the structure of one embodiment of the computer device according to this application. Detailed Implementation
[0012] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application pertains; the terminology used herein in the specification of the application is for the purpose of describing particular embodiments only and is not intended to be limiting of the application; the terms "comprising" and "having," and any variations thereof, in the specification, claims, and foregoing drawings of this application, are intended to cover non-exclusive inclusion. The terms "first," "second," etc., in the specification, claims, or foregoing drawings of this application are used to distinguish different objects, not to describe a particular order.
[0013] In this document, the term "embodiment" means that a particular feature, structure, or characteristic described in connection with an embodiment may be included in at least one embodiment of this application. The appearance of this phrase in various places throughout the specification does not necessarily refer to the same embodiment, nor is it a separate or alternative embodiment mutually exclusive with other embodiments. It will be explicitly and implicitly understood by those skilled in the art that the embodiments described herein can be combined with other embodiments.
[0014] To enable those skilled in the art to better understand the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings.
[0015] like Figure 1 As shown, system architecture 100 may include terminal device 101, network 102, and server 103. Terminal device 101 may be a laptop 1011, tablet 1012, or mobile phone 1013. Network 102 is used as a medium to provide a communication link between terminal device 101 and server 103. Network 102 may include various connection types, such as wired, wireless communication links, or fiber optic cables.
[0016] Users can use terminal device 101 to interact with server 103 via network 102 to receive or send messages, etc. Various communication client applications can be installed on terminal device 101, such as web browser applications, shopping applications, search applications, instant messaging tools, email clients, social media platform software, etc.
[0017] Terminal device 101 can be various electronic devices with a display screen and support web browsing. In addition to laptops 1011, tablets 1012, or mobile phones 1013, terminal device 101 can also be e-book readers, MP3 players (Moving Picture Experts Group Audio Layer III), MP4 players (Moving Picture Experts Group Audio Layer IV), laptops, and desktop computers.
[0018] Server 103 can be a server that provides various services, such as a backend server that provides support for the pages displayed on terminal device 101.
[0019] It should be noted that the dialogue recording processing method provided in this application embodiment is generally executed by a server / terminal device, and correspondingly, the dialogue recording processing device is generally located in the server / terminal device.
[0020] It should be understood that Figure 1 The number of terminal devices, networks, and servers shown is merely illustrative. Depending on implementation needs, any number of terminal devices, networks, and servers can be included.
[0021] Continue to refer to Figure 2 A flowchart illustrating an embodiment of a method for processing recorded conversations according to this application is shown. The method for processing recorded conversations includes the following steps: Step S201: Receive the dialogue recording data stream of the target service in real time and generate continuous new recording segments according to preset rules.
[0022] The target service refers to specific service scenarios in regulatory and compliance-sensitive industries such as fintech, insurance, and healthcare, where service providers and customers interact based on specific business needs. Examples include bank sales consultation services for personal financial products, insurance company guidance services for critical illness insurance applications, and online consultation and prescription consultation services from internet hospitals. The recorded dialogue data stream refers to a continuous data sequence composed of voice signals from the customer and service provider, continuously captured and transmitted by audio acquisition equipment during the real-time execution of the target service. This data stream possesses the characteristics of real-time performance, continuity, and completeness, and can fully record the voice information of the entire interaction process. Examples include real-time audio data streams of conversations between an insurance customer service platform and policyholders, and complete voice transmission data sequences of bank customer service hotlines.
[0023] Among them, the preset rules refer to the pre-configured recording segment splitting logic and standards based on the business characteristics of the target service, data processing efficiency requirements, and risk detection granularity requirements. These rules are used to transform continuous, unbounded dialogue recording data streams into discrete units that can be independently analyzed for risk. New recording segments refer to individual audio data units with independent time intervals and complete interactive scenarios generated after splitting the real-time transmitted dialogue recording data stream using the preset rules.
[0024] Step S202: For each newly added recording segment and the corresponding newly added time period, calculate the negative sentiment index and target semantic hit frequency of the corresponding server within the newly added time period based on the server-side text data in the newly added recording segment.
[0025] The newly added time period refers to the actual voice interaction time interval corresponding to a single newly added recording segment, and its start and end times are completely consistent with the start and end times of the audio acquisition of the newly added recording segment. Server-side text data refers to the text information obtained after recognizing, transcribing, and structuring the server's voice signal in the newly added recording segment using speech-to-text technology. It contains all the server's statements during the interaction and is the core data carrier for extracting server-side semantic violations and analyzing emotional states. The server refers to the entity that, in the target service scenario, represents the service provider (such as financial institutions, insurance companies, medical institutions, etc.) in providing services such as business consultation, product recommendations, and problem-solving to customers, including individual customer service personnel and intelligent customer service systems.
[0026] The negative sentiment index refers to a numerical indicator used to quantitatively characterize the intensity of negative emotions on the server side. It is calculated based on server-side text data, using natural language processing techniques to extract emotion-related features (such as negative vocabulary, modal particles, and logical expression), and combined with a pre-defined emotion quantification algorithm (such as a sentiment tendency analysis model). For example, by analyzing expressions like "So annoying," "I've said it so many times," and "Can't you read it yourself?" in the server-side text, the calculated negative sentiment index is 0.85 (ranging from 0 to 1), indicating strong impatience on the server side. The target semantic hit frequency refers to the cumulative number of times specific semantic words related to regulatory taboos and violations are identified by matching server-side text data with a pre-defined regulatory compliance risk semantic database. It is a core indicator quantitatively reflecting the frequency of violations on the server side. For example, in the server-side text of a newly added audio clip, violations such as "zero risk," "guaranteed principal and returns," "absolutely profitable," and "no possibility of loss" are identified a total of 4 times, corresponding to a target semantic hit frequency of 4.
[0027] Step S203: Based on the historical risk data of the server, the negative sentiment index of the current newly added time period, and the frequency of target semantic hits, continuously build a long-term risk profile of the server.
[0028] Historical risk data refers to all risk-related cumulative data stored in the database up to the current newly added time period, associated with the corresponding server-side identifier (such as employee ID). This includes negative sentiment indices and target semantic hit frequencies for each newly added historical time period, as well as historical cumulative risk values obtained through algorithms such as integral calculations and weighted averages based on this basic data. This data is the core foundation for building a long-term risk profile of the server and can reflect the historical risk behavior trends of the server.
[0029] The negative sentiment index and target semantic hit frequency for the newly added time period refer to real-time risk indicators calculated using the corresponding server-side text data for the newly added audio segment to be analyzed, employing negative sentiment quantification and target semantic matching algorithms, respectively. These indicators include the current server-side negative sentiment intensity and the number of instances of violation semantic recognition. This data serves as an immediate input for updating historical risk data and refining the long-term risk profile, reflecting the server's real-time risk status in the current interaction scenario. For example, the server-side negative sentiment index for the currently added time period (2024-05-20 16:00:00-16:02:00) is 0.78, and the target semantic hit frequency is 3 times.
[0030] Among them, the long-term risk profile refers to a dynamically updated server-side risk characteristic model constructed through algorithms such as data fusion, trend analysis, and integral calculation, based on historical risk data from the server and real-time risk indicators for newly added time periods. This profile can comprehensively characterize the server's long-term risk behavior patterns, risk preferences, and trends in risk intensity, providing historical support for risk level determination.
[0031] Step S204: Calculate the semantic risk score of the newly added recording segment based on the newly added recording segment, the preset regulatory compliance semantic knowledge graph, and the demand-side text data in the newly added recording segment.
[0032] The regulatory compliance semantic knowledge graph refers to a structured knowledge system built based on regulatory policies in industries such as fintech, insurance, and healthcare. It includes semantic nodes, semantic relationships, and risk weights related to regulatory compliance. Its core components may include a set of compliance taboo semantics (such as "guaranteed returns" and "zero risk"), mapping relationships between vague expressions and synonyms (such as "principal guaranteed" corresponding to "no possibility of principal loss"), and risk weight assignments for violation semantics (such as a weight of 1.0 for direct violations and a weight of 0.8 for vague violations).
[0033] Among them, demand-side text data refers to the text information obtained by recognizing, transcribing, and structuring the voice signals of customers (demand side) in newly added recording segments using speech-to-text technology. It includes key information such as the customer's business consultation content, expression of demands, and feedback on questions. It is important contextual data for analyzing whether the service-side response matches the customer's demands and whether there is any leading response.
[0034] The semantic risk score refers to a quantitative compliance risk indicator obtained by combining server-side text data and demand-side text data, and by performing multi-level matching between server-side text and regulatory compliance semantic knowledge graphs to calculate the semantic conflict degree between the server-side response and the customer's demands, such as when the customer is concerned about risk but the server promises zero risk. The score is obtained by weighting and summing the risk weights and conflict degree.
[0035] Step S205: When the semantic risk score is greater than the preset compliance threshold, the comprehensive risk level of the newly added recording segment is determined based on the semantic risk score, the negative sentiment index in the current newly added time period, and the cumulative risk score of the long-term risk profile.
[0036] The compliance threshold refers to a semantic risk score threshold preset through statistical analysis and expert review, based on industry regulatory policy requirements, enterprise risk management standards, and historical violation case data. This threshold is used to determine whether newly added audio clips pose a compliance risk requiring intervention. The cumulative risk score is a core quantitative indicator in the long-term risk profile. It is the total risk value obtained by weighted integration of historical risk data from the server and real-time risk indicators for the currently added time period. This score comprehensively reflects the server's long-term risk accumulation level and real-time risk status.
[0037] The comprehensive risk level refers to the risk classification result calculated using a multi-dimensional weighted assessment model when the semantic risk score exceeds the compliance threshold. This model combines the semantic risk score, the negative sentiment index of the current newly added time period, and the cumulative risk score. This level is used to accurately distinguish the severity of risks and provide a basis for matching corresponding handling methods. It can be divided into multiple levels, such as Level 1 risk (minor conflict), Level 2 risk (suspected misleading), and Level 3 risk (serious violation).
[0038] In one embodiment, when the semantic risk score is less than or equal to a preset compliance threshold, it indicates that the newly added audio segment has not triggered a compliance risk warning. The basic information of the newly added audio segment (such as time period and server identifier), the semantic risk score, and related data can be associated and stored with the server's historical risk data for continuous updates to the long-term risk profile. Simultaneously, the segment undergoes routine archiving operations, such as being categorized and stored in the audio database by service type and time dimension. This does not require triggering intervention from the target intelligent agent; the data is simply retained for subsequent periodic compliance audits or service quality retrospective analysis. For example, if a newly added audio segment for a bank's wealth management consultation has a semantic risk score of 0.4 (with a compliance threshold set to 1.0), the system synchronizes the segment's timestamp, server identifier "BK-KF-205," and score of 0.4 to the server's historical risk data pool. After updating the basic data for its long-term risk profile, the audio and text of the segment are archived in the "2024-12 Wealth Management Consultation Compliance Segment" directory, which can be accessed and viewed during monthly compliance audits.
[0039] Step S206: Based on the comprehensive risk level, determine the corresponding target agent so that the newly added recording segment can be processed by the target agent.
[0040] The target intelligent agent refers to a software module or system component with specific automated handling functions, matched to different comprehensive risk levels based on preset handling mapping rules, enabling rapid response and targeted processing of risk scenarios. The functions of the target intelligent agent need to be set according to the risk level, including real-time alerts, manual intervention transfer, violation record archiving, and rectification notification push notifications. For example, for Level 1 risk (minor conflict): it is dispatched to the "automated feedback intelligent agent," which automatically generates an email with a timestamp and the violation statement within 5 minutes after the call ends, sending it to the customer service representative for immediate correction. For Level 2 risk (suspected misleading): it is dispatched to the "customer follow-up intelligent agent," which contacts the customer within 24 hours to reconfirm key information and eliminate the risk. For Level 3 risk (serious violation): it is dispatched to the "legal reporting intelligent agent," which automatically generates a compliance report and notifies legal and senior management.
[0041] This application's embodiments employ a streaming processing design that receives real-time dialogue recording data streams and splits them into continuously added recording segments. This eliminates the traditional manual sampling model, performing full capture and automated processing of all recording data. This overcomes limitations in manpower costs and efficiency, achieving full coverage detection of dialogue recordings and resolving the missed detection issues caused by insufficient coverage in existing technologies. Relying on a regulatory compliance semantic knowledge graph and combining correlation analysis between demand-side and server-side text data, it not only achieves accurate matching of regulatory terms but also captures evasion behaviors such as ambiguous expressions and synonym substitutions in complex contexts. Simultaneously, by calculating the server-side negative sentiment index and the degree of conflict between demands and responses, it accurately identifies emotionally induced risks, significantly improving the accuracy of risk identification. By accumulating historical risk data from the server, a dynamically updated long-term risk profile is constructed. This profile, combined with real-time semantic risk scores and negative sentiment indices, quantifies the comprehensive risk level and then matches it with the corresponding target intelligent agent to complete automated handling without human intervention. This significantly improves the speed of risk handling response, enabling timely intervention in high-risk scenarios, effectively reducing the regulatory compliance pressure on enterprises, and strengthening the overall effect of service quality and compliance control.
[0042] In some optional implementations of this embodiment, step 202, for each newly added recording segment corresponding to a newly added time period, calculates the negative sentiment index and target semantic hit frequency of the corresponding server within the newly added time period based on the server-side text data in the newly added recording segment, specifically including the following steps: For each newly added recording segment, voiceprint separation processing is performed to obtain server-side voice data and demand-side voice data. The server-side voice data and demand-side voice data are then converted into text to obtain server-side text data and demand-side text data, respectively. The server-side text data and demand-side text data are then bound to their corresponding server-side identifiers and demand-side identifiers, respectively, to obtain bound server-side text data and bound demand-side text data. For each newly added recording segment and the corresponding newly added time period, based on the bound server-side text data, the negative sentiment index and target semantic hit frequency of the server-side identifier within the newly added time period are calculated.
[0043] Voiceprint separation processing refers to the process of separating a newly recorded audio segment containing mixed voice data from both the server and client sides. Using voiceprint recognition and speech separation algorithms, and based on differences in speaker voiceprint characteristics such as fundamental frequency, formants, and speech rate, the mixed voice signal is split into independent speech streams belonging to different speakers. Server-side audio data refers to the independent voice signal data belonging to the server side, obtained by voiceprint separation processing from the mixed voice data of the newly recorded audio segment; it is the original data source for server-side text data. Client-side audio data refers to the independent voice signal data belonging to the client side (customer), obtained by voiceprint separation processing from the mixed voice data of the newly recorded audio segment; it is the original data source for client-side text data.
[0044] The server-side identifier is a structured information uniquely identifying the server-side entity. It serves as the core index for linking related server-side data (voice data, text data, risk indicators, historical risk data), ensuring that risky behavior can be traced back to the specific responsible party. The demand-side identifier is a structured information uniquely identifying the demand-side (customer) entity. It serves as the core index for linking related customer data (voice data, text data, request information), facilitating the tracing of a specific customer's interaction history and risk scenarios. Binding refers to the process of establishing a unique mapping relationship between server-side text data and its corresponding server-side identifier, and between demand-side text data and its corresponding demand-side identifier, using data association technology. Its core is to ensure that each piece of text data can be accurately attributed to a specific server or demand-side entity.
[0045] In one example, taking critical illness insurance underwriting consultation service (target service) in the insurance industry as an example, after the system receives the real-time call recording data stream between the customer and the insurance customer service, it generates new recording segments according to the preset rule of "splitting every 2 minutes", such as "call segment from 09:00:00 to 09:02:00 on June 10, 2024". For this segment, a speaker separation model based on deep learning is used for voiceprint separation processing. Based on the differences in voiceprint features such as fundamental frequency and formants between the customer service representative and the customer, the system splits the data into server-side voice data (customer service voice signal) and demand-side voice data (customer voice signal). Using speech-to-text technology based on deep full-sequence convolutional neural networks, the two types of voice data are transcribed into text respectively. The server-side text data is "This critical illness insurance pays out immediately upon diagnosis, there will be absolutely no refusal to pay, zero-risk protection", and the demand-side text data is "I am worried about the difficulty of claims, and I am afraid that I will have no protection after paying the premiums". Subsequently, the server-side text data was bound to the customer service employee ID "BX-KF-086" (server-side identifier), and the demand-side text data was bound to the customer's anonymized mobile phone number "135***7890" (demand-side identifier) to ensure unique data ownership. Based on the bound server-side text data, sentiment-related expressions such as "absolutely not" and "zero risk" were extracted using a sentiment analysis model, resulting in a negative sentiment index of 0.1. Simultaneously, matching with the insurance industry regulatory compliance semantic knowledge graph identified two instances of non-compliant semantics: "zero risk" and "absolutely no refusal to pay," determining the target semantic hit frequency to be 2, providing real-time risk indicators for subsequent long-term risk profiling.
[0046] This application embodiment performs voiceprint separation processing on newly added recording segments to accurately separate server-side and demand-side voice data, avoiding identity confusion caused by mixed speech and laying the foundation for accurate risk attribution in subsequent applications. Speech-to-text technology is used to convert both types of voice data into structured text, ensuring that semantic and emotional features can be quantified and analyzed. By binding text data with corresponding identifiers, a unique association between risk data and responsible entities is achieved, ensuring the accuracy of risk tracing. Based on the bound server-side text data, a quantitative algorithm calculates the negative emotion index and the frequency of target semantic hits, accurately capturing the server-side emotional state risk and quantifying the frequency of violation semantic occurrences. This provides real and accurate real-time risk indicators for building long-term risk profiles, effectively solving the problems of insufficient correlation between speaker identity and risk characteristics and inaccurate risk identification in existing technologies.
[0047] In some optional implementations, step 203 involves continuously building a long-term risk profile of the server based on the server's historical risk data, the negative sentiment index of the current newly added time period, and the frequency of target semantic hits. This specifically includes the following steps: Obtain the historical risk data corresponding to the server. The historical risk data refers to the cumulative value of the negative sentiment index and target semantic hit frequency of all newly added time periods before the current newly added time period, which are stored in association with the server identifier. It also includes the long-term risk profile value obtained by integral calculation based on the cumulative data, up to the end time of the previous newly added time period. The negative sentiment index and target semantic hit frequency of the current newly added time period are superimposed on the historical risk data, and the long-term risk profile of the server is continuously constructed through integral calculation.
[0048] The accumulated value of the cumulative data refers to the sum of the negative sentiment index and target semantic hit frequency for all historical newly added time periods, stored in association with the server-side identifier up to the current newly added time period. This is the foundational data source for building long-term risk profiles. It is obtained by successively summing the risk indicators of each historically added audio segment, directly reflecting the total intensity and frequency of historical risk behaviors on the server side. For example, if a customer service representative's negative sentiment index for the past 10 newly added time periods was 0.3, 0.5, and 0.4, the corresponding accumulated value would be 1.2; and the target semantic hit frequency would be 2, 1, and 3, corresponding to an accumulated value of 6. This value is used for subsequent integration calculations, providing a quantified historical risk base for long-term risk profiles.
[0049] Among them, the long-term risk profile value refers to the risk quantification indicator obtained by integral calculation based on the cumulative value of server-side data, up to the cutoff time of the previous newly added time period. It is the core quantitative representation of the long-term risk profile. It integrates the risk intensity and frequency of all newly added time periods in the history of the server, and can reflect the long-term accumulation degree and trend of risk behavior of the server.
[0050] Among them, integral calculation refers to a mathematical operation that integrates the accumulated value of historical data from the server with the negative sentiment index and target semantic hit frequency of the currently added time period. It is the core algorithm for constructing a long-term risk profile. By assigning weight coefficients to different risk indicators, it performs integral calculations on the sequence of risk indicators changing over time to obtain the updated long-term risk profile value.
[0051] In one example, taking a customer service representative (server identifier "BX-KF-086") in an insurance application consultation service as an example, the process of constructing a long-term risk profile is illustrated. First, the historical risk data corresponding to this customer service representative is obtained, specifically the negative sentiment index E for the two historical time periods prior to the current newly added time period (2025-08-10 10:00-10:02). Negative (τ) represents 0.3 (corresponding to time τ1) and 0.4 (corresponding to time τ2), respectively, and the accumulated value of the accumulated data is 0.3 + 0.4 = 0.7; E Violation (τ) are 2(τ1) and 1(τ2), respectively, and the cumulative value of the accumulated data is 2+1=3. Based on this accumulated data, the long-term risk profile value up to the cutoff time t0 of the previous time period is obtained through integration. The specific integration process is as follows: P Risk (t0)= =(0.4×0.3+0.6×2)+(0.4×0.4+0.6×1)=1.32+0.76=2.08. Then, the negative sentiment index E for the currently added time period is... Negative (τ3)=0.5, target semantic hit frequency E Violation (τ3)=2 is superimposed onto historical risk data, and integration is performed again to obtain the long-term risk profile value at the current time t1: P Risk (t1)=P Risk (t0) + (0.4 × 0.5 + 0.6 × 2) = 2.08 + 1.4 = 3.48. Based on this updated value, the system synchronously updates the customer service representative's long-term risk profile: updating the profile features to "historical average negative sentiment index 0.4, total frequency of target semantic hits 5 times, cumulative risk value 3.48", and adding trend tags "current time period negative sentiment index rising, frequency of violation semantics increasing", thus completing the continuous construction of the long-term risk profile. Subsequently, every time a new recording segment is generated, the system will repeat the above process of "overlaying real-time indicators - integral calculation - updating profile features" to achieve long-term dynamic tracking of server-side risk behavior.
[0052] This application embodiment obtains historical risk data associated with the server identifier, providing a complete data foundation for tracing server-side risk characteristics and avoiding the problem of risk data being disconnected from the subject's identity in existing technologies. Then, the negative sentiment index and target semantic hit frequency of the currently added time period are superimposed onto the historical risk data, and the long-term risk profile is updated through integral calculation. This achieves dynamic accumulation of risk data and iterative profile development based on server-side interactions, accurately depicting long-term risk trends on the server side and establishing a deep correlation between speaker identity and risk characteristics. This addresses the shortcomings of existing technologies, such as vague risk level determination and lack of long-term risk tracking capabilities, providing a reliable basis for subsequent precise risk control.
[0053] In some optional implementations, step S204, based on the newly added recording segment, the preset regulatory compliance semantic knowledge graph, and the demand-side text data in the newly added recording segment, calculates the semantic risk score of the newly added recording segment, specifically including the following steps: The system matches the server-side text data corresponding to the newly added audio segment with a pre-defined regulatory compliance semantic knowledge graph to obtain a server-side sensitive word violation score; identifies risk-related request tags in the demand-side text data corresponding to the newly added audio segment; calculates the contextual conflict degree based on the server-side text data corresponding to the newly added audio segment; calculates the correlation conflict degree between the server-side response and the demand-side request based on the server-side text data corresponding to the newly added audio segment, contextual conflict, and risk-related request tags to obtain the total conflict degree of the newly added audio segment; and weights and sums the server-side sensitive word violation score, the total conflict degree, and the pre-defined weight coefficient to obtain the semantic risk score of the newly added audio segment.
[0054] Among them, the server-side sensitive word violation score refers to the quantitative score calculated by matching the server-side text data of the newly added audio clip with a preset regulatory compliance semantic knowledge graph, based on the risk weight and frequency of the matched sensitive words, such as regulatory taboo expressions and misleading words. This score is the core indicator representing the degree of direct violation of the server-side text. Risk-related request tags refer to the classification tags extracted from the request-side text data through natural language processing technology, which are related to the customer's risk concerns and potential worries. These tags are used to clarify the core direction of the customer's requests during the interaction process, providing a scenario-based basis for judging whether the server-side response is appropriate for the customer's needs.
[0055] Among them, the contextual conflict degree refers to the quantitative conflict index calculated based on the server-side text data of the newly added recording segment, analyzing whether there are contradictions in the server-side expression logic within the same interaction scenario, and whether there are inconsistencies in the semantics before and after, and is used to characterize the logical compliance of the server-side text itself. The correlation conflict degree refers to the quantitative index calculated based on the server-side text data of the newly added recording segment, the contextual conflict degree, and risk-related request tags, analyzing whether the server-side response content deviates from the customer's risk request, whether it fails to address the customer's concerns specifically, or even contains misleading responses, and is used to characterize the suitability risk of the server-side response to the customer's request.
[0056] In one example, taking online consultation services in the medical industry (target service) as an example, the server-side text data corresponding to the newly added audio clip is "This antihypertensive drug can absolutely cure hypertension, and there are no side effects even with long-term use," while the demand-side text data is "I'm worried about the side effects of antihypertensive drugs; can they completely cure my hypertension?" First, the server-side text is matched with a pre-defined medical regulatory compliance semantic knowledge graph. "Absolutely cures" and "no side effects" are considered sensitive words with corresponding risk weights of 1.0 and 0.9 respectively, resulting in a server-side sensitive word violation score of 1.9. Second, the demand-side text is identified using an intent recognition model, generating two risk-related request tags: "side effect concern" and "curative effect consultation." Next, the server-side text is analyzed; there are no contradictory statements, resulting in a contextual conflict degree of 0.1. Combining the server-side text, contextual conflict degree, and request tags, the server fails to address side effect concerns and exaggerates efficacy, resulting in a correlation conflict degree of 0.9. The total conflict degree is 0.1 + 0.9 = 1.0. Finally, the semantic risk score is calculated by weighting the results according to the preset weight coefficients (0.6 for sensitive word violation score and 0.4 for total conflict score). The semantic risk score is 1.9 × 0.6 + 1.0 × 0.4 = 1.54, thus completing the semantic risk quantification of the current segment.
[0057] This application's embodiment obtains sensitive word violation scores by matching server-side text data with a regulatory compliance semantic knowledge graph, accurately capturing explicit violation expressions. It identifies risk-related demand tags on the demand side, clarifying core customer concerns. The contextual conflict degree of the server-side text is calculated, exposing logical contradictions in the expression. Combining the above information to calculate the correlation conflict degree and total conflict degree reveals implicit risks such as "irrelevant answers." Finally, a semantic risk score is obtained through weighted summation using preset weights, achieving multi-dimensional risk quantification. The entire process overcomes the limitations of existing technologies that only match literal words, covering both explicit and implicit violations and relating them to customer demand scenarios, making semantic risk identification more comprehensive and accurate.
[0058] In some optional implementations, the step "matching the server-side text data corresponding to the newly added audio segment with the preset regulatory compliance semantic knowledge graph to obtain the server-side sensitive word violation score" specifically includes the following steps: The server-side text data corresponding to the newly added audio segment is split into multiple semantic units; multiple semantic units are matched with target semantic entities in the preset regulatory compliance semantic knowledge graph to obtain a set of matched sensitive words; the sensitivity weight of each sensitive word in the set of sensitive words in the regulatory compliance semantic knowledge graph is obtained; based on the sensitivity weight corresponding to each sensitive word, the sensitive words in the set of sensitive words are weighted and summed to obtain the server-side sensitive word violation score.
[0059] Here, a semantic unit refers to the smallest linguistic unit with independent expressive function obtained after semantic segmentation of the server-side text data of the newly added audio clip. A target semantic entity refers to a semantic object pre-defined in the regulatory compliance semantic knowledge graph that is directly related to industry regulatory prohibitions (such as misleading statements and illegal promises), and is the core reference standard for determining whether the server-side text violates regulations. Its forms include illegal words, illegal phrases, and illegal expression paradigms, and each entity is associated with clear regulatory basis.
[0060] The sensitive word set refers to the set of all semantic units with a violation association that are selected after matching the semantic units of the server-side text with the target semantic entities in the regulatory compliance semantic knowledge graph. The sensitivity weight refers to the quantitative coefficient that is pre-assigned to each target semantic entity (i.e., sensitive word) in the regulatory compliance semantic knowledge graph, representing the severity of its violation.
[0061] In one example, taking the financial industry's bank wealth management product consultation service as an example, the server-side text data for the newly added audio clip (corresponding to the newly added time period 2025-10-20 09:15-09:17) is "This stable wealth management product guarantees an annualized return of 4.5%, with no risk of principal loss, suitable for all investors." The first step uses a text segmentation algorithm based on semantic dependency analysis to split the server-side text into four semantic units: "this stable wealth management product," "guaranteed annualized return of 4.5%," "no risk of principal loss," and "suitable for all investors." This ensures that each unit has an independent semantic function, avoiding the omission of illegal information due to long sentence matching. The second step performs semantic similarity matching between the above semantic units and target semantic entities (including illegal expressions such as "guaranteed return" and "no risk of principal loss") in a pre-set bank regulatory compliance semantic knowledge graph. "Guaranteed annualized return of 4.5%" matches the "guaranteed return" entity, and "no risk of principal loss" matches the "no risk of principal loss" entity, forming a sensitive word set containing two elements. The third step is to retrieve the sensitivity weights of each element within the sensitive word set from the knowledge graph: "Guaranteed returns" has a sensitivity weight of 0.9 because it violates the rule prohibiting promises of returns; "No risk of principal loss" has a sensitivity weight of 0.8 because it exaggerates the product's safety. The fourth step is to calculate the server-side sensitive word violation score for the current segment using the formula: "Server-side sensitive word violation score = Σ (sensitive word × corresponding sensitivity weight)," i.e., 0.9 + 0.8 = 1.7 (out of 2.0), thus completing the quantification of explicit violation information.
[0062] This application's embodiment breaks down server-side text data into multiple semantic units with independent semantic functions, avoiding the omission of violation information caused by matching long sentences as a whole, thus laying the foundation for accurate matching. The semantic units are then matched with target semantic entities in the regulatory compliance semantic knowledge graph, accurately locating explicit violation expressions and forming a sensitive word set. Subsequently, the sensitivity weights of each sensitive word within the set are retrieved, and differentiated quantification is achieved based on the severity of the violation. Finally, a weighted summation is used to obtain the server-side sensitive word violation score, completing the accurate quantification of the violation severity. This entire process overcomes the limitations of existing technologies that only perform literal matching, covering different types of violation expressions and differentiating the severity of violations through weights, making sensitive word violation identification more accurate and providing a reliable basis for the subsequent calculation of semantic risk scores and explicit violation quantification.
[0063] In some optional implementations, the step "calculating the correlation conflict degree between the server response and the client's request based on the server-side text data, contextual conflicts, and risk-related request tags corresponding to the currently added recording segment, to obtain the total conflict degree of the currently added recording segment" specifically includes the following steps: Semantically match the server-side text data corresponding to the newly added recording segment with the risk-related request tags to determine the degree of contradiction between the server-side request and the corresponding client-side response in the newly added recording segment; calculate the product between the degree of contradiction and the degree of contextual conflict to obtain the degree of related conflict; obtain the preset first weight coefficient and second weight coefficient, and sum the product between the degree of contextual conflict and the first weight coefficient and the product between the degree of related conflict and the second weight coefficient to obtain the total conflict degree.
[0064] Among them, the degree of contradiction refers to the extent to which the server-side response deviates from, contradicts, or fails to address the customer's request by semantically matching the server-side text data of newly added audio clips with risk-related request tags. It is a core indicator for quantifying the risk of mismatch between requests and responses. The correlation conflict degree is used to integrate two types of risks: "request-response mismatch" and the logical consistency of the server-side expression, to comprehensively characterize the compliance deficiencies of the server-side response. This indicator avoids the one-sidedness of single-dimensional risk assessment and makes the quantification of conflict more closely reflect the characteristics of compliance risks in actual interaction scenarios.
[0065] The first weighting coefficient is a pre-set quantitative coefficient used to adjust the proportion of contextual conflict in the total conflict level. Its value needs to be determined in conjunction with industry compliance priorities and historical risk cases to reflect the importance of the logical risks in the server-side expression in the overall risk assessment. The second weighting coefficient is a pre-set quantitative coefficient used to adjust the proportion of related conflict in the total conflict level. Its sum with the first weighting coefficient is 1. Its value is set according to the impact of the risk of mismatch between the request and the response on the customer's rights and interests, reflecting the priority of this type of risk in the overall risk assessment.
[0066] In one example, taking critical illness insurance underwriting consultation services in the insurance industry as an example, the server-side text data for the newly added audio clip is "This critical illness insurance claims are processed very quickly; you can get the money as soon as you are diagnosed, without needing to provide medical records." The demand-side text data is identified with two risk-related demand tags: "claim material requirements" and "claim timeliness." The contextual conflict degree of this clip has been calculated to be 0.7 (the server initially states "no medical records are needed," but later mentions "a diagnosis certificate is required," which is contradictory). The first step uses a semantic similarity matching model to match the server-side text data with the "claim material requirements" and "claim timeliness" demand tags: the server does not explicitly state the claim timeliness, and the statement "no medical records are needed" deviates from the actual demand for required materials, thus determining the conflict degree to be 0.8 (range 0-1). The second step calculates "association conflict degree = conflict degree × contextual conflict degree," resulting in an association conflict degree of 0.8 × 0.7 = 0.56. The third step is to obtain the first weight coefficient (contextual conflict weight) of 0.3 and the second weight coefficient (related conflict weight) of 0.7 preset by the insurance industry. The total conflict degree is calculated according to the formula "total conflict degree = contextual conflict degree × first weight coefficient + related conflict degree × second weight coefficient", that is, the total conflict degree = 0.7 × 0.3 + 0.56 × 0.7 = 0.21 + 0.392 = 0.602, thus completing the quantification of the total conflict degree of the current segment.
[0067] This application's embodiment accurately identifies the degree of contradiction between the server-side response and the customer's request by semantically matching server-side text data with risk-related request tags. The degree of contradiction is then multiplied by the contextual conflict level to obtain the associated conflict level, integrating two types of risks: "request-response fit" and "expression logic," avoiding the one-sidedness of a single-dimensional assessment. Finally, using preset first and second weighting coefficients, the contextual conflict level and the associated conflict level are weighted and summed to obtain the total conflict level, allowing for adjustments to the proportion of the two types of risks based on industry compliance priorities. This entire process overcomes the limitations of existing technologies that only match literal text, comprehensively covering both implicit and explicit conflict risks, making conflict quantification more closely aligned with actual interaction scenarios.
[0068] In some optional implementations, step 206, based on the comprehensive risk level, determines the corresponding target agent to process the newly added audio segment, specifically including the following steps: Based on the newly added audio segment, semantic risk score, corresponding negative sentiment index, and comprehensive risk level, a risk information set for the newly added audio segment is generated. Based on the comprehensive risk level, the corresponding target agent is determined, and the risk information set is sent to the target agent so that the target agent can process the newly added audio segment.
[0069] Among them, the risk information set refers to a structured data set that integrates the newly added recording segment, its corresponding semantic risk score, negative sentiment index and comprehensive risk level, and contains complete risk-related information of the segment. It is the core basis for the target intelligent agent to carry out risk processing.
[0070] In one example, taking health insurance underwriting consultation services in the insurance industry as an example, the newly added audio segment corresponds to the newly added time period of 2025-12-09 15:00-15:02, and the server identifier is "BX-KF-312". Preliminary calculations show that this segment has a semantic risk score of 1.2 (out of 2.0) and a negative sentiment index of 0.3. Combined with the long-term risk profile's cumulative risk score of 2.1, the overall risk level is determined to be "Level 2 Risk (Suspected Misleading)". The first step is to generate a risk information set, integrating data in a structured format, including basic segment information (time period, server identifier), risk indicators (semantic risk score 1.2, negative sentiment index 0.3, cumulative risk score 2.1), the overall risk level "Level 2 Risk", and key text (the server statement "This product provides compensation if you get sick, no need to read the health declaration"). The second step is to match the target agent. According to the preset "risk level - agent" mapping rule, Level 2 risk corresponds to the "customer follow-up agent", therefore the target agent is determined to be the customer follow-up agent, and the risk information set is encrypted and pushed to this agent. The third step involves the customer follow-up intelligent agent receiving the risk information set and automatically generating a follow-up script under the name of "satisfaction survey". It embeds key information such as "health disclosure requirements" and "claim conditions" that need to be confirmed, and triggers a covert follow-up to the demand side within 24 hours to complete the secondary confirmation of risk information and risk elimination.
[0071] This application's embodiments generate a risk information set by integrating newly added audio clips, semantic risk scores, negative sentiment indices, and comprehensive risk levels. This transforms scattered risk data into structured and complete risk data, avoiding processing biases caused by fragmented information. Then, based on the comprehensive risk level, a corresponding target intelligent agent is matched, achieving precise correspondence between risk levels and the responsible party. After the risk information set is sent to the target intelligent agent, the agent can quickly initiate the processing flow based on the complete information, eliminating the need for manual data collection and overcoming the response lag problem caused by reliance on manual intervention in existing technologies.
[0072] It should be emphasized that, in order to further ensure the privacy and security of the aforementioned continuously added audio clips, server-side text data, negative sentiment index, target semantic hit frequency, demand-side text data, and semantic risk score, these data can also be stored in a blockchain node.
[0073] The blockchain referred to in this application is a novel application model of computer technologies such as distributed data storage, peer-to-peer transmission, consensus mechanisms, and encryption algorithms. Essentially, a blockchain is a decentralized database, a chain of data blocks linked together using cryptographic methods. Each data block contains information about a batch of network transactions, used to verify the validity of the information (anti-counterfeiting) and generate the next block. A blockchain can include an underlying blockchain platform, a platform product service layer, and an application service layer.
[0074] The embodiments of this application can acquire and process relevant data based on artificial intelligence technology. Artificial intelligence (AI) refers to the theories, methods, technologies, and application systems that use digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use that knowledge to obtain optimal results.
[0075] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by instructing related hardware through computer-readable instructions. These computer-readable instructions can be stored in a computer-readable storage medium. When the program is executed, it can include the processes of the embodiments of the above methods. The aforementioned storage medium can be a non-volatile storage medium such as a magnetic disk, optical disk, or read-only memory (ROM), or random access memory (RAM).
[0076] Further reference Figure 3 As a response to the above Figure 2 The implementation of the method shown in this application provides an embodiment of a dialogue recording processing apparatus, which is similar to... Figure 2 Corresponding to the method embodiments shown, this device can be specifically applied to various electronic devices.
[0077] like Figure 3 As shown, the dialogue recording processing device 400 of this embodiment includes: a receiving module 401, a first calculation module 402, a construction module 403, a second calculation module 404, a determining module 405, and a processing module 406. Wherein: The receiving module 401 is used to receive the dialogue recording data stream of the target service in real time and generate continuous new recording segments according to preset rules. The first calculation module 402 is used to calculate the negative sentiment index and target semantic hit frequency of the corresponding server in the newly added time period for each newly added recording segment based on the server text data in the newly added recording segment. Module 403 is used to continuously build a long-term risk profile of the server based on the historical risk data corresponding to the server, the negative sentiment index of the current newly added time period, and the frequency of target semantic hits. The second calculation module 404 is used to calculate the semantic risk score of the newly added recording segment based on the newly added recording segment, the preset regulatory compliance semantic knowledge graph, and the demand-side text data in the newly added recording segment. The determination module 405 is used to determine the comprehensive risk level of the newly added recording segment based on the semantic risk score, the negative sentiment index in the current newly added time period, and the cumulative risk score of the long-term risk profile when the semantic risk score is greater than the preset compliance threshold. The processing module 406 is used to determine the corresponding target intelligent agent based on the comprehensive risk level, so as to process the newly added recording segment through the target intelligent agent.
[0078] This application's embodiments employ a streaming processing design that receives real-time dialogue recording data streams and splits them into continuously added recording segments. This eliminates the traditional manual sampling model, performing full capture and automated processing of all recording data. This overcomes limitations in manpower costs and efficiency, achieving full coverage detection of dialogue recordings and resolving the missed detection issues caused by insufficient coverage in existing technologies. Relying on a regulatory compliance semantic knowledge graph and combining correlation analysis between demand-side and server-side text data, it not only achieves accurate matching of regulatory terms but also captures evasion behaviors such as ambiguous expressions and synonym substitutions in complex contexts. Simultaneously, by calculating the server-side negative sentiment index and the degree of conflict between demands and responses, it accurately identifies emotionally induced risks, significantly improving the accuracy of risk identification. By accumulating historical risk data from the server, a dynamically updated long-term risk profile is constructed. This profile, combined with real-time semantic risk scores and negative sentiment indices, quantifies the comprehensive risk level and then matches it with the corresponding target intelligent agent to complete automated handling without human intervention. This significantly improves the speed of risk handling response, enabling timely intervention in high-risk scenarios, effectively reducing the regulatory compliance pressure on enterprises, and strengthening the overall effect of service quality and compliance control.
[0079] In one embodiment, the first computing module 402 includes: The separation submodule is used to perform voiceprint separation processing on each newly added recording segment to obtain server-side audio data and demand-side audio data. The conversion submodule is used to convert server-side audio data and client-side audio data into text respectively, resulting in server-side text data and client-side text data. The binding submodule is used to bind server-side text data and request-side text data to the corresponding server-side identifier and request-side identifier, respectively, to obtain the bound server-side text data and bound request-side text data. The first calculation submodule is used to calculate the negative sentiment index and target semantic hit frequency of the server in the newly added time period corresponding to each newly added audio segment, based on the bound server text data.
[0080] In one embodiment, the construction module 403 includes: The acquisition submodule is used to acquire the historical risk data corresponding to the server. The historical risk data refers to the cumulative value of the negative sentiment index and target semantic hit frequency of all newly added time periods before the current newly added time period, which are stored in association with the server identifier, as well as the long-term risk profile value up to the end time of the previous newly added time period, which is obtained by integral calculation based on the cumulative data. The calculation submodule is used to overlay the negative sentiment index and target semantic hit frequency of the current newly added time period onto the historical risk data, and continuously build a long-term risk profile on the server through integral calculation.
[0081] In one embodiment, the second computing module 404 includes: The matching submodule is used to match the server-side text data corresponding to the newly added audio segment with the preset regulatory compliance semantic knowledge graph to obtain the server-side sensitive word violation score; The identification submodule is used to identify risk-related request tags in the demand-side text data corresponding to the currently added audio clip; The second calculation submodule is used to calculate the context conflict degree based on the server-side text data corresponding to the newly added recording segment. The third calculation submodule is used to calculate the degree of correlation between the server response and the client's request based on the server-side text data, contextual conflicts, and risk-related request tags corresponding to the newly added recording segment, and to obtain the total conflict degree of the newly added recording segment. The summation submodule is used to perform a weighted summation of the server-side sensitive word violation score, total conflict degree, and preset weight coefficient to obtain the semantic risk score of the newly added recording segment.
[0082] In one embodiment, the matching submodule is further configured to split the server-side text data corresponding to the newly added recording segment into multiple semantic units; match the multiple semantic units with target semantic entities in a preset regulatory compliance semantic knowledge graph to obtain a set of matched sensitive words; obtain the sensitivity weight of each sensitive word in the set of sensitive words in the regulatory compliance semantic knowledge graph; and, based on the sensitivity weight corresponding to each sensitive word, perform a weighted summation of the sensitive words in the set of sensitive words to obtain a server-side sensitive word violation score.
[0083] In one embodiment, the third calculation submodule is further configured to semantically match the server-side text data corresponding to the currently added recording segment with the risk-related request tags to determine the degree of contradiction between the server-side request and the corresponding demand side response in the currently added recording segment; calculate the product between the degree of contradiction and the contextual conflict degree to obtain the associated conflict degree; obtain the preset first weight coefficient and second weight coefficient, and sum the product between the contextual conflict degree and the first weight coefficient and the product between the associated conflict degree and the second weight coefficient to obtain the total conflict degree.
[0084] In one embodiment, the processing module 406 includes: The generation submodule is used to generate a risk information set for the newly added audio segment based on the semantic risk score, the corresponding negative sentiment index, and the comprehensive risk level. The processing submodule is used to determine the corresponding target agent based on the comprehensive risk level and send the risk information set to the target agent so that the target agent can process the newly added recording segment.
[0085] To address the aforementioned technical problems, embodiments of this application also provide a computer device. Please refer to [link / reference needed]. Figure 4 , Figure 4 This is a basic structural block diagram of the computer device in this embodiment.
[0086] Computer device 6 includes a memory 61, a processor 62, and a network interface 63 that are interconnected via a system bus. It should be noted that only computer device 6 with memory 61, processor 62, and network interface 63 is shown in the figure; however, it should be understood that it is not required to implement all the components shown, and more or fewer components can be implemented alternatively. Those skilled in the art will understand that the computer device described herein is a device capable of automatically performing numerical calculations and / or information processing according to pre-set or stored instructions, and its hardware includes, but is not limited to, microprocessors, application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), digital signal processors (DSPs), embedded devices, etc.
[0087] Computer devices can include desktop computers, laptops, handheld computers, and cloud servers. These devices allow for human-computer interaction with users through keyboards, mice, remote controls, touchpads, or voice-activated devices.
[0088] The memory 61 includes at least one type of readable storage medium, including flash memory, hard disk, multimedia card, card-type memory (e.g., SD or DX memory), random access memory (RAM), static random access memory (SRAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), programmable read-only memory (PROM), magnetic memory, magnetic disk, optical disk, etc. In some embodiments, the memory 61 may be an internal storage unit of the computer device 6, such as the hard disk or memory of the computer device 6. In other embodiments, the memory 61 may also be an external storage device of the computer device 6, such as a plug-in hard disk, smart media card (SMC), secure digital (SD) card, flash card, etc., equipped on the computer device 6. Of course, the memory 61 may include both the internal storage unit and the external storage device of the computer device 6. In this embodiment, the memory 61 is typically used to store the operating system and various application software installed on the computer device 6, such as computer-readable instructions for processing dialogue recording methods. In addition, the memory 61 may also be used to temporarily store various types of data that have been output or will be output.
[0089] In some embodiments, processor 62 may be a central processing unit (CPU), controller, microcontroller, microprocessor, or other data processing chip. This processor 62 is typically used to control the overall operation of the computer device 6. In this embodiment, processor 62 is used to execute computer-readable instructions stored in memory 61 or to process data, such as computer-readable instructions for executing a conversation recording processing method.
[0090] The network interface 63 may include a wireless network interface or a wired network interface, which is typically used to establish a communication connection between the computer device 6 and other electronic devices.
[0091] This application's embodiments employ a streaming processing design that receives real-time dialogue recording data streams and splits them into continuously added recording segments. This eliminates the traditional manual sampling model, performing full capture and automated processing of all recording data. This overcomes limitations in manpower costs and efficiency, achieving full coverage detection of dialogue recordings and resolving the missed detection issues caused by insufficient coverage in existing technologies. Relying on a regulatory compliance semantic knowledge graph and combining correlation analysis between demand-side and server-side text data, it not only achieves accurate matching of regulatory terms but also captures evasion behaviors such as ambiguous expressions and synonym substitutions in complex contexts. Simultaneously, by calculating the server-side negative sentiment index and the degree of conflict between demands and responses, it accurately identifies emotionally induced risks, significantly improving the accuracy of risk identification. By accumulating historical risk data from the server, a dynamically updated long-term risk profile is constructed. This profile, combined with real-time semantic risk scores and negative sentiment indices, quantifies the comprehensive risk level and then matches it with the corresponding target intelligent agent to complete automated handling without human intervention. This significantly improves the speed of risk handling response, enabling timely intervention in high-risk scenarios, effectively reducing the regulatory compliance pressure on enterprises, and strengthening the overall effect of service quality and compliance control.
[0092] This application also provides another embodiment, namely, a computer-readable storage medium storing computer-readable instructions that can be executed by at least one processor to cause the at least one processor to perform the steps of the dialogue recording processing method described above.
[0093] This application's embodiments employ a streaming processing design that receives real-time dialogue recording data streams and splits them into continuously added recording segments. This eliminates the traditional manual sampling model, performing full capture and automated processing of all recording data. This overcomes limitations in manpower costs and efficiency, achieving full coverage detection of dialogue recordings and resolving the missed detection issues caused by insufficient coverage in existing technologies. Relying on a regulatory compliance semantic knowledge graph and combining correlation analysis between demand-side and server-side text data, it not only achieves accurate matching of regulatory terms but also captures evasion behaviors such as ambiguous expressions and synonym substitutions in complex contexts. Simultaneously, by calculating the server-side negative sentiment index and the degree of conflict between demands and responses, it accurately identifies emotionally induced risks, significantly improving the accuracy of risk identification. By accumulating historical risk data from the server, a dynamically updated long-term risk profile is constructed. This profile, combined with real-time semantic risk scores and negative sentiment indices, quantifies the comprehensive risk level and then matches it with the corresponding target intelligent agent to complete automated handling without human intervention. This significantly improves the speed of risk handling response, enabling timely intervention in high-risk scenarios, effectively reducing the regulatory compliance pressure on enterprises, and strengthening the overall effect of service quality and compliance control.
[0094] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods of the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions to cause a terminal device (which may be a mobile phone, computer, server, air conditioner, or network device, etc.) to execute the methods of the various embodiments of this application.
[0095] Obviously, the embodiments described above are only some embodiments of this application, not all embodiments. The accompanying drawings show preferred embodiments of this application, but do not limit the patent scope of this application. This application can be implemented in many different forms; rather, the purpose of providing these embodiments is to provide a more thorough and comprehensive understanding of the disclosure of this application. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art can still modify the technical solutions described in the foregoing specific embodiments, or make equivalent substitutions for some of the technical features. Any equivalent structures made using the content of this application's specification and drawings, directly or indirectly applied to other related technical fields, are similarly within the scope of patent protection of this application.
[0096] The software tools or components not belonging to our company that appear in the embodiments of this application are merely examples and do not represent actual use.
Claims
1. A method for processing recorded conversations, characterized in that, Includes the following steps: It receives the dialogue recording data stream of the target service in real time and generates continuous new recording segments according to preset rules; For each newly added recording segment corresponding to a new time period, based on the server-side text data in the newly added recording segment, calculate the negative sentiment index and target semantic hit frequency of the corresponding server within the newly added time period; Based on the historical risk data corresponding to the server, the negative sentiment index of the current newly added time period, and the frequency of target semantic hits, a long-term risk profile of the server is continuously constructed. Based on the newly added recording segment, the preset regulatory compliance semantic knowledge graph, and the demand-side text data in the newly added recording segment, calculate the semantic risk score of the newly added recording segment. When the semantic risk score is greater than the preset compliance threshold, the comprehensive risk level of the newly added recording segment is determined based on the semantic risk score, the negative sentiment index in the current newly added time period, and the cumulative risk score of the long-term risk profile. Based on the comprehensive risk level, a corresponding target agent is determined so that the newly added audio segment can be processed by the target agent.
2. The method according to claim 1, characterized in that, The step of calculating the negative sentiment index and target semantic hit frequency of the corresponding server within the newly added time period for each newly added recording segment, based on the server-side text data in the newly added recording segment, specifically includes: Each newly added recording segment is subjected to voiceprint separation processing to obtain server-side audio data and demand-side audio data; The server-side audio data and the demand-side audio data are converted into text respectively to obtain server-side text data and demand-side text data; The server-side text data and the demand-side text data are bound to the corresponding server-side identifier and demand-side identifier, respectively, to obtain the bound server-side text data and the bound demand-side text data. For each newly added recording segment corresponding to a new time period, based on the bound server-side text data, the negative sentiment index and target semantic hit frequency of the server-side representation within the newly added time period are calculated.
3. The method according to claim 2, characterized in that, The step of continuously constructing a long-term risk profile of the server based on the historical risk data corresponding to the server, the negative sentiment index of the current newly added time period, and the frequency of target semantic hits specifically includes: Obtain the historical risk data corresponding to the server. The historical risk data refers to the cumulative value of the negative sentiment index and target semantic hit frequency of all newly added time periods before the current newly added time period, which are stored in association with the server identifier, and the long-term risk profile value up to the end time of the previous newly added time period, which is obtained by integral calculation based on the cumulative data. The negative sentiment index and target semantic hit frequency of the newly added time period are superimposed on the historical risk data, and the long-term risk profile of the server is continuously constructed through integral calculation.
4. The method according to claim 1, characterized in that, The step of calculating the semantic risk score of the newly added recording segment based on the newly added recording segment, the preset regulatory compliance semantic knowledge graph, and the demand-side text data in the newly added recording segment specifically includes: The server-side text data corresponding to the newly added audio segment is matched with the preset regulatory compliance semantic knowledge graph to obtain the server-side sensitive word violation score. Identify risk-related request tags in the demand-side text data corresponding to the newly added audio segment; Calculate the context conflict degree based on the server-side text data corresponding to the newly added audio segment; Based on the server-side text data corresponding to the newly added recording segment, the contextual conflict, and the risk-related request tags, the correlation conflict degree between the server-side response and the request from the demand side is calculated to obtain the total conflict degree of the newly added recording segment. The semantic risk score of the newly added recording segment is obtained by weighting and summing the server-side sensitive word violation score, the total conflict degree, and the preset weight coefficient.
5. The method according to claim 4, characterized in that, The step of matching the server-side text data corresponding to the newly added recording segment with a preset regulatory compliance semantic knowledge graph to obtain the server-side sensitive word violation score specifically includes: The server-side text data corresponding to the newly added audio segment is split into multiple semantic units; The multiple semantic units are matched with target semantic entities in a preset regulatory compliance semantic knowledge graph to obtain a set of matched sensitive words; Obtain the sensitivity weight of each sensitive word in the set of sensitive words in the regulatory compliance semantic knowledge graph; Based on the sensitivity weight corresponding to each sensitive word, the sensitive words in the sensitive word set are weighted and summed to obtain the server-side sensitive word violation score.
6. The method according to claim 4, characterized in that, The step of calculating the correlation conflict degree between the server response and the client's request based on the server-side text data corresponding to the newly added recording segment, the contextual conflict, and the risk-related request tags, to obtain the total conflict degree of the newly added recording segment, specifically includes: Semantically match the server-side text data corresponding to the newly added recording segment with the risk-related request tags to determine the degree of contradiction between the server-side request and the corresponding client-side response in the newly added recording segment. The product of the degree of contradiction and the degree of contextual conflict is calculated to obtain the degree of association conflict. Obtain the preset first weight coefficient and second weight coefficient, and sum the product of the context conflict degree and the first weight coefficient and the product of the association conflict degree and the second weight coefficient to obtain the total conflict degree.
7. The method according to claim 1, characterized in that, The step of determining the corresponding target agent based on the comprehensive risk level, and processing the newly added audio segment through the target agent, specifically includes: Based on the newly added audio segment, the semantic risk score, the corresponding negative sentiment index, and the comprehensive risk level, a risk information set for the newly added audio segment is generated. Based on the comprehensive risk level, a corresponding target agent is determined, and the risk information set is sent to the target agent so that the target agent can process the newly added recording segment.
8. A processing apparatus for recording conversations, characterized in that, include: The receiving module is used to receive the dialogue recording data stream of the target service in real time and generate continuous new recording segments according to preset rules. The first calculation module is used to calculate the negative sentiment index and target semantic hit frequency of the corresponding server in the newly added time period for each newly added recording segment, based on the server text data in the newly added recording segment. The construction module is used to continuously build a long-term risk profile of the server based on the historical risk data corresponding to the server, the negative sentiment index of the current newly added time period, and the frequency of target semantic hits. The second calculation module is used to calculate the semantic risk score of the newly added recording segment based on the newly added recording segment, the preset regulatory compliance semantic knowledge graph, and the demand-side text data in the newly added recording segment. The determination module is used to determine the comprehensive risk level of the newly added recording segment based on the semantic risk score, the negative sentiment index in the current newly added time period, and the cumulative risk score of the long-term risk profile when the semantic risk score is greater than a preset compliance threshold. The processing module is used to determine the corresponding target agent based on the comprehensive risk level, so as to process the newly added recording segment through the target agent.
9. A computer device, characterized in that, The device includes a memory and a processor, wherein the memory stores computer-readable instructions, and the processor executes the computer-readable instructions to implement the steps of the dialogue recording processing method as described in any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer-readable instructions, which, when executed by a processor, implement the steps of the dialogue recording processing method as described in any one of claims 1 to 7.