Interactive risk warning system, method, equipment and media based on speech recognition and large language model
By using an interactive risk warning system based on speech recognition and large language models, voice information is converted and classified in real time. Combined with emotion and scene analysis, the system solves the problems of untimely and inaccurate risk warning in existing technologies, and achieves efficient risk warning in the financial and medical fields.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- CHINA PING AN PROPERTY INSURANCE CO LTD
- Filing Date
- 2026-03-05
- Publication Date
- 2026-05-26
AI Technical Summary
In the financial and medical fields, existing technologies for voice recognition are slow to respond and struggle to accurately grasp the complex and ever-changing emotions and true intentions of customers or patients, resulting in poor risk warning effects.
An interactive risk warning system based on speech recognition and a large language model is adopted, including a speech recognition module, a semantic analysis module, a risk prediction module, and a feedback module. It collects and converts speech information into text in real time, performs emotion and scene classification, combines user historical data to predict risks, and generates timely risk warning strategies.
It enables timely and accurate risk warnings in financial and medical scenarios, accurately grasps the emotions and intentions of customers or patients, and improves the accuracy of risk identification and the overall warning effect.
Smart Images

Figure CN122090849A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of artificial intelligence technology, and can be specifically applied to the financial and medical fields. In particular, it relates to an interactive risk warning system, method, device and medium based on speech recognition and large language model. Background Technology
[0002] In recent years, consumers' expectations for service experience have continued to rise, and their awareness of their rights has also significantly increased. For example, in the financial sector, insurance and banking institutions frequently encounter complaint risks in their customer service processes. Customers express their dissatisfaction through various means such as telephone, online platforms, and regulatory channels. If these issues are not handled promptly, they can easily lead to regulatory penalties, reputational damage, and legal disputes. Similarly, in the medical field, if patients' demands regarding service quality and medical outcomes are not properly addressed during their medical treatment, they will also experience dissatisfaction, which may then lead to doctor-patient conflicts and disputes. Therefore, building an efficient and accurate interactive risk early warning system is of great significance for the financial and medical industries to improve service quality, maintain a good image, and avoid potential risks.
[0003] In response to the need for interactive risk warnings, some companies in the financial sector have introduced Natural Language Processing (NLP) and machine learning technologies to analyze customer feedback in text form, attempting to identify potential complaint risks. In the medical field, some systems are also attempting to perform sentiment analysis on patients' textual feedback to determine their emotional state and potential dissatisfaction. Simultaneously, they have incorporated manual review, keyword matching, or text classification models to try and capture risk signals from customer feedback.
[0004] However, the applicant recognizes that the relevant technology has at least the following technical problems in its implementation: Whether it's customer phone feedback in financial scenarios or patient voice communication in medical scenarios, the slow response of voice recognition seriously affects the timeliness of risk warnings. Moreover, the granularity of emotion and appeal recognition is coarse, making it difficult to accurately grasp the complex and ever-changing emotions, specific scenarios, and true intentions of customers or patients, resulting in inaccurate risk identification and poor overall risk warning effect. Summary of the Invention
[0005] In view of this, the present invention provides an interactive risk warning system, method, device and medium based on speech recognition and large language model. The main purpose is to solve the problem that it is currently difficult to accurately grasp the complex and ever-changing emotions, specific scenarios and true intentions of customers or patients, resulting in inaccurate risk identification and poor overall risk warning effect.
[0006] According to a first aspect of the present invention, an interactive risk warning system based on speech recognition and a large language model is provided, the interactive risk warning system comprising a speech recognition module, a semantic analysis module, a risk prediction module, and a feedback module: The speech recognition module collects the user's input speech information in real time and converts the speech information into text information in real time. The semantic analysis module performs sentiment classification and scene classification on the text information to obtain semantic analysis results; The risk prediction module performs interactive risk prediction based on the semantic analysis results and the user's historical data, and obtains the risk prediction result. When the risk prediction result reaches the preset warning standard, the feedback module obtains the risk warning strategy generated based on the large language model, and pushes the risk prediction result and the risk warning strategy to the risk handling party for risk handling. The risk warning strategy is matched in the preset strategy library by the large language model after identifying the text information to determine the level of interaction abnormality and the cause of the abnormality.
[0007] According to a second aspect of the present invention, an interaction risk warning method based on speech recognition and a large language model is provided, the method comprising: Real-time acquisition of user-input voice information, and real-time conversion of the voice information into text information; The text information is subjected to sentiment classification and scene classification to obtain semantic analysis results; Based on the semantic analysis results and the user's historical data, interaction risk prediction is performed to obtain the risk prediction result. If the risk prediction result reaches the preset warning standard, a risk warning strategy generated based on a large language model is obtained, and the risk prediction result and the risk warning strategy are pushed to the risk handling party for risk handling. The risk warning strategy is matched in a preset strategy library by the large language model after identifying the text information to determine the level of interaction abnormality and the cause of the abnormality.
[0008] According to a third aspect of the present invention, an apparatus is provided, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the method steps performed by any module of the first aspect described above.
[0009] According to a fourth aspect of the present invention, a medium is provided having a computer program stored thereon, wherein the computer program, when executed by a processor, implements the method steps performed by any module of the first aspect described above.
[0010] By employing the above technical solutions, this invention provides an interactive risk warning system, method, device, and medium based on speech recognition and a large language model. In this invention, whether it is customer telephone feedback in a financial scenario or patient voice communication in a medical scenario, it can be converted into text in real time and classified according to emotion and scenario. The response is rapid, ensuring the timeliness of risk warning. Moreover, by classifying emotion and scenario simultaneously, it can accurately grasp the complex and ever-changing emotions and true intentions of customers or patients, further improving the accuracy of risk identification and resulting in a better overall risk warning effect.
[0011] The above description is merely an overview of the technical solution of the present invention. In order to better understand the technical means of the present invention and to implement it in accordance with the contents of the specification, and in order to make the above and other objects, features and advantages of the present invention more apparent and understandable, specific embodiments of the present invention are described below. Attached Figure Description
[0012] Various other advantages and benefits will become apparent to those skilled in the art upon reading the following detailed description of preferred embodiments. The accompanying drawings are for illustrative purposes only and are not intended to limit the invention. Furthermore, the same reference numerals denote the same parts throughout the drawings. In the drawings: Figure 1 This is a schematic diagram of an application environment for an interactive risk warning system based on speech recognition and a large language model, according to one embodiment of the present invention. Figure 2 This is a schematic diagram of the system architecture of an interactive risk warning system based on speech recognition and a large language model in one embodiment of the present invention; Figure 3 This is a schematic diagram of the training process of a risk prediction model in one embodiment of the present invention; Figure 4 This is another system architecture diagram of an interactive risk warning system based on speech recognition and a large language model, according to one embodiment of the present invention; Figure 5 This is a flowchart illustrating the risk warning strategy generation process in one embodiment of the present invention; Figure 6 This is a flowchart illustrating an interactive risk warning method based on speech recognition and a large language model in one embodiment of the present invention. Figure 7 This is a schematic diagram of the structure of a computer device according to an embodiment of the present invention; Figure 8 This is another structural schematic diagram of a computer device according to one embodiment of the present invention. Detailed Implementation
[0013] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0014] The interactive risk warning system based on speech recognition and large language models provided in this invention can be applied to, for example... Figure 1 In this application environment, the client communicates with the server via the network. The interactive risk warning system is implemented through the server, which can collect user-input voice information in real time and convert it into text information in real time. The server performs sentiment and scenario classification on the text information to obtain semantic analysis results. Based on the semantic analysis results and the user's historical data, the server performs interactive risk prediction to obtain risk prediction results. If the risk prediction results reach the preset warning criteria, the server obtains the risk warning strategy generated based on the large language model and pushes the risk prediction results and risk warning strategy to the risk handling party for risk handling.
[0015] In this invention, both customer telephone feedback in financial scenarios and patient voice communication in medical scenarios can be converted into text in real time and categorized by emotion and context. This rapid response ensures timely risk warnings. Furthermore, the simultaneous categorization of emotion and context allows for a precise understanding of the complex and ever-changing emotions and true intentions of customers or patients, further improving the accuracy of risk identification and resulting in a better overall risk warning effect. The client can be, but is not limited to, various personal computers, laptops, smartphones, tablets, and portable wearable devices. The server can be implemented using a dedicated server or a server cluster consisting of multiple servers. The invention will be described in detail below through specific embodiments.
[0016] Please see Figure 2 As shown, Figure 2 A schematic diagram of the system architecture of an interactive risk warning system based on speech recognition and a large language model provided in an embodiment of the present invention includes a speech recognition module 101, a semantic analysis module 102, a risk prediction module 103, and a feedback module 104. The speech recognition module 101 collects the user's input speech information in real time and converts the speech information into text information in real time.
[0017] The interactive risk warning system based on speech recognition and a large language model provided by this invention can be applied to intelligent customer service or intelligent assistant engines in various application scenarios. Intelligent interaction engines are typically implemented through a server-side component, which can collect user-input voice information. Specifically, the speech recognition module 101 collects voice information from users in real time through channels such as telephone, APP (Application) voice input, and online customer service voice messages, and uses Automatic Speech Recognition (ASR) technology to instantly convert these voice signals into structured text to obtain text information. Automatic Speech Recognition is a technology that converts the lexical content of human speech into computer-readable text. Its core includes an acoustic model, a language model, and a decoder, and it can adapt to voice input in multilingual, multi-dialect, and high-noise environments. In the interactive risk warning system, an ASR engine can be configured and deployed on the server-side, supporting high concurrency and low latency speech-to-text capabilities, ensuring that user speech is converted into text within seconds.
[0018] In this way, the speech recognition module 101 not only improves speech processing efficiency but also ensures the input quality for subsequent semantic analysis. Whether in a noisy car insurance claim scene or a quiet but fast-paced medical consultation, the interactive risk warning system can stably and accurately convert the original speech into analyzable text, providing a reliable data foundation for risk warning. For example, in a financial scenario, if a customer calls an insurance company hotline complaining, "It's been three days and no one has contacted me to process my claim," the interactive risk warning system can quickly convert their speech into corresponding text. In a medical scenario, if a patient describes to the intelligent consultation assistant, "I feel worse after taking the medicine," the interactive risk warning system can also generate clear text in real time for subsequent analysis.
[0019] In this embodiment of the invention, optionally, the speech recognition module 101 collects user-inputted speech information in real time and converts the speech information into text information in real time, including: collecting user-inputted speech information from multiple channels based on a speech access channel, wherein the speech access channel is constructed by deploying a cloud processor resource pool; performing frame-by-frame processing on the speech information according to a preset audio segmentation standard to obtain multiple frames of audio information, and extracting multi-dimensional audio feature vectors of each frame of audio information and performing normalization processing to generate acoustic feature vectors for each frame of audio information, so as to obtain multiple acoustic feature vectors corresponding to multiple frames of audio information; using multiple acoustic feature vectors to generate a continuous acoustic feature sequence, inputting the acoustic feature sequence into a speech recognition model to obtain an initial recognition result, wherein the speech recognition model is obtained by optimizing the loss function using the currently targeted industry-related technical terms; and combining a pre-trained language model and a weighted finite state converter to dynamically decode and optimize the initial recognition result to obtain text information, wherein the language model is a model constructed based on the fusion of n-gram and bidirectional LSTM.
[0020] In this embodiment, the speech recognition module 101 utilizes a speech access channel built upon a cloud-based processor resource pool to collect speech information input by users through various channels in real time. This speech access channel supports multiple access methods, including dedicated telephone lines (such as customer service hotlines), in-app voice input, and smart terminal voice interaction. All speech streams are aggregated through a unified cloud-based ASR service entry point, ensuring efficient access and low-latency processing of multi-source heterogeneous speech data. The cloud-based processor resource pool can employ an elastic scaling architecture, dynamically allocating GPU computing resources based on concurrent requests to ensure speech processing stability under high load scenarios.
[0021] After receiving the raw audio signal, the speech recognition module 101 performs frame segmentation according to a preset audio segmentation standard, specifically dividing the audio into frames every 25 milliseconds to form a continuous audio frame sequence. For each audio frame, a multi-dimensional audio feature vector is extracted. Specifically, Mel-Frequency Cepstral Coefficients (MFCCs) can be used as the core feature. This feature can effectively capture the energy distribution changes of human voice in different frequency bands. In this embodiment, the dimension is set to 13 dimensions, and each feature dimension is normalized to eliminate amplitude differences caused by recording equipment, environmental noise, and other factors, thereby improving the robustness of the subsequent model. After processing, each audio frame generates a standardized acoustic feature vector. Multiple vectors are concatenated to form a continuous acoustic feature sequence, which serves as the input to the speech recognition model.
[0022] The speech recognition model is trained based on a deep neural network and has been customized and optimized specifically for professional terms in industries such as finance and healthcare. During training, a loss function weight adjustment mechanism related to industry terminology is introduced, enabling the model to achieve higher accuracy in recognizing professional terms such as "claims," "original parts," "medical insurance reimbursement," and "prescription drugs." The speech recognition module 101 inputs acoustic feature sequences into the speech recognition model to obtain initial recognition results, and then further combines a pre-trained language model and a Weighted Finite-State Transducer (WFST) for dynamic decoding optimization. The language model integrates an n-gram statistical model and a Bidirectional Long Short-Term Memory (BiLSTM) network. The former captures local statistical patterns in word order, while the latter models contextual dependencies, enhancing the understanding of semantic coherence. The WFST encodes the language model, pronunciation dictionary, and grammatical constraints into a unified state graph structure, enabling efficient decoding path search and significantly improving the end-word recognition rate and overall recognition fluency.
[0023] In this way, the above process not only achieves unified access and high-precision transcription of cross-channel voice communication, but also significantly improves the accuracy of professional terminology recognition and the semantic integrity of long sentences through industry adaptation and deep decoding optimization. This is especially suitable for complex and technically-based scenarios such as financial complaints and medical consultations. For example, in a financial scenario, a customer calls the insurance customer service hotline saying, "I reported the original parts last time, but you replaced them with aftermarket parts." The interactive risk warning system accurately identifies key terms such as "original parts" and "aftermarket parts" and correctly transcribes them. In a medical scenario, a patient says during a consultation, "I've been taking aspirin recently, but my blood pressure is still unstable." The interactive risk warning system successfully identifies the drug name and symptom description, avoiding misjudgments due to confusion of terminology.
[0024] The semantic analysis module 102 performs sentiment classification and scene classification on the text information to obtain semantic analysis results.
[0025] The semantic analysis module 102 receives text information output from the speech recognition module 101. First, it uses a pre-trained language model based on BERT (Bidirectional Encoder Representations from Transformers) to perform deep semantic understanding of the text, while simultaneously performing sentiment classification and scene classification tasks. Sentiment classification aims to determine the user's emotional state, such as positive, neutral, or negative, and further subdivides it into levels such as "mild dissatisfaction" and "severely negative." Scene classification identifies the user's current business stage, such as "claims-payment timeliness," "insurance application-service notification," or "medical visit-medication feedback." BERT is a neural network architecture that captures deep semantic relationships between words through bidirectional contextual modeling, suitable for processing implicit emotions and intentions in short texts. The semantic analysis module 102 uses a joint modeling approach to synchronously output sentiment and scene labels within the same inference process, avoiding information loss caused by traditional serial processing.
[0026] In this way, based on the semantic analysis module 102, the system can accurately capture the intertwined emotions and business demands of users in complex interactions, significantly improving the ability to understand their true intentions. For example, in a financial scenario, if a customer says, "You promised the money would arrive in 24 hours, but it's been two days now!", the interaction risk warning system can identify "severely negative" emotions and the "claims-payment timeliness" scenario. In a medical scenario, if a patient says, "The medicine I was prescribed last time is completely ineffective, and my symptoms have even worsened," the interaction risk warning system can determine this as "negative emotions" and categorize it under the "follow-up visit-questioning efficacy" scenario.
[0027] In this embodiment of the invention, optionally, the semantic analysis module 102 performs sentiment classification and scene classification on the text information to obtain semantic analysis results, including: acquiring a pre-trained semantic recognition model, which is based on a multi-level neural network structure and is obtained through secondary pre-training using multiple historical dialogue data; performing named entity recognition on the text information based on the semantic recognition model to obtain multiple entity information, and dividing the multiple entity information into sentiment-based entity information groups and scene-based entity information groups; performing sentiment classification on each entity information included in the sentiment-based entity information group according to a preset multi-level sentiment classification system to obtain sentiment classification results, and simultaneously performing scene classification on each entity information included in the scene-based entity information group according to a preset multi-level scene classification system to obtain scene classification results; and integrating the sentiment classification results and scene classification results to obtain semantic analysis results.
[0028] In this embodiment, the semantic analysis module 102 performs deep semantic analysis on the text information generated by speech recognition based on a pre-trained semantic recognition model. This model uses BERT-base-chinese (a 12-layer Transformer architecture) as its core framework and has undergone secondary pre-training using multiple historical dialogue datasets, significantly enhancing its ability to understand professional expressions in the financial and medical fields and colloquial language used by users. BERT is a pre-trained language model based on bidirectional context modeling, capable of capturing the dependencies between words in a sentence, thus more accurately understanding complex semantic structures. In practical applications, the model first performs Named Entity Recognition (NER) on the input text, extracting entity information with specific meanings from the text, such as "claim amount," "service personnel number," "drug name," and "department." Named entity recognition can be implemented using the BERT-Softmax model, combining the output of BERT with a classification layer to predict the category of each word or phrase, supporting multi-granularity entity extraction in Chinese and ensuring that no key information is missed.
[0029] The extracted entity information is automatically divided into two categories: emotional entity information groups (such as emotion-related words like "complaint," "dissatisfaction," and "rejection") and scenario entity information groups (such as business scenario keywords like "reimbursement process," "follow-up appointment," and "policy change"). Subsequently, the semantic analysis module 102 classifies emotional entities item by item according to a preset multi-level emotional classification system. For example, "very angry" is classified as "severely negative," and "a little anxious" as "mildly negative," with an emotional intensity quantification mechanism introduced. Simultaneously, scenario entities are matched according to the multi-level scenario classification system, such as classifying "medical insurance reimbursement failure" as a "medical-reimbursement anomaly" scenario and "delayed payment" as a "financial-claims timeliness" scenario. The classification process can integrate a multi-type scenario knowledge base and a multi-type emotional appeal combination rule base built internally by the enterprise, supporting accurate mapping across domains and multiple levels. Finally, the emotional classification results and scenario classification results are integrated into a unified semantic analysis result, forming a structured output that includes emotional level, core appeal, and the relevant business scenario.
[0030] Thus, based on the semantic analysis module 102, by combining large-scale industry corpus training with a fine-grained classification system, the system achieves synchronous and accurate identification of emotions and scenarios, avoiding the risk of misjudgment caused by single-dimensional analysis. This is especially suitable for service interaction scenarios where users have complex emotions and express themselves implicitly. For example, in a financial scenario, if a customer says, "I clearly paid on time, but you say I owe money, which is unreasonable," the interaction risk warning system can identify "owe money" as a scenario entity and "unreasonable" as an emotional entity, and determine it as a scenario of "severe negative emotions + claims dispute." In a medical scenario, if a patient says, "The medicine I was prescribed last time didn't work after three days, and now I'm dizzy," the interaction risk warning system can extract entities such as "ineffective medicine" and "dizziness," and determine it as a scenario of "mild negative emotions + doubts about efficacy," providing a clear basis for subsequent risk intervention.
[0031] The risk prediction module 103 performs interactive risk prediction based on semantic analysis results and the user's historical data, and obtains the risk prediction results.
[0032] The risk prediction module 103, based on the sentiment and scenario tags output by the semantic analysis module 102, and combined with the user's historical data (including past complaint records, policy status, medical history, service interaction frequency, etc.), constructs a multi-dimensional feature vector. This vector is then input into a high-performance gradient boosting tree model such as XGBoost (eXtreme Gradient Boosting, an optimized distributed gradient boosting library) to calculate the probability of a complaint or dispute risk in the current interaction. XGBoost is an ensemble learning algorithm that excels at handling structured features and nonlinear relationships, maintaining high prediction accuracy even with high-dimensional sparse data. The interaction risk warning system can set a dynamic warning threshold (e.g., complaint probability ≥ 80%). Once the prediction result exceeds this threshold, it is determined to be a high-risk interaction event.
[0033] In this way, based on the risk prediction module 103, not only is the immediate emotion of the current conversation considered, but also long-term behavioral trajectories are integrated, achieving a leap from "single-point emotional outburst" to "full life-cycle risk profile," thus improving the foresight and accuracy of early warnings. For example, in a financial scenario, a customer with two previous complaint records calls again to express dissatisfaction with the assessed loss amount. The interactive risk warning system, combining historical data with current emotions, quickly determines it as high-risk. In a medical scenario, a chronic disease patient expresses strong dissatisfaction for the first time after multiple follow-up visits. The interactive risk warning system, combining the frequency of visits with the current negative emotions, triggers a risk warning.
[0034] In this embodiment of the invention, optionally, the risk prediction module 103 performs interaction risk prediction based on semantic analysis results and the user's historical data to obtain risk prediction results, including: acquiring the user's historical data, which includes user risk level data and interaction duration data; extracting keywords from the user's historical data using a word segmentation algorithm, and constructing a co-occurrence matrix based on the extracted keywords, the co-occurrence matrix being used to indicate the frequency and association of different keywords appearing simultaneously in historical abnormal cases; constructing a multi-dimensional feature vector based on the semantic analysis results and the co-occurrence matrix, inputting the multi-dimensional feature vector into a pre-constructed risk prediction model for risk prediction, and obtaining the risk prediction results output by the risk prediction model indicating the probability of anomalies occurring, wherein the risk prediction model is trained using multiple abnormal samples, each abnormal sample including sample voice information and sample historical abnormal cases that are associated with the sample voice information.
[0035] In this embodiment, the risk prediction module 103 constructs a multi-dimensional feature representation by fusing semantic analysis results with user historical data to achieve dynamic risk assessment of current interactive behavior. The risk prediction module 103 first acquires the user's historical data, including user risk level data (such as high, medium, and low risk labels based on underwriting scores or complaint frequency), interaction duration data (such as the duration of a single call), and further, it can acquire speech volume ratio data (i.e., the proportion of time the user actively speaks in the conversation, usually in the range of 0.4 to 0.6, the higher the ratio, the more intense the emotion). This data serves as a baseline for user behavior, providing background support for subsequent risk modeling.
[0036] Subsequently, the risk prediction module 103 uses the TF-IDF (term frequency–inverse document frequency) algorithm to extract keywords from users' historical complaint texts, identifying frequently occurring complaint terms such as "rejection of payment," "delay," "poor service," and "ineffective medication," and constructs a co-occurrence matrix based on these keywords. A co-occurrence matrix is a statistical model used to record the frequency of different keywords appearing simultaneously in historical abnormal cases, thereby revealing potential correlation patterns. For example, "rejection of payment" and "assessed loss amount" often co-occur, and "ineffective medication" and "side effects" are highly correlated. This correlation can serve as an important basis for judging whether the current conversation may escalate into a serious incident.
[0037] Based on this, the risk prediction module 103 integrates the sentiment classification results (such as "severely negative"), scenario classification results (such as "claims dispute"), and keyword association strength extracted from the co-occurrence matrix output by the semantic analysis module 102 with the structured fields in the user's historical data to construct a multi-dimensional feature vector containing multi-dimensional (e.g., 89-dimensional) features. This multi-dimensional feature vector comprehensively depicts the emotional state, business scenario, historical behavioral tendencies, and potential risk paths of the current interaction. The risk prediction module 103 inputs this multi-dimensional feature vector into a pre-trained risk prediction model. This model can be built based on the XGBoost classifier, whose training samples come from a large number of real-world abnormal cases. Each sample contains original voice information and its corresponding historical abnormal cases (such as records of complaints, escalation work orders, or medical disputes), and establishes the association between voice and historical events through unique identifiers such as customer number and case number. The XGBoost model automatically learns the contribution weight of each feature to the occurrence of risk through a gradient boosting mechanism, ultimately outputting a probability value between 0 and 1, representing the possibility of abnormal risk in the current interaction. In practical applications, the training process of risk prediction models can be as follows: Figure 3 As shown, the system uses root cause analysis factors and voice text provided by the direct sales platform as input. The voice text originates from the user's speech recognition results obtained by the ASR speech-to-text module. The voice text is first fed into a sound model (BERT+Softmax), which is based on a BERT pre-trained language model combined with a Softmax classification layer for semantic feature extraction and preliminary classification of the voice text. The semantic features output by the sound model, together with the root cause analysis factors provided by the direct sales platform, constitute the input feature vector of the XGBoost classifier. The XGBoost classifier is trained under supervision using a large number of historical samples. Each sample contains voice text and its corresponding anomaly label (such as whether a complaint or dispute has occurred). Through a gradient boosting mechanism, the decision tree ensemble is continuously optimized, learning the predictive ability of different feature combinations for risk events. Finally, the trained XGBoost model outputs the probability or risk level of customer complaints, forming an early warning list to achieve accurate identification and early intervention of high-risk interactions.
[0038] Thus, based on the risk prediction module 103, not only is the judgment shifted from a single emotion or scenario to the fusion of multi-source heterogeneous data, but the system also significantly improves the ability to predict potential crises by mining hidden risk chains through co-occurrence matrices. This is particularly suitable for high-value customer groups with large emotional fluctuations and complex demands. For example, in a financial scenario, a customer with a historical risk level of "high" calls and says, "You said the money would arrive in 24 hours last time, but it's been three days now." The interactive risk warning system, combined with the customer's records of multiple complaints about delayed payments, finds a strong correlation between "delay" and "money arrival" in the co-occurrence matrix, thus predicting a complaint probability as high as 85%. In a medical scenario, a chronic disease patient expresses during a follow-up visit that "the medicine is not working and makes me feel worse." The interactive risk warning system, combined with the co-occurrence frequency of "poor treatment effect" and "adverse reaction" in the patient's historical medical records, judges that there is a high risk of escalation of doctor-patient conflict and triggers an early warning.
[0039] When the risk prediction result reaches the preset warning standard, the feedback module 104 obtains the risk warning strategy generated based on the big language model, and pushes the risk prediction result and the risk warning strategy to the risk handling party for risk handling. The risk warning strategy is matched in the preset strategy library by the big language model after determining the level of interaction anomaly and the cause of the anomaly by recognizing text information.
[0040] Feedback module 104 is responsible for packaging the risk prediction results and the acquired risk warning strategies, and pushing them in real time to the corresponding processing parties through the task distribution platform, such as dedicated customer service of insurance companies, medical staff coordinators of hospitals, or intelligent work order systems. After processing, it receives feedback results (such as "soothed," "escalated processing," "customer withdrew complaint," etc.) and feeds the closed-loop data back to the model training pool. Feedback module 104 relies on message queues and API (Application Programming Interface) gateways to achieve low-latency and highly reliable task delivery, and supports fine-grained routing by institution, channel, and personnel role.
[0041] In this way, a complete closed loop of "early warning-handling-feedback-optimization" is constructed based on feedback module 104. This not only ensures that risk events are responded to in a timely manner, but also iteratively optimizes the model by continuously accumulating effective and ineffective cases, making the interactive risk warning system more accurate and intelligent with use. For example, in financial scenarios, warning tasks can be pushed to the claims specialist's mobile app. After the specialist communicates according to the strategy, they mark it as "successfully processed," and the interactive risk warning system automatically archives it. In medical scenarios, warning information is pushed to the large screen at the outpatient nurse station. Nurses proactively call patients to explain, and the processing results are simultaneously updated to the electronic medical record system, forming a service record.
[0042] In this embodiment of the invention, optionally, the feedback module 104, when the risk prediction result reaches a preset warning standard, acquires a risk warning strategy generated based on a large language model, and pushes the risk prediction result and the risk warning strategy to the risk handling party for risk handling. This includes: determining a preset warning standard, comparing the probability of anomaly occurrence indicated by the risk prediction result with the probability threshold indicated by the warning standard, and determining that the risk prediction result has reached the warning standard when the probability of anomaly occurrence exceeds the probability warning; receiving the risk warning strategy, calling a task distribution platform, determining the risk handling party based on the task distribution platform, and pushing the risk warning strategy to the risk handling party for risk handling through the task distribution platform; correspondingly, the feedback module is also used to receive risk handling records fed back by the risk handling party, the risk handling records including risk handling results and risk handling process information, and using the risk handling records to iteratively optimize the risk prediction model used for risk prediction and the large language model used to generate the risk warning strategy.
[0043] In this embodiment, the feedback module 104 initiates a closed-loop response process when the risk prediction result reaches a preset warning standard. The feedback module 104 pre-sets a dynamic probability threshold as the warning standard; for example, interactions with an abnormal occurrence probability greater than or equal to 80% are judged as high-risk events. This threshold can be flexibly adjusted according to the business scenario. For instance, the financial sector, which is highly sensitive to complaints, can set a lower threshold, while the medical sector, which focuses more on serious adverse reaction events, can set a higher trigger threshold. When the probability value output by the risk prediction model exceeds this threshold, the feedback module 104 determines that the current interaction has reached the warning standard and enters the intervention process. Subsequently, the feedback module 104 calls the task distribution platform to receive the risk warning strategy generated by the large language model. This strategy includes specific communication scripts, handling suggestions, and reassurance plans.
[0044] The task distribution platform intelligently matches the optimal risk handling party based on multi-dimensional information such as the client's organization, service type, and customer service personnel's skill tags. For example, it pushes insurance claim disputes to claims specialists and transfers patient medication questions to clinical pharmacists or doctor-patient coordinators, ensuring that tasks accurately reach qualified personnel. Through API interfaces or message queues, the feedback module 104 pushes risk prediction results and early warning strategies to the corresponding handling party's workbench in the form of structured task orders, enabling rapid response. Meanwhile, the feedback module 104 is also responsible for receiving risk handling records submitted by the processing party. These records include not only the final handling results (such as "customer has been appeased", "complaint has been withdrawn", "solution has been adjusted"), but also complete handling process information, such as communication duration, key response content, and the trajectory of customer emotional changes. This closed-loop data is fed back to the feedback module 104 for iterative optimization of the risk prediction model and the large language model. On the one hand, by labeling the differences between the actual results and the predicted results, the feature weights of the XGBoost model are corrected. On the other hand, the large language model is trained using real handling cases to make the generated strategies more targeted and compliant. This constructs a complete data closed loop of "identification-early warning-handling-feedback-optimization", which significantly improves the self-learning ability and long-term stability of the interactive risk early warning system.
[0045] For example, in a financial scenario, the feedback module 104 detected a high-risk warning for a customer due to a delay in compensation and pushed a strategy suggestion to "explain the expedited review process and promise a response within today." After the customer service implemented the suggestion, the customer was satisfied, and this successful case was used to optimize the generation of strategies for similar scenarios in the future. In a medical scenario, a patient expressed strong dissatisfaction due to drug side effects, and the feedback module 104 pushed a strategy to "suggest a follow-up visit with the doctor and change to an alternative drug." After the doctor adopted the suggestion, the symptoms were relieved, and this handling path was incorporated into the knowledge base to improve the efficiency of dealing with similar issues in the future.
[0046] In embodiments of the present invention, optionally, such as Figure 4 As shown, the interactive risk warning system also includes a warning strategy generation module 105: when the risk prediction result reaches the preset warning standard, the warning strategy generation module 105 generates a risk warning strategy based on a large language model and transmits the risk warning strategy to the feedback module.
[0047] In this embodiment, the early warning strategy generation module 105 is activated after the risk prediction module 103 determines that the risk has reached the threshold. It then invokes a Large Language Model (LLM) to perform comprehensive reasoning on the original text, sentiment tags, scene tags, and user profile, automatically generating a personalized risk early warning strategy. The Large Language Model refers to a deep neural network trained on massive amounts of text and possessing contextual understanding and generation capabilities. Here, it is used to analyze the user's core needs, match solutions from a pre-set strategy library, and output a structured strategy that includes reference phrases, suggested solutions, and reassurance measures. The strategy generation process can integrate rule constraints with generation freedom to ensure that the suggestions are both compliant and authentic.
[0048] In this way, based on the early warning strategy generation module 105, the limitations of traditional template-based responses can be overcome, enabling intelligent intervention and significantly improving customer satisfaction and risk mitigation efficiency. For example, in a financial scenario, in response to customer complaints about slow claims processing, the interactive risk early warning system generates a message: "Your case has been expedited and is expected to be reviewed today. We will have a dedicated person follow up." In a medical scenario, in response to a patient's question about the efficacy of medication, the interactive risk early warning system suggests that the doctor reply: "We understand your concerns. We can arrange a follow-up visit this afternoon to adjust the medication plan and give you priority access to an appointment."
[0049] In an embodiment of the present invention, optionally, the early warning strategy generation module 105 generates a risk early warning strategy based on a large language model when the risk prediction result reaches a preset early warning standard. This includes: determining the preset early warning standard; comparing the probability of anomaly occurrence indicated by the risk prediction result with the probability threshold indicated by the early warning standard; determining that the risk prediction result reaches the early warning standard when the probability of anomaly occurrence exceeds the probability warning threshold; acquiring a pre-trained large language model, which is trained with an industry-specific interactive intelligent agent for the current industry and a case-handling intelligent agent for a non-current industry; calling the large language model to perform secondary recognition on text information using the industry-specific interactive intelligent agent and the case-handling intelligent agent to obtain the interaction anomaly level and the cause of the anomaly; and matching the interaction anomaly level and the cause of the anomaly in a preset strategy library to obtain a risk early warning strategy. The strategy library includes various anomaly scenario information, each of which is associated with an anomaly level and a cause of occurrence, and each anomaly scenario information corresponds to an anomaly handling strategy.
[0050] In this embodiment, after the risk prediction result reaches the preset warning standard, the early warning strategy generation module 105 initiates the intelligent strategy generation process based on the large language model. The early warning strategy generation module 105 first confirms whether the probability of an anomaly exceeds a preset threshold (e.g., 80%). Once a high-risk interaction is determined, the strategy generation mechanism is triggered. In practical applications, the early warning strategy generation module 105 calls an open-source large language model with more than 10 bytes of parameters. This model internally constructs a dual-Agent intelligent agent system: one is an "industry interaction intelligent agent" customized for the current industry (e.g., finance or healthcare), possessing a deep understanding of specific domain terminology, service processes, and customer behavior patterns; the other is a highly generalizable "non-current industry case handling intelligent agent" used to address cross-scenario or complex composite demands, improving the generalization ability of the strategy. Upon receiving the text information to be processed, the large language model, through the collaborative work of these two intelligent agents, performs secondary recognition and analysis of the input content. The industry-specific interactive agent is responsible for parsing the business logic in user expressions, such as determining that "disputes over claim amounts" are typical disputes in the insurance field. The non-industry-specific agent extracts potential points of conflict from a broader perspective, such as identifying "long-standing unresolved issues" as potentially indicating service process defects. Combining the outputs of both, the early warning strategy generation module 105 further determines the level of interaction anomaly (e.g., mild, moderate, severe) and the cause of the anomaly (e.g., "poor communication," "unclear policy explanation," "ineffective treatment," etc.). Subsequently, the early warning strategy generation module 105 uses the identified anomaly level and cause as query conditions for precise matching in a pre-set strategy library. This strategy library contains information on various anomaly scenarios, each associated with a corresponding anomaly level, cause, and standardized handling strategy. For example, "delayed claims + emotional agitation" corresponds to a handling path of "rapid response + compensation suggestion + dedicated follow-up."
[0051] In practical applications, the early warning strategy generation module 105 can also incorporate RAG (Retrieval-Augmented Generation) technology to retrieve similar cases from the historical complaint voice knowledge base, providing contextual support for strategy generation and ensuring that the recommendations are both compliant with standards and feasible. Finally, the early warning strategy generation module 105 outputs a structured risk warning strategy, including complaint reduction scripts, processing steps, and reference cases, encapsulated in JSON (JavaScript Object Notation) format for subsequent push notifications.
[0052] In this way, based on the early warning strategy generation module 105, through dual-agent collaboration and knowledge base linkage, the accuracy and professionalism of strategy generation are guaranteed, avoiding the bias of single-model reasoning and significantly improving the ability to deal with complex customer service scenarios. For example, in a financial scenario, if a customer becomes agitated due to "failure to replace original parts," the interactive risk warning system identifies it as "serious negative + service commitment violation" and matches it with the strategy of "priority compensation + written apology + after-sales follow-up." In a medical scenario, if a patient reports "dizziness after taking medication," the interactive risk warning system judges it as "moderate abnormality + adverse reaction" and recommends the strategy of "suspending medication + immediate follow-up visit + providing alternative solutions," effectively reducing the risk of escalation of disputes.
[0053] In practical applications, the process of generating risk warning strategies can be as follows: Figure 5 As shown, the user's voice is input as the warning signal. This voice, as an input parameter, is passed through the AE model and enters the warning strategy generation module 105. The warning strategy generation module 105 first calls the large model and obtains user personality and historical complaint-related tag plugins to identify the severity, scenario, and demands of the complaint in the user's voice. Simultaneously, it combines RAG enhancement technology to retrieve relevant cases and information from the historical complaint voice knowledge base to improve recognition accuracy and contextual understanding. After obtaining the initial recognition results, the warning strategy generation module 105 further incorporates structured data from the knowledge base, such as the historical complaint lifecycle, stages, and scenarios, and integrates and analyzes user complaint reduction strategies, resolution techniques, and historical cases through the workflow module. At the same time, the warning strategy generation module 105 calls module plugins to score the severity of the customer's complaint, comprehensively assessing its emotional intensity and potential risk level. Finally, the warning strategy generation module 105 integrates and processes all information to generate a structured output containing the severity of the user complaint, complaint scenario, complaint stage, complaint target, and specific complaint reduction strategies. This output is standardized in JSON format, completing the risk warning strategy generation process.
[0054] The interactive risk warning system provided in this invention can convert customer phone feedback in financial scenarios and patient voice communication in medical scenarios into text in real time and classify them by emotion and scenario. It responds quickly and ensures the timeliness of risk warning. Moreover, by classifying emotions and scenarios at the same time, it can accurately grasp the complex and ever-changing emotions and true intentions of customers or patients, further improving the accuracy of risk identification and achieving a good overall risk warning effect.
[0055] Please see Figure 6 As shown, Figure 6 A flowchart illustrating the interactive risk warning method based on speech recognition and a large language model provided in this embodiment of the invention includes the following steps: S10: Real-time acquisition of user-input voice information and real-time conversion of voice information into text information.
[0056] In this embodiment of the invention, the interactive risk warning system collects raw voice signals input by users through multiple channels such as telephone, mobile APP, and smart customer service terminals in real time via a voice access channel deployed in the cloud, and uses automatic speech recognition (ASR) technology to convert the voice stream into structured text in real time. The voice access channel can be built on a scalable cloud processor resource pool, dynamically allocating GPU computing power to handle high concurrency requests; the voice signal can be processed in 25-millisecond frames, and then the Mel-frequency cepstral coefficients (MFCC) of each frame are extracted as acoustic feature vectors. This feature effectively characterizes the energy distribution of human voice in different frequency bands and is normalized to eliminate differences between devices and the environment. The acoustic feature sequence is input into a deep neural network speech recognition model optimized for industry terms such as finance and healthcare. During the training phase, the model introduces a weighted loss function based on industry keywords to improve the recognition accuracy of professional terms such as "claims" and "prescription drugs". To further improve the transcription quality, the interactive risk warning system also combines a language model based on the fusion of n-gram and bidirectional LSTM with a weighted finite state converter (WFST) to dynamically decode and optimize the initial recognition results, ensuring semantic coherence and terminology accuracy.
[0057] In this way, the above process not only achieves low-latency, high-precision transcription of cross-channel voice, but also significantly improves the ability to restore key information in professional scenarios, laying a high-quality data foundation for subsequent risk analysis. For example, in a financial scenario, a customer complains on the phone, "You promised to pay in three days, but it's been five days and the money hasn't arrived yet." The interactive risk warning system accurately identifies keywords such as "payment" and "money arrived" and transcribes them completely. In a medical scenario, a patient describes in their voice that "my stomach pain worsened after taking aspirin." The interactive risk warning system accurately transcribes the drug name and symptoms, avoiding misjudgments due to misidentification.
[0058] S20: Perform sentiment and scene classification on the text information to obtain semantic analysis results.
[0059] In this embodiment of the invention, the interactive risk warning system invokes a BERT-based semantic recognition model, which has been pre-trained on multiple historical dialogue datasets, to perform deep semantic parsing on the output text. This semantic recognition model first performs Named Entity Recognition (NER) to extract entities with business or emotional significance from the text, such as "claim refusal," "follow-up visit," and "poor service attitude," and classifies these entities into two groups: sentiment-based (e.g., "dissatisfaction," "anger") and scenario-based (e.g., "claim processing timeliness," "medication feedback"). Subsequently, the interactive risk warning system performs fine-grained labeling of sentiment-based entities according to a pre-defined multi-level sentiment classification system (e.g., mildly negative, severely negative), and simultaneously matches and categorizes scenario-based entities according to a multi-level scenario classification system (e.g., finance-claim disputes, medical-efficacy questioning).
[0060] In practical applications, the semantic analysis process can integrate multiple scenario knowledge bases and emotional appeal rules, supporting the simultaneous output of structured semantic analysis results. This joint modeling of emotion and scenario avoids the limitations of traditional single-dimensional analysis, enabling a more comprehensive and accurate capture of users' true intentions and emotional states. For example, in a financial scenario, if a customer says, "I've submitted all the materials, but you keep delaying processing," the interactive risk warning system identifies "delay" as a serious negative emotion, and categorizes "complete materials but not processed" as "abnormal claims process." In a medical scenario, if a patient states, "The medication I was prescribed last time didn't work at all; my symptoms have actually worsened," the interactive risk warning system identifies this as a combination of "negative emotion + doubt about efficacy," providing crucial input for risk prediction.
[0061] S30: Based on the semantic analysis results and the user's historical data, perform interaction risk prediction to obtain the risk prediction result.
[0062] In this embodiment of the invention, the interactive risk warning system constructs a multi-dimensional feature vector based on semantic analysis results and user historical data (including risk level tags, historical complaint counts, interaction duration, and speech volume ratio). Specifically, the interactive risk warning system can use the TF-IDF algorithm to extract keywords from the user's historical complaint text and construct a co-occurrence matrix based on these keywords to quantify the correlation strength of different keywords (such as "claim denial" and "loss amount") in historical abnormal cases. The co-occurrence matrix, together with the current semantic tags and user behavior features, forms a multi-dimensional feature vector (e.g., 89 dimensions), which is then input into the XGBoost risk prediction model. The risk prediction model is trained using a large number of labeled samples, each containing voice text and its corresponding real abnormal event tags (such as whether the complaint was escalated or whether a dispute arose). Through a gradient boosting mechanism, it learns the contribution weight of each feature to the occurrence of risk and finally outputs the probability of abnormal occurrence between 0 and 1.
[0063] In this way, by integrating real-time semantics with long-term behavioral trajectories and combining keyword co-occurrence relationships to uncover potential risk paths, the predictive power and accuracy can be significantly improved. For example, in a financial scenario, a customer with two previous complaint records expresses dissatisfaction with the speed of compensation again. The interactive risk warning system, combining historical data with current sentiment, predicts a risk probability of 87%. In a medical scenario, a chronic disease patient strongly questions the efficacy of medication for the first time. The interactive risk warning system, combining the patient's frequent medical visits with this negative expression, triggers a high-risk warning.
[0064] S40: If the risk prediction results meet the preset warning criteria, obtain the risk warning strategy generated based on the large language model, and push the risk prediction results and risk warning strategy to the risk handling party for risk handling.
[0065] In this embodiment of the invention, when the probability of an anomaly exceeds a preset threshold (e.g., 80%), the interactive risk warning system determines that the warning standard has been met and then initiates the risk warning strategy generation and push process. The interactive risk warning system calls a large language model (with more than 10 bytes of parameters). This model embeds an industry-specific interactive intelligent agent (e.g., an agent specifically for the insurance or medical field) and a general case processing intelligent agent to perform secondary recognition on the original text, determine the level of interactive anomaly (e.g., severe, moderate) and the cause of occurrence (e.g., "unclear policy explanation", "poor treatment effect").
[0066] Subsequently, the interactive risk warning system matches corresponding abnormal scenario information in a pre-set strategy library. This library contains various scenario templates, each associated with an anomaly level, cause, and standardized handling strategy. It also utilizes RAG technology to retrieve similar cases from a historical complaint voice knowledge base, generating a structured risk warning strategy that includes complaint reduction scripts, handling suggestions, and reference solutions. The risk warning strategy, along with the risk prediction results, is pushed to the matched risk handling party (such as a claims specialist or medical coordinator) through a task distribution platform. Upon completion of the handling, the system receives risk handling records (including handling results and process information) for iterative model optimization.
[0067] In this way, through the above process, a closed-loop intelligent intervention system of "identification + early warning + handling + feedback + optimization" is constructed, realizing the personalization, professionalization and traceability of strategies. For example, in the financial scenario, the interactive risk early warning system pushes the strategy of "expedited review + dedicated follow-up visit" for customers with "delayed compensation" issues, and the customer withdraws the complaint after customer service implements it; in the medical scenario, in response to patients' feedback of "dizziness after taking medication", the interactive risk early warning system recommends the solution of "suspending medication + priority follow-up visit", and the doctor's adoption effectively resolves potential disputes.
[0068] The method provided in this invention can convert customer telephone feedback in financial scenarios or patient voice communication in medical scenarios into text in real time and classify it according to emotion and scenario. It responds quickly and ensures timely risk warning. Moreover, the simultaneous classification of emotion and scenario can accurately grasp the complex and ever-changing emotions and true intentions of customers or patients, further improving the accuracy of risk identification and the overall risk warning effect is good.
[0069] Specific limitations regarding the interactive risk warning method based on speech recognition and large language models can be found in the limitations of the interactive risk warning system based on speech recognition and large language models mentioned above, and will not be repeated here. Each module in the aforementioned interactive risk warning system based on speech recognition and large language models can be implemented entirely or partially through software, hardware, or a combination thereof. These modules can be embedded in or independent of the processor in a computer device, or stored in the memory of a computer device as software, so that the processor can call and execute the corresponding operations of each module.
[0070] In one embodiment, a computer device is provided, which may be a server, and its internal structure diagram may be as follows: Figure 7 As shown, the computer device includes a processor, memory, network interface, and database connected via a system bus. The processor provides computing and control capabilities. The memory includes non-volatile and / or volatile storage media and internal memory. The non-volatile storage media stores the operating system, computer programs, and database. The internal memory provides an environment for the operation of the operating system and computer programs stored in the non-volatile storage media. The network interface is used to communicate with external clients via a network connection. When the computer program is executed by the processor, it implements the server-side functions or steps of an interactive risk warning method based on speech recognition and a large language model.
[0071] In one embodiment, a computer device is provided, which may be a client, and its internal structure diagram may be as follows: Figure 8 As shown, the computer device includes a processor, memory, network interface, display screen, and input devices connected via a system bus. The processor provides computing and control capabilities. The memory includes non-volatile storage media and internal memory. The non-volatile storage media stores the operating system and computer programs. The internal memory provides an environment for the operation of the operating system and computer programs stored in the non-volatile storage media. The network interface is used to communicate with an external server via a network connection. When the computer program is executed by the processor, it implements client-side functions or steps of an interactive risk warning method based on speech recognition and a large language model.
[0072] In one embodiment, a computer device is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, it implements method steps for executing any module: The speech recognition module collects the user's input speech information in real time and converts the speech information into text information in real time. The semantic analysis module performs sentiment classification and scene classification on the text information to obtain semantic analysis results; The risk prediction module performs interactive risk prediction based on the semantic analysis results and the user's historical data, and obtains the risk prediction result. When the risk prediction result reaches the preset warning standard, the feedback module obtains the risk warning strategy generated based on the large language model, and pushes the risk prediction result and the risk warning strategy to the risk handling party for risk handling. The risk warning strategy is matched in the preset strategy library by the large language model after identifying the text information to determine the level of interaction abnormality and the cause of the abnormality.
[0073] In one embodiment, a computer-readable storage medium is provided having a computer program stored thereon, wherein the computer program, when executed by a processor, implements method steps for any module to be executed. The speech recognition module collects the user's input speech information in real time and converts the speech information into text information in real time. The semantic analysis module performs sentiment classification and scene classification on the text information to obtain semantic analysis results; The risk prediction module performs interactive risk prediction based on the semantic analysis results and the user's historical data, and obtains the risk prediction result. When the risk prediction result reaches the preset warning standard, the feedback module obtains the risk warning strategy generated based on the large language model, and pushes the risk prediction result and the risk warning strategy to the risk handling party for risk handling. The risk warning strategy is matched in the preset strategy library by the large language model after identifying the text information to determine the level of interaction abnormality and the cause of the abnormality.
[0074] It should be noted that the functions or steps that can be implemented by the computer-readable storage medium or computer device described above can be referred to the relevant descriptions on the server side and client side in the foregoing method embodiments. To avoid repetition, they will not be described one by one here.
[0075] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in this invention are all information and data authorized by the user or fully authorized by all parties.
[0076] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium. When executed, the computer program can include the processes of the embodiments of the above methods. Any references to memory, storage, databases, or other media used in the embodiments provided by this invention can include non-volatile and / or volatile memory. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), dual data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), Rambus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM), etc.
[0077] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the above-described division of functional units and modules is used as an example. In practical applications, the above functions can be assigned to different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above.
[0078] The above-described embodiments are only used to illustrate the technical solutions of the present invention, and are not intended to limit it. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention, and should all be included within the protection scope of the present invention.
Claims
1. An interactive risk warning system based on speech recognition and a large language model, characterized in that, The interactive risk warning system includes a voice recognition module, a semantic analysis module, a risk prediction module, and a feedback module. The speech recognition module collects the user's input speech information in real time and converts the speech information into text information in real time. The semantic analysis module performs sentiment classification and scene classification on the text information to obtain semantic analysis results; The risk prediction module performs interactive risk prediction based on the semantic analysis results and the user's historical data, and obtains the risk prediction result. When the risk prediction result reaches the preset warning standard, the feedback module obtains the risk warning strategy generated based on the large language model, and pushes the risk prediction result and the risk warning strategy to the risk handling party for risk handling. The risk warning strategy is matched in the preset strategy library by the large language model after identifying the text information to determine the interaction anomaly level and the cause of the anomaly.
2. The interactive risk early warning system according to claim 1, characterized in that, The speech recognition module collects user-input speech information in real time and converts the speech information into text information in real time, including: The system collects voice information from users inputting from multiple channels in real time based on a voice access channel, wherein the voice access channel is constructed by deploying a cloud processor resource pool. According to the preset audio segmentation standard, the speech information is processed into frames to obtain multiple frames of audio information. The multi-dimensional audio feature vector of each frame of audio information is extracted and normalized to generate an acoustic feature vector for each frame of audio information, so as to obtain multiple acoustic feature vectors corresponding to the multiple frames of audio information. Using the multiple acoustic feature vectors, a continuous acoustic feature sequence is generated. The acoustic feature sequence is then input into a speech recognition model to obtain an initial recognition result. The speech recognition model is then optimized using a loss function based on the relevant technical terminology of the target industry. By combining a pre-trained language model and a weighted finite state converter, the initial recognition result is dynamically decoded and optimized to obtain the text information. The language model is a model built based on the fusion of n-gram and bidirectional LSTM.
3. The interactive risk early warning system according to claim 1, characterized in that, The semantic analysis module performs sentiment classification and scene classification on the text information to obtain semantic analysis results, including: A pre-trained semantic recognition model is obtained, which is based on a multi-level neural network structure and is obtained by secondary pre-training using multiple historical dialogue data. Based on the semantic recognition model, named entity recognition is performed on the text information to obtain multiple entity information, and the multiple entity information is divided into an emotion-type entity information group and a scene-type entity information group. According to a preset multi-level emotion classification system, each entity information included in the emotion-type entity information group is classified according to emotion to obtain emotion classification results. At the same time, according to a preset multi-level scene classification system, each entity information included in the scene-type entity information group is classified according to scene to obtain scene classification results. The semantic analysis result is obtained by integrating the emotion classification result and the scene classification result.
4. The interactive risk early warning system according to claim 1, characterized in that, The risk prediction module performs interactive risk prediction based on the semantic analysis results and the user's historical data, obtaining risk prediction results, including: Obtain the user's historical data, which includes user risk level data and interaction duration data; The user's historical data is extracted using a word segmentation algorithm, and a co-occurrence matrix is constructed based on the extracted keywords. The co-occurrence matrix is used to indicate the frequency and association of different keywords appearing simultaneously in historical abnormal cases. Based on the semantic analysis results and the co-occurrence matrix, a multidimensional feature vector is constructed. The multidimensional feature vector is input into a pre-constructed risk prediction model for risk prediction. The risk prediction result, which indicates the probability of an anomaly, is obtained from the output of the risk prediction model. The risk prediction model is trained using multiple anomaly samples. Each anomaly sample includes sample voice information and historical anomaly cases that are related to the sample voice information.
5. The interactive risk early warning system according to claim 1, characterized in that, When the risk prediction result reaches a preset warning standard, the feedback module acquires a risk warning strategy generated based on a large language model, and pushes the risk prediction result and the risk warning strategy to the risk handling party for risk handling, including: A preset warning standard is determined, and the probability of an anomaly indicated by the risk prediction result is compared with the probability threshold indicated by the warning standard. If the probability of an anomaly exceeds the probability warning, it is determined that the risk prediction result has reached the warning standard. The system receives the risk warning strategy, invokes the task distribution platform, determines the risk handling party based on the task distribution platform, and pushes the risk warning strategy to the risk handling party for risk handling through the task distribution platform. Accordingly, the feedback module is also used to receive risk processing records from the risk processing party. The risk processing records include risk processing results and risk processing process information. The risk processing records are used to iteratively optimize the risk prediction model used for risk prediction and the large language model used to generate the risk warning strategy.
6. The interactive risk early warning system according to claim 1, characterized in that, The interactive risk warning system also includes a warning strategy generation module: When the risk prediction result reaches the preset warning standard, the early warning strategy generation module generates the risk early warning strategy based on the large language model and transmits the risk early warning strategy to the feedback module.
7. The interactive risk early warning system according to claim 6, characterized in that, When the risk prediction result reaches the preset warning standard, the early warning strategy generation module generates the risk early warning strategy based on the large language model, including: A preset warning standard is determined, and the probability of an anomaly indicated by the risk prediction result is compared with the probability threshold indicated by the warning standard. If the probability of an anomaly exceeds the probability warning, it is determined that the risk prediction result has reached the warning standard. Obtain the pre-trained large language model, which is trained with an industry-specific interactive agent for the current industry and a case-handling agent for non-current industries. The large language model is invoked to perform secondary recognition on the text information using the industry interaction agent and the case processing agent to obtain the interaction anomaly level and the cause of the anomaly. The risk warning strategy is obtained by matching the interaction anomaly level and the anomaly cause in the strategy library. The strategy library includes a variety of anomaly scenario information, each of which is associated with an anomaly level and an anomaly cause, and each of which corresponds to an anomaly handling strategy.
8. An interactive risk warning method based on speech recognition and a large language model, characterized in that, The method includes: Real-time acquisition of user-input voice information, and real-time conversion of the voice information into text information; The text information is subjected to sentiment classification and scene classification to obtain semantic analysis results; Based on the semantic analysis results and the user's historical data, interaction risk prediction is performed to obtain the risk prediction result. If the risk prediction result reaches the preset warning standard, a risk warning strategy generated based on a large language model is obtained, and the risk prediction result and the risk warning strategy are pushed to the risk handling party for risk handling. The risk warning strategy is matched in a preset strategy library by the large language model after identifying the text information to determine the level of interaction abnormality and the cause of the abnormality.
9. A device comprising a memory and a processor, the memory storing a computer program, characterized in that, When the processor executes the computer program, it implements the method steps performed by any module of claims 1 to 7.
10. A medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the method steps performed by any of the modules in claims 1 to 7.