Intelligent call management system and method for multi-modal semantic understanding and dynamic strategy arrangement

The intelligent call management system, which utilizes multimodal semantic understanding and dynamic strategy orchestration, addresses the shortcomings of traditional systems in semantic understanding and resource scheduling. It enables efficient identification and dynamic response to complex call scenarios, improves identification accuracy and user satisfaction, and resolves the disconnect between the business layer and the network layer.

CN122053749APending Publication Date: 2026-05-15EASTERN COMM
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
EASTERN COMM
Filing Date
2026-01-27
Publication Date
2026-05-15

AI Technical Summary

Technical Problem

Traditional call management systems are inadequate in terms of semantic understanding, policy response, and network resource coordination, making it difficult to meet the refined management and control needs in complex communication scenarios. In particular, they suffer from low recognition accuracy and delayed response when facing dynamically changing call contexts, and the business layer is disconnected from the network layer, with a lack of linkage mechanism for resource scheduling.

Method used

The intelligent call management system, which employs multimodal semantic understanding and dynamic strategy orchestration, achieves deep perception of call semantics, dynamic adjustment of handling strategies, and collaborative linkage between the business layer and the network layer through a multimodal feature perception module, a semantic deep understanding engine, a dynamic strategy orchestration center, and a resource scheduling and execution module. It integrates acoustic, semantic, and behavioral features to dynamically generate handling strategies and adjust network resources.

Benefits of technology

It improves the recognition accuracy and response efficiency in complex call scenarios, enhances the fraud detection rate and resource utilization, achieves efficient identification of variant fraudulent tactics, improves user satisfaction and evidence preservation efficiency, and breaks down the disconnect between the business layer and the network layer.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122053749A_ABST
    Figure CN122053749A_ABST
Patent Text Reader

Abstract

The invention relates to an intelligent call management system and method for multi-mode semantic understanding and dynamic strategy arrangement. The system comprises a multi-modal feature perception module, a semantic deep understanding engine, a dynamic strategy arrangement center and a resource scheduling execution module, and deep mining of a call intention and a risk level is realized by collecting and aligning voice, text and behavior features in a call process in real time and utilizing a dynamic semantic state machine; and dynamically generating and arranging a disposal strategy instruction set based on the identified semantic state, and synchronously driving the underlying communication network to carry out QoS priority adjustment, media stream redirection and resource collaborative scheduling. The technical problems of single dimension, strategy solidification and service and network resource disjunction of traditional call management are solved, and the performance of a communication system in the aspects of variation fraud identification, scene self-adaptive control and network resource utilization rate is remarkably improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to an intelligent call management system and method with multimodal semantic understanding and dynamic strategy orchestration. Background Technology

[0002] As communication networks continue to evolve towards intelligence and service, and as abnormal communication behaviors such as telecommunications network fraud and malicious marketing become more complex and covert, traditional call management systems have gradually revealed significant deficiencies in semantic understanding, policy response, and network resource coordination capabilities, making it difficult to meet the refined management and control needs of new communication service scenarios.

[0003] First, in terms of semantic perception, existing call management systems mostly adopt a single text modality-based processing approach, typically relying on keyword matching, rule filtering, or static feature library comparison after speech-to-text (ASR) to identify call intent. However, actual calls are generally characterized by complex factors such as colloquial expressions, ambiguity, emotional fluctuations, and environmental noise, making it difficult to fully depict the true intent of both parties using only text semantic analysis. Furthermore, traditional solutions generally neglect the correlation between acoustic paralinguistic features in speech (such as intonation variations, abnormal speech rate, and pause patterns) and user call behavior features (such as call frequency, time distribution, and conversation sequence patterns). This results in low recognition accuracy and insufficient robustness in complex scenarios such as face-to-face dialogue distortion, semantic avoidance, or strategic inducement, making it difficult to achieve a deep understanding of call risks.

[0004] Secondly, regarding the strategy handling logic, most existing call management systems employ predefined decision trees or static IF-THEN rules for call handling, exhibiting a clear "hard-coded" characteristic in their strategy execution logic. These solutions typically only respond to predefined scenarios. When the call context dynamically changes during the session (e.g., from a simple business inquiry gradually evolving into a high-risk inducement behavior), the system struggles to adjust its handling strategy online based on real-time semantic changes. Offline analysis is often required after the call ends, leading to delayed responses and failing to meet the business needs for real-time prevention and dynamic intervention.

[0005] Secondly, regarding communication resource systems, the semantic analysis results of the service layer and the resource scheduling mechanism of the underlying communication network have long been disconnected in existing technologies. Currently, the quality of service (QoS) guarantee, bandwidth allocation, and routing selection on the network side are usually configured based on signaling parameters or static policies, lacking a linkage mechanism with the semantics of call content and the level of service risk. When the service layer identifies high-value (such as VIP customers) or high-risk (such as ongoing fraud) calls, the underlying network resources cannot be adjusted in real time based on the semantic results, such as differentiated bandwidth guarantees, link isolation, or media stream redirection. This "two-layer disconnect" between services and the network severely limits the communication network's ability to support intelligent service scenarios.

[0006] Therefore, how to integrate multimodal information during a call to achieve a deep understanding of the call's semantics, emotions, and risks, and on this basis, realize the dynamic orchestration of handling strategies and the coordinated scheduling of communication network resources, has become a key technical problem that urgently needs to be solved in the field of intelligent call management and intelligent communication networks. Summary of the Invention

[0007] In view of the above-mentioned problems in the prior art, the purpose of this invention is to provide an intelligent call management system and method with multimodal semantic understanding and dynamic strategy orchestration, so as to realize deep perception of call semantics, dynamic adjustment of handling strategies, and collaborative linkage between the business layer and the network layer, thereby improving the recognition accuracy, response efficiency and resource utilization in complex call scenarios.

[0008] The intelligent call management system with multimodal semantic understanding and dynamic strategy orchestration is characterized by including a multimodal feature perception module, a semantic deep understanding engine, a dynamic strategy orchestration center, and a resource scheduling and execution module; the multimodal feature perception module is used to collect raw voice streams, text sequences, and signaling behavior data in the call link in real time, and perform cross-modal feature alignment based on a unified clock source to generate multi-dimensional time-series feature vectors; The semantic deep understanding engine is used to perform feature fusion and association mining on the multi-dimensional temporal feature vectors, identify the real-time intent, emotional polarity and business scenario tags of the call subject, and construct a dynamic semantic state machine that represents the call process. The dynamic strategy orchestration center is used to generate orchestration instruction sets in real time based on the state transition of the dynamic semantic state machine and in combination with preset business logic operators, so as to realize online adjustment and scenario adaptation of call handling strategies. The resource scheduling and execution module is used to parse the orchestration instruction set and map it to the resource configuration parameters of the underlying communication network, so as to realize the coordinated scheduling of business layer semantic goals and network layer physical resources.

[0009] The intelligent call management system is characterized in that: the multimodal feature perception module includes a sub-language feature extraction unit, which is used to identify intonation fluctuations, abnormal speech rate and pause patterns in the speech signal, and, in combination with the text semantic recognition results, determine the fraud risk level of the call content.

[0010] The intelligent call management system is characterized in that: the semantic deep understanding engine supports contextual semantic tracking, and by establishing a call state cache pool, it realizes continuous monitoring of key information points and intent drift trends in long conversations.

[0011] The intelligent call management system is characterized in that: the dynamic strategy orchestration center adopts a hierarchical response mechanism; when the semantic state machine identifies a high-risk intent, it automatically retrieves an emergency interception instruction from the preset strategy library and triggers a connection interruption on the signaling side or a speech intervention on the media side.

[0012] The intelligent call management system is characterized in that: the resource scheduling execution module issues QoS control policies through the logic orchestration interface to dynamically adjust the bandwidth priority, latency guarantee parameters or media stream forwarding path of a specific call link.

[0013] The multimodal semantic understanding and dynamic strategy orchestration method is characterized by the following steps: Step 1: Data Acquisition and Alignment: Real-time acquisition of raw voice stream, ASR transcribed text sequence and signaling behavior data in the call link; timestamp and align the above multimodal data based on the NTP network time protocol clock source; and generate time-series feature vectors through multimodal feature extraction algorithms. Step 2: Semantic Deep Understanding and State Construction: Input the temporal feature vector into the semantic understanding model for correlation analysis, identify the current intent label and emotional state of the call, and construct a dynamic semantic state machine based on contextual information to monitor the migration path from the normal call state to the risk state in real time. Step 3: Dynamic strategy orchestration: Based on the current state and state transition rate of the dynamic semantic state machine, calculate the comprehensive risk score of the call, and generate a comprehensive orchestration instruction set including media intervention, signaling control and network resource scheduling in combination with preset business logic; Step 4: Resource scheduling execution: Distribute the orchestration instruction set to the network control plane and media execution plane, dynamically adjust network slice resources and QoS priorities through the SDN controller, and provide real-time feedback on the execution results to correct the strategy for the next moment.

[0014] The method is characterized in that: in step 1, the specific process of generating the temporal feature vector is as follows: the original speech stream is collected according to preset sampling parameters, and acoustic sub-language features are extracted using time-frequency transformation and acoustic feature extraction algorithms. The acoustic sub-language features include at least energy, fundamental frequency fluctuation and short-time zero-crossing rate; text semantic features are extracted using a deep neural network model, and call behavior features are extracted using a statistical probability model; the above features are weighted and fused through an attention mechanism, and the weight of the acoustic features is automatically increased when the scenario is determined to be a specific risk scenario, and finally a multi-dimensional temporal fusion feature vector is generated.

[0015] The method is characterized in that: in step 2, the specific process of semantic deep understanding and state construction is as follows: inputting the temporal fusion feature vector into a deep temporal neural network model to capture long-distance semantic dependencies; setting a time sliding window of a preset length to establish a call state cache pool, calculating the similarity between the intent vector in the current window and the intent vector in the historical window; when the intent change rate is detected to exceed a preset threshold or a specific high-risk intent label is detected, controlling the dynamic semantic state machine to migrate from "normal call state" to "risk identification state" and triggering an intent drift warning.

[0016] The method is characterized in that: in step 3, the specific process of strategy orchestration is as follows: assess the risk value of the current state based on the value assessment model, and generate an emergency interception strategy when the calculated comprehensive risk score exceeds the preset interception threshold.

[0017] The method is characterized in that: in step 4, the resource scheduling execution includes: the media execution plane uses a speech synthesis engine to synthesize intervention speech and mix it to insert it into the call link; the network control plane issues a flow table through the SDN controller to upgrade the QoS level of the call link to a high priority and reserves a preset proportion of bandwidth resources to ensure the real-time delivery and execution of intervention instructions.

[0018] This invention breaks through the limitations of traditional single-modal recognition, integrating acoustic, semantic, and behavioral features to improve the recognition rate of variant fraudulent statements by over 40%, achieving an overall fraud recognition accuracy of 95%. Through a dynamic semantic state machine and a hierarchical response mechanism, it enables long-term session tracking and online policy adjustment, improving user satisfaction by 25%. In particular, this invention solves the long-standing technical problem of the disconnect between application-layer semantic understanding and underlying network resource scheduling. By establishing a dynamic mapping mechanism between semantic states and physical resources, it achieves on-demand allocation of network resources based on real-time semantic content. Leveraging technologies such as SDN to deeply integrate services and the network layer, it not only improves evidence preservation efficiency by 50% but also provides a highly efficient and scalable solution for intelligent anti-fraud and communication resource management, breaking down hierarchical barriers. Attached Figure Description

[0019] Figure 1 This is a system architecture diagram of the intelligent call management system of the present invention; Figure 2 This is a schematic diagram of the multimodal feature perception and fusion process in an embodiment of the present invention; Figure 3 This is a schematic diagram of the top-level transition logic of the dynamic semantic state machine described in this invention; Figure 4 This is a schematic diagram of the dynamic strategy orchestration center logic described in this invention; Figure 5 This is a multi-dimensional collaborative timing diagram of the resource scheduling execution module in an embodiment of the present invention. Detailed Implementation

[0020] The present invention will be further described below with reference to the accompanying drawings: This invention includes a multimodal feature perception module, a semantic deep understanding engine, a dynamic policy orchestration center, and a resource scheduling and execution module (attached). Figure 1 The above module division is only a logical functional division. In terms of physical implementation, the modules can run together on the same computing device or be deployed in a distributed manner. The modules interact with each other and transmit instructions through standardized interfaces (such as RESTful APIs and distributed message queues), forming a closed-loop control process.

[0021] This invention collects and fuses speech acoustic features, text semantic features, and call behavior features in a converged call link in real time to construct a semantic state machine that can dynamically evolve with the call process, enabling continuous tracking of call intent, emotional state, and risk level. Based on changes in semantic state, it dynamically generates and arranges a set of handling strategy instructions, further driving the underlying communication network to adjust service quality parameters, media stream paths, and network resources in real time. This forms a closed-loop collaborative mechanism between business layer semantic understanding and network layer resource scheduling, solving the problems of low single-modal recognition accuracy, delayed policy response, and disconnect between the business layer and the network layer in traditional systems.

[0022] The multimodal feature perception module is used to collect raw speech streams, text sequences, and signaling behavior data in the call link in real time, and performs cross-modal feature alignment based on a unified clock source to generate multi-dimensional temporal feature vectors. Furthermore, this module includes a paralinguistic feature extraction unit, used to extract acoustic paralinguistic features such as intonation changes, abnormal speech rate, stress distribution, and pause patterns from the speech signal, and to correlate these features with text semantic analysis results and call behavior features. This breaks through the limitations of traditional methods that rely solely on text keyword filtering, effectively improving the ability to identify colloquial expressions, noise interference, and variant fraudulent statements.

[0023] A semantic deep understanding engine is used to perform feature fusion and association mining on the multi-dimensional temporal feature vectors, identify the real-time intent, emotional polarity, and business scenario tags of the call subject, and construct a dynamic semantic state machine representing the call process. In a preferred embodiment, the semantic deep understanding engine supports a contextual semantic tracking mechanism. By establishing a call state cache pool, it continuously monitors key information points and intent change trends in long-term conversations, thereby achieving real-time perception of the call context shift and solving the problem that traditional static feature matching methods cannot cope with changes in the conversation state.

[0024] The dynamic policy orchestration center is used to generate and adjust the call handling policy instruction set in real time based on the state changes of the dynamic semantic state machine and in combination with preset business logic operators, thereby realizing online orchestration and scenario-based adaptation of call control policies. Furthermore, the dynamic policy orchestration center adopts a hierarchical response mechanism. When high-risk intentions or abnormal behaviors are identified, it automatically calls the corresponding intervention or interception instructions from the preset policy library and triggers signaling-side connection control or media-side scripted intervention, thereby overcoming the rigidity of traditional hard-coded decision logic and achieving smooth switching and scenario adaptation of handling policies during the call.

[0025] The resource scheduling execution module parses the orchestration instruction set and maps it to resource configuration parameters of the underlying communication network, enabling coordinated scheduling between service layer semantic objectives and network layer physical resources. Furthermore, the resource scheduling execution module issues quality of service control policies through the network control interface to dynamically adjust the bandwidth priority, latency guarantee parameters, or media stream forwarding paths of specific call links. This achieves differentiated resource allocation for high-risk or high-value calls, enhancing the communication network's support for intelligent call services.

[0026] This invention realizes multimodal semantic understanding and dynamic strategy orchestration, including the following steps: 1) Collect voice, text, and call behavior features during the call, and extract fused feature vectors using a multimodal fusion algorithm; 2) Based on the fused feature vector, semantic recognition is performed to establish a dynamically changing intent probability distribution model and identify variant fraudulent language features; 3) Based on the identified intent and risk probability, arrange call control commands in real time and calculate the network resource configuration items required to execute the command; 4) Synchronize orchestration instructions to the network control plane and media plane, and achieve dynamic closed-loop coordination between call services and underlying resources by dynamically adjusting network QoS priorities and routing policies.

[0027] In step 2), the intent recognition process introduces a voiceprint fingerprint and speech template comparison mechanism to perform real-time matching and classification of fraud patterns in dynamic migration.

[0028] In step 3), the strategy orchestration process supports the insertion of intermediate states, that is, without hanging up the call, AI voice assistants, transcription nodes or third-party manual takeover nodes are dynamically attached to the call link according to semantic changes.

[0029] In step 4), the coordinated adjustment of network resources includes dynamically rerouting or allocating sliced ​​resources to the bearer network path through the SDN control plane based on the importance of the call intent.

[0030] like Figure 2 As shown, the multimodal feature perception module serves as the data entry point and is deployed at the edge node of the communication network (such as an operator's base station server, using a high-performance processor with ≥32GB of memory). The specific processing flow of this part is divided into three levels: data acquisition, feature extraction, and feature fusion. 1. Edge Node Data Acquisition Layer: The system first accesses three types of raw data from the call link in real time: raw voice stream with a sampling rate of 8kHz and a bit depth of 16bit (compliant with telecommunications voice standards); text sequences transcribed by the ASR engine; and signaling behavior data from the core network (including call frequency and timing patterns). To ensure the consistency of multimodal data on the time axis, the system introduces an NTP clock synchronization mechanism to perform millisecond-level timestamp marking and alignment on the above three data streams (time deviation controlled within 10ms).

[0031] 2. Feature extraction layer (three-way parallel processing): (1) Acoustic sub-language feature extraction: The speech stream is processed using STFT and MFCC algorithms to extract multidimensional acoustic parameters containing 12 MFCC coefficients and 1 energy term. At the same time, the fundamental frequency fluctuation (50Hz-800Hz range) and short-term zero-crossing rate (threshold 0.05) are monitored, and abnormal pauses lasting more than 1.5 seconds are detected by combining VAD technology. The results are fused with the text semantic results to determine the fraud risk (e.g., when the speech rate is >80 words / minute and the keyword "emergency transfer" appears, the risk score increases by 20%).

[0032] (2) Text semantic feature extraction: An ASR model based on a deep neural network architecture with attention mechanism (trained on a dataset containing 50,000 labeled samples in the field of telecommunications fraud) is used to transcribe speech into text and input it into a pre-trained language model intent classifier (confidence threshold 0.7) to identify 8 preset intent labels (impersonating customer service, impersonating acquaintances, inducing transfer / remittance, inducing download / clicking links, obtaining sensitive privacy, marketing / harassment intent, normal / compliant communication).

[0033] (3) Signaling behavior feature extraction: HMM is used to model historical call behavior (such as high-frequency calls and late-night calls), with 1 million labeled calls as training data, and a risk base score of 0-100 is calculated.

[0034] 3. Feature Fusion Layer: The three sets of features mentioned above converge to the feature fusion unit. This unit uses an attention mechanism to weight the features. Specifically, in scenarios initially identified by the system as "fraud risk," the weight allocation strategy is dynamically adjusted, automatically increasing the weight of the acoustic features to 0.4 to capture the tension or pressure in the scammer's tone. Finally, the system outputs an aligned 48-dimensional temporal fusion feature vector (time step of 1 second) as input for downstream inference.

[0035] like Figure 3 As shown, the semantic deep understanding engine receives a 48-dimensional feature vector from upstream, constructs and maintains a dynamic semantic state machine, and monitors the evolution of the call state in real time. This state machine contains four core state nodes, S0 to S3, and its transition logic is as follows: S0 Normal Call Status: In the initial stage of a call, all semantic and behavioral indicators are within the baseline range.

[0036] S1 Suspicious Warning State: Trigger Condition 1: When sensitive keywords (such as "transfer", "secure account") or a local increase in the rate of change of intent are detected, the state machine transitions from S0 to S1. At this time, the system starts the key monitoring mode.

[0037] S2 Risk Identification Status: Transition Logic: In state S1, the system uses BiLSTM to capture long-distance semantic dependencies (e.g., the association between "consultation" in the first 30 seconds and "inducement" in the last 60 seconds). If the cross-modal attention mechanism detects the simultaneous occurrence of "abnormal drop in tone (<100Hz)" and "negative sentiment words," it confirms the existence of a deceptive intent combination, and the state machine transitions to state S2. Rollback mechanism: If the indicator returns to normal within a long sliding window (60 seconds), the state can be rolled back to S0.

[0038] S3 confirms fraud / intervention status: Triggering condition 2 (normal path): In state S2, if a typical graph sequence pattern is matched (such as "consultation" - "inducement" - "request for verification code"), and accompanied by an abnormal speech rate (>80 words / minute), the state transitions to state S3, triggering the final interception.

[0039] Triggering condition 3 (emergency jump path): such as Figure 3 As shown on the right, if the comprehensive risk base score calculated by the system is directly ≥80 points, or if the system detects an extremely high-risk combination of "emergency transfer + high-speed speech", the state machine will skip the intermediate state and jump directly from S0 or S1 to S3.

[0040] The dynamic policy orchestration center operates on computing nodes at the network edge. To achieve lightweight operation and rapid startup, the system is deployed using containerization technology. By leveraging lightweight virtualization at the edge, the system can process data and respond to policies close to the user end, significantly reducing network latency for data transmission back to the cloud.

[0041] like Figure 4 As shown, the system mainly consists of four core functional modules: core decision engine, orchestration core module, orchestration instruction set generation module, and strategy adaptation and optimization closed-loop module.

[0042] The core decision engine is the "brain" of the system. It adopts a three-level decision logic architecture and aims to solve the technical pain points of traditional hard-coded rules, which lack flexibility and cannot cope with the dynamic migration of fraud methods.

[0043] Level 1: Strategy Evaluation This layer is responsible for predicting the value of macro-level strategies. The system models the value of the current session state through the strategy evaluation module. This embodiment uses a dynamic scoring algorithm based on a multi-dimensional value evaluation model to construct the value function. To maximize risk control capabilities while also considering user experience and resource efficiency, the value function... The calculation formula is configured as follows:

[0044] in: This represents the normalized risk reduction rate (range [0,1], the higher the value, the better the strategy's risk suppression effect). This represents the normalized predicted user satisfaction value (range [0,1], the higher the value, the higher the user acceptance). This represents the normalized resource consumption rate (value range [0,1], including computational resources and network bandwidth consumption).

[0045] Introduced in the formula This item represents resource conservation benefits, aiming to ensure that, when risk and experience indicators are similar, the system prioritizes lightweight strategies with lower computational overhead and less bandwidth consumption, thereby optimizing the overall system performance.

[0046] Level Two: Action Selection and Risk Assessment This level includes a rule base unit and a risk score calculation unit, which work in parallel or serially and execute key hierarchical response mechanisms based on the risk score.

[0047] 1. Rule Base Unit: It has at least 12 pre-set scenario templates. For example, when the scenario is identified as "financial fraud", the standard action template matched is "hold call + voice prompt + human intervention".

[0048] 2. Risk score calculation unit: Calculates the risk score of the current session in real time based on the log-likelihood ratio algorithm.

[0049] 3. Tiered response mechanism (key feature): The system has a risk threshold set (80 points in this embodiment).

[0050] (1) High-risk bypass response (emergency channel): When the risk score is ≥80, the system determines that it is a high-risk intent and immediately triggers the "emergency response channel". This channel bypasses the third-level multi-objective optimization through the bypass mechanism and directly sends interception instructions (such as signaling interruption or strong speech intervention) to the instruction set generation module, thereby ensuring that the end-to-end response delay is less than 500ms.

[0051] (2) Normal response: When the risk score is <80, the decision flow is transferred to the third-level multi-objective optimization module.

[0052] Level 3: Multi-objective optimization For non-emergency situations, this module uses a weighted summation method to balance multiple conflicting indicators. The system has dynamic weight adjustment capabilities: when the environment is detected to be in a medium-to-high risk range, the weight of the blocking rate indicator is automatically increased to 0.5 to prioritize security while also considering user experience.

[0053] The decision intent output by the core decision engine is transmitted to the orchestration core module. This module has a built-in semantic state machine to maintain the context state of the session. The state machine performs state transitions based on the input decision intent and the current semantic state, transforming abstract decision logic into concrete, executable orchestration primitives, thus solving the "semantic gap" problem between intent and execution.

[0054] The instruction set generation module receives input from either the semantic state machine (normal channel) or the emergency response channel (high-risk channel) and generates an instruction set containing four dimensions of parameters: 1. Media Parameters: Used to control the behavior of the media server. For example, selecting a specific script template (such as "Beware of financial security") and specifying the fundamental frequency range of the synthesized speech to be between 200-250Hz to enhance the warning effect or simulate specific character characteristics.

[0055] 2. Signaling parameters: Used to control the communication link. For example, executing call forwarding to anti-fraud agents, setting the maximum hold threshold to 30 seconds; in an emergency, generating an immediate hang-up (signaling interruption) command.

[0056] 3. Network parameters: Used to ensure communication quality. For example, the QoS of relevant data packets is marked as the Accelerated Forwarding (EF) level of Differential Service Code Point (DSCP) to ensure high-priority network transmission.

[0057] 4. Management Parameters: Used for auditing and evidence preservation. For example, generating a JSON report containing full-chain data for work orders, pushing it to the BOSS system, uploading the data, and performing judicial evidence preservation.

[0058] To achieve self-evolution of the system, this embodiment constructs a policy adaptation and optimization module within the system, which includes two parallel feedback mechanisms: Mechanism A: Historical Case Transfer Based on K-NN The system maintains a historical case database of 100,000 entries and uses the K-Nearest Neighbors (K-NN) algorithm to calculate the similarity between the current scenario and historical cases. When the similarity is greater than 0.8, the system automatically transfers the parameters of historical successful cases to the current strategy.

[0059] Example: If a shorter prompt sound is more effective for historically similar cases, the system can dynamically adjust the prompt duration parameter from the default 15 seconds to 12 seconds.

[0060] Mechanism B: Real-time feedback optimization The system uses the gradient descent algorithm to fine-tune the operating parameters in real time. The system monitors real-time interactive data (such as user interruption rate) and updates the model parameters accordingly.

[0061] Example: When the user interruption rate is detected to be >30%, the system automatically shortens the speech playback duration to 8 seconds to reduce user aversion and maintain the session.

[0062] like Figure 5 As shown, once the dynamic policy orchestration center generates a comprehensive instruction set containing media, signaling, network, and management parameters, the resource scheduling execution module will achieve closed-loop coordination across planes through the following timing process: 1. Command issuance and synchronization ( Figure 5 (Steps 1-2) The orchestration center publishes the instructions to the distributed message queue. The media plane, signaling plane, network plane, and management plane act as consumers, receiving their respective instruction fragments during the synchronous execution phase. The system requires that the instruction reception time deviation of each plane be less than 100ms.

[0063] 2. Parallel execution of actions ( Figure 5 Step 3): Media-facing actions (3a): After receiving the instruction, the TTS engine based on the deep neural network architecture synthesizes the intervention voice message "Pay attention to fund security" and mixes it into the current call channel in real time with a volume ratio of 60%.

[0064] Signaling plane action (3b): Log through the SIP / TCAP protocol stack, or trigger call transfer / hang-up under extremely high risk.

[0065] Network plane action (3c): The SDN controller enables a "fast response channel," directly matching the risk level parameters in the current command based on a pre-loaded QoS policy template, thus avoiding the complex network-wide route recalculation process of the control plane. The controller issues a flow table through the OpenFlow southbound interface, instantly elevating the QoS priority of the call link to DSCP EF level and implementing network slicing isolation (reserving 10% bandwidth). Thanks to the localized matching mechanism of the policy template, the end-to-end QoS adjustment latency from command reception to flow table activation is controlled within 50ms, fully meeting the real-time transmission requirements of the call voice stream and ensuring that the intervention voice is uninterrupted and without delay.

[0066] Management actions (3D): Generate work orders and push them to the BOSS system; upload data and perform judicial evidence preservation.

[0067] 3. Feedback and closed loop ( Figure 5 (Steps 4-5): After each execution plane completes its action, it returns an execution status ACK (4a, 4b, 4c) to the orchestration center. The orchestration center summarizes the feedback results. If the overall response latency exceeds 500ms or the blocking rate does not meet the standard, the next round of strategy optimization will be triggered (such as shortening the prompt duration).

[0068] The workflow of this invention in combating telecom fraud is as follows: After a normal call is initiated, the multimodal feature perception module collects data in real time ( Figure 2 As shown, this illustrates the process of acquiring, aligning, extracting, and fusing multimodal data; the semantic deep understanding engine builds a state machine to monitor fraudulent intent (e.g., when the intent probability > 0.7, it migrates to a high-risk area). Figure 3 The evolution of the call status from normal to risky; the dynamic strategy orchestration center generates instructions (such as inserting an AI voice assistant to intervene in a "cold calling" scenario); the resource scheduling execution module adjusts synchronously ( Figure 5 The sequence diagram prioritizes QoS with a latency of <50ms. If the risk level is ≥80, a human node is automatically attached or the connection is terminated, while recording evidence is stored. The entire process does not interrupt the user experience. Verified through actual deployment (based on 10,000 real call data), the overall anti-fraud efficiency is improved by 35%, and the EER is less than 2.5% (20-second voice sample).

[0069] Note: ASR: Automatic Speech Recognition, a speech-to-text technology used to extract sequences of spoken text.

[0070] SDN: Software-Defined Networking, used for dynamic resource scheduling.

[0071] DSCP: Differentiated Services Code Point, a 6-bit identifier used in the IP packet header.

[0072] EF: Expedited Forwarding, is a very high-level identifier in DSCP.

[0073] HMM: Hidden Markov Model, used for behavioral risk prediction.

[0074] BiLSTM: Bidirectional Long Short-Term Memory, a network used for temporal semantic parsing.

[0075] TTS: Text-to-Speech, used to generate intervention scripts.

[0076] NTP: Network Time Protocol, used for clock synchronization.

[0077] EER: Equal Error Rate, a metric used to measure the accuracy of a system.

[0078] MFCC: Mel-Frequency Cepstral Coefficients, used for speech feature extraction.

Claims

1. An intelligent call management system based on multimodal semantic understanding and dynamic strategy orchestration, characterized in that... It includes a multimodal feature perception module, a semantic deep understanding engine, a dynamic policy orchestration center, and a resource scheduling and execution module; The multimodal feature perception module is used to collect raw voice streams, text sequences and signaling behavior data in the call link in real time, and perform cross-modal feature alignment based on a unified clock source to generate multi-dimensional time-series feature vectors. The semantic deep understanding engine is used to perform feature fusion and association mining on the multi-dimensional temporal feature vectors, identify the real-time intent, emotional polarity and business scenario tags of the call subject, and construct a dynamic semantic state machine that represents the call process. The dynamic strategy orchestration center is used to generate orchestration instruction sets in real time based on the state transition of the dynamic semantic state machine and in combination with preset business logic operators, so as to realize online adjustment and scenario adaptation of call handling strategies. The resource scheduling and execution module is used to parse the orchestration instruction set and map it to the resource configuration parameters of the underlying communication network, so as to realize the coordinated scheduling of business layer semantic goals and network layer physical resources.

2. The intelligent call management system according to claim 1, characterized in that: The multimodal feature perception module includes a sub-language feature extraction unit, which is used to identify intonation fluctuations, abnormal speech rate and pause patterns in speech signals, and, in combination with text semantic recognition results, determine the fraud risk level of the call content.

3. The intelligent call management system according to claim 1, characterized in that: The semantic deep understanding engine supports contextual semantic tracking and establishes a call state cache pool to achieve continuous monitoring of key information points and intent drift trends in long sessions.

4. The intelligent call management system according to claim 1, characterized in that: The dynamic policy orchestration center adopts a hierarchical response mechanism; when the semantic state machine identifies a high-risk intent, it automatically retrieves an emergency interception instruction from the preset policy library and triggers a connection interruption on the signaling side or verbal intervention on the media side.

5. The intelligent call management system according to claim 1, characterized in that: The resource scheduling execution module issues QoS control policies through the logical orchestration interface, dynamically adjusting the bandwidth priority, latency guarantee parameters, or media stream forwarding path of specific call links.

6. A multimodal semantic understanding and dynamic strategy orchestration method, characterized in that, Includes the following steps: Step 1: Data Acquisition and Alignment: Real-time acquisition of raw voice stream, ASR transcribed text sequence and signaling behavior data in the call link; timestamp and align the above multimodal data based on the NTP network time protocol clock source; and generate time-series feature vectors through multimodal feature extraction algorithms. Step 2: Semantic Deep Understanding and State Construction: Input the temporal feature vector into the semantic understanding model for correlation analysis, identify the current intent label and emotional state of the call, and construct a dynamic semantic state machine based on contextual information to monitor the migration path from the normal call state to the risk state in real time. Step 3: Dynamic strategy orchestration: Based on the current state and state transition rate of the dynamic semantic state machine, calculate the comprehensive risk score of the call, and generate a comprehensive orchestration instruction set including media intervention, signaling control and network resource scheduling in combination with preset business logic; Step 4: Resource scheduling execution: Distribute the orchestration instruction set to the network control plane and media execution plane, dynamically adjust network slice resources and QoS priorities through the SDN controller, and provide real-time feedback on the execution results to correct the strategy for the next moment.

7. The method according to claim 6, characterized in that: In step 1, the specific process of generating the temporal feature vector is as follows: the original speech stream is collected according to the preset sampling parameters, and acoustic sub-language features are extracted using time-frequency transformation and acoustic feature extraction algorithms. The acoustic sub-language features include at least energy, fundamental frequency fluctuation and short-time zero-crossing rate; text semantic features are extracted using a deep neural network model, and call behavior features are extracted using a statistical probability model. The above features are weighted and fused using an attention mechanism. The weight of acoustic features is automatically increased when a specific risk scenario is identified, and finally a multi-dimensional temporal fusion feature vector is generated.

8. The method according to claim 6, characterized in that: In step 2, the specific process of semantic deep understanding and state construction is as follows: input the temporal fusion feature vector into the deep temporal neural network model to capture long-distance semantic dependencies; set a time sliding window of a preset length to establish a call state cache pool, and calculate the similarity between the intent vector in the current window and the intent vector in the historical window; when the intent change rate is detected to exceed a preset threshold or a specific high-risk intent label is detected, control the dynamic semantic state machine to migrate from "normal call state" to "risk identification state" and trigger intent drift warning.

9. The method according to claim 6, characterized in that: In step 3, the specific process of strategy orchestration is as follows: the risk value of the current state is evaluated based on the value assessment model, and an emergency interception strategy is generated when the calculated comprehensive risk score exceeds the preset interception threshold.

10. The method according to claim 6, characterized in that: In step 4, the resource scheduling execution includes: the media execution plane uses a speech synthesis engine to synthesize intervention speech and mix it to insert it into the call link; the network control plane issues a flow table through the SDN controller to upgrade the QoS level of the call link to high priority and reserve a preset proportion of bandwidth resources to ensure the real-time delivery and execution of intervention instructions.