A method and system for intelligent telephone answering and filtering

CN122554566APending Publication Date: 2026-08-11ZHEJIANG SHENGYI OPTICAL SENSING TECH CO LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-04-30
Publication Date
2026-08-11

AI Technical Summary

Technical Problem

[0007]本申请实施例提供了一种智能电话接听与过滤方法、系统、计算机设备和计算机可读存储介质,以至少解决相关技术中难以在突破移动设备严格音频权限限制的前提下执行智能化接听与反诈骗协同应对的问题

Benefits of technology

[0019]相比于相关技术,本申请实施例提供的智能电话接听与过滤方法和系统,通过在手机终端与真实通话对端之间引入独立运作的蓝牙通话中继终端,利用标准的蓝牙免提通信规范实现音频链路的硬件级强制接管,规避底层移动操作系统的音频总线越权读取限制,解耦底层操作系统的跨硬件平台与跨通信制式;进一步地,结合云端大模型服务端对上行与下行分离音频流执行并发的深层语义解析与声纹级突变特征比对,输出精准的动态反诈风险评估;并通过高逼真用户真实声纹克隆技术及多路硬件开关总线瞬态无缝切换机制,在AI自动代答形态与机主物理语音接管模式的交替过程中维持听感平滑连续;同时提供非侵入式的隐蔽AI监听形态,在全程不干扰通话对端感知的前提下,配合穿戴式组件向机主输出多维风险预警,进而在不改变用户常规接听动作与通信习惯的前提下,为易受骗群体及大众用户构建具备强兼容性、高安全性及实时智能对抗能力的全方位通信过滤方案。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122554566A_ABST
    Figure CN122554566A_ABST
Patent Text Reader

Abstract

This application relates to an intelligent call answering and filtering method, implemented collaboratively by a Bluetooth call relay terminal and a cloud-based big data model server. The method includes: after establishing a Bluetooth call link with a mobile terminal, the Bluetooth call relay terminal, acting as an audio input / output device, takes over the audio link of the current call on the mobile terminal; the Bluetooth call relay terminal collects call voice data from the Bluetooth call link and uploads the call voice data to the cloud-based big data model server; the cloud-based big data model server identifies and processes the received call voice data to output feedback information to the Bluetooth call relay terminal, which then responds or provides risk warnings. This application, without altering users' conventional call answering actions and communication habits, constructs a comprehensive communication filtering solution with strong compatibility, high security, and real-time intelligent countermeasure capabilities for vulnerable groups and the general public.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of mobile communications, and in particular to a method, system, computer device, and computer-readable storage medium for intelligent telephone answering and filtering. Background Technology

[0002] With the continuous evolution of mobile communication networks and internet voice technology, user terminals receive a large number of harassing calls daily from external sources, including marketing, sales pitches, and potential scams. These calls not only include traditional cellular baseband voice calls provided by operators but also widely distributed across voice call services based on IP network data transmission (such as voice calls within instant messaging applications and VoIP client network calls). This massive volume of low-quality, repetitive calls severely disrupts the daily lives of end-users, and some scams employ threatening or enticing tactics, making it extremely easy for users with weak awareness of security (such as the elderly) to have their personal privacy leaked or even suffer direct financial losses.

[0003] Currently, the following types of solutions are commonly used in related technologies to address the aforementioned communication security and harassment filtering issues; however, all of them have significant technical bottlenecks: The first type of solution involves caller ID marking and direct blocking based on carrier networks or third-party applications. This type of solution primarily relies on locally pre-built blacklists of harassing numbers or cloud-based crowdsourced marking data. During the communication handshake phase, it directly blocks calls according to preset rules or displays a rough risk warning label to the user. However, this type of blocking logic based on caller ID authentication has extremely low ability to identify new numbers generated by dynamic number spoofing software or highly spoofed numbers; furthermore, this solution cannot obtain substantive internal call interaction information after connection, making it highly susceptible to mistakenly blocking critical normal business communications due to a "one-size-fits-all" blocking strategy.

[0004] The second type of solution involves the built-in artificial intelligence (AI) voice assistant function at the bottom layer of the mobile phone operating system. This mechanism uses a system-level voice assistant built into the smartphone firmware to take over the audio focus, interacting with the caller on behalf of the user and outputting structured transcribed text on the screen. However, this architecture is forcibly tied to a specific operating system version and a hardware manufacturer's customized underlying framework, exhibiting extremely poor cross-brand compatibility and cross-platform adaptability. Furthermore, its technological ecosystem is highly closed, making it difficult for third-party wearable devices or peripheral terminals to integrate and reuse it. Users cannot obtain a consistent protection experience when switching to different brands of devices.

[0005] The third type of solution is a cloud call center solution that relies on cloud-based virtual communication links and robotic operators. This solution requires the caller to dial a specifically assigned virtual front-end number, which is then routed to the real terminal by the cloud gateway after business flow and intent recognition. This architecture mainly serves enterprise-level business distribution scenarios. For personal communication, it faces high access costs and complex call forwarding configuration barriers, and cannot be integrated with users' existing native device communication habits.

[0006] Currently, no effective solution has been proposed to address the problem that related technologies rely too heavily on the underlying control of specific mobile operating systems and are unable to perform intelligent answering and anti-fraud collaborative responses while overcoming the strict audio permission restrictions of mobile devices. Summary of the Invention

[0007] This application provides a method, system, computer device, and computer-readable storage medium for intelligent telephone answering and filtering, in order to at least solve the problem in the related art that it is difficult to perform intelligent answering and anti-fraud collaborative response without breaking through the strict audio permission restrictions of mobile devices.

[0008] In a first aspect, embodiments of this application provide an intelligent telephone answering and filtering method, implemented collaboratively by a Bluetooth call relay terminal and a cloud-based large model server, the method comprising: After the Bluetooth call relay terminal establishes a Bluetooth call link with the mobile terminal, it takes over the audio link of the current call of the mobile terminal as a call audio input / output device. The Bluetooth call relay terminal collects call voice data in the Bluetooth call link and uploads the call voice data to the cloud big model server; The cloud-based large model server processes the received voice data and outputs feedback information to the Bluetooth call relay terminal, which then responds or provides risk warnings.

[0009] In some embodiments, the Bluetooth call relay terminal collects call voice data in the Bluetooth call link and uploads the call voice data to a cloud-based big data model server, including: In response to the instruction to select the AI-assisted answering mode, the Bluetooth call relay terminal collects downlink voice data from the other end of the call, streams the downlink voice data to the cloud big model server, and reduces or turns off the uplink gain of the host's microphone. In response to the instruction to select the AI ​​listening mode, the Bluetooth call relay terminal collects downlink voice data from the other end of the call and uplink voice data input by the microphone of the host end, and synchronously uploads the voice data of both parties to the cloud big model server for dialogue context analysis.

[0010] In some embodiments, in the AI-assisted answering mode, the cloud-based large model server performs recognition processing on the received call voice data to output feedback information to the Bluetooth call relay terminal, including: The cloud-based large model server performs speech recognition and intent recognition on the downlink speech data; Based on the results of intent recognition, a fraud risk assessment is performed using preset rules to determine the risk level of the current call. Based on the risk level and preset user policy, a response text is generated. The response text is then synthesized into corresponding synthetic speech data using a pre-configured user voice file and sent to the Bluetooth call relay terminal. The Bluetooth call relay terminal feeds the received synthesized voice data back into the uplink injection channel of the Bluetooth call link to output response voice to the other end of the call.

[0011] In some embodiments, the response text is synthesized into corresponding synthesized speech data using a pre-configured user voice file, including: A timbre configuration file is constructed based on the voice samples pre-recorded by the owner, and voiceprint features are extracted from the voice samples to generate timbre parameters corresponding to the owner's real timbre. The response text is synthesized into synthetic voice data that matches the actual voice characteristics of the phone owner based on the timbre parameters, so as to switch the current call between AI-assisted answering mode and owner-controlled mode for the other end of the call to obtain.

[0012] In some embodiments, in the AI ​​monitoring mode, the cloud-based large model server performs recognition processing on the received call voice data to output feedback information to the Bluetooth call relay terminal, including: The cloud-based large model server performs speech recognition and intent recognition on the downlink speech data; Based on the results of intent recognition, a fraud risk assessment is performed using preset rules to determine the risk level of the current call. Based on the risk level, a risk warning message containing the risk level and a description of suspicious points is generated, and the risk warning message is returned to the Bluetooth call relay terminal; Without interfering with the other end of the call, the Bluetooth call relay terminal outputs the risk warning information to the owner through its local risk warning module.

[0013] In some embodiments, the cloud-based large model server performs recognition processing on the received call voice data and outputs feedback information including: Semantic and acoustic features are extracted from the call voice data. The semantic features include: sensitive keywords, intent slots, and specific expressions in dialogue turns. The acoustic features include: sudden changes in speech rate, abnormal pauses, and background noise. The semantic features are matched with preset anti-fraud rules and fraud script knowledge base to obtain rule hit results and script matching results. The semantic features are input into a large language model for intent recognition to obtain the intent recognition results and corresponding confidence levels. A comprehensive risk score is calculated based on the rule hit result, the speech matching result, the confidence level, and the anomaly score corresponding to the acoustic feature. When the comprehensive risk score reaches the corresponding preset risk threshold, the corresponding risk fraud level is output, and the hit semantic features are fed back to the strategy engine in the form of structured fields.

[0014] In some embodiments, in the AI-assisted answering mode, the method further includes: a human-machine takeover switching step: In response to a preset takeover operation command triggered by the owner, the Bluetooth call relay terminal controls the local audio interface module to switch the uplink input source from the synthesized speech data to local microphone speech data. The Bluetooth call relay terminal sends a takeover signal to the cloud-based large model server, so that the cloud-based large model server switches from response generation mode to pure listening and recording mode. The call audio data is uploaded to the cloud-based large model server for real-time transcription and content storage, and the writing of new synthesized speech data to the Bluetooth call link is stopped.

[0015] In some embodiments, the method for establishing a network session connection between the Bluetooth call relay terminal and the cloud-based large model server includes: The Bluetooth call relay terminal can independently access the public network and establish a network session connection with the cloud-based large model server through its built-in wireless LAN module or cellular data module; or... The short-range data communication unit within the Bluetooth call relay terminal establishes a data forwarding channel with the mobile client application module, and establishes a network session connection with the cloud-based large model server via the public network of the mobile terminal.

[0016] Secondly, embodiments of this application provide an intelligent telephone answering and filtering system, which is implemented collaboratively by a Bluetooth call relay terminal and a cloud-based large model server, wherein: The Bluetooth call relay terminal is used to take over the audio link of the current call of the mobile terminal as a call audio input / output device after establishing a Bluetooth call link with the mobile terminal. In addition, the system collects voice data from the Bluetooth call link and uploads the voice data to the cloud-based big data model server. The cloud-based large model server is used to identify and process the received call voice data, and output feedback information to the Bluetooth call relay terminal, through which the Bluetooth call relay terminal responds or provides risk warnings.

[0017] Thirdly, embodiments of this application provide a computer device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the method described in the first aspect above.

[0018] Fourthly, embodiments of this application provide a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the method described in the first aspect above.

[0019] Compared to related technologies, the intelligent call answering and filtering method and system provided in this application introduces an independently operating Bluetooth call relay terminal between the mobile terminal and the real caller. It utilizes standard Bluetooth hands-free communication specifications to achieve hardware-level forced takeover of the audio link, circumventing the unauthorized reading restrictions of the underlying mobile operating system's audio bus, and decoupling the underlying operating system across hardware platforms and communication standards. Furthermore, it combines cloud-based large-model servers to perform concurrent deep semantic analysis and voiceprint-level mutation feature comparison of the uplink and downlink separated audio streams, outputting accurate dynamic anti-fraud risk assessments. Through highly realistic user voiceprint cloning technology and a multi-channel hardware switch bus transient seamless switching mechanism, it maintains a smooth and continuous listening experience during the alternation between AI automatic answering mode and the owner's physical voice takeover mode. Simultaneously, it provides a non-intrusive, covert AI monitoring mode, outputting multi-dimensional risk warnings to the owner in conjunction with wearable components without interfering with the caller's perception. Thus, without changing the user's regular answering actions and communication habits, it constructs a comprehensive communication filtering solution with strong compatibility, high security, and real-time intelligent countermeasure capabilities for vulnerable groups and the general public. Attached Figure Description

[0020] The accompanying drawings, which are included to provide a further understanding of this application and form part of this application, illustrate exemplary embodiments of this application and are used to explain this application, but do not constitute an undue limitation of this application. In the drawings: Figure 1 This is a schematic diagram of an intelligent telephone answering and filtering system according to an embodiment of this application; Figure 2 This is a flowchart of an intelligent telephone answering and filtering method according to an embodiment of this application; Figure 3 This is a schematic diagram of the internal structure of an electronic device according to an embodiment of this application. Detailed Implementation

[0021] To make the objectives, technical solutions, and advantages of this application clearer, the application is described and illustrated below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application. All other embodiments obtained by those skilled in the art based on the embodiments provided in this application without inventive effort are within the scope of protection of this application.

[0022] Obviously, the accompanying drawings described below are merely some examples or embodiments of this application. Those skilled in the art can apply this application to other similar scenarios based on these drawings without any inventive effort. Furthermore, it is understood that although the efforts made in this development process may be complex and lengthy, for those skilled in the art related to the content disclosed in this application, any changes to design, manufacturing, or production based on the technical content disclosed in this application are merely conventional technical means and should not be construed as insufficient disclosure of the content of this application.

[0023] In this application, the reference to "embodiment" means that a specific feature, structure, or characteristic described in connection with an embodiment may be included in at least one embodiment of this application. The appearance of this phrase in various places in the specification does not necessarily refer to the same embodiment, nor is it a separate or alternative embodiment that is mutually exclusive with other embodiments. It will be explicitly and implicitly understood by those skilled in the art that the embodiments described in this application may be combined with other embodiments without conflict.

[0024] Unless otherwise defined, the technical or scientific terms used in this application shall have the ordinary meaning understood by one of ordinary skill in the art to which this application pertains. The terms “a,” “an,” “an,” “the,” and similar words used in this application do not indicate quantity limitation and may indicate singular or plural. The terms “comprising,” “including,” “having,” and any variations thereof used in this application are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or device that includes a series of steps or modules (units) is not limited to the listed steps or units, but may also include steps or units not listed, or may include other steps or units inherent to these processes, methods, products, or devices. The terms “connected,” “linked,” “coupled,” and similar words used in this application are not limited to physical or mechanical connections, but may include electrical connections, whether direct or indirect. “Multiple” used in this application refers to two or more. “And / or” describes the relationship between related objects, indicating that three relationships may exist; for example, “A and / or B” can represent: A alone, A and B simultaneously, and B alone. The character " / " generally indicates that the preceding and following objects are in an "or" relationship. The terms "first," "second," and "third" used in this application are merely to distinguish similar objects and do not represent a specific ordering of the objects.

[0025] The various technologies described in this application can be used in a variety of communication environments. This embodiment provides an intelligent telephone answering and filtering system. Figure 1 This is a schematic diagram of an intelligent telephone answering and filtering system according to an embodiment of this application, such as... Figure 1 As shown, the system includes at least a mobile terminal 10, a Bluetooth call relay terminal 20, and a cloud-based large model server 30.

[0026] The mobile terminal operates as any smart terminal device that supports Bluetooth calling or Bluetooth audio output, and it has the ability to simultaneously support cellular voice calls from operators and voice call services based on IP networks (such as voice calls in instant messaging applications and VoIP client calls). During the call execution phase, the mobile terminal switches the call audio input and output channels to the Bluetooth call relay terminal, and it only serves as a gateway for maintaining call signaling and accessing the cellular / data network.

[0027] A Bluetooth call relay terminal serves as a physical audio relay node between a mobile terminal and a cloud-based large-scale model server. Its physical form includes, but is not limited to, a standalone relay hub, a Bluetooth headset with integrated relay functionality, smart glasses, or an in-vehicle infotainment system. Internally, a Bluetooth call relay terminal mainly comprises: a Bluetooth call module, an audio interface module, a main control processing unit, a network communication module, a human-computer interaction module, and a risk warning module.

[0028] Specifically, the Bluetooth calling module supports calling protocols such as HFP (Hands-Free Profile) or HSP (Headset Profile) of the Bluetooth Classic specification to establish a call audio link between the hands-free or headset user and the mobile terminal. When the mobile terminal switches the current audio device to this relay terminal, the microphone and speaker hardware of the mobile phone are disabled, and the two-way call audio stream is completely controlled by the relay terminal. The audio interface module is connected to the Bluetooth calling module and the main control processing unit via I²S or PCM bus, and includes a downlink acquisition channel, an uplink injection channel, and a local monitoring channel. The main control processing unit (MCU or SoC) schedules and moves the PCM data frames in each channel through DMA (Direct Memory Access) or hardware interrupts. The input of the uplink injection channel is connected to a hardware multiplexer (MUX), corresponding to two candidate audio sources: the actual speech from the local microphone and the TTS synthesized speech returned from the cloud. The network communication module provides independent access capabilities for Wi-Fi and cellular data, or connects to the cloud-based large model server via a public network data link attached to the mobile terminal through a short-range data communication unit (Bluetooth RFCOMM, PAN, or BLE). The risk warning module uses a bone conduction sound unit, vibration motor, or micro-display component to output concealed alarms.

[0029] The cloud-based large-scale model server is deployed on a distributed cloud platform and includes an Automatic Speech Recognition (ASR) module, a call understanding and large-scale model dialogue module, an anti-fraud identification module, a policy engine and dialogue management module, a text-to-speech (TTS) module, and a data storage and summarization module. The ASR module performs streaming transcription of the calling and called party voices by channel; the large-scale model dialogue module performs deep semantic parsing and intent slot extraction; the anti-fraud identification module integrates acoustic mutation features and semantic high-risk features to output a dynamic risk assessment score; the policy engine outputs a dynamic response strategy based on the assessment score; and the TTS module loads preset user voiceprint configuration parameters to generate high-fidelity cloned voice and transmits it back to the terminal.

[0030] Based on the aforementioned system hardware and cloud software collaborative architecture, this embodiment details the complete processing flow of an intelligent call answering and filtering method. This execution flow covers the entire lifecycle of processing actions, from pre-call preparation to subsequent structured data accumulation, to address the technical pain points of intelligent anti-harassment and age-appropriate anti-fraud measures. Figure 2 This is a flowchart of an intelligent telephone answering and filtering method according to an embodiment of this application, such as... Figure 2 As shown, the process includes the following steps: S201, Establish network session connection; Before or at the initial stage of a call, the Bluetooth call relay terminal obtains public network access permissions and establishes a two-way data transmission session with the cloud-based big data model server.

[0031] In the first implementation, the Bluetooth call relay terminal activates its built-in wireless LAN module or cellular data module, independently obtains an IP address, and directly establishes a TCP / WebSocket long connection with the cloud-based big data server. This method reduces the dependence on the network status of the mobile terminal.

[0032] In the second implementation, the Bluetooth relay terminal activates its short-range data communication unit to establish a data forwarding channel with the mobile client application module. This data forwarding channel is established via Bluetooth Serial Port Emulation (RFCOMM), Personal Area Network (PAN), or Bluetooth Low Energy (BLE). The mobile client application module acts as a network gateway, transmitting the received relay terminal data packets to the cloud-based large-scale model server via the mobile terminal's 4G / 5G or Wi-Fi network.

[0033] Step S202, call link takeover and status triggering; When a mobile terminal detects an incoming or outgoing voice call, the Bluetooth call relay terminal establishes a Bluetooth call link with the mobile terminal and then takes over the audio link of the current call as an audio input / output device.

[0034] Furthermore, during the call, the Bluetooth call relay terminal monitors the input signals from the human-computer interaction module (such as preset key combinations, touch gestures, or specific wake words). Upon receiving the input signal, in response to the received intelligent processing trigger command, the main control processing unit generates a global session identifier (SessionID) for the current call and sends a session initiation control frame containing the session identifier to the cloud-based large model server via the network communication module. Simultaneously, the main control processing unit reads the mode parameters carried in the trigger command to determine whether to enter "AI response mode" or "AI listening mode".

[0035] Step S203: Audio acquisition and uploading from relay terminal; The Bluetooth call relay terminal collects call voice data from the Bluetooth call link and uploads the call voice data to the cloud-based large model server in a streaming, frame-by-frame format. Based on the communication pattern determined in step S202, this step executes differentiated collection and upload logic, specifically including: In response to the command to select the AI-assisted answering mode, the main control processing unit of the Bluetooth call relay terminal controls the audio interface module to extract only the downlink PCM voice data sent by the other end of the call from the downlink acquisition channel. After performing frame segmentation, buffering, and lossless / lossy compression processing according to a preset time window, the main control processing unit streams the data to the cloud. At the same time, the main control processing unit sends control levels to the hardware multiplexer (MUX) to cut off or greatly reduce the gain of the uplink channel from the user's microphone input channel to the Bluetooth call module, ensuring that the user's ambient sound is not accidentally transmitted to the other end of the call.

[0036] In response to the command to select AI listening mode, the Bluetooth call relay terminal activates a dual-track recording mechanism. The main control processing unit simultaneously extracts the voice from the other end of the call from the downlink acquisition channel and extracts the uplink voice from the microphone on the host end. Furthermore, after the two audio data are tagged with channel identifiers, clock synchronization and packet processing are performed, and the voice data from both parties are simultaneously uploaded to the cloud-based large model server to ensure that the cloud has the data foundation to perform complete multi-turn dialogue context coherence analysis.

[0037] In addition, in AI monitoring mode, the main control processing unit retains a loop audio cache of a preset duration (such as the most recent 30 seconds) locally to maintain the context state in subsequent mode switching that may be triggered.

[0038] Step S204: Cloud-based multi-dimensional feature extraction and intent recognition; The cloud-based large model server receives frame-by-frame uploaded call voice data, which is then transcribed into stream text by the speech recognition module. Subsequently, semantic and acoustic features are extracted from the received transcribed text and the original audio stream.

[0039] Specifically, the large model dialogue module extracts semantic features from the dialogue data. The semantic features cover information elements in multiple dimensions: including but not limited to sensitive entity keywords (such as "ID card", "bank account", "verification code"), intent slot data (such as express delivery tracking number, pickup address) and specific pragmatic expressions in the dialogue rounds (such as imperative sentences containing urging, threats, and incentives).

[0040] Simultaneously, the feature engineering module extracts acoustic features from the raw audio stream. These acoustic features include: speech rate abrupt change rate (e.g., abnormally fast speech rate in a short period of time), frequency of abnormal pauses, background noise type classification (e.g., whether there is typical call center dense keyboard sound or multiple human voice background), and mechanical repetitive broadcast-style acoustic prosodic features.

[0041] Furthermore, the large-scale model dialogue module inputs the extracted semantic features into a pre-trained large language model for deep intent recognition and classification, and outputs the intent recognition results (such as logistics notifications, business promotions, and suspected telecommunications fraud) and the corresponding probability confidence scores.

[0042] Step S205, Comprehensive Anti-Fraud Risk Assessment; The anti-fraud identification module on the cloud-based big data model server combines preset rules to assess fraud risk and determine the risk level of the current call. Specifically, the semantic features extracted in step S204 are traversed and matched with the cloud-based pre-built anti-fraud rule engine and fraud script knowledge graph to obtain rule hit results and script matching results based on regular expressions or node similarity.

[0043] Subsequently, the anti-fraud identification module performs a weighted calculation, generating a global comprehensive risk score based on the score of the rule hit result, the similarity score of the speech matching result, the confidence of the large model intent recognition, and the abnormal deviation score corresponding to the acoustic features, using a preset weighted evaluation algorithm formula.

[0044] When the comprehensive risk score reaches the corresponding preset risk threshold range at each level, the anti-fraud identification module outputs the corresponding risk level label (such as low risk, suspicious, suspected fraud or high-risk fraud), and feeds back the core semantic features and corresponding suspicious points hit in this assessment to the strategy engine and dialogue management module in the form of structured JSON fields, providing a basis for decision-making for subsequent actions.

[0045] Step S206: Strategy generation and voice feedback in AI-assisted answering mode; In the current AI-assisted response mode, the cloud-based large model server generates response text based on the risk level obtained from the assessment and the preset user strategy. It then uses pre-configured user voice files to synthesize the response text into corresponding synthetic speech data, which is then sent to the Bluetooth call relay terminal for re-feeding.

[0046] Furthermore, after receiving the risk level label, the policy engine retrieves the corresponding user's security protection policy profile (for example, if the user is identified as an elderly person, the most stringent and defensive blocking policy is activated). For routine sales calls, the policy engine generates a polite rejection response text; for calls determined to be high-risk scams, it generates an adversarial inquiry text refusing to provide any confirmation information or an alert text to immediately terminate the call.

[0047] The Text-to-Speech (TTS) module receives the response text. Prior to this, the timbre configuration submodule has performed voiceprint feature extraction and vocoder parameter training based on several standard speech samples pre-recorded by the user, constructing a timbre configuration file corresponding to the physical characteristics of the user's actual vocal organs. The TTS module loads the timbre configuration file, renders the current response text into synthesized PCM speech data containing the user's specific timbre, rhythm, and emotional features, and then sends it out.

[0048] Finally, after receiving the synthesized voice data, the main control processing unit of the Bluetooth call relay terminal writes it into the uplink buffer. The main control processing unit controls the hardware MUX of the audio interface module to accurately divert and feed back the synthesized voice data to the HFP / HSP uplink injection channel of the Bluetooth call link, and finally the mobile terminal sends the radio frequency signal of the analog machine to the other end of the call.

[0049] Step S207: Real-time monitoring and covert alarms in AI eavesdropping mode. In the current AI eavesdropping mode, the cloud-based large model server and the Bluetooth call relay terminal perform risk alert functions. During AI eavesdropping, both parties in the call are in normal communication. The cloud-based anti-fraud identification module outputs a dynamic risk level in real time according to step S205. When the comprehensive risk score exceeds the warning threshold, the cloud generates a risk alert message containing the risk level and a description of suspicious points (such as "the other party is asking for a verification code"), and quickly returns it to the Bluetooth call relay terminal via network connection.

[0050] The main control processing unit of the Bluetooth call relay terminal parses the risk warning information and, under the premise of preventing the warning tone from being mixed into the uplink channel (i.e. not interfering with the hearing of the other end of the call), drives its local risk warning module to output an alarm to the owner.

[0051] For example, the audio DSP can be controlled to overlay specific frequency band alert tones onto the downlink audio stream of the local monitoring channel, or the micro-display component of the externally connected smart glasses can be controlled to highlight risk structured fields, or the body vibration motor can be driven to perform pulse vibrations at specific frequencies. After receiving an alarm, the owner can maintain the current dialogue state or actively trigger the mode switch in step S208.

[0052] Step S208: Seamless switching between human-machine takeover and form.

[0053] It should be noted that, regardless of whether it is in AI-assisted answering mode or AI-monitoring mode, this application solution supports the transfer of call control by responding to preset takeover operation commands triggered by the phone owner.

[0054] When the user triggers a takeover command by pressing a physical button or using a specific voice command, the main control processing unit immediately sends an interrupt signal to the cloud, controlling the cloud-based large model server to stop actively generating subsequent TTS response voices, and instructing the cloud state machine to downgrade to a monitoring mode that only performs silent transcription and recording for subsequent calls.

[0055] Simultaneously, at the terminal hardware level, the main control processing unit controls the local audio interface module to perform channel switching of the multiplexer (MUX). It cuts off the synthetic speech data injection path from the uplink buffer and switches the uplink input source to the local physical microphone with a nanosecond-level response speed. Furthermore, to eliminate potential level jumps, pops, or semantic breaks during switching, the main control processing unit inserts extremely short silence compensation windows at the audio frame boundaries before and after the switch, applies a smooth fade-in / fade-out envelope algorithm, and performs echo cancellation reset.

[0056] Understandably, because the early step S206 uses high-fidelity user voice cloning technology, when the actual control is switched from AI synthesized voice to the actual physical voice of the phone owner, the other end of the call experiences a smooth transition in acoustic perception, making it extremely difficult to detect the change in the responding subject.

[0057] Step S209, Session Termination and Structured Data Processing.

[0058] When the underlying baseband of the mobile terminal detects a call hang-up signal, or when the relay terminal actively initiates a hang-up control command, the main control processing unit releases the local audio buffer and sends a session end synchronization packet to the cloud-based big model server.

[0059] Upon receiving the termination command, the cloud-based large model server halts all computational tasks for the current session. The data summary module retrieves the complete transcribed text of the call, interim intent classification tags, the final fraud risk level assessment report, and triggered prevention and control strategy records. Furthermore, the large model performs chapter-level information extraction and summarization on the aforementioned discrete data, generating a structured call summary file containing the caller ID, accurate call time, topic type definition, final risk level, a summary of core interaction content, and future handling suggestions (such as adding the caller to a blacklist).

[0060] Ultimately, the structured call summary file is pushed to the mobile client application module or the corresponding management platform, and presented to the terminal interface of the owner or their designated guardian in the form of a card view, so that they can make subsequent security policy adjustments or trace and review.

[0061] Through steps S201 to S209 above, the method disclosed in this embodiment utilizes a collaborative architecture of separate Bluetooth relay hardware and a cloud-based large-scale model to bypass the sandbox permission control of the mobile terminal operating system. This constructs a call security filter supported by large-scale model computing power without altering the user's regular call answering habits. Furthermore, by employing voice cloning technology and a smooth switching mechanism combining hardware and software, the integration of AI-assisted answering and human intervention is enhanced. Finally, by combining real-time dual-track monitoring and a concealed risk warning mechanism, an invisible security barrier is provided for the elderly and vulnerable groups.

[0062] In one embodiment, Figure 3 This is a schematic diagram of the internal structure of an electronic device according to an embodiment of this application, such as... Figure 3 As shown, an electronic device is provided, which can be a server, and its internal structure diagram can be as follows. Figure 3 As shown, the electronic device includes a processor, a network interface, internal memory, and non-volatile memory connected via an internal bus. The non-volatile memory stores an operating system, computer programs, and a database. The processor provides computing and control capabilities, the network interface communicates with external terminals via a network, the internal memory provides an environment for the operating system, the computer programs are executed by the processor to implement an intelligent telephone answering and filtering method, and the database stores data.

[0063] Those skilled in the art will understand that Figure 3 The structure shown is merely a block diagram of a portion of the structure related to the present application and does not constitute a limitation on the electronic device to which the present application is applied. The specific electronic device may include more or fewer components than shown in the figure, or combine certain components, or have different component arrangements.

[0064] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. This computer program can be stored in a non-volatile computer-readable storage medium. When executed, the computer program can include the processes of the embodiments of the above methods. Any references to memory, storage, databases, or other media used in the embodiments provided in this application can include non-volatile and / or volatile memory. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), dual data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), Rambus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM), etc.

[0065] The above embodiments merely illustrate several implementation methods of this application, and while the descriptions are relatively specific and detailed, they should not be construed as limiting the scope of the invention patent. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of this application, and these all fall within the protection scope of this application. Therefore, the protection scope of this patent application should be determined by the appended claims.

Claims

1. A smart phone answering and filtering method, characterized by, The method, implemented through collaboration between a Bluetooth call relay terminal and a cloud-based large-scale model server, includes: After the Bluetooth call relay terminal establishes a Bluetooth call link with the mobile terminal, it takes over the audio link of the current call of the mobile terminal as a call audio input / output device. The Bluetooth call relay terminal collects call voice data in the Bluetooth call link and uploads the call voice data to the cloud big model server; The cloud-based large model server processes the received voice data and outputs feedback information to the Bluetooth call relay terminal, which then responds or provides risk warnings.

2. The method of claim 1, wherein, The Bluetooth call relay terminal collects call voice data in the Bluetooth call link and uploads the call voice data to the cloud big data model server, including: In response to the instruction to select the AI-assisted answering mode, the Bluetooth call relay terminal collects downlink voice data from the other end of the call, streams the downlink voice data to the cloud big model server, and reduces or turns off the uplink gain of the host's microphone. In response to the instruction to select the AI ​​listening mode, the Bluetooth call relay terminal collects downlink voice data from the other end of the call and uplink voice data input by the microphone of the host end, and synchronously uploads the voice data of both parties to the cloud big model server for dialogue context analysis.

3. The method of claim 2, wherein, In the AI-assisted answering mode, the cloud-based large model server processes the received call voice data to output feedback information to the Bluetooth call relay terminal, including: The cloud-based large model server performs speech recognition and intent recognition on the downlink speech data; Based on the results of intent recognition, a fraud risk assessment is performed using preset rules to determine the risk level of the current call. Based on the risk level and preset user policy, a response text is generated. The response text is then synthesized into corresponding synthetic speech data using a pre-configured user voice file and sent to the Bluetooth call relay terminal. The Bluetooth call relay terminal feeds the received synthesized voice data back into the uplink injection channel of the Bluetooth call link to output response voice to the other end of the call.

4. The method of claim 3, wherein, The response text is synthesized into corresponding synthesized speech data using a pre-configured user voice file, including: A timbre configuration file is constructed based on the voice samples pre-recorded by the owner, and voiceprint features are extracted from the voice samples to generate timbre parameters corresponding to the owner's real timbre. The response text is synthesized into synthetic voice data that matches the actual voice characteristics of the phone owner based on the timbre parameters, so as to switch the current call between AI-assisted answering mode and owner-controlled mode for the other end of the call to obtain.

5. The method of claim 2, wherein, In the AI ​​monitoring mode, the cloud-based large model server processes the received call voice data to output feedback information to the Bluetooth call relay terminal, including: The cloud-based large model server performs speech recognition and intent recognition on the downlink speech data; Based on the results of intent recognition, a fraud risk assessment is performed using preset rules to determine the risk level of the current call. Based on the risk level, a risk warning message containing the risk level and a description of suspicious points is generated, and the risk warning message is returned to the Bluetooth call relay terminal; Without interfering with the other end of the call, the Bluetooth call relay terminal outputs the risk warning information to the owner through its local risk warning module.

6. The method according to any one of claims 3 and 5, characterized in that, The cloud-based large model server processes the received call voice data and outputs feedback information including: Semantic and acoustic features are extracted from the call voice data. The semantic features include: sensitive keywords, intent slots, and specific expressions in dialogue turns. The acoustic features include: sudden changes in speech rate, abnormal pauses, and background noise. The semantic features are matched with preset anti-fraud rules and fraud script knowledge base to obtain rule hit results and script matching results. The semantic features are input into a large language model for intent recognition to obtain the intent recognition results and corresponding confidence levels. A comprehensive risk score is calculated based on the rule hit result, the speech matching result, the confidence level, and the anomaly score corresponding to the acoustic feature. When the comprehensive risk score reaches the corresponding preset risk threshold, the corresponding risk fraud level is output, and the hit semantic features are fed back to the strategy engine in the form of structured fields.

7. The method of claim 3, wherein, In the AI-assisted answering mode, the method further includes: a human-machine takeover switching step: In response to a preset takeover operation command triggered by the owner, the Bluetooth call relay terminal controls the local audio interface module to switch the uplink input source from the synthesized speech data to local microphone speech data. The Bluetooth call relay terminal sends a takeover signal to the cloud-based large model server, so that the cloud-based large model server switches from response generation mode to pure listening and recording mode. The call audio data is uploaded to the cloud-based large model server for real-time transcription and content storage, and the writing of new synthesized speech data to the Bluetooth call link is stopped.

8. The method of claim 1, wherein, The methods for establishing a network session connection between the Bluetooth call relay terminal and the cloud-based large model server include: The Bluetooth call relay terminal can independently access the public network and establish a network session connection with the cloud-based large model server through its built-in wireless LAN module or cellular data module; or... The short-range data communication unit within the Bluetooth call relay terminal establishes a data forwarding channel with the mobile client application module, and establishes a network session connection with the cloud-based large model server via the public network of the mobile terminal.

9. An intelligent telephone answering and filtering system, characterized in that, This is achieved through collaboration between a Bluetooth call relay terminal and a cloud-based large-scale model server, where: The Bluetooth call relay terminal is used to take over the audio link of the current call of the mobile terminal as a call audio input / output device after establishing a Bluetooth call link with the mobile terminal. In addition, the system collects voice data from the Bluetooth call link and uploads the voice data to the cloud-based big data model server. The cloud-based large model server is used to identify and process the received call voice data, and output feedback information to the Bluetooth call relay terminal, through which the Bluetooth call relay terminal responds or provides risk warnings.

10. A computer device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the method as described in any one of claims 1 to 8.