Information processing system
Patent Information
- Application Number
- CN202610250697.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2025-03-19
- Filing Date
- 2026-03-03
- Publication Date
- 2026-09-22
AI Technical Summary
[0003]现有的通信应答系统通常依赖人工进行接听和判断,对于包含诱导性陈述、虚假陈述或涉嫌欺诈、伪证的通话内容,用户往往难以及时、客观地进行识别和应对,容易在情绪压力或信息不对称的情况下作出不利决定
服务器在获得虚假性评价和情感信息后,可以生成行为建议信息。行为建议信息包括具体应对动作的自然语言描述,例如“建议暂缓转账”“建议通过官方渠道核实对方资质”“建议联系法律专业人员”等。服务器将这些建议存入评价表,并通过终端展示给用户。
Smart Images

Figure CN122802622A_ABST
Abstract
Description
Technical Field
[0001] The technology disclosed herein relates to an information processing system. Background Technology
[0002] Japanese Patent Application Publication No. 2022-180282 discloses a method for controlling a role-based chatbot executed by at least one processor. The method includes the following steps: receiving a user's speech; adding the user's speech to a prompt word, the prompt word containing instruction statements associated with an explanation of the chatbot's role; encoding the prompt word; and inputting the encoded prompt word into a language model to generate a chatbot speech in response to the user's speech.
[0003] Existing communication response systems typically rely on human intervention for answering and judgment. Users often struggle to promptly and objectively identify and respond to calls containing leading statements, false statements, or suspected fraud or perjury, making them susceptible to making unfavorable decisions under emotional pressure or information asymmetry. Particularly in voice call scenarios, current technologies generally only offer simple recording functions, lacking mechanisms for automatic transcription, semantic analysis, and intelligent perjury analysis of call content. Consequently, they cannot provide users with structured judgment criteria and suggestions during or after the call.
[0004] In addition, although some AI customer service systems exist in the existing technology, these systems mostly focus on business consultation and simple Q&A, lacking a mechanism to flexibly switch from human answering to AI automatic response for incoming calls. They also lack the function of making perjury judgments and analyzing emotional states based on the content of the call and generating suggestions for the next action accordingly. Therefore, they cannot effectively assist users in making calm and reasonable decisions in high-risk or high-stress call situations.
[0005] Furthermore, in complex or adversarial call scenarios, users' emotions fluctuate greatly, and they often overlook logical contradictions, unusual demands, or potential risks of perjury in the other party's words. Existing systems fail to dynamically adjust behavioral suggestions based on sentiment analysis results, resulting in system prompts that are rather mechanical and difficult to match with the user's actual psychological state and coping ability.
[0006] Therefore, the technical problem to be solved by the present invention is to provide a system that can switch communication from manual response to artificial intelligence automatic response, and can automatically transcribe, perform perjury analysis and emotional state analysis on the voice of the call, thereby generating prompt information for perjury judgment and next action suggestions, so as to improve the user's risk identification ability and response decision quality in various call scenarios. Summary of the Invention
[0007] To address the aforementioned technical challenges, the present invention provides an information processing system comprising a processor, wherein the processor is configured to: upon receiving a communication, switch the communication from human response to artificial intelligence automatic response, so that when user selection or preset conditions are met, artificial intelligence can answer and process the call content on behalf of the user.
[0008] The processor is further configured to: acquire voice data from automatically answered or transferred calls, and convert the voice data into text data using speech recognition technology, so that the call content is represented in a calculable, storable and retrievalable text form, providing a basis for subsequent semantic analysis and perjury determination.
[0009] The processor is further configured to: parse the text data using natural language processing technology, extract key information and potential risk factors through semantic understanding, keyword recognition, pattern matching, logical relationship analysis, etc., and generate prompt information for judging perjury based on the parsing results, thereby providing structured input for the perjury judgment model inside the system or external review.
[0010] To further assist users in taking appropriate countermeasures after identifying the risk of false evidence, the processor is also configured to generate prompt information based on the result of the false evidence judgment to suggest the next action. The action suggestion may include, but is not limited to, suggesting that the user terminate the call, suggesting that the user suspend the operation of funds or privacy information, suggesting that the user verify with relevant institutions or report to the police, thereby providing users with clear behavioral references at the system level.
[0011] Furthermore, to improve the personalization and adaptability of the system's suggestions, the processor is also configured to: analyze the user's emotional state using a sentiment analysis engine, for example, by recognizing emotional features in the user's voice or text to determine whether the user is in a state of tension, fear, anger, or anxiety; and generate prompts based on the analysis results to adjust the proposed action suggestions, so that when the user's emotions fluctuate greatly or they are under great pressure, the system can provide more cautious, gradual, or more explicit coping solutions.
[0012] Through the above-mentioned technical means, the system of the present invention can flexibly switch from human response to artificial intelligence automatic response in the call scenario, realize automatic transcription of call content and intelligent analysis of perjury, and combine synchronous or post-event emotional state analysis to generate targeted perjury judgment prompts and action suggestions for users, thereby effectively improving users' ability to identify potential perjury or fraudulent calls and the quality of their decision-making.
[0013] "System" refers to a combination of hardware and software that includes at least one processor and optional memory, communication interfaces and input / output devices, used to execute program instructions to perform functions such as communication response, voice processing, text analysis and prompt information generation.
[0014] A processor is an electronic component or arithmetic unit that can execute computer program instructions, process input data, and output processing results. It can be a single central processing unit, multiple processing cores, application-specific integrated circuits, digital signal processors, or any combination thereof.
[0015] "Communication" refers to voice, data or signals transmitted through wired or wireless networks, telephone lines or other communication media, including but not limited to telephone calls, voice calls and related signaling information.
[0016] "Human response" refers to the way in which human users directly answer and process communications, including listening to incoming calls, giving verbal replies, asking questions, or making operational decisions, rather than being automatically performed by an artificial intelligence system.
[0017] "AI-powered auto-response" refers to a response method in which a system running artificial intelligence algorithms automatically answers and processes communications without real-time human intervention, generates response content based on preset rules or models, and interacts with the other party.
[0018] “Voice data” refers to data in analog or digital form that represents human speech signals collected during communication, including but not limited to audio streams, recording files, or real-time audio clips from telephone calls.
[0019] "Speech recognition technology" refers to the technical process of converting input speech data into corresponding text data, which typically includes steps such as acoustic feature extraction, acoustic modeling, language modeling, and decoding.
[0020] "Text data" refers to textual information generated from speech data by speech recognition technology, represented in the form of characters, words, sentences, etc., and can be processed by computers.
[0021] Natural Language Processing (NLP) refers to the techniques used to perform linguistic and semantic analysis on text data, including but not limited to word segmentation, part-of-speech tagging, syntactic analysis, semantic understanding, entity recognition, relation extraction, and text classification.
[0022] "Parsing" refers to the process of using natural language processing technology to perform structured processing and semantic analysis on text data, extracting key information, semantic relationships, contextual features, or logical structures.
[0023] "Perjury" refers to the tendency and possibility of false statements, misleading statements, or obvious inconsistencies with the facts in the text content. It is used to indicate whether the communication content may contain deception, fraud, or untrue information.
[0024] "Prompt information" refers to structured or unstructured information generated by the system based on the results of speech recognition, natural language processing, perjury detection, or sentiment analysis, used to guide subsequent judgments or actions. This includes, but is not limited to, internal prompts for perjury detection and suggestive text presented to the user.
[0025] "Perfunctory judgment result" refers to the evaluation result obtained by the system based on the parsing and perfunctory analysis of text data regarding whether the text has a risk of perfunctory verification. It may include risk level, confidence level, relevant labels, or judgment conclusion.
[0026] "Next Steps Recommendations" refers to the suggestions generated by the system based on the results of the perjury determination, which are used to guide users' subsequent actions. These suggestions include, but are not limited to, terminating communication, remaining vigilant, not transferring funds, not providing sensitive information, contacting official agencies for verification, or reporting to the police.
[0027] A “sentiment analysis engine” refers to a software module or algorithm component used to identify and evaluate the emotions, attitudes, or psychological states of a user’s voice or text, and can output results such as the user’s emotion category, emotion intensity, or emotion change trend.
[0028] "Emotional state" refers to a user's emotional or psychological state during communication, including but not limited to calm, tension, fear, anger, anxiety, hesitation, and other states that can be identified by an emotion analysis engine. Attached Figure Description
[0029] Figure 1 This is a conceptual diagram illustrating an example of the configuration of the data processing system according to the first embodiment.
[0030] Figure 2 This is a conceptual diagram illustrating an example of the main functions of the data processing apparatus and smart device according to the first embodiment.
[0031] Figure 3 This is a conceptual diagram illustrating an example of the configuration of the data processing system according to the second embodiment.
[0032] Figure 4 This is a conceptual diagram illustrating an example of the main functions of the data processing device and smart glasses according to the second embodiment.
[0033] Figure 5 This is a conceptual diagram illustrating an example of the configuration of the data processing system according to the third embodiment.
[0034] Figure 6 This is a conceptual diagram illustrating an example of the main functions of the data processing apparatus and head-mounted terminal according to the third embodiment.
[0035] Figure 7 This is a conceptual diagram illustrating an example of the configuration of the data processing system according to the fourth embodiment.
[0036] Figure 8 This is a conceptual diagram illustrating an example of the main functions of the data processing device and robot according to the fourth embodiment.
[0037] Figure 9 This represents an emotion map that maps multiple emotions.
[0038] Figure 10 This represents an emotion map that maps multiple emotions.
[0039] Figure 11 This is a sequence diagram illustrating the processing flow of the data processing system of the first embodiment.
[0040] Figure 12 This is a sequence diagram illustrating the processing flow of the data processing system in Application Example 1.
[0041] Figure 13 This is a sequence diagram illustrating the processing flow of the data processing system of the second embodiment.
[0042] Figure 14 This is a sequence diagram illustrating the processing flow of the data processing system in Application Example 2. Detailed Implementation
[0043] Hereinafter, an example of an implementation of the system according to the present disclosure will be described with reference to the accompanying drawings.
[0044] First, let me explain the terminology used in the following instructions.
[0045] In the following embodiments, the processor (hereinafter referred to as "processor") with reference numerals may be a single computing device or a combination of multiple computing devices. Furthermore, the processor may be a single computing device or a combination of multiple computing devices. Examples of computing devices include CPU (Central Processing Unit), GPU (Graphics Processing Unit), GPGPU (General-Purpose computing on Graphics Processing Units), APU (Accelerated Processing Unit), etc.
[0046] In the following embodiments, RAM (Random Access Memory), as indicated in the figures, is a memory that temporarily stores information and is used as working memory by the processor.
[0047] In the following embodiments, the memory, as indicated by the reference numerals, is one or more non-volatile storage devices that store various programs and parameters. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), disks (e.g., hard disks), or magnetic tapes.
[0048] In the following embodiments, the communication I / F (Interface) with reference numerals is an interface that includes a communication processor and an antenna, etc. The communication I / F is responsible for communication between multiple computers. As an example of a communication specification applicable to the communication I / F, wireless communication specifications such as 5G (5th Generation Mobile Communication System), Wi-Fi (wireless fidelity) (registered trademark), or Bluetooth (registered trademark) can be listed.
[0049] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B". That is, "A and / or B" means that it can be only A, only B, or a combination of A and B. Furthermore, in this specification, when "and / or" connects to express more than three items, the same interpretation as "A and / or B" applies.
[0050] First Implementation Method Figure 1 An example of the configuration of the data processing system 10 according to the first embodiment is shown.
[0051] like Figure 1 As shown, the data processing system 10 includes a data processing device 12 and an intelligent device 14. A server can be cited as an example of the data processing device 12.
[0052] The data processing apparatus 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" as understood in this disclosure. The computer 22 includes a processor 28, RAM 30, and memory 32. The processor 28, RAM 30, and memory 32 are connected to a bus 34. Furthermore, the database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0053] The smart device 14 includes a computer 36, a receiving device 38, an output device 40, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, RAM 48, and memory 50. The processor 46, RAM 48, and memory 50 are connected to a bus 52. In addition, the receiving device 38, output device 40, camera 42, and communication I / F 44 are also connected to the bus 52.
[0054] The receiving device 38 includes a touchscreen 38A and a microphone 38B, and receives user input. The touchscreen 38A receives user input via touch by detecting contact with an indicator (e.g., a pen or finger). The microphone 38B receives user input via sound by detecting the user's voice. The control unit 46A in the processor 46 sends data representing the user input received by the touchscreen 38A and microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data representing the user input.
[0055] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting data in a form perceptible to the user 20 (e.g., sound and / or text). The display 40A displays visual information such as text and images according to instructions from the processor 46. The speaker 40B outputs sound according to instructions from the processor 46. The camera 42 is a small digital camera equipped with an optical system such as a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.
[0056] Communication I / F44 is connected to network 54. Communication I / F44 and 26 are responsible for sending and receiving various information between processor 46 and processor 28 via network 54.
[0057] Figure 2The diagram shows an example of the main functions of the data processing device 12 and the smart device 14.
[0058] like Figure 2 As shown, in the data processing apparatus 12, specific processing is performed by the processor 28. A specific processing program 56 is stored in the memory 32. The specific processing program 56 is an example of a "program" as understood in this disclosure. The processor 28 reads the specific processing program 56 from the memory 32 and executes the read specific processing program 56 on the RAM 30. Specific processing is implemented by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.
[0059] The memory 32 stores a data generation model 58 and an emotion-specific model 59. The data generation model 58 and the emotion-specific model 59 are used by the specific processing unit 290. The specific processing unit 290 can use the emotion-specific model 59 to infer the user's emotions and perform specific processing based on the user's emotions. In the emotion inference function (emotion-specific function) using the emotion-specific model 59, various inferences and predictions related to the user's emotions are performed, including inferences and predictions of the user's emotions, but this is not limited to this example. Furthermore, emotion inference and prediction may also include, for example, emotion analysis (parsing).
[0060] In the smart device 14, the processor 46 performs the acceptance output processing. The memory 50 stores the acceptance output program 60. The acceptance output program 60 is used in conjunction with the data processing system 10 and the specific processing program 56. The processor 46 reads the acceptance output program 60 from the memory 50 and executes the read acceptance output program 60 on the RAM 48. The acceptance output processing is implemented by the processor 46 acting as the control unit 46A according to the acceptance output program 60 executed on the RAM 48. Furthermore, the smart device 14 has the same data generation model and emotion-specific model as the data generation model 58 and the emotion-specific model 59, and these models can also be used to perform the same processing as the specific processing unit 290. The acceptance output processing is implemented by the processor 46 acting as the control unit 46A according to the acceptance output program 60 executed on the RAM 48.
[0061] Alternatively, other devices besides the data processing device 12 may also have the data generation model 58. For example, a server device (e.g., a generation server) may have the data generation model 58. In this case, the data processing device 12 obtains the processing results (prediction results, etc.) using the data generation model 58 by communicating with the server device that has the data generation model 58. Furthermore, the data processing device 12 may be a server device or a user-held terminal device (e.g., a mobile phone, robot, home appliance, etc.). Next, an example of the processing of the data processing system 10 of the first embodiment will be described.
[0062] Example 1 The flow of a specific process in Example 1 will be described. Each part of the system described below is implemented by the data processing device 12 and the smart device 14. Furthermore, the data processing device 12 is referred to as the "server," and the smart device 14 is referred to as the "terminal."
[0063] In traditional communication response technology, when a communication terminal receives an incoming call, the user typically answers it manually. Although existing technologies have incorporated speech recognition, natural language processing, and simple rule engines to record call content or detect keywords, they still have significant shortcomings in the following aspects, thus failing to fully utilize the processing capabilities of computer systems in complex voice communication scenarios.
[0064] Firstly, in terms of call answering switching control, existing technologies are mostly fixed automatic voice navigation or voicemail modes, lacking a mechanism for fine-grained control between "human answering" and "automatic answering" based on real-time user operation and call status. This results in the computer system being unable to take over the session at the appropriate time, the control logic of the call access path being scattered between the terminal and the back-end system, the switching process being opaque and unprogrammable, and resource scheduling lacking unified control, thus limiting the improvement of the overall processing capacity of the communication system.
[0065] Secondly, regarding the recording and structured processing of call content, existing technologies mostly only record and archive voice messages, or convert voice messages into text through speech recognition and then perform keyword matching. They lack a phased conversation history management mechanism designed for subsequent intelligent evaluation and reasoning, based on conversation turns. The inability of computer systems to uniformly maintain "turn-based" dialogue context and multi-dimensional risk features at the conversation level means that subsequent risk assessments, credibility judgments, and behavioral suggestions can only rely on simple rules or single-turn text, making it difficult to utilize the contextual information of the entire conversation. This limits the system's expressive and reasoning capabilities at the algorithmic level.
[0066] Furthermore, in terms of automated computer assessment of the credibility of the other party's statements, existing technologies typically employ fixed rules or traditional models, matching only certain sensitive words or predefined patterns. This fails to accurately identify potential fraud, harassment, or other high-risk calls in complex and ever-changing contexts. Especially in long dialogue scenarios, the system lacks a mechanism to comprehensively model multi-turn conversations, semantic associations, and historical behavior. This results in a coarse internal computer representation of the credibility of the other party's statements, failing to generate structured evaluation information that can be reused by other modules, and thus hindering the support for refined subsequent action decisions.
[0067] Furthermore, in the process of providing system evaluation results and suggesting next steps to users, existing technologies mostly present these feedbacks as static text prompts or simple alarms. They lack a mechanism for automatically generated "prompt information" that can dynamically adjust based on call context and risk assessment results, failing to effectively convert complex backend calculations into user-friendly interactive output. In particular, existing technologies have failed to organically integrate generative artificial intelligence models into the communication response system's processing chain. They lack a standardized framework for generating and invoking "analyzed input information," "behavioral candidate information," and "summary information," resulting in fragmented model calls, inconsistent contexts, and difficulty in forming a stable and scalable computer technology solution at the system architecture level.
[0068] Furthermore, in terms of post-call analysis and knowledge accumulation, existing systems mostly just archive call recordings or text records, lacking a mechanism to link and store conversation history with credibility evaluation results, and to further utilize generative AI models to generate summary information for users or backend analysis. Consequently, computer systems cannot form structured "conversation objects" at the "single call" level, nor can they build learnable and reusable corpora and labeled data at the "multiple calls" level, limiting the system's ability to continuously optimize risk control strategies, model parameters, and interaction strategies.
[0069] In summary, existing technologies have failed to establish a unified, end-to-end computer processing framework for communication scenarios in areas such as call answering handover control, structured management of call content, credibility assessment calculation models, integration with generative artificial intelligence models, and post-event summarization and knowledge accumulation. This results in: (1) The call response path cannot be uniformly arranged and dynamically switched by the computer; (2) Session-level data structures and risk characteristics are not systematically modeled within the computer; (3) The capabilities of generative artificial intelligence models cannot be injected into the communication response process in a standardized manner; (4) Users cannot obtain action suggestions and post-event summaries based on in-depth analysis and dynamic adjustment.
[0070] Therefore, it is necessary to provide a new system and method to improve the communication response system from the perspectives of system architecture and data processing flow, enabling the computer to: uniformly control the "human / automatic response switching" in the end-to-end data flow, transform the call content into structured conversation history and risk characteristics, use generative artificial intelligence models to conduct in-depth assessment of the credibility of the other party's speech, and automatically generate various prompts and summary information for users and the backend, thereby substantially improving the computer's processing capabilities and technical effectiveness in voice communication risk control and interaction.
[0071] The specific processing performed by the specific processing unit 290 of the data processing apparatus 12 in Embodiment 1 is achieved by the following means.
[0072] In this invention, the server includes components for receiving incoming call signals and operation information from a communication terminal and controlling the switching of manual response processing performed in the communication terminal to automatic voice response processing performed in an information processing device; components for generating a greeting voice based on pre-stored fixed response information by performing speech synthesis processing to convert text information into speech information, and sending the greeting voice to the called party's communication device via the communication terminal, based on pre-stored fixed response information, at the start of automatic voice response processing; components for acquiring call voice data from the communication terminal, performing speech recognition processing to convert the call voice data into text information, and recording the text information as a conversation history in chronological order; and components for performing language parsing processing and risk feature extraction processing on the other party's spoken text information included in the conversation history. The system includes: a component for generating prompt information as input information for parsing credibility evaluation information related to the other party's speech; a component for inputting the parsing input information into a dialogue engine including a generative artificial intelligence model to obtain output information containing credibility evaluation results related to the other party's speech and action candidate information based on the evaluation results; a component for generating prompt information containing proposal content related to the next action of the user in communication based on the output information and sending the prompt information to the communication terminal for display or notification on the communication terminal; and a component for storing the conversation history and evaluation results in a recording area, generating prompt information for generating summary information based on the conversation history and evaluation results after the call ends, and inputting the prompt information into the generative artificial intelligence model. This allows for the construction of an end-to-end processing chain within the computer, encompassing call detection, switching between human and automated responses, speech-to-text conversion, sequential recording of conversation history, extraction of risk features, and the use of generative artificial intelligence models for credibility assessment, behavior candidate generation, and conversation summary generation. This enables unified management of data flow and control flow within the communication response system, allowing the server to automatically generate various prompts based on structured conversation history and multidimensional risk features. It provides users with dynamically adjusted behavioral suggestions and generates reusable summaries and labeled data after the call ends, thereby substantially improving the computer's processing efficiency, accuracy, and scalability in voice communication risk identification and interactive decision-making.
[0073] "Communication terminal" refers to user-side electronic devices used for voice or data communication with external communication networks, including but not limited to mobile communication devices, fixed-line telephone devices, tablet devices, computing devices, etc.
[0074] "Incoming call signal" refers to a control signal or call setup signal sent from the communication network side to the communication terminal to indicate the existence of a new call request, including call notification information carrying caller ID information.
[0075] "Operation information" refers to the information generated by the user's input operations on the communication terminal, including button clicks, touch input, voice commands, etc., which are used to indicate whether to switch to automatic response or perform other control actions.
[0076] "Information processing device" refers to an electronic computing device that has processing capabilities and is connected to a communication terminal via a network to perform automatic voice response processing, data analysis and processing, and generate prompt information, including server equipment, cloud computing platforms, or local gateway equipment.
[0077] "Manual response processing" refers to the call response process in which a person directly answers, converses with, and operates the incoming call, with voice acquisition and response content controlled by a person through a communication terminal.
[0078] "Automatic voice response processing" refers to the call response process controlled and automatically executed by an information processing device, in which the system automatically interacts with the other party through preset logic, speech synthesis and speech recognition without real-time human intervention.
[0079] "Fixed response information" refers to standardized text content that is pre-stored in an information processing device and is played to the other party at the start of automatic voice response processing. It includes greetings, identification information, and explanations of purpose.
[0080] "Speech synthesis processing" refers to the process of converting text information into playable audio signals, which is usually generated by a speech synthesis engine based on the input text to produce corresponding digital speech data.
[0081] "Greeting voice" refers to the voice information generated by speech synthesis and played to the other party at the beginning of automatic voice response processing. It is used to indicate that the system has entered automatic response mode and to inquire about the other party's needs.
[0082] "Call voice data" refers to the digital voice signals generated by the other party or both parties during a call and transmitted in the communication channel, used to represent the content of the call.
[0083] "Speech recognition processing" refers to the process of converting voice data into text information, including voice signal analysis, feature extraction, and speech-to-text conversion.
[0084] “Text information” refers to the content of a call, represented in character form, obtained through speech recognition processing or other input methods, and is used for subsequent language parsing and risk assessment.
[0085] "Conversation history" refers to the collection of textual information recorded in chronological order during a call, used to represent the time-series structured data of a complete or partial conversation.
[0086] "The other party's spoken text information" refers to the text information obtained by converting the voice of the other party in the conversation history, which is used to analyze the other party's intentions and credibility.
[0087] "Language parsing and processing" refers to the natural language processing operations performed on text information, such as word segmentation, syntactic analysis, and semantic recognition, in order to extract semantic structure, entities, and relationships.
[0088] "Risk Feature Extraction Processing" refers to the process of extracting feature data from the other party's spoken text information and conversation context to determine the risk or credibility of the call, including the extraction of keyword features, semantic pattern features, behavioral pattern features, etc.
[0089] "Prompt information" refers to structured or semi-structured information generated to invoke the model or provide input to the user or other modules, including input information for parsing, behavior suggestion information, and summary generation information.
[0090] "Input information for parsing" refers to the set of input data prepared for evaluating the credibility of another party's speech in a dialogue engine or generative artificial intelligence model, including conversation fragments, context summaries, feature descriptions, etc.
[0091] "Generative AI models" refer to AI models that are trained on large-scale data and can automatically generate text, speech or other content based on input, including but not limited to large-scale language models and multimodal generative models.
[0092] A "dialogue engine" refers to a combination of hardware and software modules used to manage multi-turn dialogue states, invoke generative artificial intelligence models, and generate dialogue responses or evaluation results.
[0093] "Credibility evaluation results" refer to evaluation data obtained based on the analysis of the other party's spoken text information and risk characteristics, used to indicate the authenticity, reliability, or risk level of the other party's speech, including scores, grades, and reasons.
[0094] "Behavioral candidate information" refers to information generated based on the credibility evaluation results, representing one or more subsequent action options available in the current call context, such as ending the call, continuing the call, or transferring to a human operator.
[0095] "Output information" refers to the set of results data output by the dialogue engine or generative artificial intelligence model, which includes at least credibility evaluation results and behavioral candidate information, and may also include explanatory text or supplementary explanations.
[0096] "Proposal content related to next steps" refers to user-oriented, output-based explanatory content that suggests what actions the user should take during the current call.
[0097] The “record area” refers to the data storage space used to store session history, evaluation results, and related metadata. It can be a database, file storage, or other data storage medium.
[0098] "Summary information" refers to simplified text or structured data generated based on conversation history and evaluation results, used to summarize the conversation topic, risk situation and key information.
[0099] "Overall risk level" refers to an aggregated indicator that is calculated based on the results of multiple rounds of credibility evaluation and the overall characteristics of the conversation, and is used to represent the overall risk level of the entire call.
[0100] "Emotional analysis processing" refers to the process of analyzing spoken or written information to identify the emotional state or attitude tendencies contained therein, including emotion category determination and emotion intensity estimation.
[0101] "Dynamic adjustment of behavioral candidate information and prompt content" refers to the process of automatically changing the type, priority, and presentation of behavioral candidate information, as well as the corresponding prompt text description, based on the credibility evaluation results and sentiment analysis results updated in real time within the information processing device.
[0102] In various embodiments of the present invention, the server, terminal, and user work collaboratively as the main participating entities. The system of the present invention can be deployed on computer equipment including a processor, storage device, and network interface. The server can be a cloud server, data center server, or local gateway device, and the terminal can be a smartphone, tablet device, or computing device with communication applications installed.
[0103] At the hardware level, servers can use multi-core general-purpose processors (such as server-grade central processing units), at least several GB of memory, and persistent storage devices (such as solid-state drives), and connect to terminals and communication networks via wired or wireless network interfaces. At the software level, servers can run Unix-like operating systems and run subsystems on them, including web server software, middleware frameworks, and database management software. For example, they can use application service frameworks based on HTTP or WebSocket, as well as relational database management systems or key-value caching systems.
[0104] At the hardware level, the terminal includes a processor, memory, audio acquisition and playback module, display device, and user input device; at the software level, it includes an operating system, communication applications, and an audio interface library. The terminal's communication applications can utilize the call framework interface provided by the operating system to obtain incoming call signals and call audio, and interact with the server through secure network protocols.
[0105] Users receive incoming calls and make action choices using the terminal. Users can switch the call from human answering to automatic voice answering by clicking the "AI Response" button or a similar button on the terminal interface, at which point the information processing device on the server side takes over the processing of the conversation content and risk assessment.
[0106] After receiving call information and user operation information from the terminal, the server internally constructs a session data structure for each call. This session data structure can be in the form of a record containing fields such as session identifier, user identifier, other party identifier, timestamp, session status, round list, and risk assessment result. The round list can be stored as an array or linked list structure, and each round record includes the caller type (user side or other party side), timestamp, voice data reference, transcribed text, and intermediate calculation features.
[0107] When an automated voice response system takes over the call, the server uses pre-stored fixed response text in its storage device to perform speech synthesis processing. The server can convert preset text such as "Hello, this is the automated response system, how can I help you?" into digitized voice data by calling the speech synthesis engine. The server then sends the synthesized voice to the terminal via a media transmission module. The terminal plays the greeting voice through its local audio playback interface, thus activating the automated response mode.
[0108] The terminal continuously collects the other party's spoken voice data in automatic response mode. After collecting the voice sample, the terminal encodes and compresses the voice, and sends it to the server in real-time or near real-time via the network. Upon receiving the voice data, the server passes it to the speech recognition module to perform speech-to-text conversion. The server can use speech recognition algorithms based on acoustic and language models to map audio feature vectors into text sequences. The server writes the recognized text as "the other party's spoken text information" into the corresponding round record in the conversation data structure.
[0109] After receiving the spoken text, the server performs language parsing and risk feature extraction. The server can use natural language processing libraries to segment the text, perform part-of-speech tagging, and dependency parsing to generate a structured semantic representation. Based on this, the server extracts various risk features, such as whether it contains words related to money transfer, whether it contains threatening statements, whether it contains fields related to personal privacy, and whether there is a semantic pattern of emergency pressure. The server can represent these features as vectors and combine them with the session context to form the "parsing input information."
[0110] After generating the input information for parsing, the server constructs a prompt statement, combining the text of the other party's speech, summaries of previous rounds of conversation, and the currently extracted feature description into natural language input, which is then used to invoke the generative artificial intelligence model. This prompt statement can take the following example form: "Based on the caller's statement below, please assess its credibility (0-100 points) and indicate whether it may be a scam or harassment:" The message reads: "This is [Loan Center Name]. You have been selected as a priority loan customer by the system. Simply provide your bank card number now for immediate loan disbursement." Please provide: 1) a credibility rating from 0 to 100; 2) your judgment on whether it is suspected to be fraud; 3) the reasons for your judgment; and 4) your suggested next steps for the user. The server inputs the aforementioned prompts and structured features into the generative AI model. In one implementation, the server can use a sequence-to-sequence generation model based on a multi-layer attention mechanism, or a transformer architecture model based on multi-head self-attention and a multi-layer feedforward network as the generative AI model. During the training phase, the model can undergo supervised learning using large-scale dialogue corpora and labeled risk classification data, optimize model parameters using cross-entropy loss or weighted loss functions, and update weights through backpropagation and gradient descent. The server can employ data augmentation techniques during training, such as paraphrasing, sentence transformation, and noise injection into the original text, to improve the model's robustness to diverse expressions.
[0111] During the inference phase, the server encodes the prompts into vector representations, which are then subjected to nonlinear transformations through several layers of attention and feedforward networks to generate outputs containing credibility scores, risk level labels, and explanatory text. At the model output layer, the server can use regression units to predict continuous credibility values from 0 to 100, while simultaneously using classification units to output several predefined risk categories (such as "high-risk fraud," "moderate risk," and "low risk"), and generating natural language explanations. The server then organizes these outputs into "credibility evaluation results" and "behavioral candidate information," and writes them into the session data structure.
[0112] In generating behavioral candidate information, the server employs a combination strategy that does not entirely rely on manual rules. The server can define, for example, based on threshold rules and multi-dimensional feature weighting: when the credibility is below a certain threshold and the risk category is "high-risk fraud," behavioral candidates may include "suggest ending the call immediately" and "suggest marking the number as high-risk"; when the credibility is in the medium range and the other party's tone is calm, it may include "suggest transferring to a human for further verification," etc. The server can combine the explanatory text generated in the previous step with these behavioral suggestions and use templates to generate user-facing prompts, such as: "System analysis results: The credibility of this call is 25 points, which is considered a high-risk call. The main risk factors include: requesting immediate bank card number and promising quick loan disbursement. We suggest you end the call immediately and do not provide any personal account information." The server sends the prompt as a notification to the terminal. Upon receiving the notification, the terminal displays a risk score, explanation, and recommended action buttons on the interface, such as "Hang Up," "Transfer to Human Agent," or "Continue Call but Remain Vigilant." The user selects one of these options based on their own judgment. If the user selects "Hang Up," the terminal terminates the call via the communication interface; if the user selects "Transfer to Human Agent," the terminal switches the call media channel from automated voice response to the user's microphone and speakerphone.
[0113] During the call, the server continuously updates the session history and corresponding evaluation results. After the call ends, the server constructs a summary and generates a prompt statement based on all rounds of information stored in the session data structure and the corresponding trustworthiness evaluation results. For example, the server can use the following prompt statement: "Based on the complete phone conversation below, please generate a summary of no more than 100 words, explaining whether this call poses a risk of fraud and providing the main evidence:" Dialogue transcript: ... (List key statements by round). The server inputs the prompt into a generative artificial intelligence model, which generates a summary of the call, including the topic, risk assessment, and key evidence. The server stores this summary along with the conversation history and evaluation results in a record area for later querying, statistical analysis, and model retraining.
[0114] The system of this invention does not merely automate the simple replacement of human responses; rather, it introduces a unified data structure oriented towards the session level, a vectorized representation of risk features based on multi-turn dialogues, and a decision-making mechanism driven by a generative artificial intelligence model within the server. By providing a structured internal representation of the session history, the server improves data management efficiency, enabling each round of speech and its corresponding risk assessment to be quickly indexed and reused, thereby significantly reducing the computational burden of repetitive analysis in multi-call scenarios.
[0115] In terms of technical effectiveness, by unifying the judgment results of speech recognition, natural language processing, feature extraction, and deep generative models into a single session data structure, the server can reduce intermediate data conversion between different modules, thereby reducing memory copying and serialization overhead and improving overall processing speed. When calling generative artificial intelligence models, the server uses high-information-density prompts and feature vectors as input, avoiding the direct transmission of lengthy raw data, which reduces network communication load and speeds up model inference response time.
[0116] In terms of accuracy, the server employs multi-dimensional labels (including credibility scores, risk categories, and explanatory text) to jointly train the model during the training phase. This enables the model to simultaneously generate numerical scores and explanatory conclusions at output, thereby improving the consistency and interpretability of the judgment results. The server also incorporates sentiment analysis (a module that can be added in other implementations) to analyze the tone and emotion categories of the other party's speech. By combining sentiment features with semantic risk features, risk identification achieves higher accuracy in complex scenarios compared to traditional methods based solely on keyword matching.
[0117] Through the aforementioned architecture and processing flow, the server enables the computer system to make comprehensive judgments using methods that transcend human thought. For example, during the risk feature extraction stage, the server can introduce a rule base based on statistical patterns and a retrieval mechanism based on vector similarity. This allows it to make multi-dimensional decisions, combining contextual structure, word order, and numerical features, rather than relying solely on single keywords. This processing method, based on vector space and deep network internal representations, is significantly different from human judgments based solely on experience and explicit rules, representing a substantial extension of computer technology capabilities.
[0118] In different implementations, the server can adopt different generative artificial intelligence model structures. For example, in one implementation, the server can use a one-way transformer model for decision-making; in another implementation, the server can use a combination of a bidirectional encoder network and a decoder network; in yet another implementation, the server can introduce a multi-task learning framework, enabling the same model to simultaneously output credibility scores, category labels, and summary text, thereby sharing the underlying representation and reducing redundant computations during inference.
[0119] The server can be modularly divided into: a session management module, a speech processing module, a language parsing and feature extraction module, a generative artificial intelligence invocation module, a behavior suggestion generation module, and a storage and retrieval module. These modules interact through clearly defined data structure interfaces, avoiding the repetitive parsing of large amounts of unstructured text and unnecessary encoding / decoding processes common in traditional systems. This improves the data and control paths of the computer system in voice communication applications at the architectural level.
[0120] By employing the aforementioned structure and processing methods, the system of this invention can automatically monitor and evaluate real-world telephone calls in real-world communication scenarios, providing users with structured risk alerts and suggestions for next steps in real time. After the call concludes, it provides high-quality session and tagging data for further analysis and model improvement. This closed-loop system, encompassing hardware acquisition, software processing, model inference, and user interaction, not only enhances user security in terms of business operations but also achieves an overall improvement in processing speed, judgment accuracy, and system scalability at the computer technology level.
[0121] use Figure 11 The processing procedure is explained.
[0122] Step 1: Terminal detects incoming calls and reports them. The terminal receives incoming call signals from the communication network through the call interface provided by the operating system. The input includes metadata such as the caller ID, caller ID, and timestamp. The terminal parses and encapsulates this raw signaling data, converting it into a structured data record containing fields such as call_id, caller_id, timestamp, and call_state. This structured call notification is then sent to the server via a network protocol, outputting a "call notification" message with a unique session identifier.
[0123] Step 2: The server generates session records and initializes the state. The server takes the "incoming call notification" message sent by the terminal as input and parses out the session identifier, caller information, and time information. The server creates a new session record in the session management module, initializing fields including session status (set to "waiting for user selection"), session round list (empty list), and risk assessment result (empty or default value). The server writes this session record to a database or in-memory storage, outputting a persistent session data entity in the storage system, and returning a "session created" confirmation message to the terminal.
[0124] Step 3: Users can choose whether to enable automatic voice response. When a user sees an "AI Answer" option on the incoming call screen of the terminal, the system uses the current call screen state as input information for judgment. The user taps the "AI Answer" button, and the terminal captures this tap event as an input event, associating it with the current session identifier. The terminal encodes this event, generating a control message containing fields such as call_id, user_id, and action="start_ai", and sends it to the server over the network, outputting an "AI Answer Request" message.
[0125] Step 4: The server controls the switch from human response to automated voice response. The server takes the "AI Response Request" message as input and parses the session identifier and user action type from it. The server updates the session status from "Waiting for User Selection" to "Automatic Response" in the session record and internally initiates the automatic voice response control process. The server stores the updated session status and outputs the session record with the changed status, along with an internal control command to notify the media control module to begin allowing the automatic response device to take over the call.
[0126] Step 5: The server generates a greeting message and sends it to the terminal. The server takes the session state and pre-stored fixed response text as input, and reads the greeting text corresponding to the activation of automatic response (e.g., "Hello, this is the automatic response system, how can I help you?") from storage. The server feeds this greeting text into the speech synthesis engine, performs text-to-speech conversion, and generates digital audio data containing sampling rate, encoding format, and a sequence of speech samples. The server packages the synthesized audio data into a media data packet and sends it to the terminal through a media transmission channel, outputting a directly playable greeting voice stream. The terminal uses this voice stream as input, calls the local audio playback interface to drive the speaker to play it, thus outputting the greeting voice to the other party through the physical audio channel.
[0127] Step 6: The terminal collects and uploads the other party's voice recordings. The terminal takes the currently active call media channel as input and acquires the voice signal emitted by the other party from the system audio interface. The terminal performs framing and encoding processing on the continuous voice signal, converting the raw PCM samples into a compressed format (such as generating compressed frames using a specific encoding standard), and adding a timestamp and session identifier to each audio segment. The terminal uploads these encoded voice data blocks sequentially to the server via the network, outputting a series of "call voice data packets" with session context information.
[0128] Step 7: The server performs speech recognition and generates text information. The server takes the voice data packets uploaded by the terminal as input and sends the audio frame stream to the speech recognition module. The server performs feature extraction operations on the audio data (such as calculating acoustic features like Mel-frequency cepstral coefficients), then inputs the feature sequence into the speech recognition model to calculate the probability distribution of characters or sub-words at each time step. Finally, it uses decoding algorithms (such as Viterbi decoding and beam search) to obtain the most probable text sequence. The server treats the recognized text as "the other party's spoken text information" and writes it, along with the corresponding timestamp and speaker identifier, into the conversation round list, outputting an updated conversation history data structure.
[0129] Step 8: The server performs language parsing and risk feature extraction on the text information. The server takes the latest conversation from the other party and existing conversation history as input, and calls the language parsing module to perform word segmentation, part-of-speech tagging, and syntactic dependency analysis. From the parsing results, the server extracts semantic units such as entities, amounts, times, and action verbs, and performs pattern matching and statistical analysis operations in conjunction with contextual information. The server constructs a feature vector for this round of conversation, including risk features such as keyword counts, sensitive word markers, semantic category distribution, and contextual similarity. The server encapsulates these features along with the original text into "parsing input information," and outputs a structured risk feature data object.
[0130] Step 9: The server constructs a prompt and invokes a generative artificial intelligence model. The server takes the parsed input information and conversation context as input and constructs a natural language prompt to guide the generative AI model in making credibility judgments and risk assessments. The server integrates the caller's speech content, a brief summary of the current conversation, and extracted key risk points into a text prompt, such as: "Based on the caller's speech content below, judge its credibility (0-100 points) and explain whether it may be fraud or harassment. Speech content: ...". The server inputs this prompt into the generative AI model, performs forward propagation calculations on the model's encoder, generates intermediate semantic representations, and obtains credibility scores, risk categories, and explanatory text through the decoder or output header, outputting "credibility evaluation result data" containing multiple fields.
[0131] Step 10: Server-generated behavioral candidate information and user-oriented risk warnings The server takes the credibility evaluation results and conversation history as input, performs rule calculations or threshold judgments based on the score and risk category, and selects a set of suitable subsequent action candidates, such as "End the call," "Transfer to human verification," and "Keep the call going but be cautious." The server combines these action candidates with explanatory text provided by the model and generates a user-facing prompt statement through template filling, such as: "System analysis result: The credibility of this call is 30 points, suspected of being high risk. It is recommended that you end the call immediately and not provide any personal account information." The server encodes this prompt statement into a message and sends it to the terminal, outputting a displayable risk warning message.
[0132] Step 11: The terminal displays risk warnings and receives user decisions. The terminal takes the risk warning message sent by the server as input, and parses the risk score, text description, and behavioral candidate options from it. The terminal displays the risk information in text and graphics on the user interface, and generates corresponding buttons or selection controls (such as "Hang Up," "Transfer to Human Agent," and "Continue"). After reading the information, the user makes an action selection based on the prompts, using the user's click or touch event as input. The terminal encodes the user's selected action into a control message (containing call_id and user_action fields), sends it back to the server, and outputs a "user decision message."
[0133] Step 12: The server performs call control and updates the session state. The server takes user decision messages as input and extracts the session identifier and user-selected actions. The server updates the session state in the session record, for example, setting it to "user hung up," "transferring to human agent," or "continue auto-answering." If the server needs to end the call, it sends an end-of-call command to the communication control module; if it needs to transfer to a human agent, it sends media channel switching and agent assignment commands. After completing the state update and necessary call control operations, the server writes the final state to the storage system and outputs a session record marked as having completed control actions and terminated.
[0134] Step 13: The server generates and stores a session summary after the call ends. The server takes the complete conversation history and all credibility evaluation results as input, constructs a prompt for summary generation, and compresses the multi-turn dialogue content into contextual description text that the model can process. For example: "Please generate a brief summary based on the following complete phone conversation record, and explain whether there is a fraud risk and the main evidence. Conversation record: ...". The server inputs this prompt into the generative artificial intelligence model, performs forward generation operations, and obtains a summary text containing the call topic, risk conclusion, and key evidence. The server stores this summary text, along with the original conversation history, risk characteristics, and final state, in the record area, and outputs a conversation archive record with complete tags and summary information.
[0135] Application Example 1 The process flow corresponding to the specific processing in Use Case 1 will be described below. The various parts of the system described below are implemented by the data processing device 12 and the intelligent device 14. Furthermore, the data processing device 12 is referred to as the "server" and the intelligent device 14 is referred to as the "terminal".
[0136] In the modern communication environment, suspicious phone calls, fraudulent messages, and other unusual communications are constantly increasing. Current technology typically involves humans directly answering or reading the communications and then relying on their experience to determine if there is a risk of fraud or perjury. This approach presents the following technical problems: First, the communication processing flow heavily relies on manual intervention, making it impossible to establish a unified and automated judgment logic on the server side, thus limiting the overall system throughput and concurrent processing capabilities. Second, existing automatic response technologies are mostly based on fixed scripts or rule bases, unable to dynamically generate highly relevant response content and protection suggestions based on call context and risk assessment results, resulting in limited automation effectiveness. Third, when utilizing technologies such as speech recognition and natural language processing, existing systems typically only perform simple transcription or keyword filtering, lacking the ability to link "pseudo-evidence assessment results" with "prompt statements from generative AI models," causing the generative model to fail to fully utilize structured risk information, resulting in unstable and uncontrollable model output. Fourth, existing systems lack a comprehensive analysis and feedback adjustment mechanism for the emotional state of users and communication partners during multi-turn dialogues, failing to balance security with interactive experience and tone control, easily leading to false alarms, excessive user alarm, or insufficient intervention.
[0137] Therefore, there is an urgent need for a computer implementation solution that improves upon the system architecture and computational process: This solution should be able to uniformly receive suspicious communication data from terminals on the server side, automatically perform multimodal analysis of voice and text, generate structured perjury assessment results, and use these as the core control signal to dynamically construct prompt statements and model parameters input to the generative artificial intelligence model, thereby improving the accuracy, controllability, and personalization of automatic response content and action suggestions, and fundamentally improving the computer's technical capabilities and resource utilization efficiency in processing suspicious communications.
[0138] The specific processing performed by the specific processing unit 290 of the data processing apparatus 12 in Application Example 1 is achieved by the following means.
[0139] In this invention, the server includes a unit for switching received information transmission from human response to AI-automated response; a unit for initiating AI-automated response based on a start instruction from a terminal and storing the dialogue content with the communication object as recorded information; a unit for acquiring voice information and converting it into text information through voice recognition processing, and performing natural language processing analysis on the text information and the recorded information to generate evaluation information for assessing the falsity of evidence; a unit for generating prompt statements input to a generative AI model based on the evaluation information and the dialogue content, and using the generative AI model to generate response information for suspicious information transmission; and a unit for generating prompt statements input to a generative AI model based on the evaluation information and the dialogue content, and using the generative AI model to generate response information for suspicious information transmission. The system includes a unit for generating proposal information to instruct the user on the next action to take, and a unit for sending the response information to a terminal to automatically send it to the communication object and for sending the proposal information to the terminal for the user to view. It may also include a unit for generating control information based on the falsity assessment information to change the strength or warning level of the response content for suspicious information transmission and dynamically adjusting the prompt statements and / or model parameters input to the generative AI model according to the control information; and a unit for analyzing the user's emotional state and the communication object's emotional state through sentiment analysis processing and adjusting the content and / or tone of the prompt statements input to the generative AI model based on the analysis results and the falsity assessment information, thereby regulating the expression of the proposal information and the response information. This allows for the formation of an end-to-end automated processing chain within the computer server, encompassing multi-source communication data acquisition, speech-to-text conversion, natural language falsification assessment, dynamic construction of prompt statements, and invocation of generative artificial intelligence models. This enables the server to adaptively control the input and output of the generative artificial intelligence model based on structured assessment and emotional information, thereby improving the accuracy and response speed of suspicious communication identification, reducing the need for human intervention, enhancing the system's processing capabilities in high-concurrency environments, and improving user experience while ensuring security. Ultimately, this results in a substantial improvement in the overall performance and resource utilization efficiency of computer technology.
[0140] A "system" refers to a collection of technical solutions consisting of one or more information processing devices and the programs running on them, used for receiving, analyzing, generating, and sending communication data.
[0141] "Terminal" refers to an information processing device operated by a user and communicating with a server, including but not limited to mobile terminals, fixed terminals, or other electronic devices with communication and display functions.
[0142] A "server" is a computing device equipped with a processor and storage devices and that communicates with terminals via a network. It is used to perform program processing such as speech recognition, natural language processing, generative artificial intelligence model invocation, and control logic.
[0143] "Information transmission" refers to the process of exchanging data, such as voice, text, images, or other forms, between a terminal and an external communication object through a communication network.
[0144] "Human response" refers to the behavior of human users directly responding to information transmission, including answering phone calls, replying to messages, or engaging in interactive dialogues.
[0145] "AI-powered auto-response" refers to response information that is automatically generated and sent by a computer program without relying on real-time human intervention. The generation process is based on artificial intelligence algorithms and related models.
[0146] "Activation instruction" refers to the control information sent by the terminal to the server to request that communication processing be switched from human response to automatic response by artificial intelligence.
[0147] "Conversation content" refers to the collection of voice or text information sent back and forth between the user, the communication partner, and the AI-generated automatic response.
[0148] "Recorded information" refers to the data formed after the server or terminal stores the content of the conversation, including the content of the speech, time information, role identification, etc.
[0149] “Voice information” refers to spoken content expressed in audio form, including sound data generated by the communication partner or user during a call.
[0150] "Speech recognition processing" refers to the data processing process that uses computer algorithms to convert speech information into corresponding text information.
[0151] “Textual information” refers to readable data consisting of a sequence of characters, used to represent spoken content obtained through speech recognition or direct reception.
[0152] "Natural Language Processing Analysis" refers to the process of using natural language processing algorithms to perform semantic parsing, intent recognition, feature extraction, and classification of textual information.
[0153] "Perfunctoriness" refers to the attribute or degree to which the information contained in the statements of a communication subject has a tendency to be false, misleading, or deceptive.
[0154] "Assessment information" refers to structured data generated based on natural language processing analysis results, used to represent the risk or credibility of falsehoods, including scores, ratings, and explanations.
[0155] "Generative AI models" refer to AI models that can automatically generate text or other content based on input prompts and contextual information, including language generation models based on deep learning.
[0156] "Prompt statements" refer to the set of information used as input to generative artificial intelligence models. They are usually given in the form of natural language text and are used to instruct the model on what type and style of output content to generate.
[0157] "Response information" refers to the automatic response content generated by a generative artificial intelligence model based on prompts and dialogue content, used to reply to the communication object.
[0158] "Proposal information" refers to suggestions generated by generative artificial intelligence models based on evaluation information and dialogue content, used to prompt users to take the next action.
[0159] "Control information" refers to a set of control parameters generated based on false evidence assessment information, used to adjust the intensity of response information or warning level, and to change model inputs or parameters.
[0160] "Warning level" refers to the level of risk warning given for suspicious information transmission, which is used to indicate the severity of the system's judgment on potential fraud or perjury.
[0161] "Model parameters" refer to adjustable settings used to control the behavior of generative artificial intelligence models, including but not limited to temperature parameters, sampling strategies, length limits, and other configuration values that affect the output content.
[0162] "Sentiment analysis processing" refers to the technical process of analyzing textual information or speech features using algorithms to identify the emotional categories and intensity contained therein.
[0163] "User's emotional state" refers to information obtained through sentiment analysis that represents the type and intensity of a user's emotions during communication.
[0164] "The emotional state of the communication partner" refers to information obtained through sentiment analysis that indicates the type and intensity of emotions of the party communicating with the user during the conversation.
[0165] "Expression style" refers to the form in which the response or proposal information is presented in terms of language content, wording style, tone strength, and level of politeness.
[0166] In the embodiments of the present invention, the server, client, and user each assume different functional roles. Through the collaboration of hardware resources and software modules, the system can automatically take over suspicious communications, evaluate the falsity of evidence, and generate responses and output behavioral suggestions based on generative artificial intelligence models and prompt statements.
[0167] In the following description, "server" refers to a data processing device deployed on the network side, "end" refers to a communication device operated by a user, and "user" refers to a natural person who uses an end to communicate.
[0168] I. Overall System Composition A server includes at least one processor, storage device, and network interface. The server's processor can be a multi-core general-purpose processor, such as a central processing unit based on x86 or ARM architecture. The server's storage device includes main memory (random access memory) and secondary storage (hard disk drive or solid-state drive). The server runs an operating system (e.g., a Unix-like operating system) and application servers (e.g., a Python-based backend framework or a JavaScript-based backend framework) on the above hardware. The server communicates bidirectionally with end-users via the network interface.
[0169] The terminal includes a processor, memory, communication module, audio acquisition and playback module, and display device. The terminal can be a smartphone or a tablet. The terminal's communication module establishes a data connection with the server via a cellular network, wireless LAN, or wired network. The terminal runs a local calling application and a security application. The security application sends initiation instructions to the server, uploads conversation content, and presents the server's response and proposal information.
[0170] Users can perform input actions through the terminal operation interface, issue an instruction to start the AI automatic response, read the proposal information provided by the server, and control the call according to the proposal information, such as hanging up, blocking, or further confirmation.
[0171] II. Program Modules and Data Structures The server sets up multiple functional modules in the application layer. These modules can be physically executed by the same processor, but they have different logical responsibilities.
[0172] 1. Session Management Module The server is equipped with a session management module to generate a session identifier for each suspicious incoming communication and maintain a structured record associated with the session identifier in the storage device. The server creates a session record data structure in the storage device, and the session record includes at least the following: a session ID field, a counterparty identifier field (e.g., a normalized representation of a communication number or user identifier), a start time field, an end time field, a status field (e.g., "AI takeover in progress" or "Completed"), a risk score field, and multiple attribute fields related to the perjury assessment results.
[0173] By storing the aforementioned session records as row records in a relational data storage system, the server enables subsequent modules to retrieve the required records within a constant time based on the session ID, thereby reducing random access latency.
[0174] 2. Dialogue storage module The server is equipped with a dialogue storage module to store dialogue content uploaded by each client as structured message records. The server defines message records in the storage device, including message ID, session ID, role field (e.g., user, communication object, AI auto-responder), type field (voice, text, transcribed text), timestamp field, and content field.
[0175] After receiving the dialogue data uploaded by each client, the server converts each message into the aforementioned message record format and stores them in groups by session ID. This specific data structure allows the server to quickly reconstruct the dialogue context in chronological order, providing an efficient data access path for subsequent natural language processing analysis and generative artificial intelligence model invocation.
[0176] 3. Speech Recognition Module The server is equipped with a speech recognition module to convert speech information into text information. The server can call speech recognition software libraries or external speech recognition services. These software implementations are typically based on deep neural networks (e.g., end-to-end recognition structures based on convolutional neural networks and recurrent neural networks, or speech-to-text models based on attention mechanisms). In this embodiment of the invention, the server segments the speech stream into short audio segments, extracts acoustic features (e.g., Mel-frequency cepstral coefficients, spectral power features), inputs these features into a trained acoustic model to obtain phoneme or sub-word probability distributions, and then combines these with a language model to decode and output a character sequence.
[0177] The server aligns the transcription results of different segments with time and combines them with the timestamps in the message records. It then fills the corresponding message record's content field with the transcribed text and sets the type field to transcribed text. This achieves both the normalization of speech content to a unified text representation and provides a unified input format for subsequent natural language processing.
[0178] 4. Natural Language Processing Module The server is equipped with a natural language processing module to analyze textual and recorded information and generate false evidence assessment information. Within this module, the server can use various algorithm combinations, including feature extraction based on word embeddings, text classification based on convolutional neural networks, sequence encoding based on bidirectional long short-term memory networks, or sequence encoding based on self-attention structures.
[0179] The server can pre-train a text classification model. The model's input consists of the text sequence of the other party's speech and a summary feature of the conversation context. Internally, the model uses word embedding layers to map words to a vector space, models long-distance dependencies through a multi-layer self-attention network, and finally uses fully connected layers and a soft maximum function to output a probability distribution related to falsehood. The server generates a falsehood score based on this probability distribution and writes the score, risk level, and key triggering features into an evaluation information structure.
[0180] The server can also execute a rule engine within the natural language processing module, such as detection rules based on regular expression matching or pattern matching. When the risk score output by the model approaches the boundary threshold, the server fine-tunes the score by combining the results of the rule engine, thereby controlling the false positive rate while ensuring recall.
[0181] 5. Sentiment Analysis Module The server is equipped with a sentiment analysis module to analyze the emotional state of both the user and the communication partner. Within this module, the server can employ a sentiment classification model. The model structure can be a bidirectional recurrent network with an attention mechanism or a sentiment recognition network based on a transformer structure. The input is the spoken text and its context, and the output is the emotion category (e.g., calm, tense, angry, fear) and its intensity score.
[0182] The server stores sentiment analysis results and false evidence assessment information in the session log. By attaching sentiment vectors to the structured data, the downstream module can comprehensively consider the emotional state of both parties when generating response and proposal information, thereby implementing tone control and mitigation strategies.
[0183] 6. Generative Artificial Intelligence Module Interface The server is configured with a generative artificial intelligence module interface for interacting with the generative artificial intelligence model. In this implementation, the generative artificial intelligence model can be a language generation network based on a transformer architecture, which includes multi-layer self-attention encoders and decoders, or a multi-layer self-attention decoding structure in dialogue generation scenarios. The model acquires language modeling and task-related capabilities through pre-training and fine-tuning.
[0184] When the server invokes the generative AI model, it provides not only the original dialogue text but also falsification assessment information and sentiment analysis results output by the natural language processing module. At the interface layer, the server transforms this structured information into natural language or specific tagged forms and embeds it into the prompts, enabling the generative AI model to internally recognize this information and adjust its output accordingly.
[0185] The server maintains a set of prompt statement templates in the storage device. Different templates serve different purposes, such as generating the first automatic response, generating detailed warning suggestions, or generating a milder explanation. Before invoking a template, the server selects an appropriate template from the template set based on the session state and risk level, and fills the placeholders in the template with specific session information.
[0186] For example, the server can generate the following prompt statement for suspicious phone call opening responses: "You are a customer service robot from a security service center. A user is receiving a suspicious call. Please generate an opening statement to inform the caller that: 1. You represent the security service center; 2. This call is being recorded; 3. The caller needs to explain their purpose. Be polite but firm in your language." The server can also generate prompts for action suggestions: "The user is receiving a potentially scam call. Based on the following dialogue and the scam risk score provided by the system, please generate a clear and easy-to-understand suggestion for the user, including: 1. Should the user hang up immediately? 2. Should the user report the incident to the police or contact official customer service? 3. What dangerous actions should the user avoid (such as revealing verification codes, passwords, etc.)? Dialogue content: '...', Risk score: 0.9." The server uses the aforementioned prompts to convert internal numerical scores and detection results into natural language descriptions, enabling the generative AI model to follow specific constraints and intentions when generating text, thereby avoiding uncontrollable output caused by the model's complete freedom of generation.
[0187] 7. Output control and intensity adjustment module The server is equipped with an output control and intensity adjustment module, which generates control information based on the falsity assessment information and uses this control information to dynamically adjust the content of the prompt statements and the parameters of the generative artificial intelligence model. Within this module, the server can define multiple warning levels and map them to different threshold ranges based on the risk score.
[0188] For example, when the risk score is higher than a preset high-risk threshold, the server can lower the temperature parameter of the generative AI model to make the output more conservative and certain, and add instructions such as "strongly recommend ending the call immediately" to the prompt. When the risk score is in the middle range, the server can appropriately increase the temperature parameter, allowing the generative AI model to provide multiple suggestion options, and requiring the model to provide conditional suggestions in the prompt.
[0189] In this way, the server applies "soft constraints" to the generative artificial intelligence model using specific numerical parameters and template content, achieving fine-grained control from structured evaluation information to natural language output. This control does not rely on manual decision-making, but is formed through algorithmic computation, thereby enabling automated risk-adaptive response generation within the computer.
[0190] 8. End-to-End Presentation and Device Control Module The terminal includes a presentation module for displaying response and proposal information to the user on a display device. It also includes a device control module for invoking system-provided call control interfaces, such as hanging up, muting, or blocking communication targets, based on user selection.
[0191] The presentation module displays the proposal information returned by the server in the form of pop-ups, banners, or dialog boxes, and provides interactive controls on the interface, allowing users to end the call or block the number with a single click. In this way, the system not only makes judgments at the logical level, but also directly affects the physical call process by controlling the communication module, realizing a closed loop from data analysis to device control.
[0192] III. Technical Effects and Improvements in Computer Technology The processing achieved by the server through the above modular structure is not just a simple automation of the manual judgment process, but rather forms an optimized data processing link within the computer.
[0193] First, the server uses specific data structures for session records and message records to achieve sequential access and efficient retrieval by session ID index, reducing the computing resources required for random disk access and multi-table joins, thus improving performance under high concurrency conditions.
[0194] Secondly, in the natural language processing module, the server uses the false evidence assessment as an explicit numerical output and tightly couples this output with the generative AI model invocation, rather than using the generative AI model in isolation for general text generation. In this way, when invoking the generative AI model, the server actually implements a "feature injection and parameter tuning" control scheme, making the generated output more aligned with risk control objectives, reducing irrelevant or overly arbitrary content, and lowering the risk of misjudgment and misleading statements.
[0195] Furthermore, the server simultaneously analyzes the emotional states of both the user and the communication partner in the sentiment analysis module, and uses these sentiment vectors in the output control and intensity adjustment module to adjust the prompt statements and model parameters. For example, when a user is detected to be highly stressed, the server will design the prompt statements to be more soothing rather than overly frightening, and control the output length and word intensity in the model parameters. This ensures accurate risk warnings while reducing unnecessary psychological burden; this effect is achieved through joint modeling of internal sentiment features and control parameters.
[0196] Furthermore, the server employs templates and structured input methods when generating prompts, rather than generating input entirely freely on each call. This unconventional approach significantly reduces output bias by protecting the stability of the prompt structure and ensuring that key constraints (such as "do not ask the user for a verification code" and "must remind the user that the call is being recorded") are not ignored by the generative AI model. This structured management of prompts improves the controllability and reliability of generative AI model deployment at the system architecture level.
[0197] When training natural language processing and sentiment analysis models, servers can employ supervised learning methods, using cross-entropy as the loss function and updating weights through gradient descent or its variants (such as adaptive moment estimation). Furthermore, data augmentation strategies (such as synonym substitution, random deletion, and order perturbation) can be combined to expand the diversity of training samples, thereby improving the model's generalization ability to different fraudulent statements and sentiment expressions. These training strategies and weight update processes directly improve the model's accuracy and robustness during the inference phase, thus enhancing the server's overall discriminative ability.
[0198] Through the aforementioned structural design and algorithm arrangement, the server in this embodiment of the invention achieves automatic identification and response control of suspicious communications. This process involves specific data structures, specific model architectures, and parameter adjustment strategies, improving upon the limitations of traditional systems that rely solely on keyword filtering or fixed script responses. It substantially enhances the information processing speed, risk identification accuracy, and resource utilization efficiency of computer systems in communication security scenarios, representing an improvement to computer technology itself.
[0199] use Figure 12 The processing procedure is explained.
[0200] Step 1: The terminal receives the communication and generates an AI takeover request.
[0201] The terminal acts as the subject: It receives voice calls or text messages from external sources via its communication module, acquiring caller ID, other party ID, timestamp, and other communication metadata. The terminal displays the incoming call screen on its interface and presents an "AI Takeover" button. When the user deems the communication suspicious, they click this button on the terminal interface. The terminal encapsulates the current call ID, other party ID, timestamp, and "AI Takeover Flag" into structured data. The input is the system call event and the user's click action; the output is the AI takeover request data sent to the server. Based on the input call metadata and user instructions, the terminal combines multiple fields into a single request record through data encapsulation operations and sends this record to the server via the network stack.
[0202] Step 2: The server creates a session and records the AI takeover status.
[0203] The server acts as the subject: It receives AI takeover requests from terminals via a network interface, parses the structured data in the request, and extracts the call identifier, the other party identifier, the timestamp, and the AI takeover flag. The server generates or looks up a session ID for the call in the session management module and creates or updates a session record in the data storage, setting the status field to "AI_AUTO_REPLY_ACTIVE". The input is the AI takeover request data sent by the terminal, and the output is a session record containing the session ID and status information. The server performs parsing and database insert / update operations based on the input fields, transforming the original request into an indexable, persistent session entry.
[0204] Step 3: The server generates a prompt message for the first automatic response.
[0205] The server acts as the subject: After recognizing a session entering AI auto-response mode, the server selects a template for "suspicious communication opening response" from the prompt management module. The server inserts contextual information such as the other party's identifier, communication type, and time into placeholders in the template to form a specific prompt statement, for example: "You are a customer service robot from a security service center. The user is receiving a suspicious call. Please generate an opening statement to indicate to the other party on the phone: 1. You represent the security service center; 2. This call is being recorded; 3. The other party needs to explain their purpose. The language should be polite but firm." The input is session metadata and template definition; the output is a complete initial response prompt statement. The server transforms structured context data into a natural language prompt statement through string replacement and concatenation operations.
[0206] Step 4: The server invokes a generative artificial intelligence model to generate the first response.
[0207] The server acts as the subject: In the generative AI model interface module, the server constructs a request message using the prompt obtained in step 3 as part of the model input, setting call parameters such as model name, temperature parameters, and maximum generation length. The server encodes this content into a request message and sends it to the generative AI model's server endpoint via the network interface. Internally, the generative AI model uses a multi-layer self-attention network based on a transformer structure to encode the prompt and generates response text word by word through a decoder. The server receives the response returned by the model, parses it to obtain the automatic response text, for example: "This is the security service center. This call is undergoing security monitoring. What is the reason for your request?" The input is the prompt and model parameters, and the output is the text content of the first response. Based on the input, the server performs network request and response parsing operations, converting the natural language prompt into natural language output suitable as a call response.
[0208] Step 5: The server sends the first response message to the terminal.
[0209] The server acts as the subject: In its output module, the server packages the generated response text and session ID into response data, setting the response type to "INITIAL_REPLY". The server sends the response to the corresponding session's terminal via an application layer protocol (such as HTTP or WebSocket). The input is the response text and session identifier output by the generative AI model; the output is the response message sent to the terminal. The server performs serialization and packetization operations on the input data to make it suitable for transmission over the network and for correct parsing by the terminal.
[0210] Step 6: The terminal plays or sends the first response to the communication object.
[0211] The terminal acts as the subject: After receiving a response message from the server, the terminal parses the response text and response type from the message. For voice calls, the terminal calls its local speech synthesis engine to convert the response text into a speech signal and plays it to the communication target through the call audio channel; for text communication, the terminal encapsulates the response text into an SMS or chat message and sends it to the other party through the communication module. The input is the response text returned by the server, and the output is the actual voice or text signal received by the communication target. The terminal converts high-level text data into audio or text frames adapted to the underlying communication channel through text parsing and speech synthesis or message encapsulation operations.
[0212] Step 7: The terminal collects and uploads the content of the ongoing dialogue.
[0213] The terminal acts as the subject: During AI takeover, the terminal continuously monitors calls and message transmissions. For voice calls, the terminal divides the call audio into time segments and generates audio data blocks according to a local buffering strategy; for text communication, the terminal intercepts each newly sent or received text message. The terminal adds a role identifier, timestamp, and session ID to each voice segment or text message, forming a message record, and sends it to the server periodically or in real time via the network interface. The input is the call audio stream or text message stream, and the output is the message record with metadata uploaded to the server. Based on the input, the terminal performs segmentation, annotation, and encapsulation operations, converting continuous raw signals into structured message units.
[0214] Step 8: The server converts voice messages into text information and stores them uniformly.
[0215] The server acts as the subject: After receiving message records uploaded by the terminal, the server distinguishes the type field. For voice messages, it invokes the speech recognition module. The server extracts acoustic features from the audio data, inputs these features into a pre-trained acoustic and language model, and obtains the text sequence through a decoding algorithm. The recognized text is then written back to the content field of the corresponding message record, and the type is marked as transcribed text. For text messages, the server directly stores the content in the message record. The input is message records with audio or text content, and the output is a collection of message records uniformly represented in text and persisted in the database. Through feature extraction and sequence decoding operations, the server transforms voice data that cannot be directly analyzed into text data that can be used by the natural language processing module.
[0216] Step 9: The server performs natural language processing on the other party's statements and generates false evidence assessment information.
[0217] The server acts as the subject: In its natural language processing module, the server reads the opposing party's spoken text and context sentences within a specific time window, converting them into vector representations. For example, it maps each word to a dense vector space using word embedding or sub-word embedding, and then feeds this vector representation into a self-attention-based encoding network. The server extracts sentence-level feature vectors from the network's output, inputs them into a classification layer, and calculates the probability that the speech is fraudulent or perjury. Based on the probability values and preset thresholds, the server calculates a perjury score and risk level, and identifies key sentences containing sensitive patterns. The input is a sequence of relevant text messages from the opposing party, and the output is perjury assessment information including a risk score, risk level, and a list of key sentences. The server maps the original text into measurable risk indicator data through vectorization, attention weight calculation, and classification operations.
[0218] Step 10: The server performs sentiment analysis and updates the session state.
[0219] The server acts as the subject: In the sentiment analysis module, the server selects the user's and communication partner's speech texts, encodes them into word vector sequences, and inputs them into the sentiment recognition network, such as a model containing bidirectional recurrent units and attention layers. The server obtains the sentiment category and sentiment intensity scores from the model output and associates these values with the session ID, writing them into the session record. The input is the role-distinguished speech text, and the output is sentiment state data containing sentiment category and intensity. The server converts subjective sentiment information into structured sentiment vectors through sentiment feature extraction and classification operations for use by subsequent control logic.
[0220] Step 11: The server constructs prompts for the generated responses based on the assessment information and emotional state.
[0221] The server acts as the subject: It reads falsified assessment information (e.g., risk score of 0.87, high risk level) and the emotional states of both parties (e.g., user is tense, other party is assertive) from the prompt management module, and selects an applicable response template, such as a template emphasizing security monitoring. The server inserts information such as "risk score," "key suspicious sentences," and "user emotions" into the template to generate prompt statements with control intent, such as: "The following dialogue shows that the caller has repeatedly requested bank card information. The system assesses the fraud risk as 0.87. You need to respond as the security service center, using a firm but not provocative tone, clearly stating that the user will not provide any verification codes or account information over the phone." The inputs are falsified assessment information, emotional state data, and template content; the output is a customized prompt statement for a generative artificial intelligence model. The server synthesizes numerical values and symbolic information into natural language control instructions through template selection and placeholder replacement operations.
[0222] Step 12: The server invokes a generative artificial intelligence model to generate subsequent response information.
[0223] The server acts as the subject: Taking the prompt statement generated in step 11 and the current dialogue context as input, the server sends a request to the generative AI model and adjusts the model parameters based on the falsity assessment results, such as lowering the temperature parameter and shortening the maximum output length in high-risk situations. Internally, the generative AI model generates new response text based on the constraints in the prompt statement through a multi-layered self-attention decoding process, for example: "This call is for security verification only. We cannot process any transfers over the phone; please use official channels." The server receives and parses the generated result, storing the response text along with the session ID. The input is the prompt statement, dialogue context, and control parameters; the output is subsequent response information containing policy constraints. The server transforms the evaluation information into specific dialogue output through parameter control and model invocation operations.
[0224] Step 13: The server constructs prompts for action recommendations and generates proposal information.
[0225] The server acts as the subject: When generating action suggestions, the server reads false evidence assessment information, user emotional state, and other party behavioral characteristics, and selects a "behavioral suggestion" template. The server constructs a prompt statement, such as: "The user is receiving a potentially fraudulent phone call. Based on the following dialogue and the fraud risk score given by the system, please generate a clear and easy-to-understand suggestion for the user, including: 1. Should the user hang up immediately? 2. Should the user report the incident to the police or contact official customer service? 3. What dangerous actions should the user avoid (such as revealing verification codes, passwords, etc.)? Dialogue content: '...', Risk score: 0.9." The server inputs this prompt statement along with the context into a generative AI model to obtain suggestion text, such as: "This call is highly suspected of being fraudulent. We suggest you hang up immediately, do not provide any verification codes or bank card information, and verify through the official customer service hotline." The inputs are assessment information, emotional data, and the behavioral suggestion template; the output is a proposal information in natural language form. The server transforms structured risk data into easily understandable action suggestions through template filling and model generation operations.
[0226] Step 14: The server dynamically adjusts the prompts and model parameters based on the risk level and user sentiment.
[0227] The server acts as the subject: In the output control module, the server generates control information based on the false evidence rating range and the user's emotional intensity. This control information includes warning level, language intensity coefficient, and tone softening coefficient. The server uses this control information to modify the prompts in subsequent calls, such as adding "strongly recommend ending the call immediately" in high-risk scenarios and adding reassuring statements when the user is highly stressed. Simultaneously, the server sets the temperature, lexical diversity parameters, and output length limits for the generative AI model based on the control information. The inputs are risk scores and emotion vectors; the outputs are updated prompts and a set of model call parameters. The server transforms abstract risk and emotional signals into concrete control over model behavior through range mapping and parameter adjustment operations.
[0228] Step 15: The server sends response and proposal information to the terminal.
[0229] The server acts as the subject: In its output module, the server encapsulates the newly generated response and proposal information along with the corresponding session ID into a response message. This message explicitly distinguishes between "response content oriented towards the communication object" and "suggestion content oriented towards the user." The server sends the message to the terminal via the network interface. The input is the response text and suggestion text; the output is composite response data containing both types of content. The server performs serialization and channel selection operations to ensure that the two types of information can be used for call output and interface display on the terminal side, respectively.
[0230] Step 16: The terminal sends a follow-up response to the communication object and displays the proposal information to the user.
[0231] The terminal acts as the subject: After receiving the composite response from the server, the terminal parses out the response text and the proposal text. The terminal repeats the speech synthesis or text transmission process described in step 6 for the response text, transmitting the response to the communication target. Simultaneously, the terminal presents the proposal information in the display interface as a pop-up window or prompt bar, such as: "This call is suspected of being a scam; we suggest you hang up immediately and contact official customer service for confirmation." The input is the response and proposal text returned by the server, and the output is the actual response received by the communication target and the visual suggestion content in the user interface. Through data parsing, audio or message output, and graphical interface rendering calculations, the terminal maps abstract text into real call control and visual prompts.
[0232] Step 17: Users select operations based on the proposal information, and the terminal executes the device control.
[0233] User as the subject: After reading the proposal information on the terminal interface, the user chooses a specific action based on their own judgment, such as clicking "Hang up immediately" or "Hang up and block". Upon receiving the user's action, the terminal calls the system call control interface to end the current call and updates the local contacts or block list if necessary. The terminal then sends the user's final action as feedback data to the server. The input is the user's click action and the proposal content provided by the server; the output is the change in call status (e.g., call ended) and the feedback record. The terminal performs event responses and system API call calculations based on the input, achieving a physical-level effect from information prompts to communication module control.
[0234] Step 18: The server archives session results and uses them for subsequent model and rule optimization.
[0235] The server acts as the subject: After receiving user feedback from the terminal, the server updates the session record with the end time, final state, and user actions. The server periodically reads archived sessions along with corresponding false evidence assessment results, sentiment analysis results, and generated output effects from the background statistics module, performing aggregate statistics and performance evaluations, such as calculating the adoption rate and false positive rate of suggestions within different risk ranges. Based on these statistical results, the server can adjust the thresholds of the natural language processing model and sentiment analysis model, update prompt templates, or optimize the parameters for generative AI model calls. The input is the final session state and historical dataset; the output is the updated model configuration and rule parameters. Through aggregate statistics, performance analysis, and parameter update calculations, the server forms a closed-loop optimization process, continuously improving false evidence recognition accuracy, response rationality, and computational resource utilization efficiency across multiple runs.
[0236] Alternatively, an emotion engine for inferring user emotions can be combined. That is, the specific processing unit 290 can also use the emotion-specific model 59 to infer user emotions and perform specific processing using user emotions.
[0237] Example 2 The flow of a specific process in Example 2 will be described. Each part of the system described below is implemented by the data processing device 12 and the smart device 14. The data processing device 12 will be referred to as the "server," and the smart device 14 as the "terminal."
[0238] Existing automatic response systems, when processing voice communication, typically only convert speech to text and provide responses under fixed rules. Their technical solutions have the following shortcomings: First, in the computer's internal data processing, speech recognition results are often used only as simple text input, lacking structured extraction of deep semantic elements such as time expressions, event expressions, and negation expressions. This makes it difficult for subsequent computer-based judgment logic to accurately distinguish specific contradictions between the spoken content and external objective facts, thus limiting the computer system's ability to accurately judge perjury. Second, traditional systems rely heavily on preset rules or single algorithms for matching and judgment, lacking a unified prompt statement construction mechanism for generative artificial intelligence models. This prevents the automatic generation of high-quality input data containing textual semantic information and objective factual information within the computer, resulting in insufficient utilization of the generative artificial intelligence model's computing power on the server side. Third, existing technologies typically only provide keyword-level risk warnings for the text, without forming a complete data processing pipeline on the server to integrate communication switching, speech recognition, semantic parsing, external fact acquisition, prompt statement generation, and the inference results of the generative artificial intelligence model. Therefore, information loss and redundant calls are prone to occur in the computational chain, affecting system performance and response speed. Fourth, existing systems mostly use static templates when generating user behavior suggestions, lacking the technical means to dynamically adjust prompts and responses based on the results of perfunctory judgments and the user's emotional state. This makes it impossible to optimize the input and output of generative artificial intelligence models at the server level, thus limiting the quality of interaction and the effectiveness of human-computer collaboration.
[0239] In summary, how to improve the processing flow from voice data to text data and then to semantic and objective factual information within the server, how to effectively drive generative artificial intelligence models to make perjury judgments and generate behavioral suggestions through structured data processing and prompt statement generation mechanisms, and how to improve the performance of computer systems in terms of accuracy, real-time performance, and resource utilization throughout the entire processing chain have become urgent technical problems to be solved in this field.
[0240] The specific processing performed by the specific processing unit 290 of the data processing apparatus 12 in Embodiment 2 is achieved by the following means.
[0241] In this invention, the server includes: a device for switching received communication from manual response to automatic voice response; a device for acquiring acoustic signals from a terminal device or communication device and converting the acoustic signals into character information using audio processing technology and speech recognition technology; a device for storing the character information, along with user identification information and time information associated with the character information, in an information management device; a device for applying natural language processing technology to the character information to extract semantic information, including time expressions, event expressions, and negation expressions; a device for acquiring objective fact information from an external information providing device based on the semantic information and determining the consistency between the objective fact information and the character information to generate a prompt statement related to perfunctoriness evaluation information; and a device for generating input data for using a generative artificial intelligence model when generating the evaluation information. The device is configured to construct a prompt statement containing the character information, the semantic information, and the objective fact information, and to provide the prompt statement as input data to the generative artificial intelligence model, and to generate a perfunctoriness determination result based on the evaluation result output from the generative artificial intelligence model and to send the determination result as output data to an output device. This allows for the formation of an end-to-end data processing chain within the server, encompassing communication switching, speech recognition, semantic extraction, external fact acquisition, prompt statement generation, and generative artificial intelligence model reasoning. This enables the computer to efficiently and accurately perform perjury judgments and generate behavioral suggestions based on structured semantic information and objective factual information, thereby improving the overall technical effectiveness of the automatic response system in terms of computational accuracy, processing performance, and interaction quality.
[0242] A "system" refers to a collection of comprehensive information processing devices consisting of one or more server devices, terminal devices, and communication networks, used to perform processes such as communication switching, speech recognition, semantic parsing, fact comparison, prompt generation, and generative artificial intelligence model reasoning.
[0243] A "server" refers to an electronic computing device that executes program instructions to perform functions such as voice data processing, text data processing, semantic information extraction, acquisition of objective factual information, generation of prompt statements, and invocation of generative artificial intelligence models for reasoning operations.
[0244] "Terminal device" refers to a user-side electronic device used to collect user voice, send data to the server, and display the server's output results to the user, including but not limited to smartphones, tablet computers, personal computers, or other devices with communication and display functions.
[0245] "Communication" refers to the process of exchanging voice, text or other forms of data between users and systems or between different devices via wired or wireless networks.
[0246] "Human response" refers to a response method in which a human operator listens to, understands and responds to received communications.
[0247] "Automatic voice response" refers to a response method that does not rely on human operators, is controlled by computer programs, and automatically responds to received communications through preset rules or speech synthesis technology.
[0248] "Acoustic signal" refers to the analog signal or its digitized audio data that represents air vibration and is collected by input devices such as microphones, used to characterize the user's voice or environmental sounds.
[0249] "Audio processing technology" refers to signal processing techniques that perform operations such as sampling, quantization, encoding, noise reduction, format conversion, framing, and feature extraction on acoustic signals.
[0250] "Speech recognition technology" refers to the computational processing technology that converts acoustic signals representing speech into corresponding text information. It recognizes speech content through acoustic models, language models, and other methods.
[0251] “Character information” refers to digital text data obtained from acoustic signals through speech recognition technology to represent semantic content, including sentences, words, and symbols.
[0252] "Information management device" refers to storage and management components used to store, retrieve, update, and manage character information, user identification information, time information, etc., including database servers or other data storage systems.
[0253] "User identification information" refers to identifying data used to distinguish different users, including but not limited to user ID numbers, account IDs, or other code values that can uniquely or quasi-uniquely identify users.
[0254] "Time information" refers to time stamps associated with communication or speech, including timestamp data that records the start time, end time, or text generation time of the communication.
[0255] Natural Language Processing (NLP) technology refers to computer processing techniques used to process character information through word segmentation, part-of-speech tagging, syntactic analysis, entity recognition, and semantic analysis in order to extract syntactic structure and semantic information.
[0256] "Semantic information" refers to structured or semi-structured data extracted from character information through natural language processing technology, which represents meanings such as time, event, negation, subject, and object.
[0257] "Time expression" refers to linguistic components in character information used to indicate points in time or time intervals, such as words or phrases indicating dates, times, past or future.
[0258] "Event expression" refers to linguistic components in character information used to represent a certain behavior, state, or situation, such as verbs or phrases describing a specific event.
[0259] "Negative expression" refers to linguistic components in character information used to indicate non-existence, non-occurrence, or the opposite state, such as negative expressions like "no," "not," and "not yet."
[0260] "External information providing device" refers to a device or service that provides objective factual data related to time, place, environment or business outside the system, including but not limited to external databases, network service interfaces or third-party information platforms.
[0261] "Objective factual information" refers to structured or unstructured data provided by external information providers that describes actual events, such as weather records, transaction records, and location records.
[0262] "Consistency determination" refers to the process of comparing the content represented by character information with objective factual information based on predetermined rules or algorithms to determine whether the two are consistent or contradictory.
[0263] "Perfumery" refers to the nature or degree to which a statement is inconsistent or contradictory with objective facts, thus possessing the potential to be false, misleading, or untrue.
[0264] "Evaluation information" refers to the result data that describes and quantifies the perjury or related risks, including probability level, explanation of reasons, and related markings.
[0265] "Prompt statements" refer to input text that is constructed to drive generative artificial intelligence models and contains character information, semantic information, objective factual information, and task descriptions, used to guide the model to perform specific analysis or generation tasks.
[0266] "Generative artificial intelligence models" refer to artificial intelligence models trained through machine learning that can automatically generate text or other data outputs based on input prompts, including language models based on deep learning.
[0267] "Input data" refers to the various data sets provided to generative artificial intelligence models, including at least prompts and task-related parameters or contextual information.
[0268] "Evaluation results" refer to the textual or structured information output by a generative artificial intelligence model based on input data, used to reflect perfunctoriness, risk, or other analytical conclusions.
[0269] "Perfumeability determination result" refers to the final judgment data made on whether the content of the speech may be perfunctory based on the evaluation results and consistency judgment, including the judgment conclusion and corresponding explanatory information.
[0270] "Output device" refers to a device used to present server output data to a user, including a display, speaker, terminal screen or other human-computer interaction interface.
[0271] "Suggested information" refers to the suggested content or guiding text provided to users for their next course of action based on the results of the perjury determination and objective factual information.
[0272] "Sentiment analysis technology" refers to computational processing techniques that process character information to identify and classify the emotions or attitudes reflected within it (such as positive, negative, neutral, tense, angry, etc.).
[0273] "Emotional state" refers to the category and intensity of a user's emotions when speaking, inferred from character information through sentiment analysis technology.
[0274] "Behavioral suggestions" refer to specific or general action guidelines generated by the system for users based on the results of the perjury determination and their emotional state, regarding whether to continue interacting, whether to verify, and whether to take preventive measures.
[0275] The embodiments of this invention will focus on servers, terminals, and users, and will describe the specific usage of the system structure, data structure, algorithm flow, and generative artificial intelligence model, so that those skilled in the art can implement this invention on a general-purpose computer platform.
[0276] I. System Overall Structure A server comprises a central processing unit, main memory, solid-state storage, a network interface, and an operating system. The server loads an application program into the main memory. This application program consists of modules such as a communication management module, an audio processing module, a speech recognition interface module, a text storage module, a natural language processing module, an external information acquisition module, a prompt generation module, a generative artificial intelligence invocation module, a sentiment analysis module, and a result output module. The server can run on a general-purpose server operating system, such as a Unix-based server operating system.
[0277] The terminal includes a processor, memory, microphone, speaker, touchscreen display, and wireless communication module. The terminal runs client applications on its operating system, which include a voice acquisition unit, a network communication unit, and a results display unit. The terminal can be a smartphone, tablet, or personal computer, etc.
[0278] Users operate through the terminal to trigger voice recording and view results.
[0279] II. Data Structures and Storage Structures The server stores key data in a structured format within a database management system, such as a relational database. For each message, the server generates a record in the database, which may include the following fields: 1. Character information field: Used to store the text string obtained through speech recognition.
[0280] 2. User identification information field: Used to store the user's identification number.
[0281] 3. Session Identifier Field: Used to distinguish different communication sessions.
[0282] 4. Time information field: Used to store the timestamp when the record was generated.
[0283] 5. Semantic Information Field: Used to store structured information extracted by the natural language processing module, such as time expression, event expression, negation expression, etc., which can be in key-value pair or JSON tree structure.
[0284] 6. Objective Fact Information Field: Used to store a summary of factual data obtained from external information providing devices and verification results, including fact source identifier, time, location and key values (such as precipitation).
[0285] 7. Perfumeability Judgment Result Field: Used to store the system's final judgment value on perfumeability, such as multi-level confidence score.
[0286] 8. Sentiment Status Field: Used to store sentiment analysis results, such as sentiment category and intensity score.
[0287] 9. Behavior suggestion field: Used to store a summary or identifier of the suggestion text generated by the generative artificial intelligence model.
[0288] This structured storage makes subsequent retrieval, statistics, and reuse more efficient, reduces redundant analysis and calculations, and thus improves overall computing efficiency.
[0289] III. Specific Implementation of Servers in Speech and Text Processing When the server switches received communication from manual response to automatic voice response in the communication management module, it maintains a current call status table in memory, assigns a session identifier to each call, and sets the response mode field in the status table. When the server detects preset conditions, such as a busy agent or a call type that meets the automatic processing rules, it updates the response mode field to automatic voice response and controls the communication device to switch to automatic playback and recording mode through the communication interface.
[0290] The server performs sampling rate resampling, channel conversion, and noise reduction on the acoustic signal from the terminal in the audio processing module. The server can use short-time Fourier transform to extract spectral features to improve the input quality of the speech recognition interface module. After completing the format conversion, the server sends the audio buffer to the speech recognition interface module in blocks.
[0291] The server calls an external speech recognition service, such as a cloud-based speech recognition application programming interface (API), within its speech recognition interface module. The server packages and sends pre-processed audio segments along with information such as language encoding, sampling rate, and acoustic model parameters. Internally, the speech recognition service obtains character information through acoustic and language modeling and returns it as a confidence score. The server then concatenates and chronologically arranges the recognition results from multiple segments to generate a complete character information string, which is then written to the character information field of the database.
[0292] IV. Server Implementation in Natural Language Parsing and Semantic Extraction The server loads a pre-trained language analysis model into its natural language processing module. Through word segmentation, part-of-speech tagging, and dependency parsing, the server decomposes character information into word sequences and syntactic tree structures. On this structure, the server runs a rule-matching and statistical model to identify semantic units such as time expressions (e.g., "yesterday," "last Wednesday"), event expressions (e.g., "rain," "transfer," "login"), and negation expressions (e.g., "no," "did not happen").
[0293] The server constructs a semantic tag with a uniform format for each semantic unit and combines these tags into a semantic information object. The semantic information object can be stored in a tree structure, with nodes including type (time / event / negation), start and end positions, and normalized values (e.g., specific date, event category). The server stores this object in the semantic information field to support subsequent consistency checks and prompt statement generation.
[0294] In this way, the server processes data that was originally just unstructured text into an internal representation with defined fields and computable features, so that it can use indexes and logical operations to achieve efficient matching and reasoning in subsequent processing stages.
[0295] V. Server Implementation in External Fact Acquisition and Consistency Determination In the external information acquisition module, the server constructs an external query request based on the time expression and possible location information in the semantic information. For example, when the character information is "It didn't rain yesterday," the server extracts the time expression "yesterday" from the semantic information object, and then maps "yesterday" to a specific date through the time conversion module. At the same time, the server obtains location information from user profiles or communication metadata. The server uses the date and location as parameters to call the service interface provided by the external information providing device to obtain weather record data for that time and location.
[0296] In its consistency determination logic, the server compares the event expression in the character information with the objective fact information based on a predefined determination rule table. For example, when the event expression is "rain" and there is a negative expression, the server determines that the fact expressed by the user's statement is "it did not rain on this date." If the precipitation field in the objective fact information is greater than a threshold and the weather type field indicates rainfall, the server marks the consistency determination result as inconsistent and generates internal information describing the reason, such as "the user's statement contradicts the weather record." This internal information will be written into the evaluation information for subsequent prompt generation and as a reference for the computation of the generative artificial intelligence model.
[0297] This structured rule-based decision-making on the server side can significantly reduce irrelevant information at the data level, improve the input quality of subsequent models, and thus enhance overall decision-making accuracy and computational efficiency.
[0298] VI. Server Implementation in Generative Artificial Intelligence Models In the prompt statement generation module, the server combines character information, semantic information, and objective factual information into prompt statements according to a predefined template. The server constructs prompt statements using a structure that includes role settings, task descriptions, input content, and output format requirements, enabling the generative artificial intelligence model to perform reasoning under a unified standard.
[0299] For example, the server can generate the following prompt: You are an analytical assistant who helps determine the likelihood of a statement being fabricated.
[0300] User comment: "It didn't rain yesterday." Semantic parsing results: Time = Yesterday (mapped to date 2025-05-01); Event = Rain; Negation = Yes.
[0301] Objective weather data: Location = a city; Date = 2025-05-01; Observation results: Multiple rainfalls occurred that day, with a cumulative precipitation of 15 mm.
[0302] Please answer in Simplified Chinese: 1. The perjury probability level of this statement (low / medium / high); 2. Explain your reasoning (no more than 100 words); 3. Provide the user with a suggested next step (no more than 50 words).
[0303] The server invokes a generative AI model based on a deep neural network within the generative AI invocation module. This model employs a multi-layered self-attention sequence transformation network architecture, including an embedding layer, a self-attention layer, a feedforward network layer, and an output layer. During the invocation, the server encodes the aforementioned prompt statement into a discrete token sequence. The model calculates the semantic relationships between the tokens using an attention mechanism, then obtains the output token probability distribution through multi-layer transformations, and finally decodes it into natural language text output.
[0304] During model training, the server can utilize a large-scale corpus to perform supervised and self-supervised learning. It measures the difference between the model output and the target text using a cross-entropy loss function and updates the weight parameters using a gradient descent-based optimization algorithm. In the fine-tuning phase, the server can use a small, specialized dataset containing factual data and pseudo-factual annotations to reduce errors on specific tasks through continued training. The server can also employ data augmentation techniques, such as synonym substitution, sentence transformation, or the addition of noisy samples, to improve the model's robustness to noisy data.
[0305] By combining structured prompts with deep model reasoning, the server can perform high-dimensional feature modeling and nonlinear judgment of complex semantics and factual relationships without increasing the workload of rule configuration. This surpasses simple rule matching in computation and significantly improves the accuracy and generalization ability of perjury detection.
[0306] VII. Server Implementation in Sentiment Analysis and Behavioral Suggestion Generation In the sentiment analysis module, the server performs sentiment classification on character information. The server can use convolutional neural networks, recurrent neural networks, or self-attention-based text classification models to extract sentiment feature vectors from the text. The server inputs the feature vectors into a fully connected layer and uses a soft maximum function to output the probabilities of different emotion categories, such as tension, anger, fear, and calm. The server then writes the emotion category and its probability into the sentiment state field.
[0307] In the prompt generation module, the server incorporates emotional state as an additional condition when generating behavior-oriented prompts. For example, the server can construct the following prompt: You are a risk control and user care assistant.
[0308] User comment: "It didn't rain yesterday." The results of the perjury analysis show that there are obvious contradictions with the weather records.
[0309] Sentiment analysis results: The user's current emotion is tension, and the intensity of the emotion is high.
[0310] Please provide the following in Simplified Chinese, taking into account user sentiment: 1. A brief evaluation of the authenticity of the statement (no more than 50 words); 2. A reassuring statement to the user (no more than 50 words); 3. Suggestions for the user's next steps (no more than 50 words), in a calm tone that does not cause panic.
[0311] The server, in its generative AI invocation module, generates output text that considers both factual contradictions and user emotions based on the prompt. The server stores the behavioral suggestions portion of the output text in a behavioral suggestion field and then sends it to the terminal via the result output module.
[0312] By simultaneously integrating semantic information, objective factual information, and emotional state in the server to drive the generative artificial intelligence model, the system technically achieves differentiated processing of the same text under different emotional contexts, which helps reduce the risk of misunderstanding and improve the quality of human-computer interaction.
[0313] VIII. Interaction Patterns Between Terminals and Users The terminal utilizes the operating system's microphone interface in its voice acquisition unit to perform preliminary compression and encoding of the acoustic signals locally, reducing the amount of data uploaded. In its network communication unit, the terminal employs a secure transmission protocol to send the audio data and necessary metadata to the server. In its results display unit, the terminal receives the server's results regarding the perjury assessment, sentiment analysis, and behavioral suggestions, displaying these to the user through a graphical interface. For example, the terminal can use color coding to indicate the perjury level, icons to indicate the risk level, and textual suggestions to enable the user to quickly understand the system's output.
[0314] Users can choose to view detailed analysis on the terminal interface, or they can directly perform recommended operations based on the prompts, such as calling the official customer service number or terminating the current suspicious communication.
[0315] IX. Technical Effects and Improvements in Computer Technology The server internally implements a complete data pipeline from audio acquisition, speech recognition, semantic extraction, fact comparison, prompt generation to generative AI model inference. This reduces redundant parsing and data format conversion between different functional modules, and caches intermediate results in a structured form, significantly reducing overall computational complexity and communication load. The server pre-filters key conflict points using semantic extraction and consistency checks, allowing the generative AI model to focus on high-value inputs, reducing unnecessary computations during model inference, and contributing to shorter response times.
[0316] The server employs multi-source information fusion and a multi-level decision-making mechanism during the decision-making process, combining rule-based decision-making with deep model-based decision-making. Rule-based decision-making quickly eliminates cases with obvious consistency or inconsistency, while the deep model handles complex or boundary cases. This layered architecture technically improves average computational efficiency and overall accuracy.
[0317] The specific machine learning techniques employed by the server, such as deep neural networks, attention mechanisms, feature extraction, and loss function optimization, belong to data processing methods within the computer. By designing the structure of the prompt statements, feature representation methods, and training methods, this invention improves the model's performance under defined tasks, thereby achieving performance improvements in text-to-fact consistency judgment and emotion-sensitive behavior suggestion generation at the computer technology level, rather than simply automating the manual judgment process.
[0318] 10. Other Implementation Forms and Variations In other implementations, the server can use different natural language processing libraries or self-developed models to replace the existing language analysis module, as long as the module can extract time expressions, event expressions, and negation expressions from character information, similar technical effects can be achieved.
[0319] In terms of acquiring external information, servers can be extended beyond weather data to include transaction record services, location record services, or access control log services, thereby expanding the applicable scenarios for perjury detection.
[0320] The server can employ language models of varying sizes and architectures for generative artificial intelligence model selection, as long as they can generate evaluation and suggestion texts based on prompts, thus achieving the objectives of this invention. The server can also deploy lightweight language models locally to reduce network call latency.
[0321] The terminal can be any computing device with a microphone and display in different implementation forms, such as in-vehicle terminals, smart home devices, or self-service terminals. The core technical feature of the system lies in the processing logic and data flow structure on the server side, which is not limited by the specific form of the terminal.
[0322] Through the above-described embodiments, this invention enables the server to achieve intelligent judgment of falsification and generation of behavioral suggestions on voice communication content through a collaborative approach of specific data structures, algorithmic processes, and generative artificial intelligence models within the computer, thereby technically improving the processing performance and reliability of the automatic response system.
[0323] use Figure 13 The processing procedure is explained.
[0324] Step 1: The user initiates a voice input operation on the terminal. The input is the user's natural language speech. After the user clicks the record button, the terminal calls the operating system's microphone interface to collect acoustic signals, and converts the analog audio into digital audio data through the audio driver, buffering it in memory at a fixed sampling rate and bit width (e.g., 16kHz, 16bit). The output is a segment of encoded (e.g., PCM or compressed format) digital audio data.
[0325] Step 2: The terminal sends the acquired digital audio data to the server. The input consists of the digital audio data obtained in step 1, along with the accompanying user identifier, session identifier, and start and end time information. The terminal encapsulates the audio data into a network packet in its network communication module and sends it to the designated interface of the server via a secure transmission protocol. The output is an audio data packet indicating successful upload to the server.
[0326] Step 3: The server receives audio data uploaded by the terminal. The input consists of audio data packets and related metadata from the terminal. The server parses the packets in its communication management module, extracts the audio byte stream and user-related information, and creates temporary files or memory buffers to store the audio content. The output is a standardized audio data object accessible on the server side.
[0327] Step 4: The server preprocesses the received audio data. The input is a standardized audio data object. The server performs resampling, channel format conversion, and noise reduction in its audio processing module, removing background noise and standardizing audio parameters using digital signal processing algorithms (such as short-time Fourier transform). The server writes the processed audio to a new buffer. The output is preprocessed audio data that meets the requirements of the speech recognition interface.
[0328] Step 5: The server converts the preprocessed audio data into character information. The input is the preprocessed audio data. The server constructs a recognition request in its speech recognition interface module, sending the audio data and language parameters to the speech recognition service and receiving the returned recognition results. The server concatenates and rearranges the returned segments to generate the corresponding text string. The output is character information representing the user's speech content and the recognition confidence level.
[0329] Step 6: The server stores character information and metadata in the information management device. Inputs include character information, user identification information, session identifier, time information, and identification confidence level. The server constructs an insert statement in the database interface, writing the above data into the corresponding fields of the relational database and recording the generated record identifier. Output is a speech record stored in the database and its unique identifier.
[0330] Step 7: The server performs natural language parsing on the character information and extracts semantic information. The input is character information read from the database. In its natural language processing module, the server performs word segmentation, part-of-speech tagging, and syntactic analysis. It identifies time expressions, event expressions, and negation expressions using rules and statistical models, and generates standardized attributes (such as specific date, event type, and negation tag) for each expression. The server combines these attributes into a semantic information object. The output is semantic information containing tags such as time, event, and negation.
[0331] Step 8: The server retrieves objective factual information based on semantic information. The input includes semantic information and user location or context information. The server parses key parameters (e.g., date, location, event type) from the semantic information, sends a query request to an external information provider in the external information acquisition module, and receives the returned factual data (e.g., weather records, transaction records). The server formats and filters the returned data, retaining only fields relevant to the current statement. The output is a structured object of objective factual information.
[0332] Step 9: The server performs a consistency check between character information and objective factual information. The input includes semantic information, character information, and objective factual information. In its consistency check logic, the server compares the facts described in the statement (e.g., "no rain," "no transfer") with objective facts (e.g., rainfall amount, transfer records) according to preset rules. It uses logical operations to determine if there are any contradictions and generates preliminary evaluation data containing contradiction markers and explanations. The output is an evaluation message indicating whether the statement is consistent or contradictory.
[0333] Step 10: The server constructs prompts for the generative AI model. Inputs include character information, semantic information, objective factual information, and evaluation information. In the prompt generation module, the server concatenates this information into natural language text according to a predefined template, including role settings, task descriptions, factual descriptions, and output requirements. During string processing, the server ensures the information is complete and clearly structured so that the generative AI model can parse it correctly. The output is the prompt used to drive the generative AI model.
[0334] Step 11: The server invokes a generative artificial intelligence model to generate perfunctory evaluation results. The input is the prompt statement generated in step 10. Within the generative artificial intelligence invocation module, the server encodes the prompt statement as a labeled sequence and inputs it into the generative artificial intelligence model. Internally, the model uses embedding layers, self-attention layers, and feedforward network layers to vectorize and correlate the prompt statement. Based on the trained parameters, it calculates the probability distribution of each output label and decodes it into natural language text. The server receives this text output and can, as needed, parse the perfunctory level, reasons, and suggestions into structured fields. The output is the perfunctory evaluation result text and the corresponding structured data.
[0335] Step 12: The server generates a perjury determination result and updates the database record. Inputs include consistency evaluation information and the evaluation result provided by the generative artificial intelligence model. In the result integration module, the server merges the two according to predetermined weights or rules, determines the final perjury level and description, and writes this result into the perjury determination result field of the corresponding record in the database. Output is the updated database record with the final determination result.
[0336] Step 13: The server generates behavioral suggestion prompts based on the perjury determination result and objective factual information. The inputs are the perjury determination result, objective factual information, and (optionally) emotional state information. In the prompt generation module, the server constructs prompts tailored to the behavioral suggestion task, embedding risk level, factual background, and user emotion into a natural language description, and providing explicit output format requirements. The output is the prompt text used to generate behavioral suggestions.
[0337] Step 14: The server invokes a generative artificial intelligence model to generate behavioral suggestion information. The input is the prompt statement generated in step 13. The server then inputs the prompt statement back into the generative artificial intelligence model, which generates text containing risk warnings, reassuring explanations, and specific action suggestions according to the task requirements in the prompt. The server performs necessary truncation or formatting on the output text, extracting the core suggestion content. The output is the behavioral suggestion information text and its summary.
[0338] Step 15: The server sends the perjury determination result and behavioral recommendations to the terminal. The input includes the perjury determination result, evaluation reasons, and behavioral recommendations. The server constructs a response message in its output module, encoding the above data into a format suitable for network transmission, and sends it to the client application on the terminal via the network interface. The output is a response data packet containing the analysis results and recommendations.
[0339] Step 16: The terminal receives and displays the results returned by the server. The input is the response data packet sent by the server. The terminal parses the message in the network communication unit, extracting the perjury level, explanation, and behavioral recommendations. This information is then presented to the user in the results display unit using a combination of icons, colors, and text; for example, high-risk statements are highlighted and recommendations are displayed in a dialog box. The output is a visual analysis results interface on the terminal screen.
[0340] Step 17: Users make decisions based on the results displayed on the terminal. The input consists of the perjury assessment results and behavioral suggestions displayed on the terminal. After reading the relevant information, users select specific actions based on system prompts, such as terminating the current communication, contacting official agencies, requesting evidence from the other party, or ignoring suspicious requests. The output is the actual action taken by the user in the real world, which is influenced by the data processing and analysis results performed by the server, thereby achieving technical assistance in controlling risk scenarios.
[0341] Application Example 2 The process flow corresponding to the specific processing in Use Case 2 will be described below. The various parts of the system described below are implemented by the data processing device 12 and the intelligent device 14. In addition, the data processing device 12 is referred to as the "server" and the intelligent device 14 is referred to as the "terminal".
[0342] In modern communication environments, a large number of voice calls, online conferences, and remote consultations are conducted through communication networks. While existing technologies exist that convert speech to text and perform simple keyword matching or rule-based judgments, the following unresolved computer technology problems remain: First, traditional systems struggle to provide timely and accurate structured evaluations of the falsity or authenticity of statements. Existing technologies often perform only superficial text matching, lacking the ability to compare natural language statements with multi-source factual data and output falsity indicators and justifications in a machine-processable manner. This limits the effectiveness of computers in assisting users in identifying misinformation.
[0343] Second, existing communication systems typically only provide fixed automatic response procedures, lacking the ability to automatically construct prompts based on speech content analysis and intelligently invoke generative artificial intelligence models to generate high-quality judgments and suggestions. In other words, computers cannot automatically integrate speech recognition results, natural language parsing results, and fact retrieval results into prompts internally, and then hand them over to generative artificial intelligence models to generate structured outputs for decision support, thus limiting the intelligence level of automatic response systems in complex scenarios.
[0344] Third, existing systems typically only perform simple emotion label recognition when processing user emotions, failing to integrate emotion analysis results with false positive evaluations into a unified decision-making process. The computer system cannot dynamically adjust the input prompts of the generative AI model based on the user's emotional state to generate adaptive behavioral suggestions in terms of tone, content, and urgency. This results in a lack of sophisticated human-computer interaction capabilities when dealing with high-pressure or high-risk call scenarios (such as suspected fraud or emergency calls).
[0345] Fourth, in traditional architectures, speech recognition, natural language processing, fact comparison, sentiment analysis, and automatic response are often implemented in a fragmented manner by independent modules, lacking an integrated data flow and control flow design. The absence of a unified server-side control solution—from "switching communication to automatic response," to "multi-stage data processing," to "constructing prompt statements and invoking generative AI models," and "persisting results in a structured form and sending them back to the front end"—makes it difficult to significantly improve the efficiency, accuracy, and intelligence of computer processing of communication content at the system level.
[0346] Therefore, it is necessary to propose a new system architecture and processing method that enables the server to: automatically switch the response subject upon receiving communication, uniformly perform speech recognition, natural language parsing, fact retrieval, and falsity assessment, automatically generate prompt statements and invoke generative artificial intelligence models to complete high-quality reasoning, and simultaneously adjust the output behavior suggestions based on sentiment analysis results. In this way, the overall process of communication content parsing and response generation is improved from a computer technology perspective, enhancing the accuracy and real-time performance of falsified information identification and improving the system's intelligent response capabilities in complex human-computer interaction scenarios.
[0347] The specific processing performed by the specific processing unit 290 of the data processing apparatus 12 in Application Example 2 is achieved by the following means.
[0348] In this invention, the server includes means for switching received communication from human response to automatic response device response; means for acquiring voice information from a communication terminal or information processing device and converting the voice information into text information using speech recognition technology; means for parsing the text information using natural language processing technology and extracting target events, evaluation expressions, and numerical expressions related to the speech content; means for retrieving factual information from external information sources or storage devices based on the extraction results and generating comparison information indicating the correspondence between the speech content and the factual information; means for generating a prompt statement containing the speech content, the comparison information, and indication information related to the evaluation of falsity or authenticity, and inputting the prompt statement into a generative artificial intelligence model so that the generative artificial intelligence model evaluates the falsity or authenticity of the speech content; means for parsing the evaluation results output by the generative artificial intelligence model and generating evaluation information containing at least falsity indicators and their rationale; and means for associating the evaluation information with conversation information, storing it in a recording device, and prompting the user with the evaluation information through a display device or notification device. This allows for the formation of an end-to-end data processing chain within the server, encompassing communication takeover, speech recognition, semantic parsing, fact comparison, and the invocation of generative artificial intelligence models based on prompt statements. The computer can output indicators and reasons for the falsity of statements in a structured manner, achieving efficient and automatic determination of the authenticity of communication content and improving the system's ability to identify and process false information.
[0349] In this invention, the server may further include means for generating a prompt statement containing instructions for the user's next action based on the evaluation information and the session information, inputting the prompt statement into a generative artificial intelligence model to generate candidate actions for the response, and outputting at least a portion of the candidate actions as behavioral suggestion information. This allows the server to continue using the generative artificial intelligence model for decision-level reasoning after receiving a false evaluation, automatically outputting user-oriented action plan candidates. This enables continuous processing from "content understanding" to "action suggestion generation" within the computer, reducing manual analysis intervention and improving the intelligence level of the automatic response system.
[0350] In this invention, the server may further include means for inputting user-related voice or text information into an emotion analysis device to obtain emotion information representing emotional state, generating behavioral suggestion prompts adjusted in tone, content, or urgency based on the emotion information and the evaluation information, and inputting the prompts into a generative artificial intelligence model to generate behavioral suggestion information adjusted according to the emotional state. This allows the server to simultaneously consider the user's emotional state and the falsity of evaluations of the communication content when generating behavioral suggestions, dynamically adjusting prompts and output strategies. This enables the generative artificial intelligence model to generate responses that are more context-appropriate in terms of tone and urgency, improving the human-computer interaction experience and more effectively assisting users in making decisions under high-risk or high-stress scenarios. Thus, at the computer technology level, this achieves comprehensive optimization of the functionality and interactivity of the automatic response process.
[0351] "Communication" refers to the process of exchanging voice, text, or data information between a terminal and an information processing device via wired or wireless networks.
[0352] "Human response" refers to a response method in which human operators directly listen to, understand, and respond to received communications via voice or text.
[0353] "Automatic response device" refers to an information processing device or software module that automatically receives communications and generates response content using a pre-set program without the need for human operator intervention.
[0354] "Communication terminal" refers to a user-side device used to receive or send communications, including telephone terminals, mobile terminals, computing terminals, or other electronic devices with communication functions.
[0355] "Information processing device" refers to an electronic computing device that has the function of processing, storing and forwarding communication data, including servers, computers or embedded processing devices.
[0356] “Voice information” refers to audio data containing the speaker’s speech content, which is acquired by the voice acquisition unit and represented in digital form.
[0357] "Speech recognition technology" refers to signal processing and pattern recognition technology that converts input speech information into corresponding text information through acoustic feature analysis and language model calculation.
[0358] “Text information” refers to text data obtained by speech recognition technology or other means, represented in the form of character sequences, that can be parsed by a natural language processing module.
[0359] Natural Language Processing (NLP) refers to computer language processing technologies that perform word segmentation, part-of-speech tagging, syntactic analysis, and semantic recognition on textual information to obtain semantic structure and pragmatic information.
[0360] "Target event" refers to an event-related information unit that is extracted from the content of a speech through natural language processing technology and is related to objective facts or behaviors.
[0361] "Evaluative expression" refers to the linguistic expression in which one evaluates, judges, or asserts a target event, object, or state in the content of a speech.
[0362] "Numerical expression" refers to the expression used in speech to represent numerical information such as quantity, proportion, frequency, time, and amount.
[0363] "External information sources" refers to factual data provision devices or services that are independent of this system and accessible via network or interface, including public databases or third-party data service systems.
[0364] "Storage device" refers to a data storage component used to store data such as session information, factual information, and evaluation information, including disk storage, semiconductor storage, or database systems.
[0365] "Factual information" refers to data that is pre-recorded or obtained from external information sources, reflecting objective situations or historical records, and is used to compare with the content of the speech.
[0366] "Comparison information" refers to associated data generated based on the correspondence between spoken content and factual information, used to indicate consistency or difference.
[0367] "Prompt statements" refer to textual instructions that are constructed to enable generative artificial intelligence models to perform specific processing tasks, and include task descriptions, input conditions, and constraints.
[0368] "Generative artificial intelligence models" refer to artificial intelligence models trained using machine learning or deep learning methods that can automatically generate text output based on input prompts.
[0369] "Falsehood" refers to the degree to which the content of a statement is inconsistent with, exaggerated, or fabricated relative to factual information.
[0370] "Authenticity assessment" refers to the process of judging whether the content of a statement conforms to objective facts and providing a result.
[0371] "Evaluation results" refer to the analytical conclusions output by the generative artificial intelligence model regarding the falsity or authenticity of the speech content, including qualitative descriptions or quantitative indicators.
[0372] "Falsehood index" refers to numerical or graded information used to quantify the degree of falsehood in the content of a statement.
[0373] "Reasoning information" refers to textual explanations based on factual information and semantic analysis results that interpret falsehood indicators or truthfulness judgments.
[0374] "Evaluation information" refers to comprehensive data that includes at least indicators of falsity and reasons, used to represent the analysis results of the authenticity of the statements.
[0375] "Session information" refers to the collection of voice information, text information, and related metadata generated and recorded during communication.
[0376] "Recording device" refers to a hardware or software component used to store and manage session information, evaluation information, and related processing results.
[0377] "Display device" refers to an output device used to present evaluation information or behavioral suggestions to users in a visual form, including a display screen or other graphic output device.
[0378] "Notification device" refers to an output device used to send prompts or warnings to users through sound, vibration, pop-ups, push notifications, etc.
[0379] "User" refers to an individual or organization that uses the system to communicate, receive evaluation information and behavioral suggestions.
[0380] "Session information associated storage" refers to the operation of establishing an association between evaluation information and corresponding session content through identifiers or indexes and storing them together in the recording device.
[0381] "Conversation information prompts" refer to the output behavior of display devices or notification devices that enable users to identify and understand evaluation information or behavioral suggestions.
[0382] "Response actions" refer to the specific operations or processing steps that users can choose to perform after receiving evaluation information or behavioral suggestions.
[0383] "Behavioral suggestion information" refers to output information generated based on evaluation information, which instructs users on the next action they should take.
[0384] "Emotion analysis device" refers to a computing device or software module used to infer and output the user's emotional state based on voice or text information.
[0385] "Emotional information" refers to structured data output by an emotion analysis device that represents the category and intensity of a user's emotions.
[0386] "Emotional state" refers to an emotional state identified by an emotion analysis device, such as anxiety, anger, fear, happiness, or neutrality.
[0387] "Behavioral suggestion prompts" refer to prompts that are specifically designed and input into generative artificial intelligence models to take emotional states into account when generating behavioral suggestion information.
[0388] "Adjustment of tone, content, or urgency" refers to the automatic modification of behavioral suggestions and their output in terms of expression, level of detail, and urgency based on emotional and evaluative information.
[0389] In the following description, the server, client, and user are each referred to as subjects to specifically describe one or more embodiments of the system of the present invention. The present invention is not limited to the specific configurations and software names described below; these names are merely illustrative examples.
[0390] I. Overall System Composition A server may include one or more central processing units (CPUs), graphics processing units (GPUs, such as general-purpose graphics processing chips with parallel computing capabilities), main memory, non-volatile memory, a network interface, and an operating system (such as a Unix-like system) running on it. The server installs application server programs, a database management system, and artificial intelligence inference services on this hardware.
[0391] The terminal can be a mobile terminal, a fixed-line telephone terminal, a personal computing device, etc. The terminal includes a microphone, a speaker, a display device, local storage and a communication module, runs a mobile operating system or a general operating system, and communicates with the server through a network.
[0392] Users use the terminal to make voice calls or communicate online, and when needed, they can switch the call to be handled by the automatic response device on the server side through the terminal interface.
[0393] II. Program Generation and Module Structure The server can generate and execute programs to implement the various functions of this invention. These programs can be logically divided into the following modules: a communication management module, a speech processing module, a natural language processing module, a fact comparison module, a prompt statement generation module, a generative artificial intelligence model inference module, a sentiment analysis interface module, an evaluation and suggestion generation module, and a result storage and display module. The modules interact through predefined data structures, including session records, message records, comparison records, evaluation records, and sentiment records.
[0394] The server can create multiple data tables in the database management system, such as: a session table to store session identifiers, user identifiers, start and end times, etc.; a message table to store timestamps, speakers, and transcribed text for each voice or text message; a fact table to store factual data for products, services, events, etc.; an evaluation table to store falsehood indicators and reasoning information; and a sentiment table to store sentiment types and intensities.
[0395] III. Implementation Forms of Voice Data Acquisition and Conversion During a call, the terminal can continuously collect voice samples, which are digitized at a preset sampling rate (e.g., 16 kHz mono) and quantization precision (e.g., 16-bit depth). The terminal can acquire raw PCM data through the audio acquisition interface provided by the operating system, and then send it to the server in segments via a secure transmission protocol.
[0396] After receiving the speech segments, the server can choose to process them using a cloud-based speech recognition service or a local speech recognition library, depending on its configuration. For example, the server can call a speech recognition service interface to upload the audio segments, perform acoustic feature extraction, acoustic model inference, and language model decoding remotely, and return the corresponding text string and confidence value. Alternatively, the server can use a locally running speech recognition library to convert the speech frame sequence into a text sequence using algorithms such as Mel-frequency cepstral coefficient extraction and Hidden Markov Model decoding.
[0397] After receiving the text information, the server stores it in a message table using UTF-8 encoding, and appends a speaker identifier (e.g., the other party, the user) and a timestamp to the record. This structured storage allows subsequent modules to perform semantic analysis and spoofing checks without accessing the original audio, thus reducing storage and communication load.
[0398] IV. Implementation Forms of Natural Language Parsing and Fact Data Comparison The server can use a natural language processing library to process text information. This library can perform word segmentation, part-of-speech tagging, dependency parsing, and named entity recognition. By calling this library, the server can parse each text message into an internal representation composed of words, phrases, and dependency relations, and extract target events (e.g., "a certain product is the best-selling product on the market"), evaluative expressions (e.g., "best-selling," "100% safe"), and numerical expressions (e.g., percentage, frequency, quantity).
[0399] The server can use the extracted entity name and numerical fields as search keys to query the corresponding factual information in the factual data store. Factual data may include historical sales records, market share, number of accidents or complaints, regulatory announcements, etc. The server can quickly locate the corresponding record using an index structure (such as an inverted index or a B-tree index) and compare the differences between the numerical values and order relationships in the statements and the factual records to generate comparison information. The comparison information may include fields such as matching degree, numerical deviation, and time interval differences, and is stored in a structured format in the comparison record table.
[0400] By parsing natural language and comparing it with facts, the server internally forms an intermediate representation characterized by words and numbers, so that subsequent evaluations do not need to parse natural language again, which greatly improves the overall processing speed and reduces redundant calculations.
[0401] V. Prompt Statement Generation and the Structure and Training of Generative Artificial Intelligence Models After obtaining textual and comparison information, the server can construct prompts. These prompts consist of a task description, spoken text, a factual summary, and output format requirements, and are written in natural language to facilitate understanding by generative artificial intelligence models.
[0402] For example, the server can generate the following prompt: "Please assess whether this statement is likely exaggerated or false based on the following factual data."
[0403] Statement: 'This product is the best-selling on the market and is 100% safe.' Factual data: 1. This product had a market share of 8% over the past 12 months.
[0404] 2. The market share of competing products is 15%.
[0405] 3. There have been 3 safety complaints related to this product in the past 3 years.
[0406] Please output: 1. A false rating between 0 and 1; 2. Briefly explain your reasoning (no more than 100 words). The generative AI model used by the server can be a sequence-to-sequence model based on a multi-layer Transformer structure. This model consists of word embedding layers, multi-head self-attention layers, feedforward network layers, and layer normalization, with parameters obtained through large-scale unsupervised pre-training and supervised fine-tuning. During training, the server can use a large number of samples containing factual background and speech pairs, setting the loss function to the cross-entropy of the generated text and the reference text, and adding a mean squared error term for regressing the falsehood score on some tasks. The model weights are updated using backpropagation and gradient descent.
[0407] Training data can be obtained through data augmentation, such as word order transformation and synonym replacement of factual descriptions, and noise injection into spoken text, thereby improving the model's robustness under different expressions. The server utilizes GPUs to perform batch training of the model to improve training speed.
[0408] During the inference phase, the server converts the prompt into a sequence of sub-word units, performs forward computation through an embedding layer and multi-head self-attention, and generates the output word sequence. The server can employ a beam search strategy during decoding to strike a balance between output quality and computational cost. The output includes a falsehood index and justification information, which the server can parse using rules, such as extracting scores within numerical ranges from the text and treating the remaining content as justification.
[0409] Because the prompt statements contain structured fact comparison results, generative AI models can utilize these features during reasoning to compare the consistency between statements and facts, thereby more accurately outputting falsehood indicators. This approach of processing "fact-constrained prompt statements" differs from traditional models that rely solely on language patterns, and helps improve the reliability of judgments.
[0410] VI. Sentiment Analysis and Adaptive Emotion Prompt Statements The server can acquire user emotions through a sentiment analysis device. This device can be an acoustic sentiment recognition model based on a combination of convolutional neural networks and recurrent neural networks, or a text sentiment classifier based on a pre-trained language model. The server converts the user's speech signal into acoustic features (e.g., Mel spectrograms), inputs them into the sentiment model, and outputs the sentiment category and intensity; or it inputs the user's most recent text messages into a text sentiment network, outputs a sentiment vector for each message, and then obtains the overall sentiment through pooling.
[0411] The server can combine sentiment information with false ratings to generate behavioral suggestions using prompts. For example, when a false rating is high and the user is in an anxious state, the server can generate the following prompt: Scenario: A user is on the phone with a salesperson, discussing an investment project.
[0412] The AI rated the falsity of the other party's statement as 0.85 (high), based on the fact that the project's promised returns were far higher than the industry average and that there was a lack of publicly available audit reports.
[0413] The user's current sentiment analysis result is 'anxious', with an intensity of 0.78.
[0414] Task: 1. Use a calm and reassuring tone to provide the user with 2-3 suggestions for the next steps; 2. Suggestions may include: suspending transfers, verifying company qualifications, and consulting professionals; 3. Answers should not exceed 150 words. The server inputs the prompt into the same generative AI model or a model fine-tuned for the behavioral suggestion task, enabling the model to automatically adjust the tone, content, and urgency when generating suggestions. Internally, this adjustment is achieved by injecting emotion labels and intensity vectors into the input sequence, causing the attention mechanism to give higher weight to emotion-related prompts during the decoding phase, thus reflecting different styles such as soothing, warning, or urgent guidance in the output text.
[0415] Because emotional information directly participates in the construction of prompts and model inputs, the system does not simply attach emotion labels to the output, but allows the model to select different language templates and strategies based on emotional signals during the generation process, thus forming a unique processing path within the computer that is different from the human decision-making process.
[0416] VII. Generation of Behavioral Suggestion Information and its Technical Effects After receiving false reviews and sentiment information, the server can generate behavioral suggestions. These suggestions include natural language descriptions of specific actions to take, such as "suggest postponing the transfer," "suggest verifying the other party's qualifications through official channels," and "suggest contacting legal professionals." The server stores these suggestions in an evaluation table and displays them to the user through the terminal.
[0417] Unlike traditional warnings triggered by fixed rules or simple keywords, this invention uses a server that processes data at multiple levels and uses generative artificial intelligence models to generate suggestions in a structured manner that balance factual comparison and sentiment adaptation. This approach has the following technical advantages: By unifying speech recognition, natural language parsing, fact retrieval, and generative model inference into a single data stream, the server reduces the overhead of redundant parsing and data format conversion between multiple systems, thereby improving overall processing speed.
[0418] By using fact-based prompts and inputs with emotional signals, the server generates results that more accurately match the current communication scenario in terms of semantics and tone, thereby improving accuracy in identifying false information and providing risk warnings, and reducing false positives and false negatives.
[0419] The server manages session, comparison, and evaluation information through a unified data structure, which facilitates subsequent statistics and model retraining, thereby continuously improving model performance and overall system reliability.
[0420] 8. Integration with real-world technology applications After receiving the false alarm indicators and behavioral suggestions from the server, the terminal can graphically display the risk level, explanation, and suggested actions on the screen. For example, the terminal can display a red banner on the call interface indicating "The current message carries a high risk of being a false alarm," followed by specific suggestions. The terminal can also trigger different levels of alerts based on the urgency parameters output by the server, such as pop-ups, sound alerts, and vibration reminders.
[0421] In security scenarios, the terminal can interact with other devices. For example, when the server suggests an alarm and the urgency level exceeds a threshold, the terminal can provide a one-click dialing option and automatically fill in relevant information; in enterprise call center scenarios, the terminal can automatically transfer the call to a specific department or automatically record it as a high-risk event based on the server's suggestion.
[0422] The server compresses and transmits data on demand over the communication network, significantly reducing bandwidth consumption by sending only transcribed text and structured evaluation information, rather than long streams of raw audio. The terminal's local caching of some results also reduces the server load.
[0423] Because this invention employs a GPU with parallel computing capabilities and an efficient neural network structure on the server side, compared with traditional systems that rely entirely on manual review or simple rule matching, it can significantly shorten response time while improving judgment accuracy, and return risk warnings to the terminal while communication is still in progress, achieving near real-time technical response.
[0424] IX. Alternative Implementation Forms and Variations When implementing natural language processing, servers can use different combinations of algorithms, such as using dependency parsing models or graph neural networks to model sentence structure; when comparing facts, knowledge graphs can also be introduced to identify implicit relationships between statements and facts through graph traversal algorithms.
[0425] The generative artificial intelligence model used by the server can be replaced with other neural network models, such as recurrent neural networks with encoder-decoder structures, or text generation networks based on convolutional structures. As long as the model can generate text output containing falsehood indicators and behavioral suggestions based on prompts, it can be applied to this invention.
[0426] Servers can design multi-head outputs for generative AI models according to different tasks. For example, one part of the output is used to generate fake values, and another part is used to generate natural language explanations. By sharing intermediate representations through multi-task learning, the model's ability to represent complex semantics can be improved.
[0427] Terminals can also offload some functions. For example, when network conditions are poor, the terminal can perform speech recognition locally and send the obtained text to the server through a low-bandwidth channel. The server only needs to process the text and compare the facts, thereby reducing the network bandwidth requirements while ensuring the integrity of the functions.
[0428] In certain scenarios, users can choose to enable only false ratings and disable sentiment-adaptive suggestions. In this case, the server can omit the sentiment analysis step and directly construct prompts and output suggestions based on factual comparisons to adapt to deployment environments with limited processing resources.
[0429] Through the above-mentioned various implementation forms and variations, it can be seen that the system of the present invention not only realizes the structured analysis and risk assessment of communication content, but also forms an efficient and scalable processing link inside the computer through the technical configuration of specific data structures, prompt statement construction methods and generative artificial intelligence models, thereby improving the performance and quality of communication content analysis and automatic response in terms of technology.
[0430] use Figure 14 The processing procedure is explained.
[0431] Step 1: Users can initiate or receive communications using the terminal and choose whether to switch to automatic answer mode.
[0432] Input: incoming call signal or call setup signal, user operation events on the terminal interface.
[0433] Output: Control command indicating whether to switch to automatic response mode.
[0434] Specific data processing and data operations: Users answer calls or join voice conversations through the terminal interface. The terminal receives the incoming call event in the operating system and displays the call interface on the screen. If the user clicks the "AI Auto-Answer" button, the terminal converts the click event into an internal flag (e.g., a Boolean value) and generates a control message containing a session identifier, user identifier, and switching instructions. The terminal then sends this control message to the server via a secure communication protocol, thus triggering the subsequent auto-answer processing.
[0435] Step 2: The terminal collects the voice recordings of the call and sends the voice data to the server.
[0436] Input: Raw audio signal from the microphone and call session identifier.
[0437] Output: Segmented digital voice data packets.
[0438] Specific data processing and data operations: The terminal converts analog speech into a digital PCM stream via an audio acquisition interface at a preset sampling rate. Locally, the terminal frames and segments the audio stream (e.g., every 2-3 seconds as a segment), and adds a session identifier and timestamp to each segment. The terminal can compress the audio segments according to an encoding algorithm (e.g., compressing them to an audio format with a specific bit rate) to reduce bandwidth consumption. The terminal then sends these audio data packets to the server via a network protocol, forming a continuous audio data stream input.
[0439] Step 3: The server receives voice data and performs speech recognition to convert the speech into text.
[0440] Input: Voice data packets from the terminal (digital audio containing session identifier and timestamp).
[0441] Output: Text information with timestamps and speaker identifiers.
[0442] Specific data processing and data operations: The server receives audio data packets through the network interface, categorizes the data streams according to session identifiers, and arranges them in chronological order in a buffer queue. The server preprocesses the audio data, including resampling, noise suppression, and energy normalization, to generate a standard audio format for speech recognition. The server invokes the speech recognition engine to convert audio segments into text sequences and confidence scores. Based on call routing information, the server labels the speaker as "the other party" or "the user" and writes the transcribed text along with the corresponding timestamp into the message log. In this process, the audio waveform is mapped to feature vectors, and these feature vectors are used in matrix operations of acoustic and language models to generate text output, achieving the information transformation from speech to text.
[0443] Step 4: The server performs natural language processing on text information, extracting target events, evaluation expressions, and numerical expressions.
[0444] Input: Text information (sentence-level text and its metadata) generated by the speech recognition module.
[0445] Output: Structured data containing target events, evaluation expressions, numerical expressions, and syntactic relations.
[0446] Specific data processing and data operations: The server invokes a natural language processing library to segment and tag each text, dividing continuous character sequences into words and labeling them with their part-of-speech tags. The server further performs dependency parsing to create a dependency graph between words in the sentence. Based on predefined pattern rules and classification models, the server identifies predicate phrases declaring facts (as target events), adjectives and adverbs with superlative, absolute, or affirmative tones (as evaluative expressions), and numerical expressions such as percentages, quantities, and frequencies (as numerical expressions). The server encapsulates these results into structured records, including text fragments, their position in the original sentence, and semantic role information. In this way, the raw natural language is converted into a machine-computable feature set, facilitating subsequent fact comparison.
[0447] Step 5: The server retrieves factual data based on the extraction results and generates comparison information.
[0448] Input: Structured data containing the target event, entity name, and numerical representation.
[0449] Output: A record of the comparison between the content of the statement and the factual data.
[0450] Specific data processing and data operations: The server extracts entity names (such as product categories and service types), time ranges, and numerical fields from the extraction results, and uses these fields as search keys to query local or external fact databases. The server quickly locates relevant fact records, such as historical sales volume, market share, and number of complaints, using an index structure. The server compares the numerical values in the statements with those in the database, calculating differences, proportional deviations, or the degree to which they exceed thresholds, while also comparing the time intervals in the statements with the timestamp ranges of the fact records. The server generates comparison information, including matching metrics, difference type identifiers (such as exaggeration, understatement, or inconsistency), and a summary of reference facts, using this comparison information as the foundational features for subsequent falsehood assessments.
[0451] Step 6: The server generates a message to indicate whether the evaluation is fraudulent.
[0452] Input: Speech text, comparison information, and a description of the false positive assessment task.
[0453] Output: Prompt statements constructed in natural language.
[0454] Specific data processing and data operations: The server combines the original speech content, key facts retrieved from the fact base, and comparison results into a coherent explanatory text, along with task instructions and output format requirements. The server inserts various fields into the prompt statement according to a preset template through string concatenation and template filling, including: the sentence to be evaluated, a list of relevant facts, the range of falsehoods the model needs to output, and a length limit for the explanation. When generating the prompt statement, the server can also select different template variations based on the speech language and domain to improve the model's understanding, thus forming a complete text instruction that meets the input requirements of a generative artificial intelligence model.
[0455] Step 7: The server inputs the prompt statement into the generative artificial intelligence model to obtain a false evaluation result.
[0456] Input: A prompt containing task descriptions, speech content, and factual data.
[0457] Output: Falsehood indicators and explanations.
[0458] Specific data processing and data operations: The server segments the prompt statement into words or sub-words, mapping the text to a sequence of word vectors or sub-word embeddings. A pre-trained generative AI model is loaded onto the server's GPU. This model employs a multi-layered self-attention structure, internally transforming the input vector through matrix multiplication and non-linear activation. During forward inference, the model predicts the probability distribution of each word in the output sequence based on the context provided in the prompt statement. The server uses a beam search or greedy strategy to select output words from the probability distribution, progressively generating complete evaluation text. The server parses the falsehood index (e.g., a value between 0 and 1 or a "high / medium / low" rating) and the corresponding reasoning paragraphs from the generated text, saving them as the evaluation result.
[0459] Step 8: The server generates behavioral suggestions using prompt statements based on evaluation results and session information.
[0460] Input: falsehood indicator, reason information, and relevant session context.
[0461] Output: Prompt statements for the behavior suggestion generation task.
[0462] Specific data processing and data operations: The server reads the falsehood score and brief explanation from the evaluation results and combines them with the conversation context information (such as "investment project introduction" or "unsolicited call asking for money transfer"). Based on preset decision rules, the server categorizes the falsehood index into different risk levels and explicitly informs the generative AI model of the current risk level and the type of suggestion to be given in the prompt statement. The new prompt statement constructed by the server typically includes: a scenario description, a falsehood score and its explanation, and the expected number and style of suggestions to be output (e.g., calm, reassuring, strong warning), providing the model with sufficient contextual guidance when generating suggestions.
[0463] Step 9: The server invokes a generative artificial intelligence model to generate behavioral suggestion information.
[0464] Input: A prompt statement used to suggest behavior.
[0465] Output: One or more paragraphs of behavior suggestion text in natural language form.
[0466] Specific data processing and data operations: The server re-encodes the behavioral suggestions into a vector sequence using prompts and inputs it into the same generative AI model or a model optimized for behavioral suggestion tasks. Internally, the model uses a self-attention mechanism to focus on key phrases and risk markers in the scene description, thereby selecting appropriate coping strategy vocabulary during decoding. The server post-processes the model output, including removing redundancy, splitting it into entries, checking for length and content constraints, and combining the processed suggestions into behavioral suggestion information, such as "It is recommended not to transfer money for now," "It is recommended to verify the other party's identity through official channels," and "It is recommended to consult professionals before making a decision."
[0467] Step 10: The server adjusts behavioral suggestion information based on emotional information (optional).
[0468] Input: Emotional information output by the emotion analysis device and unadjusted behavioral suggestions.
[0469] Output: Behavioral suggestions adjusted in terms of tone, content, or urgency.
[0470] Specific data processing and data operations: The server obtains the user's current emotion category and intensity from the sentiment analysis module, such as "anxiety 0.7" or "fear 0.9". The server adjusts the prompt parameters based on the emotion intensity; for example, it increases the clarity and urgency of the suggested language under high fear, and increases the proportion of reassuring statements under moderate anxiety. The server can reconstruct prompts with emotion tags, using the original suggestions along with the emotion information as input, requesting the generative AI model to output suggestions more suited to the emotional state, such as adding "Please answer in a reassuring tone" to the text. After receiving the new output, the server replaces or supplements the original suggestions, thus technically matching the final text presented to the user with the emotional state.
[0471] Step 11: The server sends false evaluations and behavioral suggestions to the terminal, which then displays them.
[0472] Input: Server-generated fake evaluation information and behavioral suggestions.
[0473] Output: A risk warning and action suggestion interface displayed on the terminal display device.
[0474] Specific data processing and data operations: The server integrates false positive indicators, explanations, sentiment information, and behavioral suggestions into a unified data structure and sends it to the terminal as a message via a network interface. Upon receiving the message, the terminal uses a parsing library to convert the data into the format required by the interface components; for example, converting ratings into graphical progress bars and displaying text suggestions in line breaks in a list area. The terminal controls the interface color (e.g., red indicates high risk) and whether to display warning dialog boxes or play alert sounds based on the risk level. In this way, abstract evaluation results are transformed into graphical and textual information that users can directly perceive on the terminal, facilitating informed decision-making.
[0475] Step 12: Users refer to the information displayed on the terminal and select subsequent operations.
[0476] Input: The interface displayed on the terminal showing false evaluations and behavioral suggestions.
[0477] Output: The user's action selection on the terminal (e.g., hang up the call, confirm, initiate an alarm, etc.).
[0478] Specific data processing and data operations: Users can read the risk levels and suggestions displayed on the terminal and click on operation buttons on the interface, such as "End Call," "Contact Official Customer Service," and "One-Click Emergency Alarm." The terminal translates the user's choices into operation commands, which are then sent back to the server or invoked local functions (such as dialing or opening other applications) when necessary. Upon receiving new operation commands, the server can further record user behavior for subsequent statistical analysis and model adjustments. Through this closed loop, the system connects the output of the generative artificial intelligence model with real-world call control and device operation, realizing a complete technical link from data processing to device behavior.
[0479] The specific processing unit 290 sends the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires sound representing user input regarding the result of the specific processing. The control unit 46A sends the sound data representing user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the sound data.
[0480] Data generation model 58 is a so-called generative AI (Artificial Intelligence). Examples of data generation models 58 include ChatGPT (registered trademark) (accessible via the internet (URL: https: / / openai.com / blog / chatgpt)). Data generation model 58 is obtained through deep learning on a neural network. Input to data generation model 58 are prompt words containing instructions, and inference data such as sound data representing sound, text data representing text, and image data representing images (e.g., still image data or animation data). Data generation model 58 infers from the input inference data based on the instructions represented by the prompt words and outputs the inference result in one or more data forms, such as sound data, text data, and image data. Data generation model 58 includes, for example, text generation AI, image generation AI, and multimodal generation AI. Here, inference refers to, for example, analysis, classification, prediction, and / or induction. The specific processing unit 290 performs the aforementioned specific processing while using data generation model 58. The data generation model 58 can also be a model finely tuned to output inference results from prompts that do not contain instructions. In this case, the data generation model 58 can output inference results based on prompts that do not contain instructions. The data processing apparatus 12, etc., includes various data generation models 58, including AI other than the generation AI. AI other than the generation AI can be, for example, linear regression, logistic regression, decision trees, random forests, support vector machines (SVM), k-means clustering, convolutional neural networks (CNN), recurrent neural networks (RNN), generative adversarial networks (GAN), or Naive Bayes, and can perform various processes, but is not limited to this example. Furthermore, the AI can also be an AI agent. Furthermore, when the processing of the above-mentioned parts is performed by AI, the processing can be partially or entirely performed by AI, but is not limited to this example. Furthermore, the processing performed by AI including the generation AI can be replaced by processing in the rule base, and the processing in the rule base can also be replaced by processing performed by AI including the generation AI.
[0481] Furthermore, the processing of the aforementioned data processing system 10 is performed by the specific processing unit 290 of the data processing device 12 or the control unit 46A of the smart device 14, but it can also be performed by both the specific processing unit 290 of the data processing device 12 and the control unit 46A of the smart device 14. Additionally, the specific processing unit 290 of the data processing device 12 acquires or collects information required for processing from the smart device 14 or external devices, and the smart device 14 acquires or collects information required for processing from the data processing device 12 or external devices.
[0482] For example, the collection unit is implemented by the control unit 46A of the smart device 14 or the specific processing unit 290 of the data processing device 12. For example, the acquisition unit uses the camera 42 or communication I / F 44 of the smart device 14 to acquire step data, which is then processed by the specific processing unit 290 of the data processing device 12. For example, the analysis unit is implemented by the specific processing unit 290 of the data processing device 12, which analyzes the data from the collection unit and the acquisition unit. For example, the generation unit is implemented by the specific processing unit 290 of the data processing device 12, which uses a generation AI to generate a menu. For example, the serving unit is implemented by the output device 40 of the smart device 14 or the specific processing unit 290 of the data processing device 12, which provides the generated menu to the user. The correspondence between each unit and the device or control unit is not limited to the above examples and various changes can be made.
[0483] In the above embodiments, examples of specific processing by the data processing device 12 are given, but the technology disclosed herein is not limited to this, and specific processing may also be performed by the smart device 14.
[0484] Second Implementation Method Figure 3 An example of the configuration of the data processing system 210 according to the second embodiment is shown.
[0485] like Figure 3 As shown, the data processing system 210 includes a data processing device 12 and smart glasses 214. A server can be cited as an example of the data processing device 12.
[0486] The data processing apparatus 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" as understood in this disclosure. The computer 22 includes a processor 28, RAM 30, and memory 32. The processor 28, RAM 30, and memory 32 are connected to a bus 34. Furthermore, the database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0487] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, RAM 48, and memory 50. The processor 46, RAM 48, and memory 50 are connected to a bus 52. Furthermore, the microphone 238, speaker 240, camera 42, and communication I / F 44 are also connected to the bus 52.
[0488] Microphone 238 receives instructions from user 20 by receiving sounds emitted by user 20. Microphone 238 captures sounds emitted by user 20 and converts the captured sounds into sound data, which is then output to processor 46. Speaker 240 outputs sound according to instructions from processor 46.
[0489] Camera 42 is a small digital camera equipped with an optical system such as a lens, aperture and shutter, and imaging elements such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, which captures images of the user 20's surroundings (e.g., the field of view defined by an angle equivalent to the field of vision of an average healthy person).
[0490] Communication I / F44 is connected to network 54. Communication I / F44 and 26 are responsible for the transmission and reception of various information between processor 46 and processor 28 via network 54. The transmission and reception of various information between processor 46 and processor 28 using communication I / F44 and 26 is performed in a secure state.
[0491] Figure 4 This illustrates an example of the main functions of the data processing device 12 and the smart glasses 214. For example... Figure 4 As shown, in the data processing device 12, specific processing is performed by the processor 28. The specific processing program 56 is stored in the memory 32.
[0492] The specific processing program 56 is an example of a "program" involved in the technology of this disclosure. The processor 28 reads the specific processing program 56 from the memory 32 and executes the read specific processing program 56 on the RAM 30. Specific processing is implemented by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.
[0493] The memory 32 stores a data generation model 58 and an emotion-specific model 59. The data generation model 58 and the emotion-specific model 59 are used by the specific processing unit 290. The specific processing unit 290 can use the emotion-specific model 59 to infer the user's emotions and perform specific processing based on the user's emotions. In the emotion inference function (emotion-specific function) using the emotion-specific model 59, various inferences and predictions related to the user's emotions are performed, including inferences and predictions of the user's emotions, but this is not limited to this example. Furthermore, emotion inference and prediction may also include, for example, emotion analysis (parsing).
[0494] In the smart glasses 214, the processor 46 performs reception and output processing. The memory 50 stores the reception and output program 60. The processor 46 reads the reception and output program 60 from the memory 50 and executes the read reception and output program 60 on the RAM 48. The reception and output processing is implemented by the processor 46 operating as a control unit 46A according to the reception and output program 60 executed on the RAM 48. Furthermore, the smart glasses 214 has the same data generation model and emotion-specific model as the data generation model 58 and the emotion-specific model 59, and these models can also be used to perform the same processing as the specific processing unit 290.
[0495] Next, the specific processing of the specific processing unit 290 of the data processing device 12 will be described. Each part of the system described below is implemented by the data processing device 12 and the smart glasses 214. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".
[0496] Example 1 The process is the same as that of the specific process described in Embodiment 1 in the first embodiment above, so the description is omitted.
[0497] Application Example 1 The process is the same as that in the specific processing described in Application Example 1 of the first embodiment above, so the description is omitted.
[0498] Example 2 The process is the same as that of the specific process in Embodiment 2 described in the first embodiment above, so the description is omitted.
[0499] Application Example 2 The process is the same as that in the specific processing described in Application Example 2 of the first embodiment above, so the description is omitted.
[0500] The specific processing unit 290 sends the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A outputs the result of the specific processing to the speaker 240. The microphone 238 acquires sound input representing the user's input regarding the result of the specific processing. The control unit 46A sends the sound data representing the user's input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the sound data.
[0501] Data generation model 58 is a so-called generative AI (Artificial Intelligence). Examples of data generation models 58 include ChatGPT (registered trademark) (accessible via the internet (URL: https: / / openai.com / blog / chatgpt)). Data generation model 58 is obtained through deep learning on a neural network. Input to data generation model 58 are prompt words containing instructions, and inference data such as sound data representing sound, text data representing text, and image data representing images (e.g., still image data or animation data). Data generation model 58 infers from the input inference data based on the instructions represented by the prompt words and outputs the inference result in one or more data forms, such as sound data, text data, and image data. Data generation model 58 includes, for example, text generation AI, image generation AI, and multimodal generation AI. Here, inference refers to, for example, analysis, classification, prediction, and / or induction. The specific processing unit 290 performs the aforementioned specific processing while using data generation model 58. The data generation model 58 can also be a model finely tuned to output inference results from prompts that do not contain instructions. In this case, the data generation model 58 can output inference results based on prompts that do not contain instructions. The data processing apparatus 12, etc., includes various data generation models 58, including AI other than the generation AI. AI other than the generation AI can be, for example, linear regression, logistic regression, decision trees, random forests, support vector machines (SVM), k-means clustering, convolutional neural networks (CNN), recurrent neural networks (RNN), generative adversarial networks (GAN), or Naive Bayes, and can perform various processes, but is not limited to this example. Furthermore, the AI can also be an AI agent. Furthermore, when the processing of the above-mentioned parts is performed by AI, the processing can be partially or entirely performed by AI, but is not limited to this example. Furthermore, the processing performed by AI including the generation AI can be replaced by processing in the rule base, and the processing in the rule base can also be replaced by processing performed by AI including the generation AI.
[0502] Furthermore, the processing of the aforementioned data processing system 10 is performed by the specific processing unit 290 of the data processing device 12 or the control unit 46A of the smart glasses 214, but it can also be performed by both the specific processing unit 290 of the data processing device 12 and the control unit 46A of the smart glasses 214. Additionally, the specific processing unit 290 of the data processing device 12 acquires or collects information required for processing from the smart glasses 214 or external devices, and the smart glasses 214 acquires or collects information required for processing from the data processing device 12 or external devices.
[0503] For example, the collection unit is implemented by the control unit 46A of the smart glasses 214 or the specific processing unit 290 of the data processing device 12. For example, the acquisition unit uses the camera 42 or communication I / F 44 of the smart glasses 214 to acquire step data, which is then processed by the specific processing unit 290 of the data processing device 12. For example, the analysis unit is implemented by the specific processing unit 290 of the data processing device 12, which analyzes the data from the collection unit and the acquisition unit. For example, the generation unit is implemented by the specific processing unit 290 of the data processing device 12, which uses a generation AI to generate a menu. For example, the serving unit is implemented by the speaker 240 of the smart glasses 214 or the specific processing unit 290 of the data processing device 12, which provides the generated menu to the user. The correspondence between each unit and the device or control unit is not limited to the above examples and various changes can be made.
[0504] In the above embodiments, examples of specific processing by the data processing device 12 are given, but the technology disclosed herein is not limited to this, and specific processing may also be performed by the smart glasses 214.
[0505] Third Implementation Method Figure 5 An example of the configuration of the data processing system 310 according to the third embodiment is shown.
[0506] like Figure 5 As shown, the data processing system 310 includes a data processing device 12 and a head-mounted terminal 314. A server can be cited as an example of the data processing device 12.
[0507] The data processing apparatus 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" as understood in this disclosure. The computer 22 includes a processor 28, RAM 30, and memory 32. The processor 28, RAM 30, and memory 32 are connected to a bus 34. Furthermore, the database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0508] The head-mounted terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a display 343. The computer 36 includes a processor 46, RAM 48, and memory 50. The processor 46, RAM 48, and memory 50 are connected to a bus 52. Furthermore, the microphone 238, speaker 240, camera 42, display 343, and communication I / F 44 are also connected to the bus 52.
[0509] Microphone 238 receives instructions from user 20 by receiving sounds emitted by user 20. Microphone 238 captures sounds emitted by user 20 and converts the captured sounds into sound data, which is then output to processor 46. Speaker 240 outputs sound according to instructions from processor 46.
[0510] Camera 42 is a small digital camera equipped with an optical system such as a lens, aperture and shutter, and imaging elements such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, which captures images of the user 20's surroundings (e.g., the field of view defined by an angle equivalent to the field of vision of an average healthy person).
[0511] Communication I / F44 is connected to network 54. Communication I / F44 and 26 are responsible for the transmission and reception of various information between processor 46 and processor 28 via network 54. The transmission and reception of various information between processor 46 and processor 28 using communication I / F44 and 26 is performed in a secure state.
[0512] Figure 6 This illustrates an example of the main functions of the data processing device 12 and the head-mounted terminal 314. For example... Figure 6 As shown, in the data processing device 12, specific processing is performed by the processor 28. The specific processing program 56 is stored in the memory 32.
[0513] The specific processing program 56 is an example of a "program" involved in the technology of this disclosure. The processor 28 reads the specific processing program 56 from the memory 32 and executes the read specific processing program 56 on the RAM 30. Specific processing is implemented by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.
[0514] The memory 32 stores the data generation model 58 and the emotion-specific model 59. The data generation model 58 and the emotion-specific model 59 are used by the specific processing unit 290.
[0515] In the head-mounted terminal 314, the processor 46 performs the acceptance / output processing. The memory 50 stores the acceptance / output program 60. The processor 46 reads the acceptance / output program 60 from the memory 50 and executes the read acceptance / output program 60 on the RAM 48. The acceptance / output processing is implemented by the processor 46 operating as a control unit 46A according to the acceptance / output program 60 executed on the RAM 48.
[0516] Next, the specific processing of the specific processing unit 290 of the data processing device 12 will be described. Each part of the system described below is implemented by the data processing device 12 and the head-mounted terminal 314. In the following description, the data processing device 12 will be referred to as the "server" and the head-mounted terminal 314 will be referred to as the "terminal".
[0517] Example 1 The process is the same as that of the specific process described in Embodiment 1 in the first embodiment above, so the description is omitted.
[0518] Application Example 1 The process is the same as that in the specific processing described in Application Example 1 of the first embodiment above, so the description is omitted.
[0519] Example 2 The process is the same as that of the specific process in Embodiment 2 described in the first embodiment above, so the description is omitted.
[0520] Application Example 2 The process is the same as that in the specific processing described in Application Example 2 of the first embodiment above, so the description is omitted.
[0521] The specific processing unit 290 sends the result of the specific processing to the head-mounted terminal 314. In the head-mounted terminal 314, the control unit 46A outputs the result of the specific processing to the speaker 240 and the display 343. The microphone 238 acquires sound input representing the user's input regarding the result of the specific processing. The control unit 46A sends the sound data representing the user's input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the sound data.
[0522] Data generation model 58 is a so-called generative AI (Artificial Intelligence). Examples of data generation models 58 include ChatGPT (registered trademark) (accessible via the internet (URL: https: / / openai.com / blog / chatgpt)). Data generation model 58 is obtained through deep learning on a neural network. Input to data generation model 58 includes prompt words containing instructions, and inference data such as sound data representing sound, text data representing text, and image data representing images (e.g., still image data or animation data). Data generation model 58 infers the input inference data based on the instructions represented by the prompt words and outputs the inference result in one or more data forms such as sound data, text data, and image data. Data generation model 58 includes, for example, text generation AI, image generation AI, and multimodal generation AI. Here, inference refers to, for example, analysis, classification, prediction, and / or induction. The specific processing unit 290 performs the aforementioned specific processing while using data generation model 58. The data generation model 58 can also be a model finely tuned to output inference results from prompts that do not contain instructions. In this case, the data generation model 58 can output inference results based on prompts that do not contain instructions. The data processing apparatus 12, etc., includes various data generation models 58, including AI other than the generation AI. AI other than the generation AI can be, for example, linear regression, logistic regression, decision trees, random forests, support vector machines (SVM), k-means clustering, convolutional neural networks (CNN), recurrent neural networks (RNN), generative adversarial networks (GAN), or Naive Bayes, and can perform various processes, but is not limited to this example. Furthermore, the AI can also be an AI agent. Furthermore, when the processing of the above-mentioned parts is performed by AI, the processing can be partially or entirely performed by AI, but is not limited to this example. Furthermore, the processing performed by AI including the generation AI can be replaced by processing in the rule base, and the processing in the rule base can also be replaced by processing performed by AI including the generation AI.
[0523] Furthermore, the processing of the aforementioned data processing system 10 is performed by the specific processing unit 290 of the data processing device 12 or the control unit 46A of the head-mounted terminal 314, but it can also be performed by both the specific processing unit 290 of the data processing device 12 and the control unit 46A of the head-mounted terminal 314. Additionally, the specific processing unit 290 of the data processing device 12 acquires or collects information required for processing from the head-mounted terminal 314 or external devices, and the head-mounted terminal 314 acquires or collects information required for processing from the data processing device 12 or external devices.
[0524] For example, the collection unit is implemented by the control unit 46A of the head-mounted terminal 314 or the specific processing unit 290 of the data processing device 12. For example, the acquisition unit uses the camera 42 or communication I / F 44 of the head-mounted terminal 314 to acquire step data, which is then processed by the specific processing unit 290 of the data processing device 12. For example, the analysis unit is implemented by the specific processing unit 290 of the data processing device 12 to analyze the data from the collection unit and the acquisition unit. For example, the generation unit is implemented by the specific processing unit 290 of the data processing device 12 to generate a menu using a generation AI. For example, the serving unit is implemented by the speaker 240 and display 343 of the head-mounted terminal 314 or the specific processing unit 290 of the data processing device 12 to provide the generated menu to the user. The correspondence between each unit and the device or control unit is not limited to the above examples and various changes can be made.
[0525] In the above embodiments, examples of specific processing by the data processing device 12 are given, but the technology disclosed herein is not limited to this, and specific processing may also be performed by the head-mounted terminal 314.
[0526] Fourth Implementation Method Figure 7 An example of the configuration of the data processing system 410 according to the fourth embodiment is shown.
[0527] like Figure 7 As shown, the data processing system 410 includes a data processing device 12 and a robot 414. A server can be cited as an example of the data processing device 12.
[0528] The data processing apparatus 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" as understood in this disclosure. The computer 22 includes a processor 28, RAM 30, and memory 32. The processor 28, RAM 30, and memory 32 are connected to a bus 34. Furthermore, the database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0529] Robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a controlled object 443. The computer 36 includes a processor 46, RAM 48, and memory 50. The processor 46, RAM 48, and memory 50 are connected to a bus 52. Furthermore, the microphone 238, speaker 240, camera 42, controlled object 443, and communication I / F 44 are also connected to the bus 52.
[0530] Microphone 238 receives instructions from user 20 by receiving sounds emitted by user 20. Microphone 238 captures sounds emitted by user 20 and converts the captured sounds into sound data, which is then output to processor 46. Speaker 240 outputs sound according to instructions from processor 46.
[0531] Camera 42 is a small digital camera equipped with an optical system such as a lens, aperture and shutter, and imaging elements such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, to photograph the area around robot 414 (e.g., the field of view defined by an angle equivalent to the field of vision of an average healthy person).
[0532] Communication I / F44 is connected to network 54. Communication I / F44 and 26 are responsible for the transmission and reception of various information between processor 46 and processor 28 via network 54. The transmission and reception of various information between processor 46 and processor 28 using communication I / F44 and 26 is performed in a secure state.
[0533] The controlled object 443 includes a display device, LEDs (light-emitting diodes) for the eyes, and motors for driving the arms, hands, and feet. The posture or movement of the robot 414 is controlled by controlling the motors in the arms, hands, and feet. Some of the robot 414's emotions can be expressed by controlling these motors. In addition, the robot 414's facial expressions can also be expressed by controlling the illumination state of the LEDs in its eyes.
[0534] Figure 8 This illustrates an example of the main functions of the data processing device 12 and the robot 414. For example... Figure 8 As shown, in the data processing device 12, specific processing is performed by the processor 28. The specific processing program 56 is stored in the memory 32.
[0535] The specific processing program 56 is an example of a "program" involved in the technology of this disclosure. The processor 28 reads the specific processing program 56 from the memory 32 and executes the read specific processing program 56 on the RAM 30. Specific processing is implemented by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.
[0536] The memory 32 stores the data generation model 58 and the emotion-specific model 59. The data generation model 58 and the emotion-specific model 59 are used by the specific processing unit 290.
[0537] In robot 414, the processor 46 performs the acceptance and output processing. The memory 50 stores the acceptance and output program 60. The processor 46 reads the acceptance and output program 60 from the memory 50 and executes the read acceptance and output program 60 on RAM 48. The acceptance and output processing is implemented by the processor 46 acting as the control unit 46A according to the acceptance and output program 60 executed on RAM 48.
[0538] Next, the specific processing of the specific processing unit 290 of the data processing device 12 will be described. Each part of the system described below is implemented by the data processing device 12 and the robot 414. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 will be referred to as the "terminal".
[0539] Example 1 The process is the same as that of the specific process described in Embodiment 1 in the first embodiment above, so the description is omitted.
[0540] Application Example 1 The process is the same as that in the specific processing described in Application Example 1 of the first embodiment above, so the description is omitted.
[0541] Example 2 The process is the same as that of the specific process in Embodiment 2 described in the first embodiment above, so the description is omitted.
[0542] Application Example 2 The process is the same as that in the specific processing described in Application Example 2 of the first embodiment above, so the description is omitted.
[0543] The specific processing unit 290 sends the result of the specific processing to the robot 414. In the robot 414, the control unit 46A outputs the result of the specific processing to the speaker 240 and the controlled object 443. The microphone 238 acquires sound input representing the result of the specific processing. The control unit 46A sends the sound data representing the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the sound data.
[0544] Data generation model 58 is a so-called generative AI (Artificial Intelligence). Examples of data generation models 58 include ChatGPT (registered trademark) (accessible via the internet (URL: https: / / openai.com / blog / chatgpt)). Data generation model 58 is obtained through deep learning on a neural network. Input to data generation model 58 are prompt words containing instructions, and inference data such as sound data representing sound, text data representing text, and image data representing images (e.g., still image data or animation data). Data generation model 58 infers from the input inference data based on the instructions represented by the prompt words and outputs the inference result in one or more data forms, such as sound data, text data, and image data. Data generation model 58 includes, for example, text generation AI, image generation AI, and multimodal generation AI. Here, inference refers to, for example, analysis, classification, prediction, and / or induction. The specific processing unit 290 performs the aforementioned specific processing while using data generation model 58. The data generation model 58 can also be a model finely tuned to output inference results from prompts that do not contain instructions. In this case, the data generation model 58 can output inference results based on prompts that do not contain instructions. The data processing apparatus 12, etc., includes various data generation models 58, including AI other than the generation AI. AI other than the generation AI can be, for example, linear regression, logistic regression, decision trees, random forests, support vector machines (SVM), k-means clustering, convolutional neural networks (CNN), recurrent neural networks (RNN), generative adversarial networks (GAN), or Naive Bayes, and can perform various processes, but is not limited to this example. Furthermore, the AI can also be an AI agent. Furthermore, when the processing of the above-mentioned parts is performed by AI, the processing can be partially or entirely performed by AI, but is not limited to this example. Furthermore, the processing performed by AI including the generation AI can be replaced by processing in the rule base, and the processing in the rule base can also be replaced by processing performed by AI including the generation AI.
[0545] Furthermore, the processing of the aforementioned data processing system 10 is performed by the specific processing unit 290 of the data processing device 12 or the control unit 46A of the robot 414, but it can also be performed by both the specific processing unit 290 of the data processing device 12 and the control unit 46A of the robot 414. Additionally, the specific processing unit 290 of the data processing device 12 acquires or collects information required for processing from the robot 414 or external devices, and the robot 414 acquires or collects information required for processing from the data processing device 12 or external devices.
[0546] For example, the collection unit is implemented by the control unit 46A of the robot 414 or the specific processing unit 290 of the data processing device 12. For example, the acquisition unit uses the camera 42 or communication I / F 44 of the robot 414 to acquire step data, which is then processed by the specific processing unit 290 of the data processing device 12. For example, the analysis unit is implemented by the specific processing unit 290 of the data processing device 12, which analyzes the data from the collection unit and the acquisition unit. For example, the generation unit is implemented by the specific processing unit 290 of the data processing device 12, which uses a generation AI to generate a menu. For example, the serving unit is implemented by the speaker 240 of the robot 414 and the control object 443 or the specific processing unit 290 of the data processing device 12, which provides the generated menu to the user. The correspondence between each unit and the device or control unit is not limited to the above examples and various changes can be made.
[0547] In the above embodiments, examples of specific processing by the data processing device 12 are given, but the technology disclosed herein is not limited to this, and specific processing may also be performed by the robot 414.
[0548] Furthermore, the emotion-specific model 59, acting as an emotion engine, can determine a user's emotion based on a specific mapping. Specifically, the emotion-specific model 59 can determine a user's emotion based on an emotion graph that serves as a specific mapping (see [reference]). Figure 9 The emotion-specific model 59 can also determine the robot's emotion, and the specific processing unit 290 performs specific processing based on the robot's emotions.
[0549] Figure 9 This is a diagram representing an emotion map 400 that maps multiple emotions. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. The closer to the center of the concentric circles, the more primitive the emotion is. On the outer side of the concentric circles, emotions representing states or behaviors arising from mood are arranged. Emotions are concepts that include feelings and mental states. Emotions generated by reactions occurring in the brain are arranged roughly to the left of the concentric circles. Emotions derived from situational judgments are arranged roughly to the right of the concentric circles. Emotions generated by reactions occurring in the brain and derived from situational judgments are arranged roughly above and below the concentric circles. Furthermore, "pleasant" emotions are arranged above the concentric circles, and "unpleasant" emotions are arranged below them. Thus, in the emotion map 400, multiple emotions are mapped based on the structure that generates emotions, and emotions that are likely to occur simultaneously are mapped close to each other.
[0550] These emotions are distributed at the three o'clock position of the emotion map 400, typically fluctuating between peace and anxiety. In the right half of the emotion map 400, situational awareness dominates over internal sensation, thus resulting in an impression of calm.
[0551] The inner side of the emotion map 400 represents the inner state, while the outer side represents behavior. Therefore, the further outward you are from the emotion map 400, the more visible the emotion becomes (manifested in behavior).
[0552] Here, human emotions are based on various balances such as posture and blood sugar levels. When these balances deviate from an ideal state, it indicates an unpleasant state; when they approach the ideal state, it indicates a pleasant state. Emotions in robots, cars, motorcycles, etc., can also be created in the following way: based on various balances such as posture and remaining battery power, when these balances deviate from an ideal state, it indicates an unpleasant state; when they approach the ideal state, it indicates a pleasant state. Emotion maps can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on a Brain Physiological Signal Analysis System for Voice Emotion Recognition and Emotion, Tokushima University, Doctoral Dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). In the left half of the emotion map, emotions belonging to the sensory-dominated region, called "response," are arranged. Furthermore, in the right half of the emotion map, emotions belonging to the situational cognition-dominated region, called "situation," are arranged.
[0553] In the emotion map, two types of emotions that promote learning are defined. One is a negative emotion on the situational side, in the middle or peripheral region of "repentance" or "reflection." This occurs when the robot experiences negative emotions such as "I don't want to experience this feeling again" or "I don't want to be blamed again." The other is a positive emotion on the response side, near the "desire" region. This occurs when there are positive feelings such as "wanting more" or "wanting to know more."
[0554] The emotion-specific model 59 inputs user input into a pre-trained neural network to obtain emotion values representing each emotion shown in the emotion map 400, thereby determining the user's emotion. This neural network is pre-trained based on multiple training data sets, which are combinations of user input and emotion values representing each emotion shown in the emotion map 400. Furthermore, this neural network... Figure 10 As shown in the sentiment map 900, it was trained in a way that sentiments configured close to each other have similar values. Figure 10 The text shows examples of emotions such as "peace of mind", "stability", and "reassurance" that have similar emotion values.
[0555] The above description focuses on the functions of the data processing device 12, but the system of this disclosure is not necessarily installed on a server. The system of this disclosure can also be installed as a general information processing system. This disclosure can also be installed, for example, as a software program running on a personal computer, an application running on a smartphone, etc. The method of this disclosure can also be provided to users in the form of SaaS (Software as a Service).
[0556] In the above embodiments, an example of a specific process being performed by a single computer 22 is given. However, the technology disclosed herein is not limited to this, and the specific process can also be distributed among multiple computers, including computer 22. For example, the data generation model 58 can be located on an external device of the data processing apparatus 12, where data is generated based on the input data.
[0557] In the above embodiments, examples of storing a specific processing program 56 in the memory 32 have been described, but the technology disclosed herein is not limited thereto. For example, the specific processing program 56 may also be stored in a portable computer-readable non-transitory storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-transitory storage medium is installed into the computer 22 of the data processing device 12. The processor 28 executes specific processing according to the specific processing program 56.
[0558] Alternatively, a specific processing program 56 may be pre-stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 according to the requirements of the data processing device 12.
[0559] In addition, it is not necessary to store all the specific processing program 56 in the storage device such as the server connected to the data processing device 12 via the network 54 or in the memory 32; a portion of the specific processing program 56 may be stored in advance.
[0560] As hardware resources for performing specific processes, various processors, as shown below, can be used. For example, a CPU can be listed as a processor, which functions as a general-purpose processor that executes specific processes by executing software, i.e., a program. Furthermore, processors can include, for example, FPGAs (Field-Programmable Gate Arrays), PLDs (Programmable Logic Devices), or ASICs (Application-Specific Integrated Circuits), which are dedicated circuits with circuitry specifically designed to perform specific processes. Each processor has built-in or connected memory, and each processor executes specific processes using that memory.
[0561] The hardware resources for performing a specific process can consist of one of these various processors, or a combination of two or more processors of the same or different types (e.g., a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Alternatively, the hardware resources for performing a specific process can be a single processor.
[0562] As an example of a single processor, there are two approaches: First, a processor is composed of a combination of one or more CPUs and software, which functions as a hardware resource to perform a specific process; second, as represented by a SoC (System-on-a-chip), a processor is used to implement the functionality of the entire system, which includes multiple hardware resources for performing a specific process, using a single IC (Integrated Circuit) chip. In this way, the specific process is implemented by using one or more of the aforementioned processors as hardware resources.
[0563] Furthermore, the hardware architecture of these various processors, more specifically, can utilize circuits that combine semiconductor elements and other circuit components. Moreover, the specific process described above is just one example. Therefore, without departing from the main point, unnecessary steps can certainly be deleted, new steps added, or the processing order changed.
[0564] The descriptions and illustrations above are detailed explanations of a portion of the technology disclosed herein, and are merely one example of the technology disclosed herein. For example, the above descriptions of the structure, function, effect, and results are just one example of the structure, function, effect, and results of a portion of the technology disclosed herein. Therefore, without departing from the spirit of the technology disclosed herein, unnecessary parts may be deleted, new elements added, or replacements may be made to the descriptions and illustrations above. Furthermore, to avoid confusion and facilitate understanding of a portion of the technology disclosed herein, explanations of common technical knowledge that do not require special explanation under the premise of being able to implement the technology disclosed herein have been omitted from the descriptions and illustrations above.
[0565] All documents, patent applications and technical specifications set forth in this specification are incorporated herein by reference to the same extent that each document, patent application and technical specification is specifically and individually described therein and referenced by reference.
[0566] In addition, the following notes are provided in response to the above explanation.
[0567] Example 1 (Note 1) An information processing system, characterized in that it comprises: A device for receiving incoming call signals and operation information from a communication terminal, and controlling the switching of manual response processing performed in the communication terminal to automatic voice response processing performed in an information processing device. A means for generating a greeting voice based on pre-stored fixed response information by performing speech synthesis processing that converts text information into speech information, and sending the greeting voice to the called party's communication device through the communication terminal when the automatic voice response processing begins in the information processing device. A device for acquiring call voice data from the communication terminal, performing voice recognition processing to convert the call voice data into text information, and recording the text information as conversation history in chronological order; A device for performing language parsing and risk feature extraction processing on the spoken text information of the other party contained in the conversation history, and generating prompt information as input information for parsing to generate credibility evaluation information related to the other party's spoken text. A means for inputting the parsing input information into a dialogue engine including a generative artificial intelligence model, thereby obtaining output information containing a credibility evaluation result related to the other party's speech and behavioral candidate information based on the evaluation result; An apparatus for generating, based on the output information, a prompt message containing a proposal for the next action of the user communicating, and sending the prompt message to the communication terminal for display or notification on the communication terminal; A device for associating and storing the conversation history and the evaluation results in a recording area, generating a prompt message for generating summary information based on the conversation history and the evaluation results after the call ends, and inputting the prompt message into the generative artificial intelligence model.
[0568] (Note 2) According to the information processing system described in Appendix 1, the information processing device is used to calculate the overall risk level related to the entire call based on the credibility evaluation result related to the other party's speech, and select behavioral candidate information including at least one of the following: call end, call continuation, or switching back to human response, based on the overall risk level, and generate prompt information containing the behavioral candidate information.
[0569] (Note 3) According to the information processing system described in Appendix 1, the information processing device is used to perform sentiment analysis processing based on spoken voice information or text information contained in the conversation history, and to generate prompt information for dynamically adjusting the behavioral candidate information and prompt content to be prompted to the user based on the credibility evaluation result related to the other party's speech and the analysis result of the sentiment analysis processing.
[0570] Application Example 1 (Note 1) An information processing system, characterized in that it comprises: A unit used to switch the received information transmission from human response to artificial intelligence automatic response; A unit used to initiate an AI-powered automatic response based on a start instruction from the terminal, and to store the dialogue content with the communication object as recorded information; A unit for acquiring voice information and converting the voice information into text information through voice recognition processing, and performing natural language processing analysis on the text information and the recorded information to generate assessment information for evaluating the falsity of evidence. A unit for generating prompt statements that are input to a generative artificial intelligence model based on the evaluation information and the dialogue content, and using the generative artificial intelligence model to generate response information for suspicious information transmission; A unit for generating prompt statements to be input into a generative artificial intelligence model based on the evaluation information and the dialogue content, and using the generative artificial intelligence model to generate proposal information to instruct the user on the next action to be taken; A unit for sending the response information to a terminal so that it is automatically sent to the communication object, and for sending the proposal information to a terminal for the user to view.
[0571] (Note 2) The information processing system according to Appendix 1 is characterized in that the system further includes: A unit for generating control information based on the evaluation information of falsity to change the strength of the response content or the warning level for suspicious information transmission, and dynamically adjusting the prompt statements and / or model parameters input to the generative artificial intelligence model according to the control information.
[0572] (Note 3) The information processing system according to Appendix 1 is characterized in that the system further includes: A unit for analyzing the emotional state of users and communication partners through sentiment analysis processing, and adjusting the content and / or tone of prompt statements input to the generative artificial intelligence model based on the analysis results and the falsity assessment information, thereby regulating the expression of the proposal information and the response information.
[0573] Example 2 (Note 1) An information processing system, characterized in that it comprises: A device for switching received communications from human response to automatic voice response; A device for acquiring acoustic signals from a terminal device or communication device and converting the acoustic signals into character information using audio processing technology and speech recognition technology; A means for storing the character information, along with user identification information and time information associated with the character information, into an information management device; A device for applying natural language processing techniques to the character information to extract semantic information, including time expressions, event expressions, and negation expressions; An apparatus for obtaining objective factual information from an external information providing device based on the semantic information, and determining the consistency between the objective factual information and the character information, so as to generate prompt statements for evaluation information related to perfunctoriness. An apparatus for generating input data for using a generative artificial intelligence model when generating the evaluation information, the apparatus being configured to construct a prompt statement containing the character information, the semantic information, and the objective fact information, and to provide the prompt statement as input data to the generative artificial intelligence model; An apparatus for generating a perfunctory determination result based on the evaluation result output from the generative artificial intelligence model, and sending the determination result as output data to an output device.
[0574] (Note 2) The information processing system according to Appendix 1 is characterized in that, It also includes a device for generating a prompt statement suggesting the next action a user should take based on the perjury determination result and the objective factual information, and for generating a response message containing the suggestion message using the generative artificial intelligence model.
[0575] (Note 3) The information processing system according to Appendix 1 is characterized in that, It also includes a device for applying sentiment analysis technology to the character information to analyze the user's emotional state, and adjusting the content and tone of the prompt statement based on the emotional state and the perfunctory determination result, thereby changing the content or expression of the behavioral suggestions output by the generative artificial intelligence model.
[0576] Application Example 2 (Note 1) An information processing system, characterized in that it comprises: A means of switching received communications from human response to automatic response from an automatic response device; Means for acquiring voice information from a communication terminal or information processing device and converting the voice information into text information using voice recognition technology; This is a means of parsing the text information using natural language processing technology, and extracting target events, evaluation expressions and numerical expressions related to the speech content from the text information; Means for retrieving factual information from external information sources or storage devices based on the extraction results, and generating comparison information indicating the correspondence between the speech content and the factual information; Means for generating a prompt statement containing the speech content, the comparison information, and indication information related to the evaluation of falsity or authenticity, and inputting the prompt statement into a generative artificial intelligence model so that the generative artificial intelligence model evaluates the falsity or authenticity of the speech content; Means for analyzing the evaluation results output by the generative artificial intelligence model and generating evaluation information that includes at least falsehood indicators and their rationale. Means for storing the evaluation information in a recording device in association with the session information, and for prompting the user with the evaluation information through a display device or a notification device.
[0577] (Note 2) The information processing system according to Appendix 1 is characterized in that, Further, it includes: generating a prompt statement containing instructions for the user's next action based on the evaluation information and the session information, inputting the prompt statement into a generative artificial intelligence model to generate candidate actions for the action, and outputting at least a portion of the candidate actions as behavioral suggestion information.
[0578] (Note 3) The information processing system according to Appendix 1 is characterized in that, Further including: means for inputting user-related voice or text information into an emotion analysis device to obtain emotion information representing an emotional state, generating a prompt statement for behavioral suggestions adjusted in tone, content, or urgency based on the emotion information and the evaluation information, and inputting the prompt statement into a generative artificial intelligence model to generate behavioral suggestion information adjusted according to the emotional state.
Claims
1. An information processing system, characterized in that, include: processor; The processor is configured to switch received communications from human responses to AI-powered automatic responses. Acquire speech data and convert the speech data into text data using speech recognition technology; parse the text data using natural language processing technology to generate prompt information for judging the falsity of evidence.
2. The information processing system according to claim 1, characterized in that, The processor is also configured to generate a prompt message based on the result of the false evidence judgment to suggest the next course of action.
3. The information processing system according to claim 1, characterized in that, The processor is also configured to: analyze the user's emotional state using a sentiment analysis engine, and generate prompts based on the analysis results to adjust the proposed action suggestions.
Citation Information
Patent Citations
Persona chatbot control method and system
JP2022180282A