Information processing system

CN122802177APending Publication Date: 2026-09-22SOFTBANK GROUP CORP
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202610319397.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2025-03-19
Filing Date
2026-03-16
Publication Date
2026-09-22

AI Technical Summary

Technical Problem

基于该评估结果,处理器可以初步判断该邮件是否为危险邮件或可疑邮件,从而克服仅依赖固定规则而导致的适应性不足问题

Benefits of technology

[0010]通过以上手段的综合应用,本发明能够实现:基于生成式人工智能模型的内容风险评估、基于数据库比对的识别信息可信性判定以及基于用户情感状态的危险度与优先级动态调整,从而有效提高危险邮件检测的准确性与针对性,增强用户在不同情绪状态下的安全防护效果,较好地解决了现有技术中存在的适应性不足、来源可信判定不精确以及缺乏情感维度考量等问题。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122802177A_ABST
    Figure CN122802177A_ABST
Patent Text Reader

Abstract

The application provides an information processing system. An information processing system comprises a processor configured to: generate prompt information for indicating that a generative artificial intelligence model evaluates the danger of a mail, so as to judge whether the mail is a dangerous mail based on the content of the mail body; compare identification information contained in the mail with a database to judge the source of the identification information, and evaluate the credibility of the identification information according to the comparison result; analyze the expression and voice of a user by using a sentiment recognition technology to identify the emotion of the user, and adjust the danger degree of the mail according to the identified emotion of the user.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The technology disclosed herein relates to an information processing system. Background Technology

[0002] Japanese Patent Application Publication No. 2022-180282 discloses a method for controlling a role-based chatbot executed by at least one processor. The method includes the following steps: receiving a user's speech; adding the user's speech to a prompt word, the prompt word containing instruction statements associated with an explanation of the chatbot's role; encoding the prompt word; and inputting the encoded prompt word into a language model to generate a chatbot response to the user's speech.

[0003] Existing email security detection technologies primarily rely on fixed rules, blacklists / whitelists, or traditional machine learning models to assess email content, sender information, and suspicious links. These technologies have several shortcomings: First, relying solely on rules or static features makes it difficult to respond promptly to constantly evolving phishing techniques and fraudulent messages, resulting in insufficient detection rates or high false positive rates for dangerous emails. Second, when assessing the credibility of identification information (such as accounts, links, and identity verification information) contained in emails, existing systems often only perform format checks or simple source checks, lacking comprehensive utilization of multi-dimensional credibility information in the database, making it difficult to accurately determine the authenticity and reliability of the identified information. Third, existing technologies generally do not consider the user's real-time emotional state when conducting email risk assessments. They cannot adaptively adjust risk warnings and priorities based on different user emotions such as tension, anxiety, or relaxation, making it difficult to promptly increase the level of alerts and reduce security risks when users are vulnerable or inattentive.

[0004] Therefore, the problem to be solved by this invention is to provide a system that can comprehensively utilize generative artificial intelligence models to assess the danger of emails, make credibility judgments on identification information based on databases, and dynamically adjust the danger level and priority of emails in combination with user emotion recognition results, so as to improve the accuracy and flexibility of dangerous email detection, while enhancing the personalized security protection capabilities for users. Summary of the Invention

[0005] To address the aforementioned issues, the present invention provides an information processing system, including a processor, wherein the processor is configured to perform the following processing means.

[0006] First, the processor generates a prompt message to instruct the generative artificial intelligence model to assess the danger of the email. Specifically, when the system receives an email, the processor extracts the email body content and constructs a prompt message containing the email body, contextual information, and risk assessment requirements. This prompt message is then input into the generative artificial intelligence model, requesting the model to perform a comprehensive analysis based on its natural language understanding and reasoning capabilities to determine whether the email content involves phishing, fraud, or malicious manipulation, and outputs an email danger assessment result or risk score. Based on this assessment result, the processor can preliminarily determine whether the email is dangerous or suspicious, thus overcoming the problem of insufficient adaptability caused by relying solely on fixed rules.

[0007] Secondly, the processor compares the identification information contained in the email with a database to determine the source of the identification information and assesses its credibility based on the comparison results. Specifically, the processor extracts identification information from the email body, title, or attachments. This identification information may include account numbers, link addresses, phone numbers, identification numbers, organization identifiers, etc. The processor uses the extracted identification information as query conditions to access a pre-built database and retrieve whether the identification information corresponds to a registered legitimate organization, trusted service, or known malicious source. When the identification information matches a known trusted record, the processor marks the identification information as trustworthy or low-risk. When the identification information matches a known malicious record or shows an abnormal combination, the processor marks the identification information as high-risk. When the identification information is unknown, the processor can assign a medium or uncertain risk level based on the context and model evaluation results. In this way, the system can perform a refined assessment of the source credibility of the identification information itself, beyond content analysis.

[0008] Furthermore, the processor utilizes emotion recognition technology to analyze the user's facial expressions and voice to identify their emotions and adjust the email's risk level accordingly. Specifically, when a user views or prepares to interact with an email, the processor acquires the user's facial expression image and voice signal through an emotion recognition module connected to a camera and microphone. It then invokes the emotion recognition algorithm to determine if the user is currently in a state of tension, anxiety, confusion, relaxation, or alertness. When the processor detects that the user is in a relatively tense, anxious, or inattentive emotional state, it can increase the email's risk level or strengthen the risk warning based on the original email risk assessment. When the user is calm and alert, the processor can maintain or appropriately reduce the warning intensity. Thus, the system can adaptively adjust the risk presentation method and degree based on the user's real-time emotional state, reducing the possibility of users mistakenly believing dangerous emails when in a vulnerable emotional state.

[0009] Furthermore, the processor can utilize natural language processing technology to parse the email body, extracting specific keywords and phrases to assist in the generation of prompts and database queries, thereby further improving the input quality and analysis accuracy of generative artificial intelligence models. The processor can also prioritize emails based on sentiment analysis results. Emails deemed high-risk and where the user's emotions are easily influenced are given high priority and immediately notified to the user. Emails deemed low-risk and where the user's state is stable are given normal priority or delayed notification, thus achieving intelligent management of email processing order and reminder strategies.

[0010] By comprehensively applying the above methods, this invention can achieve: content risk assessment based on generative artificial intelligence models, determination of the credibility of identification information based on database comparison, and dynamic adjustment of danger level and priority based on user emotional state. This effectively improves the accuracy and targeting of dangerous email detection, enhances the security protection effect for users in different emotional states, and better solves the problems of insufficient adaptability, inaccurate determination of source credibility, and lack of consideration of emotional dimension in the existing technology.

[0011] "System" refers to an overall device or platform consisting of hardware and / or software used to perform functions such as email risk assessment, identification of information credibility judgment, and adjustment of email risk level and priority based on user sentiment. It can be deployed on servers, terminal devices, or a combination of both.

[0012] "Processor" refers to hardware or its virtualization unit capable of executing instruction code and performing calculations, control, and logical processing on data, such as a central processing unit (CPU), a graphics processing unit (GPU), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), a neural network accelerator chip, or a computing and control module composed of any one or more of the above.

[0013] "Generative artificial intelligence model" refers to an artificial intelligence model that is trained on a large-scale dataset and has the ability to generate text, code, images or other content based on input prompts. In this invention, it is mainly used as a model to understand and reason about email content based on prompt information and output email risk assessment results, such as a large language model or a multimodal generative model.

[0014] "Prompt information" refers to structured or unstructured text instructions generated by the processor and input into the generative artificial intelligence model. These instructions describe the content, context, and expected task (such as risk analysis) of the email to be evaluated, thereby guiding the generative artificial intelligence model to output the corresponding evaluation or judgment results.

[0015] "Email risk" refers to the degree of risk associated with a particular email that may contain phishing, fraud, malicious links, malicious attachments, privacy theft, or other security threats. It can be expressed using qualitative descriptions or quantitative scoring methods.

[0016] The "email body" refers to the main text content of an email excluding header information such as subject, recipient, and sender. It includes plain text, rich text, and embedded links and tags, and is used to convey specific information to the recipient.

[0017] "Identification information" refers to information that can uniquely or partially identify an entity, account, resource, or communication object, including but not limited to account, URL link address, telephone number, identification number, organization identifier, user ID, bank card number, or their anonymized form.

[0018] A "database" is a logical collection used to store and manage data records. It can be a relational database, a non-relational database, or a distributed data storage system. It is used to store source information, credibility tags, historical behavior records, and blacklist / whitelist data related to identification information.

[0019] "Reliability" refers to the comprehensive assessment of the authenticity and security of identification information or its associated entities, and is usually determined based on a variety of factors such as source, historical records, blacklists and whitelists, and third-party reputation data.

[0020] "Emotion recognition technology" refers to the technology that uses computer vision, speech signal processing and / or physiological signal analysis to process a user's facial expressions, voice features, posture or other behavioral data to identify the user's current emotional state (such as tension, anxiety, anger, happiness, calmness, etc.).

[0021] "User facial expressions" refer to the visual characteristics displayed by a user through changes in facial muscles while viewing emails or interacting with the system, including changes in eyebrows, eyes, mouth, and overall facial posture, used to reflect the user's emotional state.

[0022] "User's voice" refers to the sound signals emitted by a user when reading, replying to, or discussing emails, including the voice content itself and its acoustic characteristics, such as pitch, volume, speaking speed, tone, and pauses.

[0023] "User's emotions" refers to the psychological or emotional state inferred from user facial expressions, voice, and other behavioral data, including but not limited to tension, anxiety, confusion, anger, joy, calmness, composure, and alertness.

[0024] "Email risk level" refers to the risk level determined for a specific email after considering factors such as the evaluation results of a generative artificial intelligence model, the credibility of the identified information, and the user's emotional state. It can be expressed as a level, probability value, or numerical score, and is used to indicate the intensity of the potential threat that emails pose to user security.

[0025] Natural Language Processing (NLP) technology refers to a class of technologies that enable computers to understand, analyze, generate, and transform natural language text, including methods such as word segmentation, part-of-speech tagging, syntactic analysis, semantic understanding, entity recognition, keyword extraction, and text classification.

[0026] "Specific keywords and phrases" refer to words or phrases that are predefined in this invention or learned through data analysis and are highly relevant to email security, such as text fragments that are directly or indirectly related to passwords, payments, account verification, remittances, prize collection, emergency notifications, etc.

[0027] "Sentiment analysis results" refer to the structured output information obtained after analyzing user facial expressions and voice data through sentiment recognition technology. This includes the user's emotion category, emotion intensity, and emotion change trend, and is used to guide the adjustment of email risk level and priority.

[0028] "Email priority" refers to the processing order and alert level assigned to different emails when the system processes and displays them. It is used to determine the order of emails in the list, the notification method, and the alert intensity, such as high priority, normal priority, or low priority.

[0029] "Notification messages to users" refers to messages presented to users on the terminal interface in the form of text, icons, pop-ups, sounds, or a combination thereof, to remind users of the risk status, importance, or suggested actions of a particular email. Attached Figure Description

[0030] Figure 1 This is a conceptual diagram illustrating an example of the configuration of the data processing system according to the first embodiment.

[0031] Figure 2 This is a conceptual diagram illustrating an example of the main functions of the data processing apparatus and smart device according to the first embodiment.

[0032] Figure 3 This is a conceptual diagram illustrating an example of the configuration of the data processing system according to the second embodiment.

[0033] Figure 4 This is a conceptual diagram illustrating an example of the main functions of the data processing device and smart glasses according to the second embodiment.

[0034] Figure 5This is a conceptual diagram illustrating an example of the configuration of the data processing system according to the third embodiment.

[0035] Figure 6 This is a conceptual diagram illustrating an example of the main functions of the data processing device and head-mounted terminal according to the third embodiment.

[0036] Figure 7 This is a conceptual diagram illustrating an example of the configuration of the data processing system according to the fourth embodiment.

[0037] Figure 8 This is a conceptual diagram illustrating an example of the main functions of the data processing device and robot according to the fourth embodiment.

[0038] Figure 9 This represents an emotion map that maps multiple emotions.

[0039] Figure 10 This represents an emotion map that maps multiple emotions.

[0040] Figure 11 This is a sequence diagram illustrating the processing flow of the data processing system of the first embodiment.

[0041] Figure 12 This is a sequence diagram illustrating the processing flow of the data processing system in Application Example 1.

[0042] Figure 13 This is a sequence diagram illustrating the processing flow of the data processing system of the second embodiment.

[0043] Figure 14 This is a sequence diagram illustrating the processing flow of the data processing system in Application Example 2. Detailed Implementation

[0044] Hereinafter, an example of an implementation of the system according to the present disclosure will be described with reference to the accompanying drawings.

[0045] First, let me explain the terminology used in the following instructions.

[0046] In the following embodiments, the processor (hereinafter referred to as "processor") with reference numerals may be a single computing device or a combination of multiple computing devices. Furthermore, the processor may be a single computing device or a combination of multiple computing devices. Examples of computing devices include CPU (Central Processing Unit), GPU (Graphics Processing Unit), GPGPU (General-Purpose computing on Graphics Processing Units), APU (Accelerated Processing Unit), etc.

[0047] In the following embodiments, RAM (Random Access Memory), as indicated in the figures, is a memory that temporarily stores information and is used as working memory by the processor.

[0048] In the following embodiments, the memory, as indicated by the reference numerals, is one or more non-volatile storage devices that store various programs and parameters. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), disks (e.g., hard disks), or magnetic tapes.

[0049] In the following embodiments, the communication I / F (Interface) with reference numerals is an interface that includes a communication processor and an antenna, etc. The communication I / F is responsible for communication between multiple computers. As an example of a communication specification applicable to the communication I / F, wireless communication specifications such as 5G (5th Generation Mobile Communication System), Wi-Fi (wireless fidelity) (registered trademark), or Bluetooth (registered trademark) can be listed.

[0050] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B". That is, "A and / or B" means that it can be only A, only B, or a combination of A and B. Furthermore, in this specification, when "and / or" connects to express more than three items, the same interpretation as "A and / or B" applies.

[0051] First Implementation Method Figure 1 An example of the configuration of the data processing system 10 according to the first embodiment is shown.

[0052] like Figure 1 As shown, the data processing system 10 includes a data processing device 12 and an intelligent device 14. A server can be cited as an example of the data processing device 12.

[0053] The data processing apparatus 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" as understood in this disclosure. The computer 22 includes a processor 28, RAM 30, and memory 32. The processor 28, RAM 30, and memory 32 are connected to a bus 34. Furthermore, the database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0054] The smart device 14 includes a computer 36, a receiving device 38, an output device 40, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, RAM 48, and memory 50. The processor 46, RAM 48, and memory 50 are connected to a bus 52. In addition, the receiving device 38, output device 40, camera 42, and communication I / F 44 are also connected to the bus 52.

[0055] The receiving device 38 includes a touchscreen 38A and a microphone 38B, and receives user input. The touchscreen 38A receives user input via touch by detecting contact with an indicator (e.g., a pen or finger). The microphone 38B receives user input via sound by detecting the user's voice. The control unit 46A in the processor 46 sends data representing the user input received by the touchscreen 38A and microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data representing the user input.

[0056] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting data in a form perceptible to the user 20 (e.g., sound and / or text). The display 40A displays visual information such as text and images according to instructions from the processor 46. The speaker 40B outputs sound according to instructions from the processor 46. The camera 42 is a small digital camera equipped with an optical system such as a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.

[0057] Communication I / F44 is connected to network 54. Communication I / F44 and 26 are responsible for sending and receiving various information between processor 46 and processor 28 via network 54.

[0058] Figure 2 The diagram shows an example of the main functions of the data processing device 12 and the smart device 14.

[0059] like Figure 2 As shown, in the data processing apparatus 12, specific processing is performed by the processor 28. A specific processing program 56 is stored in the memory 32. The specific processing program 56 is an example of a "program" as understood in this disclosure. The processor 28 reads the specific processing program 56 from the memory 32 and executes the read specific processing program 56 on the RAM 30. Specific processing is implemented by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.

[0060] The memory 32 stores a data generation model 58 and an emotion-specific model 59. The data generation model 58 and the emotion-specific model 59 are used by the specific processing unit 290. The specific processing unit 290 can use the emotion-specific model 59 to infer the user's emotions and perform specific processing based on the user's emotions. In the emotion inference function (emotion-specific function) using the emotion-specific model 59, various inferences and predictions related to the user's emotions are performed, including inferences and predictions of the user's emotions, but this is not limited to this example. Furthermore, emotion inference and prediction may also include, for example, emotion analysis (parsing).

[0061] In the smart device 14, the processor 46 performs the acceptance output processing. The memory 50 stores the acceptance output program 60. The acceptance output program 60 is used in conjunction with the data processing system 10 and the specific processing program 56. The processor 46 reads the acceptance output program 60 from the memory 50 and executes the read acceptance output program 60 on the RAM 48. The acceptance output processing is implemented by the processor 46 acting as the control unit 46A according to the acceptance output program 60 executed on the RAM 48. Furthermore, the smart device 14 has the same data generation model and emotion-specific model as the data generation model 58 and the emotion-specific model 59, and these models can also be used to perform the same processing as the specific processing unit 290. The acceptance output processing is implemented by the processor 46 acting as the control unit 46A according to the acceptance output program 60 executed on the RAM 48.

[0062] Alternatively, other devices besides the data processing device 12 may also have the data generation model 58. For example, a server device (e.g., a generation server) may have the data generation model 58. In this case, the data processing device 12 obtains the processing results (prediction results, etc.) using the data generation model 58 by communicating with the server device that has the data generation model 58. Furthermore, the data processing device 12 may be a server device or a user-held terminal device (e.g., a mobile phone, robot, home appliance, etc.). Next, an example of the processing of the data processing system 10 of the first embodiment will be described.

[0063] Example 1 The flow of a specific process in Example 1 will be described. Each part of the system described below is implemented by the data processing device 12 and the smart device 14. Furthermore, the data processing device 12 is referred to as the "server," and the smart device 14 is referred to as the "terminal."

[0064] In modern communication environments, users receive a large amount of information via email and other electronic communication methods, including potentially harmful malicious communications such as phishing and fraud. Current technologies typically rely on fixed keyword rules, simple blacklists and whitelists, or single machine learning models to determine the danger of emails. These approaches suffer from the following problems: First, the rules and feature sets are outdated, making it difficult to respond promptly to evolving attack methods, resulting in insufficient detection accuracy. Second, systems often only perform static analysis of email text, failing to incorporate multi-dimensional contextual information such as the credibility of the source of the identified information, making it difficult to distinguish between trusted service notifications and fraudulent attacks. Third, traditional systems generally ignore users' emotional states and interactive behavior characteristics, failing to adjust prompt strategies and alarm intensity based on dynamic factors such as the user's current level of tension or confusion, easily leading to users ignoring important warnings or being overwhelmed by excessive false alarms. Fourth, the application of existing generative artificial intelligence models in security scenarios is usually limited to simple classification or content summarization, lacking a systematic architecture that finely controls model tasks, output formats, and explanatory content through prompt statements, making it difficult to directly embed model outputs into security alarm processes, increasing system integration complexity.

[0065] Therefore, in the field of computer technology, there is an urgent need for a comprehensive processing mechanism that can integrate natural language processing, credibility assessment of identification information, generative artificial intelligence model reasoning based on prompt statements, and user emotion perception on the server side. This mechanism would enable the server to conduct more accurate and interpretable risk assessment of document information based on multi-dimensional features, and automatically generate personalized notification control strategies for terminal devices based on the assessment results, thereby substantially improving the ability to detect malicious communication and the effect of human-computer interaction.

[0066] The specific processing performed by the specific processing unit 290 of the data processing apparatus 12 in Embodiment 1 is achieved by the following means.

[0067] In this invention, the server includes: a processing unit for receiving user communication records containing document information and extracting text information as the parsing object; a processing unit for segmenting and parsing the text information using language processing technology, extracting multiple word information and recognition information from the text information, and preprocessing the text information to generate input data for a generative artificial intelligence model; a processing unit for generating and sending query information containing the text information and prompt statements related to risk assessment for the generative artificial intelligence model, thereby obtaining risk assessment information about the document information; a processing unit for comparing the recognition information with an information set in a storage device and calculating confidence information about the document information based on the comparison result; a processing unit for fusing the obtained assessment information, the calculated confidence information, and sentiment information obtained based on user state information to calculate the risk information of the document information; and a processing unit for generating corresponding notification content based on the risk information, generating notification control prompt statements for sending the notification content to a terminal device, and sending the notification content to the terminal device. This allows for high-precision risk assessment and dynamic notification control of malicious communications on the server side through multi-source feature fusion and generative AI reasoning driven by prompt statements, reducing false alarms and false negatives, and improving the processing performance and user interaction experience of computer systems in the field of secure communication.

[0068] "Communication processing device" refers to a computer device or computer system used to send, receive, forward and process electronic communication data in a network environment, and may include servers, gateways, routing devices or combinations thereof.

[0069] The "information processing unit" refers to the processing unit in a communication processing device that executes program instructions to receive, parse, calculate, and control the output of data. It can be implemented by one or more processors and the software modules running on them.

[0070] "User communication records" refer to the collection of communication data sent or received by users and recorded by the system during electronic communication, including email data, message data and their related header information and metadata.

[0071] "Document information" refers to readable or parsable text-based information contained in a user's communication records, including email body, message body, or other content expressed in text form.

[0072] "Text information" refers to the text content extracted from the document information that serves as the primary object of parsing, excluding protocol headers, control information, or additional data unrelated to the content.

[0073] "Language processing technology" refers to computer technology used to analyze and understand natural language text, including processing methods such as text segmentation, word segmentation, part-of-speech tagging, syntactic analysis, and semantic analysis.

[0074] "Word information" refers to the basic language units obtained from the text information through language processing technology, including words, phrases or sub-words, which are used for subsequent feature extraction and model input.

[0075] "Identification information" refers to structured data in the text that can indicate a specific entity or resource, including network addresses, telephone numbers, account identifiers, email addresses, etc.

[0076] "Preprocessing" refers to the process of converting, marking, encoding, filtering, or normalizing the text information before it is input into the generative artificial intelligence model, in order to generate structured data that meets the model's input requirements.

[0077] "Generative artificial intelligence models" refer to artificial intelligence models trained through machine learning or deep learning methods that can generate text or other output results based on input data, including autoregressive language models, encoder-decoder models, etc.

[0078] "Prompt statements" refer to natural language text provided to generative artificial intelligence models to specify the task objectives, constraints, and output formats, and to guide the models to perform specific reasoning or generative operations.

[0079] "Query information" refers to the set of request data constructed by the information processing department and sent to the generative artificial intelligence model, which includes at least the main text information and prompt statements, and is used to obtain evaluation results about document information.

[0080] "Risk assessment information" refers to the evaluation results made by generative artificial intelligence models on whether document information poses risks or malicious behavior, including risk category, risk level, and reasoning.

[0081] "Storage device" refers to a physical or logical storage unit used to store information sets, programs and intermediate data, including semiconductor memory, magnetic storage medium, optical storage medium or a combination thereof.

[0082] An "information set" refers to a data set stored in a storage device used to compare and judge the credibility of information, including lists of known trustworthy sources, blacklists, and reputation score records.

[0083] "Confidence information" refers to a quantitative indicator or level of information that represents the credibility of document information, calculated based on the comparison results between the identification information and the information set.

[0084] "User status information" refers to data related to a user's current physiological or behavioral state, including facial expression data, voice data, operational behavior data, or their derived feature information.

[0085] "Emotional information" refers to the user's emotional state inferred from user status information through emotion recognition or emotion analysis technology, including emotion categories or intensity values ​​such as tension, anxiety, confusion, and calmness.

[0086] "Hazard information" refers to an indicator used to represent the overall risk level of document information, obtained by integrating hazard assessment information, confidence information, and sentiment information. It includes hazard level, hazard score, or classification results.

[0087] "Notification content" refers to text, graphics, or multimedia information generated based on risk information to inform users about the security status of document information, including warnings, risk descriptions, and operational suggestions.

[0088] "Notification control prompts" refer to prompts used to guide generative artificial intelligence models or other processing logic to generate notification content that is suitable for the display method and order of the terminal device, in order to control the notification strategy.

[0089] "Terminal device" refers to a computing device that enables users to receive and view notifications and operate document information, including computer terminals, mobile terminals, wearable terminals or similar devices.

[0090] "Feature information" refers to the structured representation extracted from textual information and related data through language processing techniques and other processing steps, which is used as input to generative artificial intelligence models. It includes vector representations, label sequences, or other feature sets.

[0091] "Priority information" refers to the level or weight determined based on emotional and risk information, used to indicate the order and importance of document information in a notification or processing flow.

[0092] The embodiments of this invention will be described in detail, in conjunction with the appendix, regarding the system architecture, data structure, algorithm flow, and hardware and software configuration. In the following description, the subject is limited to "server," "terminal," or "user."

[0093] I. Overall System Composition In this invention, a server is configured as the core processing node. The server includes: at least one general-purpose processor (CPU), an optional graphics processing unit (GPU) or dedicated accelerator (e.g., a general-purpose parallel computing device), main memory (random access memory), a mass storage device (disk drive or solid-state storage device), and a network interface device. The server runs applications and middleware on an operating system (e.g., a general-purpose server operating system).

[0094] The server's software architecture includes at least the following functional modules: a communication receiving module, a text extraction module, a natural language processing module, a feature construction module, a generative artificial intelligence inference module interface, an identification information comparison module, a confidence calculation module, a sentiment analysis module, a risk fusion module, a notification generation module, and a log and model management module. These modules run on the server's processor as processes or threads, and interact with each other through shared data structures in memory or by calling interfaces.

[0095] In this invention, the terminal is configured as the user interaction and display interface. The terminal can be a desktop computing device, a mobile terminal, or a wearable terminal. The terminal runs an email client or a dedicated security client to receive notifications from the server and present warning messages and operation options to the user.

[0096] Users can view emails and notifications through their terminals and provide feedback to the server regarding email security through a graphical user interface, thus providing labeled data for subsequent model training and parameter updates.

[0097] II. Program and Data Structure Configuration In implementing this invention, the server uses a general-purpose programming language environment (such as a Python runtime environment) to implement the interfaces for the natural language processing module, feature construction module, and generative artificial intelligence inference module. For natural language processing, the server can use open-source language processing libraries (such as general-purpose word segmentation and syntactic analysis libraries), and for deep learning, it can use deep learning frameworks (such as tensor-based frameworks or other deep learning frameworks).

[0098] The server defines various data structures in memory to support technical processing: In the communication receiving module, the server represents the raw email data as a structured record, which includes a header field set (an array of key-value pairs), the raw string of the body, the attachment list, etc.

[0099] In the text extraction module, the server stores the text information as a string field and adds format type tags (such as plain text or hypertext markup language). When needed, the server removes the markup language tags using a hypertext parsing library, retaining only the visible text, and stores it in the "Text Body" field.

[0100] In the natural language processing module, the server segments the main text into a list of sentences, with each sentence element stored in an array. The server then performs word segmentation on each sentence, generating a word list and attaching metadata such as part-of-speech tags and position indexes to each word, forming a word information structure.

[0101] When extracting identification information, the server uses regular expressions and named entity recognition algorithms to extract identification information from word sequences, such as network addresses, phone numbers, account identifiers, etc., and creates a list of identification information. Each element contains a type field, a raw string field, a position field in the text, etc.

[0102] In the feature construction module, the server transforms sentences, words, and recognition information into feature information required for model input, including tokenized tag sequences, word embedding indexes corresponding to the tags, positional encoding, and recognition information tags (e.g., attaching special labels to tags containing network addresses), ultimately forming multidimensional arrays or tensors for use by generative artificial intelligence models.

[0103] In the identification information comparison module, the server compares the identification information with a pre-built "information set" in the storage device. The "information set" can exist in the form of a key-value storage structure or a relational data table, where each record contains fields such as entity identifier, reputation score, and historical risk marker. The server calculates confidence information based on the matching results, for example, assigning a higher confidence score to network addresses that match high-reputation sources and a lower confidence score to identification information that is listed in the blacklist.

[0104] III. Generative Artificial Intelligence Model Structure and Learning Methods The server deploys a generative AI model in the generative AI inference module. The server can employ a neural network model based on a multi-layer self-attention network structure, which may include word embedding layers, multi-layer encoder layers, optional decoder layers, and an output layer. The server uses multi-head attention mechanisms, feedforward network layers, and layer normalization modules in the model to improve its ability to model dependencies and contextual semantics in long texts.

[0105] During the training phase, the server uses a large amount of historical email data and its annotations (such as "safe," "dangerous," and danger level scores) as training samples. For each email, the server constructs an input sequence, including body text tags, feature tags (such as identification information location tags), and task description tags (such as danger assessment task tags). During the forward propagation phase, the server calculates the output probability distribution or danger score based on the network structure. During the backpropagation phase, the server calculates the gradient based on a pre-defined loss function (such as cross-entropy loss or mean squared error loss) and updates the model weights using optimization algorithms (such as stochastic gradient descent, momentum optimizers, or adaptive learning rate optimizers).

[0106] During training, the server employs data augmentation techniques, such as synonym replacement for similar email types, minor sentence order adjustments, or the addition of slight noise, to improve the model's robustness to diverse expressions. The server can utilize batch training and regularization techniques (such as dropout layers or weight decay) to prevent overfitting.

[0107] During the inference phase, according to the features of this invention, the server does not directly input the email text into the model. Instead, it first generates a prompt statement that explicitly specifies the task content, and then concatenates the prompt statement with the email body as the model input. This prompt statement-driven inference method enables the server to flexibly perform multiple tasks such as risk assessment, risk justification explanation, and suspicious feature enumeration on the same model, thereby reducing the computational resource overhead of maintaining multiple models separately for different tasks.

[0108] IV. Construction and Technical Function of Prompt Statements When generating query information, the server constructs prompts for interaction with the generative AI model. Through program logic, the server generates different natural language prompts based on the current task type and target output. The design of these prompts enhances the model's understanding of the task and the consistency of the output format, thereby technically improving the controllability and parsability of the model's output.

[0109] The server can use the following example prompt statement: The server constructs prompt statements based on a detailed analysis of the scenario: "As a cybersecurity analyst, please analyze the content of the following email to determine if it exhibits characteristics of phishing, fraud, or other dangerous communication. List all suspicious points (such as requests for money transfers, password requests, suspicious links, and impersonation of organizations). Finally, give it a risk score from 0 to 100. The email content is as follows: '[Email Body]'."

[0110] The server constructs prompt statements in a fast binary classification scenario: "Based on the email content below, determine whether it is dangerous. Please only output one of the two words 'safe' or 'dangerous', and provide one Chinese sentence explaining your reasoning. Email content: '[Email Body]'."

[0111] The server constructs the prompt statement when generating a user-readable explanation: "You are now an email security assistant. Please explain to ordinary users why the following email might be dangerous and provide three specific preventative suggestions. Email content: '[Email body]'."

[0112] When constructing these prompts, the server inserts the email body text into a pre-defined template. This template-based prompt generation mechanism is implemented on the server side with minimal overhead and reduces the need to modify the internal structure of the model through text-level task descriptions.

[0113] The output obtained by the server through this prompt-driven mechanism is more structured and easier to interpret, enabling the server to extract the danger level, danger cause and recommended measures from the model output more efficiently, thereby reducing complex post-processing rules and improving overall computational efficiency.

[0114] V. The Logic of Integrating Emotional Information and Risk Level In the sentiment analysis module, the server calculates sentiment information based on user state information (such as facial expression features, voice features, or operational behavior features uploaded from the terminal). The server can use a specialized sentiment analysis neural network model to map the input multimodal features into emotion categories and intensity values. For example, the server inputs the user's facial feature vector from the terminal into a convolutional neural network and maps its output to emotion categories such as "nervous," "anxious," and "calm," along with their corresponding probabilities.

[0115] In its risk fusion module, the server employs a weighted fusion or non-linear fusion strategy to combine risk assessment information (such as scores from generative AI models), confidence information (such as reputation scores obtained through comparison with identification information), and sentiment information. The server can set a set of parameters to adjust the risk calculation method under different emotional states. For example, when a user's emotion is indicated as tension, the server can lower the alarm trigger threshold, causing emails with the same risk score to trigger higher-level alarms earlier, thus reducing the probability of accidental actions by the user under pressure.

[0116] Through this fusion logic, the server integrates multi-source data at the algorithm level, so that the risk information not only reflects the objective risk of the email content, but also takes into account the credibility of the information source and the user's current status, thereby achieving more refined risk control in terms of technology.

[0117] VI. Notification Content Generation and Terminal Control In the notification generation module, the server constructs notification content suitable for terminal display based on the risk level information. The server-generated notification content includes the alarm level, a brief description of the cause, a list of suspicious characteristics, and recommended actions. In some scenarios, the server also generates notification control prompts to control the presentation format and order of notifications, guiding the subsequent display logic on the terminal.

[0118] After receiving a notification from the server, the terminal adjusts the display method in its local application based on the level of danger. For example, in a high-risk situation, the terminal highlights the email in a prominent color in the list and displays a full-screen warning when the user attempts to open it; in a medium-risk situation, the terminal only displays a warning icon near the email title. The terminal can also adjust the display order of multiple suspicious emails based on priority information provided by the server, ensuring that high-risk emails reach the user's attention earlier.

[0119] After receiving the notification, the terminal records the user's actions (such as "delete", "mark as safe", "continue to open") as behavioral data, and can upload some of the data to the server for further analysis of the user's reaction patterns and false alarms, supporting the optimization of subsequent models and rules.

[0120] VII. Technical Effects and Improvements in Computer Technology The server, through the aforementioned specific data structures, feature construction methods, generative artificial intelligence model architecture, prompt statement design, and multi-source information fusion algorithms, achieves multi-dimensional and high-precision dangerous communication detection. Compared with traditional solutions based solely on keywords or single classification models, the server achieves the following technical improvements: By comparing identified information with the information set, the server filters and scores suspicious links and contact information in advance at the data level, reducing the computational burden on generative artificial intelligence models on irrelevant text, thereby reducing inference time and computational resource consumption and improving computational efficiency.

[0121] The server uses prompt-driven generative AI model calls to perform different tasks such as risk scoring, suspicious feature enumeration, and user interpretation generation on the same model. This reduces the overhead of deploying and maintaining independent models, lowers storage resources and model loading time, and simplifies the data preprocessing process through a unified input format.

[0122] By incorporating emotional information into risk calculation and notification control, the server not only improves the matching accuracy of the final alarm to user behavior, but also technically enables dynamic threshold adjustment based on user status. This helps reduce false alarms and missed alarms, thereby reducing unnecessary network communication and interface refreshes, and improving the overall resource utilization of the system.

[0123] During the model training phase, the server employs data augmentation and rigorous loss function and optimization algorithm settings to enhance the model's generalization ability to diverse email texts. In actual operation, the server is thus able to achieve higher detection accuracy and robustness without significantly increasing computational load.

[0124] Through a clear data flow design, the server defines the data structure at each stage, from the original email to feature information, model output, and fusion results. This makes the system easy to expand and replace modules, and facilitates distributed deployment and load balancing across multiple physical servers or virtual instances, thereby improving the system's processing power and scalability.

[0125] VIII. Optional Implementation Forms and Variations Servers can choose different generative AI model structures depending on their implementation. For example, a server can use an encoder-decoder structure, where the encoder focuses on modeling the semantics of the email content, and the decoder generates a structured risk report based on prompts. Alternatively, a server can use a decoder-only autoregressive language model to achieve flexible inference with a unified sequence input. Servers can also build lightweight models using distillation techniques for deployment on edge servers, reducing latency.

[0126] The server can employ a multi-level caching structure for identification information comparison, maintaining a cache of frequently used identification information in memory and a complete set of information in a large-capacity storage device. During comparison, the server prioritizes querying the memory cache, and only accesses the large-capacity storage if a match is missed, thus reducing disk accesses and comparison time.

[0127] In acquiring emotional information, the server can choose different data sources based on the actual deployment environment. In scenarios with high privacy requirements, the server only uses abstract behavioral features uploaded by the terminal (such as click frequency and reading time) to estimate the user's level of tension; in scenarios that allow the collection of multimodal information, the server can process facial expression image features and voice features simultaneously to obtain more accurate emotional information.

[0128] In different embodiments, the terminal can select different display strategies based on device performance. When computing resources are limited, the terminal only prompts the user with simple text based on the danger level given by the server; when resources are sufficient, the terminal can display alarm information in a hierarchical manner using various presentation formats such as graphics, colors, and animations, based on the notification control prompts generated by the server.

[0129] Through the above-described embodiments and their variations, the system of the present invention not only achieves high-precision judgment and dynamic alarm control of malicious communication, but also improves computational efficiency, resource utilization and scalability at the levels of algorithm structure, data structure and system architecture, thereby constituting a substantial improvement to computer technology itself.

[0130] use Figure 11 The processing flow is explained.

[0131] Step 1: The server receives and parses the communication records.

[0132] The server's input consists of raw email data or other document-type communication records forwarded from a communication network. This raw data includes header fields, body content, attachment data, and transmission metadata. The server performs protocol and format parsing operations on the input data, converting the raw bitstream into a structured record. Specifically, the server parses header fields such as sender address, recipient address, subject, and timestamp, identifying the communication type requiring risk analysis. After parsing, the server separates the raw body string and attachment list from the raw record, using the raw body string as the basis for subsequent text processing. The server's output is an intermediate data object containing structured header information and the raw body string.

[0133] Step 2: The server extracts and standardizes the main text information.

[0134] The server's input is the structured record output from step 1. Based on the text format type (e.g., plain text or Hypertext Markup Language), the server selects the appropriate parsing method, calling a hypertext parsing library to remove tags, styles, and script code from the Hypertext Markup Language content, retaining only the visible text content. The server performs character set conversion and encoding normalization operations on the text, unifying different encoding formats into an internal unified encoding format, and filtering control characters and invisible characters. At this stage, the server also removes signature blocks that are clearly irrelevant to the content or automatically appended advertising text (achieved through fixed pattern matching) to reduce noise interference in subsequent analysis. The server's output is the cleaned and normalized text string.

[0135] Step 3: The server performs natural language preprocessing and sentence segmentation on the main text.

[0136] The server's input is the normalized text output from step 2. The server calls a natural language processing library to perform sentence segmentation on the text, dividing the long text into a list of sentences based on punctuation, language rules, and length limits. The server performs basic text cleaning on each sentence, including case neutralization, removal of extra spaces, and duplicate punctuation. Based on the segmentation results, the server constructs a sentence array data structure, where each element contains the sentence content and its position in the original text. The server's output is a list of sentences with sentence boundary annotations.

[0137] Step 4: The server performs word segmentation and word information annotation.

[0138] The server's input is the list of sentences output in step 3. For each sentence, the server calls a word segmentation algorithm to split it into a sequence of words or sub-word tags, and performs part-of-speech tagging and basic syntactic role tagging on each word. During word segmentation, the server utilizes a custom dictionary to enhance its ability to identify industry terminology, financial terminology, and common fraudulent terms. The server encapsulates each word, its part of speech, and its index position in the sentence into a word information structure, and organizes it into a two-dimensional structure (sentence dimension and word dimension). The server's output is a set of word information containing word information and preliminary syntactic labels.

[0139] Step 5: The server extracts the identification information and builds a list of identification information.

[0140] The server's input is the set of word information output from step 4. The server uses regular expression matching rules and a named entity recognition model to identify network addresses, phone numbers, account identifiers, email addresses, and other identification information within the word sequence. When detecting identified information, the server records its string content, type label, start and end positions in the text, and the index of the sentence it belongs to. The server merges and removes duplicate or overlapping identified information to avoid redundant records. The server organizes all identified information into a list structure for subsequent comparison with the information set. The server's output is a list of identified information and annotation information associated with text positions.

[0141] Step 6: The server builds feature information for generative artificial intelligence models.

[0142] The server's input consists of the sentence list output from step 3, the word information set output from step 4, and the recognition information list output from step 5. Based on the vocabulary of the selected generative AI model, the server maps each word to a vocabulary index or sub-word tag index and constructs a tag sequence. The server generates a positional code for each tag and adds special labels to tags containing recognition information, such as adding a specific type tag to tags containing network addresses. The server encodes sentence boundaries, paragraph structure, and other information into separator tags so that the model can distinguish contextual structures. Finally, the server synthesizes all the above information into a multi-dimensional tensor structure, including a tag index tensor, a type tag tensor, and a positional code tensor, which are used as input features for the generative AI model. The server's output is a set of feature information that meets the model's input requirements.

[0143] Step 7: The server generates risk assessment-related prompts and constructs query text.

[0144] The server's input consists of the normalized text output from step 2 and the current task configuration parameters. Based on the task type (e.g., detailed analysis, rapid binary classification, generating user explanations), the server selects a preset prompt template and embeds the text text into the template. For example, the server generates the following prompt: "As a cybersecurity analyst, please analyze the content of the following email to determine if it exhibits characteristics of phishing, fraud, or other dangerous communication. List all suspicious points (such as requests for money transfers, password requests, suspicious links, and impersonation of organizations). Finally, give it a risk score from 0 to 100. The email content is as follows: '[Email Body]'."

[0145] Or generate a short prompt message: "Based on the email content below, determine whether it is dangerous. Please only output one of the two words 'safe' or 'dangerous', and provide one sentence of reasoning in Chinese. Email content: '[email body]'."

[0146] When constructing the query text, the server concatenates the prompt statement and the body text according to the format required by the model to form a single sequence input, or, in multi-input mode, inputs the prompt statement and the body text separately. The server's output is the query text or text pairs used to invoke the generative artificial intelligence model.

[0147] Step 8: The server invokes a generative artificial intelligence model and obtains hazard assessment information.

[0148] The server's input consists of the feature information output from step 6 and the query text output from step 7. The server inputs the query text into the text encoding module of the generative AI model via the model interface, using the feature information for auxiliary encoding or as additional channel input. During computation, the server sequentially performs neural network operations such as word embedding mapping, multi-layer self-attention calculation, feedforward network transformation, and normalization to obtain the semantic representation of the text. At the output layer, the server generates hazard classification results (e.g., "safe" or "dangerous"), hazard scores (e.g., continuous values ​​from 0 to 100), and descriptive text explaining suspicious features, based on the defined task. The server receives the model's returned results in text or structured output format and performs preliminary parsing. The server's output is a hazard assessment information containing hazard category, hazard score, and descriptive content.

[0149] Step 9: The server will compare the identification information with the information set and calculate the confidence level.

[0150] The server's input consists of the list of identification information output in step 5 and the information set in the storage device. The server first queries the high-frequency cache table in memory to quickly determine if the identification information exists in the trusted source list or blacklist. For identification information that does not match the cache, the server accesses the information set in the large-capacity storage device and performs a matching query using an index or hash structure. The server calculates the source reputation score and historical risk marker for each piece of identification information, and calculates the overall confidence level based on the combined results of multiple pieces of identification information, such as using a weighted average, maximum risk priority, or logical combination rules. The server's output is the confidence score or level (e.g., "high confidence," "medium confidence," "low confidence") for the document information.

[0151] Step 10: The server integrates risk assessment information, confidence information, and sentiment information to generate risk information.

[0152] The server's inputs include the hazard assessment information output in step 8, the confidence level information output in step 9, and the sentiment information provided by the sentiment analysis module. The server performs numerical calculations and nonlinear transformations on these three types of inputs according to a preset fusion formula or a small neural network. For example, the server calculates the overall hazard level based on the hazard score, confidence score, and user anxiety level. If the hazard score is high and the confidence level is low, the final hazard level is increased; if the user's emotions indicate anxiety, the alarm trigger threshold is appropriately relaxed. The server can output a hazard level (such as "low," "medium," or "high") or a specific numerical value. The server's output is hazard information representing the overall risk level.

[0153] Step 11: The server generates notification content and notification control prompts.

[0154] The server's input consists of the danger level information output in step 10 and the suspicious feature description text output in step 8. Based on the danger level, the server selects a notification template, inserts key suspicious features and suggested actions, and generates notification content text suitable for user understanding. For example, the server might generate a message such as, "The system has detected that this email may be a phishing scam. Danger level: High. Suspicious points include: requests for money transfers, and the inclusion of suspicious links. Recommendations: Do not click the links, and do not provide personal information." The server also generates notification control prompts to guide the terminal on how to display the notification, such as, "High-risk email, a full-screen warning should pop up with the title highlighted in red." The server's output is notification data containing the notification body, display level, and control instructions.

[0155] Step 12: The server sends notification data to the terminal, which then displays and guides the user's actions.

[0156] The server's input consists of the notification data output from step 11 and the terminal identification information. The server sends the notification data to the designated terminal in structured message format via an application-layer communication protocol. Upon receiving the notification data, the terminal parses the notification content and control commands locally and adjusts the interface presentation based on the risk level. In high-risk situations, the terminal displays warning text to the user via pop-ups or full-screen prompts; in low- to medium-risk situations, it alerts the user in the email list using icons or color-coded indicators. Based on the notification content, the terminal provides the user with operation buttons such as "Delete Email," "Move to Spam," and "Continue Viewing," and records the user's selections as behavioral data. The terminal's output includes visual prompts for the user and potential user feedback data.

[0157] Step 13: The server receives user feedback and uses it for subsequent model and rule optimization (optional).

[0158] The server's input consists of user feedback data uploaded by the terminal, including email identifiers, user selection results, and optional user description text. The server stores this feedback data, along with the original risk level information and identification information, in a training data warehouse. During subsequent training or fine-tuning of the generative AI model, the server adds these samples with real user feedback labels to the training set, updates the loss function calculation, and adjusts the model weights. Simultaneously, the server uses feedback statistical analysis to determine whether the identification information comparison rules and threshold settings need optimization to reduce false positives or false negatives. The server's output consists of updated model parameters and rule configuration files, used to improve the accuracy and efficiency of subsequent processing.

[0159] Application Example 1 The process flow corresponding to the specific processing in Use Case 1 will be described below. The various parts of the system described below are implemented by the data processing device 12 and the intelligent device 14. Furthermore, the data processing device 12 is referred to as the "server" and the intelligent device 14 is referred to as the "terminal".

[0160] In the modern communication environment, the number of malicious communications is constantly increasing, and phishing, fraud, and account theft using email and other communication methods are becoming increasingly complex. Traditional methods for detecting dangerous emails often rely on fixed rules or single machine learning classifiers. These technologies suffer from the following technical problems: First, rule-based matching methods are poorly adaptable to new and variant language, requiring frequent manual rule updates. While computational resources are controllable, detection accuracy is difficult to guarantee. Second, single classification models are usually trained based on static features, limiting their ability to understand context, tone, and implicit threatening intent, leading to false negatives or missed detections. Third, while servers can centrally process emails, there is a lack of systematic technical solutions for efficiently organizing input data, constructing prompts suitable for generative AI models, and effectively integrating the outputs of generative AI models with traditional statistical learning results, resulting in wasted computational resources and reduced real-time performance. Fourth, existing systems rarely link user emotional states with risk assessments, making it difficult to dynamically adjust warning strategies and presentation methods for users in high-stress or high-anxiety states, thereby improving the comprehensibility of security prompts and the effectiveness of interventions.

[0161] Therefore, how to improve the parsing of communication information, feature extraction, prompt statement generation, and generative artificial intelligence model invocation on the server side, and efficiently integrate them with statistical learning judgment results, while adaptively adjusting the risk level judgment and warning presentation based on user emotion recognition results, so as to improve the accuracy and robustness of dangerous communication detection while ensuring real-time performance, has become an urgent computer technology issue to be addressed.

[0162] The specific processing performed by the specific processing unit 290 of the data processing apparatus 12 in Application Example 1 is achieved by the following means.

[0163] In this invention, the server includes: a processing device for generating prompt statements for risk assessment based on language and structural information extracted from communication information and inputting them into a generative artificial intelligence model to obtain communication risk assessment results; a processing device for comparing the identification information contained in the communication with the set of recorded information and calculating a credibility index, and correcting the risk assessment based on the index; a processing device for extracting specific words and expressions from the communication text using language information processing technology and embedding them into the prompt statements; a processing device for learning the relationship between the characteristics and risk of past and isolated communications using statistical learning and outputting a numerical risk, and fusing the numerical result with the evaluation result of the generative artificial intelligence model to determine the final risk; a processing device for generating and sending warning information to the terminal based on the final risk, and performing isolation processing on communications exceeding a threshold; a processing device for inferring the user's emotional state through emotion recognition processing and adjusting the content and presentation of the warning information accordingly, and correcting the final risk when necessary; and a processing device for updating the judgment rules and prompt statement components based on isolated communications and user operation records to continuously optimize the accuracy of risk estimation and generative artificial intelligence model evaluation. This allows for the formation of a collaborative processing chain within the server, encompassing raw communication parsing, feature extraction, prompt statement generation, generative artificial intelligence model inference, statistical learning fusion judgment, and adaptive warning presentation. This significantly improves the accuracy and robustness of malicious communication identification while ensuring controllable processing latency, and enhances the overall processing performance of the computer in the automatic judgment of dangerous communications through a continuous online learning mechanism.

[0164] "Communication information" refers to digital data used to express content transmitted between a terminal and a server via a network, including but not limited to email messages, text messages, and the body, title, header information, and additional information contained therein.

[0165] "Language information" refers to the sequence of symbols and their grammatical and semantic features extracted from the text content of communication information, such as the body of the information, to represent the content of natural language, including words, sentences, phrases and their relationships.

[0166] "Structure information" refers to metadata and format information related to the organization of communication information, including but not limited to sender information, recipient information, time information, subject information, message header fields, body segmentation structure, and layout structure of links and attachments.

[0167] "Prompt statements" refer to request content, written in natural language or structured text, generated by the server and input into the generative artificial intelligence model, used to instruct the generative artificial intelligence model to perform risk assessment or related analysis on the communication information.

[0168] "Generative artificial intelligence models" refer to artificial intelligence models that are trained on a large amount of sample data and can automatically generate text or structured output based on input prompts, including but not limited to deep learning models used for natural language understanding and generation.

[0169] "Assessing the degree of danger" refers to the process of quantifying or classifying whether a communication message has malicious, fraudulent, or other harmful intent based on its content, structure, and related characteristics.

[0170] "Identification information" refers to identifying data contained in communication information that is used to indicate a specific entity or resource, including but not limited to network addresses, telephone numbers, account identifiers, and other contact information or access information.

[0171] "Record information set" refers to a structured or semi-structured dataset stored in a storage medium for comparison with identification information, including but not limited to whitelists, blacklists, reputation databases, and historical access record sets.

[0172] "Confidence index" refers to a parameter or value that quantifies the credibility of the source of the identification information or related entities based on the comparison results between the identification information and the set of recorded information.

[0173] "Language information processing technology" refers to computational techniques used to analyze and transform text data, including but not limited to methods such as word segmentation, part-of-speech tagging, syntactic analysis, keyword extraction, named entity recognition, and text feature extraction.

[0174] "Specific word set and expression set" refers to the set of words, phrases and sentence patterns extracted from the text of a communication that are related to risk assessment, and are used to reflect possible fraudulent intentions, requests for sensitive information or other abnormal expressions.

[0175] "Statistical learning processing" refers to the computational process of building and training data models based on probabilistic and statistical methods, thereby outputting prediction results or scores given input features, including but not limited to regression analysis, classification analysis, and cluster analysis.

[0176] "Features" refer to numerical or symbolic parameters extracted from communication information and its context that can be used as input to statistical learning models, including text features, structural features, and behavioral features.

[0177] "Judgment rules" refer to the set of algorithmic logic or parameters obtained by statistical learning or expert setting, used to classify or score the risk of communication information based on feature quantities.

[0178] "Final risk level" refers to the final quantitative or graded evaluation of the risk level of communication information after integrating the output results of generative artificial intelligence models and the inference results of statistical learning.

[0179] "Warning information" refers to output content generated based on the final level of danger, used to alert users to potential risks associated with communication information. This includes text prompts, graphic markers, or sound prompts.

[0180] "Isolation processing" refers to security control operations performed on communication information that has been determined to exceed a predetermined danger threshold, including moving the communication information to an independent storage area, restricting access or operation permissions, etc.

[0181] "Independent storage area" refers to a region on the storage medium that is logically separated from the ordinary communication storage area. It is used to store isolated communication information and subject to access control.

[0182] "Access restrictions" refer to permission control measures imposed on communication information stored in a separate storage area, including but not limited to prohibiting direct opening, prohibiting downloading attachments, or requiring additional verification steps.

[0183] "Emotion recognition processing" refers to the computational process of inferring a user's emotional state by analyzing input information (including voice, text, images, or interactive behavior) related to the user.

[0184] "Emotional state" refers to the state information obtained through emotion recognition processing that indicates the user's current psychological or emotional tendency, including but not limited to tension, anxiety, calmness, and alertness.

[0185] "Terminal" refers to electronic devices that allow users to send, receive, and view communication information, including but not limited to mobile terminals, fixed terminals, and their application programming interfaces (APIs).

[0186] "User operation logs" refer to the historical information of viewing, deleting, marking, or restoring communication information using a terminal or server. They are used to reflect the user's actual level of trust in the communication information and their processing behavior.

[0187] This invention will describe in detail how to implement the system described in the claims in a specific computer hardware and software environment by combining the collaborative work of the server, terminal, and user. This embodiment is merely an example; the server can be implemented using different hardware configurations, operating systems, and software components, as long as it implements the technical features defined in the claims.

[0188] I. Overall System Composition The server is installed on an electronic information processing device equipped with a network interface and large-capacity storage. The server may employ a multi-core general-purpose processor (such as a central processing unit based on x86 or ARM architecture), optional graphics processing unit, or dedicated accelerator to perform deep learning inference and training operations. The server runs a general-purpose operating system, such as a Unix-like system, and installs the following software components: - Execution environment: such as the Python interpreter; - Natural Language Processing Libraries: such as word segmentation and syntactic analysis libraries; - Machine learning libraries: such as libraries for statistical learning and feature vectorization; - Deep learning frameworks: such as frameworks for building and reasoning generative artificial intelligence models; - Database management system: such as relational database or document database, used to store email records, features, model parameters and user operation records.

[0189] A terminal is an electronic device used by a user, which can be a mobile terminal or a fixed terminal. The terminal runs communication applications or email clients, interacts with the server through network protocols, sends email information, and receives risk assessments and warnings from the server. Users can view email content, read warning prompts, and perform operations such as deletion or restoration through the terminal.

[0190] II. Server-side modules and data structures 1. Communication acquisition and storage module The server includes a communication acquisition module for receiving email information from terminals. The server abstracts each email into a structured data record, containing at least the following fields: - Email Identifier Field: A globally unique identifier; - Header information fields: sender identifier, receiver identifier, timestamp, subject, etc.; - Body text field: Plain text format; - Attachment metadata fields: filename, size, type; - Processing status fields: pending analysis, analyzed, isolated, etc.; - Hazard field: Numerical or graded hazard results (to be filled in after subsequent processing).

[0191] The server stores the aforementioned structured records in a database and organizes them in an indexed manner on the storage medium for quick retrieval and updates based on email identifiers.

[0192] 2. Language Information Processing Module The server includes a language information processing module, which runs in an interpreted execution environment and uses a natural language processing library to convert the email body into a data structure suitable for subsequent analysis. The server generates the following intermediate data for each email: - Word sequence list: An array of word units obtained after segmenting the main text; - Sentence list: An array of sentences divided according to punctuation and grammar; - Parts of speech and dependency relation structure: a diagram or table representing the internal grammatical structure of a sentence; - List of identified information: Network addresses, phone numbers, account identifiers, etc., extracted using regular expressions and rules.

[0193] The server stores the aforementioned intermediate data as independent tables or nested structures, with email identifiers used as foreign keys for association. This module explicitly utilizes specific algorithms (word segmentation, part-of-speech tagging, dependency parsing) to convert the original string into structured language features. This conversion allows subsequent algorithms to perform efficient calculations based on word frequency, syntactic patterns, etc., thereby reducing repeated scanning of the original long text and improving processing speed.

[0194] 3. Feature Extraction and Statistical Learning Module The server includes a feature extraction module and a statistical learning module. The server uses feature vectorization tools from a machine learning library to construct a numerical feature vector for each email. Features include, but are not limited to: - Text statistical features: word frequency, character length, number of sentences, proportion of interrogative sentences, etc.; - Keyword characteristics: Does it contain words related to account, password, identity information, or emergency operations? - Pattern characteristics: Count whether combinations such as "urgent + link + enter password" appear; - Identification information characteristics: number of suspicious network addresses, proportion of unknown domain names; - Language anomaly features: grammatical error count, abnormal word order pattern count; - Historical correlation characteristics: the degree of correlation between the sender and past dangerous email records, etc.

[0195] The server utilizes a pre-trained classification model in the statistical learning module, such as a gradient boosting, random forest, or logistic regression model, to map the aforementioned feature vectors to numerical hazard scores. This model is trained using supervised learning methods. During the training phase, the server uses labeled historical email data, measures the error between the predicted hazard score and the true label using cross-entropy loss or log loss functions, and updates the model parameters using gradient descent or its variations (such as adaptive learning rate algorithms). In this way, the server compresses complex textual features into computable low-dimensional vectors and quickly outputs hazard scores with limited computational resources, thereby improving processing latency and resource utilization.

[0196] 4. Generative Artificial Intelligence Model Module The server also includes a generative artificial intelligence model module for deep semantic-level risk assessment of emails. The server can use a Transformer-based neural network model, which includes multi-layered self-attention encoders and decoders, capturing long-distance dependencies between words through a multi-head attention mechanism. The model parameters are initialized during pre-training using large-scale corpus semantic prediction tasks (such as masked language modeling and next-sentence prediction), and then fine-tuned on security-related domain data.

[0197] The server interacts with the model in the following way: - The server generates a text-based prompt statement based on the keywords, suspicious expressions, identification information list obtained from the feature extraction module, and the preliminary results from the statistical learning module; - The server encodes the prompt statement as an input sequence into a vector and sends it to the neural network model for forward propagation computation; - The model outputs a structured text (e.g., describing the danger level and reasons in natural language), from which the server then extracts classification results such as "high-risk / suspicious / safe" and corresponding key phrases.

[0198] Generative AI models no longer rely on fixed rules for internal processing; instead, they automatically learn complex semantic patterns through multi-layered non-linear mappings and attention weights. For example, the model can recognize the implicit intent of "subtly requesting passwords," rather than being limited to obvious keyword matching, providing a technological foundation for identifying variant scam tactics.

[0199] III. Prompt Statement Generation and Model Calling Before invoking the generative AI model, the server needs to generate appropriate prompts. Based on the email body, identification information, dangerous phrases, and statistical learning results, the server constructs natural language text containing context and task descriptions. Example prompts include: Example prompt statement 1: "You are an email security analysis assistant."

[0200] The following is the content of an email. Please analyze whether this email is a dangerous email or harmful communication.

[0201] Based on the identification information in the email (such as phone number, URL), whether the language is natural, and whether it asks users to enter passwords or identity information, determine its risk level (safe / suspicious / high risk) and list the main dangerous phrases.

[0202] Email body: 'Dear user, your bank account has been frozen. Please click the following link immediately and enter your account password and ID number to restore normal access: https: / / example-bad-link.com' Please output the hazard level and provide the main causes of the hazard. Example prompt statement 2: "Please use the key information given below as a generative artificial intelligence model to assess the danger of an email."

[0203] The extracted list of dangerous keywords and phrases is as follows: ['Account has been frozen', 'Click the link below now', 'Enter your account password', 'ID number'] The email body is as follows: 'Dear user, your bank account has been frozen. Please click the following link immediately and enter your account password and ID number to restore normal access: https: / / example-bad-link.com' Please determine whether this email is malicious. If so, what type of harmful communication it is (e.g., phishing, scam, account theft, etc.), and explain its main dangers. The server passes the aforementioned prompt statement as input to the generative artificial intelligence model. The model outputs analysis results based on the agreed-upon task description and contextual information in the prompt statement. The server then transforms the model's judgment results into internal structured data by parsing keywords or fixed-format fragments in the output text.

[0204] IV. Emotion Recognition and Warning Adjustment The server also includes an emotion recognition processing module. The terminal can upload user speech snippets, facial images captured by the camera, or behavioral information such as the user's interaction rhythm with emails to the server. The server analyzes these inputs using emotion recognition algorithms (such as convolutional neural networks combined with time series models) to infer whether the user is currently experiencing an emotional state such as tension, anxiety, or calmness.

[0205] The server adaptively adjusts warning messages based on the user's emotional state. For example, when a user is inferred to be anxious or unfamiliar with cybersecurity, the server generates more detailed and explicit warning text and requires it to be displayed on the terminal in a more prominent visual format (such as a full-screen notification or high-contrast colors). When a user is inferred to be familiar with the system and in a calm state, the server can display only a concise danger label to avoid excessive disturbance. This emotion-driven adjustment is not simply business logic, but rather optimizes the interface presentation based on predictions of the user's information processing load. This helps reduce the probability of users making mistakes due to ignoring warnings, thereby improving the overall system's effective security from a technical perspective.

[0206] V. Isolation and Online Learning After integrating the statistical learning results with the output of the generative artificial intelligence model, the server determines the final risk level. For emails exceeding a preset threshold, the server marks them as objects requiring isolation and moves the corresponding storage records to a separate storage area. The server stores the email content in an isolated directory at the file system level and sets access control flags in the database. When a terminal requests to view the email, it must pass the server's verification. The server can require secondary confirmation or restrict attachment downloads.

[0207] The server simultaneously records the user's operational history on the terminal, such as whether the user marked emails deemed safe by the system as spam, or restored emails deemed dangerous to normal status. These operation records are periodically collected by the server and used as incremental training samples for the statistical learning model. The server uses error analysis results to update feature weights, optimizes classification model parameters using new training epochs, or updates the explanatory content and set of dangerous phrases in the warning statements based on new fraud patterns. Through this online learning mechanism, the server's judgment ability gradually strengthens over time, thereby reducing false positives and false negatives and improving detection and accuracy.

[0208] VI. Technical Effects and Improvements in Computer Technology The server is implemented through the above modularization, and internally employs specific data structures (structured email records, feature vector matrices, prompt text, sentiment status tags, etc.) and algorithmic processes to produce the following computer technology-level effects: 1. By pre-segmenting, syntactically analyzing, and extracting information from the email body, the server reduces the direct processing length of the original long text by the generative artificial intelligence model, lowers the input sequence length and attention computation complexity of the neural network, thereby shortening the inference time and improving the processing throughput under the same hardware conditions; 2. By integrating statistical learning with generative artificial intelligence models, the server establishes a combined decision-making mechanism between numerical risk level and semantic inference results, reducing misjudgments caused by the bias of a single model and significantly improving the accuracy and robustness of dangerous email identification. 3. By isolating storage areas and access restriction flags, the server performs physical and logical dual isolation of high-risk communication at the storage system level, effectively reducing the chance of malicious content spreading on terminals or networks and achieving fine-grained control over information flow paths; 4. By adjusting the warning presentation method based on the emotion recognition results, the server dynamically adjusts the interaction strategy according to the user's understanding and response to the information, thereby improving the final actual protection effect while keeping the calculation results unchanged. This emotion-driven interface control is an improvement of computer technology at the human-computer interaction level. 5. Through an online learning mechanism, the server continuously updates the judgment rules and prompt statement components with the help of new samples, enabling the system to adapt to the evolution of malicious communication patterns without frequent manual intervention to update the rule base, thereby improving the system's automatic adaptability and maintenance efficiency.

[0209] VII. Alternative Implementation Methods The server can adopt different generative artificial intelligence model structures according to the actual application scenario, for example: - Use a language model that only contains the encoder and output the hazard level directly through an attached classification head; - Deploy lightweight models in environments with limited computing resources to achieve fast inference by reducing the number of layers and parameters; - Using a multi-model combination architecture, dedicated sub-models are trained for different types of communication (work emails, financial emails, social emails), and then the server selects the appropriate model during the inference stage to improve accuracy and efficiency.

[0210] The server can also use different algorithms in the statistical learning part, such as gradient boosting trees, linear support vector machines, or deep feedforward networks, as long as this part can achieve numerical risk estimation of feature vectors and can be fused with the output of generative artificial intelligence models.

[0211] The terminal can flexibly implement interface control based on the control information returned by the server. For example, it can change the sorting rules in the email list, add special icons to high-risk emails, and pop up a confirmation dialog box before the user clicks on a high-risk email. These controls are all driven by the priority and danger labels issued by the server.

[0212] Through the above implementation, the server does not simply replace manual email viewing, but internally constructs a complete technical chain of multi-level feature extraction, statistical learning and generative artificial intelligence model collaborative judgment, emotion-driven interaction adjustment, isolated storage and online learning, thereby achieving substantial improvements in communication security processing capabilities and resource utilization efficiency within the computer system.

[0213] use Figure 12 The processing flow is explained.

[0214] Step 1: The server receives communication information from the terminal and generates basic records.

[0215] Input: Raw email data sent by the terminal (including header, body, attachment metadata, etc.).

[0216] The server parses network protocol data packets, restoring the email content from the transmission format to the internal unified format; the server extracts the sender identifier, recipient identifier, subject, timestamp, original body text, and attachment metadata, and assigns a unique email identifier to the email; the server combines these fields into a structured basic record and writes it to the database.

[0217] Output: "Raw mail records" with unique email identifiers.

[0218] Step 2: The server preprocesses the email body and extracts language information.

[0219] Input: The body text field of the original email record output in step 1.

[0220] The server calls a text cleaning program to remove redundant HTML tags, script fragments, and control characters, and converts the main text into plain text and uniform encoding. The server uses a word segmentation tool to split the main text into word sequences and uses a sentence segmentation algorithm to divide sentences according to punctuation and grammar rules. The server generates word lists and sentence lists and stores them in association with email identifiers.

[0221] Output: Language information data structures such as word sequences and sentence sequences associated with email identifiers.

[0222] Step 3: The server extracts identification information and constructs structural information.

[0223] Input: The word sequence and original text output from step 2.

[0224] The server uses predefined regular expressions and rules to match identification information such as network addresses, phone numbers, and account identifiers from the text; the server records each piece of data identified along with type tags (such as "URL", "phone", "account") into an identification information list; the server also reads the sender's domain name, recipient's domain name, etc. in the header as part of the structure information; through this data processing, the server converts key identifiers in unstructured text into structured entries.

[0225] Output: A list of identification information and a set of structured information including sender, receiver, etc.

[0226] Step 4: The server constructs numerical feature vectors for statistical learning.

[0227] Input: Language information from step 2, recognition information from step 3, and structural information.

[0228] The server calculates text statistical features, including word count, sentence count, and frequency of specific words; the server uses a dangerous phrase dictionary to count the frequency of words related to passwords, accounts, identity information, and emergency operations; the server counts the number of suspicious network addresses, the ratio of unknown domain names, and the number of language anomalies (e.g., the number of grammatical errors or abnormal word order patterns); the server uses a feature vectorization tool to encode these discrete features into fixed-length numerical vectors, forming a row of a feature matrix that can be computed by the model.

[0229] Output: A numerical feature vector corresponding to the email identifier.

[0230] Step 5: The server uses a statistical learning model to make a preliminary risk assessment of the emails.

[0231] Input: The numerical feature vector output from step 4.

[0232] The server inputs the feature vector into a pre-trained statistical learning model (e.g., a classifier based on a tree model or a linear model); the server internally performs operations such as matrix multiplication, nonlinear transformation, or tree structure traversal to calculate the probability corresponding to each type of label; the server obtains a set of probability distributions for categories such as "safe", "suspicious", and "dangerous", and calculates preliminary danger values ​​(e.g., danger probability or classification results) based on these distributions.

[0233] Output: Preliminary risk assessment results (including category labels and probability values).

[0234] Step 6: The server compares the identified information with the set of recorded information and calculates a credibility index.

[0235] Input: The list of identification information output from step 3 and the collection of record information stored on the server (whitelist, blacklist, reputation database, etc.).

[0236] The server performs a database query or hash lookup for each piece of identification information, matching the identification information with entries in the record information set; the server calculates a reputation score for each piece of identification information based on the matching results, the number of historical records, and the type of associated events; the server combines the reputation scores of each piece of identification information into an overall credibility index for the email by weighted averaging or other aggregation methods.

[0237] Output: The credibility index value or level corresponding to the email identifier.

[0238] Step 7: The server generates prompts for invoking generative artificial intelligence models.

[0239] Inputs: Language information from step 2, identification information list from step 3, preliminary risk assessment results from step 5, and credibility index from step 6.

[0240] The server selects representative sentences and key phrases to form a simplified email content summary; the server selects words appearing in the email from a dangerous phrase dictionary to form a dangerous phrase list; the server embeds this information, along with preliminary danger labels, danger probability, and credibility indicators, into a pre-designed natural language template to generate a complete prompt text; the server performs data processing operations such as string concatenation and placeholder replacement to form a clearly structured prompt with a clear task description.

[0241] Output: Prompt text for generative artificial intelligence models.

[0242] Step 8: The server invokes a generative artificial intelligence model and obtains semantic-level risk assessment results.

[0243] Input: The prompt message output in step 7.

[0244] The server encodes the prompt statement into a text input sequence and passes it to a generative artificial intelligence model deployed locally or remotely. The multi-layer neural network inside the model performs vector operations and attention weight calculations based on the prompt statement to generate output text containing a hazard level judgment and a reason description. After receiving the output text, the server extracts classification results such as "high-risk / suspicious / safe" as well as the corresponding main hazard phrases and reason explanations through pattern matching or simple parsing rules.

[0245] Output: Risk assessment results of the generative artificial intelligence model (including risk level and reason).

[0246] Step 9: The server integrates statistical learning results with generative artificial intelligence model results to determine the final risk level.

[0247] Input: The preliminary risk assessment results from step 5 and the risk evaluation results from step 8.

[0248] The server combines the numerical risk level given by statistical learning with the level given by the generative artificial intelligence model according to the pre-set fusion rules or weighting function. The server can appropriately increase the overall risk level according to the "high risk" output of the generative artificial intelligence model, or adopt a conservative strategy when the two are inconsistent. After completing the numerical weighting, threshold comparison or rule judgment internally, the server determines a unified final risk level and corresponding value.

[0249] Output: Final hazard result (including final rating and overall hazard value).

[0250] Step 10: The server generates and sends warning messages based on the final level of danger.

[0251] Input: The final risk level result output from step 9 and basic information from the original email record (subject, sender, etc.).

[0252] The server selects different warning templates based on the level of danger, fills in the email subject, sender introduction, and main reasons for danger, and forms a warning text for the user. When necessary, the server attaches a list of dangerous phrases extracted by a generative artificial intelligence model so that the user can understand the source of the risk. The server encapsulates the warning information into a notification message and sends it to the terminal through a communication protocol.

[0253] Output: A warning message data packet sent to the terminal.

[0254] Step 11: The terminal receives warning messages and presents security prompts to the user.

[0255] Input: The warning message data packet output from step 10 and its associated email identifier.

[0256] The terminal parses the warning information sent by the server, reads the danger level, prompt text, and control instructions (such as whether to mark it as high-risk); the terminal adds an icon or color mark to the corresponding email in the email list, and displays a pop-up window or top warning bar when the user opens the email; the terminal adjusts the sorting position or notification method of the email according to the control instructions, and displays a concise risk description to the user on the interface.

[0257] Output: Safety prompts and visual markers displayed on the terminal screen.

[0258] Step 12: The server isolates high-risk emails and sets access restrictions.

[0259] Input: The final risk level result and email identifier output from step 9.

[0260] The server determines whether the risk level exceeds the isolation threshold. If it does, the server updates the status field of the email to "isolated" in the database and moves its physical storage location to the isolation zone directory. The server sets an access control flag for the email, restricting terminals from directly downloading attachments or forwarding content. The server records the isolation result in the isolation list for subsequent review and online learning.

[0261] Output: Updated email status logs and quarantine list entries.

[0262] Step 13: The server collects user operation records and updates the model online.

[0263] Input: User actions on the terminal (such as marking emails as spam / false alarms, restoring quarantined emails, etc.), and the quarantine list and historical risk results generated in step 12.

[0264] The server periodically collects and aggregates user operation records, compares them with the final danger level results of the corresponding emails, and identifies possible misjudgments by the statistical learning model and generative artificial intelligence model. The server adds these samples to the training dataset, recalculates or incrementally updates the parameters of the statistical learning model, and updates the dangerous phrase dictionary and prompt statement templates based on newly emerging fraudulent rhetoric. By performing this data retraining and template adjustment process, the server improves the accuracy and adaptability of the model in subsequent email analysis.

[0265] Output: Updated statistical learning model parameters, updated dangerous phrase dictionary, and new prompt statement template configuration.

[0266] Alternatively, an emotion engine for inferring user emotions can be combined. That is, the specific processing unit 290 can also use the emotion-specific model 59 to infer user emotions and perform specific processing using user emotions.

[0267] Example 2 The flow of a specific process in Example 2 will be described. Each part of the system described below is implemented by the data processing device 12 and the smart device 14. The data processing device 12 will be referred to as the "server," and the smart device 14 as the "terminal."

[0268] In recent years, with the rapid increase in electronic communication volume, phishing, spam, and other malicious electronic communications have become increasingly diversified. Attackers not only forge website addresses and phone numbers but also evade traditional detection methods by disguising text content and combining multiple communication records. Current technologies typically employ the following methods to assess the security of electronic communications: first, relying solely on static data sets such as blacklists and whitelists for comparison; second, simply using keyword matching or rule engines to filter text; and third, applying machine learning models or generative artificial intelligence models in isolation to analyze communication content. However, these methods suffer from the following technical problems.

[0269] First, relying solely on static database comparison methods is insufficient to promptly address emerging malicious communications. For suspicious URLs or identification information not yet registered in the database, the system often only provides an "unknown" conclusion, lacking effective risk assessment capabilities and resulting in a high false negative rate.

[0270] Second, natural language processing rules that rely solely on fixed keywords or patterns are not robust enough when faced with attackers' constantly changing wording. The rules are costly to maintain and are easily bypassed by adversarial examples, resulting in insufficient detection accuracy of the system in complex scenarios.

[0271] Third, when machine learning models or generative AI models are used alone, the lack of an effective mechanism for integrating multi-source information such as database comparison results and rule features often makes it difficult to provide interpretable and quantifiable comprehensive risk indicators, resulting in insufficient precision and reliability in system strategy decisions and user prompts. Especially for generative AI models, without task-specific optimized prompts, model outputs are prone to instability and uncontrollability, reducing the overall reliability of the system.

[0272] Fourth, existing systems mostly focus on static analysis of single communications and lack a mechanism for adaptively updating the judgment logic and prompt statements during online service operation. This makes it impossible to fully utilize the continuously accumulating real detection data and user interaction data to continuously optimize model performance, thus limiting the performance improvement potential of the system in the long-term operation process.

[0273] Therefore, there is an urgent need for an improved solution based on computer technology. This solution would organically combine standardized processing of identification information, database comparison, machine learning discrimination, and generative artificial intelligence model analysis on the server side. Furthermore, it would introduce task-oriented prompts for generative artificial intelligence models and an online learning and update mechanism to achieve a high-precision, interpretable, and adaptive comprehensive trust evaluation of electronic communication security. This would substantially improve the technical performance of electronic communication security detection systems.

[0274] The specific processing performed by the specific processing unit 290 of the data processing apparatus 12 in Embodiment 2 is achieved by the following means.

[0275] In this invention, the server includes a processing unit for receiving electronic communication content and / or identification information from a terminal device; a processing unit for format correction of the identification information to generate standardized identification information and comparing it with a storage device storing an information set containing confidence level information and risk level information to obtain a confidence evaluation result; a processing unit for applying natural language processing technology to extract features from the text data in the electronic communication content and performing discriminative processing based on the features and the identification information using machine learning technology to generate a risk level index representing the risk of electronic communication; a processing unit for generating instruction text for a generative artificial intelligence model, the instruction text including prompt statements and input information containing at least one of the electronic communication content and identification information, to prompt the generative artificial intelligence model to analyze the possibility of whether the electronic communication is phishing communication and / or spam communication, and to integrate the obtained analysis results with the risk level index and the confidence evaluation result to form a comprehensive confidence evaluation of the electronic communication; and a processing unit for generating response data containing risk level and its basis information and / or control information based on the comprehensive confidence evaluation and sending it to the terminal device. This allows for the construction of a multi-source information fusion security assessment pipeline on the server side. Through the collaborative work of standardized comparison, machine learning discrimination, and generative artificial intelligence model analysis based on prompt statements, the detection capability for unknown identification information and complex text content can be significantly improved. This yields a comprehensive trustworthiness evaluation result with quantitative indicators and interpretable evidence. Furthermore, during system operation, the system continuously learns and updates by utilizing accumulated comparison results, model outputs, and user interaction data, thereby substantially improving the accuracy, robustness, and adaptability of electronic communication security detection and achieving a performance enhancement of the computer technology itself.

[0276] "Electronic communication" refers to messages or data sent and / or received in digital form through a communication network, including but not limited to emails, text messages, instant messages, web page access requests, and text, links, images, and other information associated with these messages.

[0277] "Identification information" refers to structured information used to identify a specific electronic communication or its communication object, including but not limited to network addresses, telephone numbers, email addresses, account identifiers, and equivalent identification data.

[0278] "Information processing device" refers to a hardware device or combination thereof that has computing power and is able to execute program code to process input data, including but not limited to servers, terminal equipment, computer equipment and processing units thereon.

[0279] A “processing unit” refers to a processing unit that runs on an information processing device and is used to execute predetermined programs to perform processing functions such as data reception, parsing, analysis, calculation, generation, and transmission. It can consist of one or more processors and their software modules.

[0280] "Format correction" refers to the preprocessing of input identification information to eliminate redundant characters, unify encoding forms, standardize representation methods, and / or complete missing information to make it meet the predetermined standard format.

[0281] "Standardized identification information" refers to the identification information representation format that conforms to predetermined representation rules after format correction, and is used as standard input in database comparison and subsequent analysis and processing.

[0282] "Information set" refers to a data set maintained in a storage device that records attribute information related to identification information, including but not limited to a list of records, a data table or its equivalent data structure containing confidence information and risk information.

[0283] "Storage device" refers to a storage medium and its control unit used to store information sets, program code and intermediate data, including but not limited to database systems, magnetic storage media, semiconductor memories and combinations thereof.

[0284] "Confidence information" refers to data used to indicate the degree to which a particular piece of identification information or electronic communication is considered trustworthy, secure, or normal, including but not limited to confidence scores, trust levels, or category labels.

[0285] "Hazard information" refers to data used to indicate the degree to which a particular piece of identification information or electronic communication is considered malicious, suspicious, or unsafe, including but not limited to hazard level, risk score, or risk category label.

[0286] "Comparison results" refers to the output information obtained after querying or matching the standardized identification information with the records in the information set, including whether a record is matched, the confidence level information and risk level information corresponding to the matching item, etc.

[0287] "Confidence assessment results" refer to the evaluation data on the credibility of the identified information based on the comparison results, including qualitative or quantitative judgments on security, reliability, or suspiciousness.

[0288] “Text data” refers to sequence data extracted from electronic communication content, which is based on characters as the basic unit, including but not limited to natural language text, tokenized text sequences, and the representations derived therefrom.

[0289] Natural Language Processing (NLP) technology refers to a class of computer technologies used to parse, understand, and extract features from natural language text, including but not limited to word segmentation, part-of-speech tagging, syntactic analysis, semantic analysis, and feature representation generation.

[0290] "Features" refer to numerical or symbolic representations extracted from text data and / or identification information and used as inputs to machine learning models, including but not limited to vectors, scalars, class labels, or combinations thereof.

[0291] Machine learning technology refers to a class of computer technologies that automatically build models for tasks such as classification, regression, or clustering by statistically learning from training data in order to predict or discriminate input data. These technologies include, but are not limited to, supervised learning, semi-supervised learning, and unsupervised learning methods.

[0292] "Discrimination processing" refers to the process of calculating input features based on machine learning models or rule systems to output categories, probabilities, or scores, thereby determining whether electronic communication is dangerous or belongs to a certain safety category.

[0293] "Risk index" refers to numerical or grade information used to quantify the degree of danger of electronic communication, including but not limited to risk scores, risk levels, probability values, or combinations thereof.

[0294] "Generative AI models" refer to AI models that can generate text, code, or other forms of output based on input prompts. They are usually obtained through training on large-scale data and have the ability to understand and generate natural language.

[0295] "Prompt statements" refer to natural language or structured text used to specify tasks, limit the output range, or guide the direction of analysis for generative artificial intelligence models. They are used as one of the inputs to the model to influence its generated results.

[0296] "Instruction text" refers to the complete input text constructed for a generative artificial intelligence model, including prompts and specific data containing at least one of electronic communication content and identification information, used to trigger the model to perform a specific analysis task.

[0297] "Analysis results" refers to the output information of the generative artificial intelligence model after receiving the instruction text, including but not limited to the judgment of whether the electronic communication is phishing communication and / or spam communication, the reasoning, and the relevant score.

[0298] "Comprehensive confidence assessment" refers to the integrated judgment of the overall credibility and risk level of electronic communication obtained by fusing the confidence assessment results, risk indicators, and the analysis results of generative artificial intelligence models.

[0299] "Response data" refers to response information generated by the processing unit and sent to the terminal device, including the degree of danger related to electronic communication, confidence evaluation results, instructions, and / or control information used to control the display or notification method.

[0300] "Terminal device" refers to a computing device that is used by users, can communicate with a server, and display response data, including but not limited to mobile terminals, desktop terminals, wearable terminals, and other network access devices.

[0301] "Display method" refers to the visual presentation of electronic communication or related prompts in the output interface of a terminal device, including but not limited to color, font, layout, icons, and pop-ups.

[0302] "Notification method" refers to the means by which a terminal device prompts users with security information or warnings related to electronic communication, including but not limited to pop-up notifications, sound alerts, vibration alerts, and status bar prompts.

[0303] "Control information" refers to data used to instruct terminal devices on how to display or notify electronic communications, including but not limited to instruction parameters such as restricting display, marking as dangerous, and highlighting warning information.

[0304] In a preferred embodiment of the present invention, the server, terminal, and user each assume different functional roles, working together to achieve a comprehensive confidence assessment of electronic communications. The system of the present invention implements a specific combination of software modules on the server side, forming a dedicated security analysis pipeline on general-purpose computer hardware. This allows for multi-stage data processing and computation of the identification information and text content of electronic communications, obtaining high-precision and interpretable risk indicators and confidence assessment results.

[0305] In terms of hardware architecture, servers can adopt general-purpose server platforms as information processing devices, such as rack-mounted or virtualized servers deployed in data centers or cloud platforms. At the hardware level, servers may include multi-core central processing units (CPUs, such as multi-core processors based on x86 architecture), graphics processing units (GPUs, used to accelerate deep learning inference and training), main memory (RAM), non-volatile memory (such as solid-state drives, SSDs), and network interface controllers (NICs). At the software level, servers can run general-purpose operating systems (such as Linux kernel-based server operating systems) and use middleware and application frameworks, including but not limited to: web server software acting as reverse proxies and static resource servers (such as implementations based on common HTTP servers), backend frameworks hosting application logic (such as Python-based web frameworks or Java-based application frameworks), and relational database management systems (such as implementations based on MySQL or PostgreSQL databases).

[0306] When generating the program required by the system of this invention, the server can deploy multiple functional modules on the operating system, including: a request receiving module, an identification information normalization module, a database comparison module, a natural language processing module, a machine learning discrimination module, a generative artificial intelligence model interface module, a result fusion module, a response generation module, and a log and learning module. Each module works collaboratively through specific software components and data structures to achieve layered analysis of electronic communications.

[0307] After receiving electronic communication content and identification information from the terminal, the server uses an identification information normalization module to correct the format of the identification information. In this module, the server employs a string processing library and a Uniform Resource Locator (URL) parsing library to perform operations such as protocol unification (e.g., unifying to lowercase "http" or "https"), hostname normalization, removal of trailing forward slashes from paths, and parameter sorting for URL-type identification information; and operations such as space removal, symbol standardization, and country code completion for telephone number-type identification information. During this process, the server uses hash tables or dictionary structures to cache common normalization rules to reduce redundant calculations and improve processing speed. Through this normalization process, the server maps logically equivalent identification information with different written forms to a unified standard form, thereby avoiding missed detections caused by format differences during database comparison. This normalization step directly reduces the uncertainty of the matching space at the algorithm level, improving the comparison hit rate and overall retrieval efficiency.

[0308] In the database comparison module, the server establishes a connection with the relational database management system using a database access driver, and calls predefined SQL query statements to perform exact matching or prefix matching between the standardized identification information and records in the information set. The server can employ a standardized table structure in the database, such as creating primary key or composite indexes for URLs and B-tree indexes for telephone numbers, to achieve efficient retrieval of identification information. After obtaining the comparison results, the server encapsulates relevant fields (such as confidence level, risk level, risk category, historical occurrence count, etc.) into a unified data record object for use by subsequent modules. By designing appropriate index structures and caching mechanisms at the database level, the server reduces the number of I / O accesses and query latency, thereby reducing the overall response time of the communication detection process.

[0309] In its natural language processing module, the server employs language processing techniques such as Chinese word segmentation, stop word removal, and lemmatization for textual data from electronic communications (e.g., email bodies, SMS messages, webpage fragments). The server can use dedicated natural language processing libraries (e.g., deep learning-based word segmenters or traditional statistical word segmenters) to annotate the text and construct feature vectors. For feature representation, the server can use bag-of-words models, TF-IDF vectors, or embedding representations based on pre-trained word vectors, such as bidirectional encoder representation (BiLSTM encoding) or text encoding based on self-attention mechanisms. The server uses these numerical features as input to the machine learning discrimination module, avoiding the rigid limitations of traditional simple keyword matching and enabling the system to capture more complex malicious patterns from context and combined features.

[0310] In the machine learning discrimination module, the server can employ gradient boosting tree models (such as XGBoost or LightGBM types), support vector machines, or neural network models. A typical implementation uses a gradient boosting tree model as a lightweight and efficient classifier, jointly inputting URL structural features (such as the number of subdomains, whether it contains IP addresses, path depth, and whether it contains obfuscated characters) with text feature vectors. During the training phase, the server uses labeled historical communication data as the training set, defining loss functions such as cross-entropy or logarithmic loss, and iteratively updates the weights of the splitting nodes and leaf nodes of each tree using the gradient boosting algorithm. During the inference phase, the server inputs the feature vectors into the model, sequentially calculates the corresponding leaf node outputs based on the splitting conditions of each decision tree, and then sums the outputs of multiple trees to obtain the final risk probability value. This tree model, compared to simple threshold rules, can analyze a higher-dimensional feature space, thus significantly improving the accuracy of identifying dangerous communications.

[0311] In another implementation, the server can employ a discriminative model based on deep neural networks, such as a multilayer perceptron or convolutional neural network, to encode URL character sequences, or a bidirectional long short-term memory network (BiLSTM) to model SMS text sequences. When training these neural networks, the server uses mini-batch gradient descent and its variants (such as the Adam optimization algorithm), with cross-entropy as the loss function. Based on the error between the true labels in the training data and the model's output probabilities, it backpropagates to update the weights and bias parameters of each layer. By introducing anti-overfitting techniques such as L2 regularization and Dropout, the server improves the model's generalization ability on unknown samples. This machine learning model utilizes high-dimensional features and nonlinear decision boundaries to replace manual rules, thereby forming a discriminative logic within the computer that differs from human experience-based judgment, enabling automatic abstraction and efficient identification of complex malicious patterns.

[0312] The server, within its generative AI model interface module, invokes externally or locally deployed generative AI models. These models can employ large-scale language models based on a Transformer architecture, containing multiple layers of self-attention networks, feedforward networks, and normalization layers. The server constructs specialized prompts for this model, enabling it to focus on security analysis tasks rather than generalized generation during inference, thereby reducing irrelevant output and enhancing decision consistency. Examples of server-generated prompts include: "You are a network security analysis assistant. Please analyze the following URL to determine its security, and provide your judgment and reasons from two perspectives: 'Is it likely a phishing website?' and 'Is it likely used for malware distribution?' Also, please give an overall risk level (low / medium / high). The URL is as follows:" https: / / example.com or: "Please determine whether the following Chinese text message has phishing or fraud characteristics, list your judgment criteria, and finally give a risk score between 0 and 1."

[0313] SMS content: Your XX Bank account has been flagged for suspicious login activity. Please verify your account at https: / / abc-bank-secure.com within 30 minutes; otherwise, your account will be frozen. When generating the aforementioned instruction text, the server concatenates the prompt statement with specific electronic communication content or identification information into a continuous text input, clearly defining the task boundaries and output format requirements in a natural language-based manner. The server sends this instruction text to the generative AI model service via HTTP or RPC protocol, receives the analysis result text output by the model, and uses regular expressions or lightweight parsing logic to extract key fields such as "risk level," "whether it's phishing," "risk score," and "main reason" from the results. This method of finely controlling the behavior of a large model through prompt statements offers higher controllability and stability compared to traditional black-box calls, and effectively reduces the complexity of unstructured output on subsequent logic, thereby improving the overall computational efficiency and result parsability of the system algorithmically.

[0314] In the result fusion module, the server weighted and fused the confidence assessment results obtained from database comparison, the risk index output by the machine learning model, and the analysis conclusions returned by the generative artificial intelligence model. The server can employ rule-based fusion strategies, such as pre-setting the priority of different information sources (e.g., prioritizing results when the database has marked a URL as high-risk), or it can use a meta-learning-based fusion model, taking the multiple scores and labels as input features and training a lightweight classifier to output the final comprehensive risk level. The server can set different threshold and confidence intervals in the fusion logic. When the conclusions from multiple information sources are inconsistent, the server can dynamically adjust the weights of certain information sources based on historical performance. This multi-source information fusion mechanism is mathematically equivalent to ensemble learning of different discriminators, helping to reduce the bias and noise of individual models and improve the stability and accuracy of the overall judgment.

[0315] In the response generation module, the server constructs response data based on the comprehensive confidence assessment results. The server's response may include structured fields such as "Final Risk Level," "Is Interception Recommended," "Main Basis Explanation," and "Recommended Action," and optionally include control information for influencing terminal display behavior. For example, when the comprehensive assessment result is high-risk, the server instructs the terminal to display the communication with a red warning box, hide the content preview, and only display the warning text; when the assessment result is medium-risk, the server allows the terminal to display the content normally but adds a risk warning icon; when the assessment result is low-risk, the server only records in the background and does not exert significant intervention on the terminal interface. By centrally deciding on display and notification strategies on the server side and distributing structured control information, the system achieves unified management of terminal presentation methods, thereby helping to reduce the logical complexity on the terminal side and lower the technical costs of cross-platform development.

[0316] In embodiments of this invention, the terminal can be a smartphone, tablet computer, desktop terminal, or other computing device with network communication capabilities. At the operating system level, the terminal can run a mobile operating system or a desktop operating system and provide input and display functions through a graphical user interface framework. After receiving user-input identification information or text content, the terminal encapsulates it into structured request data and sends it to the server via a network protocol (e.g., HTTPS). After receiving the response data returned by the server, the terminal updates the interface according to the danger level instructions and control information contained therein, such as changing the font color, displaying a mask layer, or popping up a warning dialog box. The terminal provides security prompts to the user through hardware output components such as a display screen, speaker, and vibration motor, thereby transforming the server-side calculation results into real-world human-computer interaction behavior and improving the user's ability to perceive malicious communication.

[0317] In the usage scenario of this invention, the user initiates a detection request through the graphical interface of the terminal. For example, the user types "https: / / example.com-login-secure.com" into the input box of the terminal application and clicks the "Detect Security" button. The terminal then sends the URL to the server for analysis. After completing data processing within the aforementioned modules, the server returns the comprehensive confidence evaluation result to the terminal. The terminal displays a red warning bar on the interface stating "This URL is a high-risk phishing link; please do not access it," and can restrict link clicking behavior based on control information issued by the server. This allows users to be aware of potential risks before accessing the site, preventing the theft of account passwords or sensitive personal information.

[0318] In the log and learning module, the server records intermediate and final results of each detection process in persistent storage. The server can generate a unique identifier for each request and write fields such as normalized identification information, database comparison results, machine learning output probabilities, generative AI model risk levels, and final comprehensive conclusions into the log table. During periodic training, the server extracts samples from the log table as new training data and updates the parameters of the machine learning model using batch training or incremental training methods. Simultaneously, the server can adjust and optimize the structure of prompt statements based on the performance of the generative AI model under different prompt statements, such as modifying task description language, adding output format constraints, or injecting typical examples. Through this closed-loop online learning and prompt statement optimization mechanism, the system of this invention continuously uses new data to improve the decision boundaries and output stability of the internal model during actual operation, achieving continuous adaptive optimization of the computer's internal algorithm behavior, rather than simply statically automating manual rules.

[0319] In terms of technical effectiveness, this invention improves computer performance through the following causal relationships: The server reduces redundant state space in database matching by standardizing identification information, directly shortening the query path and reducing the omission rate; the server significantly improves the detection accuracy of variant malicious samples by replacing traditional keyword and rule matching with high-dimensional feature modeling and nonlinear machine learning discrimination; the server constrains the originally highly free language generation process within a specific analysis task space by designing prompts on the input side and extracting structured information on the output side of the generative artificial intelligence model, thereby improving the stability and effective information density of the inference results and reducing post-processing complexity and the overall computational burden of the system; the server suppresses single-model misjudgments through a multi-source result fusion strategy, significantly reducing the comprehensive evaluation error; and the server enables the system to dynamically adapt to the latest attack patterns over time through a log-driven online learning and update mechanism, continuously optimizing classification boundaries and improving detection performance in long-term operation. These improvements all occur at the levels of data representation, model structure, parameter updates, and decision logic within the computer, constituting a concrete improvement to computer technology itself, rather than simply mechanically transferring manual tasks to a computer.

[0320] In other implementations, the server can use generative AI models of different scales depending on the deployment environment. For example, in resource-constrained scenarios, a small-to-medium scale Transformer encoder model can be used to judge only key information; in resource-sufficient scenarios, a large-scale pre-trained model can be used to obtain more detailed semantic analysis capabilities. The server can also adjust the database structure according to business needs, such as adding a timestamp field to distinguish between historical and recent risks, or introducing a distributed database to support the horizontal scaling of large-scale information sets. Regarding machine learning models, the server can choose to integrate multiple heterogeneous models (such as tree models and neural networks) to build a more complex ensemble learning structure to further improve robustness. These variations and alternatives, without departing from the technical spirit of this invention, are all within the scope of the embodiments of this invention.

[0321] use Figure 13 The processing procedure is explained.

[0322] Step 1: Users input identification information and / or electronic communication content on the terminal.

[0323] Users open the security detection application in the graphical interface of the terminal, click the input box and enter the identification information to be detected (such as URL, phone number, email address) and text content (such as SMS content, email body) if necessary.

[0324] Input: Raw string data entered by the user via keyboard or touchscreen, such as "https: / / example.com-login-secure.com" and the corresponding SMS text.

[0325] Output: The raw input data object in the terminal's internal memory, containing identification information fields and text content fields.

[0326] Step 2: The terminal sends a detection request to the server.

[0327] The terminal reads the raw data input by the user from the interface components, encapsulates the identification information, text content and metadata (such as timestamps and terminal type) into a request message, and uses a network communication library (such as an HTTP / HTTPS-based client library) to establish a connection to the server and send the request in JSON or other structured formats.

[0328] Input: The raw input data object stored in the terminal memory.

[0329] Output: A request message sent to the server via the network interface, which is encapsulated into a TCP / IP data packet at the transport layer.

[0330] Step 3: The server receives and parses the request.

[0331] The server receives HTTP / HTTPS requests from the terminal through the network interface. The web server forwards the requests to the application process, which uses a JSON parsing library to parse the request body into an internal data structure and extract identification information, text content, and other tag fields.

[0332] Input: The raw request message received from the network stack.

[0333] Output: The request object in the server's memory, containing the parsed identification information string, text content string, and metadata.

[0334] Step 4: The server performs format correction and standardization on the identification information.

[0335] In the identification information normalization module, the server performs string cleanup and structure parsing operations on the identified information. For example, it uses a URL parsing library to decompose the protocol, hostname, path, and query parameters, removes extra spaces, converts uppercase letters to lowercase, and removes extra forward slashes; for phone numbers, it removes spaces and dashes and completes the country code. The server matches and replaces the identified information according to a preset rule table and regular expressions, mapping different writing styles to a unified format.

[0336] Input: The original identification information field in the server request object.

[0337] Output: Normalized identification information string and its structured representation (e.g., an internal structure object containing protocol, hostname, and path).

[0338] Step 5: The server compares the standardized identification information with the information collection database.

[0339] The server constructs an SQL query using a database access library, sending the normalized identification information as query conditions to the relational database management system. The database searches for matches in the index structure (such as a B-tree index) and returns the associated records, including confidence information, hazard information, and risk category. The server parses and encapsulates the returned results, storing the hit status and relevant fields in a comparison result object.

[0340] Input: Standardized identification information and database connection configuration.

[0341] Output: Database comparison result object, including whether it is a match, the corresponding confidence level, risk level, risk category, and historical statistical information.

[0342] Step 6: The server performs natural language processing on the text data and extracts features.

[0343] The server reads the text content from the request object in the natural language processing module, calls the word segmentation library to segment the text, removes stop words, and performs part-of-speech tagging and key phrase extraction through dictionary mapping or model inference. The server further calculates TF-IDF weights, counts the frequency of suspicious keywords, or uses a pre-trained word vector model to map words to a vector space, and then constructs sentence vectors through averaging, pooling, or sequence encoding.

[0344] Input: The string containing the electronic communication text content.

[0345] Output: A set of feature vectors for subsequent machine learning model processing, including word frequency-based vectors and / or word embedding-based vector representations.

[0346] Step 7: The server performs machine learning discrimination processing based on standardized identification information and feature quantities.

[0347] In the machine learning discrimination module, the server concatenates URL structural features (such as the number of subdomains, domain length, whether it contains IP addresses, and path depth) with feature vectors extracted from the text to form a unified feature vector. This vector is then input into the loaded machine learning model (such as a gradient boosting tree model or a neural network model) for forward computation on the CPU or GPU. The model performs multi-level nonlinear transformations or tree-like splits on the features based on its internal parameters, outputting risk probabilities or discrete risk levels. The server converts the probabilities into hazard indicators based on preset thresholds and records the model's internal contribution scores for each feature for interpretation.

[0348] Input: Normalized recognition information structure features and text feature vectors.

[0349] Output: Risk index of the machine learning model (such as risk probability between 0 and 1) and intermediate explanatory information (such as a list of important features).

[0350] Step 8: The server generates prompts and constructs instruction text for the generative artificial intelligence model.

[0351] In the generative artificial intelligence model interface module, the server selects a predefined prompt template based on the request type (URL only, SMS only, or both), and inserts the prompt statement along with the actual recognition information and / or text content into the template to generate complete instruction text. For example: "You are a network security analysis assistant. Please analyze the following URL to determine its security, and provide your judgment and reasons from two perspectives: 'Is it likely a phishing website?' and 'Is it likely used for malware distribution?' Also, please give an overall risk level (low / medium / high). The URL is as follows:" https: / / example.com-login-secure.com” The server constructs the above instruction text through operations such as string concatenation and placeholder replacement.

[0352] Input: Identification information, text content, and predefined prompt templates from the request object.

[0353] Output: Complete instruction text for generative artificial intelligence models, including prompts and content to be analyzed.

[0354] Step 9: The server sends the instruction text to the generative artificial intelligence model and receives the analysis results.

[0355] The server sends instruction text to the generative AI model service deployed locally or externally via HTTP or RPC interfaces, setting parameters such as maximum output length and temperature. The generative AI model, within its internal Transformer network, performs word segmentation, embedding, inter-layer self-attention calculations, and feedforward network operations on the input text, progressively generating analytical conclusions. The server receives the output text from the model service and caches it in memory.

[0356] Input: Instruction text string and model call parameters.

[0357] Output: Analyzed text from the generative AI model, such as natural language descriptions including "high risk," "suspected phishing," and "risk score."

[0358] Step 10: The server extracts structured analysis results from the output of generative artificial intelligence models.

[0359] The server uses regular expression matching, keyword search, or a simple rule engine to parse the analysis text returned by the generative artificial intelligence model, locating risk level words (such as "high," "medium," and "low"), risk score values ​​(such as "0.92"), and key reason sentences. The server maps these elements to structured data fields, such as "llm_risk_level," "llm_score," and "llm_reason."

[0360] Input: The original analytical text output by the generative artificial intelligence model.

[0361] Output: A structured generative artificial intelligence model analysis result object, including risk level, score, and explanation fields.

[0362] Step 11: The server integrates database comparison results, machine learning risk indicators, and generative artificial intelligence model analysis results to generate a comprehensive confidence assessment.

[0363] In the result fusion module, the server reads database comparison results, machine learning outputs, and generative AI model analysis results, and performs weighted calculations or logical judgments based on preset weights or rules. For example, the server can perform a weighted average of risk scores from different sources, or adopt a rule: "If the database marks it as high risk, then directly determine it as high risk; otherwise, make a joint judgment based on machine learning probability and generative model score." Based on the fused results, the server determines the final risk level (e.g., high, medium, low), whether to recommend interception, whether to display a strong warning, and other indicators, and generates a comprehensive evaluation object containing fields such as "final_risk_level," "final_is_dangerous," and "final_reason."

[0364] Inputs: Database comparison results, machine learning risk indicators, and generative artificial intelligence model analysis results.

[0365] Output: Comprehensive confidence assessment object, including final hazard level, hazard status indicator, and comprehensive justification.

[0366] Step 12: The server generates response data and sends it to the terminal.

[0367] In the response generation module, the server constructs a response message based on the comprehensive confidence evaluation object. This message includes prompt text for display on the terminal interface (such as "This URL is a high-risk phishing link, please do not access it") and control information for controlling the terminal's display behavior (such as flags for "Content needs to be masked," "Use red warning style," and "Disable automatic redirection"). The server serializes this response object into JSON or other formats and returns it to the terminal via HTTP / HTTPS.

[0368] Input: Comprehensive confidence evaluation object and display control strategy.

[0369] Output: A response message sent to the terminal, containing result fields and control information fields.

[0370] Step 13: The terminal parses the server's response and updates the interface display and notifications.

[0371] The terminal receives the server's response messages via a network communication library, uses a JSON parser to convert them into internal data structures, and reads the danger level and control information. Based on the control information, the terminal adjusts its UI components; for example, when the danger level is high, the terminal displays a red warning bar on the screen and prohibits the user from clicking the URL; when the danger level is medium, the terminal displays a yellow icon and allows the user to continue operating after confirmation. The terminal can also trigger system notifications, vibrations, or sound alerts.

[0372] Input: A response message from the server.

[0373] Output: Visual feedback on the terminal screen (text prompts, color changes, icon display) and possible sound or vibration notifications.

[0374] Step 14: Users follow the prompts on the terminal to perform subsequent operations.

[0375] Users can view the risk warnings and alerts returned by the server on the terminal interface and decide whether to continue accessing the link, delete the SMS, or mark it as spam based on their own needs. Users can click buttons such as "Back," "Delete," and "Continue Access (Known Risks)," and the terminal will submit feedback to the server or terminate the relevant operation accordingly.

[0376] Input: The overall confidence assessment results and control elements (buttons, links) displayed on the terminal.

[0377] Output: User operation commands, such as canceling access, confirming deletion, or reporting false alarms. These commands can be sent back to the server by the terminal for subsequent learning and system optimization.

[0378] Application Example 2 The process flow corresponding to the specific processing in Use Case 2 will be described below. The various parts of the system described below are implemented by the data processing device 12 and the intelligent device 14. In addition, the data processing device 12 is referred to as the "server" and the intelligent device 14 is referred to as the "terminal".

[0379] In existing electronic communication security technologies, servers typically rely on fixed rules or a single machine learning model to statically assess the risk level of emails or other electronic communications. This approach suffers from several problems: First, server analysis of the text often remains at the level of keyword matching or simple feature classification, failing to leverage the contextual understanding capabilities of generative artificial intelligence models. It cannot dynamically generate appropriate analysis strategies and warning statements based on different communication scenarios, making it difficult to accurately and promptly identify complex phishing techniques and sophisticated fraudulent communications. Second, servers often only consider the objective characteristics of content and identification information during risk assessment, failing to incorporate subjective information such as the user's emotional state and interactive behavior. This results in the inability to reflect the true risks posed by the same communication to users under different psychological states, leading to over-warning or under-warning issues. Third, servers lack a systematic utilization of historical user responses, failing to feed back actual user reactions to warnings into models and parameters. This hinders the development of adaptive security strategies that evolve over time, making it difficult to further reduce false positive and false negative rates. In summary, how to improve the overall accuracy, personalization, and interpretability of electronic communication risk assessment by improving the collaborative processing flow between the server side and the user terminal side within the computer architecture, introducing a generative artificial intelligence model-driven prompt generation mechanism, an emotion recognition-driven dynamic risk adjustment mechanism, and a parameter adaptive update mechanism based on historical behavior, has become an urgent technical issue to be addressed in this field.

[0380] The specific processing performed by the specific processing unit 290 of the data processing apparatus 12 in Application Example 2 is achieved by the following means.

[0381] In this invention, the server includes a module for acquiring electronic communication text data and performing natural language processing to extract grammatical structures, lexical features, and expressions related to risk assessment; a module for extracting identification information and comparing it with an information set to calculate a risk index; a module for automatically generating prompt statements based on the text data and the risk index and inputting them into a generative artificial intelligence model, and receiving the risk assessment results from the model to update the risk; a module for receiving image data and voice data from a user terminal and performing emotion recognition to obtain the user's emotional state; a module for fusing the emotional state with the updated risk and dynamically adjusting the electronic communication risk; and a module for recording the user's responses on the terminal and adaptively updating the risk calculation and prompt statement generation parameters based on historical information. This allows for an end-to-end, multi-source information fusion risk assessment process within the computer. This enables the server to not only utilize generative artificial intelligence models for in-depth analysis of complex semantics and context, but also to dynamically adjust the risk level based on the user's real-time emotional state. Furthermore, by continuously collecting actual user operation results, the assessment model and prompting strategies can be iteratively optimized. This significantly improves the accuracy, robustness, and actual protection effect for users in electronic communication security assessments at the computer technology level.

[0382] "Electronic communication" refers to digital information units transmitted over a network, including emails, instant messages, text messages, and other text-based communication content.

[0383] "Text data" refers to the text portion of electronic communication that carries the main semantic information, excluding protocol header information used only for transmission control or routing purposes.

[0384] Natural Language Processing (NLP) refers to the technical process of performing word segmentation, syntactic analysis, and semantic analysis on human language text in electronic computing devices to obtain structured language features.

[0385] "Syntactic structure analysis" refers to the process of analyzing the syntactic relationships between words and phrases in a text in order to determine syntactic components such as subject, verb, and object, as well as dependency relationships.

[0386] "Lexical structure analysis" refers to the process of classifying, labeling, and identifying relationships among words in a text in order to obtain information at the lexical level, such as parts of speech, word frequency, and collocations.

[0387] "Hazard assessment related expressions" refer to words, phrases, or sentence patterns in the main text data that are related to fraud, phishing, impersonation, and other harmful behaviors. These expressions have an indicative role in judging the hazard level of electronic communications.

[0388] "Identification information" refers to string information used to identify electronically associated entities or resources, including network addresses, contact information, account identifiers, etc.

[0389] "Information set" refers to a pre-built and stored structured data group in a storage device, used to record credible and / or harmful information related to identification information for comparison and evaluation.

[0390] "Reliability" refers to the degree to which the source or target of the identified information is reliable and secure based on a set of information. It is usually expressed in the form of a rating or numerical value.

[0391] "Risk index" refers to a parameter or score used to quantitatively represent the degree of potential risk in electronic communications. This parameter can be calculated by combining multiple characteristics.

[0392] "Generative AI model" refers to an AI model that can automatically generate text output based on input prompts. This model acquires language understanding and generation capabilities through large-scale data training.

[0393] "Prompt statements" refer to the instructional text input into a generative artificial intelligence model, which specifies the tasks the model should perform, the content it should focus on, and the output format.

[0394] "Risk assessment processing" refers to the process by which generative artificial intelligence models judge and analyze whether electronic communication is harmful and the degree of risk based on prompts and input data.

[0395] "Image data" refers to digital images or image sequences that contain facial or posture information of a user and are collected by a user terminal.

[0396] “Voice data” refers to digital audio signals collected by user terminals that contain information about the user’s speech or vocalization.

[0397] "Emotion recognition processing" refers to the analytical process of inferring the category and intensity of a user's emotions based on image data and / or voice data.

[0398] "Emotional state" refers to the type of a user's current psychological emotion obtained through emotion recognition processing, such as anger, fear, tension, joy, calmness, etc.

[0399] "Negative emotions" refer to emotional types that may be associated with stress, worry, or negative experiences, including emotional states such as anger, fear, anxiety, and sadness.

[0400] "Affirmative feelings" refer to emotional types that are usually associated with security, satisfaction, or positive experiences, including emotional states such as joy, peace of mind, and relaxation.

[0401] "Dynamic adjustment" refers to the process of updating the calculated risk index in real time or periodically based on new input information (including emotional state, historical behavior, etc.).

[0402] "Warning content" refers to informational text used to alert users to potential risks in electronic communications, including explanations of the causes of the risks and preventative advice.

[0403] "User interface display style" refers to the visual layout and presentation of warning content or danger information on the user terminal, including colors, icons, font size, pop-up window format, etc.

[0404] "Notification information" refers to a data unit sent by a server to a user terminal to indicate the dangers of electronic communication and suggest behaviors. It usually contains text and structured fields.

[0405] "User action content" refers to the record of user interactions on the terminal regarding electronic communication, including actions such as opening, deleting, marking as junk, and continuing to access links.

[0406] "Actual response outcome" refers to the final handling method adopted by the user in response to electronic communication, derived from the user's operation content, and is used to reflect the user's response to the danger warning.

[0407] "Historical information" refers to data such as risk assessment results, notification information, and actual user responses that have been stored cumulatively over a period of time, which are used for subsequent analysis and model updates.

[0408] "Risk Calculation Processing" refers to the process by which a server quantifies the degree of risk in electronic communications based on various features such as text data, identification information, and model output.

[0409] "Parameter update" refers to the operation of adjusting the weights, thresholds or other control parameters used in the hazard calculation and prompt statement generation processes based on historical information and statistical results of new data.

[0410] A "machine learning model" is a computational model that, by training on both harmful and non-harmful communication data, can output a probability of harmfulness or a classification result based on input features.

[0411] "Harmful probability value" refers to the quantitative result of a machine learning model on the likelihood that electronic communication belongs to a harmful category, usually represented by a value between 0 and 1.

[0412] "Explanatory information" refers to the part that constitutes the content of the prompt statement and is used to provide context to the generative artificial intelligence model, including text data, identification information evaluation results, harmfulness probability values, and emotional states.

[0413] "Explanatory text" refers to user-oriented natural language explanations generated by generative artificial intelligence models based on prompt statements, used to explain the basis for risk assessment and provide prevention suggestions.

[0414] In this invention, the server acts as the core computing node, implementing the functions of each module through an electronic processor, memory, and network interface. In one specific embodiment, the server uses general-purpose computer hardware, such as a multi-core central processing unit, main memory, and solid-state storage. The server runs an application program on an operating system, implemented using a general-purpose programming language, such as a scripting language as the server-side logic implementation language. Within the application program, the server invokes natural language processing libraries, machine learning libraries, and deep learning frameworks. For example, it invokes a natural language processing library (e.g., an alternative word segmentation and syntactic analysis library), a traditional machine learning library (e.g., an alternative classification algorithm library), and a deep learning framework (e.g., an alternative tensor computation framework) to achieve text feature extraction and model inference functions.

[0415] In one embodiment of the present invention, the server constructs a unified data structure for the electronic communication text data. The server represents each electronic communication as a data record containing multiple fields, including a sender identifier field, a recipient identifier field, a timestamp field, a title field, a text field, an attachment field, and a metadata field. The server establishes an index table for this data record in a storage device for rapid retrieval and model training.

[0416] The server performs multi-level feature extraction on the text field within the natural language processing module. First, the server segments the text into sentences, assigning a sentence number to each sentence and storing it in a sentence table. Next, the server performs word segmentation and part-of-speech tagging on each sentence, representing each word using a vector structure containing the word literal string, part-of-speech tag, and its position index within the sentence. The server further performs dependency parsing, recording the parent node and dependency relationship tags for each word, thus forming a directed tree structure. The server organizes these results into structured features, such as using sparse matrices or dense vector representations, for subsequent use in machine learning and generative artificial intelligence models.

[0417] The server employs a character-level parsing combined with regular expression matching in its identification information extraction module. The server performs pattern matching in the text and metadata fields to identify information such as network addresses, contact information, and account identifiers. Each piece of identification information is stored in an identification information table, which includes the identification information string, type identifier (network address, phone number, mailing address, etc.), sentence number, and offset position in the original text. The server then performs database comparison processing on the identification information. It accesses a pre-built set of security information, which can be implemented using a relational database and includes a trusted list and a risk list. The server performs a query operation for each piece of identification information, assigns a trust score and category label based on the matching results, and writes the results back to the identification information table.

[0418] The server fuses multi-source features in its hazard index calculation module. It obtains text structure features from the natural language processing module, credibility features from the identification information extraction module, and time and domain name features from metadata. The server then uses a machine learning model to vectorize and combine these features. In one implementation, the server uses a linear classification model, such as logistic regression; in another, it uses a non-linear model, such as a gradient boosting tree model or a multilayer perceptron. During training, the server uses a training set consisting of harmful and non-harmful communication samples, employing cross-entropy as the loss function to optimize the model parameters. The server updates the model weights using gradient descent or its variants (e.g., adaptive learning rate optimization algorithms). During training, the server performs feature standardization and regularization to reduce overfitting and improve generalization ability. In the inference phase, the server calculates the harmfulness probability of newly arrived electronic communications and maps this probability to a standardized hazard index.

[0419] The server constructs prompt statements in the generative AI model integration module. The server combines the current electronic communication text data, the credibility evaluation results of the identified information, the harmfulness probability value output by the machine learning model, and the emotional state information uploaded by the terminal to form a contextual description. The server inserts this description into a predefined prompt template to form prompt statements for the generative AI model. The server can use prompt statements in the following forms: "Please analyze the following email for potential phishing risks, paying particular attention to any suspicious links: Subject: Account Abnormality, Please Take Immediate Action Body: Dear User, your account has been detected to have an abnormal login. Please visit http: / / example-phishing.com within 24 hours to verify your identity; otherwise, your account will be permanently frozen." In another example, the server generates the following prompt: "Please explain the potential risks in the following email from the perspective of an average user, and provide three preventative suggestions: ... (email body)..." The server can also generate prompts that include emotional states, such as: "The user exhibited 'fear' while reading this email. Based on the email content and this emotion, please provide a brief security tip, reminding the user of potential risks and suggesting safer behaviors." The server sends prompts to an external generative AI model service via a network interface. This generative AI model is implemented on the server side using a deep neural network structure, such as a multi-layer encoder-decoder structure based on self-attention. During training, the model adjusts its parameters by minimizing the loss function between the predicted and reference texts (e.g., cross-entropy loss) and updates the multi-layer weights using backpropagation. During inference, the server sends prompts and receives the analysis and suggestion texts generated by the model. The server parses the returned text, extracts the risk conclusions and their rationale from the model, and records this information in an analysis results table, which can be used as a basis for risk adjustment and explanation to the user.

[0420] In this invention, the terminal is responsible for collecting user physiological and behavioral data and performing some local analysis. In one embodiment, the terminal is a smart mobile device with a camera and microphone, running a local application to access the camera and microphone. The terminal uses an image processing library to extract facial regions from image frames captured by the camera. After detecting a face, the terminal uses a pre-trained facial expression recognition model to infer the corresponding emotion category. In one embodiment, this model is a convolutional neural network structure, containing multiple convolutional layers, pooling layers, and fully connected layers. The terminal uses the model to perform forward propagation on the image input and outputs the probability distribution of each emotion category. Optionally, the terminal extracts acoustic features from the speech data, such as Mel-frequency cepstral coefficients, and uses an emotion classification model to perform emotion recognition on the speech. The terminal combines the facial expression emotion results with the speech emotion results to form a unified emotion state vector, which is then reported to a server via a network interface or fused locally with danger information.

[0421] The terminal receives hazard indicators, risk levels, and explanatory text from the server in its dynamic hazard adjustment and display module. The terminal combines the server-provided basic hazard level with a local emotional state vector, adjusting it using rule-based mapping or a lightweight model. For example, when the server indicates a high hazard level and the terminal detects fear or anxiety in the user, it raises the display level to "extremely high risk" and displays a full-screen warning interface with a highlighted color. When the server indicates a moderate hazard level and the user is calm, the terminal displays a milder prompt. By adjusting the interface style, pop-up frequency, and interaction flow, the terminal makes the warnings more perceptible and actionable in the user's current psychological state.

[0422] Users interact with the system through a terminal interface. On the terminal, users can view email content, risk levels, and explanatory text provided by the generative artificial intelligence model. Users can choose to delete the email, mark it as spam, continue reading without clicking the link, or access the link despite confirming the risk. User operation information is recorded as a structured operation log on the terminal side, including the electronic communication identifier, timestamp, the user's selected operation type, and the current risk level indication. The terminal reports the operation log to the server, which writes it into a historical database for subsequent model selection and parameter adjustment.

[0423] The server performs statistical analysis of user behavior in the historical information utilization module. It extracts records from the historical database within a specific time window and calculates the actual deletion rate, report rate, and false click rate for users under different risk levels. Based on the statistical results, the server adjusts the risk classification threshold and feature weights; for example, it lowers the high-risk threshold to reduce false negatives or adjusts the weights of key phrases to reduce false positives. The server can also construct personalized parameter sets based on long-term user behavior characteristics, employing different prompting strategies for different user groups, thus forming an adaptive configuration mechanism at the computer level that caters to individual user differences.

[0424] In this invention, the server employs the aforementioned modules to collaboratively implement a processing flow that differs from traditional technologies. Traditional systems typically rely on fixed rule tables for hazard assessment. In this invention, the server deeply integrates natural language processing with machine learning models, using structured text features, credibility of identified information, and sentiment vectors as joint feature inputs, and leveraging an end-to-end probabilistic model for hazard estimation. In its interaction with the generative AI model, the server uses context-constrained prompts instead of simple keyword queries, enabling the model to output explanatory text with causal explanations and fine-grained suggestions. By incorporating user sentiment states and historical response results, the server integrates human-computer interaction into the computational pathway, achieving a non-linear, dynamic adjustment process. These technical features collectively improve the system's detection accuracy and robustness in complex scenarios, reduce the transmission and storage of redundant warnings in network and storage resources, thereby improving overall computational efficiency and communication load.

[0425] In another implementation, the server can utilize different deep network structures and training methods. For example, the server can employ a bidirectional sequence encoder to encode the email body, and then use an attention layer at the top to aggregate important word features. This structure allows the server to automatically focus on local segments highly correlated with fraud or phishing patterns. During training, the server can incorporate data augmentation methods, such as randomly replacing non-keywords or randomly shuffling sentence order, to improve the model's robustness against adversarial examples. The server can define a multi-task loss function to simultaneously optimize both "harmfulness prediction" and "risk category interpretation," thereby enhancing the interpretability of the generative AI model's output.

[0426] In another implementation, the terminal can offload some computation locally, such as executing a lightweight sentiment classification model locally and caching several near-real-time danger thresholds locally, thus enabling it to quickly issue initial warnings even with significant network latency. The terminal can choose to perform certain steps locally or on the server side depending on network conditions to reduce communication load and improve user experience.

[0427] Users can access the system through various terminals in different implementation forms, such as desktop terminals or wearable devices. Users receive a unified risk warning interface on these terminals, and the system maintains consistent risk calculation and display logic across different hardware platforms through common data structures and protocols.

[0428] Through the aforementioned embodiments, the system of this invention achieves a technical solution within the computer that features multi-module collaboration, a clearly defined data structure, and specific processing steps. The server, through specific feature construction, model structure, and parameter update strategies, achieves efficient and precise risk assessment; the terminal, through real-time emotion recognition and dynamic interface adjustment, converts complex internal calculation results into warning information that is understandable and operable to the user; and the user provides feedback to the system through natural operational behaviors, enabling the model to continuously optimize in a closed loop. Therefore, this invention not only automates human review work but also substantially improves computer technology itself on multiple levels, including computational accuracy, processing speed, resource utilization, and interactive intelligence.

[0429] use Figure 14 The processing procedure is explained.

[0430] Step 1: The server receives and stores electronic communication data.

[0431] The server's input consists of electronic communication data sent by the terminal, including fields such as sender identifier, recipient identifier, title, timestamp, body text, and header metadata.

[0432] The server receives the data packet through the network interface and parses the data packet in the application, mapping each field to the internally defined record structure.

[0433] The server uses a data storage device to write the parsed electronic communication records into a database table and assigns a unique identifier to each record for subsequent association processing.

[0434] The server's output consists of electronic communication records stored in the database, along with their corresponding unique record identifiers.

[0435] Step 2: The server performs natural language preprocessing on the main text.

[0436] The server's input consists of the text fields and record identifiers from the electronic communication records stored in step 1.

[0437] The server calls a natural language processing library to segment the text into sentences, breaking down the long text into multiple sentences and assigning a sentence number to each sentence.

[0438] The server performs word segmentation and part-of-speech tagging on each sentence, representing each word using a structure containing the word string, part-of-speech tag, and position index.

[0439] The server further performs dependency parsing to determine the parent node and dependency relation tags for each word, and constructs a tree structure based on sentences.

[0440] The server stores these analysis results in a structured feature table associated with the record identifier, in the form of a list or matrix.

[0441] The server outputs a list of sentences, a list of words, and syntactic features corresponding to each electronic communication.

[0442] Step 3: The server extracts the identification information and performs type labeling.

[0443] The server's input consists of the main text obtained in step 2 and the header metadata.

[0444] The server uses character parsing and regular expression matching to perform pattern searches on the body and header, identifying information strings such as network addresses, phone numbers, and mailing addresses.

[0445] The server assigns a type label (such as "network address", "telephone number", "mailing address") to each piece of identification information and records its position in the original text (sentence number, offset).

[0446] The server writes all identification information into the identification information table and associates it with the corresponding electronic communication record identifier.

[0447] The server output is a list of identification information, including identification information strings, type labels, and location information.

[0448] Step 4: The server will compare the identified information with the information set and calculate the credibility.

[0449] The server's input consists of the identification information list generated in step 3 and a set of information pre-stored in the database (including a trusted list and a risk list).

[0450] For each piece of identification information, the server constructs a query condition access information set table to determine whether the identification information appears in trusted entries, risky entries, or neither.

[0451] The server assigns a credibility score and category label to the identification information based on the query results, such as trustworthy, high-risk, unknown, etc.

[0452] The server writes these confidence scores back to the identification information table and provides input features for subsequent risk calculations.

[0453] The server outputs a list of identification information, including confidence scores and category labels.

[0454] Step 5: The server builds text and metadata features and calls a machine learning model to calculate the probability of harmfulness.

[0455] The server's input consists of the text structure features obtained in step 2, the credibility features of the identification information obtained in step 4, and metadata features (such as the sending domain name, sending time, etc.).

[0456] The server uses a feature extraction module to convert text fields into numerical features, such as constructing text vectors through word frequency statistics or TF-IDF methods.

[0457] The server appends the credibility score of the identified information as a numerical feature to the text vector, and encodes metadata such as time period information and domain name type as numerical or categorical features.

[0458] The server inputs the combined feature vector into a trained machine learning model (such as logistic regression, gradient boosting tree, or multilayer perceptron) to perform forward computation and obtain the probability value that the electronic communication belongs to the harmful category.

[0459] The server normalizes or linearly maps the probability value, converting it into an initial risk score.

[0460] The server outputs the probability value of harm and the initial risk score for each electronic communication.

[0461] Step 6: The server detects unnatural language and adds risk indicators.

[0462] The server's input consists of the syntactic structure features generated in step 2 and the original text.

[0463] The server identifies grammatical errors, sentences that do not conform to common word order, and abnormally repeated keywords by analyzing dependency structures and part-of-speech sequences.

[0464] The server defines scoring rules for each type of unnaturalness, such as the number of grammatical errors, the density of odd collocations, and the degree of keyword stuffing, and calculates an unnaturalness score for the language.

[0465] The server incorporates the unnaturalness score of the language as an additional feature into the risk calculation, and adjusts the initial risk score according to the preset weights.

[0466] The server's output includes a basic risk index corrected for language unnaturalness, along with corresponding intermediate features.

[0467] Step 7: The server generates prompts for generative artificial intelligence models.

[0468] The server's input consists of the hazard index, main text, identification information credibility results, and metadata summary obtained in steps 5 and 6.

[0469] The server selects a prompt template based on the current task objective, such as "Assess phishing risk" or "Generate user instructions".

[0470] The server fills the prompt template with the main text, key identification information and its credibility, probability of harm, and optional contextual information to generate a coherent natural language prompt.

[0471] The server can generate the following types of prompts: "Please analyze the following email for potential phishing risks, paying particular attention to any suspicious links: Subject: Account Abnormality, Please Take Immediate Action Body: Dear User, your account has been detected to have an abnormal login. Please visit http: / / example-phishing.com within 24 hours to verify your identity; otherwise, your account will be permanently frozen." The server can also generate interpreted prompts: "Please explain the potential risks in the following email from the perspective of an average user, and provide three preventative suggestions: ... (email body)..." The server's output is at least one prompt text associated with the current electronic communication.

[0472] Step 8: The server invokes a generative artificial intelligence model and obtains a hazard assessment and explanatory text.

[0473] The server's input is the prompt statement generated in step 7.

[0474] The server sends prompts to the generative artificial intelligence model service via a network interface, using them as input to the model.

[0475] The server triggers the text generation process at the generative artificial intelligence model. Inside the model, the prompts are encoded through a multi-layer self-attention network, and the analysis text and suggestion text are generated through a decoding layer.

[0476] The server receives the complete output text returned by the model and performs structured parsing on the text, such as identifying the judgment conclusion of "whether it is harmful communication", "a list of risk reasons" and "several prevention suggestions".

[0477] The server writes the parsing results into the analysis results table as supplementary information for the risk index.

[0478] The server output is a record of analysis results, including model determinations and explanations.

[0479] Step 9: The server integrates machine learning output with generative artificial intelligence model output to update the final danger level.

[0480] The server's input consists of the basic hazard indices calculated in steps 5 and 6, and the generative artificial intelligence model analysis results parsed in step 8.

[0481] The server quantifies and maps the conclusions in the model analysis results. For example, when the model text explicitly states "highly suspected phishing," this conclusion is converted into a numerical weighting factor.

[0482] The server synthesizes the basic risk index and the weights derived from the text conclusions according to a predefined fusion formula to obtain an updated comprehensive risk score.

[0483] The server classifies risk levels based on a comprehensive risk score, such as safe, suspicious, high-risk, and extremely high-risk, and records the risk level and score together in the database.

[0484] The server outputs the final hazard score and corresponding risk level for each electronic communication.

[0485] Step 10: The terminal collects user emotion-related data and performs emotion recognition.

[0486] Terminal input is the user's action event of opening a certain electronic communication on the terminal.

[0487] With the user's authorization, the terminal activates the camera and microphone to collect image and voice data from the user during the reading process.

[0488] The terminal performs face detection on the image data, extracts image frames containing face regions, and inputs them into a pre-trained expression recognition model to obtain the probability distribution of each emotion category.

[0489] The terminal can optionally extract acoustic features from the speech data and input them into the speech emotion model to obtain the speech emotion probability distribution.

[0490] The terminal performs weighted fusion of facial expression and voice emotion results to generate a unified emotion state label and intensity value.

[0491] The terminal outputs emotional state information in a structured form, such as emotion category and intensity coefficient.

[0492] Step 11: The terminal receives the risk information from the server and adjusts the risk level locally.

[0493] The terminal's input consists of the final danger score and risk level sent by the server in step 9, and the emotional state information generated in step 10.

[0494] The terminal combines emotional state with server risk level according to pre-set mapping rules. For example, when the emotion is fear or anger and the intensity is high, the displayed risk level is increased.

[0495] The terminal performs simple numerical calculations to adjust the risk rating and selects different display strategies accordingly, such as adjusting the color, font size, and pop-up format.

[0496] The terminal uses the adjusted hazard level as input for the subsequent warning display module.

[0497] The terminal outputs locally adjusted risk information that matches the current user state.

[0498] Step 12: The terminal displays warning messages and explanatory text provided by a generative artificial intelligence model to the user.

[0499] The terminal input consists of the locally adjusted risk information obtained in step 11, as well as the explanatory text and suggestions provided by the server in steps 8 and 9.

[0500] The terminal generates a warning area in the user interface, displaying a brief description of the risk level, hazard score, and a summary of the main risk causes.

[0501] When needed, the terminal provides an "expand details" function to display explanatory text generated by the generative artificial intelligence model, including why it is identified as risky communication, which links are suspicious, and recommended preventive measures.

[0502] The terminal selects different interactive elements based on the level of danger. For example, in high-risk situations, a secondary confirmation dialog box is added, requiring the user to explicitly choose whether to continue accessing the link.

[0503] The terminal outputs a comprehensive security alert interface that is visually presented to the user.

[0504] Step 13: Users perform specific operations according to the prompts on the terminal.

[0505] The user's input consists of the hazard information, explanatory text, and optional operation buttons displayed on the terminal interface.

[0506] Users can choose to delete the electronic communication, mark it as spam, ignore the warning and continue reading, or access the links within it after confirming the risk, based on their own judgment and the content of the prompts.

[0507] Users trigger selected commands on the terminal by clicking, swiping, or touching.

[0508] The user's output consists of processing decisions for the current electronic communication, which are stored on the terminal in the form of operation logs.

[0509] Step 14: The terminal records user actions and sends feedback to the server.

[0510] The terminal input consists of the operation type selected by the user in step 13, the current electronic communication identifier, and the danger level information displayed at that time.

[0511] The terminal compiles the above information into an operation log, which includes communication identifier, timestamp, operation type, displayed risk level, and related parameters.

[0512] The terminal reports operation logs to the server via the network interface as part of historical behavior data.

[0513] The terminal outputs operation log data transmitted to the server.

[0514] Step 15: The server uses historical information to update the risk calculation parameters and the prompt statement generation strategy.

[0515] The server's input consists of the user operation logs received in step 14 and previously stored historical risk assessment records.

[0516] The server performs statistical analysis on the data within a certain time window, calculating the deletion rate, error rate, and user feedback trends within each risk level range.

[0517] The server adjusts the feature weights and threshold parameters in the risk calculation module based on statistical results, for example, by increasing the weight of certain emerging fraud features or decreasing the weight of features with a high number of false alarms.

[0518] The server simultaneously analyzes the relationship between the explanatory text output by the generative artificial intelligence model and the user's response to determine which prompt templates are more effective in prompting the user to take safe actions.

[0519] Based on this, the server modifies the prompt generation template and filling strategy, making the subsequently generated prompts clearer and easier to understand in terms of content structure and wording, and more persuasive.

[0520] The server outputs updated model parameters and prompt generation rules, which will be applied in subsequent electronic communication processing, thus forming a closed loop of continuous optimization.

[0521] The specific processing unit 290 sends the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires sound representing user input regarding the result of the specific processing. The control unit 46A sends the sound data representing user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the sound data.

[0522] Data generation model 58 is a so-called generative AI (Artificial Intelligence). Examples of data generation models 58 include ChatGPT (registered trademark) (accessible via the internet (URL: https: / / openai.com / blog / chatgpt)). Data generation model 58 is obtained through deep learning on a neural network. Input to data generation model 58 are prompt words containing instructions, and inference data such as sound data representing sound, text data representing text, and image data representing images (e.g., still image data or animation data). Data generation model 58 infers from the input inference data based on the instructions represented by the prompt words and outputs the inference result in one or more data forms, such as sound data, text data, and image data. Data generation model 58 includes, for example, text generation AI, image generation AI, and multimodal generation AI. Here, inference refers to, for example, analysis, classification, prediction, and / or induction. The specific processing unit 290 performs the aforementioned specific processing while using data generation model 58. The data generation model 58 can also be a model finely tuned to output inference results from prompts that do not contain instructions. In this case, the data generation model 58 can output inference results based on prompts that do not contain instructions. The data processing apparatus 12, etc., includes various data generation models 58, including AI other than the generation AI. AI other than the generation AI can be, for example, linear regression, logistic regression, decision trees, random forests, support vector machines (SVM), k-means clustering, convolutional neural networks (CNN), recurrent neural networks (RNN), generative adversarial networks (GAN), or Naive Bayes, and can perform various processes, but is not limited to this example. Furthermore, the AI ​​can also be an AI agent. Furthermore, when the processing of the above-mentioned parts is performed by AI, the processing can be performed partially or entirely by AI, but is not limited to this example. Furthermore, the processing performed by the AI ​​including the generation AI can be replaced by processing in the rule base, and the processing in the rule base can also be replaced by processing performed by the AI ​​including the generation AI.

[0523] Furthermore, the processing of the aforementioned data processing system 10 is performed by the specific processing unit 290 of the data processing device 12 or the control unit 46A of the smart device 14, but it can also be performed by both the specific processing unit 290 of the data processing device 12 and the control unit 46A of the smart device 14. Additionally, the specific processing unit 290 of the data processing device 12 acquires or collects information required for processing from the smart device 14 or external devices, and the smart device 14 acquires or collects information required for processing from the data processing device 12 or external devices.

[0524] For example, the collection unit is implemented by the control unit 46A of the smart device 14 or the specific processing unit 290 of the data processing device 12. For example, the acquisition unit uses the camera 42 or communication I / F 44 of the smart device 14 to acquire step data, which is then processed by the specific processing unit 290 of the data processing device 12. For example, the analysis unit is implemented by the specific processing unit 290 of the data processing device 12, which analyzes the data from the collection unit and the acquisition unit. For example, the generation unit is implemented by the specific processing unit 290 of the data processing device 12, which uses a generation AI to generate a menu. For example, the serving unit is implemented by the output device 40 of the smart device 14 or the specific processing unit 290 of the data processing device 12, which provides the generated menu to the user. The correspondence between each unit and the device or control unit is not limited to the above examples and various changes can be made.

[0525] In the above embodiments, examples of specific processing by the data processing device 12 are given, but the technology disclosed herein is not limited to this, and specific processing may also be performed by the smart device 14.

[0526] Second Implementation Method Figure 3 An example of the configuration of the data processing system 210 according to the second embodiment is shown.

[0527] like Figure 3 As shown, the data processing system 210 includes a data processing device 12 and smart glasses 214. A server can be cited as an example of the data processing device 12.

[0528] The data processing apparatus 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" as understood in this disclosure. The computer 22 includes a processor 28, RAM 30, and memory 32. The processor 28, RAM 30, and memory 32 are connected to a bus 34. Furthermore, the database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0529] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, RAM 48, and memory 50. The processor 46, RAM 48, and memory 50 are connected to a bus 52. Furthermore, the microphone 238, speaker 240, camera 42, and communication I / F 44 are also connected to the bus 52.

[0530] Microphone 238 receives instructions from user 20 by receiving sounds emitted by user 20. Microphone 238 captures sounds emitted by user 20 and converts the captured sounds into sound data, which is then output to processor 46. Speaker 240 outputs sound according to instructions from processor 46.

[0531] Camera 42 is a small digital camera equipped with an optical system such as a lens, aperture and shutter, and imaging elements such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, which captures images of the user 20's surroundings (e.g., the field of view defined by an angle equivalent to the field of vision of an average healthy person).

[0532] Communication I / F44 is connected to network 54. Communication I / F44 and 26 are responsible for the transmission and reception of various information between processor 46 and processor 28 via network 54. The transmission and reception of various information between processor 46 and processor 28 using communication I / F44 and 26 is performed in a secure state.

[0533] Figure 4 This illustrates an example of the main functions of the data processing device 12 and the smart glasses 214. For example... Figure 4 As shown, in the data processing device 12, specific processing is performed by the processor 28. The specific processing program 56 is stored in the memory 32.

[0534] The specific processing program 56 is an example of a "program" involved in the technology of this disclosure. The processor 28 reads the specific processing program 56 from the memory 32 and executes the read specific processing program 56 on the RAM 30. Specific processing is implemented by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.

[0535] The memory 32 stores a data generation model 58 and an emotion-specific model 59. The data generation model 58 and the emotion-specific model 59 are used by the specific processing unit 290. The specific processing unit 290 can use the emotion-specific model 59 to infer the user's emotions and perform specific processing based on the user's emotions. In the emotion inference function (emotion-specific function) using the emotion-specific model 59, various inferences and predictions related to the user's emotions are performed, including inferences and predictions of the user's emotions, but this is not limited to this example. Furthermore, emotion inference and prediction may also include, for example, emotion analysis (parsing).

[0536] In the smart glasses 214, the processor 46 performs reception and output processing. The memory 50 stores the reception and output program 60. The processor 46 reads the reception and output program 60 from the memory 50 and executes the read reception and output program 60 on the RAM 48. The reception and output processing is implemented by the processor 46 operating as a control unit 46A according to the reception and output program 60 executed on the RAM 48. Furthermore, the smart glasses 214 has the same data generation model and emotion-specific model as the data generation model 58 and the emotion-specific model 59, and these models can also be used to perform the same processing as the specific processing unit 290.

[0537] Next, the specific processing of the specific processing unit 290 of the data processing device 12 will be described. Each part of the system described below is implemented by the data processing device 12 and the smart glasses 214. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".

[0538] Example 1 The process is the same as that of the specific process described in Embodiment 1 in the first embodiment above, so the description is omitted.

[0539] Application Example 1 The process is the same as that in the specific processing described in Application Example 1 of the first embodiment above, so the description is omitted.

[0540] Example 2 The process is the same as that of the specific process in Embodiment 2 described in the first embodiment above, so the description is omitted.

[0541] Application Example 2 The process is the same as that in the specific processing described in Application Example 2 of the first embodiment above, so the description is omitted.

[0542] The specific processing unit 290 sends the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A outputs the result of the specific processing to the speaker 240. The microphone 238 acquires sound input representing the user's input regarding the result of the specific processing. The control unit 46A sends the sound data representing the user's input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the sound data.

[0543] Data generation model 58 is a so-called generative AI (Artificial Intelligence). Examples of data generation models 58 include ChatGPT (registered trademark) (accessible via the internet (URL: https: / / openai.com / blog / chatgpt)). Data generation model 58 is obtained through deep learning on a neural network. Input to data generation model 58 are prompt words containing instructions, and inference data such as sound data representing sound, text data representing text, and image data representing images (e.g., still image data or animation data). Data generation model 58 infers from the input inference data based on the instructions represented by the prompt words and outputs the inference result in one or more data forms, such as sound data, text data, and image data. Data generation model 58 includes, for example, text generation AI, image generation AI, and multimodal generation AI. Here, inference refers to, for example, analysis, classification, prediction, and / or induction. The specific processing unit 290 performs the aforementioned specific processing while using data generation model 58. The data generation model 58 can also be a model finely tuned to output inference results from prompts that do not contain instructions. In this case, the data generation model 58 can output inference results based on prompts that do not contain instructions. The data processing apparatus 12, etc., includes various data generation models 58, including AI other than the generation AI. AI other than the generation AI can be, for example, linear regression, logistic regression, decision trees, random forests, support vector machines (SVM), k-means clustering, convolutional neural networks (CNN), recurrent neural networks (RNN), generative adversarial networks (GAN), or Naive Bayes, and can perform various processes, but is not limited to this example. Furthermore, the AI ​​can also be an AI agent. Furthermore, when the processing of the above-mentioned parts is performed by AI, the processing can be performed partially or entirely by AI, but is not limited to this example. Furthermore, the processing performed by the AI ​​including the generation AI can be replaced by processing in the rule base, and the processing in the rule base can also be replaced by processing performed by the AI ​​including the generation AI.

[0544] Furthermore, the processing of the aforementioned data processing system 10 is performed by the specific processing unit 290 of the data processing device 12 or the control unit 46A of the smart glasses 214, but it can also be performed by both the specific processing unit 290 of the data processing device 12 and the control unit 46A of the smart glasses 214. Additionally, the specific processing unit 290 of the data processing device 12 acquires or collects information required for processing from the smart glasses 214 or external devices, and the smart glasses 214 acquires or collects information required for processing from the data processing device 12 or external devices.

[0545] For example, the collection unit is implemented by the control unit 46A of the smart glasses 214 or the specific processing unit 290 of the data processing device 12. For example, the acquisition unit uses the camera 42 or communication I / F 44 of the smart glasses 214 to acquire step data, which is then processed by the specific processing unit 290 of the data processing device 12. For example, the analysis unit is implemented by the specific processing unit 290 of the data processing device 12, which analyzes the data from the collection unit and the acquisition unit. For example, the generation unit is implemented by the specific processing unit 290 of the data processing device 12, which uses a generation AI to generate a menu. For example, the serving unit is implemented by the speaker 240 of the smart glasses 214 or the specific processing unit 290 of the data processing device 12, which provides the generated menu to the user. The correspondence between each unit and the device or control unit is not limited to the above examples and various changes can be made.

[0546] In the above embodiments, examples of specific processing by the data processing device 12 are given, but the technology disclosed herein is not limited to this, and specific processing may also be performed by the smart glasses 214.

[0547] Third Implementation Method Figure 5 An example of the configuration of the data processing system 310 according to the third embodiment is shown.

[0548] like Figure 5 As shown, the data processing system 310 includes a data processing device 12 and a head-mounted terminal 314. A server can be cited as an example of the data processing device 12.

[0549] The data processing apparatus 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" as understood in this disclosure. The computer 22 includes a processor 28, RAM 30, and memory 32. The processor 28, RAM 30, and memory 32 are connected to a bus 34. Furthermore, the database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0550] The head-mounted terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a display 343. The computer 36 includes a processor 46, RAM 48, and memory 50. The processor 46, RAM 48, and memory 50 are connected to a bus 52. Furthermore, the microphone 238, speaker 240, camera 42, display 343, and communication I / F 44 are also connected to the bus 52.

[0551] Microphone 238 receives instructions from user 20 by receiving sounds emitted by user 20. Microphone 238 captures sounds emitted by user 20 and converts the captured sounds into sound data, which is then output to processor 46. Speaker 240 outputs sound according to instructions from processor 46.

[0552] Camera 42 is a small digital camera equipped with an optical system such as a lens, aperture and shutter, and imaging elements such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, which captures images of the user 20's surroundings (e.g., the field of view defined by an angle equivalent to the field of vision of an average healthy person).

[0553] Communication I / F44 is connected to network 54. Communication I / F44 and 26 are responsible for the transmission and reception of various information between processor 46 and processor 28 via network 54. The transmission and reception of various information between processor 46 and processor 28 using communication I / F44 and 26 is performed in a secure state.

[0554] Figure 6 This illustrates an example of the main functions of the data processing device 12 and the head-mounted terminal 314. For example... Figure 6 As shown, in the data processing device 12, specific processing is performed by the processor 28. The specific processing program 56 is stored in the memory 32.

[0555] The specific processing program 56 is an example of a "program" involved in the technology of this disclosure. The processor 28 reads the specific processing program 56 from the memory 32 and executes the read specific processing program 56 on the RAM 30. Specific processing is implemented by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.

[0556] The memory 32 stores the data generation model 58 and the emotion-specific model 59. The data generation model 58 and the emotion-specific model 59 are used by the specific processing unit 290.

[0557] In the head-mounted terminal 314, the processor 46 performs the acceptance / output processing. The memory 50 stores the acceptance / output program 60. The processor 46 reads the acceptance / output program 60 from the memory 50 and executes the read acceptance / output program 60 on the RAM 48. The acceptance / output processing is implemented by the processor 46 operating as a control unit 46A according to the acceptance / output program 60 executed on the RAM 48.

[0558] Next, the specific processing of the specific processing unit 290 of the data processing device 12 will be described. Each part of the system described below is implemented by the data processing device 12 and the head-mounted terminal 314. In the following description, the data processing device 12 will be referred to as the "server" and the head-mounted terminal 314 will be referred to as the "terminal".

[0559] Example 1 The process is the same as that of the specific process described in Embodiment 1 in the first embodiment above, so the description is omitted.

[0560] Application Example 1 The process is the same as that in the specific processing described in Application Example 1 of the first embodiment above, so the description is omitted.

[0561] Example 2 The process is the same as that of the specific process in Embodiment 2 described in the first embodiment above, so the description is omitted.

[0562] Application Example 2 The process is the same as that in the specific processing described in Application Example 2 of the first embodiment above, so the description is omitted.

[0563] The specific processing unit 290 sends the result of the specific processing to the head-mounted terminal 314. In the head-mounted terminal 314, the control unit 46A outputs the result of the specific processing to the speaker 240 and the display 343. The microphone 238 acquires sound input representing the user's input regarding the result of the specific processing. The control unit 46A sends the sound data representing the user's input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the sound data.

[0564] Data generation model 58 is a so-called generative AI (Artificial Intelligence). Examples of data generation models 58 include ChatGPT (registered trademark) (accessible via the internet (URL: https: / / openai.com / blog / chatgpt)). Data generation model 58 is obtained through deep learning on a neural network. Input to data generation model 58 includes prompt words containing instructions, and inference data such as sound data representing sound, text data representing text, and image data representing images (e.g., still image data or animation data). Data generation model 58 infers the input inference data based on the instructions represented by the prompt words and outputs the inference result in one or more data forms such as sound data, text data, and image data. Data generation model 58 includes, for example, text generation AI, image generation AI, and multimodal generation AI. Here, inference refers to, for example, analysis, classification, prediction, and / or induction. The specific processing unit 290 performs the aforementioned specific processing while using data generation model 58. The data generation model 58 can also be a model finely tuned to output inference results from prompts that do not contain instructions. In this case, the data generation model 58 can output inference results based on prompts that do not contain instructions. The data processing apparatus 12, etc., includes various data generation models 58, including AI other than the generation AI. AI other than the generation AI can be, for example, linear regression, logistic regression, decision trees, random forests, support vector machines (SVM), k-means clustering, convolutional neural networks (CNN), recurrent neural networks (RNN), generative adversarial networks (GAN), or Naive Bayes, and can perform various processes, but is not limited to this example. Furthermore, the AI ​​can also be an AI agent. Furthermore, when the processing of the above-mentioned parts is performed by AI, the processing can be performed partially or entirely by AI, but is not limited to this example. Furthermore, the processing performed by the AI ​​including the generation AI can be replaced by processing in the rule base, and the processing in the rule base can also be replaced by processing performed by the AI ​​including the generation AI.

[0565] Furthermore, the processing of the aforementioned data processing system 10 is performed by the specific processing unit 290 of the data processing device 12 or the control unit 46A of the head-mounted terminal 314, but it can also be performed by both the specific processing unit 290 of the data processing device 12 and the control unit 46A of the head-mounted terminal 314. Additionally, the specific processing unit 290 of the data processing device 12 acquires or collects information required for processing from the head-mounted terminal 314 or external devices, and the head-mounted terminal 314 acquires or collects information required for processing from the data processing device 12 or external devices.

[0566] For example, the collection unit is implemented by the control unit 46A of the head-mounted terminal 314 or the specific processing unit 290 of the data processing device 12. For example, the acquisition unit uses the camera 42 or communication I / F 44 of the head-mounted terminal 314 to acquire step data, which is then processed by the specific processing unit 290 of the data processing device 12. For example, the analysis unit is implemented by the specific processing unit 290 of the data processing device 12 to analyze the data from the collection unit and the acquisition unit. For example, the generation unit is implemented by the specific processing unit 290 of the data processing device 12 to generate a menu using a generation AI. For example, the serving unit is implemented by the speaker 240 and display 343 of the head-mounted terminal 314 or the specific processing unit 290 of the data processing device 12 to provide the generated menu to the user. The correspondence between each unit and the device or control unit is not limited to the above examples and various changes can be made.

[0567] In the above embodiments, examples of specific processing by the data processing device 12 are given, but the technology disclosed herein is not limited to this, and specific processing may also be performed by the head-mounted terminal 314.

[0568] Fourth Implementation Method Figure 7 An example of the configuration of the data processing system 410 according to the fourth embodiment is shown.

[0569] like Figure 7 As shown, the data processing system 410 includes a data processing device 12 and a robot 414. A server can be cited as an example of the data processing device 12.

[0570] The data processing apparatus 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" as understood in this disclosure. The computer 22 includes a processor 28, RAM 30, and memory 32. The processor 28, RAM 30, and memory 32 are connected to a bus 34. Furthermore, the database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0571] Robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a controlled object 443. The computer 36 includes a processor 46, RAM 48, and memory 50. The processor 46, RAM 48, and memory 50 are connected to a bus 52. Furthermore, the microphone 238, speaker 240, camera 42, controlled object 443, and communication I / F 44 are also connected to the bus 52.

[0572] Microphone 238 receives instructions from user 20 by receiving sounds emitted by user 20. Microphone 238 captures sounds emitted by user 20 and converts the captured sounds into sound data, which is then output to processor 46. Speaker 240 outputs sound according to instructions from processor 46.

[0573] Camera 42 is a small digital camera equipped with an optical system such as a lens, aperture and shutter, and imaging elements such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, to photograph the area around robot 414 (e.g., the field of view defined by a perspective equivalent to the field of vision of an average healthy person).

[0574] Communication I / F44 is connected to network 54. Communication I / F44 and 26 are responsible for the transmission and reception of various information between processor 46 and processor 28 via network 54. The transmission and reception of various information between processor 46 and processor 28 using communication I / F44 and 26 is performed in a secure state.

[0575] The controlled object 443 includes a display device, LEDs (light-emitting diodes) for the eyes, and motors for driving the arms, hands, and feet. The posture or movement of the robot 414 is controlled by controlling the motors in the arms, hands, and feet. Some of the robot 414's emotions can be expressed by controlling these motors. In addition, the robot 414's facial expressions can also be expressed by controlling the illumination state of the LEDs in its eyes.

[0576] Figure 8 This illustrates an example of the main functions of the data processing device 12 and the robot 414. For example... Figure 8 As shown, in the data processing device 12, specific processing is performed by the processor 28. The specific processing program 56 is stored in the memory 32.

[0577] The specific processing program 56 is an example of a "program" involved in the technology of this disclosure. The processor 28 reads the specific processing program 56 from the memory 32 and executes the read specific processing program 56 on the RAM 30. Specific processing is implemented by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.

[0578] The memory 32 stores the data generation model 58 and the emotion-specific model 59. The data generation model 58 and the emotion-specific model 59 are used by the specific processing unit 290.

[0579] In robot 414, the processor 46 performs the acceptance and output processing. The memory 50 stores the acceptance and output program 60. The processor 46 reads the acceptance and output program 60 from the memory 50 and executes the read acceptance and output program 60 on RAM 48. The acceptance and output processing is implemented by the processor 46 acting as the control unit 46A according to the acceptance and output program 60 executed on RAM 48.

[0580] Next, the specific processing of the specific processing unit 290 of the data processing device 12 will be described. Each part of the system described below is implemented by the data processing device 12 and the robot 414. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 will be referred to as the "terminal".

[0581] Example 1 The process is the same as that of the specific process described in Embodiment 1 in the first embodiment above, so the description is omitted.

[0582] Application Example 1 The process is the same as that in the specific processing described in Application Example 1 of the first embodiment above, so the description is omitted.

[0583] Example 2 The process is the same as that of the specific process in Embodiment 2 described in the first embodiment above, so the description is omitted.

[0584] Application Example 2 The process is the same as that in the specific processing described in Application Example 2 of the first embodiment above, so the description is omitted.

[0585] The specific processing unit 290 sends the result of the specific processing to the robot 414. In the robot 414, the control unit 46A outputs the result of the specific processing to the speaker 240 and the controlled object 443. The microphone 238 acquires sound input representing the result of the specific processing. The control unit 46A sends the sound data representing the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the sound data.

[0586] Data generation model 58 is a so-called generative AI (Artificial Intelligence). Examples of data generation models 58 include ChatGPT (registered trademark) (accessible via the internet (URL: https: / / openai.com / blog / chatgpt)). Data generation model 58 is obtained through deep learning on a neural network. Input to data generation model 58 are prompt words containing instructions, and inference data such as sound data representing sound, text data representing text, and image data representing images (e.g., still image data or animation data). Data generation model 58 infers from the input inference data based on the instructions represented by the prompt words and outputs the inference result in one or more data forms, such as sound data, text data, and image data. Data generation model 58 includes, for example, text generation AI, image generation AI, and multimodal generation AI. Here, inference refers to, for example, analysis, classification, prediction, and / or induction. The specific processing unit 290 performs the aforementioned specific processing while using data generation model 58. The data generation model 58 can also be a model finely tuned to output inference results from prompts that do not contain instructions. In this case, the data generation model 58 can output inference results based on prompts that do not contain instructions. The data processing apparatus 12, etc., includes various data generation models 58, including AI other than the generation AI. AI other than the generation AI can be, for example, linear regression, logistic regression, decision trees, random forests, support vector machines (SVM), k-means clustering, convolutional neural networks (CNN), recurrent neural networks (RNN), generative adversarial networks (GAN), or Naive Bayes, and can perform various processes, but is not limited to this example. Furthermore, the AI ​​can also be an AI agent. Furthermore, when the processing of the above-mentioned parts is performed by AI, the processing can be performed partially or entirely by AI, but is not limited to this example. Furthermore, the processing performed by the AI ​​including the generation AI can be replaced by processing in the rule base, and the processing in the rule base can also be replaced by processing performed by the AI ​​including the generation AI.

[0587] Furthermore, the processing of the aforementioned data processing system 10 is performed by the specific processing unit 290 of the data processing device 12 or the control unit 46A of the robot 414, but it can also be performed by both the specific processing unit 290 of the data processing device 12 and the control unit 46A of the robot 414. Additionally, the specific processing unit 290 of the data processing device 12 acquires or collects information required for processing from the robot 414 or external devices, and the robot 414 acquires or collects information required for processing from the data processing device 12 or external devices.

[0588] For example, the collection unit is implemented by the control unit 46A of the robot 414 or the specific processing unit 290 of the data processing device 12. For example, the acquisition unit uses the camera 42 or communication I / F 44 of the robot 414 to acquire step data, which is then processed by the specific processing unit 290 of the data processing device 12. For example, the analysis unit is implemented by the specific processing unit 290 of the data processing device 12, which analyzes the data from the collection unit and the acquisition unit. For example, the generation unit is implemented by the specific processing unit 290 of the data processing device 12, which uses a generation AI to generate a menu. For example, the serving unit is implemented by the speaker 240 of the robot 414 and the control object 443 or the specific processing unit 290 of the data processing device 12, which provides the generated menu to the user. The correspondence between each unit and the device or control unit is not limited to the above examples and various changes can be made.

[0589] In the above embodiments, examples of specific processing by the data processing device 12 are given, but the technology disclosed herein is not limited to this, and specific processing may also be performed by the robot 414.

[0590] Furthermore, the emotion-specific model 59, acting as an emotion engine, can determine a user's emotion based on a specific mapping. Specifically, the emotion-specific model 59 can determine a user's emotion based on an emotion graph that serves as a specific mapping (see [reference]). Figure 9 The emotion-specific model 59 can also determine the robot's emotion, and the specific processing unit 290 performs specific processing based on the robot's emotions.

[0591] Figure 9 This is a diagram representing an emotion map 400 that maps multiple emotions. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. The closer to the center of the concentric circles, the more primitive the emotion is. On the outer side of the concentric circles, emotions representing states or behaviors arising from mood are arranged. Emotions are concepts that include feelings and mental states. Emotions generated by reactions occurring in the brain are arranged roughly to the left of the concentric circles. Emotions derived from situational judgments are arranged roughly to the right of the concentric circles. Emotions generated by reactions occurring in the brain and derived from situational judgments are arranged roughly above and below the concentric circles. Furthermore, "pleasant" emotions are arranged above the concentric circles, and "unpleasant" emotions are arranged below them. Thus, in the emotion map 400, multiple emotions are mapped based on the structure that generates emotions, and emotions that are likely to occur simultaneously are mapped close to each other.

[0592] These emotions are distributed at the three o'clock position of the emotion map 400, typically fluctuating between peace and anxiety. In the right half of the emotion map 400, situational awareness dominates over internal sensation, thus resulting in an impression of calm.

[0593] The inner side of the emotion map 400 represents the inner state, while the outer side represents behavior. Therefore, the further outward you are from the emotion map 400, the more visible the emotion becomes (manifested in behavior).

[0594] Here, human emotions are based on various balances such as posture and blood sugar levels. When these balances deviate from an ideal state, it indicates an unpleasant state; when they approach the ideal state, it indicates a pleasant state. Emotions in robots, cars, motorcycles, etc., can also be created in the following way: based on various balances such as posture and remaining battery power, when these balances deviate from an ideal state, it indicates an unpleasant state; when they approach the ideal state, it indicates a pleasant state. Emotion maps can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on a Brain Physiological Signal Analysis System for Voice Emotion Recognition and Emotion, Tokushima University, Doctoral Dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). In the left half of the emotion map, emotions belonging to the sensory-dominated region, called "response," are arranged. Furthermore, in the right half of the emotion map, emotions belonging to the situational cognition-dominated region, called "situation," are arranged.

[0595] In the emotion map, two types of emotions that promote learning are defined. One is a negative emotion on the situational side, in the middle or peripheral region of "repentance" or "reflection." This occurs when the robot experiences negative emotions such as "I don't want to experience this feeling again" or "I don't want to be blamed again." The other is a positive emotion on the response side, near the "desire" region. This occurs when there are positive feelings such as "wanting more" or "wanting to know more."

[0596] The emotion-specific model 59 inputs user input into a pre-trained neural network to obtain emotion values ​​representing each emotion shown in the emotion map 400, thereby determining the user's emotion. This neural network is pre-trained based on multiple training data sets, which are combinations of user input and emotion values ​​representing each emotion shown in the emotion map 400. Furthermore, this neural network... Figure 10 As shown in the sentiment graph 900, it was trained in a way that sentiments that are configured close to each other have similar values. Figure 10 The text shows examples of emotions such as "peace of mind", "stability", and "reassurance" that have similar emotion values.

[0597] The above description focuses on the functions of the data processing device 12, but the system of this disclosure is not necessarily installed on a server. The system of this disclosure can also be installed as a general information processing system. This disclosure can also be installed, for example, as a software program running on a personal computer, an application running on a smartphone, etc. The method of this disclosure can also be provided to users in the form of SaaS (Software as a Service).

[0598] In the above embodiments, an example of a specific process being performed by a single computer 22 is given. However, the technology disclosed herein is not limited to this, and the specific process can also be distributed among multiple computers, including computer 22. For example, the data generation model 58 can be located on an external device of the data processing apparatus 12, where data is generated based on the input data.

[0599] In the above embodiments, examples of storing a specific processing program 56 in the memory 32 have been described, but the technology disclosed herein is not limited thereto. For example, the specific processing program 56 may also be stored in a portable computer-readable non-transitory storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-transitory storage medium is installed into the computer 22 of the data processing device 12. The processor 28 executes specific processing according to the specific processing program 56.

[0600] Alternatively, a specific processing program 56 may be pre-stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 according to the requirements of the data processing device 12.

[0601] In addition, it is not necessary to store all the specific processing program 56 in the storage device such as the server connected to the data processing device 12 via the network 54 or in the memory 32; a portion of the specific processing program 56 may be stored in advance.

[0602] As hardware resources for performing specific processes, various processors, as shown below, can be used. For example, a CPU can be listed as a processor, which functions as a general-purpose processor that performs specific processes by executing software, i.e., a program. Furthermore, processors can be listed as special-purpose circuits such as FPGAs (Field-Programmable Gate Arrays), PLDs (Programmable Logic Devices), or ASICs (Application-Specific Integrated Circuits), which are processors with circuitry specifically designed to perform specific processes. Each processor has built-in or connected memory, and each processor executes specific processes using that memory.

[0603] The hardware resources for performing a specific process can consist of one of these various processors, or a combination of two or more processors of the same or different types (e.g., a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Alternatively, the hardware resources for performing a specific process can be a single processor.

[0604] As an example of a single processor, there are two approaches: First, a processor is composed of a combination of one or more CPUs and software, which functions as a hardware resource to perform a specific process; second, as represented by a SoC (System-on-a-chip), a processor is used to implement the functionality of the entire system, which includes multiple hardware resources for performing a specific process, using a single IC (Integrated Circuit) chip. In this way, the specific process is implemented by using one or more of the aforementioned processors as hardware resources.

[0605] Furthermore, the hardware architecture of these various processors, more specifically, can utilize circuits that combine semiconductor elements and other circuit components. Moreover, the specific process described above is just one example. Therefore, without departing from the main point, unnecessary steps can certainly be deleted, new steps added, or the processing order changed.

[0606] The descriptions and illustrations above are detailed explanations of a portion of the technology disclosed herein, and are merely one example of the technology disclosed herein. For example, the above descriptions of the structure, function, effect, and results are just one example of the structure, function, effect, and results of a portion of the technology disclosed herein. Therefore, without departing from the spirit of the technology disclosed herein, unnecessary parts may be deleted, new elements added, or replacements may be made to the descriptions and illustrations above. Furthermore, to avoid confusion and facilitate understanding of a portion of the technology disclosed herein, explanations of common technical knowledge that do not require special explanation under the premise of being able to implement the technology disclosed herein have been omitted from the descriptions and illustrations above.

[0607] All documents, patent applications and technical specifications set forth in this specification are incorporated herein by reference to the same extent that each document, patent application and technical specification is specifically and individually described therein and referenced by reference.

[0608] In addition, the following notes are provided in response to the above explanation.

[0609] Example 1 (Note 1) An information processing system, characterized in that it comprises: A device for receiving user communication records containing document information by an information processing unit within a communication processing apparatus, and extracting text information as the object of parsing from the document information; An apparatus for the information processing unit to segment and parse the text information using language processing technology, extract multiple word information and recognition information from the text information, and preprocess the text information to generate input data for a generative artificial intelligence model. An apparatus for obtaining risk assessment information about the document information by having the information processing unit generate and send query information containing the main text information and prompt statements related to risk assessment based on the generative artificial intelligence model; An apparatus for comparing the identification information with an information set in a storage device by the information processing unit, and for calculating confidence information about the document information based on the comparison result; An apparatus for fusing the acquired assessment information, the calculated confidence information, and the sentiment information obtained based on user state information by the information processing unit, thereby calculating the risk information of the document information; An apparatus for generating corresponding notification content by the information processing unit based on the risk information, generating a notification control prompt statement for sending the notification content to the terminal device, and sending the notification content to the terminal device.

[0610] (Note 2) The information processing system according to Appendix 1 is characterized in that, The information processing unit uses the language processing technology to segment, divide, analyze language elements, and extract recognition information from the text information, thereby generating feature information for the generative artificial intelligence model.

[0611] (Note 3) The information processing system according to Appendix 1 is characterized in that, The information processing unit sets the priority information of the document information based on the emotional information, and generates prompt statements corresponding to the priority information and the danger information, so as to control the order and form of notification to the terminal device.

[0612] Application Example 1 (Note 1) An information processing system, characterized in that it comprises: An apparatus for generating an evaluation request prompt statement for assessing the degree of danger based on language information and structural information extracted from communication information by a processing circuit in an electronic information processing device, and inputting the prompt statement into a generative artificial intelligence model to obtain a danger assessment result for the communication information; An apparatus for comparing the identification information contained in the communication information with a set of recorded information, calculating a credibility index related to the source of the identification information, and correcting the risk assessment of the communication information based on the credibility index; A device for using language information processing technology to parse the main text of the communication information, divide the text into word sequences, extract a specific set of words and a set of expressions for risk determination from the word sequences, and include the specific set of words and expressions in the prompt statements input to the generative artificial intelligence model; An apparatus for using statistical learning to learn the correspondence between feature quantities and risk levels related to previous and isolated communication information, to make a numerical estimation of the risk level of newly received communication information based on the learned judgment rules, and to fuse the numerical estimation result with the evaluation result of the generative artificial intelligence model to determine the final risk level. An apparatus for generating and sending warning messages containing visual or auditory information to a terminal user based on the final level of danger, and for isolating communication messages whose level of danger exceeds a predetermined threshold by moving them to an independent storage area and imposing access restrictions. An apparatus for parsing user-related input information through emotion recognition processing to infer the user's emotional state, adjusting the content and presentation of the warning information according to the emotional state, and correcting the final danger level when necessary; An apparatus for updating the learned judgment rules and the constituent elements of the prompt statements based on communication information as the object of isolation processing and subsequent user operation records, thereby improving the accuracy of subsequent risk estimation and generative artificial intelligence model evaluation.

[0613] (Note 2) According to the information processing system described in Appendix 1, the processing circuit is configured to receive the communication information as email information, and take the body of the email information, sender information and receiver information as input, execute the language information processing technology and the statistical learning processing in real time, and classify the email information into at least one of dangerous communication information, suspicious dangerous communication information or safe communication information based on the processing results.

[0614] (Note 3) According to the information processing system described in Appendix 1, the processing circuit is configured to set the presentation priority of the email information based on the final danger level and the result of the emotion recognition processing, and control the display order or notification method in the user terminal by sending control information indicating the presentation priority to the user terminal.

[0615] Example 2 (Note 1) An information processing system, characterized in that it comprises: A device for receiving the content and / or identification information of electronic communications by a processing unit in an information processing apparatus; An apparatus for performing format correction on received identification information by the processing unit to generate standardized identification information, sending the standardized identification information as a comparison request to a storage device storing an information set containing confidence information and risk information, and evaluating the confidence of the identification information based on the comparison results obtained from the storage device. An apparatus for generating a hazard index representing the danger of electronic communication by having the processing unit apply natural language processing technology to extract features from text data contained in electronic communication content, and perform discrimination processing using machine learning technology based on the features and the identification information. An apparatus for generating instruction text for a generative artificial intelligence model by the processing unit, the instruction text including prompt statements and input information containing at least one of the content of the electronic communication and the identification information, so that the generative artificial intelligence model analyzes the possibility of whether the electronic communication is phishing communication and / or spam communication, and integrates the analysis results obtained from the generative artificial intelligence model with the risk index and the confidence evaluation results, thereby performing a comprehensive confidence evaluation of the electronic communication; And means for generating response data containing information about the risk level and its basis related to the electronic communication by the processing unit based on the comprehensive confidence assessment, and sending the response data to the terminal device.

[0616] (Note 2) The information processing system according to Appendix 1 is characterized in that, The processing unit is configured to store at least a portion of the comparison results in the information set corresponding to the electronic communication content and identification information received from the terminal device, the risk index obtained by the machine learning technology, and the analysis results obtained from the generative artificial intelligence model as learning data, and to perform learning processing to update the discrimination processing of the machine learning technology and / or the generation processing of the prompt statement based on the learning data.

[0617] (Note 3) The information processing system according to Appendix 1 is characterized in that, The processing unit is configured to determine the display mode and / or notification mode of electronic communication on the terminal device based on the comprehensive confidence evaluation, and when the risk level is not lower than a predetermined threshold, generate control information for restricting the display of the electronic communication and / or highlighting the warning information on the terminal device, and send response data containing the control information to the terminal device.

[0618] Application Example 2 (Note 1) An information processing system, characterized in that it comprises: An apparatus for acquiring text data of electronic communications through an electronic computing device, and for performing segmentation, syntactic structure analysis and lexical structure analysis on the text data using natural language processing, thereby identifying expressions related to risk determination. An apparatus for extracting identification information from the electronic communication by parsing character sequences, comparing the identification information with an information set to evaluate the credibility of the sending source or connection target corresponding to the identification information, and calculating the risk index of the electronic communication based on the credibility evaluation result; An apparatus for generating prompt statements based on the text data of the electronic communication and the hazard index, inputting them into a generative artificial intelligence model, instructing the generative artificial intelligence model to perform hazard assessment processing, and reflecting the assessment results obtained from the generative artificial intelligence model back to the hazard index to update the hazard level of the electronic communication; Apparatus for performing emotion recognition processing based on image data and voice data acquired from a user terminal, thereby estimating the user's emotional state and its intensity; An apparatus for dynamically adjusting the risk level of electronic communication by integrating the emotional state with the risk level, increasing the risk level of electronic communication when the emotional state is negative and decreasing the risk level of electronic communication when the emotional state is positive. A device for determining warning content and user interface display style based on the dynamically adjusted risk level, and sending notification information containing the risk level and corresponding suggested behavior to the user terminal; An apparatus for obtaining the actual response result of the electronic communication based on user operation content obtained from the user terminal, storing the response result as historical information, and using the historical information to update the parameters of the hazard calculation process and / or the prompt statement generation process.

[0619] (Note 2) The information processing system according to Appendix 1 is characterized in that, The electronic computing device is configured to: apply a machine learning model to the text data of the electronic communication, calculate a probability value of harmfulness based on features learned from stored harmful and non-harmful communication data, and estimate the danger level of the electronic communication by adding the probability value of harmfulness to the danger index.

[0620] (Note 3) The information processing system according to Appendix 1 is characterized in that, The electronic computing device is configured to: when generating the prompt statement, embed the text data of the electronic communication, the credibility evaluation result of the identification information, the harmfulness probability value obtained by the machine learning model, and the emotional state inferred by the emotion recognition processing as explanatory information into the prompt statement; simultaneously instruct the generative artificial intelligence model to perform risk assessment and user-oriented explanatory text generation; and include the explanatory text obtained from the generative artificial intelligence model in the notification information and provide it to the user terminal.

Claims

1. An information processing system, characterized in that, Includes a processor, the processor being configured to: Generate prompts to instruct generative artificial intelligence models to assess the danger of emails, so as to determine whether an email is dangerous based on the content of the email body; The identification information contained in the email is compared with the database to determine the source of the identification information, and the credibility of the identification information is evaluated based on the comparison results. Emotion recognition technology is used to analyze users' facial expressions and voice to identify their emotions, and the risk level of emails is adjusted based on the identified emotions.

2. The information processing system according to claim 1, characterized in that, The processor is also configured to parse the email body using natural language processing technology and extract specific keywords and phrases from the email body.

3. The information processing system according to claim 1, characterized in that, The processor is also configured to prioritize emails based on sentiment analysis results and generate notification messages for users.

Citation Information

Patent Citations

  • Persona chatbot control method and system

    JP2022180282A