System

By using AI to analyze and determine the content of phone calls, the system only rings during safe calls, solving the problem that elderly people have difficulty distinguishing between scam calls and enabling them to answer calls from family members and trusted organizations with peace of mind.

CN121907955APending Publication Date: 2026-04-21SOFTBANK GROUP CORP
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
SOFTBANK GROUP CORP
Filing Date
2025-10-15
Publication Date
2026-04-21

AI Technical Summary

Technical Problem

Elderly people lack awareness of fraudulent phone calls and have difficulty distinguishing between safe and unsafe calls, which increases the risk of receiving such calls.

Method used

The system uses AI to analyze call content, converts calls into text using speech recognition and natural language processing technologies, and combines the phone's address book and a database of fraud keywords to determine the call's security. It then rings when the call is deemed secure.

Benefits of technology

It effectively reduces the risk of elderly people receiving fraudulent calls, enabling them to answer calls from family members and trusted organizations with peace of mind and avoid missing important calls.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121907955A_ABST
    Figure CN121907955A_ABST
Patent Text Reader

Abstract

The system provided by the embodiment of the invention aims to enable the elderly to only answer the secure call. A system according to an embodiment includes an analysis unit, a determination unit, and a ringing unit. The analysis part is used for analyzing the telephone content. The determination unit determines whether the telephone is a secure telephone based on the content analyzed by the analysis unit. The ringing unit emits a ring tone when the determination unit determines that the call is a secure call.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The technology disclosed herein relates to a system. Background Technology

[0002] Patent Document 1 discloses a personalized chatbot control method executed by at least one processor, the method comprising: receiving user speech; adding the user speech to a prompt containing instructions related to a chatbot role; encoding the prompt; and inputting the encoded prompt into a language model to generate chatbot speech in response to the user speech.

[0003] Patent document 1: Japanese Patent Application Publication No. 2022-180282. Summary of the Invention

[0004] With current technology, the elderly are not sufficiently protected against fraudulent phone calls and find it difficult to only answer safe calls, which presents a challenge.

[0005] The system involved in this technical solution is designed to enable the elderly to only receive secure phone calls.

[0006] The system involved in this technical solution includes a parsing unit, a judgment unit, and a ringing unit. The parsing unit is used to parse the telephone content. The judgment unit determines whether the telephone is secure based on the parsed content. The ringing unit rings when the judgment unit determines that the telephone is secure.

[0007] The system involved in this technical solution enables the elderly to only receive secure phone calls. Attached Figure Description

[0008] Figure 1 This is a conceptual diagram illustrating an example of the configuration of a data processing system according to the first embodiment.

[0009] Figure 2 This is a conceptual diagram illustrating an example of the main functions of the data processing apparatus and smart device according to the first embodiment.

[0010] Figure 3 This is a conceptual diagram illustrating an example of the data processing system configuration in the second embodiment.

[0011] Figure 4 This is a conceptual diagram illustrating an example of the main functions of the data processing device and smart glasses according to the second embodiment.

[0012] Figure 5 This is a conceptual diagram illustrating an example of the data processing system configuration in the third embodiment.

[0013] Figure 6This is a conceptual diagram illustrating an example of the main functions of the data processing device and head-mounted terminal according to the third embodiment.

[0014] Figure 7 This is a conceptual diagram illustrating an example of the data processing system configuration in the fourth embodiment.

[0015] Figure 8 This is a conceptual diagram illustrating an example of the functions of the main parts of the data processing device and robot according to the fourth embodiment.

[0016] Figure 9 It represents an emotion graph that maps multiple emotions.

[0017] Figure 10 It represents an emotion graph that maps multiple emotions.

[0018] Explanation of reference numerals in the attached figures

[0019] Data processing systems 10, 210, 310, and 410

[0020] 12 Data processing device

[0021] 14 Smart devices

[0022] 214 Smart Glasses

[0023] 314 Head-mounted terminal

[0024] 414 Robot. Detailed Implementation

[0025] Hereinafter, an example of an implementation of the system involved in this disclosure will be described with reference to the accompanying drawings.

[0026] First, let's explain the terms used in the following description.

[0027] In the following embodiments, the processor (hereinafter referred to as "processor"), as indicated by the reference numerals, can be a single computing device or a combination of multiple computing devices. Furthermore, a processor can be a single computing device or a combination of multiple computing devices. Examples of computing devices include CPU (Central Processing Unit), GPU (Graphics Processing Unit), GPGPU (General-Purpose computing on Graphics Processing Units), APU (Accelerated Processing Unit), or TPU (Tensor Processing Unit), etc.

[0028] In the following embodiments, RAM (Random Access Memory), as indicated in the figures, is a memory for temporary information storage that is used by the processor as working memory.

[0029] In the following embodiments, the memory, as indicated by the reference numerals, is one or more non-volatile storage devices used to store various programs and various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), disk (e.g., hard disk), or magnetic tape, etc.

[0030] In the following embodiments, the communication I / F (Interface) with reference numerals is an interface including a communication processor and an antenna, etc. The communication I / F is responsible for communication between multiple computers. Examples of communication standards applicable to the communication I / F include wireless communication standards such as 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), or Bluetooth (registered trademark).

[0031] In the following implementation, "A and / or B" is synonymous with "at least one of A and B". That is, "A and / or B" means that it can be only A, only B, or a combination of A and B. Furthermore, in this specification, when "and / or" connects more than three items, the same approach as "A and / or B" applies.

[0032] First Implementation Method

[0033] Figure 1 An example of the configuration of the data processing system 10 according to the first embodiment is shown.

[0034] like Figure 1 As shown, the data processing system 10 includes a data processing device 12 and an intelligent device 14. An example of the data processing device 12 is a server.

[0035] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 includes a processor 28, RAM 30, and memory 32. The processor 28, RAM 30, and memory 32 are connected to a bus 34. Furthermore, the database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. An example of the network 54 includes a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0036] The smart device 14 includes a computer 36, a receiver 38, an output device 40, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, RAM 48, and memory 50. The processor 46, RAM 48, and memory 50 are connected to a bus 52. In addition, the receiver 38, output device 40, and camera 42 are also connected to the bus 52.

[0037] The receiving device 38 includes a touchscreen 38A and a microphone 38B, etc., for receiving user input. The touchscreen 38A receives user input generated by contact with an indicator (e.g., a pen or finger). The microphone 38B receives user input generated by sound by detecting the user's voice. The control unit 46A sends data representing user input received via the touchscreen 38A and microphone 38B to the data processing unit 12. In the data processing unit 12, a specific processing unit 290 (see reference...) Figure 2 Get the data that represents the user input.

[0038] The output device 40 includes a display 40A and a speaker 40B, etc., and presents data to the user by outputting data in a user-perceptible form (e.g., sound and / or text). The display 40A displays visual information such as text and images according to instructions from the processor 46. The speaker 40B outputs sound according to instructions from the processor 46. The camera 42 is a small digital camera equipped with an optical system such as a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.

[0039] Communication I / F 44 is connected to network 54. Communication I / F 44 and 26 are responsible for sending and receiving various information between processor 46 and processor 28 via network 54.

[0040] Figure 2 An example of the main functions of the data processing device 12 and the smart device 14 is shown.

[0041] like Figure 2 As shown, in the data processing apparatus 12, specific processing is performed by the processor 28. A specific processing program 56 is stored in the memory 32. The specific processing program 56 is an example of a "program" as understood in this disclosure. The processor 28 reads the specific processing program 56 from the memory 32 and executes the read specific processing program 56 on the RAM 30. Specific processing is implemented by the processor 28 operating as a specific processing unit 290 based on the specific processing program 56 executed on the RAM 30.

[0042] The memory 32 stores a data generation model 58 and an emotion-specific model 59. The data generation model 58 and the emotion-specific model 59 are used by the specific processing unit 290. The specific processing unit 290 can use the emotion-specific model 59 to infer the user's emotions and perform specific processing based on the user's emotions. The emotion inference function using the emotion-specific model 59 (emotion-specific function) includes inferring and predicting the user's emotions, performing various inferences and predictions related to the user's emotions, but is not limited to this example. Furthermore, emotion inference and prediction also include, for example, emotion analysis (analysis).

[0043] In the smart device 14, specific processing is performed by the processor 46. A specific processing program 60 is stored in the memory 50. The specific processing program 60 is used in conjunction with the data processing system 10. The processor 46 reads the specific processing program 60 from the memory 50 and executes the read specific processing program 60 on the RAM 48. Specific processing is implemented by the processor 46 operating as a control unit 46A based on the specific processing program 60 executed on the RAM 48. Furthermore, the smart device 14 may also have the same data generation model and emotion-specific model as the data generation model 58 and the emotion-specific model 59, and use these models to perform the same processing as the specific processing unit 290.

[0044] Furthermore, other devices besides the data processing device 12 may also have the data generation model 58. For example, a server device (e.g., a generation server) may have the data generation model 58. In this case, the data processing device 12 obtains the processing results (prediction results, etc.) using the data generation model 58 by communicating with the server device that has the data generation model 58. Furthermore, the data processing device 12 may be a server device or a user-owned terminal device (e.g., a mobile phone, robot, home appliance, etc.). Next, an example of the processing performed by the data processing system 10 of the first embodiment will be described.

[0045] Implementation Method 1

[0046] The telephone call system described in this invention is a system that uses AI to allow only safe calls to be connected when a call comes in at an elderly person's home. When a call comes in, the AI ​​analyzes the call content to determine if it is a safe call. Then, it only rings when the call is determined to be a safe call. This allows the elderly person to answer the phone with peace of mind. For example, when a call comes in, the AI ​​uses speech recognition technology to convert the call content into text. For example, calls containing keywords that suggest fraud are considered unsafe. This reduces the risk of fraud. Next, the AI ​​analyzes the call content to determine if it is a safe call. For example, calls from family or friends are considered safe, while calls from unknown numbers or containing keywords that suggest fraud are considered unsafe. This allows the elderly person to answer the phone with peace of mind. Furthermore, it only rings when the call is determined to be a safe call. This allows the elderly person to answer the phone with peace of mind. For example, the ringing only occurs when a family or friend is calling, allowing the elderly person to answer the phone with peace of mind. Through this mechanism, the elderly person can answer the phone with peace of mind, achieving a state where only safe calls are allowed. For example, it can reduce the risk of fraud, allowing seniors to answer the phone with peace of mind. Furthermore, seniors feel more secure answering calls only from family or friends. Therefore, the telephone call system can provide seniors with a sense of security when answering the phone.

[0047] The telephone call system described in this embodiment includes a parsing unit, a judgment unit, and a ringing unit. The parsing unit parses the telephone content. For example, the parsing unit uses speech recognition technology to convert the telephone content into text. The parsing unit can also use deep learning-based speech recognition technology to convert the telephone content into text with high accuracy. Furthermore, the parsing unit can also use HMM-based speech recognition technology to convert the telephone content into text. The parsing unit can also use speech recognition technology to convert the telephone content into text in real time. The judgment unit determines whether a call is safe based on the content parsed by the parsing unit. For example, the judgment unit determines calls from family or friends as safe. The judgment unit can also determine calls from people registered in the address book as safe. Furthermore, the judgment unit can determine calls from specific phone numbers as safe. The judgment unit can also prioritize determining calls from family or friends as safe. The judgment unit determines calls containing keywords suggesting fraud as unsafe. For example, the judgment unit can determine calls containing keywords such as "money," "remittance," or "emergency" as unsafe. In addition, the judgment unit can update its keywords suggesting fraud by referring to the latest information on fraud methods. The judgment unit can also automatically extract keywords indicating a high probability of fraud and determine them as unsafe. The ringing unit rings when the judgment unit determines the call is safe. For example, the ringing unit rings only when the call is safe. For example, the ringing unit rings only when a family member or friend calls. Furthermore, the ringing unit can also ring only when a specific phone number calls. For example, the ringing unit rings a specific ringtone when the call is determined to be safe. Therefore, the telephone call system according to this embodiment allows elderly people to answer the phone with peace of mind.

[0048] The parsing unit is used to analyze telephone conversations. For example, it uses speech recognition technology to convert the conversations into text. Specifically, by employing deep learning-based speech recognition technology, the conversations can be converted into text with high accuracy. Deep learning technology, by learning from large amounts of speech data, can accurately capture speech features, even in noisy environments. Alternatively, HMM (Hidden Markov Model)-based speech recognition technology can also be used. HMMs, by modeling the temporal changes in speech, can efficiently analyze continuous speech data. By combining these technologies, the parsing unit can convert the conversations into text in real time. Furthermore, the parsing unit can utilize not only speech recognition technology but also natural language processing technology to analyze the meaning of the text. For example, it can extract important keywords from the text to summarize the conversation content. Thus, the parsing unit can analyze the conversation content in detail and generate high-precision data for the decision-making unit to use.

[0049] The judgment department determines whether a call is safe based on the content analyzed by the analysis department. For example, the judgment department might classify calls from family or friends as safe. Specifically, it can classify calls from people registered in the user's address book as safe. The address book is a database of pre-registered phone numbers and names; the judgment department refers to this list to confirm the caller. Furthermore, calls from specific phone numbers can be classified as safe. For example, pre-registered phone numbers of trusted institutions such as banks and medical facilities can be classified as safe. Further, the judgment department classifies calls containing keywords suggesting fraud as unsafe. For example, calls containing keywords such as "money," "remittance," or "emergency" can be classified as unsafe. The judgment department can also update its keyword list based on the latest information about fraud methods. Thus, the judgment department can always make security judgments based on the latest fraud information. In addition, the judgment department can use AI to analyze call content and automatically extract keywords suggesting fraud. By learning from past fraud call data, AI can identify fraud patterns and counter new fraud methods. Therefore, the judgment department can accurately determine security, providing users with a safe calling environment.

[0050] The ringing unit emits a ringing tone when the determination unit determines the call is a safe call. Specifically, it only rings when the call is a safe call. For example, it only rings when a family member or friend calls. The ringing unit can also ring only when a specific phone number calls. This ensures that the user does not miss important calls. Furthermore, the ringing unit can emit a specific ringing tone when a call is determined to be a safe call. For example, a specific melody can be set for a family member's call, and a different melody for a friend's call, thus intuitively identifying the caller. Furthermore, the ringing unit can customize the ringing tone according to the user's preferences. For example, a quiet ringtone can be set for a specific time period, and a louder ringtone can be used for emergencies. This allows the ringing unit to provide flexible notification methods that suit the user's lifestyle. In addition, the ringing unit can not only emit a ringing tone, but also combine it with notification methods such as vibration or flashing lights. This not only meets the needs of users with hearing impairments or those who need to answer calls in a quiet environment, but also meets diverse needs. Therefore, the telephone call system according to this embodiment not only allows the elderly to answer calls with peace of mind, but also meets the needs of various users.

[0051] The parsing unit can convert telephone content into text using speech recognition technology. For example, the parsing unit can employ deep learning-based speech recognition technology to convert telephone content into text with high accuracy. The parsing unit can also employ HMM-based speech recognition technology to convert telephone content into text. Furthermore, the parsing unit can utilize speech recognition technology to convert telephone content into text in real time. Therefore, by converting telephone content into text, parsing accuracy can be improved. Speech recognition technologies include, for example, deep learning-based speech recognition technology, HMM-based speech recognition technology, etc., but are not limited to these. Some or all of the above processing in the parsing unit can be performed using generative AI, or it can be performed without using generative AI. For example, the parsing unit can input telephone voice data into generative AI, which then converts the voice data into text data.

[0052] The determination unit can identify calls from family or friends as safe. For example, it can identify calls from people registered in the address book as safe. It can also identify calls from specific phone numbers as safe. Furthermore, it can prioritize identifying calls from family or friends as safe. Thus, by identifying calls from family or friends as safe, the elderly can answer the phone with peace of mind. Family or friends include, for example, people registered in the address book or calls from specific phone numbers, but are not limited to these. Some or all of the above processing in the determination unit can be performed using AI, or it can be performed without AI. For example, the determination unit can input caller information into the AI, which then determines whether the call is safe based on the caller information.

[0053] The judgment unit can identify phone calls containing keywords suggesting fraud as unsafe. For example, it can identify calls containing keywords such as "money," "remittance," or "urgent." The judgment unit can also update its keyword list based on the latest information on fraud methods. Furthermore, it can automatically extract keywords suggesting fraud and identify them as unsafe. This reduces the risk of fraud. Keywords suggesting fraud include, but are not limited to, "money," "remittance," and "urgent." Some or all of the above processing in the judgment unit can be performed using AI, or it can be performed without AI. For example, the judgment unit can input the call content into AI, which will then extract keywords suggesting fraud.

[0054] The ringing unit only rings when the call is a secure call. For example, it may only ring when a family member or friend calls. Furthermore, the ringing unit can also ring only when a specific phone number is being called. For example, it may ring a specific ringtone when a call is determined to be a secure call. This allows elderly people to answer the phone with peace of mind. Secure calls include, for example, calls from family members or friends, or calls from specific phone numbers, but are not limited to these. Some or all of the above processing in the ringing unit may be performed using AI, or it may not. For example, the ringing unit can input the information that the determination unit has determined the call to be secure into the AI, and the AI ​​will then ring the incoming call.

[0055] When converting telephone conversations into text using speech recognition technology, the parsing unit can handle different languages ​​or dialects. For example, the parsing unit can use AI to automatically detect different languages ​​and select an appropriate speech recognition model for text conversion. The parsing unit can also identify dialectal or regionally specific vocabulary and use a dictionary to convert them into the standard language. Furthermore, the parsing unit can adjust the speech recognition accuracy and perform text conversion based on the user-defined language. Thus, by handling different languages ​​or dialects, the parsing accuracy is improved. Different languages ​​or dialects include, for example, multilingual speech recognition engines and dialect dictionaries, but are not limited to these. Some or all of the above processing in the parsing unit can be performed using generative AI, or it can be performed without generative AI. For example, the parsing unit can input telephone voice data into generative AI, which will then perform text conversion corresponding to different languages ​​or dialects.

[0056] When parsing telephone content, the parsing unit can remove background noise and other sounds to improve parsing accuracy. For example, the parsing unit can use AI to filter background noise in real time, extracting only the main speech for parsing. The parsing unit can also utilize noise reduction techniques to make the telephone content clearer. Furthermore, the parsing unit can emphasize specific frequency bands to reduce noise and improve speech recognition accuracy. Thus, by removing background noise and other sounds, parsing accuracy is improved. Background noise and other sounds include, but are not limited to, noise reduction techniques and filtering techniques. Some or all of the above processing in the parsing unit can be performed using generative AI, or it can be performed without generative AI. For example, the parsing unit can input telephone voice data into generative AI, which will then perform the removal of background noise and other sounds.

[0057] When the parsing unit converts telephone conversations into text using speech recognition technology, it can highlight specific keywords or phrases. For example, the parsing unit can automatically detect keywords that may indicate fraud and highlight them in the text. It can also highlight important phrases such as the names of family members or friends. Furthermore, it can highlight specific keywords set by the user. Thus, by highlighting specific keywords or phrases, important information is less likely to be overlooked. Specific keywords or phrases include, but are not limited to, words such as "important," "urgent," and "confirm." Some or all of the above processing in the parsing unit can be performed using generative AI, or it can be performed without generative AI. For example, the parsing unit can input telephone voice data into generative AI, which can then perform the highlighting of specific keywords or phrases.

[0058] When analyzing phone call content, the analysis unit can refer to past call records to improve analysis accuracy. For example, the analysis unit can analyze past call records, identifying specific frequently used phrases or keywords. The analysis unit can also learn specific patterns from past call records to improve analysis accuracy. Furthermore, the analysis unit can identify high-risk fraudulent calls based on past call records to improve analysis accuracy. Thus, by referring to past call records, analysis accuracy is improved. Past call records include, but are not limited to, factors such as the retention period of call content and the frequency of reference. Some or all of the above processing in the analysis unit can be performed using generative AI, or it can be performed without generative AI. For example, the analysis unit can input past call record data into the generative AI, which will then perform call record analysis.

[0059] When determining whether a family member's or friend's phone call is safe, the judgment unit can refer to call logs and contact lists to improve the accuracy of the judgment. For example, the judgment unit can analyze call logs to identify frequently called individuals as safe calls. The judgment unit can also refer to contact lists to identify registered family members or friends as safe calls. The judgment unit can further improve the accuracy by combining call logs and contact lists. Thus, by referring to call logs and contact lists, the accuracy of the judgment is improved. Call logs and contact lists include, but are not limited to, factors such as call frequency and contact trustworthiness. Some or all of the above processing in the judgment unit can be performed using a generative AI, or it can be performed without using a generative AI. For example, the judgment unit can input call log data and contact lists into a generative AI, which will then perform call log and contact list analysis.

[0060] When the judgment department determines that a call containing keywords suggesting potential fraud is unsafe, it can update its judgment criteria by referencing the latest information on fraud methods. For example, the judgment department can periodically update the latest fraud methods and reflect this in the judgment criteria. The judgment department can also automatically update the list of keywords suggesting potential fraud and incorporate them into the judgment criteria. The judgment department can also adjust its judgment criteria by referring to news or reports about fraud methods. Thus, by referencing the latest information on fraud methods, the judgment criteria remain up-to-date. The latest information on fraud methods includes, for example, police databases, security company reports, etc., but is not limited to these. Some or all of the above processing in the judgment department can be performed using AI, or it can be performed without AI. For example, the judgment department can input the latest information on fraud methods into the AI, which will then update the judgment criteria.

[0061] When determining whether a call from family or friends is safe, the judgment unit can consider call frequency and time period to improve judgment accuracy. For example, the judgment unit can analyze call frequency to classify frequently called numbers as safe. The judgment unit can also consider call time period to classify calls received during typical time periods as safe. The judgment unit can further improve judgment accuracy by combining call frequency and time period. Thus, by considering call frequency and time period, judgment accuracy is improved. Call frequency and time period include, for example, the number of calls and the duration of calls, but are not limited to these. Some or all of the above processing in the judgment unit can be performed using a generative AI, or it can be performed without using a generative AI. For example, the judgment unit can input call record data into a generative AI, which will then perform call frequency and time period analysis.

[0062] When the judgment unit determines a call containing keywords suggesting potential fraud as unsafe, it can consider not only the call content but also the sender's information. For example, the judgment unit can analyze the sender's phone number, identifying specific high-risk numbers. It can also refer to the sender's location information, identifying calls from specific high-risk areas. The judgment unit can further improve the accuracy of its judgment by combining sender information and call content. Thus, by considering sender information, the accuracy of the judgment is improved. Sender information includes, for example, the sender's phone number and location, but is not limited to these. Some or all of the above processing in the judgment unit can be performed using generative AI, or it can be performed without generative AI. For example, the judgment unit can input sender information into generative AI, which will then perform sender information analysis.

[0063] The ringing department, which only rings for secure calls, can adjust the ringing timing by considering the user's schedule and activity level. For example, it can refer to the user's calendar information to avoid ringing during important meetings. It can also monitor the user's activity in real time to avoid ringing during exercise or sleep. Furthermore, the ringing department can select the optimal time to ring based on the user's schedule. Thus, it can ring at the best time based on the user's schedule and activity level. User schedules and activity levels include, but are not limited to, calendar applications and activity trackers. Some or all of the above processing in the ringing department can be performed using AI, or it can be performed without AI. For example, the ringing department can input the user's schedule information into the AI, which will then adjust the ringing timing.

[0064] When the ringing unit emits a ringtone, it can provide a customizable ringtone based on user preferences. For example, the ringing unit can set the user-selected music or sound effect as the ringtone. It can also provide ringtones that reflect the user's customized volume or timbre. Furthermore, the ringing unit can set different ringtones for specific contacts set by the user. Thus, by setting ringtones according to user preferences, a more comfortable user experience is provided. User preferences include, but are not limited to, user settings and past selection records. Some or all of the above processing in the ringing unit can be performed using AI generation, or it can be performed without AI generation. For example, the ringing unit can input user preference data into AI generation, which will then provide a customizable ringtone.

[0065] The ringing unit, when ringing only for secure calls, can consider user device settings and ambient sound to set the optimal volume. For example, the ringing unit can automatically set a suitable volume based on user device settings. The ringing unit can also detect ambient sound in real time and set the optimal volume accordingly. The ringing unit can also customize the volume based on user preferences. Therefore, by setting the optimal volume based on user device settings and ambient sound, a more suitable ringtone can be set. User device settings and ambient sound include, but are not limited to, device volume settings and ambient noise levels. Some or all of the above processing in the ringing unit can be performed using a generated AI, or it can be performed without using a generated AI. For example, the ringing unit can input user device settings and ambient sound data into the generated AI, which will then perform the optimal volume setting.

[0066] When the ringing unit issues a ring, it can adjust the ringing pattern by referring to the user's past response records. For example, the ringing unit can analyze the user's past response time periods and adjust the ringing pattern accordingly. The ringing unit can also learn the optimal ringing pattern from the user's past response records and apply it. The ringing unit can also set a corresponding ringing pattern when the user is likely to respond during specific time periods. Thus, by referring to the user's past response records, an optimal ringing pattern can be set. The user's past response records include, for example, response frequency, response time periods, etc., but are not limited to these. Some or all of the above processing in the ringing unit can be performed using a generation AI, or it can be performed without using a generation AI. For example, the ringing unit can input the user's past response record data into the generation AI, which will then perform the ringing pattern adjustment.

[0067] The system involved in this embodiment is not limited to the examples described above. For example, various modifications can be made as shown below.

[0068] When converting telephone conversations into text using speech recognition technology, the parsing unit can handle different languages ​​or dialects. For example, the parsing unit can use AI to automatically detect different languages ​​and select an appropriate speech recognition model for text conversion. The parsing unit can also identify dialectal or regionally specific vocabulary and use a dictionary to convert them into the standard language. Furthermore, the parsing unit can adjust the speech recognition accuracy and perform text conversion based on the user-defined language. Thus, by handling different languages ​​or dialects, the parsing accuracy is improved. Different languages ​​or dialects include, for example, multilingual speech recognition engines and dialect dictionaries, but are not limited to these. Some or all of the above processing in the parsing unit can be performed using generative AI, or it can be performed without generative AI. For example, the parsing unit can input telephone voice data into generative AI, which will then perform text conversion corresponding to different languages ​​or dialects.

[0069] When parsing telephone content, the parsing unit can remove background noise and other sounds to improve parsing accuracy. For example, the parsing unit can use AI to filter background noise in real time, extracting only the main speech for parsing. The parsing unit can also utilize noise reduction techniques to make the telephone content clearer. Furthermore, the parsing unit can emphasize specific frequency bands to reduce noise and improve speech recognition accuracy. Thus, by removing background noise and other sounds, parsing accuracy is improved. Background noise and other sounds include, but are not limited to, noise reduction techniques and filtering techniques. Some or all of the above processing in the parsing unit can be performed using generative AI, or it can be performed without generative AI. For example, the parsing unit can input telephone voice data into generative AI, which will then perform the removal of background noise and other sounds.

[0070] When the judgment department determines that a call containing keywords suggesting potential fraud is unsafe, it can update its judgment criteria by referencing the latest information on fraud methods. For example, the judgment department can periodically update the latest fraud methods and reflect this in the judgment criteria. The judgment department can also automatically update the list of keywords suggesting potential fraud and incorporate them into the judgment criteria. The judgment department can also adjust its judgment criteria by referring to news or reports about fraud methods. Thus, by referencing the latest information on fraud methods, the judgment criteria remain up-to-date. The latest information on fraud methods includes, for example, police databases, security company reports, etc., but is not limited to these. Some or all of the above processing in the judgment department can be performed using AI, or it can be performed without AI. For example, the judgment department can input the latest information on fraud methods into the AI, which will then update the judgment criteria.

[0071] The ringing department, which only rings for secure calls, can adjust the ringing timing by considering the user's schedule and activity level. For example, it can refer to the user's calendar information to avoid ringing during important meetings. It can also monitor the user's activity in real time to avoid ringing during exercise or sleep. Furthermore, the ringing department can select the optimal time to ring based on the user's schedule. Thus, it can ring at the best time based on the user's schedule and activity level. User schedules and activity levels include, but are not limited to, calendar applications and activity trackers. Some or all of the above processing in the ringing department can be performed using AI, or it can be performed without AI. For example, the ringing department can input the user's schedule information into the AI, which will then adjust the ringing timing.

[0072] When analyzing phone call content, the analysis unit can refer to past call records to improve analysis accuracy. For example, the analysis unit can analyze past call records, identifying specific frequently used phrases or keywords. The analysis unit can also learn specific patterns from past call records to improve analysis accuracy. Furthermore, the analysis unit can identify high-risk fraudulent calls based on past call records to improve analysis accuracy. Thus, by referring to past call records, analysis accuracy is improved. Past call records include, but are not limited to, factors such as the retention period of call content and the frequency of reference. Some or all of the above processing in the analysis unit can be performed using generative AI, or it can be performed without generative AI. For example, the analysis unit can input past call record data into the generative AI, which will then perform call record analysis.

[0073] The ringing unit, when ringing only for secure calls, can consider user device settings and ambient sound to set the optimal volume. For example, the ringing unit can automatically set a suitable volume based on user device settings. The ringing unit can also detect ambient sound in real time and set the optimal volume accordingly. The ringing unit can also customize the volume based on user preferences. Therefore, by setting the optimal volume based on user device settings and ambient sound, a more suitable ringtone can be set. User device settings and ambient sound include, but are not limited to, device volume settings and ambient noise levels. Some or all of the above processing in the ringing unit can be performed using a generated AI, or it can be performed without using a generated AI. For example, the ringing unit can input user device settings and ambient sound data into the generated AI, which will then perform the optimal volume setting.

[0074] The following is a brief description of the processing flow of Implementation Method 1.

[0075] Step 1: The parsing unit parses the telephone content. The parsing unit uses speech recognition technology to convert the telephone content into text. For example, deep learning-based speech recognition technology or HMM-based speech recognition technology can be used to convert the telephone content into text with high accuracy and in real time.

[0076] Step 2: The judgment department determines whether a call is safe based on the analysis provided by the analysis department. Calls from family or friends, or from people registered in the contact list, are considered safe. Calls containing keywords suggesting potential fraud are deemed unsafe, and the department may update these keywords based on the latest information on fraud methods.

[0077] Step 3: The ringing unit sounds a ringing tone when the determination unit determines the phone number to be secure. For example, the ringing tone may only sound when a family member or friend calls, or when a specific phone number calls. Furthermore, a specific ringing tone may be played when the phone number is determined to be secure.

[0078] Implementation Method 2

[0079] The telephone call system described in this invention is a system that uses AI to allow only safe calls to be connected when a call comes in at an elderly person's home. When a call comes in, the AI ​​analyzes the call content to determine if it is a safe call. Then, it only rings when the call is determined to be a safe call. This allows the elderly person to answer the phone with peace of mind. For example, when a call comes in, the AI ​​uses speech recognition technology to convert the call content into text. For example, calls containing keywords that suggest fraud are considered unsafe. This reduces the risk of fraud. Next, the AI ​​analyzes the call content to determine if it is a safe call. For example, calls from family or friends are considered safe, while calls from unknown numbers or containing keywords that suggest fraud are considered unsafe. This allows the elderly person to answer the phone with peace of mind. Furthermore, it only rings when the call is determined to be a safe call. This allows the elderly person to answer the phone with peace of mind. For example, the ringing only occurs when a family or friend is calling, allowing the elderly person to answer the phone with peace of mind. Through this mechanism, the elderly person can answer the phone with peace of mind, achieving a state where only safe calls are allowed. For example, it can reduce the risk of fraud, allowing seniors to answer the phone with peace of mind. Furthermore, seniors feel more secure answering calls only from family or friends. Therefore, the telephone call system can provide seniors with a sense of security when answering the phone.

[0080] The telephone call system described in this embodiment includes a parsing unit, a judgment unit, and a ringing unit. The parsing unit parses the telephone content. For example, the parsing unit uses speech recognition technology to convert the telephone content into text. The parsing unit can also use deep learning-based speech recognition technology to convert the telephone content into text with high accuracy. Furthermore, the parsing unit can also use HMM-based speech recognition technology to convert the telephone content into text. The parsing unit can also use speech recognition technology to convert the telephone content into text in real time. The judgment unit determines whether a call is safe based on the content parsed by the parsing unit. For example, the judgment unit determines calls from family or friends as safe. The judgment unit can also determine calls from people registered in the address book as safe. Furthermore, the judgment unit can determine calls from specific phone numbers as safe. The judgment unit can also prioritize determining calls from family or friends as safe. The judgment unit determines calls containing keywords suggesting fraud as unsafe. For example, the judgment unit can determine calls containing keywords such as "money," "remittance," or "emergency" as unsafe. In addition, the judgment unit can update its keywords suggesting fraud by referring to the latest information on fraud methods. The judgment unit can also automatically extract keywords indicating a high probability of fraud and determine them as unsafe. The ringing unit rings when the judgment unit determines the call is safe. For example, the ringing unit rings only when the call is safe. For example, the ringing unit rings only when a family member or friend calls. Furthermore, the ringing unit can also ring only when a specific phone number calls. For example, the ringing unit rings a specific ringtone when the call is determined to be safe. Therefore, the telephone call system according to this embodiment allows elderly people to answer the phone with peace of mind.

[0081] The parsing unit is used to analyze telephone conversations. For example, it uses speech recognition technology to convert the conversations into text. Specifically, by employing deep learning-based speech recognition technology, the conversations can be converted into text with high accuracy. Deep learning technology, by learning from large amounts of speech data, can accurately capture speech features, even in noisy environments. Alternatively, HMM (Hidden Markov Model)-based speech recognition technology can also be used. HMMs, by modeling the temporal changes in speech, can efficiently analyze continuous speech data. By combining these technologies, the parsing unit can convert the conversations into text in real time. Furthermore, the parsing unit can utilize not only speech recognition technology but also natural language processing technology to analyze the meaning of the text. For example, it can extract important keywords from the text to summarize the conversation content. Thus, the parsing unit can analyze the conversation content in detail and generate high-precision data for the decision-making unit to use.

[0082] The judgment department determines whether a call is safe based on the content analyzed by the analysis department. For example, the judgment department might classify calls from family or friends as safe. Specifically, it can classify calls from people registered in the user's address book as safe. The address book is a database of pre-registered phone numbers and names; the judgment department refers to this list to confirm the caller. Furthermore, calls from specific phone numbers can be classified as safe. For example, pre-registered phone numbers of trusted institutions such as banks and medical facilities can be classified as safe. Further, the judgment department classifies calls containing keywords suggesting fraud as unsafe. For example, calls containing keywords such as "money," "remittance," or "emergency" can be classified as unsafe. The judgment department can also update its keyword list based on the latest information about fraud methods. Thus, the judgment department can always make security judgments based on the latest fraud information. In addition, the judgment department can use AI to analyze call content and automatically extract keywords suggesting fraud. By learning from past fraud call data, AI can identify fraud patterns and counter new fraud methods. Therefore, the judgment department can accurately determine security, providing users with a safe calling environment.

[0083] The ringing unit emits a ringing tone when the determination unit determines the call is a safe call. Specifically, it only rings when the call is a safe call. For example, it only rings when a family member or friend calls. The ringing unit can also ring only when a specific phone number calls. This ensures that the user does not miss important calls. Furthermore, the ringing unit can emit a specific ringing tone when a call is determined to be a safe call. For example, a specific melody can be set for a family member's call, and a different melody for a friend's call, thus intuitively identifying the caller. Furthermore, the ringing unit can customize the ringing tone according to the user's preferences. For example, a quiet ringtone can be set for a specific time period, and a louder ringtone can be used for emergencies. This allows the ringing unit to provide flexible notification methods that suit the user's lifestyle. In addition, the ringing unit can not only emit a ringing tone, but also combine it with notification methods such as vibration or flashing lights. This not only meets the needs of users with hearing impairments or those who need to answer calls in a quiet environment, but also meets diverse needs. Therefore, the telephone call system according to this embodiment not only allows the elderly to answer calls with peace of mind, but also meets the needs of various users.

[0084] The parsing unit can convert telephone content into text using speech recognition technology. For example, the parsing unit can employ deep learning-based speech recognition technology to convert telephone content into text with high accuracy. The parsing unit can also employ HMM-based speech recognition technology to convert telephone content into text. Furthermore, the parsing unit can utilize speech recognition technology to convert telephone content into text in real time. Therefore, by converting telephone content into text, parsing accuracy can be improved. Speech recognition technologies include, for example, deep learning-based speech recognition technology, HMM-based speech recognition technology, etc., but are not limited to these. Some or all of the above processing in the parsing unit can be performed using generative AI, or it can be performed without using generative AI. For example, the parsing unit can input telephone voice data into generative AI, which then converts the voice data into text data.

[0085] The determination unit can identify calls from family or friends as safe. For example, it can identify calls from people registered in the address book as safe. It can also identify calls from specific phone numbers as safe. Furthermore, it can prioritize identifying calls from family or friends as safe. Thus, by identifying calls from family or friends as safe, the elderly can answer the phone with peace of mind. Family or friends include, for example, people registered in the address book or calls from specific phone numbers, but are not limited to these. Some or all of the above processing in the determination unit can be performed using AI, or it can be performed without AI. For example, the determination unit can input caller information into the AI, which then determines whether the call is safe based on the caller information.

[0086] The judgment unit can identify phone calls containing keywords suggesting fraud as unsafe. For example, it can identify calls containing keywords such as "money," "remittance," or "urgent." The judgment unit can also update its keyword list based on the latest information on fraud methods. Furthermore, it can automatically extract keywords suggesting fraud and identify them as unsafe. This reduces the risk of fraud. Keywords suggesting fraud include, but are not limited to, "money," "remittance," and "urgent." Some or all of the above processing in the judgment unit can be performed using AI, or it can be performed without AI. For example, the judgment unit can input the call content into AI, which will then extract keywords suggesting fraud.

[0087] The ringing unit only rings when the call is a secure call. For example, it may only ring when a family member or friend calls. Furthermore, the ringing unit can also ring only when a specific phone number is being called. For example, it may ring a specific ringtone when a call is determined to be a secure call. This allows elderly people to answer the phone with peace of mind. Secure calls include, for example, calls from family members or friends, or calls from specific phone numbers, but are not limited to these. Some or all of the above processing in the ringing unit may be performed using AI, or it may not. For example, the ringing unit can input the information that the determination unit has determined the call to be secure into the AI, and the AI ​​will then ring the incoming call.

[0088] When analyzing telephone content, the analysis unit can infer user emotions and adjust the analysis accuracy based on these inferences. For example, when a user is nervous, the AI ​​detects emotions and can focus on specific keywords to improve analysis accuracy. When a user is relaxed, the AI ​​detects emotions and adjusts the analysis accuracy to analyze a wider range of content. When a user is excited, the AI ​​detects emotions and adjusts the analysis accuracy to highlight keywords with a high probability of being scammed. Thus, by adjusting the analysis accuracy based on user emotions, the analysis accuracy is improved. User emotions can be inferred based on factors such as tone of voice, word choice, and speaking style, but are not limited to these. Emotion inference can be achieved through emotion inference functions such as emotion engines or generative AI. Generative AI can be text-generated AI (such as LLM) or multimodal generative AI, but is not limited to these. Some or all of the above processing in the analysis unit can be performed using generative AI, or it can be performed without generative AI. For example, the analysis unit can input telephone voice data into generative AI, which will then perform user emotion inference.

[0089] When converting telephone conversations into text using speech recognition technology, the parsing unit can handle different languages ​​or dialects. For example, the parsing unit can use AI to automatically detect different languages ​​and select an appropriate speech recognition model for text conversion. The parsing unit can also identify dialectal or regionally specific vocabulary and use a dictionary to convert them into the standard language. Furthermore, the parsing unit can adjust the speech recognition accuracy and perform text conversion based on the user-defined language. Thus, by handling different languages ​​or dialects, the parsing accuracy is improved. Different languages ​​or dialects include, for example, multilingual speech recognition engines and dialect dictionaries, but are not limited to these. Some or all of the above processing in the parsing unit can be performed using generative AI, or it can be performed without generative AI. For example, the parsing unit can input telephone voice data into generative AI, which will then perform text conversion corresponding to different languages ​​or dialects.

[0090] When parsing telephone content, the parsing unit can remove background noise and other sounds to improve parsing accuracy. For example, the parsing unit can use AI to filter background noise in real time, extracting only the main speech for parsing. The parsing unit can also utilize noise reduction techniques to make the telephone content clearer. Furthermore, the parsing unit can emphasize specific frequency bands to reduce noise and improve speech recognition accuracy. Thus, by removing background noise and other sounds, parsing accuracy is improved. Background noise and other sounds include, but are not limited to, noise reduction techniques and filtering techniques. Some or all of the above processing in the parsing unit can be performed using generative AI, or it can be performed without generative AI. For example, the parsing unit can input telephone voice data into generative AI, which will then perform the removal of background noise and other sounds.

[0091] When analyzing telephone content, the analysis unit can infer the user's emotions and adjust the display of the analysis results accordingly. For example, when the user is nervous, the analysis unit can display concise results, highlighting only important information. When the user is relaxed, the analysis unit can display detailed results and provide additional information. When the user is excited, the analysis unit can display results intuitively, highlighting important keywords. Thus, by adjusting the display of analysis results based on user emotions, the understanding of the analysis results is deepened. User emotions can be inferred based on factors such as tone of voice, word choice, and speaking style, but are not limited to these. Emotion inference can be achieved through emotion engines or generative AI functions. Generative AI can be text-based AI (such as LLM) or multimodal AI, but is not limited to these. Some or all of the above processing in the analysis unit can be performed using generative AI, or it can be performed without generative AI. For example, the analysis unit can input telephone voice data into generative AI, which will then perform user emotion inference.

[0092] When the parsing unit converts telephone conversations into text using speech recognition technology, it can highlight specific keywords or phrases. For example, the parsing unit can automatically detect keywords that may indicate fraud and highlight them in the text. It can also highlight important phrases such as the names of family members or friends. Furthermore, it can highlight specific keywords set by the user. Thus, by highlighting specific keywords or phrases, important information is less likely to be overlooked. Specific keywords or phrases include, but are not limited to, words such as "important," "urgent," and "confirm." Some or all of the above processing in the parsing unit can be performed using generative AI, or it can be performed without generative AI. For example, the parsing unit can input telephone voice data into generative AI, which can then perform the highlighting of specific keywords or phrases.

[0093] When analyzing phone call content, the analysis unit can refer to past call records to improve analysis accuracy. For example, the analysis unit can analyze past call records, identifying specific frequently used phrases or keywords. The analysis unit can also learn specific patterns from past call records to improve analysis accuracy. Furthermore, the analysis unit can identify high-risk fraudulent calls based on past call records to improve analysis accuracy. Thus, by referring to past call records, analysis accuracy is improved. Past call records include, but are not limited to, factors such as the retention period of call content and the frequency of reference. Some or all of the above processing in the analysis unit can be performed using generative AI, or it can be performed without generative AI. For example, the analysis unit can input past call record data into the generative AI, which will then perform call record analysis.

[0094] The judgment unit can infer user emotions and adjust the criteria for determining safe calls based on these inferences. For example, when a user is nervous, the judgment unit can tighten the criteria, more strictly judging calls as potentially fraudulent. When a user is relaxed, the judgment unit can relax the criteria, prioritizing calls from family or friends as safe. When a user is excited, the judgment unit can adjust the criteria based on important keywords. Thus, by adjusting the judgment criteria according to user emotions, the accuracy of the judgment is improved. User emotions can be inferred based on factors such as tone of voice, word choice, and speaking style, but are not limited to these. Emotion inference can be achieved through emotion engines or generative AI functions. Generative AI can be text-generated AI (such as LLM) or multimodal generative AI, but is not limited to these. Some or all of the above processing in the judgment unit can be performed using generative AI, or it can be performed without generative AI. For example, the judgment unit can input telephone voice data into generative AI, which will then perform user emotion inference.

[0095] When determining whether a family member's or friend's phone call is safe, the judgment unit can refer to call logs and contact lists to improve the accuracy of the judgment. For example, the judgment unit can analyze call logs to identify frequently called individuals as safe calls. The judgment unit can also refer to contact lists to identify registered family members or friends as safe calls. The judgment unit can further improve the accuracy by combining call logs and contact lists. Thus, by referring to call logs and contact lists, the accuracy of the judgment is improved. Call logs and contact lists include, but are not limited to, factors such as call frequency and contact trustworthiness. Some or all of the above processing in the judgment unit can be performed using a generative AI, or it can be performed without using a generative AI. For example, the judgment unit can input call log data and contact lists into a generative AI, which will then perform call log and contact list analysis.

[0096] When the judgment department determines that a call containing keywords suggesting potential fraud is unsafe, it can update its judgment criteria by referencing the latest information on fraud methods. For example, the judgment department can periodically update the latest fraud methods and reflect this in the judgment criteria. The judgment department can also automatically update the list of keywords suggesting potential fraud and incorporate them into the judgment criteria. The judgment department can also adjust its judgment criteria by referring to news or reports about fraud methods. Thus, by referencing the latest information on fraud methods, the judgment criteria remain up-to-date. The latest information on fraud methods includes, for example, police databases, security company reports, etc., but is not limited to these. Some or all of the above processing in the judgment department can be performed using AI, or it can be performed without AI. For example, the judgment department can input the latest information on fraud methods into the AI, which will then update the judgment criteria.

[0097] The judgment unit can infer user emotions and adjust the notification method based on the inferred user emotions. For example, when the user is nervous, the judgment unit can provide a concise and clear notification. When the user is relaxed, the judgment unit can provide a detailed notification with additional information. When the user is excited, the judgment unit can provide an intuitive and easy-to-understand notification. Thus, by adjusting the notification method according to the user's emotions, the understanding of the notification is deepened. User emotions are inferred, for example, based on tone of voice, word choice, and speaking style, but are not limited to these. Emotion inference is achieved, for example, through emotion inference functions such as emotion engines or generative AI. Generative AI can be text generation AI (such as LLM) or multimodal generation AI, but is not limited to these. Some or all of the above processing in the judgment unit can be performed using generative AI, or it can be performed without generative AI. For example, the judgment unit can input telephone voice data into generative AI, which will then perform user emotion inference.

[0098] When determining whether a call from family or friends is safe, the judgment unit can consider call frequency and time period to improve judgment accuracy. For example, the judgment unit can analyze call frequency to classify frequently called numbers as safe. The judgment unit can also consider call time period to classify calls received during typical time periods as safe. The judgment unit can further improve judgment accuracy by combining call frequency and time period. Thus, by considering call frequency and time period, judgment accuracy is improved. Call frequency and time period include, for example, the number of calls and the duration of calls, but are not limited to these. Some or all of the above processing in the judgment unit can be performed using a generative AI, or it can be performed without using a generative AI. For example, the judgment unit can input call record data into a generative AI, which will then perform call frequency and time period analysis.

[0099] When the judgment unit determines a call containing keywords suggesting potential fraud as unsafe, it can consider not only the call content but also the sender's information. For example, the judgment unit can analyze the sender's phone number, identifying specific high-risk numbers. It can also refer to the sender's location information, identifying calls from specific high-risk areas. The judgment unit can further improve the accuracy of its judgment by combining sender information and call content. Thus, by considering sender information, the accuracy of the judgment is improved. Sender information includes, for example, the sender's phone number and location, but is not limited to these. Some or all of the above processing in the judgment unit can be performed using generative AI, or it can be performed without generative AI. For example, the judgment unit can input sender information into generative AI, which will then perform sender information analysis.

[0100] The ringing unit can infer the user's emotions and adjust the type and volume of the ringtone accordingly. For example, when the user is nervous, the ringing unit can set a calm and steady ringtone. When the user is relaxed, the ringing unit can set a bright and cheerful ringtone. When the user is excited, the ringing unit can adjust the volume to set a suitable ringtone. Thus, by adjusting the type and volume of the ringtone based on the user's emotions, a more appropriate ringtone can be set. User emotions can be inferred, for example, based on tone of voice, word choice, and speaking style, but are not limited to these. Emotion inference can be achieved, for example, through emotion inference functions such as emotion engines or generative AI. Generative AI can be text-generated AI (such as LLM) or multimodal generative AI, but is not limited to these. Some or all of the above processing in the ringing unit can be performed using generative AI, or it can be performed without generative AI. For example, the ringing unit can input user emotion data into generative AI, which will then perform the adjustment of the ringtone type and volume.

[0101] The ringing department, which only rings for secure calls, can adjust the ringing timing by considering the user's schedule and activity level. For example, it can refer to the user's calendar information to avoid ringing during important meetings. It can also monitor the user's activity in real time to avoid ringing during exercise or sleep. Furthermore, the ringing department can select the optimal time to ring based on the user's schedule. Thus, it can ring at the best time based on the user's schedule and activity level. User schedules and activity levels include, but are not limited to, calendar applications and activity trackers. Some or all of the above processing in the ringing department can be performed using AI, or it can be performed without AI. For example, the ringing department can input the user's schedule information into the AI, which will then adjust the ringing timing.

[0102] When the ringing unit emits a ringtone, it can provide a customizable ringtone based on user preferences. For example, the ringing unit can set the user-selected music or sound effect as the ringtone. It can also provide ringtones that reflect the user's customized volume or timbre. Furthermore, the ringing unit can set different ringtones for specific contacts set by the user. Thus, by setting ringtones according to user preferences, a more comfortable user experience is provided. User preferences include, but are not limited to, user settings and past selection records. Some or all of the above processing in the ringing unit can be performed using AI generation, or it can be performed without AI generation. For example, the ringing unit can input user preference data into AI generation, which will then provide a customizable ringtone.

[0103] The ringing unit can infer the user's emotions and adjust the ringing time of the incoming call based on the inferred emotions. For example, the ringing unit can set a shorter ringing time when the user is nervous. A longer ringing time can be set when the user is relaxed. A suitable ringing time can be set when the user is excited. Therefore, by adjusting the ringing time of the incoming call based on the user's emotions, a more suitable ringtone can be set. User emotions can be inferred based on factors such as tone of voice, word choice, and speaking style, but are not limited to these. Emotion inference can be achieved through emotion inference functions such as emotion engines or generative AI. Generative AI can be text-generated AI (such as LLM) or multimodal generative AI, but is not limited to these. Some or all of the above processing in the ringing unit can be performed using generative AI, or it can be performed without generative AI. For example, the ringing unit can input user emotion data into generative AI, which will then adjust the ringing time of the incoming call.

[0104] The ringing unit, when ringing only for secure calls, can consider user device settings and ambient sound to set the optimal volume. For example, the ringing unit can automatically set a suitable volume based on user device settings. The ringing unit can also detect ambient sound in real time and set the optimal volume accordingly. The ringing unit can also customize the volume based on user preferences. Therefore, by setting the optimal volume based on user device settings and ambient sound, a more suitable ringtone can be set. User device settings and ambient sound include, but are not limited to, device volume settings and ambient noise levels. Some or all of the above processing in the ringing unit can be performed using a generated AI, or it can be performed without using a generated AI. For example, the ringing unit can input user device settings and ambient sound data into the generated AI, which will then perform the optimal volume setting.

[0105] When the ringing unit issues a ring, it can adjust the ringing pattern by referring to the user's past response records. For example, the ringing unit can analyze the user's past response time periods and adjust the ringing pattern accordingly. The ringing unit can also learn the optimal ringing pattern from the user's past response records and apply it. The ringing unit can also set a corresponding ringing pattern when the user is likely to respond during specific time periods. Thus, by referring to the user's past response records, an optimal ringing pattern can be set. The user's past response records include, for example, response frequency, response time periods, etc., but are not limited to these. Some or all of the above processing in the ringing unit can be performed using a generation AI, or it can be performed without using a generation AI. For example, the ringing unit can input the user's past response record data into the generation AI, which will then perform the ringing pattern adjustment.

[0106] The system involved in this embodiment is not limited to the examples described above. For example, various modifications can be made as shown below.

[0107] When analyzing telephone content, the analysis unit can infer user emotions and adjust the analysis accuracy based on these inferences. For example, when a user is nervous, the AI ​​detects emotions and can focus on specific keywords to improve analysis accuracy. When a user is relaxed, the AI ​​detects emotions and adjusts the analysis accuracy to analyze a wider range of content. When a user is excited, the AI ​​detects emotions and adjusts the analysis accuracy to highlight keywords with a high probability of being scammed. Thus, by adjusting the analysis accuracy based on user emotions, the analysis accuracy is improved. User emotions can be inferred based on factors such as tone of voice, word choice, and speaking style, but are not limited to these. Emotion inference can be achieved through emotion inference functions such as emotion engines or generative AI. Generative AI can be text-generated AI (such as LLM) or multimodal generative AI, but is not limited to these. Some or all of the above processing in the analysis unit can be performed using generative AI, or it can be performed without generative AI. For example, the analysis unit can input telephone voice data into generative AI, which will then perform user emotion inference.

[0108] When converting telephone conversations into text using speech recognition technology, the parsing unit can handle different languages ​​or dialects. For example, the parsing unit can use AI to automatically detect different languages ​​and select an appropriate speech recognition model for text conversion. The parsing unit can also identify dialectal or regionally specific vocabulary and use a dictionary to convert them into the standard language. Furthermore, the parsing unit can adjust the speech recognition accuracy and perform text conversion based on the user-defined language. Thus, by handling different languages ​​or dialects, the parsing accuracy is improved. Different languages ​​or dialects include, for example, multilingual speech recognition engines and dialect dictionaries, but are not limited to these. Some or all of the above processing in the parsing unit can be performed using generative AI, or it can be performed without generative AI. For example, the parsing unit can input telephone voice data into generative AI, which will then perform text conversion corresponding to different languages ​​or dialects.

[0109] The judgment unit can infer user emotions and adjust the criteria for determining safe calls based on these inferences. For example, when a user is nervous, the judgment unit can tighten the criteria, more strictly judging calls as potentially fraudulent. When a user is relaxed, the judgment unit can relax the criteria, prioritizing calls from family or friends as safe. When a user is excited, the judgment unit can adjust the criteria based on important keywords. Thus, by adjusting the judgment criteria according to user emotions, the accuracy of the judgment is improved. User emotions can be inferred based on factors such as tone of voice, word choice, and speaking style, but are not limited to these. Emotion inference can be achieved through emotion engines or generative AI functions. Generative AI can be text-generated AI (such as LLM) or multimodal generative AI, but is not limited to these. Some or all of the above processing in the judgment unit can be performed using generative AI, or it can be performed without generative AI. For example, the judgment unit can input telephone voice data into generative AI, which will then perform user emotion inference.

[0110] When parsing telephone content, the parsing unit can remove background noise and other sounds to improve parsing accuracy. For example, the parsing unit can use AI to filter background noise in real time, extracting only the main speech for parsing. The parsing unit can also utilize noise reduction techniques to make the telephone content clearer. Furthermore, the parsing unit can emphasize specific frequency bands to reduce noise and improve speech recognition accuracy. Thus, by removing background noise and other sounds, parsing accuracy is improved. Background noise and other sounds include, but are not limited to, noise reduction techniques and filtering techniques. Some or all of the above processing in the parsing unit can be performed using generative AI, or it can be performed without generative AI. For example, the parsing unit can input telephone voice data into generative AI, which will then perform the removal of background noise and other sounds.

[0111] When the judgment department determines that a call containing keywords suggesting potential fraud is unsafe, it can update its judgment criteria by referencing the latest information on fraud methods. For example, the judgment department can periodically update the latest fraud methods and reflect this in the judgment criteria. The judgment department can also automatically update the list of keywords suggesting potential fraud and incorporate them into the judgment criteria. The judgment department can also adjust its judgment criteria by referring to news or reports about fraud methods. Thus, by referencing the latest information on fraud methods, the judgment criteria remain up-to-date. The latest information on fraud methods includes, for example, police databases, security company reports, etc., but is not limited to these. Some or all of the above processing in the judgment department can be performed using AI, or it can be performed without AI. For example, the judgment department can input the latest information on fraud methods into the AI, which will then update the judgment criteria.

[0112] The ringing unit can infer the user's emotions and adjust the type and volume of the ringtone accordingly. For example, when the user is nervous, the ringing unit can set a calm and steady ringtone. When the user is relaxed, the ringing unit can set a bright and cheerful ringtone. When the user is excited, the ringing unit can adjust the volume to set a suitable ringtone. Thus, by adjusting the type and volume of the ringtone based on the user's emotions, a more appropriate ringtone can be set. User emotions can be inferred, for example, based on tone of voice, word choice, and speaking style, but are not limited to these. Emotion inference can be achieved, for example, through emotion inference functions such as emotion engines or generative AI. Generative AI can be text-generated AI (such as LLM) or multimodal generative AI, but is not limited to these. Some or all of the above processing in the ringing unit can be performed using generative AI, or it can be performed without generative AI. For example, the ringing unit can input user emotion data into generative AI, which will then perform the adjustment of the ringtone type and volume.

[0113] The ringing department, which only rings for secure calls, can adjust the ringing timing by considering the user's schedule and activity level. For example, it can refer to the user's calendar information to avoid ringing during important meetings. It can also monitor the user's activity in real time to avoid ringing during exercise or sleep. Furthermore, the ringing department can select the optimal time to ring based on the user's schedule. Thus, it can ring at the best time based on the user's schedule and activity level. User schedules and activity levels include, but are not limited to, calendar applications and activity trackers. Some or all of the above processing in the ringing department can be performed using AI, or it can be performed without AI. For example, the ringing department can input the user's schedule information into the AI, which will then adjust the ringing timing.

[0114] When analyzing phone call content, the analysis unit can refer to past call records to improve analysis accuracy. For example, the analysis unit can analyze past call records, identifying specific frequently used phrases or keywords. The analysis unit can also learn specific patterns from past call records to improve analysis accuracy. Furthermore, the analysis unit can identify high-risk fraudulent calls based on past call records to improve analysis accuracy. Thus, by referring to past call records, analysis accuracy is improved. Past call records include, but are not limited to, factors such as the retention period of call content and the frequency of reference. Some or all of the above processing in the analysis unit can be performed using generative AI, or it can be performed without generative AI. For example, the analysis unit can input past call record data into the generative AI, which will then perform call record analysis.

[0115] The judgment unit can infer user emotions and adjust the notification method based on the inferred user emotions. For example, when the user is nervous, the judgment unit can provide a concise and clear notification. When the user is relaxed, the judgment unit can provide a detailed notification with additional information. When the user is excited, the judgment unit can provide an intuitive and easy-to-understand notification. Thus, by adjusting the notification method according to the user's emotions, the understanding of the notification is deepened. User emotions are inferred, for example, based on tone of voice, word choice, and speaking style, but are not limited to these. Emotion inference is achieved, for example, through emotion inference functions such as emotion engines or generative AI. Generative AI can be text generation AI (such as LLM) or multimodal generation AI, but is not limited to these. Some or all of the above processing in the judgment unit can be performed using generative AI, or it can be performed without generative AI. For example, the judgment unit can input telephone voice data into generative AI, which will then perform user emotion inference.

[0116] The ringing unit, when ringing only for secure calls, can consider user device settings and ambient sound to set the optimal volume. For example, the ringing unit can automatically set a suitable volume based on user device settings. The ringing unit can also detect ambient sound in real time and set the optimal volume accordingly. The ringing unit can also customize the volume based on user preferences. Therefore, by setting the optimal volume based on user device settings and ambient sound, a more suitable ringtone can be set. User device settings and ambient sound include, but are not limited to, device volume settings and ambient noise levels. Some or all of the above processing in the ringing unit can be performed using a generated AI, or it can be performed without using a generated AI. For example, the ringing unit can input user device settings and ambient sound data into the generated AI, which will then perform the optimal volume setting.

[0117] The following is a brief description of the processing flow of Implementation Method 2.

[0118] Step 1: The parsing unit parses the telephone content. The parsing unit uses speech recognition technology to convert the telephone content into text. For example, deep learning-based speech recognition technology or HMM-based speech recognition technology can be used to convert the telephone content into text with high accuracy and in real time.

[0119] Step 2: The judgment department determines whether a call is safe based on the analysis provided by the analysis department. Calls from family or friends, or from people registered in the contact list, are considered safe. Calls containing keywords suggesting potential fraud are deemed unsafe, and the department may update these keywords based on the latest information on fraud methods.

[0120] Step 3: The ringing unit sounds a ringing tone when the determination unit determines the phone number to be secure. For example, the ringing tone may only sound when a family member or friend calls, or when a specific phone number calls. Furthermore, a specific ringing tone may be played when the phone number is determined to be secure.

[0121] The specific processing unit 290 sends the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires voice representing the user's input to the result of the specific processing. The control unit 46A sends the voice data representing the user's input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the voice data.

[0122] Data generation model 58 is what is known as generative AI (Artificial Intelligence). An example of data generation model 58 includes ChatGPT (registered trademark) (Internet search).<URL:https: / / openai.com / blog / chatgpt> Generative AI, such as data generation model 58, is obtained by deep learning through a neural network. The data generation model 58 is input with a prompt containing instructions, and with inference data such as speech data representing speech, text data representing text, and image data representing images (e.g., data of still images or data of moving images). The data generation model 58 infers the input inference data according to the instructions shown in the prompt, and outputs the inference result in one or more data forms such as speech data, text data, and image data. The data generation model 58 includes, for example, text generation AI, image generation AI, and multimodal generation AI. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization. The specific processing unit 290 performs the above-described specific processing while using the data generation model 58. The data generation model 58 can also be a fine-tuned model to output inference results from prompts without instructions; in this case, the data generation model 58 can output inference results from prompts without instructions. The data processing device 12, etc., includes multiple data generation models 58, including AI other than generative AI. AI other than generative AI includes, but is not limited to, linear regression, logistic regression, decision trees, random forests, support vector machines (SVM), k-means clustering, convolutional neural networks (CNN), recurrent neural networks (RNN), generative adversarial networks (GAN), or Naive Bayes. Furthermore, AI can also act as an AI agent. Additionally, when the processing described above is performed by AI, this processing may be partially or entirely performed by AI, but is not limited to this example. Moreover, processing performed by AI, including generative AI, can be replaced by rule-based processing, and vice versa.

[0123] Furthermore, the processing performed by the aforementioned data processing system 10 is executed by the specific processing unit 290 of the data processing device 12 or the control unit 46A of the smart device 14, but it can also be executed jointly by the specific processing unit 290 of the data processing device 12 and the control unit 46A of the smart device 14. Additionally, the specific processing unit 290 of the data processing device 12 acquires or collects information required for processing from the smart device 14 or external devices, and the smart device 14 acquires or collects information required for processing from the data processing device 12 or external devices.

[0124] Each of the elements, including the parsing unit, the determination unit, and the ringing unit, can be implemented, for example, in at least one of the smart device 14 and the data processing device 12. For instance, the parsing unit is implemented by the processor 46 of the smart device 14, which uses speech recognition technology to convert the telephone content into text. The determination unit is implemented, for example, by the specific processing unit 290 of the data processing device 12, which determines whether the telephone is secure based on the parsed content. The ringing unit is implemented, for example, by the control unit 46A of the smart device 14, which rings the incoming call when the telephone is determined to be secure. The correspondence between each element and the device or control unit is not limited to the above examples and can be varied.

[0125] Second Implementation Method

[0126] Figure 3 An example of the configuration of the data processing system 210 according to the second embodiment is shown.

[0127] like Figure 3 As shown, the data processing system 210 includes a data processing device 12 and smart glasses 214. One example of the data processing device 12 is a server.

[0128] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 includes a processor 28, RAM 30, and memory 32. The processor 28, RAM 30, and memory 32 are connected to a bus 34. Furthermore, the database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. An example of the network 54 includes a WAN and / or LAN, etc.

[0129] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, RAM 48, and memory 50. The processor 46, RAM 48, and memory 50 are connected to a bus 52. Furthermore, the microphone 238, speaker 240, and camera 42 are also connected to the bus 52.

[0130] Microphone 238 receives user commands by receiving the user's voice. Microphone 238 captures the user's voice, converts the captured sound into speech data, and outputs it to processor 46. Speaker 240 outputs sound according to commands from processor 46.

[0131] Camera 42 is a small digital camera equipped with an optical system such as a lens, aperture and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, used to capture the user's surroundings (e.g., the shooting range defined by an angle of view equivalent to the field of vision of an average healthy person).

[0132] Communication I / F 44 is connected to network 54. Communication I / F 44 and 26 are responsible for the transmission and reception of various information between processor 46 and processor 28 via network 54. The transmission and reception of various information between processor 46 and processor 28 using communication I / F 44 and 26 is performed in a secure state.

[0133] Figure 4 An example of the main functions of the data processing device 12 and the smart glasses 214 is shown. Figure 4 As shown, in the data processing device 12, specific processing is performed by the processor 28. The memory 32 stores a specific processing program 56.

[0134] The processor 28 reads a specific processing program 56 from the memory 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is implemented by the processor 28 operating as a specific processing unit 290 based on the specific processing program 56 executed on the RAM 30.

[0135] The memory 32 stores a data generation model 58 and an emotion-specific model 59. The data generation model 58 and the emotion-specific model 59 are used by the specific processing unit 290. The specific processing unit 290 can use the emotion-specific model 59 to infer the user's emotions and perform specific processing based on the user's emotions. The emotion inference function using the emotion-specific model 59 (emotion-specific function) includes inferring and predicting the user's emotions, performing various inferences and predictions related to the user's emotions, but is not limited to this example. Furthermore, emotion inference and prediction also include, for example, emotion analysis (analysis).

[0136] In the smart glasses 214, specific processing is performed by the processor 46. A specific processing program 60 is stored in the memory 50. The processor 46 reads the specific processing program 60 from the memory 50 and executes the read specific processing program 60 on the RAM 48. Specific processing is implemented by the processor 46 operating as a control unit 46A based on the specific processing program 60 executed on the RAM 48. Furthermore, the smart glasses 214 may also have the same data generation model and emotion-specific model as the data generation model 58 and the emotion-specific model 59, and use these models to perform the same processing as the specific processing unit 290.

[0137] Furthermore, other devices besides the data processing device 12 may also have the data generation model 58. For example, a server device may have the data generation model 58. In this case, the data processing device 12 obtains the processing results (prediction results, etc.) using the data generation model 58 by communicating with the server device that has the data generation model 58. In addition, the data processing device 12 may be a server device or a user-owned terminal device (e.g., a mobile phone, robot, home appliance, etc.).

[0138] The specific processing unit 290 sends the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires voice input representing the user's input to the specific processing result. The control unit 46A sends the voice data representing the user's input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the voice data.

[0139] Data generation model 58 is a so-called generative AI. An example of data generation model 58 includes generative AIs such as ChatGPT. Data generation model 58 is obtained by performing deep learning on a neural network. A prompt containing instructions is input to data generation model 58, along with inference data such as speech data representing speech, text data representing text, and image data representing images (e.g., data from still images or moving images). Data generation model 58 infers the inference data based on the instructions shown in the prompt and outputs the inference result in one or more data forms such as speech data, text data, and image data. Data generation model 58 includes, for example, text generation AI, image generation AI, and multimodal generation AI. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization. Specific processing unit 290 performs the aforementioned specific processing while using data generation model 58. Data generation model 58 can also be a fine-tuned model to output inference results from prompts without instructions; in this case, data generation model 58 can output inference results from prompts without instructions. The data processing device 12, etc., includes various data generation models 58, which include AI other than generative AI. These AIs include, for example, linear regression, logistic regression, decision trees, random forests, support vector machines (SVM), k-means clustering, convolutional neural networks (CNN), recurrent neural networks (RNN), generative adversarial networks (GAN), or Naive Bayes, and are capable of various processing methods, but are not limited to these examples. Furthermore, AI can also be an AI agent. Additionally, when the processing described above is performed by AI, this processing may be partially or entirely performed by AI, but is not limited to these examples. Moreover, processing performed by AI including generative AI can be replaced by rule-based processing, and vice versa.

[0140] The data processing system 210 of the second embodiment performs the same processing as the data processing system 10 of the first embodiment. The processing performed by the data processing system 210 is executed by the specific processing unit 290 of the data processing device 12 or the control unit 46A of the smart glasses 214, but it can also be executed jointly by the specific processing unit 290 of the data processing device 12 and the control unit 46A of the smart glasses 214. Furthermore, the specific processing unit 290 of the data processing device 12 acquires or collects information required for processing from the smart glasses 214 or external devices, and the smart glasses 214 acquires or collects information required for processing from the data processing device 12 or external devices.

[0141] Each of the elements, including the analysis unit, determination unit, and ringing unit, can be implemented, for example, in at least one of the smart glasses 214 and the data processing device 12. For instance, the analysis unit is implemented by the processor 46 of the smart glasses 214, which uses speech recognition technology to convert the telephone content into text. The determination unit is implemented, for example, by the specific processing unit 290 of the data processing device 12, which determines whether the telephone is secure based on the analyzed content. The ringing unit is implemented, for example, by the control unit 46A of the smart glasses 214, which rings the incoming call when the telephone is determined to be secure. The correspondence between each element and the device or control unit is not limited to the above example and can be varied.

[0142] Third Implementation Method

[0143] Figure 5 An example of the configuration of the data processing system 310 according to the third embodiment is shown.

[0144] like Figure 5 As shown, the data processing system 310 includes a data processing device 12 and a head-mounted terminal 314. One example of the data processing device 12 is a server.

[0145] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 includes a processor 28, RAM 30, and memory 32. The processor 28, RAM 30, and memory 32 are connected to a bus 34. Furthermore, the database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. An example of the network 54 includes a WAN and / or LAN, etc.

[0146] The head-mounted terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a display 343. The computer 36 includes a processor 46, RAM 48, and memory 50. The processor 46, RAM 48, and memory 50 are connected to a bus 52. Furthermore, the microphone 238, speaker 240, camera 42, and display 343 are also connected to the bus 52.

[0147] Microphone 238 receives user commands by receiving the user's voice. Microphone 238 captures the user's voice, converts the captured sound into speech data, and outputs it to processor 46. Speaker 240 outputs sound according to commands from processor 46.

[0148] Camera 42 is a small digital camera equipped with an optical system such as a lens, aperture and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, used to capture the user's surroundings (e.g., the shooting range defined by an angle of view equivalent to the field of vision of an average healthy person).

[0149] Communication I / F 44 is connected to network 54. Communication I / F 44 and 26 are responsible for the transmission and reception of various information between processor 46 and processor 28 via network 54. The transmission and reception of various information between processor 46 and processor 28 using communication I / F 44 and 26 is performed in a secure state.

[0150] Figure 6 An example of the main functions of the data processing device 12 and the head-mounted terminal 314 is shown. Figure 6 As shown, in the data processing device 12, specific processing is performed by the processor 28. The memory 32 stores a specific processing program 56.

[0151] The processor 28 reads a specific processing program 56 from the memory 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is implemented by the processor 28 operating as a specific processing unit 290 based on the specific processing program 56 executed on the RAM 30.

[0152] The memory 32 stores a data generation model 58 and an emotion-specific model 59. The data generation model 58 and the emotion-specific model 59 are used by the specific processing unit 290. The specific processing unit 290 can use the emotion-specific model 59 to infer the user's emotions and perform specific processing based on the user's emotions. The emotion inference function using the emotion-specific model 59 (emotion-specific function) includes inferring and predicting the user's emotions, performing various inferences and predictions related to the user's emotions, but is not limited to this example. Furthermore, emotion inference and prediction also include, for example, emotion analysis (analysis).

[0153] In the head-mounted terminal 314, specific processing is performed by the processor 46. A specific program 60 is stored in the memory 50. The processor 46 reads the specific program 60 from the memory 50 and executes the read specific program 60 on the RAM 48. Specific processing is implemented by the processor 46 operating as a control unit 46A based on the specific program 60 executed on the RAM 48. Furthermore, the head-mounted terminal 314 may also have the same data generation model and emotion-specific model as the data generation model 58 and the emotion-specific model 59, and use these models to perform the same processing as the specific processing unit 290.

[0154] Furthermore, other devices besides the data processing device 12 may also have the data generation model 58. For example, a server device may have the data generation model 58. In this case, the data processing device 12 obtains the processing results (prediction results, etc.) using the data generation model 58 by communicating with the server device that has the data generation model 58. In addition, the data processing device 12 may be a server device or a user-owned terminal device (e.g., a mobile phone, robot, home appliance, etc.).

[0155] The specific processing unit 290 sends the result of the specific processing to the head-mounted terminal 314. In the head-mounted terminal 314, the control unit 46A causes the speaker 240 and display 343 to output the result of the specific processing. The microphone 238 acquires voice representing the user's input to the specific processing result. The control unit 46A sends the voice data representing the user's input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the voice data.

[0156] Data generation model 58 is a so-called generative AI. An example of data generation model 58 includes generative AIs such as ChatGPT. Data generation model 58 is obtained by performing deep learning on a neural network. A prompt containing instructions is input to data generation model 58, along with inference data such as speech data representing speech, text data representing text, and image data representing images (e.g., data from still images or moving images). Data generation model 58 infers the inference data based on the instructions shown in the prompt and outputs the inference result in one or more data forms such as speech data, text data, and image data. Data generation model 58 includes, for example, text generation AI, image generation AI, and multimodal generation AI. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization. Specific processing unit 290 performs the aforementioned specific processing while using data generation model 58. Data generation model 58 can also be a fine-tuned model to output inference results from prompts without instructions; in this case, data generation model 58 can output inference results from prompts without instructions. The data processing device 12, etc., includes various data generation models 58, which include AI other than generative AI. These AIs include, for example, linear regression, logistic regression, decision trees, random forests, support vector machines (SVM), k-means clustering, convolutional neural networks (CNN), recurrent neural networks (RNN), generative adversarial networks (GAN), or Naive Bayes, and are capable of various processing methods, but are not limited to these examples. Furthermore, AI can also be an AI agent. Additionally, when the processing described above is performed by AI, this processing may be partially or entirely performed by AI, but is not limited to these examples. Moreover, processing performed by AI including generative AI can be replaced by rule-based processing, and vice versa.

[0157] The data processing system 310 of the third embodiment performs the same processing as the data processing system 10 of the first embodiment. The processing performed by the data processing system 310 is executed by the specific processing unit 290 of the data processing device 12 or the control unit 46A of the head-mounted terminal 314, but it can also be executed jointly by the specific processing unit 290 of the data processing device 12 and the control unit 46A of the head-mounted terminal 314. Furthermore, the specific processing unit 290 of the data processing device 12 acquires or collects information required for processing from the head-mounted terminal 314 or external devices, and the head-mounted terminal 314 acquires or collects information required for processing from the data processing device 12 or external devices.

[0158] Each of the elements, including the analysis unit, determination unit, and ringing unit, can be implemented, for example, in at least one of the headset 314 and the data processing device 12. For instance, the analysis unit is implemented by the processor 46 of the headset 314, which uses speech recognition technology to convert the telephone content into text. The determination unit is implemented, for example, by the specific processing unit 290 of the data processing device 12, which determines whether the telephone is secure based on the analyzed content. The ringing unit is implemented, for example, by the control unit 46A of the headset 314, which rings the incoming call when the telephone is determined to be secure. The correspondence between each element and the device or control unit is not limited to the above examples and can be varied.

[0159] Fourth Implementation Method

[0160] Figure 7 An example of the configuration of the data processing system 410 according to the fourth embodiment is shown.

[0161] like Figure 7 As shown, the data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.

[0162] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 includes a processor 28, RAM 30, and memory 32. The processor 28, RAM 30, and memory 32 are connected to a bus 34. Furthermore, the database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. An example of the network 54 includes a WAN and / or LAN, etc.

[0163] Robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a control object 443. Computer 36 includes a processor 46, RAM 48, and memory 50. The processor 46, RAM 48, and memory 50 are connected to a bus 52. Furthermore, the microphone 238, speaker 240, camera 42, and control object 443 are also connected to the bus 52.

[0164] Microphone 238 receives user commands by receiving the user's voice. Microphone 238 captures the user's voice, converts the captured sound into speech data, and outputs it to processor 46. Speaker 240 outputs sound according to commands from processor 46.

[0165] The camera 42 is a small digital camera equipped with an optical system such as a lens, aperture and shutter, and an imaging element such as a CMOS image sensor or a CCD image sensor, used to photograph the user's surroundings (e.g., the shooting range defined by an angle of view equivalent to the field of vision of an average healthy person).

[0166] Communication I / F 44 is connected to network 54. Communication I / F 44 and 26 are responsible for the transmission and reception of various information between processor 46 and processor 28 via network 54. The transmission and reception of various information between processor 46 and processor 28 using communication I / F 44 and 26 is performed in a secure state.

[0167] The controlled object 443 includes a display device, LEDs for the eyes, and motors for driving the arms, hands, and feet. The posture and movements of the robot 414 are controlled by controlling the motors for the arms, hands, and feet. Some of the robot 414's emotions can be expressed by controlling these motors. In addition, facial expressions of the robot 414 can also be expressed by controlling the illumination state of the LEDs for the robot 414's eyes.

[0168] Figure 8 An example of the main functions of the data processing device 12 and the robot 414 is shown. For example... Figure 8 As shown, in the data processing device 12, specific processing is performed by the processor 28. The memory 32 stores a specific processing program 56.

[0169] The processor 28 reads a specific processing program 56 from the memory 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is implemented by the processor 28 operating as a specific processing unit 290 based on the specific processing program 56 executed on the RAM 30.

[0170] The memory 32 stores a data generation model 58 and an emotion-specific model 59. The data generation model 58 and the emotion-specific model 59 are used by the specific processing unit 290. The specific processing unit 290 can use the emotion-specific model 59 to infer the user's emotions and perform specific processing based on the user's emotions. The emotion inference function using the emotion-specific model 59 (emotion-specific function) includes inferring and predicting the user's emotions, performing various inferences and predictions related to the user's emotions, but is not limited to this example. Furthermore, emotion inference and prediction also include, for example, emotion analysis (analysis).

[0171] In robot 414, specific processing is performed by processor 46. A specific program 60 is stored in memory 50. Processor 46 reads the specific program 60 from memory 50 and executes the read specific program 60 on RAM 48. Specific processing is achieved by processor 46 acting as control unit 46A based on the specific program 60 executed on RAM 48. Furthermore, robot 414 may also have the same data generation model and emotion-specific model as data generation model 58 and emotion-specific model 59, and use these models to perform the same processing as specific processing unit 290.

[0172] Furthermore, other devices besides the data processing device 12 may also have the data generation model 58. For example, a server device may have the data generation model 58. In this case, the data processing device 12 obtains the processing results (prediction results, etc.) using the data generation model 58 by communicating with the server device that has the data generation model 58. In addition, the data processing device 12 may be a server device or a user-owned terminal device (e.g., a mobile phone, robot, home appliance, etc.).

[0173] The specific processing unit 290 sends the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the controlled object 443 to output the result of the specific processing. The microphone 238 acquires voice representing the user's input regarding the result of the specific processing. The control unit 46A sends the voice data representing the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the voice data.

[0174] Data generation model 58 is a so-called generative AI. An example of data generation model 58 includes generative AIs such as ChatGPT. Data generation model 58 is obtained by performing deep learning on a neural network. A prompt containing instructions is input to data generation model 58, along with inference data such as speech data representing speech, text data representing text, and image data representing images (e.g., data from still images or moving images). Data generation model 58 infers the inference data based on the instructions shown in the prompt and outputs the inference result in one or more data forms such as speech data, text data, and image data. Data generation model 58 includes, for example, text generation AI, image generation AI, and multimodal generation AI. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization. Specific processing unit 290 performs the aforementioned specific processing while using data generation model 58. Data generation model 58 can also be a fine-tuned model to output inference results from prompts without instructions; in this case, data generation model 58 can output inference results from prompts without instructions. The data processing device 12, etc., includes various data generation models 58, which include AI other than generative AI. These AIs include, for example, linear regression, logistic regression, decision trees, random forests, support vector machines (SVM), k-means clustering, convolutional neural networks (CNN), recurrent neural networks (RNN), generative adversarial networks (GAN), or Naive Bayes, and are capable of various processing methods, but are not limited to these examples. Furthermore, AI can also be an AI agent. Additionally, when the processing described above is performed by AI, this processing may be partially or entirely performed by AI, but is not limited to these examples. Moreover, processing performed by AI including generative AI can be replaced by rule-based processing, and vice versa.

[0175] The data processing system 410 of the fourth embodiment performs the same processing as the data processing system 10 of the first embodiment. The processing performed by the data processing system 410 is executed by the specific processing unit 290 of the data processing device 12 or the control unit 46A of the robot 414, but it can also be executed jointly by the specific processing unit 290 of the data processing device 12 and the control unit 46A of the robot 414. Furthermore, the specific processing unit 290 of the data processing device 12 acquires or collects information required for processing from the robot 414 or external devices, and the robot 414 acquires or collects information required for processing from the data processing device 12 or external devices.

[0176] Each of the elements, including the analysis unit, determination unit, and ringing unit, can be implemented, for example, in at least one of the robot 414 and the data processing device 12. For instance, the analysis unit is implemented by the processor 46 of the robot 414, which uses speech recognition technology to convert the telephone content into text. The determination unit is implemented, for example, by the specific processing unit 290 of the data processing device 12, which determines whether the telephone is secure based on the analyzed content. The ringing unit is implemented, for example, by the control unit 46A of the robot 414, which rings the incoming call when the telephone is determined to be secure. The correspondence between the various units and the device or control unit is not limited to the above examples and can be modified in various ways.

[0177] Furthermore, the emotion-specific model 59, acting as an emotion engine, can determine the user's emotion based on a specific mapping. Specifically, the emotion-specific model 59 can determine the user's emotion based on an emotion graph that serves as a specific mapping (see...). Figure 9 The robot's emotions can be determined by the emotion-specific model 59. In addition, the emotion-specific model 59 can also determine the robot's emotions in the same way, and the specific processing unit 290 can also perform specific processing using the robot's emotions.

[0178] Figure 9 This is a diagram representing an emotion map 400 that maps various emotions. In the emotion map 400, emotions are arranged radially from the center in concentric circles. The closer to the center of the concentric circles, the more primitive the emotion is. Further out on the concentric circles, emotions are arranged representing states or actions arising from mood. Emotion is a concept that includes both feelings and mental states. To the left of the concentric circles, emotions generated by reactions occurring in the brain are arranged roughly. To the right of the concentric circles, emotions guided by situational judgments are arranged roughly. Above and below the concentric circles, emotions generated by reactions occurring in the brain and guided by situational judgments are arranged roughly. Furthermore, the emotion of "pleasure" is arranged above the concentric circles, and the emotion of "unpleasantness" is arranged below. Thus, in the emotion map 400, various emotions are mapped according to the structure of emotion generation, while easily generated emotions are mapped nearby.

[0179] These emotions are distributed at the 3 o'clock position on the Emotion Chart 400, and usually fluctuate between peace and unease. In the right half of the Emotion Chart 400, because situational awareness is more dominant than internal feelings, it gives a sense of calm.

[0180] The inner side of the emotion diagram 400 represents the mind, and the outer side of the emotion diagram 400 represents actions. Therefore, the further you go to the outer side of the emotion diagram 400, the more the emotion can be seen (manifested in actions).

[0181] Here, human emotions are based on a balance of various factors such as posture and blood sugar levels. When these balances deviate from the ideal, it indicates unhappiness; when they approach the ideal, it indicates pleasure. In robots, cars, and motorcycles, emotions can also be created based on a balance of factors such as posture and remaining battery power. When these balances deviate from the ideal, it indicates unhappiness; when they approach the ideal, it indicates pleasure. Emotion maps can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on Speech Emotion Recognition and Brain Physiological Signal Analysis Systems for Emotion, Tokushima University, Doctoral Dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). In the left half of the emotion map, emotions belonging to the "Reaction" domain, where sensation is dominant, are arranged. Furthermore, in the right half of the emotion map, emotions belonging to the "Situation" domain, where situational cognition is dominant, are arranged.

[0182] The emotion map defines two types of emotions that promote learning. One is a negative emotion located near the middle of "repentance" or "reflection" on the situation side. That is, when the robot experiences negative emotions such as "I never want to feel this way again" or "I never want to be scolded again." The other is a positive emotion located near "desire" on the response side. That is, when the robot experiences positive feelings such as "wanting more" or "wanting to know more."

[0183] The emotion-specific model 59 feeds user input into a pre-learned neural network to obtain emotion values ​​representing each emotion shown in the emotion graph 400, and determines the user's emotion. This neural network is pre-learned based on multiple learning data sets that combine user input with emotion values ​​representing each emotion shown in the emotion graph 400. Furthermore, this neural network is learned to... Figure 10 As shown in sentiment graph 900, sentiment values ​​in nearby configurations are similar to each other. Figure 10 Examples show that multiple emotions such as "peace of mind", "stability", and "reassurance" have similar emotional values.

[0184] In the above embodiments, a specific processing is described by a single computer 22, but the technology disclosed herein is not limited to this, and distributed processing by multiple computers, including computer 22, is also possible.

[0185] In the above embodiments, an example of storing a specific processing program 56 in memory 32 is illustrated, but the technology disclosed herein is not limited thereto. For example, the specific processing program 56 may also be stored in a portable computer-readable non-transitory storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-transitory storage medium is installed into the computer 22 of the data processing device 12. The processor 28 performs specific processing according to the specific processing program 56.

[0186] Alternatively, the specific processing program 56 can be stored in a storage device such as a server connected to the data processing device 12 via a network 54, and the specific processing program 56 can be downloaded and installed into the computer 22 upon request from the data processing device 12.

[0187] Furthermore, it is not necessary to store the entire specific process 56 in a storage device such as a server connected to the data processing device 12 via the network 54, nor is it necessary to store the entire specific process 56 in the memory 32; a portion of the specific process 56 may also be stored.

[0188] As a hardware resource for performing specific processing, various processors can be used. For example, a CPU is a general-purpose processor that functions as a hardware resource for performing specific processing by executing software, i.e., programs. Additionally, a dedicated circuit can be listed as a processor; it is a processor with a circuit structure specifically designed for performing specific processing, such as a FPGA (Field-Programmable Gate Array), PLD (Programmable Logic Device), or ASIC (Application Specific Integrated Circuit). Every processor has built-in or connected memory, and every processor executes specific processing by using that memory.

[0189] The hardware resources for performing a specific process can consist of one of these various processors, or a combination of two or more processors of the same or different types (e.g., a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Furthermore, the hardware resources for performing a specific process can also be a single processor.

[0190] As an example of a single processor, the first type consists of a combination of one or more CPUs and software, which functions as a hardware resource to perform specific processing. The second type uses a processor, such as a System-on-a-chip (SoC), which implements the entire system functionality, including multiple hardware resources for performing specific processing, using a single IC chip. In this case, the specific processing is implemented using one or more of the aforementioned processors that serve as hardware resources.

[0191] Furthermore, as the hardware architecture of these various processors, more specifically, circuits combining semiconductor elements and other circuit components can be used. Moreover, the specific process described above is merely an example. Therefore, it goes without saying that, without departing from the main point, unnecessary steps can be removed, new steps can be added, or the processing order can be changed.

[0192] Furthermore, although the above examples have been described in terms of first to fourth embodiments, some or all of these embodiments can be combined. Additionally, the smart device 14, smart glasses 214, head-mounted terminal 314, and robot 414 are just examples and can be combined separately, or other devices may be used. Furthermore, although the above examples have been described in terms of morphological example 1 and morphological example 2, these can also be combined.

[0193] The foregoing descriptions and illustrations are detailed explanations of the parts covered by this disclosure and are merely one example of this disclosure. For instance, the descriptions of the above-described structure, function, role, and effect are just one example of the structure, function, role, and effect of the parts covered by this disclosure. Therefore, it goes without saying that, without departing from the spirit of this disclosure, unnecessary parts can be deleted, new elements can be added, or replacements can be made to the foregoing descriptions and illustrations. Furthermore, to avoid confusion and facilitate understanding of the parts covered by this disclosure, explanations of technical common sense that does not require special explanation for implementing this disclosure have been omitted from the foregoing descriptions and illustrations.

[0194] All documents, patent applications and technical standards described in this specification are incorporated herein by reference as if they were specifically and individually described as incorporated by reference.

[0195] [Postscript 1]

[0196] A system, characterized in that it comprises:

[0197] The parsing unit is used to parse the content of telephone calls;

[0198] The determination unit is used to determine whether a call is secure based on the content parsed by the parsing unit.

[0199] The ringing unit is used to ring when the determination unit determines that the telephone is a secure telephone.

[0200] [Postscript 2]

[0201] The system as described in Appendix 1 is characterized in that,

[0202] The parsing unit uses speech recognition technology to convert telephone content into text.

[0203] [Postscript 3]

[0204] The system as described in Appendix 1 is characterized in that,

[0205] The determination unit identifies calls from family or friends as secure calls.

[0206] [Postscript 4]

[0207] The system as described in Appendix 1 is characterized in that,

[0208] The determination unit will classify phone calls containing keywords that suggest a possibility of fraud as unsafe phone calls.

[0209] [Postscript 5]

[0210] The system as described in Appendix 1 is characterized in that,

[0211] The ringing unit only rings when it is a security call.

[0212] [Postscript 6]

[0213] The system as described in Appendix 1 is characterized in that,

[0214] When parsing the content of a phone call, the parsing unit infers the user's emotions and adjusts the parsing accuracy based on the inferred emotions.

[0215] [Postscript 7]

[0216] The system as described in Appendix 1 is characterized in that,

[0217] When the parsing unit uses speech recognition technology to convert telephone content into text, it can correspond to different languages ​​or dialects.

[0218] [Postscript 8]

[0219] The system as described in Appendix 1 is characterized in that,

[0220] The parsing unit removes background noise and other sounds to improve parsing accuracy when parsing telephone content.

[0221] [Postscript 9]

[0222] The system as described in Appendix 1 is characterized in that,

[0223] When analyzing the content of a phone call, the analysis unit infers the user's emotions and adjusts the display method of the analysis results according to the inferred user emotions.

[0224] [Postscript 10]

[0225] The system as described in Appendix 1 is characterized in that,

[0226] When the parsing unit uses speech recognition technology to convert telephone content into text, it highlights specific keywords or phrases.

[0227] [Postscript 11]

[0228] The system as described in Appendix 1 is characterized in that,

[0229] The parsing unit refers to past call records to improve parsing accuracy when parsing telephone content.

[0230] [Postscript 12]

[0231] The system as described in Appendix 1 is characterized in that,

[0232] The judgment unit infers the user's emotions and adjusts the security call judgment criteria based on the inferred user emotions.

[0233] [Postscript 13]

[0234] The system as described in Appendix 1 is characterized in that,

[0235] When determining that a call from family or friends is safe, the determination unit refers to call logs and contact lists to improve the accuracy of the determination.

[0236] [Postscript 14]

[0237] The system as described in Appendix 1 is characterized in that,

[0238] When the determination unit determines that a phone call containing keywords that suggest a potential scam is unsafe, it updates the determination criteria by referring to the latest information on scam tactics.

[0239] [Postscript 15]

[0240] The system as described in Appendix 1 is characterized in that,

[0241] The judgment unit infers the user's emotions and adjusts the notification method of the judgment result based on the inferred user emotions.

[0242] [Postscript 16]

[0243] The system as described in Appendix 1 is characterized in that,

[0244] When determining whether a call from family or friends is safe, the determination unit considers call frequency and time period to improve the accuracy of the determination.

[0245] [Postscript 17]

[0246] The system as described in Appendix 1 is characterized in that,

[0247] When determining that a call containing keywords that may indicate fraud is unsafe, the determination unit considers not only the content of the call but also the sender's information.

[0248] [Postscript 18]

[0249] The system as described in Appendix 1 is characterized in that,

[0250] The ringing unit infers the user's emotions and adjusts the type and volume of the incoming call ringtone based on the inferred emotions.

[0251] [Postscript 19]

[0252] The system as described in Appendix 1 is characterized in that,

[0253] The ringing unit adjusts the ringing timing based on the user's schedule and activity status only when ringing for a security call.

[0254] [Postscript 20]

[0255] The system as described in Appendix 1 is characterized in that,

[0256] When the ringing unit emits a ringtone, it provides a customizable incoming call ringtone according to the user's preferences.

[0257] [Postscript 21]

[0258] The system as described in Appendix 1 is characterized in that,

[0259] The ringing unit infers the user's emotions and adjusts the ringing time of the incoming call based on the inferred user emotions.

[0260] [Postscript 22]

[0261] The system as described in Appendix 1 is characterized in that,

[0262] The ringing unit only rings when making a secure call, taking into account the user's device settings and ambient sound to set the optimal volume.

[0263] [Postscript 23]

[0264] The system as described in Appendix 1 is characterized in that,

[0265] When the ringing unit emits a ring, it refers to the user's past response records to adjust the ringing mode.

Claims

1. A system, characterized in that, include: The parsing unit is used to parse the content of telephone calls; The determination unit is used to determine whether a call is secure based on the content parsed by the parsing unit. The ringing unit is used to ring when the determination unit determines that the telephone is a secure telephone.

2. The system as described in claim 1, characterized in that, The parsing unit uses speech recognition technology to convert telephone content into text.

3. The system as described in claim 1, characterized in that, The determination unit identifies calls from family or friends as secure calls.

4. The system as described in claim 1, characterized in that, The determination unit will classify phone calls containing keywords that suggest a possibility of fraud as unsafe phone calls.

5. The system as described in claim 1, characterized in that, The ringing unit only rings when it is a security call.

6. The system as described in claim 1, characterized in that, When parsing the content of a phone call, the parsing unit infers the user's emotions and adjusts the parsing accuracy based on the inferred emotions.

7. The system as described in claim 1, characterized in that, When the parsing unit uses speech recognition technology to convert telephone content into text, it can correspond to different languages ​​or dialects.

8. The system as described in claim 1, characterized in that, The parsing unit removes background noise and other sounds to improve parsing accuracy when parsing telephone content.

9. The system as described in claim 1, characterized in that, When analyzing the content of a phone call, the analysis unit infers the user's emotions and adjusts the display method of the analysis results according to the inferred user emotions.

10. The system as claimed in claim 1, characterized in that, When the parsing unit uses speech recognition technology to convert telephone content into text, it highlights specific keywords or phrases.

Citation Information

Patent Citations

  • Persona chatbot control method and system

    JP2022180282A