system

CN122618992APending Publication Date: 2026-08-21SOFTBANK GROUP CORP
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202610126531.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2025-02-21
Filing Date
2026-01-29
Publication Date
2026-08-21

AI Technical Summary

Technical Problem

[0004]在现有技术中,存在一个课题,即难以及时地向听觉障碍者通知紧急广播

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122618992A_ABST
    Figure CN122618992A_ABST
Patent Text Reader

Abstract

The system according to the present embodiment includes a receiving unit, an identifying unit, and a notifying unit. The receiving unit receives an emergency broadcast. The identifying unit identifies a voice received by the receiving unit in real time and converts the voice into text. The notifying unit notifies the text converted by the identifying unit.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The technology disclosed herein relates to a system. Background Technology

[0002] Patent Document 1 discloses a personalized chatbot control method executed by at least one processor, the method comprising: receiving user speech; adding the user speech to a prompt containing instructions related to a chatbot role; encoding the prompt; and inputting the encoded prompt into a language model to generate chatbot speech in response to the user speech.

[0003] Patent document 1: Japanese Patent Application Publication No. 2022-180282.

[0004] In the existing technology, there is a problem that it is difficult to notify people with hearing impairments of emergency broadcasts in a timely manner. Summary of the Invention

[0005] The system described in this embodiment includes a receiving unit, an identification unit, and a notification unit. The receiving unit receives emergency broadcasts. The identification unit performs real-time identification on the voice received by the receiving unit and converts it into text. The notification unit notifies the user of the text converted by the identification unit. Attached Figure Description

[0006] Figure 1 This is a conceptual diagram illustrating an example of the configuration of a data processing system according to the first embodiment.

[0007] Figure 2 This is a conceptual diagram illustrating an example of the main functions of the data processing apparatus and smart device according to the first embodiment.

[0008] Figure 3 This is a conceptual diagram illustrating an example of the data processing system configuration in the second embodiment.

[0009] Figure 4 This is a conceptual diagram illustrating an example of the main functions of the data processing device and smart glasses according to the second embodiment.

[0010] Figure 5 This is a conceptual diagram illustrating an example of the data processing system configuration in the third embodiment.

[0011] Figure 6 This is a conceptual diagram illustrating an example of the main functions of the data processing device and head-mounted terminal according to the third embodiment.

[0012] Figure 7 This is a conceptual diagram illustrating an example of the data processing system configuration in the fourth embodiment.

[0013] Figure 8This is a conceptual diagram illustrating an example of the functions of the main parts of the data processing device and robot according to the fourth embodiment.

[0014] Figure 9 It represents an emotion graph that maps multiple emotions.

[0015] Figure 10 It represents an emotion graph that maps multiple emotions.

[0016] Explanation of reference numerals in the attached figures Data processing systems 10, 210, 310, and 410 12 Data processing device 14 Smart devices 214 Smart Glasses 314 Head-mounted terminal 414 Robot. Detailed Implementation

[0017] Hereinafter, an example of an implementation of the system involved in this disclosure will be described with reference to the accompanying drawings.

[0018] First, let's explain the terms used in the following description.

[0019] In the following embodiments, the processor (hereinafter referred to as "processor") can be a single computing device or a combination of multiple computing devices. Furthermore, a processor can be a single computing device or a combination of multiple computing devices. Examples of computing devices include CPU (Central Processing Unit), GPU (Graphics Processing Unit), GPGPU (General-Purpose computing on Graphics Processing Units), APU (Accelerated Processing Unit), or TPU (Tensor Processing Unit), etc.

[0020] In the following implementation, the labeled RAM (Random Access Memory) is a memory that temporarily stores information and is used by the processor as working memory.

[0021] In the following embodiments, the labeled memory is one or more non-volatile storage devices used to store various programs and various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), disk (e.g., hard disk) or magnetic tape, etc.

[0022] In the following implementation, the labeled Communication I / F (Interface) is an interface that includes a communication processor and an antenna, etc. The Communication I / F is responsible for communication between multiple computers. Examples of communication standards applicable to the Communication I / F include wireless communication standards such as 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), or Bluetooth (registered trademark).

[0023] In the following implementation, "A and / or B" is synonymous with "at least one of A and B". That is, "A and / or B" means that it can be only A, only B, or a combination of A and B. Furthermore, in this specification, when "and / or" connects more than three items, the same approach as "A and / or B" applies.

[0024] [First Implementation] Figure 1 An example of the configuration of the data processing system 10 according to the first embodiment is shown.

[0025] like Figure 1 As shown, the data processing system 10 includes a data processing device 12 and an intelligent device 14. An example of the data processing device 12 is a server.

[0026] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 includes a processor 28, RAM 30, and memory 32. The processor 28, RAM 30, and memory 32 are connected to a bus 34. Furthermore, the database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. An example of the network 54 includes a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0027] The smart device 14 includes a computer 36, a receiver 38, an output device 40, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, RAM 48, and memory 50. The processor 46, RAM 48, and memory 50 are connected to a bus 52. In addition, the receiver 38, output device 40, and camera 42 are also connected to the bus 52.

[0028] The receiving device 38 includes a touchscreen 38A and a microphone 38B, etc., for receiving user input. The touchscreen 38A receives user input generated by contact with an indicator (e.g., a pen or finger). The microphone 38B receives user input generated by sound by detecting the user's voice. The control unit 46A sends data representing user input received via the touchscreen 38A and microphone 38B to the data processing unit 12. In the data processing unit 12, a specific processing unit 290 (see...) Figure 2 Get the data that represents the user input.

[0029] The output device 40 includes a display 40A and a speaker 40B, etc., and presents data to the user by outputting data in a user-perceptible form (e.g., sound and / or text). The display 40A displays visual information such as text and images according to instructions from the processor 46. The speaker 40B outputs sound according to instructions from the processor 46. The camera 42 is a small digital camera equipped with an optical system such as a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.

[0030] Communication I / F 44 is connected to network 54. Communication I / F 44 and 26 are responsible for sending and receiving various information between processor 46 and processor 28 via network 54.

[0031] Figure 2 An example of the main functions of the data processing device 12 and the smart device 14 is shown.

[0032] like Figure 2 As shown, in the data processing apparatus 12, specific processing is performed by the processor 28. A specific processing program 56 is stored in the memory 32. The specific processing program 56 is an example of a "program" as understood in this disclosure. The processor 28 reads the specific processing program 56 from the memory 32 and executes the read specific processing program 56 on the RAM 30. Specific processing is implemented by the processor 28 operating as a specific processing unit 290 based on the specific processing program 56 executed on the RAM 30.

[0033] The memory 32 stores a data generation model 58 and an emotion-specific model 59. The data generation model 58 and the emotion-specific model 59 are used by the specific processing unit 290. The specific processing unit 290 can use the emotion-specific model 59 to infer the user's emotions and perform specific processing based on the user's emotions. The emotion inference function using the emotion-specific model 59 (emotion-specific function) includes various inferences and predictions related to the user's emotions, but is not limited to this example. Furthermore, emotion inference and prediction also include, for example, emotion analysis (analysis).

[0034] In the smart device 14, specific processing is performed by the processor 46. A specific processing program 60 is stored in the memory 50. The specific processing program 60 is used in conjunction with the data processing system 10. The processor 46 reads the specific processing program 60 from the memory 50 and executes the read specific processing program 60 on the RAM 48. Specific processing is implemented by the processor 46 operating as a control unit 46A based on the specific processing program 60 executed on the RAM 48. Furthermore, the smart device 14 may also have the same data generation model and emotion-specific model as the data generation model 58 and the emotion-specific model 59, and use these models to perform the same processing as the specific processing unit 290.

[0035] Furthermore, other devices besides the data processing device 12 may also have the data generation model 58. For example, a server device (e.g., a generation server) may have the data generation model 58. In this case, the data processing device 12 obtains the processing results (prediction results, etc.) using the data generation model 58 by communicating with the server device that has the data generation model 58. In addition, the data processing device 12 may be a server device or a user-owned terminal device (e.g., a mobile phone, robot, home appliance, etc.). Next, an example of the processing performed by the data processing system 10 of the first embodiment will be described.

[0036] (Example) The system involved in the embodiments of the present invention is a system that can achieve "emergency broadcast" through an application program, and the AI performs real-time speech recognition and notifies the hearing-impaired persons. This system allows users to receive emergency broadcasts through the application program. The AI performs real-time recognition on the received speech and converts it into text. The converted text is notified to the hearing-impaired persons. Through this system, the hearing-impaired persons can immediately grasp the content of the emergency broadcast. For example, when a user receives an emergency broadcast through the application program, they only need to start the application program and click the button for receiving the emergency broadcast. For example, in an emergency such as an earthquake or a fire, the user only needs to start the application program and press the button for receiving the emergency broadcast. Next, the AI performs real-time recognition on the received speech and converts it into text. The AI uses speech recognition technology to analyze the received speech and convert it into text. For example, when receiving an emergency broadcast of "地震が発生しました。避難してください。" (An earthquake has occurred. Please take shelter.), the AI will analyze this speech and convert it into text of "地震が発生しました。避難してください。". The converted text is notified to the hearing-impaired persons. The notification methods include displaying the text on the smartphone screen or notifying through vibration. For example, while the text is displayed on the smartphone screen, the user is notified through vibration. Through this method, the hearing-impaired persons can immediately grasp the content of the emergency broadcast. Through this system, the hearing-impaired persons can grasp the content of the emergency broadcast in real time and can respond quickly. For example, in the event of an earthquake, the hearing-impaired persons can immediately take shelter; in the event of a fire, they can also take shelter quickly. Thereby, the safety of the hearing-impaired persons can be ensured. Therefore, this system can notify the hearing-impaired persons of the content of the emergency broadcast in real time. Specifically, this system is composed of a dedicated application program running on a mobile terminal such as a smartphone or a tablet, and a neural network model for speech recognition (such as a convolutional neural network, a recurrent neural network, or a speech recognition model based on Transformer) deployed in the cloud or on the terminal. When the user starts the application program and clicks the emergency broadcast receiving button, this system obtains speech data in the 16kHz, 16bit, mono PCM format from the microphone of the terminal. The obtained speech data is first subjected to noise suppression (such as spectral subtraction method, Wiener filtering, etc.), volume normalization, and frame segmentation (such as a 25ms window, a 10ms shift, etc.) in the preprocessing unit, and is converted into feature vectors such as Mel-frequency cepstral coefficients (MFCC) or low-power spectra (for example: 13-dimensional MFCC + Δ + ΔΔ = 39 dimensions). These sequences of feature vectors (such as a length of 300 frames × 39 dimensions) are input into the input layer of the speech recognition model. The speech recognition model, for example, has an encoder-decoder structure. The encoder maps the sequential features to a high-dimensional latent space, and the decoder gradually generates a character sequence (such as "地震が発生しました。避難してください。") using the CTC (Connectionist Temporal Classification) or Attention mechanism.The output is obtained as a text sequence encoded in UTF-8, along with a confidence score (e.g., 0.98) and the confidence distribution of each recognized word. For a specific example, when the input speech is "ジシンガハッセイシマシタ。ヒナンシテクダサイ。", the output text is "地震が発生しました。避難してください。" with a confidence of 0.97. Another example, when the input is "カサイガハッセイシマシタ。ヒナンシテクダサイ。", the output text is "火灾が発生しました。避難してください。" with a confidence of 0.95. These text outputs are passed to the notification department, which uses the screen display API of the terminal to display the text centered in large font, and at the same time calls the vibration control API of the terminal to generate continuous vibration for 1.5 seconds. The notification department can also refer to the user's terminal settings, past notification history, and the user's emotional state (such as the stress level inferred from facial images or voices) to optimize the notification method (display color, font size, vibration mode, etc.). The speech recognition process of AI is different from traditional manual dictation tasks. It performs pattern matching, sequence annotation, and probabilistic reasoning in a hundreds-dimensional feature space and is executed on a high-speed parallel computing cluster, so it can balance real-time performance and high precision. In addition, the AI model uses pre-trained weights and adapts to the vocabulary and intonation unique to emergency broadcasts through transfer learning or online learning. In terms of technical effects, compared with traditional manual information transmission, this system can textify the content of emergency broadcasts with high precision within seconds and instantly notify the hearing-impaired, achieving the acceleration of the initial response during disasters and the improvement of safety. In addition, through the diversification of notification methods (screen display, vibration, speech synthesis, etc.) and optimization according to the user's state, it can provide the best information transmission for each user. Specific application fields include emergency broadcasts during natural disasters such as earthquakes, fires, tsunamis, typhoons, venue broadcasts in railway stations, airports, commercial facilities, etc., and evacuation instruction broadcasts in schools, hospitals, etc., and are applicable to all scenarios where it is necessary to transmit information to diverse users including the hearing-impaired and the elderly in real time.

[0037] The system according to this embodiment includes a receiving unit, an identifying unit, and a notifying unit. The receiving unit is used to receive emergency broadcasts. For example, the user receives an emergency broadcast through an application. At this time, the user starts the application and operates a button for receiving the emergency broadcast with one click. For example, in an emergency such as an earthquake or a fire, the user only needs to start the application and press the button for receiving the emergency broadcast. The identifying unit uses AI to perform real-time recognition on the received voice and convert it into text. For example, when receiving an emergency broadcast of "地震が発生しました。避難してください。", the identifying unit analyzes the voice and converts it into text of "地震が発生しました。避難してください。". The notifying unit notifies the converted text to the hearing-impaired person. The notification methods include displaying the text on the smartphone screen or notifying through vibration. For example, while the text is displayed on the smartphone screen, the user is notified through vibration. Thus, the system can notify the content of the emergency broadcast to the hearing-impaired person in real time. A part or all of the above-mentioned processing in the receiving unit can be implemented by AI or can be implemented without using AI. For example, the receiving unit can automate the operation of pressing the button for receiving the emergency broadcast through AI. A part or all of the above-mentioned processing in the identifying unit can be implemented by generative AI or can be implemented without using generative AI. For example, the identifying unit can input the received voice into generative AI, and the generative AI converts it into text. A part or all of the above-mentioned processing in the notifying unit can be implemented by AI or can be implemented without using AI. For example, the notifying unit can notify the converted text in an optimal manner through AI. Thus, the system can notify the content of the emergency broadcast to the hearing-impaired person in real time. Specifically, this system is composed of a dedicated application running on a mobile terminal such as a smartphone or a tablet, and a neural network model for speech recognition deployed in the cloud or in the terminal. When the user starts the application and clicks the emergency broadcast receiving button, this system obtains voice data (such as 16kHz, 16bit, mono PCM) from the microphone input of the terminal. After obtaining the voice data, the receiving unit performs noise suppression (such as spectral subtraction method, Wiener filtering, etc.), volume normalization, and frame segmentation (such as 25ms window, 10ms shift) through a preprocessing unit, and converts it into feature vectors such as Mel Frequency Cepstral Coefficients (MFCC) or low-power spectrum (such as 13-dimensional MFCC + Δ + ΔΔ = 39 dimensions). These sequences of feature vectors (such as length 300 frames × 39 dimensions) are input into the input layer of the speech recognition model in the identifying unit. The identifying unit, for example, adopts a Transformer-based speech recognition model with an encoder-decoder structure. The encoder maps the temporal features to a high-dimensional latent space, and the decoder gradually generates a character sequence (such as "地震が発生しました。避難してください。") using the CTC or Attention mechanism. Input examples of AI include a sequence of 300 frames × 39-dimensional MFCC or a denoised speech waveform tensor.Examples of the output of AI include text sequences encoded in UTF-8 (such as "火灾が発生しました。避難してください。"), confidence scores (such as 0.95), and confidence distributions for each recognized word (such as 0.90 - 0.99 for each word). In subsequent processing, the notification unit displays the output text centered in large font on the screen of the terminal through the screen display API of the terminal, and calls the vibration control API to generate continuous vibration for 1.5 seconds. The notification unit can also optimize the notification method (display color, font size, vibration mode, etc.) by referring to the user's terminal settings, past notification history, and the user's emotional state (such as the degree of stress inferred from facial images or voices). The speech recognition processing of AI is different from traditional manual dictation tasks. It performs pattern matching, sequence annotation, and probabilistic inference in a feature space of hundreds of dimensions and is executed on a high-speed parallel computing cluster, so it can balance real-time performance and high accuracy. In addition, the AI model uses pre-trained weights and adapts to the vocabulary and intonation unique to emergency broadcasts through transfer learning or online learning. In terms of technical effects, compared with traditional manual information transmission, this system can accurately textify the content of emergency broadcasts and immediately notify the hearing-impaired within seconds, achieving the acceleration of the initial response during disasters and the improvement of safety. In addition, through the diversification of notification methods (screen display, vibration, speech synthesis, etc.) and optimization according to the user's state, it can provide optimal information transmission for each user. Specific application fields include emergency broadcasts during natural disasters such as earthquakes, fires, tsunamis, and typhoons, venue broadcasts in railway stations, airports, commercial facilities, etc., and evacuation instruction broadcasts in schools, hospitals, etc., and are applicable to all scenarios where it is necessary to transmit information to diverse users including the hearing-impaired and the elderly in real time.

[0038] The notification unit includes a method of displaying text on the smartphone screen. For example, the content of an emergency broadcast is displayed as text on the smartphone screen. For example, the text "地震が発生しました。避難してください。" (An earthquake has occurred. Please evacuate.) is displayed. The text displayed on the smartphone screen can be adjusted in font size and display position. For example, the font can be enlarged to improve visibility, or the display position can be adjusted to highlight important information. In addition, the scroll function can be used to display longer texts. For example, when the content of an emergency broadcast is long, the entire content can be displayed by scrolling. By displaying text on the smartphone screen, the content of the emergency broadcast can be notified to hearing-impaired persons. Specifically, the notification unit uses the screen display API of the terminal and can display text in various layout modes such as centered or fixed at the top. The notification unit can automatically adjust the font size (such as 24pt to 72pt), font type (such as boldface, Mincho, etc.), and the contrast between the text color and the background color (such as black text on a white background, blue text on a yellow background, etc.) according to the user's terminal settings and visual characteristics (such as color vision abnormality, amblyopia). The notification unit has an automatic scroll function. When the text exceeds the screen size, the user can slide their finger or automatically scroll vertically to display the entire content. In addition, the notification unit can emphasize the text (such as bold, underline, coloring) or add icons (such as earthquake signs, fire signs) according to the importance and category of the emergency broadcast (such as earthquake, fire, tsunami, etc.). As an AI optimization function, the notification unit can learn the user's past browsing history and reactions during notifications (such as the number of screen clicks, scroll speed) and automatically select the most visible and least stressful display method. For example, users who prefer large fonts and high-contrast color schemes will also be given priority to the same display in the future. In terms of technical effects, through screen display optimization, the notification unit can significantly improve the visibility and readability of each user compared with traditional unified text display, which helps to instantly understand emergency information and prevent misunderstandings. Specific application areas include emergency broadcast notifications for hearing-impaired persons and the elderly, venue broadcast assistance in public facilities and transportation, and multilingual displays for foreign language users.

[0039] The notification system includes vibration-based notification methods. Vibration can be delivered in various modes, such as continuous or intermittent vibration. For example, continuous vibration can be used when the emergency broadcast content is important, while intermittent vibration can be used when the content is relatively unimportant. Furthermore, the vibration intensity can be adjusted. For instance, strong vibration can be used to notify users of urgency. Vibration notifications can immediately inform hearing-impaired individuals of emergency broadcast content. Specifically, the notification system utilizes the terminal's vibration control API to finely control the vibration mode (e.g., continuous 1.5 seconds, 0.5-second intervals x 3 times, Morse code style, etc.), vibration intensity (e.g., weak, medium, strong), number of vibrations, and interval. The notification system can automatically select the optimal vibration mode based on the type of emergency broadcast (e.g., earthquake, fire, tsunami, etc.) or importance score (e.g., continuous strong vibration for scores above 0.9, intermittent medium vibration for scores between 0.7 and 0.9). As an AI optimization function, the notification system can learn from the user's past reaction history (e.g., whether the terminal was operated after the vibration notification, and the time required for the notification to be lifted) and the user's emotional state (e.g., avoiding strong vibration when under high stress) to determine the optimal vibration mode and intensity for each user. For example, users who were previously startled by strong vibrations will now be given priority for intermittent vibrations of moderate intensity. The notification unit can also switch notification methods based on the terminal's battery level and hardware limitations (such as small terminals like smartwatches only allowing short-term vibrations). Technically, through vibration control optimization, the notification unit significantly improves the reliability and timeliness of delivering emergency information based on user conditions and terminal characteristics compared to traditional uniform vibration notifications. Specific applications include emergency broadcast notifications for hearing-impaired users or users in noisy environments, non-visual notifications for smartwatches and wearable devices, and silent notifications for public facilities and transportation.

[0040] The recognition unit uses speech recognition technology to analyze the received speech and convert it into text. The speech recognition technology can adopt technologies such as deep learning or HMM (Hidden Markov Model). For example, high-precision speech recognition can be achieved by adopting deep learning, and recognition considering the temporal variation of speech can be achieved by adopting HMM. By adopting speech recognition technology, the speech of the emergency broadcast can be accurately converted into text. Specifically, the recognition unit takes speech data (such as 16kHz, 16bit, mono PCM) as input, performs noise suppression, volume normalization, and frame segmentation in the preprocessing unit, and extracts feature vectors such as Mel Frequency Cepstral Coefficients (MFCC) or low-power spectra (such as 13-dimensional MFCC+Δ+ΔΔ = 39 dimensions). The recognition unit inputs these sequences of feature quantities (such as length 300 frames × 39 dimensions) into a Convolutional Neural Network (CNN), a Recurrent Neural Network (RNN), or a speech recognition model based on Transformer. Examples of the input to the AI include the denoised MFCC sequence or the speech spectrum image. Examples of the output of the AI include a text sequence encoded in UTF-8 (such as "地震が発生しました。避難してください。"), a confidence score (such as 0.98), and the confidence distribution of each recognized word (such as 0.90 - 0.99 for each word). The recognition unit uses CTC (Connectionist Temporal Classification) or Attention mechanism inside the model to optimize the alignment between speech and character sequences and reduce the misrecognition rate. In addition, the recognition unit can adapt to the vocabulary and intonation unique to emergency broadcasts through transfer learning or online learning. In terms of technical effects, the recognition unit introduces speech recognition technology, which achieves real-time performance, high precision, and reduced misrecognition rate compared with traditional manual dictation, greatly improving the reliability and immediacy of emergency information transmission. Specific application areas include emergency broadcasts during disasters, automatic text conversion of venue broadcasts, translation assistance for foreign language broadcasts, etc.

[0041] The recognition unit can identify speech in real time and convert it into text. To achieve real-time speech recognition, for example, it is necessary to convert speech into text within a few seconds. For example, converting the speech into text within a few seconds after receiving an emergency broadcast can achieve real-time notification. By recognizing speech in real time, the content of the emergency broadcast can be immediately converted into text. Specifically, the recognition unit processes the speech data step by step in units of frames (such as a 25ms window and a 10ms shift), extracts the feature vectors of each frame, and inputs them into the model. The recognition unit adopts a speech recognition model that supports streaming (such as a streaming Transformer, a bidirectional RNN, etc.), and gradually generates a text output while the speech input arrives. Examples of the input of the AI include a sequence of continuously arriving speech frames or a column of MFCC vectors updated in real time. Examples of the output of the AI include a partial text sequence (such as "地震が発生し…” → “地震が発生しました。”), and a gradually updated confidence score (such as 0.85 → 0.92 → 0.98). When the output text exceeds a certain confidence threshold (such as 0.95), the recognition unit immediately forwards it to the notification unit, and the notification unit immediately executes terminal screen display or vibration notification. To minimize the internal delay of the model, the recognition unit adopts a lightweight neural network structure and quantization and distillation technologies to achieve high-speed inference within the terminal. In terms of technical effects, the real-time speech recognition of the recognition unit, compared with traditional batch processing-based speech recognition, has achieved the immediate text conversion and notification of emergency broadcasts, making a significant contribution to the initial response and safety guarantee in disasters. Specific application areas include real-time caption generation for emergency broadcasts, immediate text notification for on-site broadcasts, synchronous information transmission for remote users, etc.

[0042] The receiving unit includes a method for estimating user emotions and adjusting the timing of emergency broadcast reception based on the estimated user emotions. For example, the receiving unit may delay receiving emergency broadcasts when the user is stressed, and receive them only when the user is calm; receive emergency broadcasts immediately when the user is relaxed for a rapid response; and prioritize receiving emergency broadcasts and immediately notifying the user of important information when the user is in a hurry. By adjusting the timing of emergency broadcast reception according to user emotions, notifications can be delivered at more appropriate times. Emotion estimation can be achieved through emotion estimation functions such as emotion engines or generative AI. Generative AI includes, but is not limited to, text generation AI (such as LLM) or multimodal generation AI. Specifically, the receiving unit acquires multimodal data such as the user's facial image, speech, text input, and heart rate, and inputs it into a neural network for emotion estimation (such as a multimodal Transformer, a convolutional + RNN hybrid model). Examples of AI inputs include facial images (224×224 pixel RGB), speech waveforms (3 seconds, 16kHz), text (such as "I'm busy now"), and heart rate timing data (60 seconds, 1Hz sampling). Examples of AI outputs include emotion labels (such as "stress," "relaxed," and "tense"), emotion intensity scores (such as stress 0.85 and relaxation 0.10), and presumed confidence levels (such as 0.92). The receiving unit controls the timing of emergency broadcasts based on the emotion presumption results. For example, it delays reception by 5 minutes when stress is high, receives it immediately when relaxed, and prioritizes reception during emergencies. The receiving unit also learns from users' past emotional history and emergency broadcast response history to determine the optimal timing. Technically, by controlling emotion presumption and reception timing, the receiving unit reduces users' psychological burden and maximizes information acceptability and coping abilities compared to traditional uniform information distribution. Specific application areas include medical settings that emphasize stress management, educational settings, disaster relief to prevent panic, and personalized optimized information distribution services.

[0043] The receiving department analyzes users' past emergency broadcast reception history and selects appropriate reception methods. For example, it prioritizes broadcasts based on the types of emergency broadcasts users frequently received in the past; it analyzes the time periods during which users received emergency broadcasts in the past and suggests optimal reception timing; it can also analyze users' past reception history to analyze their reactions to specific emergency broadcasts and select the optimal notification method. By analyzing past reception history, the optimal reception method can be selected. Specifically, the receiving department maintains an emergency broadcast reception history database for each user (containing structured data such as time, broadcast type, receiving terminal, notification method, and user reaction logs), and inputs this historical data as a time-series vector (e.g., 1 year of reception events × 10-dimensional attributes per event) into the AI ​​model. Examples of AI inputs include a reception history vector for the past 30 days (e.g., 30 × 10 dimensions, with each row containing reception time, broadcast type, notification method, user reaction, etc.), reception frequency distribution for specific categories (e.g., earthquake, fire), and user operation logs after past notifications (e.g., screen clicks within 30 seconds of notification, notification cancellation time, etc.). AI inputs this historical data into convolutional neural networks, recurrent neural networks dedicated to time-series processing, or Transformer-based time-series analysis models to extract users' reception preferences and response patterns. Examples of AI output include priority emergency broadcast category labels (e.g., "earthquake" priority 0.95, "fire" priority 0.85), optimal reception times (e.g., weekdays 6 PM to 10 PM), and recommended notification methods (e.g., vibration + screen display, 36pt font size). Based on the AI ​​output, the receiving unit automatically prioritizes and selects notification methods when receiving emergency broadcasts. For example, for users who previously had poor responses to nighttime notifications, optimizations include vibration notifications at night and screen display in the morning. The receiving unit continuously learns from users' reception and response history, updating model parameters online to flexibly respond to changes in users' lifestyles and preferences. Technically, through historical analysis and optimized reception methods, the receiving unit maximizes user information receptivity and responsiveness compared to traditional unified information distribution, significantly improving the efficiency of emergency broadcast delivery and user satisfaction. Specific application areas include emergency broadcast notifications for individual residences and multi-family homes, individual optimization of internal broadcasts for organizations such as enterprises and schools, and personalized information distribution services based on user attributes.

[0044] When receiving emergency broadcasts, the receiving unit filters them based on the user's current location information and situation. For example, when the user is in a specific area, the receiving unit only receives emergency broadcasts related to that area; when the user is moving, it prioritizes receiving emergency broadcasts related to the destination; and when the user is in a specific building, it receives emergency broadcasts related to that building. By filtering based on current location information and situation, highly relevant emergency broadcasts can be received. Specifically, the receiving unit acquires location information (such as vector data including latitude, longitude, altitude, building ID, floor number, etc.) from the terminal's GPS sensor, Wi-Fi / Bluetooth beacon, indoor positioning system, etc., in real time, and compares it with the metadata of the emergency broadcast (such as target area code, building ID, source information, etc.). Examples of AI inputs include the current location's latitude and longitude vector (e.g., 35.6895, 139.6917), building ID "B123", movement speed (e.g., 1.2 m / s), and movement trajectory over the past 30 minutes (e.g., a 30×2-dimensional coordinate sequence). AI combines this location and situation data with the target area information of emergency broadcasts, outputting a geographic relevance score (e.g., 0.98) or a priority reception indicator (e.g., a receptive tag). Examples of AI outputs include "Current location is within the earthquake warning target area: reception priority 1.0", "Destination is within the fire warning target area: reception priority 0.9", "Building ID matches: must receive", etc. Based on the AI ​​output, the receiving unit only distributes highly relevant emergency broadcasts to terminals, filtering out unnecessary broadcasts. Furthermore, the receiving unit can predict user movement (e.g., commuting routes, movement tendencies) using time-series AI models to prepare in advance for receiving emergency broadcasts to their destinations. Technically, by filtering based on location information and situation, the receiving unit, compared to traditional uniform distribution, can quickly and reliably convey only the truly needed emergency information to users, preventing information overload and confusion caused by misdeliveries, compared to traditional uniform distribution. Specific application areas include area-specific emergency broadcasts in cities and large facilities, personalized alerts for users on the move, and disaster information distribution linking indoor and outdoor environments.

[0045] The receiving unit includes a method for estimating user emotions and determining the priority of received emergency broadcasts based on the estimated user emotions. For example, when the user is anxious, the receiving unit prioritizes receiving emergency broadcasts of high importance; when the user is relaxed, it also receives emergency broadcasts of lower importance; and when the user is in a hurry, it only receives the most important emergency broadcasts. By prioritizing emergency broadcasts based on user emotions, important information can be communicated preferentially. Emotion estimation can be achieved through emotion estimation functions such as emotion engines or generative AI. Generative AI includes, but is not limited to, text generation AI (such as LLM) or multimodal generation AI. Specifically, the receiving unit acquires multimodal data such as the user's facial image (e.g., 224×224 pixel RGB), speech waveform (e.g., 3 seconds 16kHz), text input (e.g., "I'm busy now"), and heart rate timing (e.g., 60 seconds 1Hz sampling), and inputs it into a neural network for emotion estimation (e.g., a multimodal Transformer, a convolutional + RNN hybrid model). Examples of AI input include combinations of facial images + speech + text + biosignals, or single-modal data can also be used. AI outputs emotional labels (e.g., "nervous," "relaxed," "urgent"), emotional intensity scores (e.g., nervous 0.85, relaxed 0.10), and presumed confidence levels (e.g., 0.92). Based on the AI ​​output, the receiving unit determines the priority of emergency broadcasts. For example, when nervousness is high, only broadcasts with importance scores above 0.9 are received; when relaxed, broadcasts above 0.5 are received; and when urgent, only the most important broadcasts are received—threshold controls are implemented. The receiving unit also considers users' past emotional history and emergency broadcast response history, optimizing the priority decision-making algorithm through online learning. Technically, by using emotional presumption and priority control, the receiving unit reduces users' psychological burden and maximizes information acceptability and coping abilities compared to traditional uniform information distribution. Specific application areas include medical settings requiring stress management, educational settings, disaster relief to prevent panic, and individualized optimized information distribution services.

[0046] When receiving emergency broadcasts, the receiving unit prioritizes broadcasts with high relevance based on the user's geographic location information. For example, when the user is located in a specific area, the receiving unit prioritizes emergency broadcasts related to that area; when the user is moving, it prioritizes emergency broadcasts related to the destination; and when the user is inside a specific building, it prioritizes emergency broadcasts related to that building. By considering geographic location information, highly relevant emergency broadcasts can be prioritized. Specifically, the receiving unit acquires location information (such as vector data including latitude, longitude, altitude, building ID, floor number, etc.) from the terminal's GPS sensor, Wi-Fi / Bluetooth beacon, indoor positioning system, etc., in real time, and compares it with the metadata of the emergency broadcast (such as target area code, building ID, source information, etc.). Examples of AI inputs include the current location's latitude and longitude vector (e.g., 35.6895, 139.6917), building ID "B123", movement speed (e.g., 1.2 m / s), and movement trajectory over the past 30 minutes (e.g., a 30×2-dimensional coordinate sequence). AI combines this location and situation data with the target area information of emergency broadcasts, outputting a geographic relevance score (e.g., 0.98) or a priority reception indicator (e.g., a receptive tag). Examples of AI outputs include "Current location is within the earthquake warning target area: reception priority 1.0", "Destination is within the fire warning target area: reception priority 0.9", "Building ID matches: must receive", etc. Based on the AI ​​output, the receiving unit only distributes highly relevant emergency broadcasts to terminals, filtering out unnecessary broadcasts. Furthermore, the receiving unit can predict user movement (e.g., commuting routes, movement tendencies) using time-series AI models to prepare in advance for receiving emergency broadcasts to their destinations. Technically, by filtering based on location information and situation, the receiving unit, compared to traditional uniform distribution, can quickly and reliably convey only the truly needed emergency information to users, preventing information overload and confusion caused by misdeliveries, compared to traditional uniform distribution. Specific application areas include area-specific emergency broadcasts in cities and large facilities, personalized alerts for users on the move, and disaster information distribution linking indoor and outdoor environments.

[0047] When receiving emergency broadcasts, the receiving unit analyzes users' social media activity and receives relevant broadcasts. For example, when a user posts information related to a specific region on social media, the receiving unit receives emergency broadcasts related to that region; when a user participates in a specific event, it receives emergency broadcasts related to that event; and when a user uses a specific hashtag, it receives emergency broadcasts related to that hashtag. By analyzing social media activity, relevant emergency broadcasts can be received. Specifically, the receiving unit obtains users' publicly posted social media data (such as posted text, location tags, event participation history, tag lists, etc.) via API and inputs it into a natural language processing AI model (such as a Transformer-based text classification model or a multimodal model). Examples of AI input include the 100 most recent posted texts (such as "#Shibuya#FireworksFestival" or "Event in Shinjuku today"), location information for each post (such as latitude and longitude), and tag lists (such as #Earthquake#Evacuation#Typhoon), etc. AI outputs attention scores (e.g., Shibuya 0.92, Shinjuku 0.85, earthquake 0.95, typhoon 0.80) or relevance tags (e.g., recommending receiving earthquake-related broadcasts) from these released data. Examples of AI outputs include "Users are highly concerned about activities in the Shibuya area: Prioritize receiving emergency broadcasts related to Shibuya" and "#Earthquake hashtag is used frequently: Earthquake alerts must be received." Based on the AI ​​output, the receiving department prioritizes receiving and notifying users of emergency broadcasts related to their interests and behaviors. Furthermore, the receiving department can continuously learn from changes in users' social media activity and flexibly respond to changes in areas of interest by updating model parameters online. Technically, through social media analysis and broadcast selection, the receiving department achieves information delivery that aligns with users' actual behavior and interests, significantly improving the usefulness and acceptability of information compared to traditional uniform distribution, thus improving the effectiveness of information delivery. Specific application areas include targeted alerts for event locations and tourist destinations, SNS-linked disaster information distribution, and personalized alert services based on user attributes.

[0048] The recognition unit includes a method of estimating the user's emotion and adjusting the speech recognition accuracy based on the estimated user emotion. For example, when the user is nervous, the recognition unit increases the speech recognition accuracy to prevent misrecognition; when the user is relaxed, it maintains the normal speech recognition accuracy; when the user is in a hurry, it prioritizes the speech recognition speed and quickly converts it into text. By adjusting the speech recognition accuracy according to the user's emotion, misrecognition can be prevented. Emotion estimation can be achieved through emotion estimation functions such as an emotion engine or generative AI. The generative AI is, for example, text generation AI (such as LLM) or multimodal generation AI, but is not limited thereto. Specifically, the recognition unit obtains multimodal data such as the user's facial image (224×224 pixel RGB), speech waveform (3 seconds at 16 kHz), text input (such as "very nervous now"), and heart rate time series (sampled at 1 Hz for 60 seconds), and inputs them into a neural network for emotion estimation (such as a multimodal Transformer or a convolutional + RNN hybrid model). Input examples for the AI include combinations of facial image + speech + text + biometric signals, or single-modal data (such as only speech or only facial image) can also be utilized. The AI outputs emotion labels (such as "nervous", "relaxed", "in a hurry"), emotion intensity scores (such as nervousness 0.85, relaxation 0.10), and estimated confidence levels (such as 0.92). The recognition unit dynamically adjusts the operating parameters of the speech recognition model according to the AI output. For example, when the tension level is high, it expands the beam search width of the speech recognition model and selects the text with the highest confidence level from multiple candidates; when relaxed, it uses the normal parameter inference; when in a hurry, it simplifies the model decoding layer and prioritizes the inference speed. Input examples for the AI's speech recognition include denoised MFCC sequences (such as 300 frames × 39 dimensions), speech spectrogram images (such as 128×300 pixels), etc. Output examples for the AI's speech recognition include UTF-8 encoded text sequences (such as "An earthquake has occurred. Please evacuate."), confidence scores (such as 0.98), and the confidence distribution of each recognized word (such as 0.90 - 0.99 for each word). In subsequent processing, when the output text exceeds a certain confidence threshold (such as 0.95), the recognition unit immediately forwards it to the notification unit, and the notification unit performs terminal screen display or vibration notification. The recognition unit uses the Attention mechanism or CTC (Connectionist Temporal Classification) inside the model to optimize the alignment of speech and character sequences and reduce the misrecognition rate. In addition, the recognition unit can adapt to the vocabulary and intonation unique to emergency broadcasts through transfer learning or online learning. In terms of technical effects, through emotion estimation and speech recognition accuracy control, the recognition unit achieves optimal recognition accuracy and speed according to the user's mental state and situation compared with traditional unified speech recognition, taking into account both the reduction of misrecognition and real-time performance. Specific application areas include emergency broadcasts during disasters, information transmission accompanied by stress management in medical and educational sites, and personalized speech recognition services according to the user's state.

[0049] When performing speech recognition, the recognition unit adjusts the recognition detail level according to the importance of the emergency broadcast. For example, when the importance of the emergency broadcast is high, the recognition unit performs detailed speech recognition and converts it into accurate text; when the importance of the emergency broadcast is low, the recognition unit performs regular speech recognition; when the importance of the emergency broadcast is medium, the recognition unit performs moderately detailed speech recognition. By adjusting the recognition detail level according to the importance of the emergency broadcast, accurate text can be converted. Specifically, when receiving, the recognition unit obtains the metadata of the emergency broadcast (such as importance score 0.0 - 1.0, category label, source information, etc.), and dynamically changes the inference parameters of the speech recognition model according to the importance score. Input examples of the AI include speech feature vectors (such as a 300-frame × 39-dimensional MFCC sequence), the importance score of the emergency broadcast (such as 0.95), the broadcast category (such as earthquake, fire, typhoon), etc. When the importance is high, the recognition unit expands the beam search width of the speech recognition model, selects the text with the highest confidence from multiple candidates, and strengthens post-processing such as noise suppression, spelling check, and grammar correction. When the importance is medium, standard parameter inference is used, and when the importance is low, the model decoding layer is simplified to prioritize the inference speed. Output examples of the AI include a text sequence encoded in UTF-8 (such as "火灾が発生しました。避難してください。"), a confidence score (such as 0.95), and the confidence distribution of each recognized word (such as 0.90 - 0.99 for each word). In subsequent processing, when the output text exceeds a certain confidence threshold (such as 0.95), the recognition unit immediately forwards it to the notification unit, and the notification unit performs terminal screen display or vibration notification. The recognition unit uses the Attention mechanism or CTC (Connectionist Temporal Classification) inside the model to optimize the alignment of speech and character sequences, reducing the misrecognition rate. In addition, the recognition unit can adapt to the vocabulary and intonation unique to emergency broadcasts through transfer learning or online learning. In terms of technical effects, through importance-based recognition detail level control, the recognition unit achieves optimal resource allocation and improved recognition accuracy compared with traditional unified speech recognition, preventing misrecognition of important information and ensuring real-time performance. Specific application areas include emergency broadcasts during disasters, transmission of important information in medical and educational fields, and personalized speech recognition services according to user status.

[0050] When performing speech recognition, the recognition unit applies different recognition algorithms according to the category of the emergency broadcast. For example, when it comes to an emergency broadcast related to an earthquake, the recognition unit applies an algorithm for recognizing earthquake-related professional terms; when it comes to an emergency broadcast related to a fire, the recognition unit applies an algorithm for recognizing fire-related professional terms; when it comes to an emergency broadcast related to a typhoon, the recognition unit applies an algorithm for recognizing typhoon-related professional terms. By applying recognition algorithms according to the category, professional terms can be accurately recognized. Specifically, the recognition unit obtains the category information of the emergency broadcast (such as tags like earthquake, fire, typhoon, etc.) during reception and automatically selects a speech recognition model or a vocabulary expansion dictionary optimized for each category. Input examples of the AI include speech feature vectors (such as a 300-frame × 39-dimensional MFCC sequence) and category labels (such as "earthquake"). The recognition unit applies a model that strengthens earthquake-related vocabulary (such as "seismic intensity", "aftershock", "tsunami", etc.) in the earthquake category, a model that strengthens fire-related vocabulary (such as "spread of fire", "extinguish fire", "evacuate", etc.) in the fire category, and a model that strengthens typhoon-related vocabulary (such as "storm", "high tide", "path", etc.) in the typhoon category. Output examples of the AI include a text sequence encoded in UTF-8 (such as "An earthquake has occurred. Please evacuate."), a confidence score (such as 0.98), and the confidence distribution of each recognized word (such as 0.90 - 0.99 for each word). In subsequent processing, when the output text exceeds a certain confidence threshold (such as 0.95), the recognition unit immediately forwards it to the notification unit, and the notification unit performs terminal screen display or vibration notification. The recognition unit applies different model parameters or Attention mechanism weights for each category to reduce the misrecognition of professional terms. In addition, the recognition unit can also flexibly respond to changes in the broadcast content through category vocabulary addition or online learning. In terms of technical effects, by applying the category-adaptive recognition algorithm, the recognition unit has achieved an improvement in the recognition accuracy of professional terms and a reduction in misrecognition compared to traditional general speech recognition, greatly enhancing the reliability of emergency information transmission. Specific application areas include emergency broadcasts during disasters, information transmission containing professional terms in the medical and disaster prevention fields, and category-specific speech recognition services in industrial sites.

[0051] The recognition unit includes a method of estimating the user's emotion and adjusting the speech recognition speed based on the estimated user emotion. For example, when the user is tense, the recognition unit reduces the speech recognition speed to prioritize recognition accuracy; when the user is relaxed, it performs speech recognition at a normal speed; when the user is in a hurry, it increases the speech recognition speed to quickly convert to text. By adjusting the speech recognition speed according to the user's emotion, it can be quickly and accurately converted to text. Emotion estimation can be achieved through emotion estimation functions such as an emotion engine or generative AI. Generative AI is, for example, text generation AI (such as LLM) or multi-modal generation AI, but is not limited thereto. Specifically, the recognition unit obtains multi-modal data such as the user's facial image (RGB of 224×224 pixels), speech waveform (16 kHz for 3 seconds), text input (such as "very urgent now"), heart rate time series (sampled at 1 Hz for 60 seconds), etc., and inputs them into a neural network for emotion estimation (such as a multi-modal Transformer or a convolutional + RNN hybrid model). Input examples for the AI include combinations of facial image + speech + text + biometric signals, or single-modal data can also be utilized. The AI outputs an emotion label (such as "tense", "relaxed", "in a hurry"), an emotion intensity score (such as 0.90 for in a hurry), and an estimation confidence level (such as 0.93). The recognition unit controls the inference speed of the speech recognition model according to the AI output. For example, when the tension level is high, the decoding layer of the model is deepened to perform step-by-step confirmation or generate multiple candidates to prioritize recognition accuracy; when relaxed, it is inferred at the standard speed; when in a hurry, the model parameters are simplified to maximize the inference speed. Input examples for the AI's speech recognition include denoised MFCC sequences (such as 300 frames × 39 dimensions), speech spectrum images (such as 128×300 pixels), etc. Output examples for the AI's speech recognition include UTF-8 encoded text sequences (such as "A fire has occurred. Please evacuate."), confidence scores (such as 0.95), and the confidence distribution of each recognized word (such as 0.90 - 0.99 for each word). In subsequent processing, when the output text exceeds a certain confidence threshold (such as 0.95), the recognition unit immediately forwards it to the notification unit, and the notification unit performs terminal screen display or vibration notification. The recognition unit uses the Attention mechanism or CTC (Connectionist Temporal Classification) inside the model to optimize the alignment of speech and character sequences and reduce the misrecognition rate. In addition, the recognition unit can adapt to the vocabulary and intonation unique to emergency broadcasts through transfer learning or online learning. In terms of technical effects, through emotion estimation and recognition speed control, the recognition unit achieves the optimal recognition speed and accuracy according to the user's mental state and situation compared with traditional unified speech recognition, taking into account both real-time performance and reduction of misrecognition. Specific application fields include emergency broadcasts during disasters, information transmission accompanied by stress management in medical and educational sites, personalized speech recognition services according to the user's state, etc.

[0052] When performing speech recognition, the recognition unit determines the recognition priority based on the source of the emergency broadcast. For example, during an emergency broadcast from a government agency, the recognition unit performs speech recognition with the highest priority; during an emergency broadcast from a local government, it performs speech recognition with the second highest priority; and during an emergency broadcast from a private enterprise, it performs speech recognition with the normal priority. By determining the recognition priority based on the source, important information can be recognized quickly. Specifically, the recognition unit obtains the metadata of the emergency broadcast (such as the source ID, source category label, confidence score, etc.) during reception and assigns a priority score to each source. Input examples for the AI include speech feature vectors (such as a 300-frame × 39-dimensional MFCC sequence) and source categories (such as "government agency", "local government", "private enterprise"). The recognition unit immediately performs speech recognition model inference on broadcasts with high priority (such as those from government agencies) and expands the beam search width to maximize recognition accuracy; when the priority is medium, it uses standard parameter inference, and when the priority is low, it uses batch processing or low-resource mode inference. Output examples for the AI include a text sequence encoded in UTF-8 (such as "地震が発生しました。避難してください。"), a confidence score (such as 0.98), and the confidence distribution of each recognized word (such as 0.90 - 0.99 for each word). In subsequent processing, when the output text exceeds a certain confidence threshold (such as 0.95), the recognition unit immediately forwards it to the notification unit, and the notification unit performs terminal screen display or vibration notification. The recognition unit applies different model parameters or Attention mechanism weights to each source to reduce the misrecognition of important information. In addition, the recognition unit can also flexibly respond to changes in broadcast content through source vocabulary addition or online learning. In terms of technical effects, through the recognition priority control based on the source, compared with traditional unified speech recognition, the recognition unit has achieved the immediate recognition of important information and the prevention of misrecognition, greatly improving the reliability and real-time nature of emergency information transmission. Specific application areas include emergency broadcasts during disasters, source-specific information transmission in the medical and disaster prevention fields, and priority control type speech recognition services in industrial sites.

[0053] During speech recognition, the identification unit adjusts the recognition order based on the relevance of emergency broadcasts. For example, it prioritizes speech recognition of emergency broadcasts relevant to the user's current location; it prioritizes highly relevant emergency broadcasts based on the user's past reception history; and it can also prioritize highly relevant emergency broadcasts based on the user's social media activity. By adjusting the recognition order based on relevance, important information can be identified first. Specifically, the identification unit integrates metadata such as emergency broadcast metadata (e.g., target area code, category, source information, etc.), the user's current location (e.g., latitude, longitude, building ID), past reception history (e.g., reception time, category, response log), and social media activity (e.g., posted text, tags, event participation history), and inputs it into an AI model (e.g., a Transformer-based relevance inference model). Examples of AI inputs include the latitude and longitude vector of the current location (e.g., 35.6895, 139.6917), the reception history vector of the past 30 days (e.g., 30×10 dimension), and the 100 most recent posted texts (e.g., "#earthquake#evacuation"). The AI ​​outputs a relevance score (e.g., 0.98, 0.85, 0.60) for each emergency broadcast based on this information. The recognition unit prioritizes speech recognition processing based on the relevance score. Examples of AI outputs include "Current location is within the earthquake warning target area: relevance 0.98", "Past responses to fire broadcasts were good: relevance 0.90", and "#Typhoon tag used frequently: relevance 0.85". In subsequent processing, the recognition unit immediately infers from broadcasts with high relevance scores, while batch processing or delayed processing is used for broadcasts with low scores. The recognition unit updates the relevance inference model parameters through online learning, flexibly responding to changes in user behavior and interests. In terms of technical effectiveness, the recognition unit, through relevance-based recognition sequence control, achieves real-time delivery of information that users truly need and prevents misidentification compared to traditional unified speech recognition, significantly improving information delivery efficiency and user satisfaction. Specific application areas include area-specific emergency broadcasts in cities and large facilities, personalized alarm services, and SNS-linked disaster information distribution.

[0054] The notification department includes methods for inferring user emotions and adjusting notification presentation based on these inferred emotions. For example, when a user is stressed, the notification department provides a concise and highly visual notification; when the user is relaxed, it provides a notification with detailed information; and when the user is in a hurry, it provides a notification that gets to the point. By adjusting the notification presentation according to user emotions, more appropriate notifications can be achieved. Emotion inference can be achieved through emotion inference functions such as emotion engines or generative AI. Generative AI includes, but is not limited to, text generation AI (such as LLM) or multimodal generation AI. Specifically, the notification department acquires multimodal data such as the user's facial image (224×224 pixels RGB), speech waveform (3 seconds 16kHz), text input (such as "I am very nervous now"), and heart rate timing (60 seconds 1Hz sampling) through terminal sensors or application programming interfaces, and inputs it into a neural network for emotion inference (such as a multimodal Transformer or a convolutional + RNN hybrid model). Examples of AI input include combinations of facial images + speech + text + biosignals, or single-modal data (such as speech only, facial images only) can also be used. AI outputs emotional labels (e.g., "nervous," "relaxed," "rushed"), emotional intensity scores (e.g., nervous 0.85, relaxed 0.10), and inferred confidence levels (e.g., 0.92). Based on the AI ​​output, the notification department dynamically determines the notification presentation. For example, when nervousness is high, the notification department uses the terminal screen display API to display large fonts, high contrast, and only the minimum amount of text centered, while generating brief, continuous vibrations via the vibration control API. When relaxed, detailed explanations and supplementary information, relevant icons, and additional information links or FAQ buttons are added at the bottom of the screen. When rushed, only key points are emphasized in bold, and the notification content is summarized into no more than three lines. Compared to traditional unified notification displays, AI-optimized notification presentation achieves information delivery based on the user's psychological state and situation, significantly improving information acceptability, comprehension, and reducing stress. The AI ​​model utilizes pre-trained weights and continuously improves the accuracy of user emotion inference and the notification presentation optimization algorithm through transfer learning or online learning. The notification department also refers to the user's past notification history and reaction logs (e.g., number of screen clicks after notification, notification cancellation time) to enhance the personalization of the presentation. Specific application areas include emergency broadcast notifications during disasters, information delivery for stress management at medical and educational sites, personalized notification services based on user status, and optimized notifications for diverse user attributes on public facilities and transportation. Technically, the notification department, through affective inference and optimized notification performance, significantly improves information acceptability, comprehension, and stress reduction for each user compared to traditional unified notification displays, thereby enhancing the reliability and timeliness of emergency information delivery.

[0055] When issuing a notification, the notification department adjusts the level of detail based on the importance of the emergency broadcast. For example, it provides a detailed notification for high-importance emergency broadcasts, a concise notification for low-importance emergency broadcasts, and a notification with moderate detail for medium-importance emergency broadcasts. By adjusting the level of detail based on the importance of the emergency broadcast, appropriate information can be provided. Specifically, upon receiving the emergency broadcast, the notification department obtains metadata (e.g., importance score 0.0–1.0, category tag, source information, etc.) and dynamically determines the level of detail of the notification content based on the importance score. Examples of AI inputs include the emergency broadcast's importance score (e.g., 0.95), broadcast category (e.g., earthquake, fire, typhoon), and notification history vectors (e.g., the content of the past 30 notifications and user reactions). Based on this information, the AI ​​outputs the optimal level of notification detail (e.g., detailed, standard, simple) and display elements (e.g., body text, supplementary explanations, diagrams, FAQ links, etc.). When a notification is of high importance, the notification department utilizes the screen display API to display detailed text, evacuation procedures, maps, and related links in large font and high contrast, while generating strong, continuous vibration through the vibration control API. When importance is moderate, only a summary of key points and simple supplementary information are displayed, with moderate, intermittent vibration. When importance is low, only the notification title and a one-line summary are displayed in smaller font, with vibration omitted or brief. The AI ​​model learns from users' past notification history and reaction logs, updating the detail selection algorithm online to continuously optimize the optimal notification detail for each user. The technical effect is that, compared to traditional uniform notification displays, this importance-based notification detail control optimizes resource allocation and improves information delivery accuracy, preventing misinterpretation of important information and increasing user satisfaction. Specific application areas include emergency broadcasts during disasters, important information delivery at medical or educational sites, and personalized notification services based on user status.

[0056] When issuing notifications, the notification department applies different notification methods based on the category of the emergency broadcast. For example, in the case of an earthquake-related emergency broadcast, the department provides a notification method emphasizing earthquake-related information; in the case of a fire-related emergency broadcast, it provides a notification method emphasizing fire-related information; and in the case of a typhoon-related emergency broadcast, it provides a notification method emphasizing typhoon-related information. By applying notification methods corresponding to the category, the content of the emergency broadcast can be effectively conveyed. Specifically, upon receiving the broadcast, the notification department obtains the category information of the emergency broadcast (e.g., labels such as earthquake, fire, typhoon, etc.) and automatically selects notification templates and display elements optimized for each category. Examples of AI input include broadcast category labels (e.g., "earthquake"), notification history vectors (e.g., the content of the past 30 notifications categorized by category and user responses), and user attributes (e.g., visual impairment, device type). Based on this information, the AI ​​outputs the optimal notification method for each category (e.g., earthquake signs and evacuation route maps for the earthquake category, fire signs and firefighting steps for the fire category, and path maps and storm warning information for the typhoon category). The notification system automatically switches colors, icons, emphasis (e.g., yellow background for earthquakes, red background for fires, and blue background for typhoons), notification sounds, and vibration modes for different categories. It also automatically generates explanatory texts containing category-specific vocabulary and terminology to improve user comprehension. The AI ​​model learns from users' past notification history and reaction logs for each category, updating the notification method selection algorithm online to continuously optimize the best category notification for each user. The technical effect is that, compared to traditional general notifications, the application of category-adaptive notification methods significantly improves the accuracy of professional information delivery and prevents misinterpretation, greatly enhancing the reliability of emergency information delivery. Specific application areas include emergency broadcasts during disasters, information delivery containing professional terminology in the medical / disaster prevention field, and category-specific notification services in industrial settings.

[0057] The notification unit includes methods for estimating user emotions and adjusting notification timing based on these estimated emotions. For example, the notification unit may delay notification when the user is tense and notify when the user is calm; notify immediately when the user is relaxed; and prioritize notification of the most important information when the user is in a hurry. By adjusting notification timing based on user emotions, notifications can be delivered at more appropriate times. Emotion estimation can be achieved through emotion estimation functions such as emotion engines or generative AI. Generative AI includes, but is not limited to, text generation AI (such as LLM) or multimodal generation AI. Specifically, the notification unit acquires multimodal data such as the user's facial image (224×224 pixels RGB), speech waveform (3 seconds 16kHz), text input (e.g., "I'm very nervous now"), and heart rate timing (60 seconds 1Hz sampling) through terminal sensors or application interfaces, and inputs this data into a neural network for emotion estimation (such as a multimodal Transformer or a convolutional + RNN hybrid model). Examples of AI input include combinations of facial images, speech, text, and physiological signals, but single-modal data can also be used. AI outputs emotional labels (e.g., "nervous," "relaxed," "hurried"), emotional intensity scores (e.g., nervous 0.85, relaxed 0.10), and presumed confidence levels (e.g., 0.92) from this data. Based on the AI ​​output, the notification department dynamically applies a notification timing control algorithm. For example, notifications are delayed by 5 minutes when tension is high, and are executed when user stress decreases; notifications are sent immediately when relaxed; and only the most important information is prioritized when the user is in a hurry, with more detailed information delayed. The notification department also references the user's past notification history and reaction logs (e.g., number of screen clicks after notification, notification cancellation time) to optimize the timing control algorithm through online learning. The technical effect is that, compared to traditional uniform notification timing, the notification department's emotional presumption and notification timing control reduce the user's psychological burden and maximize information acceptance and coping ability. Specific application areas include medical settings that emphasize stress management, educational settings, disaster relief to prevent panic, and individually optimized information distribution services.

[0058] When the notification department sends a notification, it considers the user's device information and selects an appropriate notification method. For example, when the user is using a smartphone, the notification department provides a notification method that displays text on the screen; when the user is using a tablet, it provides a notification method suitable for a large screen; when the user is using a smartwatch, it provides a notification method through vibration. By considering the device information, the optimal notification method can be provided. Specifically, the notification department obtains the device information of the terminal in real time (such as: terminal type, screen size, operating system version, battery level, hardware functions, etc.), and automatically selects a notification template and display elements optimized for each device. Input examples for the AI include terminal type (such as: smartphone, tablet, smartwatch), screen resolution (such as: 1080×1920), battery level (such as: 80%), hardware functions (such as: presence or absence of vibration, ability to output voice), etc. The AI outputs the optimal notification method based on this information (such as: for a smartphone, large font in the center of the screen + vibration, for a tablet, split screen + detailed information, for a smartwatch, short text + only vibration, etc.). The notification department automatically adjusts the notification sound, vibration mode, font size, color, and layout of each device, and realizes the personalization of the notification method by referring to the user's terminal usage situation and past notification history. The technical effect is that the notification method based on device information by the notification department is optimized, and compared with the traditional unified notification display, the reliability and immediacy of information transmission according to the terminal characteristics and user usage situation are greatly improved. Specific application fields include emergency broadcast notifications in a multi-device environment, non-visual notifications for wearable devices, multi-terminal linked information distribution in public facilities or transportation, etc.

[0059] When issuing notifications, the notification department refers to the user's past notification history to select an appropriate notification method. For example, the department may prioritize notification methods the user has previously preferred; it may also recommend the optimal notification method for a specific time period based on the user's past notification history; and further analyze the user's past notification history to select the most effective notification method. By referring to past notification history, the optimal notification method can be selected. Specifically, the notification department maintains a notification history database for each user (e.g., structured data such as notification time, notification method, device type, and user response logs), and inputs this historical data as a time-series vector (e.g., 1 year of notification events × 10-dimensional attributes per event) into an AI model. Examples of AI inputs include notification history vectors for the past 30 days (e.g., 30 × 10 dimensions, with each row containing notification time, notification method, and user response), notification frequency distribution for specific categories (e.g., earthquakes, fires), and user operation logs after past notifications (e.g., screen clicks within 30 seconds of notification, notification cancellation time), etc. The AI ​​then inputs this historical data into convolutional neural networks, recurrent neural networks dedicated to time-series processing, or Transformer-based time-series analysis models to extract the user's notification tendencies and response patterns. AI output examples include priority notification method tags (e.g., screen display + vibration priority 0.95, voice notification priority 0.85), optimal notification timing (e.g., optimal weekday 6 PM to 10 PM), and recommended notification detail (e.g., 36pt font, additional detailed description). Based on the AI ​​output, the notification department automatically selects the optimal notification method, detail, and timing. For example, for users who previously had poor response to nighttime notifications, optimizations include vibration notifications at night and screen display notifications in the morning. The notification department continuously learns from users' notification and response history, flexibly responding to changes in user lifestyles and preferences by updating model parameters online. The technical effect is that, through historical analysis and notification method optimization, compared to traditional uniform notification displays, the notification department maximizes user information acceptance and responsiveness, significantly improving the delivery efficiency and user satisfaction of emergency broadcasts. Specific application areas include emergency broadcast notifications in individual residences or multi-family homes, individual optimization of broadcasts within organizations such as businesses or schools, and personalized information distribution services based on user attributes.

[0060] The system described in this embodiment is not limited to the examples above. For example, various modifications can be made. Specifically, the architecture of the speech recognition model or the sentiment inference model can be changed. For example, in the speech recognition unit, convolutional neural networks, recurrent neural networks, Transformer-based models, self-supervised learning models, or an integration of multiple models can be used. In the sentiment inference unit, any single or combined multimodal AI model from images, speech, text, and physiological signals can be applied. In the notification unit, various notification methods such as screen display, vibration, speech synthesis, LED flashing, and sending signals to IoT devices can be combined. In addition, learning methods such as cloud-based distributed inference, in-terminal edge inference, personalized user learning, transfer learning, online learning, and federated learning can be selected. Regarding data flow, voice data, sentiment data, location information, historical data, etc., can be encrypted during transmission to enhance privacy protection. The output of the AI ​​model can be in various forms such as text sequences, confidence scores, category labels, detailedness recommendation values, and notification timing recommendation values. In terms of subsequent processing, the notification method, detailedness, timing, priority, etc., can be automatically controlled based on the output results, and user reaction logs can be continuously learned to optimize the entire system. The technological benefits are that these variants, compared to traditional fixed information delivery systems, significantly improve flexibility, scalability, and adaptability, enabling advanced emergency broadcast notification services that can respond instantly to diverse user needs and changing circumstances. Specific application areas include emergency broadcasting during disasters, information delivery at medical / educational / industrial sites, IoT-linked smart home alarms, and personalized information distribution services.

[0061] The receiving unit can also acquire the user's physiological information and adjust the timing of emergency broadcast reception. For example, by monitoring the user's heart rate and blood pressure, the reception of emergency broadcasts can be temporarily delayed when abnormalities are detected; when the user is exercising, the emergency broadcast can be received after the exercise ends; when the user is sleeping, the emergency broadcast can be adjusted to be received after waking up. Thus, the timing of emergency broadcast reception can be optimized based on the user's physiological information. Specifically, the receiving unit acquires physiological information such as heart rate (e.g., 1Hz sampling), blood pressure (e.g., once per minute), activity level (e.g., accelerometer value), and sleep state (e.g., estimated by acceleration + heart rate changes) in real time from the terminal or wearable device, and inputs this time-series data into an AI model for physiological information analysis (e.g., an LSTM-based time-series anomaly detection model or a multimodal Transformer). Examples of AI inputs include the heart rate time series of the last 60 seconds (e.g., 60×1 dimension), blood pressure value (e.g., 120 / 80 mmHg), activity vector (e.g., 60-second 3-axis acceleration), and sleep status labels (e.g., "sleeping" or "awake"). The AI ​​outputs anomaly detection labels (e.g., "abnormal heart rate," "exercising," "sleeping"), anomaly scores (e.g., 0.92), and presumed confidence scores (e.g., 0.95) from this data. Based on the AI ​​output, the receiving unit controls the timing of emergency broadcasts. For example, a 10-minute delay is set for receiving an abnormal heart rate, notifications are sent after exercise ends if the user is exercising, and notifications are sent after the user wakes up if they are sleeping. The receiving unit also optimizes the timing control algorithm through online learning, referencing the user's past physiological information history and emergency broadcast response history. The technical effect is that, compared to traditional uniform information distribution, the receiving unit's timing control based on physiological information achieves optimal information delivery according to the user's health status and activity level, reducing psychological and physical burden while improving information acceptability. Specific application areas include emergency broadcast notifications in medical settings or for the elderly, personalized information distribution to users during exercise or sleep, and IoT-linked alarm services accompanying health management.

[0062] The notification department can also adjust the notification method according to the degree of the user's visual impairment. For example, when the visual impairment is severe, voice notification is prioritized; when the visual impairment is mild, visibility is improved by increasing the font size and contrast; when the user has color vision abnormalities, the color combination of the notification content can be adjusted. Thus, appropriate notification methods can be provided based on the degree of visual impairment. Specifically, the notification department obtains the user's impairment information (such as visual impairment level, color vision abnormality type, amblyopia / total blindness, etc.) from the user profile or terminal settings and inputs the notification method selection AI model. Examples of AI input include visual impairment level (e.g., severe, moderate, mild), color vision abnormality type (e.g., red-green, blue-yellow, full color vision), terminal type (e.g., smartphone, tablet), and past notification history (e.g., frequency of voice notification use). Based on this information, the AI ​​outputs the optimal notification method (e.g., severe cases use voice synthesis notifications, mild cases use large fonts with high contrast, and color vision abnormalities use automatically adjusted color combinations). The notification system utilizes a speech synthesis API to automatically read notification content and a screen display API to automatically adjust font size (e.g., 36pt–72pt), contrast ratio (e.g., black text on white background, blue text on yellow background), and color (e.g., color palettes for color vision deficiencies). Furthermore, the system learns from users' past notification response logs (e.g., number of times voice notifications were played, number of times screen zooming was performed) and updates the notification method selection algorithm online. The technical effect is that the notification system optimizes the notification method for visually impaired individuals, significantly improving the reliability and timeliness of information delivery based on impairment characteristics compared to traditional uniform notification displays. Specific application areas include emergency broadcast notifications for visually impaired individuals or the elderly, accessibility enhancements in public facilities or transportation, and personalized information distribution services for individuals with color vision deficiencies.

[0063] The recognition unit can also remove background noise when parsing emergency broadcast audio. For example, when receiving emergency broadcasts in noisy environments, noise reduction technology can be used to make the speech clearer; specific noises such as wind and traffic noise can be filtered to improve speech recognition accuracy; and emergency broadcast audio can be extracted first when multiple sound sources are mixed. This minimizes the impact of background noise and achieves accurate speech recognition. Specifically, the recognition unit takes speech data (e.g., 16kHz, 16bit, mono PCM) as input and applies spectral subtraction, Wiener filtering, and self-supervised denoising neural networks (e.g., Denoising Autoencoder, U-Net model, etc.) to reduce noise components in the preprocessing unit. Examples of AI input include noisy speech waveforms (e.g., 3 seconds 16kHz), spectral images (e.g., 128×300 pixels), and noise labels (e.g., "wind," "traffic," "crowd"). The AI ​​outputs a denoised speech waveform, a noise component estimation mask, and a noise residual score (e.g., 0.05) from this data. The recognition unit extracts features such as MFCC from the denoised speech and inputs them into the speech recognition model (e.g., based on Transformer, RNN, etc.). When multiple sound sources are mixed, sound source separation AI (e.g., blind source separation model) is applied to extract only the emergency broadcast speech. AI output examples include denoised speech waveforms, confidence scores of extracted speech (e.g., 0.98), and sound source separation labels (e.g., "emergency broadcast," "ambient sound"). The technical effect is that by introducing noise reduction and sound source separation functions, the recognition unit can achieve high-precision speech recognition even in complex environments compared to traditional simple filtering, significantly improving the reliability and timeliness of emergency information transmission. Specific application areas include emergency broadcasts in noisy environments during disasters, automatic recognition of indoor broadcasts in transportation vehicles or commercial facilities, and voice information transmission for outdoor activities.

[0064] The notification unit can also infer user emotions and customize notification content based on these inferred emotions. For example, when a user feels uneasy, reassuring information is added; when a user is excited, information promoting calmness is displayed; and when a user is fatigued, concise and easy-to-understand notification content is provided. Thus, appropriate notification content can be provided based on user emotions. Specifically, the notification unit acquires multimodal data such as the user's facial image (224×224 pixels RGB), speech waveform (3 seconds, 16kHz), text input (e.g., "I'm uneasy now"), and heart rate timing (60 seconds, 1Hz sampling), and inputs it into a neural network for emotion inference (e.g., a multimodal Transformer or a convolutional + RNN hybrid model). Examples of AI input include a combination of facial image, speech, text, and physiological signals, but single-modal data can also be used. From this data, the AI ​​outputs emotion labels (e.g., "uneasy," "excited," "fatigued"), emotion intensity scores (e.g., uneasy 0.85, excited 0.10), and inferred confidence scores (e.g., 0.92). The notification department, powered by AI, takes the emergency broadcast text and emotional tags as input and generates AI-generated notification content (e.g., LLM or template generation models). This automatically generates reassuring supplementary text (e.g., "Please act calmly"), calming messages (e.g., "Take a deep breath and follow instructions"), and concise summaries (e.g., "Only display evacuation instructions"). The department also references users' past notification history and reaction logs, using online learning to optimize its notification content customization algorithm. The technical effect is that the department's emotionally adaptive notification content generation significantly improves the reliability, reassurance, and comprehensibility of information delivery based on the user's psychological state, compared to traditional uniform notification text. Specific applications include emergency broadcasts during disasters, information delivery for stress management in medical or educational settings, and personalized notification services based on user status.

[0065] The receiving unit can also monitor the user device's remaining battery level and adjust the emergency broadcast reception method. For example, when the battery level is low, emergency broadcasts are received in low-power mode; when the battery level is insufficient, the notification method is simplified to suppress battery consumption; when the battery level is extremely low, emergency broadcast reception can be temporarily stopped and resumed after charging is complete. Thus, the optimal reception method can be provided based on the device's remaining battery level. Specifically, the receiving unit acquires the terminal's remaining battery level (e.g., in 1% increments), charging status (e.g., charging / not charging), and power consumption history (e.g., average power consumption in mAh over the past hour) in real time, and inputs this information into a battery management AI model (e.g., time-series LSTM or rule-based model). Examples of AI inputs include current battery level (e.g., 15%), charging status (e.g., not charging), power consumption history (e.g., power consumption per minute over 60 minutes), and terminal type (e.g., smartphone, tablet). Based on this information, the AI ​​outputs the optimal reception mode (e.g., low-power mode, normal mode, reception paused) and notification method (e.g., screen display only, vibration omitted, voice notification omitted). The receiving unit simplifies voice recognition and notification processing in a low-power mode when the battery level is below 20%, minimizes notification methods when it's below 10%, and pauses receiving emergency broadcasts when it's below 5%, resuming only after charging is complete. The receiving unit also optimizes its receiving method control algorithm through online learning, referencing the user's past battery history and emergency broadcast response history. The technical effect is that the receiving unit optimizes the receiving method based on battery level, balancing terminal operation continuity and information transmission reliability compared to traditional uniform information distribution, significantly improving user convenience and security. Specific application areas include emergency broadcast notifications during power outages, power-saving information distribution for users who are out or on the move, and battery management alarm services for IoT terminals or wearable devices.

[0066] The recognition unit can also infer user emotions and provide speech recognition feedback based on the inferred emotions. For example, when a user feels uneasy, the speech recognition result is explained in detail to provide reassurance; when a user is excited, feedback that promotes calmness is provided; when a user is tired, concise and easy-to-understand feedback is provided. Thus, appropriate feedback can be provided according to the user's emotions. Specifically, the recognition unit acquires multimodal data such as the user's facial image (224×224 pixels RGB), speech waveform (3 seconds 16kHz), text input (e.g., "I'm uneasy now"), and heart rate timing (60 seconds 1Hz sampling), and inputs it into a neural network for emotion inference (e.g., a multimodal Transformer or a convolutional + RNN hybrid model). Examples of AI input include a combination of facial image + speech + text + physiological signals, or single-modal data. From this data, the AI ​​outputs emotion labels (e.g., "uneasy", "excited", "fatigued"), emotion intensity scores (e.g., uneasy 0.85, excited 0.10), and inferred confidence scores (e.g., 0.92). The recognition department, based on AI output, inputs recognition results and emotion tags into the speech recognition result feedback AI (such as LLM or template generation models) to automatically generate detailed explanations (e.g., "Speech recognition result: An earthquake has occurred. Please evacuate. Confidence 0.98. Low probability of misidentification"), calming information (e.g., "Please calmly follow instructions"), and concise summaries (e.g., "Only emphasizing evacuation instructions"). The recognition department also references users' past feedback history and reaction logs, optimizing the feedback generation algorithm through online learning. The technical effect is that the recognition department's emotion-adaptive feedback generation significantly improves the reliability, reassurance, and comprehensibility of information delivery based on the user's psychological state compared to traditional uniform recognition result display. Specific application areas include emergency broadcasts during disasters, information delivery for stress management in medical or educational settings, and personalized speech recognition services based on user status.

[0067] The notification department can also learn from users' past behavior patterns to suggest optimal notification timing. For example, when a user wakes up at a specific time each morning, an emergency broadcast can be sent at that time; when a user is at a specific location on a specific week, an emergency broadcast related to that location can be prioritized; when a user is busy during a specific time period, notifications can be sent outside of that time period. This allows for optimal notification timing based on user behavior patterns. Specifically, the notification department inputs users' historical behavioral data (such as wake-up / bedtime, mobile history, device usage, calendar events, etc.) as a time-series vector (e.g., 1 year of behavioral events × 10-dimensional attributes per event) into the AI ​​model. Examples of AI inputs include wake-up time vectors for the past 30 days (e.g., 30 × 1 dimension), mobile destination history broken down by week (e.g., 7 × 10 dimensions), device usage logs (e.g., application launch time, screen on / off count), calendar events (e.g., meetings, courses, etc.), etc. AI inputs this historical data into time-series analysis models (such as LSTM and Transformer) to output optimal notification timing (e.g., 7:00 AM on weekdays, 10:00 AM on Saturdays, after meetings, etc.) and notification priority scores (e.g., 0.95). Based on the AI ​​output, the notification department automatically adjusts the timing of emergency broadcasts, optimizing information delivery to adapt to users' lifestyles and behavioral patterns. The notification department also continuously learns from users' past notification response history and changes in lifestyle patterns, maintaining optimal accuracy by updating model parameters online. The technical effect is that the notification department's behavior-pattern learning-based notification timing optimization significantly improves the reliability and convenience of information delivery based on users' lifestyles and situations compared to traditional uniform notification timing. Specific application areas include emergency broadcast notifications in individual residences or multi-family homes, individual optimization of broadcasts within organizations such as businesses or schools, and personalized information distribution services based on user attributes.

[0068] The receiving unit can also infer user emotions and summarize emergency broadcast content based on the inferred user emotions. For example, when the user is nervous, only a concise summary of important information is given; when the user is relaxed, a notification with detailed information is provided; when the user is in a hurry, the most important points are emphasized. Thus, emergency broadcasts can be delivered with an appropriate amount of information based on the user's emotions. Specifically, the receiving unit acquires multimodal data such as the user's facial image (224×224 pixels RGB), speech waveform (3 seconds 16kHz), text input (e.g., "I'm nervous now"), and heart rate timing (60 seconds 1Hz sampling), and inputs it into a neural network for emotion inference (e.g., a multimodal Transformer or a convolutional + RNN hybrid model). Examples of AI input include a combination of facial image + speech + text + physiological signals, or single-modal data can be used. From this data, the AI ​​outputs emotion labels (e.g., "nervous", "relaxed", "in a hurry"), emotion intensity scores (e.g., nervous 0.85, relaxed 0.10), and inferred confidence scores (e.g., 0.92). Based on AI output, the receiving unit inputs the emergency broadcast text and sentiment tags into a summary generation AI (such as an LLM or template generation model) to automatically generate a concise summary containing only the important information (e.g., "An earthquake has occurred"). Evacuation instructions), and additional detailed information (such as: seismic intensity). evacuation routes The receiving unit also includes notes and key points (such as "Evacuation instructions should be displayed with the highest priority"), and optimizes the summary generation algorithm through online learning by referencing users' past notification history and reaction logs. The technical effect is that the receiving unit's emotion-adaptive summary generation significantly improves the reliability, comprehensibility, and stress reduction of information delivery based on the user's psychological state and situation compared to traditional uniform information notifications. Specific application areas include emergency broadcasts during disasters, information delivery for stress management in medical or educational settings, and personalized notification services based on user status.

[0069] The notification department can also monitor user device usage and select the optimal notification method. For example, when a user is using a smartphone, a notification method that displays text on the screen can be provided; when a user is using a smartwatch, a notification method that uses vibration can be provided; and when a user is using a computer, a desktop notification can be displayed. Thus, the optimal notification method can be provided based on device usage. Specifically, the notification department acquires real-time terminal usage information (such as active applications, screen on / off status, connected device type, user operation logs, etc.) and inputs it into a usage information parsing AI model (such as a time-series LSTM or rule-based model). Examples of AI input include currently active devices (such as smartphones, smartwatches, and computers), screen status (such as ON / OFF), active application names (such as messages and browsers), and user operation history (such as the number of clicks / touches in the last 10 minutes). Based on this information, the AI ​​outputs the optimal notification method (such as screen display for smartphones, vibration for smartwatches, and desktop notification for computers). The notification department automatically adjusts the notification sound, vibration mode, font size, color, and layout for each device, and refers to the user's terminal usage and past notification history to personalize the notification method. The technical benefits include optimized notification methods based on device usage, significantly improving the reliability and timeliness of information delivery based on terminal usage and user behavior compared to traditional uniform notification displays. Specific application areas include emergency broadcast notifications in multi-device environments, non-visual notifications for wearable devices, and desktop-based information distribution in offices or homes.

[0070] The recognition unit can also infer user emotions and emphasize the speech recognition results based on the inferred emotions. For example, when a user is anxious, important information is emphasized to provide reassurance; when a user is excited, information that promotes calmness is emphasized; when a user is tired, concise and easy-to-understand information is emphasized. Thus, appropriate information can be provided based on the user's emotions. Specifically, the recognition unit acquires multimodal data such as the user's facial image (224×224 pixels RGB), speech waveform (3 seconds, 16kHz), text input (e.g., "I'm anxious now"), and heart rate timing (60 seconds, 1Hz sampling), and inputs it into a neural network for emotion inference (e.g., a multimodal Transformer or a convolutional + RNN hybrid model). Examples of AI input include a combination of facial image, speech, text, and physiological signals, but single-modal data can also be used. From this data, the AI ​​outputs emotion labels (e.g., "anxious," "excited," "fatigued"), emotion intensity scores (e.g., anxiety 0.85, excitement 0.10), and inferred confidence scores (e.g., 0.92). Based on AI output, the recognition department uses a speech recognition result-based emphasis display control algorithm to highlight important information (such as "evacuation instructions") through bolding, coloring, and adding icons. When calmness is needed, supplementary text is added to encourage calm action; when fatigued, key points are simply displayed in large font. The recognition department also references users' past notification history and reaction logs, optimizing the emphasis display algorithm through online learning. The technical effect is that the recognition department's emotion-adaptive emphasis display significantly improves the reliability, reassurance, and comprehensibility of information delivery based on the user's psychological state and condition compared to traditional uniform information display. Specific application areas include emergency broadcasts during disasters, information delivery for stress management in medical or educational settings, and personalized speech recognition services based on user status.

[0071] The following is a brief description of the processing flow of the implementation method. Specifically, this system consists of a receiving unit, a recognition unit, and a notification unit working together to achieve real-time and high-precision processing of emergency broadcasts from reception to notification on the user terminal. The receiving unit acquires voice data (e.g., 16kHz, 16bit, mono PCM) from the terminal microphone input, performs noise reduction, volume normalization, and frame segmentation in the preprocessing unit, and converts it into feature vectors such as MFCC (e.g., 13-dimensional MFCC + Delta + DeltaDelta = 39-dimensional). The recognition unit inputs these feature sequences (e.g., 300 frames × 39 dimensions) into a speech recognition model (e.g., based on Transformer, RNN, etc.), and uses CTC or Attention mechanisms to gradually generate text sequences (e.g., "Earthquake has occurred. Please evacuate."). The input example for AI is an MFCC sequence or a denoised speech waveform, and the output example for AI is a UTF-8 encoded text sequence, confidence score, and confidence distribution of each recognized word. The notification department displays the output text in a large, centered font via the terminal screen display API and generates 1.5 seconds of continuous vibration via the vibration control API. The notification department optimizes the notification method (display color, font size, vibration mode, etc.) based on the user's terminal settings, past notification history, and emotional state. AI performs speech recognition and notification optimization processing. Unlike traditional manual operation or uniform notification display, it performs pattern matching, sequence labeling, and probabilistic inference in a hundreds-dimensional feature space, executed on a high-speed parallel computing cluster, to achieve optimal information delivery tailored to each user's psychological state and terminal condition. The technical effect is that this system can convert emergency broadcast content into text with high accuracy within seconds and instantly notify hearing-impaired users and various other users, accelerating initial response and improving safety during disasters. Specific application areas include emergency broadcasts during natural disasters such as earthquakes, fires, tsunamis, and typhoons; in-building broadcasts in railways, airports, and commercial facilities; and evacuation instructions broadcast in schools or hospitals.

[0072] Step 1: The receiving unit receives emergency broadcasts. For example, the user receives an emergency broadcast through an application. The user simply launches the application and presses the button to receive the broadcast with a single click. For example, in an emergency such as an earthquake or fire, the user only needs to launch the application and press the button to receive the emergency broadcast. Step 2: The recognition unit uses AI to recognize the received speech in real time and convert it into text. For example, when receiving an emergency broadcast saying "Earthquake has occurred. Please evacuate.", the recognition unit analyzes the speech and converts it into the text "Earthquake has occurred. Please evacuate." Step 3: The notification unit notifies the hearing-impaired user of the converted text. Notification methods include displaying the text on a smartphone screen or notifying via vibration. For example, the text can be displayed on the smartphone screen while simultaneously being vibrated to notify the user. Thus, the system can notify the hearing-impaired user of emergency broadcast content in real time. Specifically, in step 1, the receiving unit acquires speech data (e.g., 16kHz, 16bit, mono PCM) from the terminal microphone input. The preprocessing unit performs noise reduction, volume normalization, and frame segmentation, converting the data into feature vectors such as MFCC (e.g., 13-dimensional MFCC + Delta + DeltaDelta = 39-dimensional). In step 2, the recognition unit inputs these feature sequences (e.g., 300 frames × 39 dimensions) into a speech recognition model (e.g., based on Transformer, RNN, etc.), using CTC or Attention mechanisms to gradually generate a text sequence (e.g., "An earthquake has occurred. Please evacuate."). AI input examples are MFCC sequences or denoised speech waveforms, and AI output examples are UTF-8 encoded text sequences, confidence scores, and the confidence distribution of each recognized word. In step 3, the notification unit displays the output text in a large, centered font via the terminal screen display API and generates 1.5 seconds of continuous vibration via the vibration control API. The notification department optimizes notification methods (display color, font size, vibration mode, etc.) by referencing users' terminal settings, past notification history, and emotional state. AI performs speech recognition and notification optimization, unlike traditional manual operation or uniform notification display. It performs pattern matching, sequence labeling, and probabilistic inference in a hundreds-dimensional feature space, executed on a high-speed parallel computing cluster, to achieve optimal information delivery tailored to each user's psychological state and terminal condition. Technically, this system can convert emergency broadcast content into text with high accuracy within seconds and instantly notify hearing-impaired users and various other users, accelerating initial response and improving safety during disasters. Specific application areas include emergency broadcasts during natural disasters such as earthquakes, fires, tsunamis, and typhoons; in-building broadcasts in railways, airports, and commercial facilities; and evacuation instructions broadcast in schools and hospitals.

[0073] The specific processing unit 290 sends the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires voice representing the user's input to the result of the specific processing. The control unit 46A sends the voice data representing the user's input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the voice data.

[0074] Data generation model 58 is what is known as generative AI (Artificial Intelligence). An example of data generation model 58 includes ChatGPT (registered trademark) (Internet search).<URL: https: / / openai.com / blog / chatgpt> Generative AI, such as data generation model 58, is obtained by deep learning through a neural network. The data generation model 58 is input with a prompt containing instructions, and with inference data such as speech data representing speech, text data representing text, and image data representing images (e.g., data of still images or data of moving images). The data generation model 58 infers the inference data based on the instructions shown in the prompt and outputs the inference result in one or more data forms such as speech data, text data, and image data. The data generation model 58 includes, for example, text generation AI, image generation AI, and multimodal generation AI. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization. The specific processing unit 290 performs the above-described specific processing while using the data generation model 58. The data generation model 58 can also be a fine-tuned model to output inference results from prompts without instructions; in this case, the data generation model 58 can output inference results from prompts without instructions. The data processing device 12, etc., includes multiple data generation models 58, including AI other than generative AI. AI other than generative AI includes, but is not limited to, linear regression, logistic regression, decision trees, random forests, support vector machines (SVM), k-means clustering, convolutional neural networks (CNN), recurrent neural networks (RNN), generative adversarial networks (GAN), or Naive Bayes. Furthermore, AI can also act as an AI agent. Additionally, when the processing described above is performed by AI, this processing may be partially or entirely performed by AI, but is not limited to this example. Moreover, processing performed by AI, including generative AI, can be replaced by rule-based processing, and vice versa.

[0075] Furthermore, the processing performed by the aforementioned data processing system 10 is executed by the specific processing unit 290 of the data processing device 12 or the control unit 46A of the smart device 14, but it can also be executed jointly by the specific processing unit 290 of the data processing device 12 and the control unit 46A of the smart device 14. Additionally, the specific processing unit 290 of the data processing device 12 acquires or collects information required for processing from the smart device 14 or external devices, and the smart device 14 acquires or collects information required for processing from the data processing device 12 or external devices.

[0076] Each of the aforementioned elements, including the receiving unit, the recognition unit, and the notification unit, can be implemented, for example, in at least one of the smart device 14 and the data processing device 12. For instance, the receiving unit may be implemented by the control unit 46A of the smart device 14, allowing the user to receive an emergency broadcast by operating a button in an application. The recognition unit may be implemented, for example, by the specific processing unit 290 of the data processing device 12, which uses AI to recognize the received speech in real time and convert it into text. The notification unit may be implemented, for example, by the control unit 46A of the smart device 14, which displays the converted text on a smartphone screen and notifies the user via vibration. The correspondence between the various units and the device or control unit is not limited to the examples described above and can be varied.

[0077] [Second Implementation] Figure 3 An example of the configuration of the data processing system 210 according to the second embodiment is shown.

[0078] like Figure 3 As shown, the data processing system 210 includes a data processing device 12 and smart glasses 214. One example of the data processing device 12 is a server.

[0079] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 includes a processor 28, RAM 30, and memory 32. The processor 28, RAM 30, and memory 32 are connected to a bus 34. Furthermore, the database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. An example of the network 54 includes a WAN and / or LAN, etc.

[0080] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, RAM 48, and memory 50. The processor 46, RAM 48, and memory 50 are connected to a bus 52. Furthermore, the microphone 238, speaker 240, and camera 42 are also connected to the bus 52.

[0081] Microphone 238 receives user commands by receiving the user's voice. Microphone 238 captures the user's voice, converts the captured sound into speech data, and outputs it to processor 46. Speaker 240 outputs sound according to commands from processor 46.

[0082] Camera 42 is a small digital camera equipped with an optical system such as a lens, aperture and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, used to capture the user's surroundings (e.g., the shooting range defined by an angle of view equivalent to the field of vision of an average healthy person).

[0083] Communication I / F 44 is connected to network 54. Communication I / F 44 and 26 are responsible for the transmission and reception of various information between processor 46 and processor 28 via network 54. The transmission and reception of various information between processor 46 and processor 28 using communication I / F 44 and 26 is performed in a secure state.

[0084] Figure 4 An example of the main functions of the data processing device 12 and the smart glasses 214 is shown. Figure 4 As shown, in the data processing device 12, specific processing is performed by the processor 28. The memory 32 stores a specific processing program 56.

[0085] The processor 28 reads a specific processing program 56 from the memory 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is implemented by the processor 28 operating as a specific processing unit 290 based on the specific processing program 56 executed on the RAM 30.

[0086] The memory 32 stores a data generation model 58 and an emotion-specific model 59. The data generation model 58 and the emotion-specific model 59 are used by the specific processing unit 290. The specific processing unit 290 can use the emotion-specific model 59 to infer the user's emotions and perform specific processing based on the user's emotions. The emotion inference function using the emotion-specific model 59 (emotion-specific function) includes various inferences and predictions related to the user's emotions, but is not limited to this example. Furthermore, emotion inference and prediction also include, for example, emotion analysis (analysis).

[0087] In the smart glasses 214, specific processing is performed by the processor 46. A specific processing program 60 is stored in the memory 50. The processor 46 reads the specific processing program 60 from the memory 50 and executes the read specific processing program 60 on the RAM 48. Specific processing is implemented by the processor 46 operating as a control unit 46A based on the specific processing program 60 executed on the RAM 48. Furthermore, the smart glasses 214 may also have the same data generation model and emotion-specific model as the data generation model 58 and the emotion-specific model 59, and use these models to perform the same processing as the specific processing unit 290.

[0088] Furthermore, other devices besides the data processing device 12 may also have the data generation model 58. For example, a server device may have the data generation model 58. In this case, the data processing device 12 obtains the processing results (prediction results, etc.) using the data generation model 58 by communicating with the server device that has the data generation model 58. In addition, the data processing device 12 may be a server device or a user-owned terminal device (e.g., a mobile phone, robot, home appliance, etc.).

[0089] The specific processing unit 290 sends the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires voice input representing the user's input to the specific processing result. The control unit 46A sends the voice data representing the user's input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the voice data.

[0090] The data generation model 58 is a so-called generative AI. An example of the data generation model 58 includes generative AIs such as ChatGPT. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with a prompt containing instructions, and also with inference data such as speech data representing speech, text data representing text, and image data representing images (e.g., data of still images or data of moving images). The data generation model 58 infers the inference data based on the instructions shown in the prompt and outputs the inference result in one or more data forms such as speech data, text data, and image data. The data generation model 58 includes, for example, text generation AI, image generation AI, and multimodal generation AI. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization. The specific processing unit 290 performs the aforementioned specific processing while using the data generation model 58. The data generation model 58 can also be a fine-tuned model to output inference results from prompts that do not contain instructions; in this case, the data generation model 58 is able to output inference results from prompts that do not contain instructions. The data processing device 12, etc., includes various data generation models 58, which include AI other than generative AI. These AIs include, for example, linear regression, logistic regression, decision trees, random forests, support vector machines (SVM), k-means clustering, convolutional neural networks (CNN), recurrent neural networks (RNN), generative adversarial networks (GAN), or Naive Bayes, and are capable of various processing methods, but are not limited to these examples. Furthermore, AI can also be an AI agent. Additionally, when the processing described above is performed by AI, this processing may be partially or entirely performed by AI, but is not limited to these examples. Moreover, processing performed by AI including generative AI can be replaced by rule-based processing, and vice versa.

[0091] The data processing system 210 of the second embodiment performs the same processing as the data processing system 10 of the first embodiment. The processing performed by the data processing system 210 is executed by the specific processing unit 290 of the data processing device 12 or the control unit 46A of the smart glasses 214, but it can also be executed jointly by the specific processing unit 290 of the data processing device 12 and the control unit 46A of the smart glasses 214. Furthermore, the specific processing unit 290 of the data processing device 12 acquires or collects information required for processing from the smart glasses 214 or external devices, and the smart glasses 214 acquires or collects information required for processing from the data processing device 12 or external devices.

[0092] Each of the aforementioned elements, including the receiving unit, recognition unit, and notification unit, can be implemented, for example, in at least one of the smart glasses 214 and the data processing device 12. For instance, the receiving unit may be implemented by the control unit 46A of the smart glasses 214, allowing the user to receive emergency broadcasts by operating a button in the application. The recognition unit may be implemented, for example, by the specific processing unit 290 of the data processing device 12, which uses AI to recognize the received speech in real time and convert it into text. The notification unit may be implemented, for example, by the control unit 46A of the smart glasses 214, which displays the converted text on the smart glasses' screen and notifies the user via vibration. The correspondence between the various units and the device or control unit is not limited to the above examples and can be varied.

[0093] [Third Implementation] Figure 5 An example of the configuration of the data processing system 310 according to the third embodiment is shown.

[0094] like Figure 5 As shown, the data processing system 310 includes a data processing device 12 and a head-mounted terminal 314. An example of the data processing device 12 is a server.

[0095] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 includes a processor 28, RAM 30, and memory 32. The processor 28, RAM 30, and memory 32 are connected to a bus 34. Furthermore, the database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. An example of the network 54 includes a WAN and / or LAN, etc.

[0096] The head-mounted terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a display 343. The computer 36 includes a processor 46, RAM 48, and memory 50. The processor 46, RAM 48, and memory 50 are connected to a bus 52. Furthermore, the microphone 238, speaker 240, camera 42, and display 343 are also connected to the bus 52.

[0097] Microphone 238 receives user commands by receiving the user's voice. Microphone 238 captures the user's voice, converts the captured sound into speech data, and outputs it to processor 46. Speaker 240 outputs sound according to commands from processor 46.

[0098] Camera 42 is a small digital camera equipped with an optical system such as a lens, aperture and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, used to capture the user's surroundings (e.g., the shooting range defined by an angle of view equivalent to the field of vision of an average healthy person).

[0099] Communication I / F 44 is connected to network 54. Communication I / F 44 and 26 are responsible for the transmission and reception of various information between processor 46 and processor 28 via network 54. The transmission and reception of various information between processor 46 and processor 28 using communication I / F 44 and 26 is performed in a secure state.

[0100] Figure 6 An example of the main functions of the data processing device 12 and the head-mounted terminal 314 is shown. Figure 6 As shown, in the data processing device 12, specific processing is performed by the processor 28. The memory 32 stores a specific processing program 56.

[0101] The processor 28 reads a specific processing program 56 from the memory 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is implemented by the processor 28 operating as a specific processing unit 290 based on the specific processing program 56 executed on the RAM 30.

[0102] The memory 32 stores a data generation model 58 and an emotion-specific model 59. The data generation model 58 and the emotion-specific model 59 are used by the specific processing unit 290. The specific processing unit 290 can use the emotion-specific model 59 to infer the user's emotions and perform specific processing based on the user's emotions. The emotion inference function using the emotion-specific model 59 (emotion-specific function) includes various inferences and predictions related to the user's emotions, but is not limited to this example. Furthermore, emotion inference and prediction also include, for example, emotion analysis (analysis).

[0103] In the head-mounted terminal 314, specific processing is performed by the processor 46. A specific program 60 is stored in the memory 50. The processor 46 reads the specific program 60 from the memory 50 and executes the read specific program 60 on the RAM 48. Specific processing is implemented by the processor 46 operating as a control unit 46A based on the specific program 60 executed on the RAM 48. Furthermore, the head-mounted terminal 314 may also have the same data generation model and emotion-specific model as the data generation model 58 and the emotion-specific model 59, and use these models to perform the same processing as the specific processing unit 290.

[0104] Furthermore, other devices besides the data processing device 12 may also have the data generation model 58. For example, a server device may have the data generation model 58. In this case, the data processing device 12 obtains the processing results (prediction results, etc.) using the data generation model 58 by communicating with the server device that has the data generation model 58. In addition, the data processing device 12 may be a server device or a user-owned terminal device (e.g., a mobile phone, robot, home appliance, etc.).

[0105] The specific processing unit 290 sends the result of the specific processing to the head-mounted terminal 314. In the head-mounted terminal 314, the control unit 46A causes the speaker 240 and display 343 to output the result of the specific processing. The microphone 238 acquires voice representing the user's input to the specific processing result. The control unit 46A sends the voice data representing the user's input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the voice data.

[0106] The data generation model 58 is a so-called generative AI. An example of the data generation model 58 includes generative AIs such as ChatGPT. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with a prompt containing instructions, and also with inference data such as speech data representing speech, text data representing text, and image data representing images (e.g., data of still images or data of moving images). The data generation model 58 infers the inference data based on the instructions shown in the prompt and outputs the inference result in one or more data forms such as speech data, text data, and image data. The data generation model 58 includes, for example, text generation AI, image generation AI, and multimodal generation AI. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization. The specific processing unit 290 performs the aforementioned specific processing while using the data generation model 58. The data generation model 58 can also be a fine-tuned model to output inference results from prompts that do not contain instructions; in this case, the data generation model 58 is able to output inference results from prompts that do not contain instructions. The data processing device 12, etc., includes various data generation models 58, which include AI other than generative AI. These AIs include, for example, linear regression, logistic regression, decision trees, random forests, support vector machines (SVM), k-means clustering, convolutional neural networks (CNN), recurrent neural networks (RNN), generative adversarial networks (GAN), or Naive Bayes, and are capable of various processing methods, but are not limited to these examples. Furthermore, AI can also be an AI agent. Additionally, when the processing described above is performed by AI, this processing may be partially or entirely performed by AI, but is not limited to these examples. Moreover, processing performed by AI including generative AI can be replaced by rule-based processing, and vice versa.

[0107] The data processing system 310 of the third embodiment performs the same processing as the data processing system 10 of the first embodiment. The processing performed by the data processing system 310 is executed by the specific processing unit 290 of the data processing device 12 or the control unit 46A of the head-mounted terminal 314, but it can also be executed jointly by the specific processing unit 290 of the data processing device 12 and the control unit 46A of the head-mounted terminal 314. Furthermore, the specific processing unit 290 of the data processing device 12 acquires or collects information required for processing from the head-mounted terminal 314 or external devices, and the head-mounted terminal 314 acquires or collects information required for processing from the data processing device 12 or external devices.

[0108] Each of the aforementioned elements, including the receiving unit, recognition unit, and notification unit, can be implemented, for example, in at least one of the head-mounted terminal 314 and the data processing device 12. For instance, the receiving unit may be implemented by the control unit 46A of the head-mounted terminal 314, allowing the user to receive emergency broadcasts by operating a button in the application. The recognition unit may be implemented, for example, by the specific processing unit 290 of the data processing device 12, which uses AI to recognize the received speech in real time and convert it into text. The notification unit may be implemented, for example, by the control unit 46A of the head-mounted terminal 314, which displays the converted text on the head-mounted terminal's screen and notifies the user via vibration. The correspondence between the various units and the device or control unit is not limited to the above examples and can be varied.

[0109] [Fourth Implementation] Figure 7 An example of the configuration of the data processing system 410 according to the fourth embodiment is shown.

[0110] like Figure 7 As shown, the data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.

[0111] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 includes a processor 28, RAM 30, and memory 32. The processor 28, RAM 30, and memory 32 are connected to a bus 34. Furthermore, the database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. An example of the network 54 includes a WAN and / or LAN, etc.

[0112] Robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a control object 443. The computer 36 includes a processor 46, RAM 48, and memory 50. The processor 46, RAM 48, and memory 50 are connected to a bus 52. Furthermore, the microphone 238, speaker 240, camera 42, and control object 443 are also connected to the bus 52.

[0113] Microphone 238 receives user commands by receiving the user's voice. Microphone 238 captures the user's voice, converts the captured sound into speech data, and outputs it to processor 46. Speaker 240 outputs sound according to commands from processor 46.

[0114] The camera 42 is a small digital camera equipped with an optical system such as a lens, aperture and shutter, and an imaging element such as a CMOS image sensor or a CCD image sensor, used to photograph the user's surroundings (e.g., the shooting range defined by an angle of view equivalent to the field of vision of an average healthy person).

[0115] Communication I / F 44 is connected to network 54. Communication I / F 44 and 26 are responsible for the transmission and reception of various information between processor 46 and processor 28 via network 54. The transmission and reception of various information between processor 46 and processor 28 using communication I / F 44 and 26 is performed in a secure state.

[0116] The controlled object 443 includes a display device, LEDs for the eyes, and motors for driving the arms, hands, and feet. The posture and movements of the robot 414 are controlled by controlling the motors for the arms, hands, and feet. Some of the robot 414's emotions can be expressed by controlling these motors. In addition, facial expressions of the robot 414 can also be expressed by controlling the illumination state of the LEDs for the robot 414's eyes.

[0117] Figure 8 An example of the main functions of the data processing device 12 and the robot 414 is shown. For example... Figure 8 As shown, in the data processing device 12, specific processing is performed by the processor 28. The memory 32 stores a specific processing program 56.

[0118] The processor 28 reads a specific processing program 56 from the memory 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is implemented by the processor 28 operating as a specific processing unit 290 based on the specific processing program 56 executed on the RAM 30.

[0119] The memory 32 stores a data generation model 58 and an emotion-specific model 59. The data generation model 58 and the emotion-specific model 59 are used by the specific processing unit 290. The specific processing unit 290 can use the emotion-specific model 59 to infer the user's emotions and perform specific processing based on the user's emotions. The emotion inference function using the emotion-specific model 59 (emotion-specific function) includes various inferences and predictions related to the user's emotions, but is not limited to this example. Furthermore, emotion inference and prediction also include, for example, emotion analysis (analysis).

[0120] In robot 414, specific processing is performed by processor 46. A specific program 60 is stored in memory 50. Processor 46 reads the specific program 60 from memory 50 and executes the read specific program 60 on RAM 48. Specific processing is achieved by processor 46 acting as control unit 46A based on the specific program 60 executed on RAM 48. Furthermore, robot 414 may also have the same data generation model and emotion-specific model as data generation model 58 and emotion-specific model 59, and use these models to perform the same processing as specific processing unit 290.

[0121] Furthermore, other devices besides the data processing device 12 may also have the data generation model 58. For example, a server device may have the data generation model 58. In this case, the data processing device 12 obtains the processing results (prediction results, etc.) using the data generation model 58 by communicating with the server device that has the data generation model 58. In addition, the data processing device 12 may be a server device or a user-owned terminal device (e.g., a mobile phone, robot, home appliance, etc.).

[0122] The specific processing unit 290 sends the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the controlled object 443 to output the result of the specific processing. The microphone 238 acquires voice representing the user's input regarding the result of the specific processing. The control unit 46A sends the voice data representing the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the voice data.

[0123] The data generation model 58 is a so-called generative AI. An example of the data generation model 58 includes generative AIs such as ChatGPT. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with a prompt containing instructions, and also with inference data such as speech data representing speech, text data representing text, and image data representing images (e.g., data of still images or data of moving images). The data generation model 58 infers the inference data based on the instructions shown in the prompt and outputs the inference result in one or more data forms such as speech data, text data, and image data. The data generation model 58 includes, for example, text generation AI, image generation AI, and multimodal generation AI. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization. The specific processing unit 290 performs the aforementioned specific processing while using the data generation model 58. The data generation model 58 can also be a fine-tuned model to output inference results from prompts that do not contain instructions; in this case, the data generation model 58 is able to output inference results from prompts that do not contain instructions. The data processing device 12, etc., includes various data generation models 58, which include AI other than generative AI. These AIs include, for example, linear regression, logistic regression, decision trees, random forests, support vector machines (SVM), k-means clustering, convolutional neural networks (CNN), recurrent neural networks (RNN), generative adversarial networks (GAN), or Naive Bayes, and are capable of various processing methods, but are not limited to these examples. Furthermore, AI can also be an AI agent. Additionally, when the processing described above is performed by AI, this processing may be partially or entirely performed by AI, but is not limited to these examples. Moreover, processing performed by AI including generative AI can be replaced by rule-based processing, and vice versa.

[0124] The data processing system 410 of the fourth embodiment performs the same processing as the data processing system 10 of the first embodiment. The processing performed by the data processing system 410 is executed by the specific processing unit 290 of the data processing device 12 or the control unit 46A of the robot 414, but it can also be executed jointly by the specific processing unit 290 of the data processing device 12 and the control unit 46A of the robot 414. Furthermore, the specific processing unit 290 of the data processing device 12 acquires or collects information required for processing from the robot 414 or external devices, and the robot 414 acquires or collects information required for processing from the data processing device 12 or external devices.

[0125] Each of the aforementioned elements, including the receiving unit, recognition unit, and notification unit, can be implemented, for example, in at least one of the robot 414 and the data processing device 12. For instance, the receiving unit may be implemented by the control unit 46A of the robot 414, allowing the user to receive emergency broadcasts by operating a button in an application. The recognition unit may be implemented, for example, by the specific processing unit 290 of the data processing device 12, which uses AI to recognize the received speech in real time and convert it into text. The notification unit may be implemented, for example, by the control unit 46A of the robot 414, which displays the converted text on the robot's display screen and notifies the user via vibration. The correspondence between the various units and the device or control unit is not limited to the examples described above and can be varied.

[0126] Furthermore, the emotion-specific model 59, serving as an emotion engine, can determine the user's emotion based on a specific mapping. Specifically, the emotion-specific model 59 can determine the user's emotion based on an emotion graph that serves as a specific mapping (see...). Figure 9 The system determines the user's emotions. In addition, the emotion-specific model 59 can also determine the robot's emotions in the same way, and the specific processing unit 290 can also perform specific processing using the robot's emotions.

[0127] Figure 9 This is a diagram representing an emotion map 400 that maps various emotions. In the emotion map 400, emotions are arranged radially from the center in concentric circles. The closer to the center of the concentric circles, the more primitive the emotion is. Further out on the concentric circles, emotions are arranged representing states or actions arising from mood. Emotion is a concept that includes both feelings and mental states. To the left of the concentric circles, emotions generated by reactions occurring in the brain are arranged roughly. To the right of the concentric circles, emotions guided by situational judgments are arranged roughly. Above and below the concentric circles, emotions generated by reactions occurring in the brain and guided by situational judgments are arranged roughly. Furthermore, the emotion of "pleasure" is arranged above the concentric circles, and the emotion of "unpleasantness" is arranged below. Thus, in the emotion map 400, various emotions are mapped according to the structure of emotion generation, while easily generated emotions are mapped nearby.

[0128] These emotions are distributed at the 3 o'clock position on the Emotion Chart 400, and usually fluctuate between peace and unease. In the right half of the Emotion Chart 400, because situational awareness is more dominant than internal feelings, it gives a sense of calm.

[0129] The inner side of the emotion diagram 400 represents the mind, and the outer side of the emotion diagram 400 represents actions. Therefore, the further you go to the outer side of the emotion diagram 400, the more the emotion can be seen (manifested in actions).

[0130] Here, human emotions are based on a balance of various factors such as posture and blood sugar levels. When these balances deviate from the ideal, it indicates unhappiness; when they approach the ideal, it indicates pleasure. In robots, cars, and motorcycles, emotions can also be created based on a balance of factors such as posture and remaining battery power. When these balances deviate from the ideal, it indicates unhappiness; when they approach the ideal, it indicates pleasure. Emotion maps can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on Speech Emotion Recognition and Brain Physiological Signal Analysis Systems for Emotion, Tokushima University, Doctoral Dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). In the left half of the emotion map, emotions belonging to the "Reaction" domain, where sensation is dominant, are arranged. Furthermore, in the right half of the emotion map, emotions belonging to the "Situation" domain, where situational cognition is dominant, are arranged.

[0131] The emotion map defines two types of emotions that promote learning. One is a negative emotion located near the middle of "repentance" or "reflection" on the situation side. That is, when the robot experiences negative emotions such as "I never want to feel this way again" or "I never want to be scolded again." The other is a positive emotion located near "desire" on the response side. That is, when the robot experiences positive feelings such as "wanting more" or "wanting to know more."

[0132] The emotion-specific model 59 feeds user input into a pre-learned neural network to obtain emotion values ​​representing each emotion shown in the emotion graph 400, and determines the user's emotion. This neural network is pre-learned based on multiple learning data sets that combine user input with emotion values ​​representing each emotion shown in the emotion graph 400. Furthermore, this neural network is learned to... Figure 10 As shown in sentiment graph 900, sentiment values ​​in nearby configurations are similar to each other. Figure 10 Examples show how emotions such as "peace of mind," "stability," and "reassurance" can result in similar emotional values.

[0133] In the above embodiments, a specific processing is described by a single computer 22, but the technology disclosed herein is not limited to this, and distributed processing by multiple computers, including computer 22, is also possible.

[0134] In the above embodiments, an example of storing a specific processing program 56 in memory 32 is illustrated, but the technology disclosed herein is not limited thereto. For example, the specific processing program 56 may also be stored in a portable computer-readable non-transitory storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-transitory storage medium is installed in the computer 22 of the data processing device 12. The processor 28 performs specific processing according to the specific processing program 56.

[0135] Alternatively, the specific processing program 56 can be stored in a storage device such as a server connected to the data processing device 12 via a network 54, and the specific processing program 56 can be downloaded and installed into the computer 22 upon request from the data processing device 12.

[0136] Furthermore, it is not necessary to store the entire specific process 56 in a storage device such as a server connected to the data processing device 12 via the network 54, nor is it necessary to store the entire specific process 56 in the memory 32; a portion of the specific process 56 may also be stored.

[0137] As a hardware resource for performing specific processing, various processors can be used. For example, a CPU is a general-purpose processor that functions as a hardware resource for performing specific processing by executing software, i.e., programs. Additionally, a dedicated circuit can be listed as a processor; it is a processor with a circuit structure specifically designed for performing specific processing, such as a FPGA (Field-Programmable Gate Array), PLD (Programmable Logic Device), or ASIC (Application Specific Integrated Circuit). Every processor has built-in or connected memory, and every processor executes specific processing by using that memory.

[0138] The hardware resources for performing a specific process can consist of one of these various processors, or a combination of two or more processors of the same or different types (e.g., a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Furthermore, the hardware resources for performing a specific process can also be a single processor.

[0139] As an example of a single processor, the first type consists of a combination of one or more CPUs and software, which functions as a hardware resource to perform specific processing. The second type uses a processor, such as a System-on-a-chip (SoC), which implements the entire system functionality, including multiple hardware resources performing specific processing, using a single IC chip. In this case, the specific processing is implemented using one or more of the aforementioned processors that serve as hardware resources.

[0140] Furthermore, as the hardware architecture of these various processors, more specifically, circuits combining semiconductor elements and other circuit components can be used. Moreover, the specific process described above is merely an example. Therefore, it goes without saying that, without departing from the main point, unnecessary steps can be removed, new steps can be added, or the processing order can be changed.

[0141] Furthermore, although the above examples have been described in terms of first to fourth embodiments, some or all of these embodiments can also be combined. Additionally, the smart device 14, smart glasses 214, head-mounted terminal 314, and robot 414 are just examples; they can be combined separately or are other devices.

[0142] The foregoing descriptions and illustrations are detailed explanations of the parts covered by this disclosure and are merely one example of this disclosure. For instance, the descriptions of the above-described structure, function, role, and effect are just one example of the structure, function, role, and effect of the parts covered by this disclosure. Therefore, it goes without saying that, without departing from the spirit of this disclosure, unnecessary parts can be deleted, new elements can be added, or replacements can be made to the foregoing descriptions and illustrations. Furthermore, to avoid confusion and facilitate understanding of the parts covered by this disclosure, explanations of technical common sense that does not require special explanation for implementing this disclosure have been omitted from the foregoing descriptions and illustrations.

[0143] All documents, patent applications and technical standards described in this specification are incorporated herein by reference as if they were specifically and individually described as incorporated by reference.

[0144] (Note 1) A system comprising: The receiving unit is used to receive emergency broadcasts; The recognition unit is used to recognize the speech received by the receiving unit in real time and convert it into text; and The notification unit is used to notify the text converted by the recognition unit.

[0145] (Note 2) The system as described in Appendix 1 is characterized in that, The notification unit includes a method for displaying text on a smartphone screen.

[0146] (Note 3) The system as described in Appendix 1 is characterized in that, The notification unit includes a method for sending notifications via vibration.

[0147] (Note 4) The system as described in Appendix 1 is characterized in that, The recognition unit uses speech recognition technology to parse the received speech and convert it into text.

[0148] (Note 5) The system as described in Appendix 1 is characterized in that, The recognition unit can recognize speech and convert it into text in real time.

[0149] (Note 6) The system as described in Appendix 1 is characterized in that, The receiving unit includes a method for estimating user emotions and adjusting the timing of receiving emergency broadcasts based on the estimated user emotions.

[0150] (Note 7) The system as described in Appendix 1 is characterized in that, The receiving unit analyzes the user's past emergency broadcast reception history and selects an appropriate reception method.

[0151] (Note 8) The system as described in Appendix 1 is characterized in that, When receiving an emergency broadcast, the receiving unit filters the information based on the user's current location and situation.

[0152] (Note 9) The system as described in Appendix 1 is characterized in that, The receiving unit includes a method for estimating user sentiment and determining the priority of received emergency broadcasts based on the estimated user sentiment.

[0153] (Postscript 10) The system as described in Appendix 1 is characterized in that, When receiving emergency broadcasts, the receiving unit prioritizes receiving broadcasts with high relevance based on the user's geographical location information.

[0154] (Postscript 11) The system as described in Appendix 1 is characterized in that, When receiving an emergency broadcast, the receiving unit analyzes the user's social media activity and receives the relevant broadcast.

[0155] (Postscript 12) The system as described in Appendix 1 is characterized in that, The recognition unit includes a method for estimating user emotions and adjusting speech recognition accuracy based on the estimated user emotions.

[0156] (Postscript 13) The system as described in Appendix 1 is characterized in that, During speech recognition, the recognition unit adjusts the level of detail based on the importance of the emergency broadcast.

[0157] (Postscript 14) The system as described in Appendix 1 is characterized in that, The recognition unit applies different recognition algorithms based on the category of the emergency broadcast during speech recognition.

[0158] (Postscript 15) The system as described in Appendix 1 is characterized in that, The recognition unit includes a method for estimating user emotions and adjusting speech recognition speed based on the estimated user emotions.

[0159] (Postscript 16) The system as described in Appendix 1 is characterized in that, During speech recognition, the recognition unit determines the recognition priority based on the source of the emergency broadcast.

[0160] (Postscript 17) The system as described in Appendix 1 is characterized in that, During speech recognition, the recognition unit adjusts the recognition order based on the relevance of the emergency broadcast.

[0161] (Postscript 18) The system as described in Appendix 1 is characterized in that, The notification unit includes a method for presuming user sentiment and adjusting the notification presentation based on the presumed user sentiment.

[0162] (Postscript 19) The system as described in Appendix 1 is characterized in that, When issuing a notification, the notification department adjusts the level of detail in the notification based on the importance of the emergency broadcast.

[0163] (Postscript 20) The system as described in Appendix 1 is characterized in that, When issuing a notification, the notification department applies different notification methods depending on the type of emergency broadcast.

[0164] (Postscript 21) The system as described in Appendix 1 is characterized in that, The notification unit includes a method for presuming user sentiment and adjusting the timing of notifications based on the presumed user sentiment.

[0165] (Postscript 22) The system as described in Appendix 1 is characterized in that, When issuing a notification, the notification department takes into account the user's device information and selects an appropriate notification method.

[0166] (Postscript 23) The system as described in Appendix 1 is characterized in that, When issuing a notification, the notification department refers to the user's past notification history and selects an appropriate notification method.

Claims

1. A system, characterized in that, include: The receiving unit is used to receive emergency broadcasts; The recognition unit uses the voice received by the receiving unit to perform real-time recognition and convert it into text; as well as The notification unit is used to notify the text converted by the recognition unit.

2. The system as described in claim 1, characterized in that, The notification unit includes a method for displaying text on a smartphone screen.

3. The system as described in claim 1, characterized in that, The notification unit includes a method for sending notifications via vibration.

4. The system as described in claim 1, characterized in that, The recognition unit uses speech recognition technology to parse the received speech and convert it into text.

5. The system as described in claim 1, characterized in that, The recognition unit can recognize speech and convert it into text in real time.

6. The system as described in claim 1, characterized in that, The receiving unit includes a method for estimating user emotions and adjusting the timing of receiving emergency broadcasts based on the estimated user emotions.

7. The system as described in claim 1, characterized in that, The receiving unit analyzes the user's past emergency broadcast reception history and selects an appropriate reception method.

8. The system as described in claim 1, characterized in that, When receiving an emergency broadcast, the receiving unit filters the data based on the user's current location and situation.

9. The system as described in claim 1, characterized in that, The receiving unit includes a method for estimating user sentiment and determining the priority of received emergency broadcasts based on the estimated user sentiment.

10. The system as claimed in claim 1, characterized in that, When receiving emergency broadcasts, the receiving unit prioritizes receiving broadcasts with high relevance based on the user's geographical location information.

Citation Information

Patent Citations

  • Persona chatbot control method and system

    JP2022180282A