system

CN122621584APending Publication Date: 2026-08-21SOFTBANK GROUP CORP
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202610195293.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2025-02-21
Filing Date
2026-02-11
Publication Date
2026-08-21

AI Technical Summary

Technical Problem

[0004]在现有技术中,存在这样的问题:当发生认知障碍时,难以准确把握本人的意向

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122621584A_ABST
    Figure CN122621584A_ABST
Patent Text Reader

Abstract

The system according to the present embodiment includes a collection unit, an analysis unit, and an estimation unit. The collection unit collects the utterance of the person in real time. The analysis unit analyzes the utterance data collected by the collection unit. The estimation unit estimates the intention of the person on the basis of the information analyzed by the analysis unit.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The technology disclosed herein relates to a system. Background Technology

[0002] Patent Document 1 discloses a personalized chatbot control method executed by at least one processor, the method comprising: receiving user speech; adding the user speech to a prompt containing instructions related to a chatbot role; encoding the prompt; and inputting the encoded prompt into a language model to generate chatbot speech in response to the user speech.

[0003] Patent document 1: Japanese Patent Application Publication No. 2022-180282.

[0004] The existing technology has the following problem: when cognitive impairment occurs, it is difficult to accurately grasp the person's intentions. Summary of the Invention

[0005] The system described in this embodiment includes a collection unit, a parsing unit, and an estimation unit. The collection unit collects the speaker's statements in real time. The parsing unit parses the statement data collected by the collection unit. The estimation unit estimates the speaker's intentions based on the information parsed by the parsing unit. Attached Figure Description

[0006] Figure 1 This is a conceptual diagram illustrating an example of the configuration of a data processing system according to the first embodiment.

[0007] Figure 2 This is a conceptual diagram illustrating an example of the main functions of the data processing apparatus and smart device according to the first embodiment.

[0008] Figure 3 This is a conceptual diagram illustrating an example of the data processing system configuration in the second embodiment.

[0009] Figure 4 This is a conceptual diagram illustrating an example of the main functions of the data processing device and smart glasses according to the second embodiment.

[0010] Figure 5 This is a conceptual diagram illustrating an example of the data processing system configuration in the third embodiment.

[0011] Figure 6 This is a conceptual diagram illustrating an example of the main functions of the data processing device and head-mounted terminal according to the third embodiment.

[0012] Figure 7 This is a conceptual diagram illustrating an example of the data processing system configuration in the fourth embodiment.

[0013] Figure 8This is a conceptual diagram illustrating an example of the functions of the main parts of the data processing device and robot according to the fourth embodiment.

[0014] Figure 9 It represents an emotion graph that maps multiple emotions.

[0015] Figure 10 It represents an emotion graph that maps multiple emotions.

[0016] Explanation of reference numerals in the attached figures Data processing systems 10, 210, 310, and 410 12 Data processing device 14 Smart devices 214 Smart Glasses 314 Head-mounted terminal 414 Robot. Detailed Implementation

[0017] Hereinafter, an example of an implementation of the system involved in this disclosure will be described with reference to the accompanying drawings.

[0018] First, let's explain the terms used in the following description.

[0019] In the following embodiments, the processor (hereinafter referred to as "processor") can be a single computing device or a combination of multiple computing devices. Furthermore, a processor can be a single computing device or a combination of multiple computing devices. Examples of computing devices include CPU (Central Processing Unit), GPU (Graphics Processing Unit), GPGPU (General-Purpose computing on Graphics Processing Units), APU (Accelerated Processing Unit), or TPU (Tensor Processing Unit), etc.

[0020] In the following implementation, the labeled RAM (Random Access Memory) is a memory that temporarily stores information and is used by the processor as working memory.

[0021] In the following embodiments, the labeled memory is one or more non-volatile storage devices used to store various programs and various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), disk (e.g., hard disk) or magnetic tape, etc.

[0022] In the following implementation, the labeled Communication I / F (Interface) is an interface that includes a communication processor and an antenna, etc. The Communication I / F is responsible for communication between multiple computers. Examples of communication standards applicable to the Communication I / F include wireless communication standards such as 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), or Bluetooth (registered trademark).

[0023] In the following implementation, "A and / or B" is synonymous with "at least one of A and B". That is, "A and / or B" means that it can be only A, only B, or a combination of A and B. Furthermore, in this specification, when "and / or" connects more than three items, the same approach as "A and / or B" applies.

[0024] [First Implementation] Figure 1 An example of the configuration of the data processing system 10 according to the first embodiment is shown.

[0025] like Figure 1 As shown, the data processing system 10 includes a data processing device 12 and an intelligent device 14. An example of the data processing device 12 is a server.

[0026] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 includes a processor 28, RAM 30, and memory 32. The processor 28, RAM 30, and memory 32 are connected to a bus 34. Furthermore, the database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. An example of the network 54 includes a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0027] The smart device 14 includes a computer 36, a receiver 38, an output device 40, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, RAM 48, and memory 50. The processor 46, RAM 48, and memory 50 are connected to a bus 52. In addition, the receiver 38, output device 40, and camera 42 are also connected to the bus 52.

[0028] The receiving device 38 includes a touchscreen 38A and a microphone 38B, etc., for receiving user input. The touchscreen 38A receives user input generated by contact with an indicator (e.g., a pen or finger). The microphone 38B receives user input generated by sound by detecting the user's voice. The control unit 46A sends data representing user input received via the touchscreen 38A and microphone 38B to the data processing unit 12. In the data processing unit 12, a specific processing unit 290 (see...) Figure 2 Get the data that represents user input.

[0029] The output device 40 includes a display 40A and a speaker 40B, etc., and presents data to the user by outputting data in a user-perceptible form (e.g., sound and / or text). The display 40A displays visual information such as text and images according to instructions from the processor 46. The speaker 40B outputs sound according to instructions from the processor 46. The camera 42 is a small digital camera equipped with an optical system such as a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.

[0030] Communication I / F 44 is connected to network 54. Communication I / F 44 and 26 are responsible for sending and receiving various information between processor 46 and processor 28 via network 54.

[0031] Figure 2 An example of the main functions of the data processing device 12 and the smart device 14 is shown.

[0032] like Figure 2 As shown, in the data processing apparatus 12, specific processing is performed by the processor 28. A specific processing program 56 is stored in the memory 32. The specific processing program 56 is an example of a "program" as understood in this disclosure. The processor 28 reads the specific processing program 56 from the memory 32 and executes the read specific processing program 56 on the RAM 30. Specific processing is implemented by the processor 28 operating as a specific processing unit 290 based on the specific processing program 56 executed on the RAM 30.

[0033] The memory 32 stores a data generation model 58 and an emotion-specific model 59. The data generation model 58 and the emotion-specific model 59 are used by the specific processing unit 290. The specific processing unit 290 can use the emotion-specific model 59 to infer the user's emotions and perform specific processing based on the user's emotions. The emotion inference function using the emotion-specific model 59 (emotion-specific function) includes various inferences and predictions related to the user's emotions, but is not limited to this example. Furthermore, emotion inference and prediction also include, for example, emotion analysis (analysis).

[0034] In the smart device 14, specific processing is performed by the processor 46. A specific processing program 60 is stored in the memory 50. The specific processing program 60 is used in conjunction with the data processing system 10. The processor 46 reads the specific processing program 60 from the memory 50 and executes the read specific processing program 60 on the RAM 48. Specific processing is implemented by the processor 46 operating as a control unit 46A based on the specific processing program 60 executed on the RAM 48. Furthermore, the smart device 14 may also have the same data generation model and emotion-specific model as the data generation model 58 and the emotion-specific model 59, and use these models to perform the same processing as the specific processing unit 290.

[0035] Furthermore, other devices besides the data processing device 12 may also have the data generation model 58. For example, a server device (e.g., a generation server) may have the data generation model 58. In this case, the data processing device 12 obtains the processing results (prediction results, etc.) using the data generation model 58 by communicating with the server device that has the data generation model 58. In addition, the data processing device 12 may be a server device or a user-owned terminal device (e.g., a mobile phone, robot, home appliance, etc.). Next, an example of the processing performed by the data processing system 10 of the first embodiment will be described.

[0036] (Example) The system described in this invention is a wearable device for understanding an individual's intentions when cognitive impairment occurs. This wearable device can collect the individual's speech in real time, and a generating AI analyzes the collected speech data to extract the language's tendency and meaning. Furthermore, based on the extracted information, the individual's intentions are inferred. The device also implements a life log recording and anti-wandering functions, comprehensively supporting the individual's life. For example, the wearable device can collect the individual's speech in real time. Subsequently, the generating AI analyzes the collected speech data to extract the language's tendency and meaning. Furthermore, based on the extracted information, the individual's intentions are inferred. The device also implements a life log recording and anti-wandering functions, comprehensively supporting the individual's life. Therefore, even when cognitive impairment occurs, the individual's intentions can be understood through the wearable device. Specifically, this wearable device is equipped with a microphone for voice input, an accelerometer, a GPS module, a vital signs sensor, and other sensor groups, enabling real-time acquisition of the individual's speech at 16kHz sampled PCM data. This system first performs spectrogram transformation on the speech data and then performs noise suppression processing before inputting it into a speech recognition engine (e.g., an end-to-end speech recognition model based on Transformer). The speech recognition engine takes a speech waveform tensor (e.g., one second of data with a length of 16000×1) as input and outputs a speech-text sequence (e.g., "I want to go for a walk today"). Next, the system inputs the speech-text into a large-scale language model (e.g., a pre-trained Transformer-based LLM) to extract the intention, emotion, and topic of the speech. Input examples include natural language sentences such as "I want to go for a walk today," "I'm hungry," and "I don't know where I am." The large-scale language model outputs structured data, such as intention labels (e.g., desire to go out, need for food, anxiety), emotion scores (e.g., joy 0.8, anxiety 0.6), and topic categories (e.g., health, movement, food). Furthermore, the inference module infers the individual's current intention (e.g., wanting to go out, wanting to eat, seeking help) based on these output values. The inference results can be used for threshold determination or rule-based branching (e.g., notifying family members when the anxiety score is greater than 0.7). The life log recording function records the individual's activity history (e.g., steps, movement path, voice content, emotional changes) as a time-series database. The anti-wandering function combines GPS location information with behavioral pattern analysis (e.g., deviation from the usual route detection) to issue an alert when abnormal behavior occurs. As a technical effect, this system, through the multi-layer feature extraction and pattern recognition capabilities of the AI ​​model, achieves high-frequency, high-precision voice analysis and intention inference, which is difficult to achieve by manual observation or recording alone, significantly improving the accuracy of grasping the intentions of individuals with cognitive impairment. In addition, through real-time processing and automated data linkage, it can reduce the burden on caregivers and achieve rapid response.The applicable fields include the care of dementia patients, home care support, understanding patient intentions at medical sites, and support facilities for people with disabilities. Furthermore, through the synergy of multiple AI models (speech recognition, natural language understanding, emotion inference, and behavior prediction), it achieves multifaceted life support that traditional single-function devices cannot provide, which is also an important technical feature of this invention.

[0037] The system according to this embodiment includes a collection unit, a parsing unit, and an estimation unit. The collection unit collects the speaker's speech in real time. For example, the collection unit can collect the speaker's speech as voice data. The collection unit can also collect the speaker's speech as text data. Furthermore, the collection unit can also collect the speaker's speech as gesture data. The parsing unit parses the speech data collected by the collection unit. For example, the parsing unit can use speech recognition technology to convert voice data into text data. The parsing unit can also use text parsing technology to parse text data. Furthermore, the parsing unit can use sentiment analysis technology to analyze the sentiment of the speech data. The estimation unit estimates the speaker's intentions based on the information parsed by the parsing unit. For example, the estimation unit can estimate the speaker's decision based on the parsing results. The estimation unit can also estimate the speaker's emotional state based on the parsing results. Furthermore, the estimation unit can estimate the speaker's behavioral intentions based on the parsing results. Therefore, the system according to this embodiment can collect and parse a speaker's speech and estimate intentions in real time. Specifically, this system, acting as the collection unit, is equipped with a high-sensitivity microphone array, IMU sensor, touch sensor, etc., enabling it to acquire speech data (e.g., 16kHz, 16bit, mono PCM), text data (e.g., input records from chat applications), and gesture data (e.g., three-axis acceleration values, pose estimation vectors) in a temporal manner. The collection unit buffers this data and transmits it to the analysis unit periodically. The analysis unit performs spectrogram transformation on the speech data and inputs it into a Transformer-type speech recognition model, outputting a speech-text sequence (e.g., "want to drink water"). For the text data, natural language processing algorithms such as morphological parsing, entity extraction, and dependency parsing are applied to extract the topic, intent, and sentiment words of the speech. For the gesture data, temporal convolutional neural networks (TCN) or LSTM are used to classify action patterns (e.g., wandering, sitting, starting to walk). The sentiment analysis technology employs a BERT-based sentiment classification model or a multimodal sentiment inference model. Inputting speech text or speech features (e.g., F0, energy, spectral envelope), it outputs sentiment labels (e.g., joy, anger, anxiety) and sentiment intensity scores (e.g., 0.7, 0.2, 0.1). The inference unit inputs the features obtained from the analysis unit, such as topic, intent, sentiment, and action patterns, into an inference algorithm such as a multilayer perceptron or decision tree to infer the individual's decisions (e.g., desire to go out, need for rest), emotional state (e.g., calm, anxiety), and behavioral intent (e.g., starting to move, seeking help). The inference results are used for threshold determination or rule-based branching (e.g., notifying family members when anxiety score is high). As a technical advantage, this system can comprehensively analyze multimodal data, achieving higher accuracy and more robust inferences than traditional intent inferences relying on a single data source. Furthermore, through real-time processing, it can respond instantly to changes in the individual's state.Applicable areas include care for dementia patients, support for people with disabilities, understanding patient intentions in medical settings, and home care. Furthermore, by disclosing the internal processing details of the AI ​​model (feature extraction, classification, and inference), the black-box uncertainty is eliminated, clarifying the basis for technological improvements.

[0038] The system includes a recording unit for keeping a daily log. This unit records the user's daily life to comprehensively support their lifestyle. For example, it can record the user's activity history, health data, and location information. Specifically, the recording unit records activity history such as steps, distance traveled, calories burned, and sleep time as time-series data in 1-minute increments. Health data such as heart rate, blood pressure, body temperature, and blood oxygen saturation are acquired via Bluetooth and recorded every 10 seconds. Location information, including latitude, longitude, altitude, and movement speed, is acquired every 5 seconds via a GPS module and recorded along with the time. The recording unit structures this data into a time-series database (e.g., tuples of time, data type, and value) and automatically performs data compression and outlier detection (e.g., automatically marking abnormally high movement speeds). Furthermore, the recording unit also associates the user's voice content and sentiment inference results (e.g., voice-text, sentiment score) with the daily log. Recorded data is automatically backed up to a cloud server or local storage, accessible to family members or medical personnel as needed. As a technological advantage, compared to traditional manual or single-data-type recording, this recording system comprehensively and frequently records various physiological, behavioral, and vocal data, providing a multifaceted understanding of an individual's life. This enables early detection of abnormal behaviors, improved accuracy in health management, and optimization of care plans. Applicable areas include the care of dementia patients, home care, health management applications, and medical data linkage platforms. Furthermore, the recording system is scalable, dynamically adjusting recording frequency and data types based on the individual's condition or external instructions, allowing for flexible adaptation to future functional expansions.

[0039] The system includes a prevention unit that implements anti-loitering functionality to ensure the user's safety. For example, the prevention unit has a location tracking function, enabling it to track the user's location in real time. It also has an alarm notification function, issuing an alarm when the user leaves a specific area. Furthermore, the prevention unit has a behavior restriction function, which can limit the user's behavior. By implementing the anti-loitering function, the user's safety is ensured. Specifically, the prevention unit uses a GPS module or Wi-Fi / Bluetooth beacon to obtain the user's current location (latitude, longitude, altitude) every 5 seconds and determines its relationship to a preset safe zone (e.g., a 200-meter radius around the residence). When the user deviates from the safe zone, the prevention unit immediately sends an alarm signal (e.g., vibration, voice guidance, smartphone notification) to both the user and their caregiver's terminals. As a behavior restriction function, when the user approaches a dangerous area (e.g., a driveway, a station), a voice synthesis alert is given, or family members are automatically notified. Furthermore, the prevention system performs time-series analysis on the individual's movement speed and patterns (e.g., deviating from usual routes, going out late at night). When abnormal behavior is detected, a loitering risk score is calculated, and additional actions are taken (e.g., automatically contacting a security company) if the score exceeds a threshold. As a technical advantage, by combining real-time location and behavior analysis with multi-stage alarms and restriction functions, the prevention system achieves higher accuracy and faster loitering prevention than traditional simple location notification systems. Applicable areas include care for dementia patients, support for people with disabilities, child safety management, and loitering countermeasures in medical and nursing facilities. In addition, the prevention system can be linked to the individual's behavioral history and emotional state, possessing scalability for applying individual optimized prevention algorithms.

[0040] The data collection unit can infer the subject's emotions and adjust the timing of speech collection based on the inferred emotions. For example, the collection unit can increase the frequency of speech collection when the subject is relaxed, and decrease the frequency when the subject is stressed. Furthermore, the collection unit can adjust the timing of speech collection when the subject is excited to avoid missing important statements. By adjusting the timing of speech collection according to the subject's emotions, important statements can be collected without omission. Emotion inference can be achieved through emotion engines or generative AI, etc. Generative AI can be text-generated AI (such as LLM) or multimodal generative AI, but is not limited to these. Specifically, the collection unit acquires the subject's speech or text data in real time and inputs it into the emotion inference module. The emotion inference module uses a BERT-based emotion classification model or a multimodal emotion inference model, inputting speech text (e.g., "I'm happy today," "I'm tired") or speech features (e.g., F0, energy, spectral envelope), and outputting emotion labels (e.g., relaxed, stressed, excited) and emotion intensity scores (e.g., relaxed 0.8, stressed 0.2). The data collection department dynamically sets collection frequency parameters based on the emotional intensity score (e.g., 1-minute intervals when relaxed, 5-minute intervals when stressed, and 30-second intervals when excited). For example, when the individual is relaxed, the focus is on natural conversation flow, with high-frequency collection of statements; when stressed, excessive intervention is avoided, and the collection frequency is reduced; when excited, statements are collected at short intervals to avoid missing important statements or signs of abnormal behavior. Input examples for emotional inference include "I feel good today," "I hope to get help," and "I don't know where I am," with output examples including emotional labels: relaxed, stressed, excited, and emotional intensity scores: 0.7, 0.2, 0.1, etc. Based on these output values, the data collection department implements algorithms to automatically adjust the timing and frequency of collection (e.g., threshold determination, PID control). As a technical effect, the data collection department can optimize the collection timing according to the individual's emotional state, collecting important statements or signs of abnormality with high accuracy and efficiency. Therefore, compared to traditional fixed-interval collection methods, the comprehensiveness of information and system responsiveness are significantly improved. Applicable areas include the care of dementia patients, the monitoring of mental illness patients, and home care support. Furthermore, by changing the type and parameters of the sentiment inference model, individual optimization and diverse use cases can be achieved.

[0041] The data collection department can analyze an individual's past speaking history and select the optimal collection method. For example, it can set collection timing based on the time periods when the individual frequently spoke in the past. It can also analyze the content of the individual's past statements, prioritizing the collection of important statements. Furthermore, the department can optimize the collection method based on the individual's past speaking patterns. Specifically, the department stores an individual's speaking history database (e.g., speaking time, content, sentiment score, and length) in a time-series format, automatically extracting peak speaking times (e.g., 9-10 AM, 7-8 PM) using temporal clustering or autoregressive models. The department increases the collection frequency during these peak times and decreases it during off-peak times. For the content, it calculates an importance score using TF-IDF or BERT embedding vectors, prioritizing the collection and recording of highly important statements (e.g., "Help!", "Want to get out!"). In terms of speech pattern analysis, time-series models such as HMM (Hidden Markov Model) or LSTM are used to detect periodic or abnormal patterns in speech (e.g., frequent speaking late at night). When abnormal patterns occur, the collection method is automatically switched (e.g., collection interval, collection mode). As a technical benefit, the collection department analyzes past speech history using AI models to dynamically optimize collection timing and methods, significantly improving the comprehensiveness and efficiency of information compared to traditional fixed-plan methods. Applicable areas include the care of dementia patients, monitoring of mental illness patients, home care support, and patient status monitoring in medical settings. Furthermore, personalized collection strategies can be achieved through customized optimization of the analysis algorithms and parameters.

[0042] When collecting comments, the collection department can filter them based on the user's lifestyle and areas of interest. For example, it can prioritize comments related to topics the user is currently interested in. The department can also adjust the content of the collected comments based on the user's lifestyle. Furthermore, the department can filter and collect comments related to the user's current activities. By filtering comments based on the user's lifestyle and areas of interest, highly relevant comments can be collected. Specifically, the collection department obtains the user's lifestyle data (e.g., current activity type, location information, health status) and area of ​​interest data (e.g., recent search history, conversation history, interest profile), and compares them with the comments. The comments are vectorized using natural language embedding models such as BERT, and cosine similarity is calculated with the area of ​​interest vectors to obtain a relevance score. The collection department only prioritizes collecting and recording comments with relevance scores higher than a threshold. For example, if the user is interested in "gardening," comments containing keywords such as "flowers," "garden," and "plants" are prioritized. When adjusting based on lifestyle, comments related to exercise are prioritized during exercise, and comments related to diet are prioritized during meals. Current activity is inferred using accelerometers or location information, and collection filters are switched for each activity type (e.g., walking, resting, eating). As a technological advantage, the collection unit, through dynamic filtering based on the individual's lifestyle and areas of interest, can acquire highly relevant information more efficiently than traditional indiscriminate collection methods. Applicable areas include care for dementia patients, personalized health management, home care support, and life support robots. Furthermore, by modifying the filtering algorithm and relevance calculation method, it can flexibly address diverse use cases.

[0043] The data collection unit can infer the individual's emotions and determine the priority order for collecting statements based on the inferred emotions. For example, when the individual is relaxed, the collection unit can prioritize collecting everyday statements. It can also prioritize collecting emotional statements when the individual is stressed. Furthermore, when the individual is excited, the collection unit can prioritize collecting important statements. By prioritizing statements based on the individual's emotions, important statements can be collected preferentially. Emotion inference can be achieved through emotion engines or generative AI, etc. Generative AI can be text-generated AI (such as LLM) or multimodal generative AI, but is not limited to these. Specifically, the collection unit acquires the individual's speech or text data in real time and inputs it into the emotion inference module. The emotion inference module uses a BERT-based emotion classification model or a multimodal emotion inference model, inputting speech text (e.g., "I'm happy today," "I'm tired") or speech features (e.g., F0, energy, spectral envelope), and outputting emotion labels (e.g., relaxed, stressed, excited) and emotion intensity scores (e.g., relaxed 0.8, stressed 0.2). The collection department calculates the priority score of statements based on emotion tags and intensity scores. For example, it prioritizes collecting everyday statements when relaxed (e.g., "The weather is nice today"), emotional statements when stressed (e.g., "Help!", "I can't take it anymore"), and important statements when excited (e.g., "I want to go out!", "I don't know where I am"). The priority score is calculated by combining the importance estimate of the TF-IDF score or BERT embedding vector of the statement content, and the top N statements are forwarded to the recording and parsing departments. As a technical advantage, the collection department can dynamically optimize the priority order of statements based on the individual's emotional state, significantly improving the comprehensiveness of important information and the system's responsiveness, which is superior to the traditional uniform collection method. Applicable areas include the care of dementia patients, the monitoring of mental illness patients, and home care support. Furthermore, by modifying the emotion inference model and priority calculation algorithm, individual optimization and diversified use cases can be achieved.

[0044] When collecting speech data, the collection unit prioritizes highly relevant speech based on the user's geographic location information. For example, when the user is in a specific location, the collection unit can prioritize speech related to that location. It can also prioritize speech related to the user's destination while the user is moving. Furthermore, when the user is in a specific area, the collection unit can prioritize speech related to that area. By considering geographic location information, highly relevant speech can be prioritized. Specifically, the collection unit uses location detection sensors such as GPS modules or Wi-Fi / Bluetooth beacons to obtain the user's current location (latitude, longitude, and altitude) every 5 seconds and records the location information in a time-series database. When collecting voice or text data, the collection unit associates location information with the spoken content. The collection unit vectorizes the spoken content using natural language embedding models such as BERT and calculates cosine similarity with topic vectors associated with the location information (e.g., place categories such as hospitals, supermarkets, and parks) to obtain a relevance score between the speech and the location. The collection unit only prioritizes collecting and recording speech with a relevance score higher than a threshold. For example, when the user is in a hospital, the system prioritizes collecting statements containing keywords such as "examination," "medicine," and "waiting time." When in a supermarket, it prioritizes collecting statements related to "shopping," "shopping cart," and "checkout." While moving, it prioritizes collecting statements related to the predicted destination (e.g., the route from home to the station). Furthermore, the collection department can link with geographic area information (e.g., city area, facility name) to prioritize collecting statements related to local events or situations (e.g., holidays, traffic control, disasters). AI input examples include location information vectors (e.g., latitude 35.6895, longitude 139.6917, altitude 10m) and spoken text (e.g., "Going to the hospital today," "Want to shop at the supermarket"). AI output includes a relevance score (e.g., 0.85), priority labels (e.g., high, medium, low), and a collection availability flag (e.g., 1 or 0). These output values ​​are used for subsequent processing by the collection department, such as transmitting data to the recording department or filtering input data for the parsing department. As a technical advantage, this data collection unit can quantitatively assess the correlation between geographic location information and spoken content in a high-dimensional vector space, achieving higher accuracy and more personalized information collection than traditional indiscriminate methods. This suppresses the accumulation of useless data and improves the efficiency of parsing, inference processing, and overall system responsiveness. Applicable areas include caregiving for dementia patients when they are out, monitoring the status of medical and nursing facilities, optimizing dialogue for life support robots, and regional collaborative health management systems. Furthermore, by modifying the location information acquisition interval and relevance calculation algorithm, it can flexibly address diverse use cases in urban areas, suburbs, and indoor / outdoor environments.

[0045] When collecting posts, the collection department can analyze an individual's social media activity to gather relevant posts. For example, it can collect posts related to topics the individual frequently discusses on social media. The department can also set collection times based on the individual's social media activity periods. Furthermore, the department can analyze the content of the individual's posts on social media, prioritizing the collection of important posts. Specifically, the collection department regularly obtains post history data (such as post time, content, and number of interactions) from multiple social media platforms used by the individual (e.g., Weibo, SNS, forums) via API. The department vectorizes the post content using natural language embedding models such as BERT or RoBERTa, calculates the cosine similarity between the individual's posts and the vectors of the social media post topics, and obtains a relevance score. The department clusters frequently posted topics (e.g., health, interests, family, news), prioritizing the collection and recording of posts with high relevance to these topics. For activity time period analysis, a histogram is created from the submission time data of the past month to automatically extract peak periods (e.g., 20:00–22:00) and reflect this in the speech collection scheduling. Important statements are determined using TF-IDF scores or BERT embedding vectors to estimate their importance, prioritizing the collection of statements similar to those with high interaction (e.g., likes, comments) on the user's social media. AI input examples include social media submission text (e.g., "Started a new interest today," "Haven't been feeling well lately"), submission time (e.g., 2024-06-01 20:15), and statement text (e.g., "Want to try a new interest"). AI output includes relevance score (e.g., 0.92), priority labels (e.g., high, medium, low), and recommended collection time (e.g., 20:00–22:00). These output values ​​are used for subsequent processing by the collection department, such as transmitting data to the recording department, filtering input data for the parsing department, and setting collection scheduler parameters. As a technological advantage, this data collection system can quantitatively assess the relevance of an individual's social media activities and posted content in a high-dimensional vector space. This achieves higher accuracy and better alignment with the individual's interests and social activities compared to traditional methods that rely solely on posted content. This allows for earlier understanding of an individual's social connections and changes in their psychological state, improving the quality of care and medical support, and enabling personalized health management. Applicable areas include preventing social isolation in dementia patients, monitoring mental illness patients, home care support, and personalized health management applications. Furthermore, by modifying the types of social media APIs and the relevance calculation algorithm, it can flexibly address multiple platforms and diverse use cases.

[0046] The analysis unit can infer the user's emotions and adjust the analysis presentation based on the inferred emotions. For example, when the user is relaxed, the analysis unit can provide detailed analysis results. When the user is stressed, the analysis unit can provide concise analysis results. In addition, when the user is excited, the analysis unit can provide visually easy-to-understand analysis results. By adjusting the analysis presentation according to the user's emotions, appropriate analysis results can be provided. Emotion inference can be achieved through emotion inference functions such as emotion engines or generative AI. Generative AI can be text generation AI (such as LLM) or multimodal generation AI, but is not limited to these. Specifically, this analysis unit acquires the user's speech or text data in real time and inputs it into the emotion inference module. The emotion inference module uses a BERT-based emotion classification model or a multimodal emotion inference model, inputting speech text (e.g., "I'm happy today," "I'm tired") or speech features (e.g., F0, energy, spectral envelope), and outputting emotion labels (e.g., relaxed, stressed, excited) and emotion intensity scores (e.g., relaxed 0.8, stressed 0.2). This analysis unit dynamically switches the presentation of analysis results based on sentiment tags and intensity scores. When relaxed, it generates a detailed report containing comprehensive semantic analysis of the speech content (e.g., topic extraction, intent inference, sentiment change maps, etc.); when stressed, it provides a concise summary extracting only key points (e.g., topic tags, sentiment tags, list of important statements); when excited, it outputs analysis results with a visually emphasized interface (e.g., color charts, icon displays, sentiment change animations, etc.). AI input examples include spoken text (e.g., "I'm in a good mood today," "I hope to get help"), speech feature vectors (e.g., F0=120Hz, energy=0.8), and sentiment intensity scores (e.g., 0.7, 0.2, 0.1). AI output includes analysis presentation type (e.g., detailed, concise, visualization) and analysis result data (e.g., topic category, sentiment change map, summary text). These output values ​​are used for subsequent processing by the analysis unit, such as selecting the user interface display format or transmitting data to the recording unit. As a technological advantage, this analytics tool optimizes the presentation of analytics results based on the individual's emotional state, providing information that is easier for both the individual and caregiver to understand and more relevant to the actual situation than traditional one-size-fits-all output methods. This improves information delivery efficiency, reduces stress, and supports rapid decision-making. Applicable areas include the care of dementia patients, monitoring of mental illness patients, home care support, and patient status reporting in healthcare settings. Furthermore, by modifying the analytics algorithm and interface design, it can flexibly address diverse use cases and individual optimizations.

[0047] During parsing, the parsing unit can adjust the level of detail based on the importance of the speech. For example, it can perform detailed parsing for important speeches, as well as concise parsing for everyday speeches. Furthermore, it can provide detailed analysis of emotional changes in emotionally charged speeches. By adjusting the level of detail based on the importance of the speech, appropriate parsing results can be provided. Specifically, this parsing unit vectorizes the speech content using natural language embedding models such as BERT, and calculates the importance score of the speech using TF-IDF scores or keyword extraction algorithms. For speeches with high importance scores (such as "help," "want to go out," etc.), it performs multi-stage detailed parsing including topic extraction, intent inference, emotional change analysis, and related event detection, and outputs structured data (such as topic categories, intent tags, and emotional change maps). For everyday speeches (such as "the weather is nice today," it only generates concise summaries such as topic tags and emotional tags). For emotionally charged speeches (such as "can't take it anymore," "help," etc.), it performs detailed parsing of the temporal changes in emotional intensity scores and outlier detection, outputting emotional change maps or anomaly detection markers. AI input examples include spoken text (e.g., "Help," "The weather is nice today"), importance scores (e.g., 0.95, 0.2), and sentiment intensity scores (e.g., 0.8, 0.1, 0.1). AI output includes parsing detail labels (e.g., detailed, concise) and parsing result data (e.g., topic category, sentiment change graph, summary text). These output values ​​are used for subsequent processing by the parsing department, such as user interface display or data transmission to the recording department. As a technical benefit, this parsing department can dynamically optimize parsing detail based on the importance of the speech, significantly improving the comprehensiveness of important information and system responsiveness, superior to traditional uniform parsing methods. This allows caregivers or medical personnel to quickly and accurately obtain the information they need. Applicable areas include the care of dementia patients, monitoring of mental illness patients, home care support, and patient status monitoring in the medical field. Furthermore, by changing the importance calculation algorithm and parsing detail threshold settings, it can flexibly address diverse use cases and individual optimizations.

[0048] During parsing, the parsing unit can apply different parsing algorithms based on the category of the speech. For example, it can apply sentiment analysis algorithms to emotional speech. It can also apply natural language processing algorithms to everyday speech. Furthermore, it can apply detailed semantic parsing algorithms to important speech. By applying different parsing algorithms according to the category of the speech, appropriate parsing results can be provided. Specifically, this parsing unit pre-inputs the speech content into a category classification model (e.g., a multi-class classifier, a BERT-based speech category classification model) to determine the speech category (e.g., emotional, everyday, important). For emotional speech (e.g., "I can't take it anymore," "Help!"), it applies a BERT-based sentiment classification model or a multimodal sentiment inference model to analyze sentiment tags and sentiment intensity scores in detail. For everyday speech (e.g., "The weather is nice today"), it applies natural language processing algorithms such as morphological parsing, entity extraction, and dependency structure parsing to generate topic tags or summary text. For important statements (such as "want to go out," "don't know where I am," etc.), semantic parsing algorithms (such as semantic similarity calculation of BERT embedding vectors, intent inference models) are applied to output detailed parsing results, including topic category, intent label, and related event detection. AI input examples include spoken text (such as "help," "the weather is nice today," "want to go out") and category labels (such as sentimentality, everyday occurrence, importance). AI output includes parsing algorithm selection labels (such as sentiment analysis, natural language processing, semantic parsing) and parsing result data (such as sentiment intensity score, topic category, intent label), etc. These output values ​​are used for subsequent processing by the parsing department, such as user interface display or data transmission to the recording department. As a technical effect, this parsing department can dynamically select and apply the optimal parsing algorithm for each statement category, significantly improving parsing accuracy and information comprehensiveness, which is superior to traditional uniform parsing methods. As a result, caregivers or medical personnel can quickly take appropriate measures according to the situation. Applicable areas include the care of dementia patients, monitoring of mental illness patients, home care support, and patient status monitoring in the medical field. In addition, by changing the category classification model and the type of parsing algorithm, it can flexibly respond to diverse use cases and individual optimizations.

[0049] The parsing unit can infer the subject's emotions and adjust the parsing length based on the inferred emotions. For example, the parsing unit can perform detailed parsing when the subject is relaxed. It can also perform concise parsing when the subject is stressed. Furthermore, it can perform visually understandable parsing when the subject is excited. By adjusting the parsing length according to the subject's emotions, appropriate parsing results can be provided. Emotion inference can be achieved through emotion engines or generative AI, etc. Generative AI can be text-generated AI (such as LLM) or multimodal generative AI, but is not limited to these. Specifically, this parsing unit acquires the subject's speech or text data in real time and inputs it into the emotion inference module. This emotion inference module uses a BERT-based emotion classification model or a multimodal emotion inference model that simultaneously inputs speech and text. AI input examples include spoken text (e.g., "I'm in a good mood today," "I hope to get help"), speech feature vectors (e.g., F0=120Hz, energy=0.8, spectral envelope=0.6), and gesture data (e.g., acceleration values ​​0.2, 0.1, 0.3). The sentiment inference model outputs sentiment labels (e.g., relaxed, stressed, excited) and sentiment intensity scores (e.g., relaxed 0.8, stressed 0.1, excited 0.1) based on these inputs. The parsing unit dynamically determines the parsing length parameters (e.g., detailed, concise, visual) based on the sentiment labels and intensity scores. When relaxed, it generates a long report containing detailed semantic parsing of the speech content (e.g., topic extraction, intent inference, sentiment change graph generation, related event detection, etc.). When stressed, it generates a short summary extracting only the key points (e.g., topic labels, sentiment labels, list of important statements). When excited, it outputs the parsing results with a visually emphasized interface (e.g., color charts, icon displays, sentiment change animations, etc.). AI output examples include parsing length labels (e.g., detailed, concise, visual) and parsing result data (e.g., topic categories, sentiment change graphs, summary text). These output values ​​are used for subsequent processing by the parsing department, such as selecting the user interface display format or transmitting data to the recording department. As a technical benefit, this parsing department can optimize the length and presentation of parsing results based on the individual's emotional state, providing information that is easier for both the individual and caregivers to understand and more relevant to the actual situation than traditional uniform output methods. This improves information delivery efficiency, reduces stress, and supports rapid decision-making. Applicable areas include the care of dementia patients, monitoring of mental illness patients, home care support, and patient status reporting in the medical field. Furthermore, by modifying the parsing length control algorithm and interface design, it can flexibly address diverse use cases and individual optimizations. By changing the type and parameters of the sentiment inference model (e.g., threshold, weighting), the parsing department can achieve optimal parsing length control based on individual characteristics and on-site needs, significantly improving the flexibility and scalability of computer technology.

[0050] During parsing, the parsing unit can determine the parsing priority based on the submission time of the speech. For example, the parsing unit can prioritize parsing the most recent speech. It can also prioritize parsing important past speeches. Furthermore, it can prioritize parsing speeches related to specific events. By determining the parsing priority based on the submission time of the speech, appropriate parsing results can be provided. Specifically, this parsing unit obtains metadata such as the timestamp (e.g., 2024-06-01 10:15:00), speech content, importance score, and event tags for each speech from the speech database. The parsing unit uses a temporal sorting algorithm or a priority queue to prioritize the parsing of the latest speech or event-related speech. AI input examples include spoken text (e.g., "Help," "The weather is nice today"), submission time (e.g., 2024-06-01 10:15:00), and event tags (e.g., "going to the doctor," "going out," "eating"). The AI ​​outputs a parsing priority score (e.g., 0.95, 0.2) and a parsing order tag (e.g., high, medium, low) based on the submission time and event tag. The parsing department performs multi-stage analysis, including detailed semantic parsing, sentiment inference, and topic extraction, based on the order of high-priority speeches. For important past speeches, importance is estimated using TF-IDF or BERT embedding vectors, prioritizing the parsing of high-importance content. For speeches related to specific events, priority is dynamically adjusted based on the proximity of the event's occurrence time to the speech's time and content relevance. These output values ​​are used in subsequent processing by the parsing department, such as the display order on the user interface or the order of data transmission to the recording department. As a technical advantage, this parsing department significantly improves the rapid extraction of important information and situational responsiveness through dynamic priority control based on the timing of speech submission and event relevance, surpassing traditional simple time-series parsing methods. This allows caregivers or medical personnel to obtain the necessary information in real time, enabling rapid decision-making and response. Applicable areas include dementia patient care, patient status monitoring in medical settings, home care support, and emergency response systems. Furthermore, by changing the parsing priority algorithm and event label assignment method, it can flexibly address diverse use cases and individual optimizations. The parsing department can adjust the weighting and threshold settings of priority scores according to on-site needs and individual characteristics, significantly improving the flexibility and scalability of computer technology.

[0051] During parsing, the parsing unit can adjust the parsing order based on the relevance of the statements. For example, it can prioritize parsing statements with high relevance and postpone parsing statements with low relevance. Furthermore, the parsing unit can optimize the parsing order based on the relevance of the statements. By adjusting the parsing order according to the relevance of the statements, appropriate parsing results can be provided. Specifically, this parsing unit vectorizes the speech content using natural language embedding models such as BERT or RoBERTa, and calculates the relevance scores between statements using cosine similarity or clustering methods (such as K-means, hierarchical clustering). AI input examples include spoken text (e.g., "want to go out," "don't know where I am," "the weather is nice today"), and spoken vectors (e.g., 768-dimensional BERT embedding). The AI ​​outputs a relevance score matrix (e.g., 0.92 between statements A and B, 0.15 between statements A and C) and relevant clustering labels (e.g., related to going out, daily conversation, health). The parsing unit prioritizes parsing groups of statements with high relevance scores and postpones the parsing of statements with low relevance. Further application of relevance-based order optimization algorithms (such as graph-based priority decision-making and dynamic programming) balances parsing efficiency and information comprehensiveness. These output values ​​are used for subsequent processing by the parsing unit, such as the display order on the user interface or the order of data transmission to the recording unit. As a technical effect, this parsing unit can quantitatively evaluate the relevance of spoken content in a high-dimensional vector space, achieving higher accuracy in information grouping and situational understanding than traditional simple time-series parsing methods. Therefore, caregivers or medical personnel can grasp relevant information at once, enabling rapid and accurate responses. Applicable areas include the care of dementia patients, patient status monitoring in medical settings, home care support, and optimization of dialogue in life support robots. Furthermore, by modifying the relevance calculation algorithm and clustering method, it can flexibly address diverse use cases and individual optimizations. The parsing unit can adjust the relevance score threshold and the number of clusters according to on-site needs and individual characteristics, significantly improving the flexibility and scalability of computer technology.

[0052] The inference unit can infer an individual's emotions and adjust its inference method based on the inferred emotions. For example, the inference unit can infer detailed intentions when the individual is relaxed. It can also infer concise intentions when the individual is stressed. Furthermore, the inference unit can consider emotional changes and infer intentions when the individual is excited. By adjusting intentions based on the individual's emotions, appropriate intentions can be inferred. Emotion inference can be achieved through emotion inference functions such as emotion engines or generative AI. Generative AI can be text-generated AI (such as LLM) or multimodal generative AI, but is not limited to these. Specifically, this inference unit takes the speech content, emotion labels, emotion intensity scores, and other features from the parsing unit as input and uses inference algorithms (such as multilayer perceptrons, decision trees, and rule-based inferrs) to infer intentions. AI input examples include spoken text (e.g., "want to go out," "help"), emotion tags (e.g., relaxed, stressed, excited), emotion intensity scores (e.g., 0.8, 0.1, 0.1), speech time, and related event tags. When relaxed, the presumption department comprehensively analyzes speech content, past behavioral history, and living situation data to presume detailed intentions (e.g., desire to go out, desire for hobbies / activities); when stressed, it extracts key points to presume concise intentions (e.g., need for rest, seeking help); when excited, it considers changes in emotion intensity and outliers to presume intentions with high urgency (e.g., seeking help, signs of abnormal behavior). AI output examples include intention tags (e.g., desire to go out, need for rest, seeking help), presumption confidence scores (e.g., 0.92), and presumption detail tags (e.g., detailed, concise, urgent). These output values ​​are used for subsequent processing by the presumption department, such as sending alarms to the notification system or caregiver's terminal, and transmitting data to the recording department. As a technological advantage, this inference unit can dynamically optimize the intention inference method based on the individual's emotional state, achieving a higher accuracy, more realistic understanding of intentions, and faster response compared to traditional uniform inference methods. This allows caregivers or medical personnel to respond instantly to changes in the individual's state, improving the quality and safety of support. Applicable areas include the care of dementia patients, monitoring of mental illness patients, home care support, and understanding patient intentions in the medical field. Furthermore, by changing the type and parameters of the inference algorithm and emotional inference model, it can flexibly address diverse use cases and individual optimizations. The inference unit can adjust the inference method switching logic and confidence thresholds according to on-site needs and individual characteristics, significantly improving the flexibility and scalability of computer technology.

[0053] During the estimation process, the estimation department can optimize the estimation algorithm by referring to past speaking data. For example, the estimation department can adjust the estimation algorithm based on past speaking data. The estimation department can also improve the accuracy of intention estimation by referring to past speaking content. In addition, the estimation department can also optimize the estimation algorithm by analyzing past speaking patterns. By referring to past speaking data, the estimation algorithm can be optimized and the estimation accuracy can be improved. Specifically, this estimation department stores a database of the speaker's speaking history over the past week to month in a time-series manner (e.g., speaking time, speaking content, sentiment score, speaking length, and estimated intention tags), and automatically extracts peak speaking frequency periods and speaking patterns using time-series clustering, autoregressive models (AR models), LSTM, and other time-series neural networks. Examples of AI input include past speech-text sequences (e.g., "Help," "Want to go out," "The weather is nice today"), speech time sequences (e.g., 2024-06-01 10:15, 2024-06-01 12:30, ...), and sentiment score sequences (e.g., 0.8, 0.2, 0.1, ...). The inference department dynamically optimizes inference algorithm parameters (e.g., weights, thresholds, feature selection) based on extracted patterns and features. For example, during periods when phrases like "Help" and "Want to go out" were frequently used, the weight of the inference for wanting to go out or seeking help is increased. The inference algorithm is automatically switched when abnormal patterns occur, combining the change patterns of the importance score and sentiment score of the speech content. Examples of AI output include optimized inference algorithm parameters, inference confidence scores (e.g., 0.95), and inference intention labels (e.g., wanting to go out, seeking help). These output values ​​are used for subsequent processing by the inference department, such as notifying the system or transmitting data to the recording department. As a technological achievement, this estimation department analyzes past speaking history using an AI model and dynamically optimizes the estimation algorithm, significantly improving information comprehensiveness and estimation accuracy compared to traditional fixed-parameter methods. This enables high-precision, personalized estimations that reflect individual changes in state and unique characteristics. Applicable areas include the care of dementia patients, monitoring of mental illness patients, home care support, and patient condition monitoring in medical settings. Furthermore, by modifying the analysis algorithm and parameter optimization methods, it can flexibly address diverse use cases and individual optimizations. The estimation department can adjust the optimization logic and feature selection criteria according to on-site needs and individual characteristics, greatly enhancing the flexibility and scalability of computer technology.

[0054] During inference, the inference unit can apply different inference methods based on the category of the speech. For example, for emotional speech, the inference unit can infer intent based on sentiment analysis. It can also infer intent for everyday speech based on natural language processing. Furthermore, it can infer intent for important speech based on detailed semantic analysis. By applying different inference methods to the category of speech, appropriate intent can be inferred. Specifically, this inference unit pre-inputs the speech content from the analysis unit into a category classification model (e.g., a multi-class classifier, a BERT-based speech category classification model) to determine the speech category (e.g., emotional, everyday, important). For emotional speech (e.g., "I can't take it anymore," "Help!"), this inference unit applies a BERT-based sentiment classification model or a multimodal sentiment inference model to analyze sentiment labels and sentiment intensity scores in detail, and inputs these sentiment features into inference algorithms such as multilayer perceptrons or decision trees to infer the speaker's intent (e.g., seeking help, wanting rest). For everyday statements (e.g., "The weather is nice today"), natural language processing algorithms such as morphological parsing, entity extraction, and dependency structure parsing are applied to extract topic tags or summary text, and then infer everyday intentions (e.g., wanting to continue the conversation, wanting to chat). For important statements (e.g., "wanting to go out," "not knowing where I am"), semantic similarity calculation using BERT embedding vectors or an intent inference model are applied, combined with detailed parsing results such as topic category, intent tags, and related event detection, to infer intentions with high urgency or importance (e.g., wanting to go out, requesting emergency support). AI input examples include spoken text (e.g., "Help," "The weather is nice today," "wanting to go out"), category tags (e.g., sentimentality, everyday, importance), speech features (e.g., F0, energy, spectral envelope), sentiment intensity scores (e.g., 0.8, 0.1, 0.1), etc. AI output includes intent tags (e.g., wanting to go out, needing rest, seeking help), inferred confidence scores (e.g., 0.92), and inferred detail tags (e.g., detailed, concise, urgent), etc. These output values ​​are used for subsequent processing by the estimation department, such as sending alarms to the notification system or caregiver terminal, and transmitting data to the recording department. As a technical advantage, this estimation department can dynamically select and apply the optimal estimation algorithm for each speech category, significantly improving parsing accuracy and information comprehensiveness, superior to traditional uniform estimation methods. This allows caregivers or medical personnel to quickly take appropriate measures based on the situation. Applicable areas include the care of dementia patients, monitoring of mental illness patients, home care support, and patient status monitoring in the medical field. Furthermore, by changing the category classification model and estimation algorithm, it can flexibly address diverse use cases and individual optimizations. The estimation department can adjust the estimation method switching logic and confidence threshold according to on-site needs and individual characteristics, greatly improving the flexibility and scalability of computer technology.

[0055] The inference unit can infer an individual's emotions and determine the priority of inferences based on the inferred emotions. For example, when the individual is relaxed, the inference unit can prioritize inferring detailed intentions. Furthermore, when the individual is stressed, the inference unit can prioritize inferring concise intentions. Moreover, when the individual is excited, the inference unit can consider changes in emotion and prioritize inferring intentions. Thus, by determining the priority of inferences based on the individual's emotions, appropriate intentions can be inferred. Emotion inference can be achieved through emotion inference functions such as emotion engines or generative AI. Generative AI can be text generation AI (e.g., LLM) or multimodal generation AI, but is not limited to these. For example, this inference unit takes features such as speech content, emotion tags, and emotion intensity scores received from the parsing unit as input and uses inference algorithms (such as multilayer perceptrons, decision trees, and rule-based inferrs) to infer intentions. Examples of AI inputs include spoken text (e.g., "want to go out," "help"), emotion tags (e.g., relaxed, stressed, excited), emotion intensity scores (e.g., 0.8, 0.1, 0.1), speaking time, and related event tags. When relaxed, the presumption department comprehensively analyzes the spoken content, past behavioral history, and life situation data to prioritize the presumption of detailed intentions (e.g., desire to go out, desire to engage in activities of interest); when stressed, it prioritizes the presumption of concise intentions that extract only the key points (e.g., need for rest, seeking help); when excited, it considers changes in emotion intensity or outliers and prioritizes the presumption of intentions with high urgency (e.g., seeking help, signs of abnormal behavior). Examples of AI outputs include intention tags (e.g., desire to go out, need for rest, seeking help), presumption confidence scores (e.g., 0.92), and presumption priority tags (e.g., high, medium, low). These output values ​​serve as subsequent processing by the presumption department and can be used for notification systems, alarm sending to nursing staff terminals, and data forwarding to the recording department. In terms of technical effectiveness, this inference unit can dynamically optimize inference priorities based on the individual's emotional state, achieving a more accurate grasp of intent and rapid response than traditional uniform inference methods. This allows caregivers or medical professionals to respond instantly to changes in the individual's state, improving the quality and safety of support. Applicable areas include dementia patient care, mental illness patient monitoring, home care support, and patient intent assessment in medical settings. Furthermore, by changing the type and parameters of the inference algorithm or emotional inference model, it can flexibly address diverse use cases and individual optimizations. The inference unit can adjust the inference priority decision logic and confidence thresholds according to on-site needs and individual characteristics, thereby significantly improving the flexibility and scalability of computer technology.

[0056] The presumption department can weight presumptions based on the submission time of the speech. For example, the presumption department can prioritize recent speeches to presume intent, or it can prioritize important past speeches. Furthermore, the presumption department can also prioritize speeches related to specific events to presume intent. Thus, by weighting presumptions based on the submission time of speeches, appropriate intents can be presumed. Specifically, this presumption department obtains metadata such as the timestamp (e.g., 2024-06-01 10:15:00), speech content, importance score, and event tags for each speech from the speech database. Using a time-series sorting algorithm or priority queue, it weights the latest speeches or event-related speeches and inputs them into the presumption algorithm. Examples of AI input include speech text (e.g., "Help," "The weather is nice today"), submission time (e.g., 2024-06-01 10:15:00), event tags (e.g., seeking medical attention, going out, dining), and importance scores (e.g., 0.95, 0.2). The AI ​​outputs a weighted score (e.g., 0.9, 0.3) and a priority label (e.g., high, medium, low) based on the submission timing and event label. The estimation department performs detailed semantic analysis, sentiment estimation, and topic extraction in multiple stages, starting with statements that have high weighted scores. For important past statements, importance is estimated using TF-IDF or BERT embedding vectors, prioritizing the re-estimation of high-importance content. For statements related to specific events, the weighting is dynamically adjusted by calculating the proximity of the event's occurrence time to the statement's time and the relevance of its content. These output values ​​serve as subsequent processing by the estimation department and can be used to notify the system or the data forwarding order of the recording department. Technically, this estimation department, through dynamic weighting control based on statement submission timing and event relevance, significantly improves the rapid extraction of important information and contextual response capabilities compared to traditional simple time-series estimation methods. This allows caregivers or medical professionals to access necessary information in real time, enabling rapid decision-making and response. Applicable areas include dementia patient care, patient status monitoring in medical settings, home care support, and emergency response systems. Furthermore, by changing the presumed weighting algorithm or the event label assignment method, diverse use cases and individual optimizations can be flexibly addressed. The presumed department can adjust the weighting score and threshold settings according to on-site needs and individual characteristics, thereby significantly improving the flexibility and scalability of computer technology.

[0057] The estimation department can make inferences by referring to relevant market data when making inferences. For example, the estimation department can estimate intentions based on relevant market data, or it can estimate intentions by considering market trends. Furthermore, the estimation department can also refer to market data to improve the accuracy of intention estimation. Thus, by referring to relevant market data, the accuracy of intention estimation can be improved. Specifically, this estimation department regularly obtains the latest market prices, demand dynamics, and trend indicators (such as stock prices, consumer indices, search trends, etc.) from external market databases or open data APIs, and correlates them with the content of the speech and the input features of intention estimation. Examples of AI inputs include speech text (such as "want to buy new home appliances" or "want to eat out"), market price data (such as average price of home appliance categories, cost index of eating out), trend scores (such as home appliance demand 1.2 times, trend of eating out 0.8), etc. The inference department calculates the quantitative relevance between spoken content and market data using BERT embedding vectors or cosine similarity. When the relevance score is high, it combines market trends to infer intentions (e.g., appliance purchase intentions, dining-out desires). AI output examples include intention tags (e.g., appliance purchase desires, dining-out desires), inference confidence scores (e.g., 0.88), and market linkage scores (e.g., 0.75). These output values ​​serve as subsequent processing by the inference department and can be used for notification systems, data forwarding in the recording department, and linkage with recommendation systems. Technically, this inference department achieves higher-precision intention inferences that better align with socioeconomic conditions than traditional inference methods based solely on spoken content by quantitatively evaluating the relevance between external market data and spoken content in a high-dimensional vector space. This allows for early understanding of changes in individual consumption behavior and lifestyle intentions, enabling personalized support and improved recommendation quality. Applicable areas include consumer behavior analysis, personalized recommendations, life support robots, and health management applications. Furthermore, by changing the type of market data API or the relevance calculation algorithm, it can flexibly address multiple industries and use cases. The estimation department can adjust the market data linkage logic and weighting parameters according to on-site needs and individual characteristics, thereby greatly improving the flexibility and scalability of computer technology.

[0058] The recording unit can infer an individual's emotions and adjust the recording method of the life log based on the inferred emotions. For example, the recording unit can record a detailed life log when the individual is relaxed, and a concise life log when the individual is stressed. Furthermore, the recording unit can also consider emotional changes when the individual is excited and record the life log accordingly. Thus, by adjusting the life log recording method based on the individual's emotions, an appropriate life log can be recorded. Emotion inference can be achieved through emotion engine or emotion inference functions such as generative AI. Generative AI can be text generation AI (such as LLM) or multimodal generation AI, but is not limited to these. Specifically, this recording unit acquires the individual's spoken audio data, text data, vital sign sensor data (such as heart rate, blood pressure, body temperature), accelerometer sensor data, GPS location information, etc. in real time, and inputs them into the emotion inference module. The emotion inference module adopts a BERT-based emotion classification model or a multimodal emotion inference model that integrates speech, text, and physiological information. AI input examples include spoken text (e.g., "I'm in a good mood today," "I'm very tired"), speech features (e.g., F0=120Hz, energy=0.8), vital signs data (e.g., heart rate 80bpm, blood pressure 120 / 80), and acceleration values ​​(e.g., 0.2, 0.1, 0.3). AI outputs emotion labels (e.g., relaxed, stressed, excited) and emotion intensity scores (e.g., relaxed 0.8, stressed 0.1, excited 0.1). Based on the emotion labels and intensity scores, this recording department dynamically determines the parameters of the life log recording method (e.g., recording granularity, number of recording items, recording format). When relaxed, a detailed life log is recorded in 1-minute increments (e.g., activity history, health data, spoken content, emotion change graphs, etc.); when stressed, only the main events and key points are extracted in 5-minute increments to generate a concise life log (e.g., event timestamps, emotion labels, main spoken words); when excited, the focus is on recording changes in emotion intensity and outliers, generating a life log with emotion change graphs and anomaly detection markers. Examples of AI outputs include recording method tags (e.g., detailed, concise, anomaly emphasis), a list of recorded items (e.g., activity history, emotional changes, abnormal events), and recording frequency (e.g., 1 minute, 5 minutes, when the event occurred). These output values, as subsequent processing by the recording department, can be used for time-series database storage, cloud backup, and data forwarding to home and healthcare professional terminals. Technically, this recording department dynamically optimizes the life log recording method based on the individual's emotional state, significantly improving the comprehensiveness, importance, and responsiveness of information compared to traditional uniform recording methods. This enables early detection of abnormal behavior and changes in health status, achieving optimal care planning and medical responses. Applicable areas include dementia patient care, mental illness patient monitoring, home care support, health management applications, and medical data linkage platforms.Furthermore, by changing the types and parameters of the emotion inference model or recording method control algorithm, diverse use cases and individual optimizations can be flexibly addressed. The recording department can adjust the recording method switching logic, recording granularity, and project selection criteria according to on-site needs and individual characteristics, thereby significantly improving the flexibility and scalability of computer technology.

[0059] The recording department can optimize its recording algorithm by referencing past life log data. For example, it can adjust the recording algorithm based on past life log data or improve recording accuracy by referencing past life log content. Furthermore, the recording department can analyze past life log patterns to optimize the recording algorithm. Thus, by referring to past life log data, the recording algorithm can be optimized and recording accuracy improved. Specifically, this recording department stores a database of the individual's life logs from the past week to the past month (e.g., activity history, health data, speech content, emotional score, recording frequency, recording granularity) in chronological order. It uses temporal clustering, autoregression models (AR models), LSTM, and other temporal neural networks to automatically extract recording patterns and outlier trends. Examples of AI input include past life log data sequences (e.g., steps, heart rate, number of speech, emotional score), recording time sequences (e.g., 2024-06-01 10:15, 2024-06-01 12:30), and recording granularity (e.g., 1 minute, 5 minutes, when the event occurred). The recording department dynamically optimizes recording algorithm parameters (such as recording frequency, recording item selection, and anomaly detection threshold) based on extracted patterns and features. For example, if abnormal values ​​frequently occur during a certain period (such as a sharp increase in heart rate or wandering at night), the recording frequency and the number of recording items are increased; conversely, if a stable pattern persists, the recording granularity is coarsened to reduce storage and communication burden. AI output examples include optimized recording algorithm parameters, recording accuracy scores (e.g., 0.95), and a list of recording items (e.g., activity history, emotional changes, and abnormal events). These output values, as subsequent processing by the recording department, can be used for time-series database storage, cloud backup, and data forwarding to home and healthcare professional terminals. Technically, this recording department analyzes past life log history using an AI model and dynamically optimizes the recording algorithm, significantly improving the comprehensiveness and recording accuracy of information compared to traditional fixed-parameter methods. This enables high-precision life log recording tailored to individual state changes and specific characteristics. Applicable areas include dementia patient care, mental illness patient monitoring, home care support, patient status monitoring in medical settings, and health management applications. Furthermore, by changing the analysis algorithm or parameter optimization method, diverse use cases and individual optimizations can be flexibly addressed. The recording department can adjust the optimization logic and feature selection criteria according to on-site needs and individual characteristics, thereby significantly improving the flexibility and scalability of computer technology.

[0060] The recording unit can infer the individual's emotions and adjust the recording frequency based on the inferred emotions. For example, the recording unit can record daily log entries frequently when the individual is relaxed, and reduce the recording frequency when the individual is stressed. Furthermore, the recording unit can adjust the recording frequency when the individual is excited to avoid missing important events. Thus, by adjusting the recording frequency according to the individual's emotions, appropriate daily log entries can be recorded. Emotion inference can be achieved through emotion engines or emotion inference functions such as generative AI. Generative AI can be text-generated AI (such as LLM) or multimodal generative AI, but is not limited to these. Specifically, this recording unit acquires the individual's spoken audio data, text data, vital sign sensor data, accelerometer sensor data, GPS location information, etc. in real time and inputs them into the emotion inference module. The emotion inference module uses a BERT-based emotion classification model or a multimodal emotion inference model that integrates speech, text, and physiological information. AI input examples include spoken text (e.g., "I'm in a good mood today," "I'm very tired"), speech features (e.g., F0=120Hz, energy=0.8), vital signs data (e.g., heart rate 80bpm, blood pressure 120 / 80), and acceleration values ​​(e.g., 0.2, 0.1, 0.3). The AI ​​outputs emotion labels (e.g., relaxed, stressed, excited) and emotion intensity scores (e.g., relaxed 0.8, stressed 0.1, excited 0.1). Based on the emotion labels and intensity scores, this recording unit dynamically sets the recording frequency parameters (e.g., 1-minute intervals for relaxation, 5-minute intervals for stress, and 30-second intervals for excitement). During relaxation, it records the log frequently, emphasizing a natural rhythm; during stress, it reduces the recording frequency to avoid the burden of excessive recording; during excitement, it records the log at short intervals to ensure no important events or signs of abnormal behavior are missed. Examples of AI outputs include recording frequency labels (e.g., high, medium, low), recording intervals (e.g., 1 minute, 5 minutes, 30 seconds), and recording availability flags (e.g., 1 or 0). These output values, as subsequent processing by the recording department, can be used for time-series database storage, cloud backup, and data forwarding to home and healthcare professional terminals. Technically, this recording department dynamically optimizes the recording frequency based on the individual's emotional state, significantly improving the comprehensiveness of information and system responsiveness compared to traditional fixed-interval recording methods. This allows for early detection of abnormal behavior and changes in health status, achieving optimal care plans and medical responses. Applicable areas include dementia patient care, mental illness patient monitoring, home care support, health management applications, and medical data linkage platforms. Furthermore, by changing the type and parameters of the emotion inference model or recording frequency control algorithm, diverse use cases and individual optimizations can be flexibly addressed. The recording department can adjust the recording frequency switching logic and recording interval setting standards according to on-site needs and individual characteristics, thereby significantly improving the flexibility and scalability of computer technology.

[0061] The recording department can weight the recorded data based on the submission time of the life log. For example, the recording department can prioritize recording recent life logs or important past events. Furthermore, the recording department can also prioritize recording life logs related to specific events. Thus, by weighting the recorded data based on the submission time of the life log, appropriate life logs can be recorded. Specifically, this recording department obtains metadata such as timestamps (e.g., 2024-06-01 10:15:00), record content, event tags (e.g., medical visit, outing, dining), and importance scores for each record from the life log database. Using a time-series sorting algorithm or priority queue, it weights the latest records or event-related records and inputs them into the recording algorithm. Examples of AI input include life log content (e.g., steps, heart rate, spoken content), recording time (e.g., 2024-06-01 10:15:00), event tags (e.g., medical visit, outing, dining), and importance scores (e.g., 0.95, 0.2). The AI ​​outputs a weighted score (e.g., 0.9, 0.3) and a priority label (e.g., high, medium, low) based on the submission time and event tags. The recording department processes records with high weighted scores sequentially through multiple stages, including detailed recording, backup, and data forwarding. For important past events, importance is estimated using TF-IDF or BERT embedding vectors, prioritizing the re-recording and backup of high-importance content. Records related to specific events are dynamically weighted based on the proximity of the event's occurrence time to the recording time and the relevance of the content. These output values ​​serve as subsequent processing data for the recording department and can be used for time-series database storage, cloud backup, and data forwarding order for home and healthcare worker terminals. Technically, this recording department, through dynamic weighting control based on the submission time and event relevance of daily logs, significantly improves the rapid extraction of important information and the ability to respond to situations compared to traditional simple time-series recording methods. This allows caregivers and healthcare professionals to access the necessary information in real time, enabling rapid decision-making and response. Applicable areas include dementia patient care, patient status monitoring in healthcare settings, home care support, and emergency response systems. Furthermore, by changing the record weighting algorithm or event label assignment method, diverse use cases and individual optimizations can be flexibly addressed. The records department can adjust the weighting score and threshold settings according to on-site needs and individual characteristics, thereby significantly improving the flexibility and scalability of computer technology.

[0062] The anti-wandering unit can infer the individual's emotions and adjust the anti-wandering method based on the inferred emotions. For example, when the individual is relaxed, the anti-wandering unit can gently prevent wandering; when the individual is stressed, it can prevent wandering by relieving stress. Furthermore, when the individual is excited, the anti-wandering unit can prevent wandering by calming emotions. Thus, by adjusting the anti-wandering method according to the individual's emotions, appropriate wandering prevention can be achieved. Emotion inference can be achieved through emotion inference functions such as emotion engines or generative AI. Generative AI can be text-generated AI (such as LLM) or multimodal generative AI, but is not limited to these. Specifically, this anti-wandering unit acquires the individual's spoken audio data, text data, vital sign sensor data, accelerometer sensor data, GPS location information, etc. in real time and inputs them into the emotion inference module. The emotion inference module adopts a BERT-based emotion classification model or a multimodal emotion inference model that integrates speech, text, and physiological information. AI input examples include spoken text (e.g., "I'm in a good mood today," "I'm very tired"), speech features (e.g., F0=120Hz, energy=0.8), vital signs data (e.g., heart rate 80bpm, blood pressure 120 / 80), and acceleration values ​​(e.g., 0.2, 0.1, 0.3). The AI ​​outputs emotional labels (e.g., relaxed, stressed, excited) and emotional intensity scores (e.g., relaxed 0.8, stressed 0.1, excited 0.1). Based on the emotional labels and intensity scores, this prevention unit dynamically determines the parameters of the loitering prevention method (e.g., alarm intensity, notification method, and behavior restriction level). When relaxed, gentle alarms such as voice guidance or mild vibration are used to prevent loitering; when stressed, to avoid increasing stress, the notification frequency and alarm intensity are reduced, and family members or caregivers are notified first; when excited, soothing voice messages or relaxation guidance are played, and emergency notifications are made when necessary. Examples of AI outputs include prevention method labels (e.g., mild, stress-reducing, emotionally stable), alarm strength (e.g., low, medium, high), and notification recipient lists (e.g., the individual, family members, security company). These output values ​​are used for subsequent processing by the prevention unit, such as alarm sending, behavior restriction enforcement, and data forwarding to the recording department. Technically, this prevention unit dynamically optimizes loitering prevention methods based on the individual's emotional state, significantly reducing the individual's psychological burden compared to traditional blanket alarm methods, achieving high-precision loitering prevention. This ensures the individual's safety and improves QOL (Quality of Life), reducing the burden on caregivers. Applicable areas include dementia patient care, mental illness patient monitoring, home care support, and loitering countermeasures in medical care facilities. Furthermore, by changing the type and parameters of the emotional inference model or prevention method control algorithm, it can flexibly address diverse use cases and individual optimizations. The prevention unit can adjust the prevention method switching logic and alarm strength setting standards according to on-site needs and individual characteristics, thereby significantly improving the flexibility and scalability of computer technology.

[0063] The prevention department can optimize its prevention algorithm by referring to past loitering data during prevention. For example, it can adjust the algorithm based on past loitering data or analyze past loitering patterns to optimize it. Furthermore, it can improve the accuracy of loitering prevention by referring to past loitering data. Thus, by referring to past loitering data, the prevention algorithm can be optimized and the accuracy of loitering prevention can be improved. Specifically, this prevention department saves a database of a person's loitering history from the past week to the past month (e.g., loitering time, location, movement path, emotional scores before and after loitering, alarm sending history, intervention results) in chronological order. This data is input into temporal neural networks such as temporal clustering, autoregression models (AR models), and LSTM to automatically extract loitering trends and patterns (e.g., late-night outings, frequent loitering on specific weeks, deviation trends from specific paths). Based on the extracted patterns and features, this prevention department dynamically optimizes the prevention algorithm parameters (e.g., alarm threshold, notification timing, behavior restriction level, alarm intensity). For example, when late-night loitering was frequent in the past, the nighttime alarm threshold was lowered and the alarm sending frequency was increased; when deviations from specific movement paths were frequent, the alarm intensity was increased only when passing through that path. Furthermore, the prevention department combines emotional scores and vital sign data (e.g., increased heart rate, increased stress score) before and after loitering to analyze the correlation between emotional changes and loitering occurrences, optimizing the timing of early warning alerts and interventions when emotional abnormalities occur. Examples of AI inputs include loitering occurrence time series (e.g., 2024-06-01 02:15, 2024-06-03 23:40), movement path vectors (e.g., GPS coordinates), emotional score sequences (e.g., 0.8, 0.2, 0.1), alert sending history (e.g., whether an alert was sent, sending time), and intervention results (e.g., successful / failed loitering prevention). AI outputs include optimized prevention algorithm parameters, loitering risk score (e.g., 0.95), recommended alert sending timing, and behavioral restriction level. These output values ​​serve as subsequent processing by the prevention department and can be used for real-time alert sending, behavioral restriction enforcement, and data forwarding to the recording department. In terms of technical effectiveness, this prevention unit analyzes past loitering history using an AI model and dynamically optimizes the prevention algorithm, significantly improving the accuracy of loitering warning detection and the speed and flexibility of prevention responses compared to traditional fixed-parameter methods. This enables high-precision loitering prevention tailored to individual changes in state and unique characteristics. Applicable areas include dementia patient care, loitering management for mental illness patients, home care support, and security management of healthcare facilities. Furthermore, by changing the analysis algorithm or parameter optimization methods, it can flexibly address diverse use cases and individual optimizations. The prevention unit can adjust the optimization logic and feature selection criteria according to on-site needs and individual characteristics, thereby greatly enhancing the flexibility and scalability of computer technology.

[0064] The prevention unit can infer the individual's emotions and determine the priority of prevention measures based on these inferences. For example, when the individual is relaxed, the prevention unit can gently prevent loitering; when the individual is stressed, it can prevent loitering by alleviating stress. Furthermore, when the individual is excited, the prevention unit can prevent loitering by calming emotions. Thus, by prioritizing prevention measures based on the individual's emotions, appropriate loitering prevention can be achieved. Emotion inference can be achieved through emotion inference functions such as emotion engines or generative AI. Generative AI can be text-generated AI (such as LLM) or multimodal generative AI, but is not limited to these. Specifically, this prevention unit acquires the individual's spoken audio data, text data, vital sign sensor data (such as heart rate, blood pressure, body temperature), accelerometer data, GPS location information, etc., in real time and inputs them into the emotion inference module. The emotion inference module uses a BERT-based emotion classification model or a multimodal emotion inference model that integrates speech, text, and physiological information. AI input examples include spoken text (e.g., "I'm in a good mood today," "I'm very tired"), speech features (e.g., F0=120Hz, energy=0.8), vital signs data (e.g., heart rate 80bpm, blood pressure 120 / 80), and acceleration values ​​(e.g., 0.2, 0.1, 0.3). The AI ​​outputs emotional labels (e.g., relaxed, stressed, excited) and emotional intensity scores (e.g., relaxed 0.8, stressed 0.1, excited 0.1). Based on the emotional labels and intensity scores, this prevention unit dynamically determines the prevention priority score (e.g., high, medium, low) and prevention method parameters (e.g., alarm intensity, notification recipients, behavioral restriction level). When relaxed, gentle voice guidance or low-stimulation alarms such as vibration are prioritized; when stressed, to avoid increasing stress, the notification frequency and alarm intensity are reduced, and family members or caregivers are prioritized; when excited, soothing voice messages or relaxation guidance are played, and emergency notifications are made when necessary. Examples of AI outputs include prevention priority labels (e.g., high, medium, low), prevention method labels (e.g., mild, stress-reducing, emotionally stable), alarm intensity (e.g., low, medium, high), and a list of notification recipients (e.g., the individual, family members, security company). These output values ​​serve as subsequent processing by the prevention department and can be used for alarm sending, behavior restriction enforcement, and data forwarding to the recording department. Technically, this prevention department dynamically optimizes prevention priorities and methods based on the individual's emotional state, significantly reducing the individual's psychological burden compared to traditional blanket alarm methods, achieving high-precision loitering prevention. This ensures the individual's safety and improves QOL (Quality of Life), while reducing the burden on caregivers. Applicable areas include dementia patient care, mental illness patient monitoring, home care support, and loitering countermeasures in healthcare facilities. Furthermore, by changing the type and parameters of the emotional inference model or prevention priority decision algorithm, diverse use cases and individual optimizations can be flexibly addressed.The prevention department can adjust the prevention priority decision logic and alarm intensity setting standards according to on-site needs and individual characteristics, thereby greatly improving the flexibility and scalability of computer technology.

[0065] The prevention department can weight prevention data based on the timing of loitering. For example, it can prioritize recent loitering data or important past loitering data. Furthermore, it can prioritize loitering data related to specific events. Thus, by weighting prevention data based on the timing of loitering, appropriate loitering prevention can be achieved. Specifically, this prevention department obtains metadata such as timestamps (e.g., 2024-06-01 02:15:00), locations, movement paths, sentiment scores, and event tags (e.g., late at night, holidays, before / after going out) for each loitering event from a loitering history database. Using a time-series sorting algorithm or a priority queue, it weights the latest loitering data or event-related loitering data before inputting it into the prevention algorithm. Examples of AI input include loitering occurrence time (e.g., 2024-06-01 02:15:00), location (e.g., near home, in front of a station), event tags (e.g., holidays, before going out), and importance scores (e.g., 0.95, 0.2). The AI ​​outputs a prevention weighted score (e.g., 0.9, 0.3) and a prevention priority label (e.g., high, medium, low) based on the timing of the occurrence and event labels. This prevention unit sequentially performs multi-stage prevention processing, including alarm sending and behavior restrictions, starting with the lingering data with the highest weighted scores. For important past lingering events, importance is estimated using TF-IDF or BERT embedding vectors, prioritizing the re-analysis and prevention of high-importance content. For lingering data related to specific events, the weighting is dynamically adjusted by calculating the proximity of the event occurrence time to the lingering time and the relevance of the content. These output values ​​serve as subsequent processing by the prevention unit and can be used for alarm sending, behavior restriction enforcement, and the order of data forwarding to the recording unit. Technically, this prevention unit, through dynamic weighted control based on the timing of lingering occurrence and event relevance, significantly improves the rapid extraction of important information and the ability to respond to situations compared to traditional simple time-series prevention methods. This allows caregivers or medical professionals to access the necessary information in real time, enabling rapid decision-making and response. Applicable areas include dementia patient care, patient status monitoring in medical settings, home care support, and emergency response systems. Furthermore, by changing the weighting algorithm or event label assignment method, diverse use cases and individual optimizations can be flexibly addressed. The prevention department can adjust the weighting score and threshold settings according to on-site needs and individual characteristics, thereby significantly improving the flexibility and scalability of computer technology.

[0066] The system described in this embodiment is not limited to the examples above. For example, various modifications can be made as described below. Specifically, the structure and data flow of each module, such as the collection unit, analysis unit, estimation unit, prevention unit, and recording unit, can be flexibly changed according to the characteristics of the user site and the users. For example, in the collection unit, in addition to the voice sensor, image sensors and environmental sensors (such as temperature, humidity, illuminance, CO2 concentration, etc.) can be added to achieve multimodal data collection. In the analysis unit, advanced AI models such as Transformer-type natural language processing models or graph neural networks can be introduced to enhance the high-dimensional feature extraction and analysis of speech content and behavioral patterns. In the estimation unit, an integrated estimation method combining multiple estimation algorithms (such as decision trees, random forests, SVM, deep learning models) can be adopted to improve estimation accuracy and robustness. In the prevention unit, in addition to loitering prevention, fall warning detection and emergency notification functions can be added to achieve comprehensive safety management. In the recording unit, security enhancement functions such as cloud linkage, distributed recording of edge devices, data encryption and anonymization can be added. Furthermore, by dynamically optimizing the training data and parameters of the AI ​​model based on on-site needs and individual characteristics, personalized support and applicability to diverse use cases can be achieved. Technically, this system achieves significant improvements over traditional fixed-function systems in scalability, adaptability, security, accuracy, and responsiveness through flexible changes to the structure, algorithms, and data flow of each module. Therefore, it can be widely applied in various fields such as dementia patient care, mental illness patient monitoring, home care support, patient status monitoring in medical settings, life support robots, and health management applications. Moreover, by diversifying the AI ​​model and data flow, it can flexibly respond to future technological evolution and new use cases.

[0067] The collection unit can collect not only the speaker's speech but also ambient sounds. For example, it can measure the ambient noise level and temporarily stop speech collection when the noise level is high. Furthermore, it can analyze ambient sounds and prioritize speech collection when specific sounds (such as ambulance sirens or doorbells) are detected. Moreover, it can analyze ambient sound data and record sound events related to the speaker's speech. Thus, by considering ambient sounds, the accuracy of speech collection can be improved. Specifically, this collection unit uses a high-sensitivity microphone array or environmental sensors to acquire ambient acoustic data (e.g., 16kHz, 16bit, mono PCM) in real time and preprocesses it using an acoustic feature extraction module (e.g., spectrogram, MFCC, zero crossover rate, etc.). Using the acoustic features as input, the collection unit determines the noise level (e.g., dB value), sound source type (e.g., alarm, doorbell, voice, background noise, etc.), and the timing of the sound event using an ambient sound classification model (e.g., a CNN-based acoustic event detection model). Examples of AI inputs include acoustic feature vectors (e.g., MFCC 13D × time-series), ambient sound labels (e.g., alarms, doorbells), and noise levels (e.g., 75dB). AI outputs include noise level indicators (e.g., high / low), sound event detection labels (e.g., alarm detection, doorbell detection), and speech collection availability indicators (e.g., 1 or 0). The collection unit pauses speech collection when noise is determined to be high and prioritizes speech collection when a sound event is detected. Furthermore, it records the speaker's speech and concurrent sound events (e.g., alarm sounds during speech) in chronological order and forwards them to the analysis or recording unit, thereby accurately grasping the correlation between speech content and ambient sound. Technically, this collection unit uses an AI model to analyze surrounding ambient sound, dynamically optimizing the timing and priority of speech collection, significantly improving noise resistance, information comprehensiveness, and contextual responsiveness compared to traditional simple speech collection methods. Therefore, important speeches can be collected even in noisy environments or emergency situations, improving practicality in nursing and medical settings. Applicable areas include dementia patient care, home care support, patient status monitoring in medical settings, and life support robots. Furthermore, by changing the type and parameters of the acoustic event detection model or collection control algorithm, it can flexibly address diverse use cases and individual optimizations. The collection unit can adjust the collection control logic and acoustic feature selection criteria according to on-site needs and individual characteristics, thereby significantly improving the flexibility and scalability of computer technology.

[0068] The recording unit can automatically back up an individual's personal log data to the cloud. For example, it can upload log data to the cloud at regular intervals. When the user's device storage capacity is insufficient, older data can also be transferred to the cloud to ensure sufficient storage space. Furthermore, the recording unit can synchronize cloud data with other devices, enabling multiple devices to access the log data. This improves the security and accessibility of the log data. Specifically, the recording unit stores personal log data such as activity history, health data, speech content, emotional changes, and location information in a time-series database and automatically uploads it to the cloud storage service at regular intervals (e.g., 5 minutes, 1 hour, 1 day). The recording unit is equipped with a storage capacity monitoring module; when the remaining local storage capacity falls below a threshold, it automatically transfers older data (e.g., data from one month ago) to the cloud to ensure sufficient local storage space. Cloud-based log data is synchronized in real-time with multiple devices, including home terminals, medical practitioner terminals, and nursing staff terminals, through a data synchronization module, and browsing and editing permissions are controlled by the user through an access permission management module. Examples of AI inputs include lifestyle log data (e.g., steps, heart rate, speech content), recording time, remaining storage capacity, and a list of target devices for synchronization. AI outputs include backup execution flags (e.g., 1 or 0), data transfer recommendation tags (e.g., cloud transfer, local retention), and synchronization execution flags (e.g., 1 or 0). These output values ​​serve as subsequent processing for the recording department, enabling cloud uploads, data transfers, synchronization execution, and access permission settings. Technically, this recording department uses an AI model to control the automatic backup, synchronization, and storage optimization of lifestyle log data, significantly improving data security, availability, accessibility, and storage efficiency compared to traditional manual backup methods. This reduces the risk of data loss, enables real-time access from multiple locations, and facilitates information sharing in medical and nursing settings. Applicable areas include dementia patient care, home care support, health management applications, and medical data linkage platforms. Furthermore, by changing backup, synchronization algorithms, or access permission management methods, diverse use cases and individual optimizations can be flexibly addressed. The recording department can adjust backup frequency and data transfer standards according to on-site needs and individual characteristics, thereby significantly improving the flexibility and scalability of computer technology.

[0069] The prevention unit can learn an individual's behavioral patterns and detect signs of loitering. For example, it can learn an individual's usual movement paths and issue an alert when abnormal movement patterns are detected. Furthermore, it can analyze an individual's behavioral history to identify specific time periods when loitering is likely to occur. Moreover, based on an individual's behavioral patterns, it can predict high-risk loitering situations and take countermeasures in advance. Thus, by detecting signs of loitering and responding early, the individual's safety can be ensured. Specifically, this prevention unit records movement path data (e.g., time-series vectors of longitude, latitude, and altitude) obtained from location detection sensors such as GPS modules or Wi-Fi / Bluetooth beacons, accelerometer data, and vital sign data (e.g., heart rate, stress score) in a time sequence and inputs them into a temporal AI model such as a temporal convolutional neural network (TCN) or LSTM. The AI ​​learns typical movement path patterns (e.g., home-park-supermarket-home) and outputs an alert when abnormal movement patterns are detected (e.g., going out late at night, deviating from the usual path, or moving a long distance in a short time). Examples of AI inputs include movement path vectors (e.g., GPS coordinates), movement speed sequences, behavioral history (e.g., the past week), and vital sign data sequences (e.g., heart rate 80 bpm, 90 bpm, etc.). AI outputs include abnormal behavior detection flags (e.g., 1 or 0), a loitering risk score (e.g., 0.85), and recommended alarm sending times. When the loitering risk score exceeds a threshold, the prevention unit automatically executes countermeasures such as sending alarms, restricting behavior, and notifying family members or caregivers. Furthermore, through temporal clustering or autoregression models of behavioral history, it identifies specific time periods prone to loitering (e.g., 1 AM to 3 AM, afternoons on holidays) and dynamically adjusts monitoring intensity and alarm thresholds only during these periods. Technically, this prevention unit, through high-dimensional temporal data analysis and pattern learning using AI models, significantly improves the accuracy, response speed, and individual optimization of loitering warning detection compared to traditional simple location monitoring methods. This ensures the safety of the individual, reduces the burden on caregivers, and prevents loitering incidents. Applicable areas include dementia patient care, loitering countermeasures for mental illness patients, home care support, and safety management of medical and nursing facilities. Furthermore, by changing the type and parameters of behavioral pattern learning models or anomaly detection algorithms, diverse use cases and individual optimizations can be flexibly addressed. Prevention departments can adjust monitoring intensity and alarm threshold settings according to on-site needs and individual characteristics, thereby significantly improving the flexibility and scalability of computer technology.

[0070] The data collection unit can infer an individual's emotions and adjust the speech collection method based on the inferred emotions. For example, when the individual is relaxed, the collection unit can collect speech in a natural conversational manner; when the individual is stressed, it can collect speech in the form of questions. Furthermore, when the individual is excited, the collection unit can collect speech at short intervals. Thus, by adjusting the speech collection method according to the individual's emotions, more natural speech can be collected. Specifically, this collection unit acquires the individual's spoken audio data, text data, vital sign sensor data (such as heart rate, blood pressure, skin potential), accelerometer data, etc., in real time, and inputs this multidimensional data into the emotion inference module. This emotion inference module uses a BERT-based emotion classification model or a multimodal emotion inference model that integrates speech and physiological information (such as a Transformer model that simultaneously inputs speech features and vital sign data). AI input examples include spoken text (e.g., "I'm in a good mood today," "I'm very tired"), speech feature vectors (e.g., F0=120Hz, energy=0.8, spectral envelope=0.6), vital sign data (e.g., heart rate 80bpm, blood pressure 120 / 80), and acceleration values ​​(e.g., 0.2, 0.1, 0.3). AI outputs emotion labels (e.g., relaxed, stressed, excited) and emotion intensity scores (e.g., relaxed 0.8, stressed 0.1, excited 0.1). Based on the emotion labels and intensity scores, this collection unit dynamically determines the collection method parameters (e.g., conversational, question-based, shortened interval). When relaxed, it emphasizes natural dialogue flow, lowers the speech detection algorithm threshold, and prioritizes collecting spontaneous speech; when stressed, AI generates questions at appropriate times (e.g., "What's been bothering you lately?") to guide speech collection; when excited, it shortens the speech detection interval, collecting speech multiple times in a short period to obtain detailed temporal data on emotion changes. Examples of AI outputs include collection method labels (e.g., natural conversation, question-guided, shortened intervals), collection frequency (e.g., 1 minute, 30 seconds), and collection availability flags (e.g., 1 or 0). These output values ​​serve as subsequent processing for the collection unit, enabling data forwarding to the recording unit, data filtering in the input parsing unit, and parameter settings for the collection scheduler. Technically, this collection unit quantitatively evaluates the individual's emotional state in a high-dimensional vector space and dynamically optimizes the collection method, significantly reducing the individual's psychological burden compared to traditional uniform collection methods, achieving high-precision speech collection that fits the context. This improves the comprehensiveness of natural speech data, enables early detection of abnormal behavior and emotional changes, and enhances practicality in nursing and medical settings. Applicable areas include dementia patient care, mental illness patient monitoring, home care support, health management applications, and life support robots. Furthermore, by changing the type and parameters of the emotion inference model or collection method control algorithm, it can flexibly address diverse use cases and individual optimizations.The collection department can adjust the collection method switching logic and collection frequency setting standards according to on-site needs and individual characteristics, thereby greatly improving the flexibility and scalability of computer technology.

[0071] The data collection department can analyze an individual's past speaking history to select the optimal time for data collection. For example, the department can set collection times based on periods of frequent speaking in the past, or analyze the content of past statements to prioritize collecting important statements. Furthermore, the department can optimize collection methods based on past speaking patterns. Thus, by analyzing past speaking history, the optimal time for data collection can be selected. Specifically, the department stores an individual's speaking history database from the past week to month (e.g., speaking time, content, sentiment score, speaking length, and speaking frequency) in chronological order, and inputs it into temporal neural networks such as temporal clustering, autoregression models (AR models), and LSTM to automatically extract peak speaking frequency periods and speaking patterns (e.g., morning chatter, anxious nighttime speech, weekend desire to go out, etc.). Examples of AI input include past speech text sequences (e.g., "Help," "Want to go out," "The weather is nice today"), speech time sequences (e.g., 2024-06-01 10:15, 2024-06-01 12:30), and sentiment score sequences (e.g., 0.8, 0.2, 0.1). Based on the extracted patterns and features, the AI ​​outputs recommended collection timing values ​​(e.g., 8 AM - 10 AM, 8 PM - 10 PM), collection priority scores (e.g., 0.95, 0.2), and collection method labels (e.g., normal, important, simple). The collection department increases collection frequency during peak hours and employs detailed collection methods during periods of high frequency of important speeches. It also automatically switches collection methods when abnormal patterns occur, based on the importance scores and sentiment score change patterns of the speech content. Examples of AI output include optimized collection algorithm parameters, collection confidence scores (e.g., 0.95), and recommended collection timing values ​​(e.g., 8 PM - 10 PM). These output values, as subsequent processing by the collection department, can be used for data forwarding to the recording department, data filtering in the input parsing department, and parameter setting of the collection scheduler. Technically, this collection department analyzes past speaking history using an AI model to dynamically optimize collection timing and methods, significantly improving the comprehensiveness and accuracy of information compared to traditional fixed-parameter methods. This enables high-precision speech collection tailored to individual state changes and characteristics. Applicable areas include dementia patient care, mental illness patient monitoring, home care support, and patient status monitoring in medical settings. Furthermore, by changing the analysis algorithm or parameter optimization methods, it can flexibly address diverse use cases and individual optimizations. The collection department can adjust the optimization logic and feature selection criteria according to on-site needs and individual characteristics, thereby significantly improving the flexibility and scalability of computer technology.

[0072] The collection department can adjust the method of collecting comments based on the user's lifestyle and areas of interest. For example, the department can prioritize collecting comments related to the user's current interests and adjust the content of the collected comments according to the user's lifestyle. Furthermore, the department can filter and collect comments related to the user's current activities. Thus, by filtering comments based on the user's lifestyle and areas of interest, highly relevant comments can be collected. Specifically, the collection department comprehensively manages the user's lifestyle data (e.g., residential status, family structure, occupation, daily routine), area of ​​interest data (e.g., interests, health, news, travel, etc.), and current activity data (e.g., walking, eating, resting), and inputs this information as feature vectors into the AI ​​model. Examples of AI inputs include lifestyle vectors (e.g., single, elderly, at home), lists of interest topics (e.g., gardening, cooking, health), activity tags (e.g., walking, watching TV), and spoken text (e.g., "want to plant new flowers," "not feeling well lately"), etc. AI calculates the relevance between quantitative input features and spoken content using BERT embedding vectors or cosine similarity, outputting a relevance score (e.g., 0.92), priority labels (e.g., high, medium, low), and a collection availability flag (e.g., 1 or 0). The collection unit prioritizes collecting statements with high relevance scores and dynamically adjusts the collection targets and frequency based on living conditions and activities. For example, when focusing on health, it prioritizes collecting health-related statements; while traveling, it prioritizes statements related to movement and outings. The AI's output is used for subsequent processing by the collection unit, such as forwarding data to the recording unit and filtering data in the input parsing unit. Technically, this collection unit, by quantitatively evaluating living conditions, areas of interest, and activities in a high-dimensional vector space, dynamically optimizes the statement collection method, significantly improving the accuracy of information collection that matches the individual's current situation compared to traditional uniform collection methods. This suppresses the accumulation of useless data, improves the efficiency of parsing and inference processing, and enhances the overall responsiveness of the system. Applicable areas include dementia patient care, monitoring the status of medical care facilities, optimizing dialogue for life support robots, and regional collaborative health management systems. Furthermore, by modifying the living conditions data or relevance calculation algorithms, diverse use cases and individual optimizations can be flexibly addressed. The collection department can adjust the collection control logic and relevance threshold settings according to on-site needs and individual characteristics, thereby significantly improving the flexibility and scalability of computer technology.

[0073] The data collection unit can infer an individual's emotions and determine the priority order for collecting statements based on these inferred emotions. For example, when the individual is relaxed, the collection unit can prioritize collecting everyday statements; when the individual is stressed, it can prioritize collecting emotional statements. Furthermore, when the individual is excited, the collection unit can prioritize collecting important statements. Thus, by prioritizing statements based on the individual's emotions, important statements can be collected preferentially. Specifically, this collection unit acquires the individual's spoken audio data, text data, vital sign sensor data, accelerometer data, etc., in real time and inputs them into the emotion inference module. The emotion inference module uses a BERT-based emotion classification model or a multimodal emotion inference model that integrates speech and physiological information. Examples of AI input include spoken text (e.g., "I'm in a good mood today," "I'm very tired"), speech features (e.g., F0=120Hz, energy=0.8), vital sign data (e.g., heart rate 80bpm, blood pressure 120 / 80), and acceleration values ​​(e.g., 0.2, 0.1, 0.3). The AI ​​outputs emotional labels (e.g., relaxed, stressed, excited) and emotional intensity scores (e.g., relaxed 0.8, stressed 0.1, excited 0.1). Based on these emotional labels and intensity scores, the collection unit dynamically determines the priority score (e.g., high, medium, low) and the label of the collection object (e.g., daily, emotional, important). When relaxed, daily statements (e.g., "The weather is nice today") are collected first; when stressed, emotional statements (e.g., "I can't take it anymore," "Help!") are collected first; when excited, important statements (e.g., "I want to go out," "I don't know where I am") are collected first. The AI's output is used for subsequent processing by the collection unit, including forwarding data to the recording unit, filtering data in the input parsing unit, and setting parameters for the collection scheduler. Technically, this collection unit quantitatively evaluates an individual's emotional state in a high-dimensional vector space and dynamically optimizes statement priority, significantly improving the comprehensiveness of important information and system responsiveness compared to traditional uniform collection methods. This provides rapid and accurate information support for nursing staff and medical professionals. Applicable areas include dementia patient care, mental illness patient monitoring, home care support, and patient status monitoring in the medical field. Furthermore, by changing the type and parameters of the sentiment inference model or priority decision-making algorithm, diverse use cases and individual optimizations can be flexibly addressed. The collection department can adjust the priority decision-making logic and threshold settings according to on-site needs and individual characteristics, thereby significantly improving the flexibility and scalability of computer technology.

[0074] The collection unit can adjust its speech collection method based on the user's geographic location information. For example, when the user is in a specific location, the collection unit can prioritize collecting speech related to that location; when the user is moving, it can also prioritize collecting speech related to the destination. Furthermore, when the user is in a specific area, the collection unit can also prioritize collecting speech related to that area. Thus, by considering geographic location information, it can prioritize collecting highly relevant speech. Specifically, this collection unit acquires geographic location information (e.g., time-series vectors of longitude, latitude, and altitude), movement path data, and current location tags (e.g., home, hospital, park, supermarket, etc.) in real time through location detection sensors such as GPS modules or Wi-Fi / Bluetooth beacons, and inputs this information as feature vectors into the AI ​​model. Examples of AI inputs include current location tags (e.g., home, hospital), movement path vectors (e.g., GPS coordinates), and spoken text (e.g., "want to go out," "want to go to the hospital"), etc. AI calculates the relevance between quantitative geographic location information and spoken content using BERT embedding vectors or cosine similarity, outputting a relevance score (e.g., 0.92), priority labels (e.g., high, medium, low), and collection availability flags (e.g., 1 or 0). The collection unit prioritizes collecting statements with high relevance scores, increasing the collection frequency of relevant statements in specific locations or while moving. For example, when in a hospital, it prioritizes collecting statements related to health and treatment; when out in parks, supermarkets, or other locations, it prioritizes collecting statements related to movement and shopping. The AI's output is used for subsequent processing by the collection unit, such as forwarding data to the recording unit or filtering data input to the parsing unit. Technically, this collection unit significantly improves the accuracy of information collection by quantitatively evaluating the relevance between geographic location information and spoken content in a high-dimensional vector space, compared to traditional indiscriminate collection methods. This suppresses the accumulation of useless data, improves the efficiency of parsing and inference processing, and enhances the overall responsiveness of the system. Applicable areas include caregiving for dementia patients while they are out, monitoring the status of medical and nursing facilities, optimizing dialogue for life support robots, and regional collaborative health management systems. Furthermore, by changing the location information acquisition interval or relevance calculation algorithm, diverse use cases such as urban areas, suburbs, and indoor / outdoor environments can be flexibly addressed. The collection department can adjust the location information linkage logic and relevance threshold settings according to on-site needs and individual characteristics, thereby significantly improving the flexibility and scalability of computer technology.

[0075] The collection department can analyze an individual's social media activity and collect relevant posts. For example, it can collect posts related to topics the individual frequently discusses on social media. Furthermore, the department can schedule collection based on the time periods of the individual's social media activity. Further, the department can analyze the content of the individual's posts on social media, prioritizing the collection of important posts. Specifically, the collection department regularly obtains post history data (such as post time, post content, and number of interactions) from multiple social media platforms used by the individual (such as Weibo, SNS, forums, etc.) via API. The department uses natural language embedding models such as BERT or RoBERTa to vectorize the post content and calculates a relevance score by measuring the cosine similarity between the individual's posts and the vectors of social media post topics. The department clusters frequently posted topics (such as health, interests, family, news, etc.) and prioritizes collecting and recording posts with high relevance to these topics. Regarding activity time period analysis, a histogram is created from the submission time data of the past month to automatically extract peak periods (e.g., 20:00–22:00) and reflect this in the speech collection scheduling. For determining important statements, TF-IDF scores or BERT embedding vectors are used for importance estimation, prioritizing the collection of statements similar to those that have garnered high engagement (e.g., likes, comments) on social media. AI input examples include social media submission text (e.g., “Started a new hobby today”, “Haven’t been feeling well lately”), submission time (e.g., 2024-06-01 20:15), and statement text (e.g., “Want to try a new hobby”). AI output includes relevance scores (e.g., 0.92), priority labels (e.g., high, medium, low), and recommended collection time values ​​(e.g., 20:00–22:00). These output values ​​are used for subsequent processing by the collection department, including data transmission to the recording department, input data filtering to the parsing department, and parameter setting for the collection scheduler. In terms of technical effectiveness, this data collection unit can quantitatively assess the relevance of an individual's social media activities and posted content in a high-dimensional vector space. Compared to traditional data collection methods that rely solely on posted content, this achieves high-precision information collection that better aligns with an individual's interests, concerns, and social activities. This allows for early understanding of changes in an individual's social connections and psychological state, leading to improved quality of care and medical support, and personalized health management. Applicable areas include preventing social isolation in dementia patients, monitoring mental illness patients, home care support, and personalized health management applications. Furthermore, by modifying the social media API type and relevance calculation algorithm, it possesses scalability to flexibly adapt to various platforms and use cases. The data collection unit can adjust the collection control logic and relevance threshold settings according to on-site needs and individual characteristics, significantly enhancing the flexibility and scalability of computer technology.

[0076] The analysis unit can infer the user's emotions and adjust the analysis presentation based on the inferred emotions. For example, when the user is relaxed, the analysis unit can provide detailed analysis results; when the user is stressed, it can provide concise analysis results; and when the user is excited, it can provide analysis results in a visually easy-to-understand way. By adjusting the analysis presentation according to the user's emotions, appropriate analysis results can be provided. Emotion inference can be achieved through emotion inference functions such as emotion engines or generative AI. Generative AI can be text generation AI (such as large-scale language models) or multimodal generation AI, but is not limited to these. Specifically, this analysis unit acquires the user's spoken speech or text data in real time and inputs it into the emotion inference module. The emotion inference module uses a BERT-based emotion classification model or a multimodal emotion inference model, inputting spoken text (such as "I'm very happy today" or "I'm very tired") or speech features (such as F0, energy, spectral envelope), and outputting emotion labels (such as relaxed, stressed, excited) and emotion intensity scores (such as relaxed 0.8, stressed 0.2). This analysis unit dynamically switches the presentation of analysis results based on sentiment tags and intensity scores. When relaxed, it generates a detailed report containing comprehensive semantic analysis of the spoken content (e.g., topic extraction, intent inference, sentiment change maps, etc.); when stressed, it presents a concise summary extracting only key points (e.g., topic tags, sentiment tags, list of important statements); when excited, it outputs analysis results with a visually enhanced interface (e.g., color-coded charts, icon displays, sentiment change animations, etc.). Examples of AI inputs include spoken text (e.g., "I'm in a good mood today," "I hope to get help"), speech feature vectors (e.g., F0=120Hz, energy=0.8), and sentiment intensity scores (e.g., 0.7, 0.2, 0.1). AI outputs include analysis presentation type (e.g., detailed, concise, visual), and analysis result data (e.g., topic category, sentiment change map, summary text). These output values ​​are used for subsequent processing by the analysis unit, for selecting the display format of the user interface, or for data transmission to the recording unit. In terms of technical effectiveness, this analysis unit can optimize the presentation of analysis results based on the individual's emotional state. Compared to traditional uniform output methods, it provides more easily understood and realistic information for both the individual and their caregiver. This improves information delivery efficiency, alleviates stress, and supports rapid decision-making. Applicable areas include care for dementia patients, monitoring of mental illness patients, home care support, and patient status reporting in medical settings. Furthermore, by modifying the analysis presentation algorithm or interface design, it can flexibly handle diverse use cases and individual optimizations. The analysis unit can adjust the analysis presentation control logic and output format selection criteria according to on-site needs and individual characteristics, significantly improving the flexibility and scalability of computer technology.

[0077] The following is a brief description of the implementation process. Specifically, this system consists of multiple modules, including a collection unit, an analysis unit, an inference unit, a prevention unit, and a recording unit. Through data flow collaboration between these modules, it achieves integrated execution of real-time collection, analysis, inference, recording, and prevention processing of the user's speech data and biometric information. The collection unit acquires speech, text, vital sign data, and location information from various sensors such as voice sensors, vital sign sensors, accelerometers, and GPS, and transmits the data to the sentiment inference module or the speech content analysis module. The analysis unit utilizes BERT, Transformer-type natural language processing models, and multimodal sentiment inference models to perform multi-stage analysis of the speech content, including semantic analysis, sentiment inference, topic extraction, intent inference, and sentiment change graph generation. The inference unit takes the output of the analysis unit (such as speech text, sentiment tags, sentiment intensity scores, topic categories, etc.) as input and uses inference algorithms such as multilayer perceptrons, decision trees, and rule-based inferencers to infer the user's intentions (such as willingness to go out, request for help, willingness to rest, etc.). The prevention unit, based on information from the presumption unit or the recording unit, performs security management procedures such as preventing loitering, detecting abnormal behavior precursors, sending alarms, and restricting behavior. The recording unit records and backs up data from all modules to a time-series database or cloud storage, and synchronizes data with family members' and medical personnel's terminals as needed. Each module can dynamically optimize AI model parameters and control algorithms according to on-site needs and individual characteristics, thus enabling personalized support and diverse use cases. Technically, this system, through inter-module data collaboration and dynamic optimization of AI models, achieves significant improvements in scalability, adaptability, accuracy, responsiveness, and security compared to traditional fixed-function systems. Therefore, it can be widely applied in areas such as dementia patient care, mental illness patient monitoring, home care support, patient status monitoring in medical settings, life assistance robots, and health management applications. Furthermore, by increasing the diversity of AI models or data streams, it can flexibly respond to future technological evolution and new use cases.

[0078] Step 1: The collection unit collects the speaker's speech in real time. For example, the collection unit can collect the speaker's speech as voice data, text data, or gesture data. Step 2: The parsing unit analyzes the speech data collected by the collection unit. For example, the parsing unit can use speech recognition technology to convert voice data into text data, or it can use text parsing technology to analyze text data. Furthermore, it can use sentiment analysis technology to analyze the sentiment of the speech data. Step 3: The inference unit infers the speaker's intentions based on the information analyzed by the parsing unit. For example, the inference unit can infer the speaker's decision-making, emotional state, or behavioral intentions based on the parsing results. Specifically, in Step 1, this system acquires various data such as speech, gestures, and heart rate in real time from voice sensors, accelerometers, and vital sign sensors. The voice data is preprocessed using acoustic feature extraction modules such as spectrograms or MFCC. The collection unit inputs these feature vectors or text data into an AI model for speech detection or speech interval extraction. In step 2, the parsing unit transcribes the speech data into text using a speech recognition model (such as Transformer-type ASR) and utilizes natural language processing models such as BERT for topic extraction, intent inference, and sentiment inference (e.g., relaxation, stress, excitement). Examples of AI inputs include speech feature vectors (e.g., MFCC 13D × temporal), spoken text (e.g., "want to go out," "help me"), and vital sign data (e.g., heart rate 80 bpm). AI outputs include speech interval labels, topic categories, sentiment labels, and sentiment intensity scores. In step 3, the inference unit uses the outputs from the parsing unit (e.g., topic categories, sentiment labels, sentiment intensity scores) as input and employs inference algorithms such as multilayer perceptrons and decision trees to infer the individual's intentions (e.g., willingness to go out, request for help, willingness to rest). Examples of AI outputs include intention labels (e.g., willingness to go out, seeking help), inference confidence scores (e.g., 0.92), and inference detail labels (e.g., detailed, concise, urgent). These output values ​​are used for subsequent processing, such as sending alarms or restricting behavior in the prevention unit and recording data transmission in the recording unit. In terms of technical effectiveness, this system, through high-dimensional data analysis and dynamic optimization of AI models at each step, significantly improves information coverage, accuracy, responsiveness, and security compared to traditional methods of single-step speech collection, analysis, and inference. Therefore, it can be widely applied in fields such as dementia patient care, mental illness patient monitoring, home care support, patient status monitoring in medical settings, and life assistance robots. Furthermore, by increasing the diversity of AI models or data streams, it can flexibly respond to future technological evolution and new use cases.

[0079] The specific processing unit 290 sends the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires voice representing the user's input to the result of the specific processing. The control unit 46A sends the voice data representing the user's input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the voice data.

[0080] Data generation model 58 is what is known as generative AI (Artificial Intelligence). An example of data generation model 58 includes ChatGPT (registered trademark) (Internet search).<URL: https: / / openai.com / blog / chatgpt> Generative AI, such as data generation model 58, is obtained by deep learning through a neural network. The data generation model 58 is input with a prompt containing instructions, and with inference data such as speech data representing speech, text data representing text, and image data representing images (e.g., data of still images or data of moving images). The data generation model 58 infers the inference data based on the instructions shown in the prompt and outputs the inference result in one or more data forms such as speech data, text data, and image data. The data generation model 58 includes, for example, text generation AI, image generation AI, and multimodal generation AI. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization. The specific processing unit 290 performs the above-described specific processing while using the data generation model 58. The data generation model 58 can also be a fine-tuned model to output inference results from prompts without instructions; in this case, the data generation model 58 can output inference results from prompts without instructions. The data processing device 12, etc., includes multiple data generation models 58, including AI other than generative AI. AI other than generative AI includes, but is not limited to, linear regression, logistic regression, decision trees, random forests, support vector machines (SVM), k-means clustering, convolutional neural networks (CNN), recurrent neural networks (RNN), generative adversarial networks (GAN), or Naive Bayes. Furthermore, AI can also act as an AI agent. Additionally, when the processing described above is performed by AI, this processing may be partially or entirely performed by AI, but is not limited to this example. Moreover, processing performed by AI, including generative AI, can be replaced by rule-based processing, and vice versa.

[0081] Furthermore, the processing performed by the aforementioned data processing system 10 is executed by the specific processing unit 290 of the data processing device 12 or the control unit 46A of the smart device 14, but it can also be executed jointly by the specific processing unit 290 of the data processing device 12 and the control unit 46A of the smart device 14. Additionally, the specific processing unit 290 of the data processing device 12 acquires or collects information required for processing from the smart device 14 or external devices, and the smart device 14 acquires or collects information required for processing from the data processing device 12 or external devices.

[0082] Each of the multiple elements, such as the collection unit, analysis unit, estimation unit, recording unit, and prevention unit, can be implemented by at least one of the smart device 14 or the data processing device 12. For example, the collection unit can be implemented by the receiving device 38 of the smart device 14, and the collected speech data can be displayed on the display 40. The analysis unit analyzes the speech data displayed on the display 40, and the estimation unit estimates the individual's intention based on the analyzed information. The recording unit records the individual's activity history and health data through the receiving device 38 of the smart device 14. The prevention unit is implemented through the receiving device 38 of the smart device 14, realizing location tracking and alarm communication functions. The correspondence between each unit and the device or control unit is not limited to the above examples and can be modified in various ways.

[0083] [Second Implementation] Figure 3 An example of the configuration of the data processing system 210 according to the second embodiment is shown.

[0084] like Figure 3 As shown, the data processing system 210 includes a data processing device 12 and smart glasses 214. One example of the data processing device 12 is a server.

[0085] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 includes a processor 28, RAM 30, and memory 32. The processor 28, RAM 30, and memory 32 are connected to a bus 34. Furthermore, the database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. An example of the network 54 includes a WAN and / or LAN, etc.

[0086] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, RAM 48, and memory 50. The processor 46, RAM 48, and memory 50 are connected to a bus 52. Furthermore, the microphone 238, speaker 240, and camera 42 are also connected to the bus 52.

[0087] Microphone 238 receives user commands by receiving the user's voice. Microphone 238 captures the user's voice, converts the captured sound into speech data, and outputs it to processor 46. Speaker 240 outputs sound according to commands from processor 46.

[0088] Camera 42 is a small digital camera equipped with an optical system such as a lens, aperture and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, used to capture the user's surroundings (e.g., the shooting range defined by an angle of view equivalent to the field of vision of an average healthy person).

[0089] Communication I / F 44 is connected to network 54. Communication I / F 44 and 26 are responsible for the transmission and reception of various information between processor 46 and processor 28 via network 54. The transmission and reception of various information between processor 46 and processor 28 using communication I / F 44 and 26 is performed in a secure state.

[0090] Figure 4 An example of the main functions of the data processing device 12 and the smart glasses 214 is shown. Figure 4 As shown, in the data processing device 12, specific processing is performed by the processor 28. The memory 32 stores a specific processing program 56.

[0091] The processor 28 reads a specific processing program 56 from the memory 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is implemented by the processor 28 operating as a specific processing unit 290 based on the specific processing program 56 executed on the RAM 30.

[0092] The memory 32 stores a data generation model 58 and an emotion-specific model 59. The data generation model 58 and the emotion-specific model 59 are used by the specific processing unit 290. The specific processing unit 290 can use the emotion-specific model 59 to infer the user's emotions and perform specific processing based on the user's emotions. The emotion inference function using the emotion-specific model 59 (emotion-specific function) includes various inferences and predictions related to the user's emotions, but is not limited to this example. Furthermore, emotion inference and prediction also include, for example, emotion analysis (analysis).

[0093] In the smart glasses 214, specific processing is performed by the processor 46. A specific processing program 60 is stored in the memory 50. The processor 46 reads the specific processing program 60 from the memory 50 and executes the read specific processing program 60 on the RAM 48. Specific processing is implemented by the processor 46 operating as a control unit 46A based on the specific processing program 60 executed on the RAM 48. Furthermore, the smart glasses 214 may also have the same data generation model and emotion-specific model as the data generation model 58 and the emotion-specific model 59, and use these models to perform the same processing as the specific processing unit 290.

[0094] Furthermore, other devices besides the data processing device 12 may also have the data generation model 58. For example, a server device may have the data generation model 58. In this case, the data processing device 12 obtains the processing results (prediction results, etc.) using the data generation model 58 by communicating with the server device that has the data generation model 58. In addition, the data processing device 12 may be a server device or a user-owned terminal device (e.g., a mobile phone, robot, home appliance, etc.).

[0095] The specific processing unit 290 sends the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires voice input representing the user's input to the specific processing result. The control unit 46A sends the voice data representing the user's input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the voice data.

[0096] The data generation model 58 is a so-called generative AI. An example of the data generation model 58 includes generative AIs such as ChatGPT. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with a prompt containing instructions, and also with inference data such as speech data representing speech, text data representing text, and image data representing images (e.g., data of still images or data of moving images). The data generation model 58 infers the inference data based on the instructions shown in the prompt and outputs the inference result in one or more data forms such as speech data, text data, and image data. The data generation model 58 includes, for example, text generation AI, image generation AI, and multimodal generation AI. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization. The specific processing unit 290 performs the aforementioned specific processing while using the data generation model 58. The data generation model 58 can also be a fine-tuned model to output inference results from prompts that do not contain instructions; in this case, the data generation model 58 is able to output inference results from prompts that do not contain instructions. The data processing device 12, etc., includes various data generation models 58, which include AI other than generative AI. These AIs include, for example, linear regression, logistic regression, decision trees, random forests, support vector machines (SVM), k-means clustering, convolutional neural networks (CNN), recurrent neural networks (RNN), generative adversarial networks (GAN), or Naive Bayes, and are capable of various processing methods, but are not limited to these examples. Furthermore, AI can also be an AI agent. Additionally, when the processing described above is performed by AI, this processing may be partially or entirely performed by AI, but is not limited to these examples. Moreover, processing performed by AI including generative AI can be replaced by rule-based processing, and vice versa.

[0097] The data processing system 210 of the second embodiment performs the same processing as the data processing system 10 of the first embodiment. The processing performed by the data processing system 210 is executed by the specific processing unit 290 of the data processing device 12 or the control unit 46A of the smart glasses 214, but it can also be executed jointly by the specific processing unit 290 of the data processing device 12 and the control unit 46A of the smart glasses 214. Furthermore, the specific processing unit 290 of the data processing device 12 acquires or collects information required for processing from the smart glasses 214 or external devices, and the smart glasses 214 acquires or collects information required for processing from the data processing device 12 or external devices.

[0098] Each of the multiple elements, such as the collection unit, analysis unit, estimation unit, recording unit, and prevention unit, can be implemented, for example, by the smart glasses 214. For instance, the collection unit can be implemented through the receiving device 38 of the smart glasses 214, and the collected speech data can be displayed on the display 40. The analysis unit analyzes the speech data displayed on the display 40, and the estimation unit infers the individual's intentions based on the analyzed information. The recording unit records the individual's activity history and health data through the receiving device 38 of the smart glasses 214. The prevention unit, implemented through the receiving device 38 of the smart glasses 214, provides location tracking and alarm communication functions. The correspondence between each unit and the device or control unit is not limited to the above example and can be modified in various ways.

[0099] [Third Implementation] Figure 5 An example of the configuration of the data processing system 310 according to the third embodiment is shown.

[0100] like Figure 5 As shown, the data processing system 310 includes a data processing device 12 and a head-mounted terminal 314. An example of the data processing device 12 is a server.

[0101] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 includes a processor 28, RAM 30, and memory 32. The processor 28, RAM 30, and memory 32 are connected to a bus 34. Furthermore, the database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. An example of the network 54 includes a WAN and / or LAN, etc.

[0102] The head-mounted terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a display 343. The computer 36 includes a processor 46, RAM 48, and memory 50. The processor 46, RAM 48, and memory 50 are connected to a bus 52. Furthermore, the microphone 238, speaker 240, camera 42, and display 343 are also connected to the bus 52.

[0103] Microphone 238 receives user commands by receiving the user's voice. Microphone 238 captures the user's voice, converts the captured sound into speech data, and outputs it to processor 46. Speaker 240 outputs sound according to commands from processor 46.

[0104] Camera 42 is a small digital camera equipped with an optical system such as a lens, aperture and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, used to capture the user's surroundings (e.g., the shooting range defined by an angle of view equivalent to the field of vision of an average healthy person).

[0105] Communication I / F 44 is connected to network 54. Communication I / F 44 and 26 are responsible for the transmission and reception of various information between processor 46 and processor 28 via network 54. The transmission and reception of various information between processor 46 and processor 28 using communication I / F 44 and 26 is performed in a secure state.

[0106] Figure 6 An example of the main functions of the data processing device 12 and the head-mounted terminal 314 is shown. Figure 6 As shown, in the data processing device 12, specific processing is performed by the processor 28. The memory 32 stores a specific processing program 56.

[0107] The processor 28 reads a specific processing program 56 from the memory 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is implemented by the processor 28 operating as a specific processing unit 290 based on the specific processing program 56 executed on the RAM 30.

[0108] The memory 32 stores a data generation model 58 and an emotion-specific model 59. The data generation model 58 and the emotion-specific model 59 are used by the specific processing unit 290. The specific processing unit 290 can use the emotion-specific model 59 to infer the user's emotions and perform specific processing based on the user's emotions. The emotion inference function using the emotion-specific model 59 (emotion-specific function) includes various inferences and predictions related to the user's emotions, but is not limited to this example. Furthermore, emotion inference and prediction also include, for example, emotion analysis (analysis).

[0109] In the head-mounted terminal 314, specific processing is performed by the processor 46. A specific program 60 is stored in the memory 50. The processor 46 reads the specific program 60 from the memory 50 and executes the read specific program 60 on the RAM 48. Specific processing is implemented by the processor 46 operating as a control unit 46A based on the specific program 60 executed on the RAM 48. Furthermore, the head-mounted terminal 314 may also have the same data generation model and emotion-specific model as the data generation model 58 and the emotion-specific model 59, and use these models to perform the same processing as the specific processing unit 290.

[0110] Furthermore, other devices besides the data processing device 12 may also have the data generation model 58. For example, a server device may have the data generation model 58. In this case, the data processing device 12 obtains the processing results (prediction results, etc.) using the data generation model 58 by communicating with the server device that has the data generation model 58. In addition, the data processing device 12 may be a server device or a user-owned terminal device (e.g., a mobile phone, robot, home appliance, etc.).

[0111] The specific processing unit 290 sends the result of the specific processing to the head-mounted terminal 314. In the head-mounted terminal 314, the control unit 46A causes the speaker 240 and display 343 to output the result of the specific processing. The microphone 238 acquires voice representing the user's input to the specific processing result. The control unit 46A sends the voice data representing the user's input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the voice data.

[0112] The data generation model 58 is a so-called generative AI. An example of the data generation model 58 includes generative AIs such as ChatGPT. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with a prompt containing instructions, and also with inference data such as speech data representing speech, text data representing text, and image data representing images (e.g., data of still images or data of moving images). The data generation model 58 infers the inference data based on the instructions shown in the prompt and outputs the inference result in one or more data forms such as speech data, text data, and image data. The data generation model 58 includes, for example, text generation AI, image generation AI, and multimodal generation AI. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization. The specific processing unit 290 performs the aforementioned specific processing while using the data generation model 58. The data generation model 58 can also be a fine-tuned model to output inference results from prompts that do not contain instructions; in this case, the data generation model 58 is able to output inference results from prompts that do not contain instructions. The data processing device 12, etc., includes various data generation models 58, which include AI other than generative AI. These AIs include, for example, linear regression, logistic regression, decision trees, random forests, support vector machines (SVM), k-means clustering, convolutional neural networks (CNN), recurrent neural networks (RNN), generative adversarial networks (GAN), or Naive Bayes, and are capable of various processing methods, but are not limited to these examples. Furthermore, AI can also be an AI agent. Additionally, when the processing described above is performed by AI, this processing may be partially or entirely performed by AI, but is not limited to these examples. Moreover, processing performed by AI including generative AI can be replaced by rule-based processing, and vice versa.

[0113] The data processing system 310 of the third embodiment performs the same processing as the data processing system 10 of the first embodiment. The processing performed by the data processing system 310 is executed by the specific processing unit 290 of the data processing device 12 or the control unit 46A of the head-mounted terminal 314, but it can also be executed jointly by the specific processing unit 290 of the data processing device 12 and the control unit 46A of the head-mounted terminal 314. Furthermore, the specific processing unit 290 of the data processing device 12 acquires or collects information required for processing from the head-mounted terminal 314 or external devices, and the head-mounted terminal 314 acquires or collects information required for processing from the data processing device 12 or external devices.

[0114] Each of the multiple elements, such as the collection unit, analysis unit, estimation unit, recording unit, and prevention unit, can be implemented, for example, by a head-mounted device 314. For instance, the collection unit can be implemented through the receiving device 38 of the head-mounted device 314, and the collected speech data can be displayed on the display 40. The analysis unit analyzes the speech data displayed on the display 40, and the estimation unit infers the individual's intentions based on the analyzed information. The recording unit records the individual's activity history and health data through the receiving device 38 of the head-mounted device 314. The prevention unit, implemented through the receiving device 38 of the head-mounted device 314, implements location tracking and alarm communication functions. The correspondence between each unit and the device or control unit is not limited to the above examples and can be modified in various ways.

[0115] [Fourth Implementation] Figure 7 An example of the configuration of the data processing system 410 according to the fourth embodiment is shown.

[0116] like Figure 7 As shown, the data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.

[0117] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 includes a processor 28, RAM 30, and memory 32. The processor 28, RAM 30, and memory 32 are connected to a bus 34. Furthermore, the database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. An example of the network 54 includes a WAN and / or LAN, etc.

[0118] Robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a control object 443. The computer 36 includes a processor 46, RAM 48, and memory 50. The processor 46, RAM 48, and memory 50 are connected to a bus 52. Furthermore, the microphone 238, speaker 240, camera 42, and control object 443 are also connected to the bus 52.

[0119] Microphone 238 receives user commands by receiving the user's voice. Microphone 238 captures the user's voice, converts the captured sound into speech data, and outputs it to processor 46. Speaker 240 outputs sound according to commands from processor 46.

[0120] The camera 42 is a small digital camera equipped with an optical system such as a lens, aperture and shutter, and an imaging element such as a CMOS image sensor or a CCD image sensor, used to photograph the user's surroundings (e.g., the shooting range defined by an angle of view equivalent to the field of vision of an average healthy person).

[0121] Communication I / F 44 is connected to network 54. Communication I / F 44 and 26 are responsible for the transmission and reception of various information between processor 46 and processor 28 via network 54. The transmission and reception of various information between processor 46 and processor 28 using communication I / F 44 and 26 is performed in a secure state.

[0122] The controlled object 443 includes a display device, LEDs for the eyes, and motors for driving the arms, hands, and feet. The posture and movements of the robot 414 are controlled by controlling the motors for the arms, hands, and feet. Some of the robot 414's emotions can be expressed by controlling these motors. In addition, facial expressions of the robot 414 can also be expressed by controlling the illumination state of the LEDs for the robot 414's eyes.

[0123] Figure 8 An example of the main functions of the data processing device 12 and the robot 414 is shown. For example... Figure 8 As shown, in the data processing device 12, specific processing is performed by the processor 28. The memory 32 stores a specific processing program 56.

[0124] The processor 28 reads a specific processing program 56 from the memory 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is implemented by the processor 28 operating as a specific processing unit 290 based on the specific processing program 56 executed on the RAM 30.

[0125] The memory 32 stores a data generation model 58 and an emotion-specific model 59. The data generation model 58 and the emotion-specific model 59 are used by the specific processing unit 290. The specific processing unit 290 can use the emotion-specific model 59 to infer the user's emotions and perform specific processing based on the user's emotions. The emotion inference function using the emotion-specific model 59 (emotion-specific function) includes various inferences and predictions related to the user's emotions, but is not limited to this example. Furthermore, emotion inference and prediction also include, for example, emotion analysis (analysis).

[0126] In robot 414, specific processing is performed by processor 46. A specific program 60 is stored in memory 50. Processor 46 reads the specific program 60 from memory 50 and executes the read specific program 60 on RAM 48. Specific processing is achieved by processor 46 acting as control unit 46A based on the specific program 60 executed on RAM 48. Furthermore, robot 414 may also have the same data generation model and emotion-specific model as data generation model 58 and emotion-specific model 59, and use these models to perform the same processing as specific processing unit 290.

[0127] Furthermore, other devices besides the data processing device 12 may also have the data generation model 58. For example, a server device may have the data generation model 58. In this case, the data processing device 12 obtains the processing results (prediction results, etc.) using the data generation model 58 by communicating with the server device that has the data generation model 58. In addition, the data processing device 12 may be a server device or a user-owned terminal device (e.g., a mobile phone, robot, home appliance, etc.).

[0128] The specific processing unit 290 sends the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the controlled object 443 to output the result of the specific processing. The microphone 238 acquires voice representing the user's input regarding the result of the specific processing. The control unit 46A sends the voice data representing the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the voice data.

[0129] The data generation model 58 is a so-called generative AI. An example of the data generation model 58 includes generative AIs such as ChatGPT. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with a prompt containing instructions, and also with inference data such as speech data representing speech, text data representing text, and image data representing images (e.g., data of still images or data of moving images). The data generation model 58 infers the inference data based on the instructions shown in the prompt and outputs the inference result in one or more data forms such as speech data, text data, and image data. The data generation model 58 includes, for example, text generation AI, image generation AI, and multimodal generation AI. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization. The specific processing unit 290 performs the aforementioned specific processing while using the data generation model 58. The data generation model 58 can also be a fine-tuned model to output inference results from prompts that do not contain instructions; in this case, the data generation model 58 is able to output inference results from prompts that do not contain instructions. The data processing device 12, etc., includes various data generation models 58, which include AI other than generative AI. These AIs include, for example, linear regression, logistic regression, decision trees, random forests, support vector machines (SVM), k-means clustering, convolutional neural networks (CNN), recurrent neural networks (RNN), generative adversarial networks (GAN), or Naive Bayes, and are capable of various processing methods, but are not limited to these examples. Furthermore, AI can also be an AI agent. Additionally, when the processing described above is performed by AI, this processing may be partially or entirely performed by AI, but is not limited to these examples. Moreover, processing performed by AI including generative AI can be replaced by rule-based processing, and vice versa.

[0130] The data processing system 410 of the fourth embodiment performs the same processing as the data processing system 10 of the first embodiment. The processing performed by the data processing system 410 is executed by the specific processing unit 290 of the data processing device 12 or the control unit 46A of the robot 414, but it can also be executed jointly by the specific processing unit 290 of the data processing device 12 and the control unit 46A of the robot 414. Furthermore, the specific processing unit 290 of the data processing device 12 acquires or collects information required for processing from the robot 414 or external devices, and the robot 414 acquires or collects information required for processing from the data processing device 12 or external devices.

[0131] Each of the various elements, such as the collection unit, analysis unit, estimation unit, recording unit, and prevention unit, can be implemented by, for example, the robot 414. For instance, the collection unit can be implemented through the receiving device 38 of the robot 414, and the collected speech data can be displayed on the display 40. The analysis unit analyzes the speech data displayed on the display 40, and the estimation unit infers the speaker's intentions based on the analyzed information. The recording unit records the speaker's activity history and health data through the receiving device 38 of the robot 414. The prevention unit, implemented through the receiving device 38 of the robot 414, provides location tracking and alarm communication functions. The correspondence between each unit and the device or control unit is not limited to the above example and can be modified in various ways.

[0132] Furthermore, the emotion-specific model 59, serving as an emotion engine, can determine the user's emotion based on a specific mapping. Specifically, the emotion-specific model 59 can determine the user's emotion based on an emotion graph that serves as a specific mapping (see...). Figure 9 The robot's emotions can be determined by the emotion-specific model 59. In addition, the emotion-specific model 59 can also determine the robot's emotions in the same way, and the specific processing unit 290 can also perform specific processing using the robot's emotions.

[0133] Figure 9 This is a diagram representing an emotion map 400 that maps various emotions. In the emotion map 400, emotions are arranged radially from the center in concentric circles. The closer to the center of the concentric circles, the more primitive the emotion is. Further out on the concentric circles, emotions are arranged representing states or actions arising from mood. Emotion is a concept that includes both feelings and mental states. To the left of the concentric circles, emotions generated by reactions occurring in the brain are arranged roughly. To the right of the concentric circles, emotions guided by situational judgments are arranged roughly. Above and below the concentric circles, emotions generated by reactions occurring in the brain and guided by situational judgments are arranged roughly. Furthermore, the emotion of "pleasure" is arranged above the concentric circles, and the emotion of "unpleasantness" is arranged below. Thus, in the emotion map 400, various emotions are mapped according to the structure of emotion generation, while easily generated emotions are mapped nearby.

[0134] These emotions are distributed at the 3 o'clock position on the Emotion Chart 400, and usually fluctuate between peace and unease. In the right half of the Emotion Chart 400, because situational awareness is more dominant than internal feelings, it gives a sense of calm.

[0135] The inner side of the emotion diagram 400 represents the mind, and the outer side of the emotion diagram 400 represents actions. Therefore, the further you go to the outer side of the emotion diagram 400, the more the emotion can be seen (manifested in actions).

[0136] Here, human emotions are based on a balance of various factors such as posture and blood sugar levels. When these balances deviate from the ideal, it indicates unhappiness; when they approach the ideal, it indicates pleasure. In robots, cars, and motorcycles, emotions can also be created based on a balance of factors such as posture and remaining battery power. When these balances deviate from the ideal, it indicates unhappiness; when they approach the ideal, it indicates pleasure. Emotion maps can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on Speech Emotion Recognition and Brain Physiological Signal Analysis Systems for Emotion, Tokushima University, Doctoral Dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). In the left half of the emotion map, emotions belonging to the "Reaction" domain, where sensation is dominant, are arranged. Furthermore, in the right half of the emotion map, emotions belonging to the "Situation" domain, where situational cognition is dominant, are arranged.

[0137] The emotion map defines two types of emotions that promote learning. One is a negative emotion located near the middle of "repentance" or "reflection" on the situation side. That is, when the robot experiences negative emotions such as "I never want to feel this way again" or "I never want to be scolded again." The other is a positive emotion located near "desire" on the response side. That is, when the robot experiences positive feelings such as "wanting more" or "wanting to know more."

[0138] The emotion-specific model 59 feeds user input into a pre-learned neural network to obtain emotion values ​​representing each emotion shown in the emotion graph 400, and determines the user's emotion. This neural network is pre-learned based on multiple learning data sets that combine user input with emotion values ​​representing each emotion shown in the emotion graph 400. Furthermore, this neural network is learned to... Figure 10 As shown in sentiment graph 900, sentiment values ​​in nearby configurations are similar to each other. Figure 10 Examples show that multiple emotions such as "peace of mind", "stability", and "reassurance" have similar emotional values.

[0139] In the above embodiments, a specific processing is described by a single computer 22, but the technology disclosed herein is not limited to this, and distributed processing by multiple computers, including computer 22, is also possible.

[0140] In the above embodiments, an example of storing a specific processing program 56 in memory 32 is illustrated, but the technology disclosed herein is not limited thereto. For example, the specific processing program 56 may also be stored in a portable computer-readable non-transitory storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-transitory storage medium is installed in the computer 22 of the data processing device 12. The processor 28 performs specific processing according to the specific processing program 56.

[0141] Alternatively, the specific processing program 56 can be stored in a storage device such as a server connected to the data processing device 12 via a network 54, and the specific processing program 56 can be downloaded and installed into the computer 22 upon request from the data processing device 12.

[0142] Furthermore, it is not necessary to store the entire specific process 56 in a storage device such as a server connected to the data processing device 12 via the network 54, nor is it necessary to store the entire specific process 56 in the memory 32; a portion of the specific process 56 may also be stored.

[0143] As a hardware resource for performing specific processing, various processors can be used. For example, a CPU is a general-purpose processor that functions as a hardware resource for performing specific processing by executing software, i.e., programs. Additionally, a dedicated circuit can be listed as a processor; it is a processor with a circuit structure specifically designed for performing specific processing, such as a FPGA (Field-Programmable Gate Array), PLD (Programmable Logic Device), or ASIC (Application Specific Integrated Circuit). Every processor has built-in or connected memory, and every processor executes specific processing by using that memory.

[0144] The hardware resources for performing a specific process can consist of one of these various processors, or a combination of two or more processors of the same or different types (e.g., a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Furthermore, the hardware resources for performing a specific process can also be a single processor.

[0145] As an example of a single processor, the first type consists of a combination of one or more CPUs and software, which functions as a hardware resource to perform specific processing. The second type uses a processor, such as a System-on-a-chip (SoC), which implements the entire system functionality, including multiple hardware resources performing specific processing, using a single IC chip. In this case, the specific processing is implemented using one or more of the aforementioned processors that serve as hardware resources.

[0146] Furthermore, as the hardware architecture of these various processors, more specifically, circuits combining semiconductor elements and other circuit components can be used. Moreover, the specific process described above is merely an example. Therefore, it goes without saying that, without departing from the main point, unnecessary steps can be removed, new steps can be added, or the processing order can be changed.

[0147] Furthermore, although the above examples have been described in terms of first to fourth embodiments, some or all of these embodiments can also be combined. Additionally, the smart device 14, smart glasses 214, head-mounted terminal 314, and robot 414 are just examples; they can be combined separately or are other devices.

[0148] The foregoing descriptions and illustrations are detailed explanations of the parts covered by this disclosure and are merely one example of this disclosure. For instance, the descriptions of the above-described structure, function, role, and effect are just one example of the structure, function, role, and effect of the parts covered by this disclosure. Therefore, it goes without saying that, without departing from the spirit of this disclosure, unnecessary parts can be deleted, new elements can be added, or replacements can be made to the foregoing descriptions and illustrations. Furthermore, to avoid confusion and facilitate understanding of the parts covered by this disclosure, explanations of technical common sense that does not require special explanation for implementing this disclosure have been omitted from the foregoing descriptions and illustrations.

[0149] All documents, patent applications and technical standards described in this specification are incorporated herein by reference as if they were specifically and individually described as incorporated by reference.

[0150] (Note 1) A system comprising: The collection department is used to collect the speaker's remarks in real time. The parsing unit is used to parse the speech data collected by the collection unit; and The estimation section is used to estimate the individual's intention based on the information parsed by the analysis section.

[0151] (Note 2) The system as described in Appendix 1 is characterized by including a recording unit for recording a life log.

[0152] (Note 3) The system as described in Appendix 1 is characterized by including a prevention unit that implements the function of preventing wandering.

[0153] (Note 4) The system as described in Appendix 1 is characterized in that, The collection department presumes the individual's emotions and, based on the presumed emotions, determines the method for adjusting the timing of speech collection.

[0154] (Note 5) The system as described in Appendix 1 is characterized in that, The collection department analyzes the individual's past speaking history and selects appropriate collection methods.

[0155] (Note 6) The system as described in Appendix 1 is characterized in that, When collecting comments, the collection department filters them based on the individual's life circumstances and areas of interest.

[0156] (Note 7) The system as described in Appendix 1 is characterized in that, The collection department presumes the individual's emotions and, based on the presumed emotions, determines the priority order of the collected statements.

[0157] (Note 8) The system as described in Appendix 1 is characterized in that, When collecting comments, the collection department prioritizes collecting comments with high relevance based on the user's geographical location information.

[0158] (Note 9) The system as described in Appendix 1 is characterized in that, When collecting comments, the collection department analyzes the individual's social media activity and collects relevant comments.

[0159] (Postscript 10) The system as described in Appendix 1 is characterized in that, The analysis unit presumes the user's emotions and, based on the presumed emotions, determines the mode of expression for the analysis.

[0160] (Postscript 11) The system as described in Appendix 1 is characterized in that, During the analysis process, the analysis unit adjusts the level of detail based on the importance of the statement.

[0161] (Postscript 12) The system as described in Appendix 1 is characterized in that, During the parsing process, the parsing unit applies different parsing algorithms based on the type of speech.

[0162] (Postscript 13) The system as described in Appendix 1 is characterized in that, The analysis unit presumes the user's emotions and determines the length of the analysis based on the presumed emotions.

[0163] (Postscript 14) The system as described in Appendix 1 is characterized in that, During parsing, the parsing unit determines the parsing priority based on the timing of the message submission.

[0164] (Postscript 15) The system as described in Appendix 1 is characterized in that, During the parsing process, the parsing unit adjusts the parsing order based on the relevance of the statements.

[0165] (Postscript 16) The system as described in Appendix 1 is characterized in that, The presumption unit presumes the individual's emotions and, based on the presumed emotions, determines the method of presuming intentions.

[0166] (Postscript 17) The system as described in Appendix 1 is characterized in that, The estimation unit optimizes the estimation algorithm by referring to past speech data during the estimation process.

[0167] (Postscript 18) The system as described in Appendix 1 is characterized in that, When making inferences, the inference department applies different inference methods depending on the type of speech.

[0168] (Postscript 19) The system as described in Appendix 1 is characterized in that, The method of presuming one's own emotions and determining the order of presumptions based on the presumed emotions of the person in question.

[0169] (Postscript 20) The system as described in Appendix 1 is characterized in that, When making a prediction, the presumption department weights the prediction based on the timing of the submission of the statement.

[0170] (Postscript 21) The system as described in Appendix 1 is characterized in that, When making inferences, the inference department refers to relevant market data from the speech.

[0171] (Postscript 22) The system as described in Appendix 2 is characterized in that, The recording department presumes the individual's emotions and, based on the presumed emotions, determines the method of recording the life log.

[0172] (Postscript 23) The system as described in Appendix 2 is characterized in that, The recording unit optimizes the recording algorithm by referring to past life log data during the recording process.

[0173] (Postscript 24) The system as described in Appendix 2 is characterized in that, The recording unit presumes the individual's emotions and determines the recording frequency based on the presumed emotions.

[0174] (Postscript 25) The system as described in Appendix 2 is characterized in that, When recording, the recording department weights the recorded data according to the submission time of the life log.

[0175] (Postscript 26) The system as described in Appendix 3 is characterized in that, The prevention department presumes the individual's emotions and, based on the presumed emotions, determines a method to prevent hesitation.

[0176] (Postscript 27) The system as described in Appendix 3 is characterized in that, When performing prevention, the prevention unit refers to past lingering data to optimize the prevention algorithm.

[0177] (Postscript 28) The system as described in Appendix 3 is characterized in that, The method of the prevention department presumes the individual's emotions and determines the priority of prevention based on the presumed emotions.

[0178] (Postscript 29) The system as described in Appendix 3 is characterized in that, When preventing incidents, the prevention unit weights the prevention data according to the timing of the incident.

Claims

1. A system, characterized in that, include: The collection department is used to collect the speaker's remarks in real time. The parsing unit is used to parse the speech data collected by the collection unit; as well as The estimation section is used to estimate the individual's intentions based on the information parsed by the analysis section.

2. The system as described in claim 1, characterized in that, This includes a record-keeping department for keeping a daily log.

3. The system as described in claim 1, characterized in that, This includes a prevention unit that prevents loitering.

4. The system as described in claim 1, characterized in that, The collection department presumes the individual's emotions and, based on the presumed emotions, determines the method for adjusting the timing of speech collection.

5. The system as described in claim 1, characterized in that, The collection department analyzes the individual's past speaking history and selects appropriate collection methods.

6. The system as described in claim 1, characterized in that, When collecting comments, the collection department filters them based on the individual's life circumstances and areas of interest.

7. The system as described in claim 1, characterized in that, The collection department presumes the individual's emotions and, based on the presumed emotions, determines the priority order of the collected statements.

8. The system as described in claim 1, characterized in that, When collecting comments, the collection department prioritizes collecting comments with high relevance based on the user's geographical location information.

9. The system as described in claim 1, characterized in that, When collecting comments, the collection department analyzes the individual's social media activity and collects relevant comments.

10. The system as claimed in claim 1, characterized in that, The analysis unit presumes the user's emotions and, based on the presumed emotions, determines the mode of expression for the analysis.

Citation Information

Patent Citations

  • Persona chatbot control method and system

    JP2022180282A