system

The system addresses privacy and inefficiency in surveillance by converting video into stylized representations with captions and detecting anomalies, reducing effort and enhancing safety.

JP2026072666APending Publication Date: 2026-05-01SOFTBANK GROUP CORP
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
SOFTBANK GROUP CORP
Filing Date
2024-10-18
Publication Date
2026-05-01

AI Technical Summary

Technical Problem

Existing surveillance systems face concerns about privacy and the inefficiency of video review, particularly in monitoring cameras.

Method used

A system that utilizes an acquisition unit to capture video, an analysis unit to convert it into a visually stylized representation, a caption insertion unit to add annotation captions, and an anomaly detection unit to detect anomalies, with an alert issuing unit to notify of incidents, thereby reducing privacy concerns and the effort required for video review.

Benefits of technology

The system alleviates privacy concerns and reduces the effort needed for video review by converting footage into stylized representations with annotation captions and promptly alerting for anomalies, enhancing safety and efficiency in monitoring.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026072666000001_ABST
    Figure 2026072666000001_ABST
Patent Text Reader

Abstract

The system according to this embodiment aims to alleviate privacy concerns regarding surveillance cameras and reduce the effort required for video review. [Solution] The system according to the embodiment comprises an acquisition unit, an analysis unit, a caption insertion unit, an anomaly detection unit, and an alert issuing unit. The acquisition unit acquires video footage captured by a camera. The analysis unit analyzes the video footage acquired by the acquisition unit and converts it into a visually stylized representation. The caption insertion unit inserts annotation captions based on the actions and situations in the video footage analyzed by the analysis unit. The anomaly detection unit detects anomalies based on the annotation captions inserted by the caption insertion unit. The alert issuing unit issues an alert based on the anomaly detected by the anomaly detection unit.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The technology of the present disclosure relates to a system.

Background Art

[0002] Patent Document 1 discloses a method for controlling a persona chatbot, which is performed by at least one processor, including steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to an explanation of a chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance.

Prior Art Documents

Patent Documents

[0003]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0004] In the prior art, there were problems such as great concerns about privacy and the trouble of video confirmation in a monitoring camera.

[0005] [[ID=第38]] The system according to the embodiment aims to reduce concerns about privacy of the monitoring camera and reduce the trouble of video confirmation.

Means for Solving the Problems

[0006] The system according to this embodiment comprises an acquisition unit, an analysis unit, a caption insertion unit, an anomaly detection unit, and an alert issuing unit. The acquisition unit acquires video footage captured by a camera. The analysis unit analyzes the video footage acquired by the acquisition unit and converts it into a visually stylized representation. The caption insertion unit inserts annotation captions based on the actions and situations in the video footage analyzed by the analysis unit. The anomaly detection unit detects anomalies based on the annotation captions inserted by the caption insertion unit. The alert issuing unit issues an alert based on the anomaly detected by the anomaly detection unit. [Effects of the Invention]

[0007] The system according to this embodiment can alleviate privacy concerns regarding surveillance cameras and reduce the effort required for video review. [Brief explanation of the drawing]

[0008] [Figure 1] This is a conceptual diagram showing an example of the configuration of a data processing system according to the first embodiment. [Figure 2] This is a conceptual diagram showing an example of the essential functions of a data processing device and a smart device according to the first embodiment. [Figure 3] This is a conceptual diagram showing an example of the configuration of a data processing system according to the second embodiment. [Figure 4] This is a conceptual diagram showing an example of the main functions of a data processing device and smart glasses according to the second embodiment. [Figure 5] This is a conceptual diagram showing an example of the configuration of a data processing system according to the third embodiment. [Figure 6] This is a conceptual diagram showing an example of the main functions of a data processing device and a headset-type terminal according to the third embodiment. [Figure 7] This is a conceptual diagram showing an example of the configuration of a data processing system according to the fourth embodiment. [Figure 8] This is a conceptual diagram showing an example of the main functions of a data processing device and a robot according to the fourth embodiment. [Figure 9]This shows an emotion map where multiple emotions are mapped. [Figure 10] This shows an emotion map where multiple emotions are mapped. [Modes for carrying out the invention]

[0009] Hereinafter, an example of an embodiment of the system relating to the technology of this disclosure will be described with reference to the attached drawings.

[0010] First, let's explain the terminology used in the following explanation.

[0011] In the following embodiments, the signed processor (hereinafter simply referred to as "processor") may be a single arithmetic unit or a combination of multiple arithmetic units. Furthermore, the processor may be a single type of arithmetic unit or a combination of multiple types of arithmetic units. Examples of arithmetic units include CPU (Central Processing Unit), GPU (Graphics Processing Unit), GPGPU (General-Purpose computing on Graphics Processing Units), APU (Accelerated Processing Unit), or TPU (Tensor Processing Unit).

[0012] In the following embodiments, signed RAM (Random Access Memory) is a memory that temporarily stores information and is used as work memory by the processor.

[0013] In the following embodiments, the signed storage is one or more non-volatile storage devices that store various programs and various parameters. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), or magnetic tapes.

[0014] In the following embodiments, the labeled communication I / F (Interface) is an interface including a communication processor, an antenna, and the like. The communication I / F manages communication between a plurality of computers. Examples of communication standards applicable to the communication I / F include wireless communication standards such as 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), or Bluetooth (registered trademark).

[0015] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B". That is, "A and / or B" means that it may be only A, only B, or a combination of A and B. Also, in this specification, when expressing three or more matters connected by "and / or", the same concept as "A and / or B" is applied.

[0016] [First Embodiment] FIG. 1 shows an example of the configuration of a data processing system 10 according to the first embodiment.

[0017] As shown in FIG. 1, the data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.

[0018] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. Also, the database 24 and the communication I / F 26 are connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0019] The smart device 14 comprises a computer 36, a receiving device 38, an output device 40, a camera 42, and a communication interface 44. The computer 36 comprises a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The receiving device 38, output device 40, and camera 42 are also connected to the bus 52.

[0020] The reception device 38 is equipped with a touch panel 38A and a microphone 38B, and accepts user input. The touch panel 38A accepts user input via touch by detecting contact with an object (e.g., a pen or finger). The microphone 38B accepts user input via voice by detecting the user's voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and microphone 38B to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 (see Figure 2) acquires the data indicating the user input.

[0021] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user by outputting the data in a form perceptible to the user (e.g., audio and / or text). The display 40A displays visible information such as text and images according to instructions from the processor 46. The speaker 40B outputs audio according to instructions from the processor 46. The camera 42 is a small digital camera equipped with an optical system such as a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.

[0022] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various types of information between processor 46 and processor 28 via network 54.

[0023] Figure 2 shows an example of the main functions of the data processing device 12 and the smart device 14.

[0024] As shown in Figure 2, in the data processing device 12, a specific processing is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" related to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.

[0025] Storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290. The identification processing unit 290 can estimate the user's emotions using the emotion identification model 59 and perform identification processing using the user's emotions. The emotion estimation function (emotion identification function) using the emotion identification model 59 performs various estimations and predictions regarding the user's emotions, including but not limited to these examples. Furthermore, emotion estimation and prediction also include, for example, emotion analysis.

[0026] In the smart device 14, specific processing is performed by the processor 46. The storage 50 stores a specific processing program 60. The specific processing program 60 is used in conjunction with the specific processing program 56 by the data processing system 10. The processor 46 reads the specific processing program 60 from the storage 50 and executes the read specific processing program 60 on the RAM 48. The specific processing is realized by the processor 46 operating as a control unit 46A according to the specific processing program 60 executed on the RAM 48. The smart device 14 also has a data generation model 58 and an emotion identification model 59, similar to the data generation model and emotion identification model 59, and can perform processing similar to that of the specific processing unit 290 using these models.

[0027] Furthermore, other devices besides the data processing device 12 may also have the data generation model 58. For example, a server device (e.g., a generation server) may have the data generation model 58. In this case, the data processing device 12 obtains processing results (such as prediction results) using the data generation model 58 by communicating with the server device having the data generation model 58. The data processing device 12 may also be a server device or a terminal device owned by a user (e.g., a mobile phone, robot, home appliance, etc.). Next, an example of processing by the data processing system 10 according to the first embodiment will be described.

[0028] (Example of form 1) The monitoring system according to an embodiment of the present invention is a system that eliminates privacy concerns and the hassle of reviewing video in existing monitoring cameras. This monitoring system reduces the effort of monitoring and reviewing video by using a generation AI to distort the video captured by the camera, analyzing what is happening in the video, and inserting annotation captions. In addition, if an accident or other incident is suspected, an alert notification is immediately issued. For example, video captured by the camera is input to the generation AI. The generation AI analyzes the video and converts it into a visually distorted representation. For example, video of an elderly person walking in a room can be converted into video that looks like an animated character walking. This allows for understanding the content of the video while protecting privacy. Next, the generation AI analyzes what is happening in the video. The generation AI recognizes the actions and situations in the video and inserts annotation captions based on that. For example, if an elderly person is eating, an annotation caption such as "eating" can be displayed. This significantly reduces the effort of reviewing the video. Furthermore, if an accident or other incident is suspected, an alert notification is immediately issued. The generation AI issues an alert when it detects abnormal actions or situations in the video. For example, if an elderly person falls, an alert notification can be sent immediately, prompting a quick response. This system eliminates privacy concerns associated with monitoring cameras and reduces the effort required to review video footage. For instance, it can provide a secure environment for family members and caregivers to monitor the elderly. It also contributes to improved safety by enabling a quick response to emergencies such as accidents. In this way, the monitoring system can reduce the effort required to review video footage while protecting privacy.

[0029] The monitoring system according to this embodiment comprises an acquisition unit, an analysis unit, a text overlay insertion unit, an anomaly detection unit, and an alert issuing unit. The acquisition unit acquires video footage captured by a camera. The acquisition unit can, for example, acquire video footage captured by a camera in real time. The acquisition unit can also acquire video footage from the camera periodically. Furthermore, the acquisition unit can acquire video footage from the camera in streaming format. For example, the acquisition unit can acquire video footage from the camera at 30 frames per second. The acquisition unit can also acquire video footage from the camera every hour. Furthermore, the acquisition unit can acquire video footage from the camera in live streaming format. The analysis unit analyzes the video footage acquired by the acquisition unit and converts it into a visually stylized representation. For example, the analysis unit can stylize the video in an anime style. Furthermore, the analysis unit can stylize the video in a manga style. Furthermore, the analysis unit can silhouette the video. For example, the analysis unit can stylize the video in an anime character style. The analysis unit can also deform the video into a comic strip-like format. Furthermore, the analysis unit can silhouette the video to protect privacy. The caption insertion unit inserts annotation captions based on the actions and situations in the video analyzed by the analysis unit. For example, if an elderly person is walking in the video, the caption insertion unit will insert an annotation caption such as "Walking." Also, if an elderly person is eating in the video, the caption insertion unit can insert an annotation caption such as "Eating." Furthermore, if an elderly person is resting in the video, the caption insertion unit can insert an annotation caption such as "Resting." For example, if an elderly person is walking in the video, the caption insertion unit will insert an annotation caption such as "Walking." Also, if an elderly person is eating in the video, the caption insertion unit can insert an annotation caption such as "Eating." Furthermore, if an elderly person is resting in the video, the caption insertion unit can insert an annotation caption such as "Resting." The anomaly detection unit detects anomalies based on the annotation captions inserted by the caption insertion unit. The anomaly detection unit detects an anomaly, for example, if an elderly person falls in the video.Furthermore, the anomaly detection unit can also detect anomalies if an elderly person remains motionless for a long period of time in the video. In addition, the anomaly detection unit can also detect anomalies if an elderly person performs suspicious actions in the video. For example, the anomaly detection unit detects an anomaly if an elderly person falls in the video. Furthermore, the anomaly detection unit can also detect an anomaly if an elderly person remains motionless for a long period of time in the video. Furthermore, the anomaly detection unit can also detect an anomaly if an elderly person performs suspicious actions in the video. The alert issuing unit issues an alert based on the anomaly detected by the anomaly detection unit. For example, the alert issuing unit issues an alert if an elderly person falls. Furthermore, the alert issuing unit can also issue an alert if an elderly person remains motionless for a long period of time. Furthermore, the alert issuing unit can also issue an alert if an elderly person performs suspicious actions. As a result, the monitoring system according to this embodiment can reduce the effort required for video confirmation while protecting privacy.

[0030] The acquisition unit acquires video footage captured by cameras. For example, the acquisition unit can acquire video footage in real time. It can also acquire camera footage periodically. Furthermore, the acquisition unit can acquire camera footage in streaming format. For example, the acquisition unit can acquire camera footage at 30 frames per second. It can also acquire camera footage every hour. Additionally, the acquisition unit can acquire camera footage in live streaming format. Specifically, the acquisition unit can acquire video from multiple cameras simultaneously, enabling wide-area surveillance. Various types of cameras are used, including fixed cameras and pan-tilt-zoom (PTZ) cameras. Fixed cameras constantly monitor a specific area, while PTZ cameras allow for remotely changing the field of view. This allows for the acquisition of detailed footage when specific events or anomalies occur. Furthermore, the acquisition unit has the ability to dynamically adjust the video resolution and frame rate, optimizing network bandwidth and storage capacity. For example, it can acquire video at a low resolution under normal circumstances and switch to high resolution when an anomaly is detected. This enables efficient data management. Furthermore, the acquisition unit ensures data security by encrypting and transmitting video data. This enhances privacy protection and prevents unauthorized access from external sources.

[0031] The analysis unit analyzes the video acquired by the acquisition unit and converts it into a visually stylized representation. For example, the analysis unit can stylize the video in an anime style. It can also stylize the video in a manga style. Furthermore, the analysis unit can silhouette the video. For example, the analysis unit can stylize the video in an anime character style. It can also stylize the video in a manga panel style. Furthermore, the analysis unit can silhouette the video to protect privacy. Specifically, the analysis unit uses AI to recognize people and objects in the video and converts them into a specific style. For example, it utilizes image recognition technology using deep learning to analyze people's movements and expressions and convert them into an anime or manga style. When silhouetting, it extracts the outline of the person and omits internal details to protect privacy. Furthermore, the analysis unit has a function to automatically blur the background in the video, hiding information other than people. This further enhances privacy protection. In addition, the analysis unit can update the video analysis results in real time and respond to dynamic changes. For example, if a person moves or a new object appears in the image, this is immediately reflected in the stylized representation. This allows for the constant provision of up-to-date information.

[0032] The caption insertion unit inserts annotation captions based on the actions and situations in the video analyzed by the analysis unit. For example, if an elderly person is walking in the video, the caption insertion unit will insert an annotation caption such as "Walking." It can also insert an annotation caption such as "Eating" if an elderly person is eating in the video. Furthermore, it can insert an annotation caption such as "Resting" if an elderly person is resting in the video. Specifically, the caption insertion unit uses AI to automatically recognize actions and situations in the video and generates corresponding captions. For example, it uses an action recognition algorithm to detect actions such as walking, eating, and resting by elderly people and inserts corresponding captions. Furthermore, the content of the on-screen text is designed to be customizable by the user, allowing for the setting of messages tailored to specific situations. In addition, the on-screen text insertion unit has the functionality to dynamically adjust the display position, font, and color of the text, optimizing the visibility of the video. For example, it adjusts the position of the text so as not to obscure important parts of the video, ensuring visibility. It also adjusts the display time of the text to provide necessary information at the appropriate time. As a result, the on-screen text insertion unit can provide appropriate annotations according to the actions and situations in the video, making the information easy for viewers to understand.

[0033] The anomaly detection unit detects anomalies based on annotation captions inserted by the caption insertion unit. For example, the anomaly detection unit detects an anomaly if an elderly person falls in the video. The anomaly detection unit can also detect an anomaly if an elderly person remains motionless for a long period of time in the video. Furthermore, the anomaly detection unit can also detect an anomaly if an elderly person performs suspicious actions in the video. Specifically, the anomaly detection unit uses AI to analyze the motion patterns in the video and detect unusual actions or situations. For example, it uses a fall detection algorithm to detect an anomaly if an elderly person suddenly falls. It also uses a motion stop detection algorithm to detect an anomaly if an elderly person remains motionless for a certain period of time or longer. Furthermore, it uses a suspicious action detection algorithm to detect an anomaly if an elderly person performs unusual actions. This allows the anomaly detection unit to detect anomalies in real time and respond quickly. Furthermore, the anomaly detection unit can learn patterns of anomalies based on past data and improve detection accuracy. For example, it can improve the accuracy of fall detection by learning from past fall incidents. The anomaly detection unit also has a function to immediately issue an alert when an anomaly is detected, supporting a rapid response. As a result, the anomaly detection unit can detect anomalies with high accuracy and enable a quick and appropriate response.

[0034] The alert unit issues alerts based on anomalies detected by the anomaly detection unit. For example, the alert unit issues an alert if an elderly person falls. The alert unit can also issue an alert if an elderly person remains motionless for an extended period. Furthermore, the alert unit can issue an alert if an elderly person exhibits suspicious behavior. Specifically, when an anomaly is detected, the alert unit immediately notifies the relevant parties. Notification methods include smartphone push notifications, SMS, email, and voice calls. For example, it notifies family members and caregivers in real time that an anomaly has occurred, prompting a quick response. The alert unit can also customize the content of notifications, sending appropriate messages according to the type and urgency of the anomaly. Furthermore, the alert unit has a function to record notification history for later review. This allows for understanding past incidents and using that information to inform future countermeasures. Furthermore, the alert unit has a function to automatically notify emergency contacts when an incident occurs, supporting a rapid response. For example, it automatically notifies registered medical institutions and police to encourage prompt rescue operations. In this way, the alert unit supports a swift and appropriate response when an incident occurs, ensuring the safety of the elderly.

[0035] The analysis unit can convert video into a visually stylized representation. For example, the analysis unit can stylize video in an anime style. The analysis unit can also stylize video in a manga style. The analysis unit can also silhouette video. By stylizing the video, the content of the video can be understood while protecting privacy. Visually stylized representations include, but are not limited to, mosaic processing, silhouettes, and color changes. Some or all of the above processing in the analysis unit is performed using a generation AI. For example, the analysis unit inputs video captured by a camera into the generation AI, which analyzes the video and converts it into a stylized representation.

[0036] The caption insertion unit can insert annotation captions based on actions and situations within the video. For example, if an elderly person is walking in the video, the caption insertion unit can insert an annotation caption such as "Walking." For example, if an elderly person is eating in the video, the caption insertion unit can also insert an annotation caption such as "Eating." For example, if an elderly person is resting in the video, the caption insertion unit can also insert an annotation caption such as "Resting." This reduces the effort required to review the video by inserting annotation captions. Actions and situations include, but are not limited to, actions such as walking, running, sitting, and standing, and situations such as crowding and silence. Some or all of the above processing in the caption insertion unit is performed using AI. For example, the caption insertion unit inputs actions and situations in the video into the AI, and the AI ​​analyzes the actions and situations to generate annotation captions.

[0037] The anomaly detection unit can detect abnormal movements or situations within the video. For example, the anomaly detection unit can detect an anomaly if an elderly person falls in the video. The anomaly detection unit can also detect an anomaly if an elderly person remains motionless for a long period of time in the video. The anomaly detection unit can also detect an anomaly if an elderly person performs suspicious movements in the video. This enables a rapid response by detecting abnormal movements or situations. Abnormal movements or situations include, but are not limited to, suspicious movements, unusual sounds, and specific patterns. Some or all of the above processing in the anomaly detection unit is performed using AI. For example, the anomaly detection unit inputs the movements and situations in the video into the AI, which analyzes the movements and situations to detect anomalies.

[0038] The alert issuing unit can issue an alert when an anomaly is detected. For example, the alert issuing unit will issue an alert if an elderly person falls. The alert issuing unit can also issue an alert if an elderly person remains motionless for a long period of time. The alert issuing unit can also issue an alert if an elderly person performs suspicious actions. This enables a rapid response by issuing an alert when an anomaly is detected. Alerts include, but are not limited to, voice alerts, visual alerts, and notification messages. Some or all of the above processing in the alert issuing unit is performed using AI. For example, the alert issuing unit inputs the anomaly detected by the anomaly detection unit to the AI, which analyzes the anomaly and generates an alert.

[0039] The acquisition unit can select the optimal acquisition method based on the user's activity pattern when acquiring video. For example, if the user is active during the day, the acquisition unit will acquire video during the daytime. For example, if the user is active at night, the acquisition unit can also acquire video at night. For example, if the user's activity pattern is irregular, the acquisition unit can use AI to learn and adjust the optimal acquisition timing. This enables efficient video acquisition by selecting the optimal acquisition method based on the user's activity pattern. The optimal acquisition method includes, but is not limited to, camera angle, resolution, and frame rate. Some or all of the above processing in the acquisition unit is performed using AI. For example, the acquisition unit inputs user activity data into the AI, and the AI ​​analyzes the activity data to select the optimal acquisition method.

[0040] The acquisition unit can filter video footage based on the user's current environment and circumstances. For example, if the user is indoors, the acquisition unit will prioritize acquiring indoor video. If the user is outdoors, the acquisition unit can also prioritize acquiring outdoor video. The acquisition unit can also automatically adjust the filtering using AI if the user's environment changes. This enables the acquisition of appropriate video footage by filtering based on the user's current environment and circumstances. Current environment and circumstances include, but are not limited to, indoor / outdoor distinctions, weather, and time of day. Some or all of the above processing in the acquisition unit is performed using AI. For example, the acquisition unit inputs the user's environmental data into the AI, which then analyzes the environmental data and performs filtering.

[0041] The acquisition unit can prioritize acquiring highly relevant video footage by considering the user's geographical location information when acquiring video. For example, if the user is at home, the acquisition unit will prioritize acquiring video footage of the area around the user's home. If the user is out, the acquisition unit can also prioritize acquiring video footage of the location they are at. If the user is traveling, the acquisition unit can also prioritize acquiring video footage of their travel destination. This allows for the acquisition of appropriate video footage by prioritizing highly relevant video footage while considering the user's geographical location information. Geographical location information includes, but is not limited to, GPS data and Wi-Fi location information. Some or all of the above processing in the acquisition unit is performed using AI. For example, the acquisition unit inputs the user's location data into the AI, which analyzes the location data and selects highly relevant video footage.

[0042] The acquisition unit can analyze the user's social media activity when acquiring video and acquire relevant videos. For example, the acquisition unit can prioritize acquiring videos of places the user has shared on social media. The acquisition unit can also prioritize acquiring videos of places the user has shown interest in on social media. For example, the acquisition unit can acquire videos of relevant events from the user's social media activity. This enables appropriate video acquisition by analyzing the user's social media activity and acquiring relevant videos. Social media activity includes, but is not limited to, posts, the number of likes, and comments. Some or all of the above processing in the acquisition unit is performed using AI. For example, the acquisition unit inputs the user's social media data into the AI, which analyzes the data and selects relevant videos.

[0043] The analysis unit can adjust the level of detail of the deformation based on specific actions or situations when analyzing video. For example, if an elderly person is walking, the analysis unit will use a deformation that emphasizes walking. For example, if an elderly person is eating, the analysis unit can also use a deformation that emphasizes eating. For example, if an elderly person is resting, the analysis unit can also use a deformation that emphasizes a relaxed posture. By adjusting the level of detail of the deformation based on specific actions or situations, a more appropriate deformation representation becomes possible. Specific actions and situations include, but are not limited to, actions such as walking, running, sitting, and standing, and situations such as crowding and silence. Some or all of the above processing in the analysis unit is performed using a generation AI. For example, the analysis unit inputs the actions and situations in the video into the generation AI, which analyzes the actions and situations and adjusts the level of detail of the deformation.

[0044] The analysis unit can apply different deformation algorithms depending on the scenario when analyzing video. For example, the analysis unit uses a simple deformation algorithm in everyday life scenarios. For example, the analysis unit can also use a detailed deformation algorithm in emergency scenarios. For example, the analysis unit can also use a deformation algorithm tailored to a specific event scenario. By applying different deformation algorithms according to different scenarios, more appropriate deformation representation becomes possible. Different scenarios include, but are not limited to, emergency situations, normal situations, and specific events. Some or all of the above processing in the analysis unit is performed using a generation AI. For example, the analysis unit inputs scenario data into the generation AI, which analyzes the scenario data and selects an appropriate deformation algorithm.

[0045] The analysis unit can determine the priority of distortion based on the time the video was shot when analyzing the video. For example, the analysis unit may prioritize bright distortion for video shot during the day. For example, the analysis unit may also prioritize muted distortion for video shot at night. For example, the analysis unit may also prioritize distortion appropriate for an event for video shot during a specific event. By determining the priority of distortion based on the time the video was shot, appropriate distortion is possible. The shooting time includes, but is not limited to, the date and time, season, and specific events. Some or all of the above processing in the analysis unit is performed using a generation AI. For example, the analysis unit inputs the video shooting time data into the generation AI, and the generation AI analyzes the shooting time data to determine the priority of distortion.

[0046] The analysis unit can adjust the order of deformation based on the relevance of the video during video analysis. For example, the analysis unit may prioritize deformation of video containing important actions. The analysis unit may also postpone deformation of video containing everyday actions. The analysis unit may also prioritize deformation of video containing emergencies. By adjusting the order of deformation based on the relevance of the video, appropriate deformation representation becomes possible. Video relevance includes, but is not limited to, similarity of content and temporal continuity. Some or all of the above processing in the analysis unit is performed using a generation AI. For example, the analysis unit inputs video relevance data into the generation AI, which analyzes the relevance data and adjusts the order of deformation.

[0047] The text overlay insertion unit can adjust the level of detail of the text overlay based on the importance of the action or situation when inserting text. For example, the text overlay insertion unit will insert detailed text when an important action is being performed. For example, the text overlay insertion unit can also insert concise text when a routine action is being performed. For example, the text overlay insertion unit can also insert detailed and rapid text when an emergency occurs. This allows for appropriate text display by adjusting the level of detail of the text overlay based on the importance of the action or situation. The importance of the action or situation includes, but is not limited to, frequency, impact, and urgency. Some or all of the above processing in the text overlay insertion unit is performed using AI. For example, the text overlay insertion unit inputs action and situation data into the AI, which analyzes the data and adjusts the level of detail of the text overlay.

[0048] The text overlay insertion unit can apply different text overlay insertion algorithms depending on the different scenarios during text overlay insertion. For example, in everyday life scenarios, the text overlay insertion unit uses a simple text overlay insertion algorithm. In emergency scenarios, for example, the text overlay insertion unit can also use a more detailed text overlay insertion algorithm. In scenarios of specific events, for example, the text overlay insertion unit can also use an event-specific text overlay insertion algorithm. This enables appropriate text overlay insertion depending on the different scenarios. Different scenarios include, but are not limited to, emergency situations, normal situations, and specific events. Some or all of the above processing in the text overlay insertion unit is performed using AI. For example, the text overlay insertion unit inputs scenario data into the AI, which analyzes the scenario data and selects an appropriate text overlay insertion algorithm.

[0049] The text overlay insertion unit can determine the priority of text overlays based on the timing of the action or situation. For example, the text overlay insertion unit can insert a text overlay immediately after an important action has occurred. It can also insert a text overlay after a routine action has been performed. For example, it can insert a text overlay immediately in the event of an emergency. This allows for appropriate text overlay display by determining the priority of text overlays based on the timing of the action or situation. The timing of the action or situation includes, but is not limited to, the date and time, season, or specific event. Some or all of the above processing in the text overlay insertion unit is performed using AI. For example, the text overlay insertion unit inputs data on the timing of the action or situation into the AI, which analyzes the data to determine the priority of text overlays.

[0050] The text overlay insertion unit can adjust the order of text overlays based on the relevance of actions and situations when inserting them. For example, the text overlay insertion unit will prioritize inserting text overlays related to important actions. For example, the text overlay insertion unit can postpone inserting text overlays related to routine actions. For example, the text overlay insertion unit can prioritize inserting text overlays related to emergencies. This allows for appropriate text overlay display by adjusting the order of text overlays based on the relevance of actions and situations. The relevance of actions and situations includes, but is not limited to, similarity of content and temporal continuity. Some or all of the above processing in the text overlay insertion unit is performed using AI. For example, the text overlay insertion unit inputs data on the relevance of actions and situations into the AI, which analyzes the data and adjusts the order of the text overlays.

[0051] The anomaly detection unit can optimize its anomaly detection algorithm by referring to past anomaly data when an anomaly is detected. For example, the anomaly detection unit adjusts the anomaly detection algorithm based on past anomaly data. The anomaly detection unit can also learn specific patterns from past anomaly data and reflect them in the anomaly detection algorithm. For example, the anomaly detection unit can analyze past anomaly data to improve the accuracy of the anomaly detection algorithm. As a result, the accuracy of anomaly detection is improved by optimizing the anomaly detection algorithm by referring to past anomaly data. Past anomaly data includes, but is not limited to, past anomaly cases, anomaly frequency, and anomaly impact. Some or all of the above processing in the anomaly detection unit is performed using AI. For example, the anomaly detection unit inputs past anomaly data into the AI, and the AI ​​analyzes the data to optimize the anomaly detection algorithm.

[0052] The anomaly detection unit can apply different anomaly detection methods depending on the different scenarios when an anomaly is detected. For example, in everyday life scenarios, the anomaly detection unit uses a simple anomaly detection method. In emergency scenarios, for example, the anomaly detection unit can also use a more detailed anomaly detection method. In scenarios of specific events, for example, the anomaly detection unit can also use an event-specific anomaly detection method. This improves the accuracy of anomaly detection by applying the appropriate anomaly detection method according to different scenarios. Different scenarios include, but are not limited to, emergency situations, normal situations, and specific events. Some or all of the above processing in the anomaly detection unit is performed using AI. For example, the anomaly detection unit inputs scenario data into the AI, which analyzes the scenario data and selects an appropriate anomaly detection method.

[0053] The anomaly detection unit can perform anomaly detection while considering the geographical distribution of the video. For example, the anomaly detection unit can perform anomaly detection based on the location where the video was shot. The anomaly detection unit can also analyze the geographical distribution of the video and adjust the anomaly detection algorithm. For example, the anomaly detection unit can prioritize anomaly detection in a specific area based on the geographical distribution of the video. This makes it possible to perform appropriate anomaly detection by considering the geographical distribution of the video. Geographical distribution includes, but is not limited to, the anomaly occurrence rate for each region and geographical characteristics. Some or all of the above processing in the anomaly detection unit is performed using AI. For example, the anomaly detection unit inputs geographical distribution data into the AI, and the AI ​​analyzes the data and performs anomaly detection.

[0054] The anomaly detection unit can improve the accuracy of anomaly detection by referring to relevant literature in the video when an anomaly is detected. For example, the anomaly detection unit adjusts the anomaly detection algorithm based on the relevant literature in the video. The anomaly detection unit can also learn specific patterns from the relevant literature in the video and reflect them in the anomaly detection algorithm. For example, the anomaly detection unit can analyze the relevant literature in the video to improve the accuracy of the anomaly detection algorithm. By improving the accuracy of anomaly detection by referring to relevant literature in the video, more accurate anomaly detection becomes possible. Relevant literature includes, but is not limited to, academic papers, technical reports, and patent documents. Some or all of the above processing in the anomaly detection unit is performed using AI. For example, the anomaly detection unit inputs the relevant literature data into the AI, and the AI ​​analyzes the data to adjust the anomaly detection algorithm.

[0055] The alert issuing unit can optimize its alert issuing algorithm by referring to past alert data when issuing an alert. For example, the alert issuing unit adjusts the alert issuing algorithm based on past alert data. The alert issuing unit can also learn specific patterns from past alert data and reflect them in the alert issuing algorithm. For example, the alert issuing unit can analyze past alert data to improve the accuracy of the alert issuing algorithm. This improves the accuracy of alert issuing by optimizing the alert issuing algorithm by referring to past alert data. Past alert data includes, but is not limited to, past alert cases, alert frequency, and alert impact. Some or all of the above processing in the alert issuing unit is performed using AI. For example, the alert issuing unit inputs past alert data into the AI, and the AI ​​analyzes the data to optimize the alert issuing algorithm.

[0056] The alert issuing unit can apply different alert issuing methods depending on the different scenarios when issuing an alert. For example, in everyday life scenarios, the alert issuing unit uses a simple alert issuing method. In emergency scenarios, for example, the alert issuing unit can also use a detailed alert issuing method. In scenarios of specific events, for example, the alert issuing unit can also use an event-specific alert issuing method. This improves the accuracy of alert issuing by applying the appropriate alert issuing method according to different scenarios. Different scenarios include, but are not limited to, emergency situations, normal situations, and specific events. Some or all of the above processing in the alert issuing unit is performed using AI. For example, the alert issuing unit inputs scenario data into the AI, and the AI ​​analyzes the scenario data to select the appropriate alert issuing method.

[0057] The alert issuing unit can adjust the timing of alert issuance based on the video recording date. For example, the alert issuing unit uses the normal alert timing for video recorded during the day. For example, the alert issuing unit can prioritize issuing high-urgency alerts for video recorded at night. For example, the alert issuing unit can use an alert timing tailored to a specific event for video recorded during a particular event. By adjusting the alert issuance timing based on the video recording date, appropriate alerts can be issued. The recording date includes, but is not limited to, the date and time, season, and specific events. Some or all of the above processing in the alert issuing unit is performed using AI. For example, the alert issuing unit inputs recording date data into the AI, which analyzes the data and adjusts the alert issuance timing.

[0058] The alert issuing unit can adjust the order in which alerts are issued based on the relevance of the video footage. For example, the alert issuing unit will issue an alert preferentially if it relates to an important action. For example, the alert issuing unit can also issue an alert later if it relates to a routine action. For example, the alert issuing unit can issue an alert with the highest priority if it relates to an emergency. This allows for appropriate alert issuance by adjusting the order in which alerts are issued based on the relevance of the video footage. Video relevance includes, but is not limited to, similarity of content and temporal continuity. Some or all of the above processing in the alert issuing unit is performed using AI. For example, the alert issuing unit inputs relevance data into the AI, and the AI ​​analyzes the data to adjust the order in which alerts are issued.

[0059] The system according to the embodiment is not limited to the example described above, and various modifications are possible, for example, as follows.

[0060] The acquisition unit can select the optimal acquisition method based on the user's activity pattern when acquiring video. For example, if the user is active during the day, the acquisition unit will acquire video during the daytime. For example, if the user is active at night, the acquisition unit can also acquire video at night. For example, if the user's activity pattern is irregular, the acquisition unit can use AI to learn and adjust the optimal acquisition timing. This enables efficient video acquisition by selecting the optimal acquisition method based on the user's activity pattern. The optimal acquisition method includes, but is not limited to, camera angle, resolution, and frame rate. Some or all of the above processing in the acquisition unit is performed using AI. For example, the acquisition unit inputs user activity data into the AI, and the AI ​​analyzes the activity data to select the optimal acquisition method.

[0061] The acquisition unit can filter video footage based on the user's current environment and circumstances. For example, if the user is indoors, the acquisition unit will prioritize acquiring indoor video. If the user is outdoors, the acquisition unit can also prioritize acquiring outdoor video. The acquisition unit can also automatically adjust the filtering using AI if the user's environment changes. This enables the acquisition of appropriate video footage by filtering based on the user's current environment and circumstances. Current environment and circumstances include, but are not limited to, indoor / outdoor distinctions, weather, and time of day. Some or all of the above processing in the acquisition unit is performed using AI. For example, the acquisition unit inputs the user's environmental data into the AI, which then analyzes the environmental data and performs filtering.

[0062] The analysis unit can adjust the level of detail of the deformation based on specific actions or situations when analyzing video. For example, if an elderly person is walking, the analysis unit will use a deformation that emphasizes walking. For example, if an elderly person is eating, the analysis unit can also use a deformation that emphasizes eating. For example, if an elderly person is resting, the analysis unit can also use a deformation that emphasizes a relaxed posture. By adjusting the level of detail of the deformation based on specific actions or situations, a more appropriate deformation representation becomes possible. Specific actions and situations include, but are not limited to, actions such as walking, running, sitting, and standing, and situations such as crowding and silence. Some or all of the above processing in the analysis unit is performed using a generation AI. For example, the analysis unit inputs the actions and situations in the video into the generation AI, which analyzes the actions and situations and adjusts the level of detail of the deformation.

[0063] The analysis unit can apply different deformation algorithms depending on the scenario when analyzing video. For example, the analysis unit uses a simple deformation algorithm in everyday life scenarios. For example, the analysis unit can also use a detailed deformation algorithm in emergency scenarios. For example, the analysis unit can also use a deformation algorithm tailored to a specific event scenario. By applying different deformation algorithms according to different scenarios, more appropriate deformation representation becomes possible. Different scenarios include, but are not limited to, emergency situations, normal situations, and specific events. Some or all of the above processing in the analysis unit is performed using a generation AI. For example, the analysis unit inputs scenario data into the generation AI, which analyzes the scenario data and selects an appropriate deformation algorithm.

[0064] The text overlay insertion unit can adjust the level of detail of the text overlay based on the importance of the action or situation when inserting text. For example, the text overlay insertion unit will insert detailed text when an important action is being performed. For example, the text overlay insertion unit can also insert concise text when a routine action is being performed. For example, the text overlay insertion unit can also insert detailed and rapid text when an emergency occurs. This allows for appropriate text display by adjusting the level of detail of the text overlay based on the importance of the action or situation. The importance of the action or situation includes, but is not limited to, frequency, impact, and urgency. Some or all of the above processing in the text overlay insertion unit is performed using AI. For example, the text overlay insertion unit inputs action and situation data into the AI, which analyzes the data and adjusts the level of detail of the text overlay.

[0065] The text overlay insertion unit can apply different text overlay insertion algorithms depending on the different scenarios during text overlay insertion. For example, in everyday life scenarios, the text overlay insertion unit uses a simple text overlay insertion algorithm. In emergency scenarios, for example, the text overlay insertion unit can also use a more detailed text overlay insertion algorithm. In scenarios of specific events, for example, the text overlay insertion unit can also use an event-specific text overlay insertion algorithm. This enables appropriate text overlay insertion depending on the different scenarios. Different scenarios include, but are not limited to, emergency situations, normal situations, and specific events. Some or all of the above processing in the text overlay insertion unit is performed using AI. For example, the text overlay insertion unit inputs scenario data into the AI, which analyzes the scenario data and selects an appropriate text overlay insertion algorithm.

[0066] The following briefly describes the processing flow for example form 1.

[0067] Step 1: The acquisition unit acquires the video captured by the camera. The acquisition unit can, for example, acquire the video captured by the camera in real time. The acquisition unit can also acquire the camera's video periodically. Furthermore, the acquisition unit can acquire the camera's video in streaming format. For example, the acquisition unit can acquire the camera's video at 30 frames per second. The acquisition unit can also acquire the camera's video every hour. Furthermore, the acquisition unit can acquire the camera's video in live streaming format. Step 2: The analysis unit analyzes the video acquired by the acquisition unit and converts it into a visually stylized representation. For example, the analysis unit can stylize the video in an anime style. The analysis unit can also stylize the video in a manga style. Furthermore, the analysis unit can silhouette the video. For example, the analysis unit can stylize the video in an anime character style. The analysis unit can also stylize the video in a manga panel style. Furthermore, the analysis unit can silhouette the video to protect privacy. Step 3: The caption insertion unit inserts annotation captions based on the actions and situations in the video analyzed by the analysis unit. For example, if an elderly person is walking in the video, the caption insertion unit will insert an annotation caption such as "Walking." The caption insertion unit can also insert an annotation caption such as "Eating" if an elderly person is eating in the video. Furthermore, the caption insertion unit can also insert an annotation caption such as "Resting" if an elderly person is resting in the video. Step 4: The anomaly detection unit detects anomalies based on the annotation captions inserted by the caption insertion unit. For example, the anomaly detection unit detects an anomaly if an elderly person falls in the video. The anomaly detection unit can also detect an anomaly if an elderly person remains motionless for a long period of time in the video. Furthermore, the anomaly detection unit can also detect an anomaly if an elderly person makes suspicious movements in the video. Step 5: The alert unit issues an alert based on the anomaly detected by the anomaly detection unit. For example, the alert unit will issue an alert if an elderly person falls. The alert unit can also issue an alert if an elderly person remains motionless for a long period of time. Furthermore, the alert unit can also issue an alert if an elderly person performs any suspicious actions.

[0068] (Example of form 2) The monitoring system according to an embodiment of the present invention is a system that eliminates privacy concerns and the hassle of reviewing video in existing monitoring cameras. This monitoring system reduces the effort of monitoring and reviewing video by using a generation AI to distort the video captured by the camera, analyzing what is happening in the video, and inserting annotation captions. In addition, if an accident or other incident is suspected, an alert notification is immediately issued. For example, video captured by the camera is input to the generation AI. The generation AI analyzes the video and converts it into a visually distorted representation. For example, video of an elderly person walking in a room can be converted into video that looks like an animated character walking. This allows for understanding the content of the video while protecting privacy. Next, the generation AI analyzes what is happening in the video. The generation AI recognizes the actions and situations in the video and inserts annotation captions based on that. For example, if an elderly person is eating, an annotation caption such as "eating" can be displayed. This significantly reduces the effort of reviewing the video. Furthermore, if an accident or other incident is suspected, an alert notification is immediately issued. The generation AI issues an alert when it detects abnormal actions or situations in the video. For example, if an elderly person falls, an alert notification can be sent immediately, prompting a quick response. This system eliminates privacy concerns associated with monitoring cameras and reduces the effort required to review video footage. For instance, it can provide a secure environment for family members and caregivers to monitor the elderly. It also contributes to improved safety by enabling a quick response to emergencies such as accidents. In this way, the monitoring system can reduce the effort required to review video footage while protecting privacy.

[0069] The monitoring system according to this embodiment comprises an acquisition unit, an analysis unit, a text overlay insertion unit, an anomaly detection unit, and an alert issuing unit. The acquisition unit acquires video footage captured by a camera. The acquisition unit can, for example, acquire video footage captured by a camera in real time. The acquisition unit can also acquire video footage from the camera periodically. Furthermore, the acquisition unit can acquire video footage from the camera in streaming format. For example, the acquisition unit can acquire video footage from the camera at 30 frames per second. The acquisition unit can also acquire video footage from the camera every hour. Furthermore, the acquisition unit can acquire video footage from the camera in live streaming format. The analysis unit analyzes the video footage acquired by the acquisition unit and converts it into a visually stylized representation. For example, the analysis unit can stylize the video in an anime style. Furthermore, the analysis unit can stylize the video in a manga style. Furthermore, the analysis unit can silhouette the video. For example, the analysis unit can stylize the video in an anime character style. The analysis unit can also deform the video into a comic strip-like format. Furthermore, the analysis unit can silhouette the video to protect privacy. The caption insertion unit inserts annotation captions based on the actions and situations in the video analyzed by the analysis unit. For example, if an elderly person is walking in the video, the caption insertion unit will insert an annotation caption such as "Walking." Also, if an elderly person is eating in the video, the caption insertion unit can insert an annotation caption such as "Eating." Furthermore, if an elderly person is resting in the video, the caption insertion unit can insert an annotation caption such as "Resting." For example, if an elderly person is walking in the video, the caption insertion unit will insert an annotation caption such as "Walking." Also, if an elderly person is eating in the video, the caption insertion unit can insert an annotation caption such as "Eating." Furthermore, if an elderly person is resting in the video, the caption insertion unit can insert an annotation caption such as "Resting." The anomaly detection unit detects anomalies based on the annotation captions inserted by the caption insertion unit. The anomaly detection unit detects an anomaly, for example, if an elderly person falls in the video.Furthermore, the anomaly detection unit can also detect anomalies if an elderly person remains motionless for a long period of time in the video. In addition, the anomaly detection unit can also detect anomalies if an elderly person performs suspicious actions in the video. For example, the anomaly detection unit detects an anomaly if an elderly person falls in the video. Furthermore, the anomaly detection unit can also detect an anomaly if an elderly person remains motionless for a long period of time in the video. Furthermore, the anomaly detection unit can also detect an anomaly if an elderly person performs suspicious actions in the video. The alert issuing unit issues an alert based on the anomaly detected by the anomaly detection unit. For example, the alert issuing unit issues an alert if an elderly person falls. Furthermore, the alert issuing unit can also issue an alert if an elderly person remains motionless for a long period of time. Furthermore, the alert issuing unit can also issue an alert if an elderly person performs suspicious actions. As a result, the monitoring system according to this embodiment can reduce the effort required for video confirmation while protecting privacy.

[0070] The acquisition unit acquires video footage captured by cameras. For example, the acquisition unit can acquire video footage in real time. It can also acquire camera footage periodically. Furthermore, the acquisition unit can acquire camera footage in streaming format. For example, the acquisition unit can acquire camera footage at 30 frames per second. It can also acquire camera footage every hour. Additionally, the acquisition unit can acquire camera footage in live streaming format. Specifically, the acquisition unit can acquire video from multiple cameras simultaneously, enabling wide-area surveillance. Various types of cameras are used, including fixed cameras and pan-tilt-zoom (PTZ) cameras. Fixed cameras constantly monitor a specific area, while PTZ cameras allow for remotely changing the field of view. This allows for the acquisition of detailed footage when specific events or anomalies occur. Furthermore, the acquisition unit has the ability to dynamically adjust the video resolution and frame rate, optimizing network bandwidth and storage capacity. For example, it can acquire video at a low resolution under normal circumstances and switch to high resolution when an anomaly is detected. This enables efficient data management. Furthermore, the acquisition unit ensures data security by encrypting and transmitting video data. This enhances privacy protection and prevents unauthorized access from external sources.

[0071] The analysis unit analyzes the video acquired by the acquisition unit and converts it into a visually stylized representation. For example, the analysis unit can stylize the video in an anime style. It can also stylize the video in a manga style. Furthermore, the analysis unit can silhouette the video. For example, the analysis unit can stylize the video in an anime character style. It can also stylize the video in a manga panel style. Furthermore, the analysis unit can silhouette the video to protect privacy. Specifically, the analysis unit uses AI to recognize people and objects in the video and converts them into a specific style. For example, it utilizes image recognition technology using deep learning to analyze people's movements and expressions and convert them into an anime or manga style. When silhouetting, it extracts the outline of the person and omits internal details to protect privacy. Furthermore, the analysis unit has a function to automatically blur the background in the video, hiding information other than people. This further enhances privacy protection. In addition, the analysis unit can update the video analysis results in real time and respond to dynamic changes. For example, if a person moves or a new object appears in the image, this is immediately reflected in the stylized representation. This allows for the constant provision of up-to-date information.

[0072] The caption insertion unit inserts annotation captions based on the actions and situations in the video analyzed by the analysis unit. For example, if an elderly person is walking in the video, the caption insertion unit will insert an annotation caption such as "Walking." It can also insert an annotation caption such as "Eating" if an elderly person is eating in the video. Furthermore, it can insert an annotation caption such as "Resting" if an elderly person is resting in the video. Specifically, the caption insertion unit uses AI to automatically recognize actions and situations in the video and generates corresponding captions. For example, it uses an action recognition algorithm to detect actions such as walking, eating, and resting by elderly people and inserts corresponding captions. Furthermore, the content of the on-screen text is designed to be customizable by the user, allowing for the setting of messages tailored to specific situations. In addition, the on-screen text insertion unit has the functionality to dynamically adjust the display position, font, and color of the text, optimizing the visibility of the video. For example, it adjusts the position of the text so as not to obscure important parts of the video, ensuring visibility. It also adjusts the display time of the text to provide necessary information at the appropriate time. As a result, the on-screen text insertion unit can provide appropriate annotations according to the actions and situations in the video, making the information easy for viewers to understand.

[0073] The anomaly detection unit detects anomalies based on annotation captions inserted by the caption insertion unit. For example, the anomaly detection unit detects an anomaly if an elderly person falls in the video. The anomaly detection unit can also detect an anomaly if an elderly person remains motionless for a long period of time in the video. Furthermore, the anomaly detection unit can also detect an anomaly if an elderly person performs suspicious actions in the video. Specifically, the anomaly detection unit uses AI to analyze the motion patterns in the video and detect unusual actions or situations. For example, it uses a fall detection algorithm to detect an anomaly if an elderly person suddenly falls. It also uses a motion stop detection algorithm to detect an anomaly if an elderly person remains motionless for a certain period of time or longer. Furthermore, it uses a suspicious action detection algorithm to detect an anomaly if an elderly person performs unusual actions. This allows the anomaly detection unit to detect anomalies in real time and respond quickly. Furthermore, the anomaly detection unit can learn patterns of anomalies based on past data and improve detection accuracy. For example, it can improve the accuracy of fall detection by learning from past fall incidents. The anomaly detection unit also has a function to immediately issue an alert when an anomaly is detected, supporting a rapid response. As a result, the anomaly detection unit can detect anomalies with high accuracy and enable a quick and appropriate response.

[0074] The alert unit issues alerts based on anomalies detected by the anomaly detection unit. For example, the alert unit issues an alert if an elderly person falls. The alert unit can also issue an alert if an elderly person remains motionless for an extended period. Furthermore, the alert unit can issue an alert if an elderly person exhibits suspicious behavior. Specifically, when an anomaly is detected, the alert unit immediately notifies the relevant parties. Notification methods include smartphone push notifications, SMS, email, and voice calls. For example, it notifies family members and caregivers in real time that an anomaly has occurred, prompting a quick response. The alert unit can also customize the content of notifications, sending appropriate messages according to the type and urgency of the anomaly. Furthermore, the alert unit has a function to record notification history for later review. This allows for understanding past incidents and using that information to inform future countermeasures. Furthermore, the alert unit has a function to automatically notify emergency contacts when an incident occurs, supporting a rapid response. For example, it automatically notifies registered medical institutions and police to encourage prompt rescue operations. In this way, the alert unit supports a swift and appropriate response when an incident occurs, ensuring the safety of the elderly.

[0075] The analysis unit can convert video into a visually stylized representation. For example, the analysis unit can stylize video in an anime style. The analysis unit can also stylize video in a manga style. The analysis unit can also silhouette video. By stylizing the video, the content of the video can be understood while protecting privacy. Visually stylized representations include, but are not limited to, mosaic processing, silhouettes, and color changes. Some or all of the above processing in the analysis unit is performed using a generation AI. For example, the analysis unit inputs video captured by a camera into the generation AI, which analyzes the video and converts it into a stylized representation.

[0076] The caption insertion unit can insert annotation captions based on actions and situations within the video. For example, if an elderly person is walking in the video, the caption insertion unit can insert an annotation caption such as "Walking." For example, if an elderly person is eating in the video, the caption insertion unit can also insert an annotation caption such as "Eating." For example, if an elderly person is resting in the video, the caption insertion unit can also insert an annotation caption such as "Resting." This reduces the effort required to review the video by inserting annotation captions. Actions and situations include, but are not limited to, actions such as walking, running, sitting, and standing, and situations such as crowding and silence. Some or all of the above processing in the caption insertion unit is performed using AI. For example, the caption insertion unit inputs actions and situations in the video into the AI, and the AI ​​analyzes the actions and situations to generate annotation captions.

[0077] The anomaly detection unit can detect abnormal movements or situations within the video. For example, the anomaly detection unit can detect an anomaly if an elderly person falls in the video. The anomaly detection unit can also detect an anomaly if an elderly person remains motionless for a long period of time in the video. The anomaly detection unit can also detect an anomaly if an elderly person performs suspicious movements in the video. This enables a rapid response by detecting abnormal movements or situations. Abnormal movements or situations include, but are not limited to, suspicious movements, unusual sounds, and specific patterns. Some or all of the above processing in the anomaly detection unit is performed using AI. For example, the anomaly detection unit inputs the movements and situations in the video into the AI, which analyzes the movements and situations to detect anomalies.

[0078] The alert issuing unit can issue an alert when an anomaly is detected. For example, the alert issuing unit will issue an alert if an elderly person falls. The alert issuing unit can also issue an alert if an elderly person remains motionless for a long period of time. The alert issuing unit can also issue an alert if an elderly person performs suspicious actions. This enables a rapid response by issuing an alert when an anomaly is detected. Alerts include, but are not limited to, voice alerts, visual alerts, and notification messages. Some or all of the above processing in the alert issuing unit is performed using AI. For example, the alert issuing unit inputs the anomaly detected by the anomaly detection unit to the AI, which analyzes the anomaly and generates an alert.

[0079] The acquisition unit can estimate the user's emotions and adjust the timing of video acquisition based on the estimated emotions. For example, if the user is relaxed, the acquisition unit will periodically acquire video. If the user is stressed, for example, the acquisition unit can reduce the frequency of video acquisition and acquire only when necessary. If the user is excited, for example, the acquisition unit can continue to acquire video in real time. This allows for more appropriate video acquisition by adjusting the timing of video acquisition according to the user's emotions. User emotions include, but are not limited to, facial recognition, voice analysis, and biometric data. Some or all of the above processing in the acquisition unit is performed using AI. For example, the acquisition unit inputs the user's facial expression data into the AI, which analyzes the facial expression data to estimate emotions.

[0080] The acquisition unit can select the optimal acquisition method based on the user's activity pattern when acquiring video. For example, if the user is active during the day, the acquisition unit will acquire video during the daytime. For example, if the user is active at night, the acquisition unit can also acquire video at night. For example, if the user's activity pattern is irregular, the acquisition unit can use AI to learn and adjust the optimal acquisition timing. This enables efficient video acquisition by selecting the optimal acquisition method based on the user's activity pattern. The optimal acquisition method includes, but is not limited to, camera angle, resolution, and frame rate. Some or all of the above processing in the acquisition unit is performed using AI. For example, the acquisition unit inputs user activity data into the AI, and the AI ​​analyzes the activity data to select the optimal acquisition method.

[0081] The acquisition unit can filter video footage based on the user's current environment and circumstances. For example, if the user is indoors, the acquisition unit will prioritize acquiring indoor video. If the user is outdoors, the acquisition unit can also prioritize acquiring outdoor video. The acquisition unit can also automatically adjust the filtering using AI if the user's environment changes. This enables the acquisition of appropriate video footage by filtering based on the user's current environment and circumstances. Current environment and circumstances include, but are not limited to, indoor / outdoor distinctions, weather, and time of day. Some or all of the above processing in the acquisition unit is performed using AI. For example, the acquisition unit inputs the user's environmental data into the AI, which then analyzes the environmental data and performs filtering.

[0082] The acquisition unit can estimate the user's emotions and determine the priority of the videos to acquire based on the estimated user emotions. For example, if the user is relaxed, the acquisition unit will prioritize acquiring normal videos. If the user is stressed, the acquisition unit may also prioritize acquiring relaxing videos. If the user is excited, the acquisition unit may also prioritize acquiring videos that alleviate excitement. By prioritizing videos according to the user's emotions, more appropriate video acquisition becomes possible. User emotions include, but are not limited to, facial recognition, voice analysis, and biometric data. Some or all of the above processing in the acquisition unit is performed using AI. For example, the acquisition unit inputs user emotion data into the AI, and the AI ​​analyzes the emotion data to determine the priority of the videos.

[0083] The acquisition unit can prioritize acquiring highly relevant video footage by considering the user's geographical location information when acquiring video. For example, if the user is at home, the acquisition unit will prioritize acquiring video footage of the area around the user's home. If the user is out, the acquisition unit can also prioritize acquiring video footage of the location they are at. If the user is traveling, the acquisition unit can also prioritize acquiring video footage of their travel destination. This allows for the acquisition of appropriate video footage by prioritizing highly relevant video footage while considering the user's geographical location information. Geographical location information includes, but is not limited to, GPS data and Wi-Fi location information. Some or all of the above processing in the acquisition unit is performed using AI. For example, the acquisition unit inputs the user's location data into the AI, which analyzes the location data and selects highly relevant video footage.

[0084] The acquisition unit can analyze the user's social media activity when acquiring video and acquire relevant videos. For example, the acquisition unit can prioritize acquiring videos of places the user has shared on social media. The acquisition unit can also prioritize acquiring videos of places the user has shown interest in on social media. For example, the acquisition unit can acquire videos of relevant events from the user's social media activity. This enables appropriate video acquisition by analyzing the user's social media activity and acquiring relevant videos. Social media activity includes, but is not limited to, posts, the number of likes, and comments. Some or all of the above processing in the acquisition unit is performed using AI. For example, the acquisition unit inputs the user's social media data into the AI, which analyzes the data and selects relevant videos.

[0085] The analysis unit can estimate the user's emotions and adjust the deformation representation method based on the estimated user emotions. For example, if the user is relaxed, the analysis unit may use a deformation representation with soft colors. If the user is stressed, the analysis unit may also use a deformation representation with calm colors. If the user is excited, the analysis unit may also use a deformation representation with bright colors. By adjusting the deformation representation method according to the user's emotions, a more appropriate deformation representation becomes possible. User emotions include, but are not limited to, facial recognition, voice analysis, and biometric data. Some or all of the above processing in the analysis unit is performed using a generative AI. For example, the analysis unit inputs the user's emotion data into the generative AI, which analyzes the emotion data and adjusts the deformation representation method.

[0086] The analysis unit can adjust the level of detail of the deformation based on specific actions or situations when analyzing video. For example, if an elderly person is walking, the analysis unit will use a deformation that emphasizes walking. For example, if an elderly person is eating, the analysis unit can also use a deformation that emphasizes eating. For example, if an elderly person is resting, the analysis unit can also use a deformation that emphasizes a relaxed posture. By adjusting the level of detail of the deformation based on specific actions or situations, a more appropriate deformation representation becomes possible. Specific actions and situations include, but are not limited to, actions such as walking, running, sitting, and standing, and situations such as crowding and silence. Some or all of the above processing in the analysis unit is performed using a generation AI. For example, the analysis unit inputs the actions and situations in the video into the generation AI, which analyzes the actions and situations and adjusts the level of detail of the deformation.

[0087] The analysis unit can apply different deformation algorithms depending on the scenario when analyzing video. For example, the analysis unit uses a simple deformation algorithm in everyday life scenarios. For example, the analysis unit can also use a detailed deformation algorithm in emergency scenarios. For example, the analysis unit can also use a deformation algorithm tailored to a specific event scenario. By applying different deformation algorithms according to different scenarios, more appropriate deformation representation becomes possible. Different scenarios include, but are not limited to, emergency situations, normal situations, and specific events. Some or all of the above processing in the analysis unit is performed using a generation AI. For example, the analysis unit inputs scenario data into the generation AI, which analyzes the scenario data and selects an appropriate deformation algorithm.

[0088] The analysis unit can estimate the user's emotions and adjust the deformation style based on the estimated emotions. For example, if the user is relaxed, the analysis unit will use a soft-touch deformation style. If the user is stressed, the analysis unit may also use a simple and calm deformation style. If the user is excited, the analysis unit may also use a bright and lively deformation style. By adjusting the deformation style according to the user's emotions, more appropriate deformation representation becomes possible. User emotions include, but are not limited to, facial recognition, voice analysis, and biometric data. Some or all of the above processing in the analysis unit is performed using a generative AI. For example, the analysis unit inputs user emotion data into the generative AI, which analyzes the emotion data and adjusts the deformation style.

[0089] The analysis unit can determine the priority of distortion based on the time the video was shot when analyzing the video. For example, the analysis unit may prioritize bright distortion for video shot during the day. For example, the analysis unit may also prioritize muted distortion for video shot at night. For example, the analysis unit may also prioritize distortion appropriate for an event for video shot during a specific event. By determining the priority of distortion based on the time the video was shot, appropriate distortion is possible. The shooting time includes, but is not limited to, the date and time, season, and specific events. Some or all of the above processing in the analysis unit is performed using a generation AI. For example, the analysis unit inputs the video shooting time data into the generation AI, and the generation AI analyzes the shooting time data to determine the priority of distortion.

[0090] The analysis unit can adjust the order of deformation based on the relevance of the video during video analysis. For example, the analysis unit may prioritize deformation of video containing important actions. The analysis unit may also postpone deformation of video containing everyday actions. The analysis unit may also prioritize deformation of video containing emergencies. By adjusting the order of deformation based on the relevance of the video, appropriate deformation representation becomes possible. Video relevance includes, but is not limited to, similarity of content and temporal continuity. Some or all of the above processing in the analysis unit is performed using a generation AI. For example, the analysis unit inputs video relevance data into the generation AI, which analyzes the relevance data and adjusts the order of deformation.

[0091] The text overlay unit can estimate the user's emotions and adjust the way the text is displayed based on those emotions. For example, if the user is relaxed, the text overlay unit will use soft fonts and colors. If the user is stressed, the text overlay unit may use simple, highly visible text. If the user is excited, the text overlay unit may use bright, eye-catching text. By adjusting the way the text is displayed according to the user's emotions, more appropriate text display becomes possible. User emotions include, but are not limited to, facial recognition, voice analysis, and biometric data. Some or all of the above processing in the text overlay unit is performed using AI. For example, the text overlay unit inputs user emotion data into the AI, which analyzes the emotion data and adjusts the way the text is displayed.

[0092] The text overlay insertion unit can adjust the level of detail of the text overlay based on the importance of the action or situation when inserting text. For example, the text overlay insertion unit will insert detailed text when an important action is being performed. For example, the text overlay insertion unit can also insert concise text when a routine action is being performed. For example, the text overlay insertion unit can also insert detailed and rapid text when an emergency occurs. This allows for appropriate text display by adjusting the level of detail of the text overlay based on the importance of the action or situation. The importance of the action or situation includes, but is not limited to, frequency, impact, and urgency. Some or all of the above processing in the text overlay insertion unit is performed using AI. For example, the text overlay insertion unit inputs action and situation data into the AI, which analyzes the data and adjusts the level of detail of the text overlay.

[0093] The text overlay insertion unit can apply different text overlay insertion algorithms depending on the different scenarios during text overlay insertion. For example, in everyday life scenarios, the text overlay insertion unit uses a simple text overlay insertion algorithm. In emergency scenarios, for example, the text overlay insertion unit can also use a more detailed text overlay insertion algorithm. In scenarios of specific events, for example, the text overlay insertion unit can also use an event-specific text overlay insertion algorithm. This enables appropriate text overlay insertion depending on the different scenarios. Different scenarios include, but are not limited to, emergency situations, normal situations, and specific events. Some or all of the above processing in the text overlay insertion unit is performed using AI. For example, the text overlay insertion unit inputs scenario data into the AI, which analyzes the scenario data and selects an appropriate text overlay insertion algorithm.

[0094] The text overlay unit can estimate the user's emotions and adjust the length of the text overlay based on the estimated emotions. For example, if the user is relaxed, the text overlay unit will use longer text. If the user is stressed, for example, the text overlay unit can also use shorter, more concise text. If the user is excited, for example, the text overlay unit can also use visually stimulating text. By adjusting the length of the text overlay according to the user's emotions, more appropriate text display becomes possible. User emotions include, but are not limited to, facial recognition, voice analysis, and biometric data. Some or all of the above processing in the text overlay unit is performed using AI. For example, the text overlay unit inputs user emotion data into the AI, which analyzes the emotion data and adjusts the length of the text overlay.

[0095] The text overlay insertion unit can determine the priority of text overlays based on the timing of the action or situation. For example, the text overlay insertion unit can insert a text overlay immediately after an important action has occurred. It can also insert a text overlay after a routine action has been performed. For example, it can insert a text overlay immediately in the event of an emergency. This allows for appropriate text overlay display by determining the priority of text overlays based on the timing of the action or situation. The timing of the action or situation includes, but is not limited to, the date and time, season, or specific event. Some or all of the above processing in the text overlay insertion unit is performed using AI. For example, the text overlay insertion unit inputs data on the timing of the action or situation into the AI, which analyzes the data to determine the priority of text overlays.

[0096] The text overlay insertion unit can adjust the order of text overlays based on the relevance of actions and situations when inserting them. For example, the text overlay insertion unit will prioritize inserting text overlays related to important actions. For example, the text overlay insertion unit can postpone inserting text overlays related to routine actions. For example, the text overlay insertion unit can prioritize inserting text overlays related to emergencies. This allows for appropriate text overlay display by adjusting the order of text overlays based on the relevance of actions and situations. The relevance of actions and situations includes, but is not limited to, similarity of content and temporal continuity. Some or all of the above processing in the text overlay insertion unit is performed using AI. For example, the text overlay insertion unit inputs data on the relevance of actions and situations into the AI, which analyzes the data and adjusts the order of the text overlays.

[0097] The anomaly detection unit can estimate the user's emotions and adjust the anomaly detection criteria based on the estimated emotions. For example, if the user is relaxed, the anomaly detection unit uses normal anomaly detection criteria. For example, if the user is stressed, the anomaly detection unit can tighten the anomaly detection criteria. For example, if the user is excited, the anomaly detection unit can loosen the anomaly detection criteria. This allows for more appropriate anomaly detection by adjusting the criteria according to the user's emotions. User emotions include, but are not limited to, facial recognition, voice analysis, and biometric data. Some or all of the above processing in the anomaly detection unit is performed using AI. For example, the anomaly detection unit inputs user emotion data into the AI, which analyzes the emotion data and adjusts the anomaly detection criteria.

[0098] The anomaly detection unit can optimize its anomaly detection algorithm by referring to past anomaly data when an anomaly is detected. For example, the anomaly detection unit adjusts the anomaly detection algorithm based on past anomaly data. The anomaly detection unit can also learn specific patterns from past anomaly data and reflect them in the anomaly detection algorithm. For example, the anomaly detection unit can analyze past anomaly data to improve the accuracy of the anomaly detection algorithm. As a result, the accuracy of anomaly detection is improved by optimizing the anomaly detection algorithm by referring to past anomaly data. Past anomaly data includes, but is not limited to, past anomaly cases, anomaly frequency, and anomaly impact. Some or all of the above processing in the anomaly detection unit is performed using AI. For example, the anomaly detection unit inputs past anomaly data into the AI, and the AI ​​analyzes the data to optimize the anomaly detection algorithm.

[0099] The anomaly detection unit can apply different anomaly detection methods depending on the different scenarios when an anomaly is detected. For example, in everyday life scenarios, the anomaly detection unit uses a simple anomaly detection method. In emergency scenarios, for example, the anomaly detection unit can also use a more detailed anomaly detection method. In scenarios of specific events, for example, the anomaly detection unit can also use an event-specific anomaly detection method. This improves the accuracy of anomaly detection by applying the appropriate anomaly detection method according to different scenarios. Different scenarios include, but are not limited to, emergency situations, normal situations, and specific events. Some or all of the above processing in the anomaly detection unit is performed using AI. For example, the anomaly detection unit inputs scenario data into the AI, which analyzes the scenario data and selects an appropriate anomaly detection method.

[0100] The anomaly detection unit can estimate the user's emotions and adjust the order in which anomaly detection results are displayed based on the estimated user emotions. For example, if the user is relaxed, the anomaly detection unit will display the anomaly detection results in the normal order. For example, if the user is stressed, the anomaly detection unit can also prioritize displaying important anomaly detection results. For example, if the user is excited, the anomaly detection unit can also postpone displaying less urgent anomaly detection results. By adjusting the order in which anomaly detection results are displayed according to the user's emotions, it becomes possible to display more appropriate anomaly detection results. User emotions include, but are not limited to, facial recognition, voice analysis, and biometric data. Some or all of the above processing in the anomaly detection unit is performed using AI. For example, the anomaly detection unit inputs the user's emotion data into the AI, which analyzes the emotion data and adjusts the order in which anomaly detection results are displayed.

[0101] The anomaly detection unit can perform anomaly detection while considering the geographical distribution of the video. For example, the anomaly detection unit can perform anomaly detection based on the location where the video was shot. The anomaly detection unit can also analyze the geographical distribution of the video and adjust the anomaly detection algorithm. For example, the anomaly detection unit can prioritize anomaly detection in a specific area based on the geographical distribution of the video. This makes it possible to perform appropriate anomaly detection by considering the geographical distribution of the video. Geographical distribution includes, but is not limited to, the anomaly occurrence rate for each region and geographical characteristics. Some or all of the above processing in the anomaly detection unit is performed using AI. For example, the anomaly detection unit inputs geographical distribution data into the AI, and the AI ​​analyzes the data and performs anomaly detection.

[0102] The anomaly detection unit can improve the accuracy of anomaly detection by referring to relevant literature in the video when an anomaly is detected. For example, the anomaly detection unit adjusts the anomaly detection algorithm based on the relevant literature in the video. The anomaly detection unit can also learn specific patterns from the relevant literature in the video and reflect them in the anomaly detection algorithm. For example, the anomaly detection unit can analyze the relevant literature in the video to improve the accuracy of the anomaly detection algorithm. By improving the accuracy of anomaly detection by referring to relevant literature in the video, more accurate anomaly detection becomes possible. Relevant literature includes, but is not limited to, academic papers, technical reports, and patent documents. Some or all of the above processing in the anomaly detection unit is performed using AI. For example, the anomaly detection unit inputs the relevant literature data into the AI, and the AI ​​analyzes the data to adjust the anomaly detection algorithm.

[0103] The alerting unit can estimate the user's emotions and adjust the alerting method based on the estimated emotions. For example, if the user is relaxed, the alerting unit will use the normal alerting method. If the user is stressed, the alerting unit may prioritize sending high-priority alerts. If the user is excited, the alerting unit may postpone sending low-priority alerts. By adjusting the alerting method according to the user's emotions, more appropriate alerts can be delivered. User emotions include, but are not limited to, facial recognition, voice analysis, and biometric data. Some or all of the above processing in the alerting unit is performed using AI. For example, the alerting unit inputs user emotion data into the AI, which analyzes the emotion data and adjusts the alerting method.

[0104] The alert issuing unit can optimize its alert issuing algorithm by referring to past alert data when issuing an alert. For example, the alert issuing unit adjusts the alert issuing algorithm based on past alert data. The alert issuing unit can also learn specific patterns from past alert data and reflect them in the alert issuing algorithm. For example, the alert issuing unit can analyze past alert data to improve the accuracy of the alert issuing algorithm. This improves the accuracy of alert issuing by optimizing the alert issuing algorithm by referring to past alert data. Past alert data includes, but is not limited to, past alert cases, alert frequency, and alert impact. Some or all of the above processing in the alert issuing unit is performed using AI. For example, the alert issuing unit inputs past alert data into the AI, and the AI ​​analyzes the data to optimize the alert issuing algorithm.

[0105] The alert issuing unit can apply different alert issuing methods depending on the different scenarios when issuing an alert. For example, in everyday life scenarios, the alert issuing unit uses a simple alert issuing method. In emergency scenarios, for example, the alert issuing unit can also use a detailed alert issuing method. In scenarios of specific events, for example, the alert issuing unit can also use an event-specific alert issuing method. This improves the accuracy of alert issuing by applying the appropriate alert issuing method according to different scenarios. Different scenarios include, but are not limited to, emergency situations, normal situations, and specific events. Some or all of the above processing in the alert issuing unit is performed using AI. For example, the alert issuing unit inputs scenario data into the AI, and the AI ​​analyzes the scenario data to select the appropriate alert issuing method.

[0106] The alerting unit can estimate the user's emotions and determine the priority of alerts based on those emotions. For example, if the user is relaxed, the alerting unit will use the normal alert priority. If the user is stressed, the alerting unit may prioritize high-urgency alerts. If the user is excited, the alerting unit may postpone low-urgency alerts. This allows for more appropriate alert delivery by prioritizing alerts according to the user's emotions. User emotions include, but are not limited to, facial recognition, voice analysis, and biometric data. Some or all of the above processing in the alerting unit is performed using AI. For example, the alerting unit inputs user emotion data into the AI, which analyzes the emotion data to determine the alert priority.

[0107] The alert issuing unit can adjust the timing of alert issuance based on the video recording date. For example, the alert issuing unit uses the normal alert timing for video recorded during the day. For example, the alert issuing unit can prioritize issuing high-urgency alerts for video recorded at night. For example, the alert issuing unit can use an alert timing tailored to a specific event for video recorded during a particular event. By adjusting the alert issuance timing based on the video recording date, appropriate alerts can be issued. The recording date includes, but is not limited to, the date and time, season, and specific events. Some or all of the above processing in the alert issuing unit is performed using AI. For example, the alert issuing unit inputs recording date data into the AI, which analyzes the data and adjusts the alert issuance timing.

[0108] The alert issuing unit can adjust the order in which alerts are issued based on the relevance of the video footage. For example, the alert issuing unit will issue an alert preferentially if it relates to an important action. For example, the alert issuing unit can also issue an alert later if it relates to a routine action. For example, the alert issuing unit can issue an alert with the highest priority if it relates to an emergency. This allows for appropriate alert issuance by adjusting the order in which alerts are issued based on the relevance of the video footage. Video relevance includes, but is not limited to, similarity of content and temporal continuity. Some or all of the above processing in the alert issuing unit is performed using AI. For example, the alert issuing unit inputs relevance data into the AI, and the AI ​​analyzes the data to adjust the order in which alerts are issued.

[0109] The system according to the embodiment is not limited to the example described above, and various modifications are possible, for example, as follows.

[0110] The acquisition unit can estimate the user's emotions and adjust the timing of video acquisition based on the estimated emotions. For example, if the user is relaxed, the acquisition unit will periodically acquire video. If the user is stressed, for example, the acquisition unit can reduce the frequency of video acquisition and acquire only when necessary. If the user is excited, for example, the acquisition unit can continue to acquire video in real time. This allows for more appropriate video acquisition by adjusting the timing of video acquisition according to the user's emotions. User emotions include, but are not limited to, facial recognition, voice analysis, and biometric data. Some or all of the above processing in the acquisition unit is performed using AI. For example, the acquisition unit inputs the user's facial expression data into the AI, which analyzes the facial expression data to estimate emotions.

[0111] The acquisition unit can select the optimal acquisition method based on the user's activity pattern when acquiring video. For example, if the user is active during the day, the acquisition unit will acquire video during the daytime. For example, if the user is active at night, the acquisition unit can also acquire video at night. For example, if the user's activity pattern is irregular, the acquisition unit can use AI to learn and adjust the optimal acquisition timing. This enables efficient video acquisition by selecting the optimal acquisition method based on the user's activity pattern. The optimal acquisition method includes, but is not limited to, camera angle, resolution, and frame rate. Some or all of the above processing in the acquisition unit is performed using AI. For example, the acquisition unit inputs user activity data into the AI, and the AI ​​analyzes the activity data to select the optimal acquisition method.

[0112] The acquisition unit can filter video footage based on the user's current environment and circumstances. For example, if the user is indoors, the acquisition unit will prioritize acquiring indoor video. If the user is outdoors, the acquisition unit can also prioritize acquiring outdoor video. The acquisition unit can also automatically adjust the filtering using AI if the user's environment changes. This enables the acquisition of appropriate video footage by filtering based on the user's current environment and circumstances. Current environment and circumstances include, but are not limited to, indoor / outdoor distinctions, weather, and time of day. Some or all of the above processing in the acquisition unit is performed using AI. For example, the acquisition unit inputs the user's environmental data into the AI, which then analyzes the environmental data and performs filtering.

[0113] The acquisition unit can estimate the user's emotions and determine the priority of the videos to acquire based on the estimated user emotions. For example, if the user is relaxed, the acquisition unit will prioritize acquiring normal videos. If the user is stressed, the acquisition unit may also prioritize acquiring relaxing videos. If the user is excited, the acquisition unit may also prioritize acquiring videos that alleviate excitement. By prioritizing videos according to the user's emotions, more appropriate video acquisition becomes possible. User emotions include, but are not limited to, facial recognition, voice analysis, and biometric data. Some or all of the above processing in the acquisition unit is performed using AI. For example, the acquisition unit inputs user emotion data into the AI, and the AI ​​analyzes the emotion data to determine the priority of the videos.

[0114] The analysis unit can estimate the user's emotions and adjust the deformation representation method based on the estimated user emotions. For example, if the user is relaxed, the analysis unit may use a deformation representation with soft colors. If the user is stressed, the analysis unit may also use a deformation representation with calm colors. If the user is excited, the analysis unit may also use a deformation representation with bright colors. By adjusting the deformation representation method according to the user's emotions, a more appropriate deformation representation becomes possible. User emotions include, but are not limited to, facial recognition, voice analysis, and biometric data. Some or all of the above processing in the analysis unit is performed using a generative AI. For example, the analysis unit inputs the user's emotion data into the generative AI, which analyzes the emotion data and adjusts the deformation representation method.

[0115] The analysis unit can adjust the level of detail of the deformation based on specific actions or situations when analyzing video. For example, if an elderly person is walking, the analysis unit will use a deformation that emphasizes walking. For example, if an elderly person is eating, the analysis unit can also use a deformation that emphasizes eating. For example, if an elderly person is resting, the analysis unit can also use a deformation that emphasizes a relaxed posture. By adjusting the level of detail of the deformation based on specific actions or situations, a more appropriate deformation representation becomes possible. Specific actions and situations include, but are not limited to, actions such as walking, running, sitting, and standing, and situations such as crowding and silence. Some or all of the above processing in the analysis unit is performed using a generation AI. For example, the analysis unit inputs the actions and situations in the video into the generation AI, which analyzes the actions and situations and adjusts the level of detail of the deformation.

[0116] The analysis unit can apply different deformation algorithms depending on the scenario when analyzing video. For example, the analysis unit uses a simple deformation algorithm in everyday life scenarios. For example, the analysis unit can also use a detailed deformation algorithm in emergency scenarios. For example, the analysis unit can also use a deformation algorithm tailored to a specific event scenario. By applying different deformation algorithms according to different scenarios, more appropriate deformation representation becomes possible. Different scenarios include, but are not limited to, emergency situations, normal situations, and specific events. Some or all of the above processing in the analysis unit is performed using a generation AI. For example, the analysis unit inputs scenario data into the generation AI, which analyzes the scenario data and selects an appropriate deformation algorithm.

[0117] The text overlay unit can estimate the user's emotions and adjust the way the text is displayed based on those emotions. For example, if the user is relaxed, the text overlay unit will use soft fonts and colors. If the user is stressed, the text overlay unit may use simple, highly visible text. If the user is excited, the text overlay unit may use bright, eye-catching text. By adjusting the way the text is displayed according to the user's emotions, more appropriate text display becomes possible. User emotions include, but are not limited to, facial recognition, voice analysis, and biometric data. Some or all of the above processing in the text overlay unit is performed using AI. For example, the text overlay unit inputs user emotion data into the AI, which analyzes the emotion data and adjusts the way the text is displayed.

[0118] The text overlay insertion unit can adjust the level of detail of the text overlay based on the importance of the action or situation when inserting text. For example, the text overlay insertion unit will insert detailed text when an important action is being performed. For example, the text overlay insertion unit can also insert concise text when a routine action is being performed. For example, the text overlay insertion unit can also insert detailed and rapid text when an emergency occurs. This allows for appropriate text display by adjusting the level of detail of the text overlay based on the importance of the action or situation. The importance of the action or situation includes, but is not limited to, frequency, impact, and urgency. Some or all of the above processing in the text overlay insertion unit is performed using AI. For example, the text overlay insertion unit inputs action and situation data into the AI, which analyzes the data and adjusts the level of detail of the text overlay.

[0119] The text overlay insertion unit can apply different text overlay insertion algorithms depending on the different scenarios during text overlay insertion. For example, in everyday life scenarios, the text overlay insertion unit uses a simple text overlay insertion algorithm. In emergency scenarios, for example, the text overlay insertion unit can also use a more detailed text overlay insertion algorithm. In scenarios of specific events, for example, the text overlay insertion unit can also use an event-specific text overlay insertion algorithm. This enables appropriate text overlay insertion depending on the different scenarios. Different scenarios include, but are not limited to, emergency situations, normal situations, and specific events. Some or all of the above processing in the text overlay insertion unit is performed using AI. For example, the text overlay insertion unit inputs scenario data into the AI, which analyzes the scenario data and selects an appropriate text overlay insertion algorithm.

[0120] The following briefly describes the processing flow for example form 2.

[0121] Step 1: The acquisition unit acquires the video captured by the camera. The acquisition unit can, for example, acquire the video captured by the camera in real time. The acquisition unit can also acquire the camera's video periodically. Furthermore, the acquisition unit can acquire the camera's video in streaming format. For example, the acquisition unit can acquire the camera's video at 30 frames per second. The acquisition unit can also acquire the camera's video every hour. Furthermore, the acquisition unit can acquire the camera's video in live streaming format. Step 2: The analysis unit analyzes the video acquired by the acquisition unit and converts it into a visually stylized representation. For example, the analysis unit can stylize the video in an anime style. The analysis unit can also stylize the video in a manga style. Furthermore, the analysis unit can silhouette the video. For example, the analysis unit can stylize the video in an anime character style. The analysis unit can also stylize the video in a manga panel style. Furthermore, the analysis unit can silhouette the video to protect privacy. Step 3: The caption insertion unit inserts annotation captions based on the actions and situations in the video analyzed by the analysis unit. For example, if an elderly person is walking in the video, the caption insertion unit will insert an annotation caption such as "Walking." The caption insertion unit can also insert an annotation caption such as "Eating" if an elderly person is eating in the video. Furthermore, the caption insertion unit can also insert an annotation caption such as "Resting" if an elderly person is resting in the video. Step 4: The anomaly detection unit detects anomalies based on the annotation captions inserted by the caption insertion unit. For example, the anomaly detection unit detects an anomaly if an elderly person falls in the video. The anomaly detection unit can also detect an anomaly if an elderly person remains motionless for a long period of time in the video. Furthermore, the anomaly detection unit can also detect an anomaly if an elderly person makes suspicious movements in the video. Step 5: The alert unit issues an alert based on the anomaly detected by the anomaly detection unit. For example, the alert unit will issue an alert if an elderly person falls. The alert unit can also issue an alert if an elderly person remains motionless for a long period of time. Furthermore, the alert unit can also issue an alert if an elderly person performs any suspicious actions.

[0122] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0123] Data generation model 58 is a form of so-called generative AI (Artificial Intelligence). An example of data generation model 58 is ChatGPT (registered trademark) (Internet search).<URL: https: / / openai.com / blog / chatgpt> Examples of generative AI include text generation AI, image generation AI, and multimodal generation AI. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images (e.g., still image data or video data). The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference result in one or more data formats from audio data, text data, and image data. The data generation model 58 includes, for example, text generation AI, image generation AI, and multimodal generation AI. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization. The specific processing unit 290 performs the specific processing described above using the data generation model 58. The data generation model 58 may be a fine-tuned model that outputs inference results from prompts that do not contain instructions, in which case the data generation model 58 can output inference results from prompts that do not contain instructions. In the data processing device 12, etc., there are multiple types of data generation models 58, and the data generation model 58 includes AI other than generative AI. AI other than generative AI includes, for example, linear regression, logistic regression, decision trees, random forests, support vector machines (SVMs), k-means clustering, convolutional neural networks (CNNs), recurrent neural networks (RNNs), generative adversarial networks (GANs), or naive Bayes, and can perform various processes, but is not limited to these examples. Also, the AI ​​may be an AI agent. Furthermore, when the processing of each of the above parts is performed by the AI, the processing may be performed by the AI ​​in part or in whole, but is not limited to this example.Furthermore, processing performed by AI, including generative AI, may be replaced with rule-based processing, and rule-based processing may be replaced with processing performed by AI, including generative AI.

[0124] Furthermore, the processing performed by the data processing system 10 described above is carried out by the specific processing unit 290 of the data processing device 12 or the control unit 46A of the smart device 14, but it may also be carried out by the specific processing unit 290 of the data processing device 12 and the control unit 46A of the smart device 14. In addition, the specific processing unit 290 of the data processing device 12 acquires or collects information necessary for processing from the smart device 14 or an external device, and the smart device 14 acquires or collects information necessary for processing from the data processing device 12 or an external device.

[0125] Each of the multiple elements described above, including the acquisition unit, analysis unit, caption insertion unit, anomaly detection unit, and alert issuing unit, is implemented in at least one of the smart device 14 and the data processing device 12. For example, the acquisition unit acquires video using the camera 42 of the smart device 14 and analyzes it using the specific processing unit 290 of the data processing device 12. The analysis unit is implemented in the specific processing unit 290 of the data processing device 12 and distorts the video. The caption insertion unit is implemented in the specific processing unit 46A of the smart device 14 and inserts annotation captions into the analyzed video. The anomaly detection unit is implemented in the specific processing unit 290 of the data processing device 12 and detects anomalies based on the inserted captions. The alert issuing unit is implemented in the specific processing unit 46A of the smart device 14 and issues an alert when an anomaly is detected. The correspondence between each unit and the device or control unit is not limited to the example described above and can be modified in various ways.

[0126] [Second Embodiment] Figure 3 shows an example of the configuration of the data processing system 210 according to the second embodiment.

[0127] As shown in Figure 3, the data processing system 210 includes a data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.

[0128] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN and / or LAN.

[0129] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication interface 44. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, and camera 42 are also connected to the bus 52.

[0130] The microphone 238 receives voice signals from the user and accepts instructions from the user. The microphone 238 captures the voice signals from the user, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0131] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, which captures images of the area around the user (for example, an imaging range defined by a field of view equivalent to the field of vision of a typical healthy person).

[0132] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0133] Figure 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Figure 4, the data processing device 12 performs specific processing by the processor 28. The storage 32 stores the specific processing program 56.

[0134] The processor 28 reads a specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 acting as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.

[0135] Storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290. The identification processing unit 290 can estimate the user's emotions using the emotion identification model 59 and perform identification processing using the user's emotions. The emotion estimation function (emotion identification function) using the emotion identification model 59 performs various estimations and predictions regarding the user's emotions, including but not limited to these examples. Furthermore, emotion estimation and prediction also include, for example, emotion analysis.

[0136] In the smart glasses 214, specific processing is performed by the processor 46. The storage 50 stores a specific processing program 60. The processor 46 reads the specific processing program 60 from the storage 50 and executes the read specific processing program 60 on the RAM 48. The specific processing is realized by the processor 46 acting as a control unit 46A according to the specific processing program 60 executed on the RAM 48. The smart glasses 214 also have a data generation model 58 and an emotion identification model 59, similar to the data generation model and emotion identification model 59, and can perform processing similar to that of the specific processing unit 290 using these models.

[0137] Furthermore, other devices besides the data processing device 12 may also have the data generation model 58. For example, a server device may have the data generation model 58. In this case, the data processing device 12 obtains processing results (such as prediction results) using the data generation model 58 by communicating with the server device that has the data generation model 58. Also, the data processing device 12 may be a server device or a terminal device owned by the user (for example, a mobile phone, robot, home appliance, etc.).

[0138] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0139] The data generation model 58 is a so-called generative AI. An example of a data generation model 58 is a generative AI such as ChatGPT. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and inference data such as audio data representing speech, text data representing text, and image data representing images (e.g., still image data or video data). The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference result in one or more data formats such as audio data, text data, and image data. The data generation model 58 includes, for example, text generation AI, image generation AI, and multimodal generation AI. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization. The specific processing unit 290 performs the specific processing described above using the data generation model 58. The data generation model 58 may be a fine-tuned model that outputs inference results from prompts that do not contain instructions, in which case the data generation model 58 can output inference results from prompts that do not contain instructions. In the data processing device 12, etc., there are multiple types of data generation models 58, and the data generation model 58 includes AI other than generative AI. AI other than generative AI includes, for example, linear regression, logistic regression, decision trees, random forests, support vector machines (SVM), k-means clustering, convolutional neural networks (CNN), recurrent neural networks (RNN), generative adversarial networks (GAN), or naive Bayes, and can perform various processes, but is not limited to these examples. Also, the AI ​​may be an AI agent. Furthermore, when the processing of each part described above is performed by the AI, the processing may be performed by the AI ​​in part or in whole, but is not limited to this example. Also, processing performed by an AI including a generative AI may be replaced by rule-based processing, and rule-based processing may be replaced by processing performed by an AI including a generative AI.

[0140] The data processing system 210 according to the second embodiment performs the same processing as the data processing system 10 according to the first embodiment. The processing by the data processing system 210 is performed by the specific processing unit 290 of the data processing device 12 or the control unit 46A of the smart glasses 214, but it may also be performed by the specific processing unit 290 of the data processing device 12 and the control unit 46A of the smart glasses 214. In addition, the specific processing unit 290 of the data processing device 12 acquires or collects information necessary for processing from the smart glasses 214 or an external device, and the smart glasses 214 acquires or collects information necessary for processing from the data processing device 12 or an external device.

[0141] Each of the multiple elements described above, including the acquisition unit, analysis unit, caption insertion unit, anomaly detection unit, and alert issuing unit, is implemented, for example, in at least one of the smart glasses 214 and the data processing device 12. For example, the acquisition unit acquires video using the camera 42 of the smart glasses 214 and analyzes it using the specific processing unit 290 of the data processing device 12. The analysis unit is implemented, for example, by the specific processing unit 290 of the data processing device 12 and distorts the video. The caption insertion unit is implemented, for example, by the control unit 46A of the smart glasses 214 and inserts annotation captions into the analyzed video. The anomaly detection unit is implemented, for example, by the specific processing unit 290 of the data processing device 12 and detects anomalies based on the inserted captions. The alert issuing unit is implemented, for example, by the control unit 46A of the smart glasses 214 and issues an alert when an anomaly is detected. The correspondence between each unit and the device or control unit is not limited to the example described above and can be modified in various ways.

[0142] [Third Embodiment] Figure 5 shows an example of the configuration of the data processing system 310 according to the third embodiment.

[0143] As shown in Figure 5, the data processing system 310 includes a data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.

[0144] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN and / or LAN.

[0145] The headset terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a display 343. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and display 343 are also connected to the bus 52.

[0146] The microphone 238 receives voice signals from the user and accepts instructions from the user. The microphone 238 captures the voice signals from the user, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0147] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, which captures images of the area around the user (for example, an imaging range defined by a field of view equivalent to the field of vision of a typical healthy person).

[0148] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0149] Figure 6 shows an example of the main functions of the data processing device 12 and the headset terminal 314. As shown in Figure 6, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0150] The processor 28 reads a specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 acting as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.

[0151] Storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290. The identification processing unit 290 can estimate the user's emotions using the emotion identification model 59 and perform identification processing using the user's emotions. The emotion estimation function (emotion identification function) using the emotion identification model 59 performs various estimations and predictions regarding the user's emotions, including but not limited to these examples. Furthermore, emotion estimation and prediction also include, for example, emotion analysis.

[0152] In the headset terminal 314, specific processing is performed by the processor 46. The storage 50 stores a specific program 60. The processor 46 reads the specific program 60 from the storage 50 and executes the read specific program 60 on the RAM 48. The specific processing is realized by the processor 46 acting as a control unit 46A according to the specific program 60 executed on the RAM 48. The headset terminal 314 also has a data generation model 58 and an emotion identification model 59, similar to the data generation model and emotion identification model 59, and can perform processing similar to that of the specific processing unit 290 using these models.

[0153] Furthermore, other devices besides the data processing device 12 may also have the data generation model 58. For example, a server device may have the data generation model 58. In this case, the data processing device 12 obtains processing results (such as prediction results) using the data generation model 58 by communicating with the server device that has the data generation model 58. Also, the data processing device 12 may be a server device or a terminal device owned by the user (for example, a mobile phone, robot, home appliance, etc.).

[0154] The specific processing unit 290 transmits the result of the specific processing to the headset terminal 314. In the headset terminal 314, the control unit 46A causes the speaker 240 and display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0155] The data generation model 58 is a so-called generative AI. An example of a data generation model 58 is a generative AI such as ChatGPT. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and inference data such as audio data representing speech, text data representing text, and image data representing images (e.g., still image data or video data). The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference result in one or more data formats such as audio data, text data, and image data. The data generation model 58 includes, for example, text generation AI, image generation AI, and multimodal generation AI. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization. The specific processing unit 290 performs the specific processing described above using the data generation model 58. The data generation model 58 may be a fine-tuned model that outputs inference results from prompts that do not contain instructions, in which case the data generation model 58 can output inference results from prompts that do not contain instructions. In the data processing device 12, etc., there are multiple types of data generation models 58, and the data generation model 58 includes AI other than generative AI. AI other than generative AI includes, for example, linear regression, logistic regression, decision trees, random forests, support vector machines (SVM), k-means clustering, convolutional neural networks (CNN), recurrent neural networks (RNN), generative adversarial networks (GAN), or naive Bayes, and can perform various processes, but is not limited to these examples. Also, the AI ​​may be an AI agent. Furthermore, when the processing of each part described above is performed by the AI, the processing may be performed by the AI ​​in part or in whole, but is not limited to this example. Also, processing performed by an AI including a generative AI may be replaced by rule-based processing, and rule-based processing may be replaced by processing performed by an AI including a generative AI.

[0156] The data processing system 310 according to the third embodiment performs the same processing as the data processing system 10 according to the first embodiment. The processing by the data processing system 310 is performed by the specific processing unit 290 of the data processing device 12 or the control unit 46A of the headset terminal 314, but may also be performed by the specific processing unit 290 of the data processing device 12 and the control unit 46A of the headset terminal 314. In addition, the specific processing unit 290 of the data processing device 12 acquires or collects information necessary for processing from the headset terminal 314 or an external device, and the headset terminal 314 acquires or collects information necessary for processing from the data processing device 12 or an external device.

[0157] Each of the multiple elements described above, including the acquisition unit, analysis unit, caption insertion unit, anomaly detection unit, and alert issuing unit, is implemented in at least one of the headset terminal 314 and the data processing unit 12. For example, the acquisition unit acquires video using the camera 42 of the headset terminal 314 and analyzes it using the specific processing unit 290 of the data processing unit 12. The analysis unit is implemented in the specific processing unit 290 of the data processing unit 12 and distorts the video. The caption insertion unit is implemented in the specific processing unit 46A of the headset terminal 314 and inserts annotation captions into the analyzed video. The anomaly detection unit is implemented in the specific processing unit 290 of the data processing unit 12 and detects anomalies based on the inserted captions. The alert issuing unit is implemented in the specific processing unit 46A of the headset terminal 314 and issues an alert when an anomaly is detected. The correspondence between each unit and the device or control unit is not limited to the example described above and can be modified in various ways.

[0158] [Fourth Embodiment] Figure 7 shows an example of the configuration of the data processing system 410 according to the fourth embodiment.

[0159] As shown in Figure 7, the data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.

[0160] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN and / or LAN.

[0161] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a controlled object 443. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and controlled object 443 are also connected to the bus 52.

[0162] The microphone 238 receives voice signals from the user and accepts instructions from the user. The microphone 238 captures the voice signals from the user, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0163] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS image sensor or CCD image sensor, which captures images of the area around the user (for example, an imaging range defined by a field of view equivalent to the field of vision of a typical healthy person).

[0164] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0165] The controlled object 443 includes a display device, LEDs in the eyes, and motors that drive the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the robot 414's emotions can be expressed by controlling these motors. The robot 414's facial expressions can also be expressed by controlling the illumination state of the LEDs in its eyes.

[0166] Figure 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Figure 8, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0167] The processor 28 reads a specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 acting as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.

[0168] Storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290. The identification processing unit 290 can estimate the user's emotions using the emotion identification model 59 and perform identification processing using the user's emotions. The emotion estimation function (emotion identification function) using the emotion identification model 59 performs various estimations and predictions regarding the user's emotions, including but not limited to these examples. Furthermore, emotion estimation and prediction also include, for example, emotion analysis.

[0169] In robot 414, specific processing is performed by processor 46. A specific program 60 is stored in storage 50. Processor 46 reads the specific program 60 from storage 50 and executes it on RAM 48. The specific processing is achieved by processor 46 acting as a control unit 46A according to the specific program 60 executed on RAM 48. Robot 414 also has data generation model 58 and emotion identification model 59, similar to those of the robot, and can perform processing similar to that of the specific processing unit 290 using these models.

[0170] Furthermore, other devices besides the data processing device 12 may also have the data generation model 58. For example, a server device may have the data generation model 58. In this case, the data processing device 12 obtains processing results (such as prediction results) using the data generation model 58 by communicating with the server device that has the data generation model 58. Also, the data processing device 12 may be a server device or a terminal device owned by the user (for example, a mobile phone, robot, home appliance, etc.).

[0171] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the controlled object 443 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0172] The data generation model 58 is a so-called generative AI. An example of a data generation model 58 is a generative AI such as ChatGPT. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and inference data such as audio data representing speech, text data representing text, and image data representing images (e.g., still image data or video data). The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference result in one or more data formats such as audio data, text data, and image data. The data generation model 58 includes, for example, text generation AI, image generation AI, and multimodal generation AI. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization. The specific processing unit 290 performs the specific processing described above using the data generation model 58. The data generation model 58 may be a fine-tuned model that outputs inference results from prompts that do not contain instructions, in which case the data generation model 58 can output inference results from prompts that do not contain instructions. In the data processing device 12, etc., there are multiple types of data generation models 58, and the data generation model 58 includes AI other than generative AI. AI other than generative AI includes, for example, linear regression, logistic regression, decision trees, random forests, support vector machines (SVM), k-means clustering, convolutional neural networks (CNN), recurrent neural networks (RNN), generative adversarial networks (GAN), or naive Bayes, and can perform various processes, but is not limited to these examples. Also, the AI ​​may be an AI agent. Furthermore, when the processing of each part described above is performed by the AI, the processing may be performed by the AI ​​in part or in whole, but is not limited to this example. Also, processing performed by an AI including a generative AI may be replaced by rule-based processing, and rule-based processing may be replaced by processing performed by an AI including a generative AI.

[0173] The data processing system 410 according to the fourth embodiment performs the same processing as the data processing system 10 according to the first embodiment. The processing by the data processing system 410 is performed by the specific processing unit 290 of the data processing device 12 or the control unit 46A of the robot 414, but it may also be performed by the specific processing unit 290 of the data processing device 12 and the control unit 46A of the robot 414. In addition, the specific processing unit 290 of the data processing device 12 acquires or collects information necessary for processing from the robot 414 or an external device, and the robot 414 acquires or collects information necessary for processing from the data processing device 12 or an external device.

[0174] Each of the multiple elements described above, including the acquisition unit, analysis unit, caption insertion unit, anomaly detection unit, and alert issuing unit, is implemented in at least one of the robot 414 and the data processing unit 12. For example, the acquisition unit acquires video using the camera 42 of the robot 414 and analyzes it using the specific processing unit 290 of the data processing unit 12. The analysis unit is implemented in the specific processing unit 290 of the data processing unit 12 and distorts the video. The caption insertion unit is implemented in the control unit 46A of the robot 414 and inserts annotation captions into the analyzed video. The anomaly detection unit is implemented in the specific processing unit 290 of the data processing unit 12 and detects anomalies based on the inserted captions. The alert issuing unit is implemented in the control unit 46A of the robot 414 and issues an alert when an anomaly is detected. The correspondence between each unit and the device or control unit is not limited to the example described above and can be modified in various ways.

[0175] Furthermore, the emotion identification model 59, acting as an emotion engine, may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to a specific mapping, which is an emotion map (see Figure 9). Similarly, the emotion identification model 59 may also determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.

[0176] Figure 9 shows the emotion map 400, in which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. The closer to the center of the concentric circles, the more primitive the emotions are located. Further out of the concentric circles, emotions representing states and actions arising from mental states are located. Emotion is a concept that includes feelings and mental states. On the left side of the concentric circles, emotions that are generally generated from reactions occurring in the brain are located. On the right side of the concentric circles, emotions that are generally induced by situational judgment are located. Above and below the concentric circles, emotions that are generally generated from reactions occurring in the brain and induced by situational judgment are located. In addition, the emotion of "pleasure" is located on the upper side of the concentric circles, and the emotion of "displeasure" is located on the lower side. Thus, in the emotion map 400, multiple emotions are mapped based on the structure in which emotions arise, and emotions that are likely to occur simultaneously are mapped close together.

[0177] These emotions are distributed at the 3 o'clock position on the Emotion Map 400, and usually fluctuate between feelings of security and anxiety. In the right half of the Emotion Map 400, situational awareness takes precedence over internal feelings, resulting in a calm impression.

[0178] The inside of the Emotion Map 400 represents inner thoughts, while the outside represents actions. Therefore, the further you go from the outside of the Emotion Map 400, the more visible (expressed in actions) your emotions become.

[0179] Here, human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. Similarly, in robots, cars, and motorcycles, emotions can be created based on various balances, such as posture and battery level. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. The emotion map can be generated based, for example, on Dr. Mitsuyoshi's emotion map (Research on a system for analyzing brain physiological signals of speech emotion recognition and emotion, Tokushima University, doctoral dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map contains emotions belonging to a region called "response," where sensation is dominant. The right half of the emotion map contains emotions belonging to a region called "situation," where situational awareness is dominant.

[0180] The emotion map defines two emotions that promote learning. One is the emotion around the middle of the negative "repentance" and "reflection" on the situation side. In other words, it is when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is the emotion around the positive "desire" on the reaction side. In other words, it is when the robot has positive feelings such as "I want more" or "I want to know more."

[0181] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values ​​representing each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple training data sets, which are combinations of user input and emotion values ​​representing each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions located close together have similar values, as shown in the emotion map 900 in Figure 10. Figure 10 shows an example where multiple emotions such as "reassured," "calm," and "confident" have similar emotion values.

[0182] In the above embodiment, an example was given in which a specific process is performed by a single computer 22. However, the technology of this disclosure is not limited thereto, and a distributed processing method for the specific process may be used, which includes computer 22 and multiple other computers.

[0183] In the above embodiment, an example was given in which the specific processing program 56 is stored in the storage 32, but the technology of this disclosure is not limited thereto. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-temporary storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-temporary storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes specific processing according to the specific processing program 56.

[0184] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.

[0185] Furthermore, it is not necessary to store the entirety of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store the entirety of the specific processing program 56 in the storage 32; it is acceptable to store only a portion of the specific processing program 56.

[0186] The following types of processors can be used as hardware resources to perform specific processing. Examples of processors include a CPU, a general-purpose processor that functions as a hardware resource to perform specific processing by executing software, i.e., a program. Other examples of processors include dedicated electrical circuits, such as FPGAs (Field-Programmable Gate Arrays), PLDs (Programmable Logic Devices), or ASICs (Application Specific Integrated Circuits), which have circuit configurations specifically designed to perform specific processing. All of these processors have built-in or connected memory, and all of them perform specific processing by using memory.

[0187] The hardware resource that performs a specific process may consist of one of these various processors, or it may consist of a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Alternatively, the hardware resource that performs a specific process may consist of a single processor.

[0188] Examples of configurations using a single processor include, firstly, a configuration in which one or more CPUs and software are combined to form a single processor, and this processor functions as a hardware resource that performs a specific process. Secondly, there is a configuration using a processor that realizes the functions of the entire system, including multiple hardware resources that perform a specific process, on a single IC chip, as exemplified by SoCs (System-on-a-chip). In this way, a specific process is realized using one or more of the above types of processors as hardware resources.

[0189] Furthermore, the hardware structure of these various processors can more specifically utilize electrical circuits that combine circuit elements such as semiconductor devices. Also, the specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps can be deleted, new steps added, or the processing order rearranged, as long as it does not deviate from the main purpose.

[0190] Furthermore, although the above-described examples were divided into four embodiments, some or all of these embodiments may be combined. Also, the smart device 14, smart glasses 214, headset terminal 314, and robot 414 are just examples, and they may be combined, or other devices may be used. Also, although the above-described examples were divided into two embodiments, Embodiment 1 and Embodiment 2, these may be combined.

[0191] The descriptions and illustrations presented above are detailed explanations of the technical aspects of this disclosure and are merely examples of the technical aspects. For example, the above descriptions of the structure, function, operation, and effect are examples of the structure, function, operation, and effect of the technical aspects of this disclosure. Therefore, it goes without saying that you may delete unnecessary parts, add new elements, or replace elements in the descriptions and illustrations presented above, as long as you do not deviate from the essence of the technical aspects of this disclosure. Furthermore, in order to avoid confusion and facilitate understanding of the technical aspects of this disclosure, explanations of common technical knowledge and other things that do not require special explanation to enable the implementation of the technical aspects of this disclosure have been omitted from the descriptions and illustrations presented above.

[0192] All documents, patent applications, and technical standards described herein are incorporated by reference to the same extent as if each individual document, patent application, and technical standard were specifically and individually noted to be incorporated by reference.

[0193] (Note 1) An acquisition unit that acquires images captured by a camera, An analysis unit analyzes the video acquired by the acquisition unit and converts it into a visually stylized representation, A caption insertion unit inserts annotation captions based on the actions and situations in the video analyzed by the aforementioned analysis unit, An anomaly detection unit detects anomalies based on annotation captions inserted by the aforementioned caption insertion unit, The system includes an alert issuing unit that issues an alert based on an anomaly detected by the anomaly detection unit. A system characterized by the following features. (Note 2) The aforementioned analysis unit, Transforming images into visually stylized representations. The system described in Appendix 1, characterized by the features described herein. (Note 3) The aforementioned text overlay insertion section is, Insert annotation captions based on actions and situations within the video. The system described in Appendix 1, characterized by the features described herein. (Note 4) The abnormality detection unit, Detecting abnormal behavior or situations within the video. The system described in Appendix 1, characterized by the features described herein. (Note 5) The alert issuing unit is An alert will be issued if an anomaly is detected. The system described in Appendix 1, characterized by the features described herein. (Note 6) The acquisition unit is, The system estimates the user's emotions and adjusts the timing of video acquisition based on those emotions. The system described in Appendix 1, characterized by the features described herein. (Note 7) The acquisition unit is, When acquiring video footage, the optimal acquisition method is selected based on the user's activity patterns. The system described in Appendix 1, characterized by the features described herein. (Note 8) The acquisition unit is, When acquiring video footage, filtering is performed based on the user's current environment and circumstances. The system described in Appendix 1, characterized by the features described herein. (Note 9) The acquisition unit is, It estimates the user's emotions and determines the priority of the videos to acquire based on the estimated user emotions. The system described in Appendix 1, characterized by the features described herein. (Note 10) The acquisition unit is, When acquiring video footage, the system prioritizes acquiring highly relevant footage by considering the user's geographical location. The system described in Appendix 1, characterized by the features described herein. (Note 11) The acquisition unit is, When acquiring video footage, the system analyzes the user's social media activity and retrieves relevant videos. The system described in Appendix 1, characterized by the features described herein. (Note 12) The aforementioned analysis unit, It estimates the user's emotions and adjusts the way the deformation is expressed based on the estimated user emotions. The system described in Appendix 1, characterized by the features described herein. (Note 13) The aforementioned analysis unit, When analyzing video footage, the level of detail in the deformation is adjusted based on specific actions or situations. The system described in Appendix 1, characterized by the features described herein. (Note 14) The aforementioned analysis unit, When analyzing video footage, different deformation algorithms are applied depending on the different scenarios. The system described in Appendix 1, characterized by the features described herein. (Note 15) The aforementioned analysis unit, It estimates the user's emotions and adjusts the style of deformation based on the estimated user emotions. The system described in Appendix 1, characterized by the features described herein. (Note 16) The aforementioned analysis unit, When analyzing video footage, the priority of deformation is determined based on when the footage was shot. The system described in Appendix 1, characterized by the features described herein. (Note 17) The aforementioned analysis unit, During video analysis, the order of deformation is adjusted based on the relevance of the images. The system described in Appendix 1, characterized by the features described herein. (Note 18) The aforementioned text overlay insertion section is, The system estimates the user's emotions and adjusts the way the text is displayed based on those estimated emotions. The system described in Appendix 1, characterized by the features described herein. (Note 19) The aforementioned text overlay insertion section is, When inserting text overlays, adjust the level of detail based on the importance of the action or situation. The system described in Appendix 1, characterized by the features described herein. (Note 20) The aforementioned text overlay insertion section is, When inserting text overlays, different text overlay insertion algorithms are applied depending on the different scenarios. The system described in Appendix 1, characterized by the features described herein. (Note 21) The aforementioned text overlay insertion section is, The system estimates the user's emotions and adjusts the length of the on-screen text based on those emotions. The system described in Appendix 1, characterized by the features described herein. (Note 22) The aforementioned text overlay insertion section is, When inserting text overlays, the priority of the text overlays is determined based on when the action or situation occurred. The system described in Appendix 1, characterized by the features described herein. (Note 23) The aforementioned text overlay insertion section is, When inserting text overlays, adjust the order of the text overlays based on the relevance of actions and situations. The system described in Appendix 1, characterized by the features described herein. (Note 24) The abnormality detection unit, The system estimates the user's emotions and adjusts the anomaly detection criteria based on the estimated emotions. The system described in Appendix 1, characterized by the features described herein. (Note 25) The abnormality detection unit, When an anomaly is detected, the anomaly detection algorithm is optimized by referring to past anomaly data. The system described in Appendix 1, characterized by the features described herein. (Note 26) The abnormality detection unit, When detecting anomalies, different anomaly detection methods are applied depending on the different scenarios. The system described in Appendix 1, characterized by the features described herein. (Note 27) The abnormality detection unit, It estimates the user's emotions and adjusts the order in which anomaly detection results are displayed based on the estimated user emotions. The system described in Appendix 1, characterized by the features described herein. (Note 28) The abnormality detection unit, When detecting anomalies, the geographical distribution of the video footage is taken into consideration. The system described in Appendix 1, characterized by the features described herein. (Note 29) The abnormality detection unit, When anomalies are detected, the accuracy of anomaly detection is improved by referring to relevant literature related to the video. The system described in Appendix 1, characterized by the features described herein. (Note 30) The alert issuing unit is It estimates the user's emotions and adjusts how alerts are sent based on those estimated emotions. The system described in Appendix 1, characterized by the features described herein. (Note 31) The alert issuing unit is When an alert is issued, the alert generation algorithm is optimized by referring to past alert data. The system described in Appendix 1, characterized by the features described herein. (Note 32) The alert issuing unit is When issuing an alert, different alerting methods are applied depending on the different scenario. The system described in Appendix 1, characterized by the features described herein. (Note 33) The alert issuing unit is It estimates the user's emotions and determines the priority of alerts based on the estimated user emotions. The system described in Appendix 1, characterized by the features described herein. (Note 34) The alert issuing unit is When an alert is issued, the timing of the alert is adjusted based on when the video was captured. The system described in Appendix 1, characterized by the features described herein. (Note 35) The alert issuing unit is When an alert is issued, the order in which the alerts are issued is adjusted based on the relevance of the video footage. The system described in Appendix 1, characterized by the features described herein. [Explanation of Symbols]

[0194] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Devices 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robots

Claims

1. An acquisition unit that acquires images captured by a camera, An analysis unit analyzes the video acquired by the acquisition unit and converts it into a visually stylized representation, A caption insertion unit inserts annotation captions based on the actions and situations in the video analyzed by the aforementioned analysis unit, An anomaly detection unit detects anomalies based on annotation captions inserted by the aforementioned caption insertion unit, The system includes an alert issuing unit that issues an alert based on an anomaly detected by the anomaly detection unit. A system characterized by the following features.

2. The aforementioned analysis unit, Transforming images into visually stylized representations. The system according to feature 1.

3. The aforementioned text overlay insertion section is, Insert annotation captions based on actions and situations within the video. The system according to feature 1.

4. The abnormality detection unit, Detecting abnormal behavior or situations within the video. The system according to feature 1.

5. The alert issuing unit is An alert will be issued if an anomaly is detected. The system according to feature 1.

6. The acquisition unit is, The system estimates the user's emotions and adjusts the timing of video acquisition based on those emotions. The system according to feature 1.

7. The acquisition unit is, When acquiring video footage, the optimal acquisition method is selected based on the user's activity patterns. The system according to feature 1.

8. The acquisition unit is, When acquiring video footage, filtering is performed based on the user's current environment and circumstances. The system according to feature 1.

Citation Information

Patent Citations

  • Persona chatbot control method and system

    JP2022180282A