System

A surveillance-based system with video analysis and directional speakers addresses passenger troubles in public transportation by automatically detecting and warning individuals, enhancing safety and comfort.

JP2026022298APending Publication Date: 2026-02-12SOFTBANK GROUP CORP
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2024123815
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-07-30
Publication Date
2026-02-12

AI Technical Summary

Technical Problem

Public transportation systems face challenges in effectively detecting and responding to passenger troubles and violations of etiquette, as current monitoring methods often lead to further conflicts and are difficult for crew members to manage in real-time.

Method used

A system utilizing surveillance cameras, video analysis, trouble detection algorithms, and directional speakers to automatically identify and address issues by generating and delivering targeted warnings to passengers, ensuring a safe and comfortable riding environment.

Benefits of technology

The system enables rapid and accurate detection of abnormal behavior, generating appropriate warnings, and ensuring passenger safety and comfort by minimizing disruptions to others.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026022298000001_ABST
    Figure 2026022298000001_ABST
Patent Text Reader

Abstract

A system is provided.SOLUTION: A system includes monitoring means, video analysis means, trouble detection means, warning message generation means, and utterance means.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The technology of the present disclosure relates to a system. [Background technology]

[0002] Patent document 1 discloses a persona chatbot control method performed by at least one processor, the method including the steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to a description of the chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance. [Prior art documents] [Patent documents]

[0003] [Patent Document 1] Japanese Patent Publication No. 2022-180282 Summary of the Invention [Problem to be solved by the invention]

[0004] Public transportation is prone to frequent troubles between passengers and violations of etiquette. Current monitoring methods pose the risk of direct warnings between passengers causing further trouble, and it is practically difficult for crew members to monitor and respond to all troubles. Therefore, in order to provide a safe and comfortable riding environment, effective means of detecting and responding to troubles are necessary. [Means for solving the problem]

[0005] In order to solve the above problems, the present invention provides the following means. A system is provided that includes a monitoring means, a video analysis means, a trouble detection means, a warning message generation means, and a speech means. The monitoring means uses a surveillance camera to capture video of the inside of a public transportation facility, and the video analysis means analyzes the video to extract features such as suspicious movements, crowds of people, and abnormal sounds. The trouble detection means detects trouble based on these features, and the warning message generation means generates an appropriate warning message in response to the detected trouble. Finally, the speech means speaks the generated warning message to the target passenger, effectively preventing trouble and realizing a safe and comfortable riding environment.

[0006] "Monitoring means" refers to devices or systems that acquire video information from inside vehicles or public transportation.

[0007] "Video analysis means" refers to a device or system that analyzes video information acquired by the monitoring means and detects suspicious activity or abnormal situations.

[0008] The "trouble detection means" is a device or system that automatically detects abnormal behavior or trouble based on the feature values ​​analyzed by the video analysis means.

[0009] The "warning message generating means" is a device or system that generates appropriate warning or caution messages in response to detected troubles.

[0010] The "utterance means" refers to a device or system that utters the generated warning message aloud and effectively conveys the warning or caution to the target person. [Brief explanation of the drawings]

[0011] [Figure 1] 1 is a conceptual diagram showing an example of the configuration of a data processing system according to a first embodiment. [Figure 2] 1 is a conceptual diagram showing an example of main functions of a data processing device and a smart device according to a first embodiment. [Figure 3] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a second embodiment. [Figure 4] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and smart glasses according to a second embodiment. [Figure 5] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a third embodiment. [Figure 6] FIG. 11 is a conceptual diagram showing an example of main functions of a data processing device and a headset-type terminal according to a third embodiment. [Figure 7] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a fourth embodiment. [Figure 8] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and a robot according to a fourth embodiment. [Figure 9] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 10] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 11] FIG. 3 is a sequence diagram showing a processing flow of the data processing system according to the first embodiment. [Figure 12] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 1. [Figure 13] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system according to the second embodiment when an emotion engine is combined. [Figure 14] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 2 when an emotion engine is combined. DETAILED DESCRIPTION OF THE INVENTION

[0012] An example of an embodiment of a system according to the technology of the present disclosure will be described below with reference to the accompanying drawings.

[0013] First, the terms used in the following description will be explained.

[0014] In the following embodiments, a coded processor (hereinafter simply referred to as a "processor") may be a single arithmetic device or a combination of multiple arithmetic devices. Furthermore, a processor may be a single type of arithmetic device or a combination of multiple types of arithmetic devices. Examples of arithmetic devices include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), and an APU (Accelerated Processing Unit).

[0015] In the following embodiments, a coded RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a working memory by a processor.

[0016] In the following embodiments, the coded storage is one or more non-volatile storage devices that store various programs, various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), and magnetic tapes.

[0017] In the following embodiments, a communication I / F (Interface) with a symbol is an interface including a communication processor, an antenna, etc. The communication I / F controls communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), Bluetooth (registered trademark), etc.

[0018] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." In other words, "A and / or B" means that it may be only A, only B, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" is also applied when three or more things are expressed connected by "and / or."

[0019] [First embodiment]

[0020] FIG. 1 shows an example of the configuration of a data processing system 10 according to the first embodiment.

[0021] 1, a data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.

[0022] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0023] The smart device 14 includes a computer 36, a reception device 38, an output device 40, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The reception device 38, the output device 40, and the camera 42 are also connected to the bus 52.

[0024] The reception device 38 includes a touch panel 38A, a microphone 38B, and the like, and receives user input. The touch panel 38A detects contact with an indicator (for example, a pen or a finger) to receive user input by the touch of the indicator. The microphone 38B detects the user's voice to receive user input by voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.

[0025] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form of expression that the user 20 can perceive (for example, audio and / or text). The display 40A displays visible information such as text and images in accordance with instructions from the processor 46. The speaker 40B outputs audio in accordance with instructions from the processor 46. The camera 42 is a compact digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.

[0026] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54.

[0027] FIG. 2 shows an example of the main functions of the data processing device 12 and the smart device 14.

[0028] 2, in the data processing device 12, a specific process is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific process is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0029] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0030] In the smart device 14, the processor 46 performs the reception output process. The storage 50 stores a reception output program 60. The reception output program 60 is used in conjunction with the specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0031] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0032] The present invention is a system that uses AI technology to monitor troubles and violations of etiquette on public transportation and automatically urges appropriate caution, and can be implemented in the following forms.

[0033] System configuration

[0034] The system includes the following elements:

[0035] 1. Terminal (surveillance camera)

[0036] 2. Server

[0037] 3. Terminal (directional speaker)

[0038] Program processing and natural language explanation

[0039] Video monitoring and transmission

[0040] The devices (surveillance cameras) capture images in real time inside vehicles and public transportation. These images are sent to a server via a network. The surveillance cameras are high-resolution and are installed to cover a wide area.

[0041] Video analysis

[0042] The server analyzes the received video data and uses an AI model (such as a convolutional neural network) to extract features from the video, such as suspicious movements, crowds of people, and unusual sounds. This allows for the rapid and accurate detection of unusual behavior and abnormal situations in the video.

[0043] Trouble detection

[0044] The server then runs a problem detection algorithm based on the extracted features. For example, if abnormally close movements of people or aggressive behavior are detected, it will be recognized as a specific problem (e.g., sexual harassment, assault, etc.) and a warning flag will be raised.

[0045] Generate warning messages

[0046] The server then generates a warning message based on the type of trouble, location, and behavior of the target. For example, if harassment is detected on a crowded train, the server generates a warning message saying, "Please stop harassing."

[0047] Caution speech

[0048] The generated warning message is sent from the server to a terminal (directional speaker), which then speaks the message and delivers the warning directly to the target person. Because the directional speaker can focus sound in a specific direction, it can deliver the message only to the target person without causing unnecessary inconvenience to surrounding passengers.

[0049] Specific examples

[0050] As a concrete example, we will explain how to detect and respond to sexual harassment on a crowded train.

[0051] 1. Video monitoring and transmission: The surveillance camera inside the vehicle captures the video and transmits it to the server. The video is high resolution and covers a wide area.

[0052] 2. Video analysis: Using an AI model, the server extracts features from the video received, such as suspicious movements and crowds of people, and analyzes people's movements and relative positions in detail.

[0053] 3. Trouble detection: AI analyzes these features to detect abnormal behavior, such as sexual harassment, which then raises a warning flag.

[0054] 4. Generation of warning message: When a problem is detected, the server generates a warning message such as "Please stop harassing the user."

[0055] 5. Caution speech: The server sends the generated caution text to the terminal (directional speaker), and the speaker speaks the caution to the target person.

[0056] In this way, this system effectively manages troubles and violations of etiquette on public transport, providing a safe and comfortable riding environment.

[0057] The processing flow will be explained below.

[0058] Step 1:

[0059] The device (surveillance camera) captures real-time images from inside a vehicle or public transport vehicle. The surveillance camera is set to capture high-resolution images and cover a wide area.

[0060] Step 2:

[0061] The video data captured by the terminal (surveillance camera) is compressed and sent over the network to the server, where it is converted into an appropriate format for efficient data transfer.

[0062] Step 3:

[0063] The server decodes the received video data and analyzes it in real time using dedicated video analysis software.

[0064] Step 4:

[0065] The server uses an AI model to extract features from the video data. These features include suspicious movements, crowds of people, and unusual sounds. Convolutional neural networks (CNNs) are used for extraction.

[0066] Step 5:

[0067] The server evaluates the extracted features to detect abnormal behavior and signs of trouble, where machine learning algorithms identify abnormal patterns based on past data.

[0068] Step 6:

[0069] The server sets a warning flag for any abnormal behavior it detects. For example, if assault or sexual harassment is detected, a warning flag for that behavior will be set.

[0070] Step 7:

[0071] The server collects details of the incident (type of behavior, location, characteristics of the victim, etc.) and uses conversational AI to generate a warning message, such as "Mobbing behavior has been detected. Please stop immediately."

[0072] Step 8:

[0073] The server converts the generated warning message into an appropriate data format and transmits it to the terminal (directional speaker) via the network.

[0074] Step 9:

[0075] The terminal (directional speaker) analyzes the received warning message and speaks it out to the target person. The directional speaker focuses the sound in a specific direction, warning the person without disturbing surrounding passengers.

[0076] Step 10:

[0077] The server continues to analyze the surveillance camera footage and monitors the subject's reaction to the attention, checking for changes in the subject's movements and behavior.

[0078] Step 11:

[0079] If the user does not follow the warning, the server generates an additional warning message and sends it to the device (directional speaker) again. If necessary, the warning message is changed to a stronger one.

[0080] Step 12:

[0081] If the server cannot resolve the problem, it automatically sends an emergency notification to the police and the driver, which includes detailed information about the problem, enabling a prompt response.

[0082] These are the specific processing steps of the program for this system. In this way, by linking each step together, troubles within public transportation systems can be effectively prevented and passenger safety can be ensured.

[0083] Example 1

[0084] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0085] There is a need to monitor troubles and violations of etiquette on public transportation and respond quickly and appropriately. However, current systems have difficulty detecting troubles in real time and automatically issuing appropriate warnings. Rapid detection and response to troubles is particularly important on packed transportation, and an effective system for this purpose is required.

[0086] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0087] In this invention, the server includes a monitoring means, a video analysis means, a trouble detection means, a warning message generation means, and a speech means, which makes it possible to detect troubles and manner violations on public transportation in real time and to promptly and appropriately warn people.

[0088] "Monitoring means" refers to devices for capturing images inside public transportation, and mainly includes cameras.

[0089] The "video analysis means" is a device that has the function of analyzing received video data and extracting features such as suspicious movements, crowds of people, and abnormal sounds.

[0090] The "trouble detection means" refers to an algorithm or device for detecting trouble based on the feature amount extracted by the video analysis means.

[0091] The "warning message generation means" is a device that has the function of using a generative AI model to generate appropriate warning messages in response to detected problems.

[0092] The "utterance means" is a device for uttering the generated warning message to a specific target person, and particularly includes a directional speaker.

[0093] This invention is a system that uses AI technology to monitor troubles and bad manners on public transportation and automatically warns passengers of appropriate precautions. This system is implemented using the following hardware and software.

[0094] System configuration

[0095] The system includes the following elements:

[0096] 1. Terminal (surveillance camera)

[0097] 2. Server

[0098] 3. Terminal (directional speaker)

[0099] Video acquisition and transmission

[0100] Terminals (surveillance cameras) are installed inside public transportation facilities and capture high-resolution video in real time. The surveillance cameras are installed to cover a wide area, making them particularly useful for detecting trouble on crowded trains. The captured video data is sent to a server via a network. The transmitted data is encrypted to ensure security.

[0101] Video analysis

[0102] After the video data is sent to the server, the server receives the data. The server is equipped with a convolutional neural network (CNN) built using TensorFlow and PyTorch. The server uses this AI model to analyze the received video data and extract features such as suspicious movements, crowds of people, and unusual sounds. The analysis is performed in real time, dividing each video frame and extracting features using a specific algorithm.

[0103] Trouble detection

[0104] The server runs a problem detection algorithm based on the features obtained from the video analysis. For example, if abnormally close movements of people or aggressive behavior are detected, it will recognize this as a specific problem (e.g., sexual harassment or assault) and raise a warning flag. For example, if the server recognizes that a specific passenger is making inappropriate contact with another passenger on a crowded train, a warning flag will be raised.

[0105] Generate warning messages

[0106] When a warning flag is raised on the server, the server generates a warning message based on the problem. The message is generated using a generative AI model (e.g., GPT-3) based on a specific prompt. The warning message is generated using the following prompt:

[0107] "Trouble type: sexual harassment

[0108] Location: A crowded train

[0109] Situation: Passenger A makes inappropriate contact with another passenger B."

[0110] Based on this prompt, an appropriate warning message such as "Please stop harassing others" is generated.

[0111] Caution speech

[0112] The warning message generated by the server is sent to a terminal (directional speaker). The terminal (directional speaker) can focus sound in a specific direction, so the message is delivered to the target person only, without causing unnecessary inconvenience to surrounding passengers. For example, the speaker can issue a voice message saying, "Please stop molesting me" to a passenger sitting in a specific seat on a crowded train.

[0113] This system quickly and accurately detects any trouble or violations of etiquette on public transport in real time and provides appropriate warnings, resulting in a safe and comfortable riding environment.

[0114] The embodiment of the invention is configured as described above, which makes it possible to effectively monitor and quickly respond to problems within public transportation.

[0115] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0116] Program processing flow

[0117] Step 1: Acquire and transmit video

[0118] Users capture video using terminals (surveillance cameras) installed within public transportation. The cameras capture high-resolution video in real time and transmit the video data to a server via a network. The video captured by the camera is continuously divided into frames and transmitted to the server. During this process, the video data is encrypted to ensure security.

[0119] Input: Real-time camera footage

[0120] Output: Video data sent to the server

[0121] Step 2: Receiving the video and preparing for analysis

[0122] The server receives video data sent from the terminal (surveillance camera) via the network. The received video data is divided into frames and saved.

[0123] Input: Video data transmitted over the network

[0124] Output: Frame-by-frame video data stored on the server

[0125] Step 3: Analyzing the footage

[0126] The server analyzes the stored video data using a convolutional neural network (CNN). Specifically, it uses AI frameworks such as TensorFlow and PyTorch to perform real-time data analysis. As a result of the analysis, features such as suspicious movements, crowds of people, and unusual sounds are extracted. These features are used as input for predicting trouble.

[0127] Input: Frame-by-frame video data stored on the server

[0128] Output: Features such as suspicious movements, crowds of people, and abnormal sounds

[0129] Step 4: Detecting the problem

[0130] The server runs a problem detection algorithm based on the features obtained through the analysis. For example, if abnormal proximity or aggressive behavior is detected, it recognizes this as a specific problem (e.g., sexual harassment or assault) and raises a warning flag. This allows specific abnormal behavior to be identified in real time.

[0131] Input: Features extracted by analysis

[0132] Output: Trouble detection flag

[0133] Step 5: Generate warning text

[0134] The server generates a warning message when a warning flag is raised. The server uses a generative AI model (e.g., GPT-3) to generate an appropriate warning message based on a specific prompt. For example, given the prompt "Trouble type: sexual harassment, location: crowded train, situation: passenger A making inappropriate contact with another passenger B," the server generates a warning message saying, "Please stop sexual harassment."

[0135] Input: Trouble detection flag, prompt text

[0136] Output: Generated warning message

[0137] Step 6: Send and speak the warning message

[0138] The server generates a warning message and sends it to the device (directional speaker). The directional speaker can focus sound in a specific direction, delivering the message only to the target person without disturbing surrounding passengers. The speaker then speaks a warning message to the target person, saying, "Please stop harassing."

[0139] Input: Generated warning text

[0140] Output: Warning message spoken from the device (directional speaker)

[0141] In this way, this system can monitor troubles and violations of etiquette on public transportation in real time and automatically provide appropriate warnings.

[0142] (Application example 1)

[0143] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0144] Systems for effectively monitoring and quickly responding to incidents and violations of etiquette on public transportation are limited, making it difficult for security staff to quickly grasp the situation and take appropriate action. Effective methods for focusing sound in a specific direction to draw attention are also lacking. Furthermore, existing systems struggle to provide real-time notifications and visual assistance, hindering security improvements.

[0145] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0146] In this invention, the server includes a monitoring means, an image analysis means, a trouble detection means, a warning message generation means, an output means, a notification means, and a support display means, which makes it possible to monitor and analyze troubles and manner violations on public transportation in real time, quickly detect troubles, generate and output appropriate warning messages, and immediately notify security staff and provide visual support information.

[0147] "Surveillance" refers to the equipment and technology used to capture footage inside public transport.

[0148] "Video analysis means" refers to devices and techniques for extracting and analyzing features from acquired video data.

[0149] "Trouble detection means" refers to devices and technologies for detecting trouble by discovering suspicious movements or abnormalities from features extracted by video analysis means.

[0150] "Warning message generation means" refers to a device and technology for generating appropriate warning messages in response to a detected problem.

[0151] "Speech means" refers to a device and technology for outputting the generated warning message as voice.

[0152] "Notification means" refers to devices and technologies used to notify security staff and other relevant parties of detected problems and generated warning messages.

[0153] "Support display means" refers to devices and technologies that visually provide necessary information to security staff when they respond.

[0154] The present invention is a system that uses AI technology to monitor troubles and bad manners on public transportation and automatically urges appropriate cautions. An embodiment of the present invention will be described.

[0155] System configuration

[0156] The system includes the following elements:

[0157] 1. Surveillance methods (e.g., built-in cameras in smart glasses)

[0158] 2. Server

[0159] 3. Speaking means (e.g., directional speaker)

[0160] 4. Means of notification

[0161] 5. Supportive display means (e.g., smart glasses display)

[0162] Program processing

[0163] Video Acquisition

[0164] The monitoring tool will capture real-time footage of public transport vehicles, for example, through cameras built into smart glasses, allowing security staff to keep up to date with the latest developments.

[0165] Video transmission

[0166] The captured images are then sent to a server via Wi-Fi or mobile data networks, where high-resolution images are transmitted for improved analysis accuracy.

[0167] Video analysis

[0168] The server analyzes the received video data using an AI model (e.g., convolutional neural network), which extracts features such as suspicious movements, crowds of people, and abnormal sounds.

[0169] Trouble detection

[0170] The server then runs a problem detection algorithm based on the extracted features. For example, if unusually close movements of people or aggressive behavior are detected, this is recognized as a specific problem and a warning flag is raised.

[0171] Generate warning messages

[0172] The server generates a warning message based on the type of incident, location, and the target's behavior. For example, if a sexual assault is detected on a crowded train, a warning message saying "Please stop sexual assault" will be automatically generated.

[0173] Caution speech

[0174] The generated warning message is sent from the server to a speech device (directional speaker), which then speaks the message and delivers the warning directly to the target person. This allows the message to be delivered only to specific targets without disturbing surrounding passengers.

[0175] Notifications and Support Displays

[0176] At the same time, the server notifies security staff of any detected problems or warning messages via notification means, and visual support information is displayed on the smart glasses' display to help them respond quickly and appropriately.

[0177] Specific examples

[0178] As a concrete example, we will explain how to detect and respond to sexual harassment on a crowded train.

[0179] 1. Image acquisition: The built-in camera in the smart glasses captures images and sends them to the server. High-resolution images cover a wide area.

[0180] 2. Video analysis: Using an AI model, the server extracts features from the video received, such as suspicious movements and crowds of people, and performs a detailed analysis of people's movements and relative positions.

[0181] 3. Trouble detection: Abnormal behavior, such as sexual harassment, is detected and a warning flag is raised.

[0182] 4. Generation of warning message: The server generates a warning message saying, "Please stop harassing."

[0183] 5. Caution message generation and notification: The server sends the generated caution message to the directional speaker and issues a message, while simultaneously providing visual support information to security staff.

[0184] Prompt Sentence Examples

[0185] "Please detect abnormal movements and crowding of people on crowded trains and notify us so that appropriate measures can be taken."

[0186] This allows any trouble or violations of etiquette on public transport to be detected immediately, allowing security staff to respond quickly and appropriately.

[0187] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[0188] Step 1:

[0189] Users wear smart glasses while in public transport, and the built-in camera captures real-time images of passengers. The input is real-time images from inside the public transport vehicle, and the output is high-resolution video data.

[0190] Step 2:

[0191] The device (smart glasses) transmits the captured video to a server via Wi-Fi or a mobile data network. The input is high-resolution video data, and the output is the video data transmitted to the server.

[0192] Step 3:

[0193] The server analyzes the received video data using an AI model (for example, a convolutional neural network). The input is the transmitted video data, and the output is the extracted features (suspicious movements, crowds of people, abnormal sounds, etc.). Specific operations include inputting the video data into the AI ​​model and extracting the features.

[0194] Step 4:

[0195] The server runs a trouble detection algorithm based on the extracted features. The input is the extracted features, and the output is the trouble detection result (e.g., detection of sexual harassment). Specifically, it compares the extracted features with pre-set criteria to detect anomalies.

[0196] Step 5:

[0197] The server generates an appropriate warning message based on the results of the trouble detection. The input is the trouble detection result, and the output is the generated warning message (e.g., "Please stop harassing others"). Specifically, the server inputs the type of trouble and location information into a predefined warning message template to generate the appropriate warning message.

[0198] Step 6:

[0199] The server transmits the generated warning message to the speech generator (directional speaker). The input is the generated warning message, and the output is the data to be transmitted to the directional speaker. Specifically, the server transmits an appropriate message to the directional speaker.

[0200] Step 7:

[0201] The speech generator (directional speaker) speaks the transmitted warning message as a voice in a specific direction. The input is the transmitted data, and the output is the actual voice warning message. Specifically, it uses a voice synthesis function to speak a specific message to the target.

[0202] Step 8:

[0203] At the same time, the server uses a notification means to notify security staff of the detected trouble and the generated warning message. The input is the trouble detection result and the warning message, and the output is visual support information displayed on the smart glasses' display. Specifically, the server sends the notification data to the smart glasses, and appropriate support information is displayed on the display.

[0204] Prompt Sentence Examples

[0205] "Detect abnormal movements and crowding of people inside a moving train and provide appropriate countermeasures."

[0206] Furthermore, an emotion engine that estimates the user's emotion may be combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59 and perform identification processing using the user's emotion.

[0207] The present invention is a system that uses AI technology to monitor troubles and violations of etiquette on public transportation and automatically urges appropriate cautions, and by combining it with an emotion engine that recognizes the user's emotions, it is intended to enhance the effectiveness of preventing troubles. It can be implemented in the following forms.

[0208] System configuration

[0209] The system includes the following elements:

[0210] 1. Terminal (surveillance camera)

[0211] 2. Server

[0212] 3. Terminal (directional speaker)

[0213] 4. Emotion Engine

[0214] Program processing and natural language explanation

[0215] Video monitoring and transmission

[0216] The devices (surveillance cameras) capture images in real time inside vehicles and public transportation. These images are sent to a server via a network. The surveillance cameras are high-resolution and are installed to cover a wide area.

[0217] Video analysis

[0218] The server analyzes the received video data and uses an AI model (such as a convolutional neural network) to extract features from the video, such as suspicious movements, crowds of people, and unusual sounds. This allows for the rapid and accurate detection of unusual behavior and abnormal situations in the video.

[0219] Trouble detection

[0220] The server then runs a problem detection algorithm based on the extracted features. For example, if abnormally close movements of people or aggressive behavior are detected, it will be recognized as a specific problem (e.g., sexual harassment, assault, etc.) and a warning flag will be raised.

[0221] Emotion recognition

[0222] When a problem is detected, the emotion engine analyzes the target's facial expressions and voice tone to recognize their emotional state (e.g., anger, anxiety, fear, etc.). The emotion engine uses an AI model to extract emotions from video and audio data.

[0223] Generate warning messages

[0224] Next, the server generates a warning message based on the emotional data of the target person recognized by the emotion engine. The conversational AI automatically creates an appropriate warning message based on the type of trouble, location, and the target person's emotional state. For example, if the target person is expressing anger, a message such as "Avoid trouble. Remain calm" will be generated.

[0225] Caution speech

[0226] The generated warning message is sent from the server to the device (directional speaker), which then directly warns the target person. The speaker also speaks in a tone and language that corresponds to the emotional state recognized by the emotion engine. This allows for more effective warning communication to the target person.

[0227] Specific examples

[0228] As a concrete example, we will explain how to detect and respond to problems on a crowded train.

[0229] 1. Video monitoring and transmission: The surveillance camera inside the vehicle captures the video and transmits it to the server. The video is high resolution and covers a wide area.

[0230] 2. Video analysis: Using an AI model, the server extracts features from the video received, such as suspicious movements and crowds of people, and analyzes people's movements and relative positions in detail.

[0231] 3. Trouble detection: AI analyzes these features to detect abnormal behavior, such as sexual harassment, which then raises a warning flag.

[0232] 4. Emotion Recognition: The emotion engine analyzes the facial expressions and voice of the target when a problem is detected and recognizes their emotional state. For example, if the target shows feelings of anxiety or anger.

[0233] 5. Generating warning messages: The server takes into account the recognized emotional data and generates warning messages such as "Please stay calm to avoid trouble" rather than "Please stop harassing others."

[0234] 6. Speech of warning: The server sends the generated warning message to the device (directional speaker), and the speaker speaks the warning to the target person. At the same time, the message is conveyed in a tone and language that matches the user's emotions.

[0235] These are the specific processing steps of the program of this system. In this way, by linking each step together, it is possible to effectively prevent trouble on public transportation and ensure the safety of passengers.

[0236] The processing flow will be explained below.

[0237] Step 1:

[0238] The device (surveillance camera) captures real-time images from inside a vehicle or public transport vehicle. The surveillance camera is set to capture high-resolution images and cover a wide area.

[0239] Step 2:

[0240] The video data captured by the terminal (surveillance camera) is compressed and sent over the network to the server, where it is converted into an appropriate format for efficient data transfer.

[0241] Step 3:

[0242] The server decodes the received video data and analyzes it in real time using dedicated video analysis software.

[0243] Step 4:

[0244] The server uses an AI model to extract features from the video data. These features include suspicious movements, crowds of people, and unusual sounds. Convolutional neural networks (CNNs) are used for extraction.

[0245] Step 5:

[0246] The server evaluates the extracted features to detect abnormal behavior and signs of trouble, where machine learning algorithms identify abnormal patterns based on past data.

[0247] Step 6:

[0248] The server sets a warning flag for any abnormal behavior it detects. For example, if assault or sexual harassment is detected, a warning flag for that behavior will be set.

[0249] Step 7:

[0250] When a problem is detected, the emotion engine analyzes the target's facial expressions and voice tone to recognize their emotional state (e.g., anger, anxiety, fear, etc.). The emotion engine uses an AI model to extract emotions from video and audio data.

[0251] Step 8:

[0252] The server organizes the details of the trouble (type of behavior, location, characteristics of the target, etc.) and the emotional state recognized by the emotion engine, and generates a warning message using conversational AI. For example, if the target shows anger, a message such as "Avoid trouble. Remain calm" will be generated.

[0253] Step 9:

[0254] The server converts the generated warning message into an appropriate data format and transmits it to the terminal (directional speaker) via the network.

[0255] Step 10:

[0256] The device (directional speaker) analyzes the received warning message and speaks it out to the target passenger. The directional speaker focuses the sound in a specific direction, warning the passenger without disturbing other passengers. The emotion engine also recognizes the passenger's emotional state and uses a tone and wording that matches that state, providing more effective warnings.

[0257] Step 11:

[0258] The server continues to analyze the surveillance camera footage and monitors the subject's reaction to the attention, checking for changes in the subject's movements and behavior.

[0259] Step 12:

[0260] If the user does not follow the warning, the server generates an additional warning message and sends it to the device (directional speaker) again. If necessary, the warning message is changed to a stronger one.

[0261] Step 13:

[0262] If the server cannot resolve the problem, it automatically sends an emergency notification to the police and the driver, which includes detailed information about the problem, enabling a prompt response.

[0263] Example 2

[0264] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0265] Trouble and bad manners on public transportation are serious issues that threaten passenger safety and comfort. Current monitoring systems have limitations in detecting suspicious movements and abnormal behavior, and are unable to respond appropriately while taking into account the emotional state of passengers. This makes it difficult to take prompt and effective measures when trouble occurs. There is also a lack of technology that can recognize passengers' emotions and psychological state and respond accordingly.

[0266] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[0267] In this invention, the server includes a monitoring means, a video analysis means, a trouble detection means, an emotion recognition means, a warning message generation means, and a speech utterance means. This makes it possible to monitor video in public transportation in real time, detect suspicious movements and crowding of people, and recognize the emotional state of passengers, and generate and speak appropriate warning messages based on those emotions.

[0268] "Monitoring means" refers to devices for capturing images inside public transportation, and primarily involves the use of surveillance cameras.

[0269] "Video analysis means" refers to technology or equipment that analyzes acquired video data and extracts features such as suspicious movements, crowds of people, and abnormal sounds.

[0270] "Trouble detection means" refers to algorithms and technologies for detecting abnormal behavior or trouble based on extracted features.

[0271] "Emotion recognition means" refers to technology or equipment that analyzes the facial expressions and tone of voice of a person when a problem is detected, and recognizes the emotional state of the person.

[0272] The "warning message generation means" refers to a technique or device for generating appropriate warning messages based on the recognized emotional state.

[0273] The "utterance means" is a device for uttering the generated warning message in a specific direction to convey it to the target person.

[0274] This invention is a system that uses AI technology to monitor troubles and bad manners on public transportation and automatically provides appropriate warnings. The system includes the following elements:

[0275] monitoring means

[0276] The terminal (surveillance camera) captures images of the inside of public transportation in real time. The surveillance camera has high resolution and is installed to cover a wide area. The captured images are sent to a server via a network.

[0277] Video analysis methods

[0278] The server analyzes the received video data, using AI models such as convolutional neural networks to extract features from the video, such as suspicious movements, crowds of people, and unusual sounds.

[0279] Trouble detection methods

[0280] The server runs a problem detection algorithm based on the extracted features. If abnormally close movements of people or aggressive behavior are detected, it will recognize it as a specific problem (e.g., assault) and raise a warning flag.

[0281] emotion recognition means

[0282] When the emotion engine detects a problem, it analyzes the target's facial expressions and vocal tone, using an AI model that extracts emotions from video and audio data to recognize the target's emotional state (e.g., anger, anxiety, fear, etc.).

[0283] Caution statement generation means

[0284] The server generates warning messages based on the emotional data of the target person recognized by the emotion engine. The conversational AI automatically creates appropriate warning messages based on the type of trouble, location, and the target person's emotional state. For example, if the target person is expressing anger, a message such as "Avoid trouble. Remain calm" will be generated.

[0285] Means of speech

[0286] The generated warning message is sent from the server to the device (directional speaker). The directional speaker then directly warns the target person. The speaker also speaks in a tone and language that corresponds to the emotional state recognized by the emotion engine. This operation allows for more effective warning communication to the target person.

[0287] Specific examples

[0288] As a concrete example, we will explain how to detect and respond to problems on a crowded train.

[0289] 1. Video monitoring and transmission: The terminal (surveillance camera) captures images from inside the train and transmits them to the server in real time. The camera has high resolution and covers a wide area.

[0290] 2. Video analysis: The server uses an AI model to analyze the video data it receives and extracts suspicious movements and crowds of people as features.

[0291] 3. Trouble detection: The server uses AI to analyze features and detect abnormal approaches or unnatural movements. This allows it to detect troubles such as sexual harassment and raise a warning flag.

[0292] 4. Emotion recognition: The emotion engine analyzes the subject's facial expressions and voice to recognize emotions such as anxiety or anger.

[0293] 5. Generating warning messages: Based on the emotional data, the server generates warning messages such as "Please stay calm to avoid trouble" rather than "Please stop harassing others."

[0294] 6. Speech warning: The server generates a warning message and sends it to the device (directional speaker), which then speaks it to the target person. By conveying the warning in a tone and using words that match the target person's emotions, a more effective response can be achieved.

[0295] Examples of prompt statements

[0296] Below is an example of a prompt sentence to input to the generative AI model.

[0297] "Please explain the system for detecting suspicious behavior on public transportation and taking appropriate action. For example, please explain in detail how to respond to sexual harassment on a crowded train."

[0298] The system serves as an effective means of improving safety within public transport.

[0299] The flow of the identification process in the second embodiment will be described with reference to FIG.

[0300] Step 1:

[0301] Video monitoring and transmission

[0302] explanation

[0303] The device (surveillance camera) captures images of the inside of public transportation in real time, and the captured image data is sent to a server via a network.

[0304] Specific actions

[0305] Input: Public transport footage

[0306] Data processing: The surveillance camera sensor captures the video and obtains detailed video data in high resolution.

[0307] Output: Video data sent to the server in real time

[0308] Step 2:

[0309] Video analysis

[0310] explanation

[0311] The server analyzes the received video data and uses an AI model (such as a convolutional neural network) to extract features such as suspicious movements, crowds of people, and unusual sounds from the video.

[0312] Specific actions

[0313] Input: Received video data

[0314] Data processing: Decoding video data, analyzing each frame, and using AI models to identify suspicious movements, crowds of people, and unusual sounds

[0315] Output: Extracted feature data

[0316] Step 3:

[0317] Trouble detection

[0318] explanation

[0319] The server runs a trouble-detection algorithm based on the extracted features. For example, if abnormally close movements of people or aggressive behavior are detected, it will recognize the problem and raise a warning flag.

[0320] Specific actions

[0321] Input: Extracted feature data

[0322] Data processing: Input feature data into the trouble detection algorithm to detect abnormal patterns

[0323] Output: Trouble detection flag

[0324] Step 4:

[0325] Emotion recognition

[0326] explanation

[0327] The emotion engine recognizes the emotional state of the person who detected the problem by analyzing facial expressions and tone of voice (e.g., anger, anxiety, fear, etc.).

[0328] Specific actions

[0329] Input: Subject's facial expression data and voice data

[0330] Data processing: The emotion engine uses AI models to analyze facial expressions and voice to identify emotional states.

[0331] Output: Recognized emotion data

[0332] Step 5:

[0333] Generate warning messages

[0334] explanation

[0335] The server generates warning messages based on the target person's emotional data recognized by the emotion engine. The conversational AI creates warning messages based on the type of trouble, location, and the target person's emotional state.

[0336] Specific actions

[0337] Input: Recognized emotion data, type of trouble, location information

[0338] Data processing: Input prompts into the generative AI model to generate appropriate warning messages

[0339] Output: Generated warning message

[0340] Step 6:

[0341] Caution speech

[0342] explanation

[0343] The generated warning message is sent from the server to the device (directional speaker), which then speaks the message to the target person.

[0344] Specific actions

[0345] Input: Generated warning text data

[0346] Data processing: A directional speaker uses a speech synthesis engine to play back warning messages and speak them in a specific direction.

[0347] Output: Audio warning to the target

[0348] In this way, each step works together to effectively prevent trouble on public transport and ensure passenger safety.

[0349] (Application example 2)

[0350] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0351] In commercial facilities, customers often violate etiquette and cause trouble, which often has a negative impact on store operations and other customers. The purpose of this invention is to maintain order and safety within commercial facilities by quickly and effectively detecting such trouble and violations of etiquette and automatically issuing appropriate warnings.

[0352] The identification process by the identification processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes a monitoring means, a video analysis means, a trouble detection means, an emotion recognition means, a warning message generation means, a speech means, and a notification means. This makes it possible to automatically detect troubles and etiquette violations within a commercial facility and issue appropriate warnings.

[0353] The "monitoring means" is a device that acquires images of the inside of a commercial facility in real time.

[0354] The "video analysis means" is a device or software for processing acquired video data and extracting specific features.

[0355] The "trouble detection means" is a device or algorithm that detects abnormal movements or behaviors based on features extracted from the video analysis means.

[0356] An "emotion recognition means" is a device or algorithm that analyzes the facial expressions and voice of a detected subject and recognizes their emotional state.

[0357] The "warning message generation means" is a device or software that generates appropriate warning messages for the target person based on the data obtained from the emotion recognition means.

[0358] The "utterance means" is a device for uttering the generated warning message as voice.

[0359] The "notification means" is a device or software that transmits the generated warning message to a smartphone or other device.

[0360] This invention is a system that detects troubles and bad manners in commercial facilities and automatically issues appropriate warnings. This system includes seven main means: "monitoring means," "video analysis means," "trouble detection means," "emotion recognition means," "warning message generation means," "speech means," and "notification means."

[0361] System configuration

[0362] The system includes the following elements:

[0363] 1. Monitoring measures

[0364] Camera devices and smartphone cameras installed within commercial facilities are used to capture images of the facility in real time.

[0365] 2. Video analysis methods

[0366] The captured video data is received and specific features are extracted using an AI model (e.g., convolutional neural network).

[0367] The features include abnormal movements, crowding of people, and abnormal behavior.

[0368] 3. Trouble detection methods

[0369] The server detects abnormal behavior based on the features obtained from the video analysis means.

[0370] In this case, a trouble detection algorithm is used, for example, to analyze theft or aggressive behavior and recognize it as trouble.

[0371] 4. Emotion recognition means

[0372] The server analyzes the subject's facial expressions and voice to recognize their emotional state.

[0373] To do this, it uses an AI model that extracts emotions from video and audio data.

[0374] 5. Caution statement generation means

[0375] The server generates appropriate warning messages based on the data obtained from the emotion recognition means.

[0376] For example, if the subject is feeling anxious, the system will automatically generate the phrase "Please remain calm."

[0377] 6. Means of speech

[0378] The warning messages sent from the server are spoken through speakers installed within the commercial facility or via a smartphone app.

[0379] Speech is delivered in a tone and language that corresponds to the emotion.

[0380] 7. Means of notification

[0381] The generated warning message is sent to store employees and the target person via a smartphone notification system.

[0382] For example, if any theft is detected, the store security team is immediately notified.

[0383] Specific examples

[0384] As a concrete example, we will explain the detection and response to theft behavior in a large shopping mall.

[0385] 1. Video monitoring and transmission: Surveillance cameras in the shopping mall capture video and transmit the data to a server. The cameras have high resolution and wide coverage.

[0386] 2. Video analysis: The server uses an AI model to extract suspicious movements and crowding of people from the received video data as features.

[0387] 3. Trouble detection: AI analyzes these features and detects, for example, theft. It raises a warning flag depending on the type of trouble.

[0388] 4. Emotion recognition: The server utilizes emotion recognition means to analyze the target person's facial expressions and voice to recognize their emotional state (e.g., anxiety, anger).

[0389] 5. Generating warning messages: The server generates appropriate warning messages (e.g., "Please stay calm") based on the emotion data.

[0390] 6. Caution notification: The generated caution message is sent directly to the target person via speech, and is also sent to employees and security teams via their smartphone notification system.

[0391] Prompt Sentence Examples

[0392] An example of a prompt to input to a generative AI model is as follows:

[0393] Prompt: If you detect any unusual behavior in the mall, create a message to encourage the person you spot to remain calm.

[0394] Response: "Remain calm and try to resolve the issue. If you need assistance, please speak to a member of staff."

[0395] The above configuration enables quick and effective detection of troubles and violations of etiquette within commercial facilities, and makes it possible to take appropriate measures when they occur.

[0396] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[0397] Step 1:

[0398] The server acquires real-time video footage of the commercial facility from the monitoring means (surveillance cameras). This video footage is high resolution and contains a wide range of data. The input is the video data acquired from the surveillance cameras, and the output is the real-time video data sent to the server.

[0399] Step 2:

[0400] The server analyzes the received video data using a video analysis method. Specifically, it uses a convolutional neural network (CNN) to extract features such as abnormal movements and crowding of people in the video. The input is the video data acquired in step 1, and the output is the extracted feature data.

[0401] Step 3:

[0402] The server uses the trouble detection means to detect troublesome behavior based on the features obtained from the video analysis means. For example, theft or aggressive behavior is analyzed using an algorithm and recognized as trouble. The input is the feature data extracted in step 2, and the output is the type of trouble detected and its location information.

[0403] Step 4:

[0404] The server uses emotion recognition means to analyze the facial expressions and voices of the people involved in the trouble and recognize their emotional state. Specifically, it identifies the emotional state using an AI model that extracts emotions. The input is a list of people involved in the trouble detected in step 3 and their video data, and the output is the emotional state of each person.

[0405] Step 5:

[0406] The server uses the warning message generation means to generate appropriate warning messages based on the emotion data obtained from the emotion recognition means. For example, it generates a message such as "Please stay calm to avoid problems." The input is the emotion data and trouble information obtained in step 4, and the output is the generated warning message.

[0407] Step 6:

[0408] The server sends the generated warning message to the speech means (directional speaker) and notification means (smartphone notification system). The warning message is spoken directly within the commercial facility through the speech means, and simultaneously notifies the target person or employee via their smartphone. The input is the warning message generated in step 5, and the output is the speech within the commercial facility and the smartphone notification.

[0409] These processing steps enable quick and effective detection of troubles and violations of etiquette within commercial facilities, and appropriate measures to be taken.

[0410] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0411] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0412] In the above embodiment, an example in which the specific process is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific process may be performed by the smart device 14.

[0413] [Second embodiment]

[0414] FIG. 3 shows an example of the configuration of a data processing system 210 according to the second embodiment.

[0415] 3, the data processing system 210 includes the data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.

[0416] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0417] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, and the camera 42 are also connected to the bus 52.

[0418] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[0419] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[0420] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[0421] Fig. 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Fig. 4, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[0422] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0423] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0424] In the smart glasses 214, the reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0425] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal."

[0426] The present invention is a system that uses AI technology to monitor troubles and violations of etiquette on public transportation and automatically urges appropriate caution, and can be implemented in the following forms.

[0427] System configuration

[0428] The system includes the following elements:

[0429] 1. Terminal (surveillance camera)

[0430] 2. Server

[0431] 3. Terminal (directional speaker)

[0432] Program processing and natural language explanation

[0433] Video monitoring and transmission

[0434] The devices (surveillance cameras) capture images in real time inside vehicles and public transportation. These images are sent to a server via a network. The surveillance cameras are high-resolution and are installed to cover a wide area.

[0435] Video analysis

[0436] The server analyzes the received video data and uses an AI model (such as a convolutional neural network) to extract features from the video, such as suspicious movements, crowds of people, and unusual sounds. This allows for the rapid and accurate detection of unusual behavior and abnormal situations in the video.

[0437] Trouble detection

[0438] The server then runs a problem detection algorithm based on the extracted features. For example, if abnormally close movements of people or aggressive behavior are detected, it will be recognized as a specific problem (e.g., sexual harassment, assault, etc.) and a warning flag will be raised.

[0439] Generate warning messages

[0440] The server then generates a warning message based on the type of trouble, location, and behavior of the target. For example, if harassment is detected on a crowded train, the server generates a warning message saying, "Please stop harassing."

[0441] Caution speech

[0442] The generated warning message is sent from the server to a terminal (directional speaker), which then speaks the message and delivers the warning directly to the target person. Because the directional speaker can focus sound in a specific direction, it can deliver the message only to the target person without causing unnecessary inconvenience to surrounding passengers.

[0443] Specific examples

[0444] As a concrete example, we will explain how to detect and respond to sexual harassment on a crowded train.

[0445] 1. Video monitoring and transmission: The surveillance camera inside the vehicle captures the video and transmits it to the server. The video is high resolution and covers a wide area.

[0446] 2. Video analysis: Using an AI model, the server extracts features from the video received, such as suspicious movements and crowds of people, and analyzes people's movements and relative positions in detail.

[0447] 3. Trouble detection: AI analyzes these features to detect abnormal behavior, such as sexual harassment, which then raises a warning flag.

[0448] 4. Generation of warning message: When a problem is detected, the server generates a warning message such as "Please stop harassing the user."

[0449] 5. Caution speech: The server sends the generated caution text to the terminal (directional speaker), and the speaker speaks the caution to the target person.

[0450] In this way, this system effectively manages troubles and violations of etiquette on public transport, providing a safe and comfortable riding environment.

[0451] The processing flow will be explained below.

[0452] Step 1:

[0453] The device (surveillance camera) captures real-time images from inside a vehicle or public transport vehicle. The surveillance camera is set to capture high-resolution images and cover a wide area.

[0454] Step 2:

[0455] The video data captured by the terminal (surveillance camera) is compressed and sent over the network to the server, where it is converted into an appropriate format for efficient data transfer.

[0456] Step 3:

[0457] The server decodes the received video data and analyzes it in real time using dedicated video analysis software.

[0458] Step 4:

[0459] The server uses an AI model to extract features from the video data. These features include suspicious movements, crowds of people, and unusual sounds. Convolutional neural networks (CNNs) are used for extraction.

[0460] Step 5:

[0461] The server evaluates the extracted features to detect abnormal behavior and signs of trouble, where machine learning algorithms identify abnormal patterns based on past data.

[0462] Step 6:

[0463] The server sets a warning flag for any abnormal behavior it detects. For example, if assault or sexual harassment is detected, a warning flag for that behavior will be set.

[0464] Step 7:

[0465] The server collects details of the incident (type of behavior, location, characteristics of the victim, etc.) and uses conversational AI to generate a warning message, such as "Mobbing behavior has been detected. Please stop immediately."

[0466] Step 8:

[0467] The server converts the generated warning message into an appropriate data format and transmits it to the terminal (directional speaker) via the network.

[0468] Step 9:

[0469] The terminal (directional speaker) analyzes the received warning message and speaks it out to the target person. The directional speaker focuses the sound in a specific direction, warning the person without disturbing surrounding passengers.

[0470] Step 10:

[0471] The server continues to analyze the surveillance camera footage and monitors the subject's reaction to the attention, checking for changes in the subject's movements and behavior.

[0472] Step 11:

[0473] If the user does not follow the warning, the server generates an additional warning message and sends it to the device (directional speaker) again. If necessary, the warning message is changed to a stronger one.

[0474] Step 12:

[0475] If the server cannot resolve the problem, it automatically sends an emergency notification to the police and the driver, which includes detailed information about the problem, enabling a prompt response.

[0476] These are the specific processing steps of the program for this system. In this way, by linking each step together, troubles within public transportation systems can be effectively prevented and passenger safety can be ensured.

[0477] Example 1

[0478] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0479] There is a need to monitor troubles and violations of etiquette on public transportation and respond quickly and appropriately. However, current systems have difficulty detecting troubles in real time and automatically issuing appropriate warnings. Rapid detection and response to troubles is particularly important on packed transportation, and an effective system for this purpose is required.

[0480] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0481] In this invention, the server includes a monitoring means, a video analysis means, a trouble detection means, a warning message generation means, and a speech means, which makes it possible to detect troubles and manner violations on public transportation in real time and to promptly and appropriately warn people.

[0482] "Monitoring means" refers to devices for capturing images inside public transportation, and mainly includes cameras.

[0483] The "video analysis means" is a device that has the function of analyzing received video data and extracting features such as suspicious movements, crowds of people, and abnormal sounds.

[0484] The "trouble detection means" refers to an algorithm or device for detecting trouble based on the feature amount extracted by the video analysis means.

[0485] The "warning message generation means" is a device that has the function of using a generative AI model to generate appropriate warning messages in response to detected problems.

[0486] The "utterance means" is a device for uttering the generated warning message to a specific target person, and particularly includes a directional speaker.

[0487] This invention is a system that uses AI technology to monitor troubles and bad manners on public transportation and automatically warns passengers of appropriate precautions. This system is implemented using the following hardware and software.

[0488] System configuration

[0489] The system includes the following elements:

[0490] 1. Terminal (surveillance camera)

[0491] 2. Server

[0492] 3. Terminal (directional speaker)

[0493] Video acquisition and transmission

[0494] Terminals (surveillance cameras) are installed inside public transportation facilities and capture high-resolution video in real time. The surveillance cameras are installed to cover a wide area, making them particularly useful for detecting trouble on crowded trains. The captured video data is sent to a server via a network. The transmitted data is encrypted to ensure security.

[0495] Video analysis

[0496] After the video data is sent to the server, the server receives the data. The server is equipped with a convolutional neural network (CNN) built using TensorFlow and PyTorch. The server uses this AI model to analyze the received video data and extract features such as suspicious movements, crowds of people, and unusual sounds. The analysis is performed in real time, dividing each video frame and extracting features using a specific algorithm.

[0497] Trouble detection

[0498] The server runs a problem detection algorithm based on the features obtained from the video analysis. For example, if abnormally close movements of people or aggressive behavior are detected, it will recognize this as a specific problem (e.g., sexual harassment or assault) and raise a warning flag. For example, if the server recognizes that a specific passenger is making inappropriate contact with another passenger on a crowded train, a warning flag will be raised.

[0499] Generate warning messages

[0500] When a warning flag is raised on the server, the server generates a warning message based on the problem. The message is generated using a generative AI model (e.g., GPT-3) based on a specific prompt. The warning message is generated using the following prompt:

[0501] "Trouble type: sexual harassment

[0502] Location: A crowded train

[0503] Situation: Passenger A makes inappropriate contact with another passenger B."

[0504] Based on this prompt, an appropriate warning message such as "Please stop harassing others" is generated.

[0505] Caution speech

[0506] The warning message generated by the server is sent to a terminal (directional speaker). The terminal (directional speaker) can focus sound in a specific direction, so the message is delivered to the target person only, without causing unnecessary inconvenience to surrounding passengers. For example, the speaker can issue a voice message saying, "Please stop molesting me" to a passenger sitting in a specific seat on a crowded train.

[0507] This system quickly and accurately detects any trouble or violations of etiquette on public transport in real time and provides appropriate warnings, resulting in a safe and comfortable riding environment.

[0508] The embodiment of the invention is configured as described above, which makes it possible to effectively monitor and quickly respond to problems within public transportation.

[0509] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0510] Program processing flow

[0511] Step 1: Acquire and transmit video

[0512] Users capture video using terminals (surveillance cameras) installed within public transportation. The cameras capture high-resolution video in real time and transmit the video data to a server via a network. The video captured by the camera is continuously divided into frames and transmitted to the server. During this process, the video data is encrypted to ensure security.

[0513] Input: Real-time camera footage

[0514] Output: Video data sent to the server

[0515] Step 2: Receiving the video and preparing for analysis

[0516] The server receives video data sent from the terminal (surveillance camera) via the network. The received video data is divided into frames and saved.

[0517] Input: Video data transmitted over the network

[0518] Output: Frame-by-frame video data stored on the server

[0519] Step 3: Analyzing the footage

[0520] The server analyzes the stored video data using a convolutional neural network (CNN). Specifically, it uses AI frameworks such as TensorFlow and PyTorch to perform real-time data analysis. As a result of the analysis, features such as suspicious movements, crowds of people, and unusual sounds are extracted. These features are used as input for predicting trouble.

[0521] Input: Frame-by-frame video data stored on the server

[0522] Output: Features such as suspicious movements, crowds of people, and abnormal sounds

[0523] Step 4: Detecting the problem

[0524] The server runs a problem detection algorithm based on the features obtained through the analysis. For example, if abnormal proximity or aggressive behavior is detected, it recognizes this as a specific problem (e.g., sexual harassment or assault) and raises a warning flag. This allows specific abnormal behavior to be identified in real time.

[0525] Input: Features extracted by analysis

[0526] Output: Trouble detection flag

[0527] Step 5: Generate warning text

[0528] The server generates a warning message when a warning flag is raised. The server uses a generative AI model (e.g., GPT-3) to generate an appropriate warning message based on a specific prompt. For example, given the prompt "Trouble type: sexual harassment, location: crowded train, situation: passenger A making inappropriate contact with another passenger B," the server generates a warning message saying, "Please stop sexual harassment."

[0529] Input: Trouble detection flag, prompt text

[0530] Output: Generated warning message

[0531] Step 6: Send and speak the warning message

[0532] The server generates a warning message and sends it to the device (directional speaker). The directional speaker can focus sound in a specific direction, delivering the message only to the target person without disturbing surrounding passengers. The speaker then speaks a warning message to the target person, saying, "Please stop harassing."

[0533] Input: Generated warning text

[0534] Output: Warning message spoken from the device (directional speaker)

[0535] In this way, this system can monitor troubles and violations of etiquette on public transportation in real time and automatically provide appropriate warnings.

[0536] (Application example 1)

[0537] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0538] Systems for effectively monitoring and quickly responding to incidents and violations of etiquette on public transportation are limited, making it difficult for security staff to quickly grasp the situation and take appropriate action. Effective methods for focusing sound in a specific direction to draw attention are also lacking. Furthermore, existing systems struggle to provide real-time notifications and visual assistance, hindering security improvements.

[0539] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0540] In this invention, the server includes a monitoring means, an image analysis means, a trouble detection means, a warning message generation means, an output means, a notification means, and a support display means, which makes it possible to monitor and analyze troubles and manner violations on public transportation in real time, quickly detect troubles, generate and output appropriate warning messages, and immediately notify security staff and provide visual support information.

[0541] "Surveillance" refers to the equipment and technology used to capture footage inside public transport.

[0542] "Video analysis means" refers to devices and techniques for extracting and analyzing features from acquired video data.

[0543] "Trouble detection means" refers to devices and technologies for detecting trouble by discovering suspicious movements or abnormalities from features extracted by video analysis means.

[0544] "Warning message generation means" refers to a device and technology for generating appropriate warning messages in response to a detected problem.

[0545] "Speech means" refers to a device and technology for outputting the generated warning message as voice.

[0546] "Notification means" refers to devices and technologies used to notify security staff and other relevant parties of detected problems and generated warning messages.

[0547] "Support display means" refers to devices and technologies that visually provide necessary information to security staff when they respond.

[0548] The present invention is a system that uses AI technology to monitor troubles and bad manners on public transportation and automatically urges appropriate cautions. An embodiment of the present invention will be described.

[0549] System configuration

[0550] The system includes the following elements:

[0551] 1. Surveillance methods (e.g., built-in cameras in smart glasses)

[0552] 2. Server

[0553] 3. Speaking means (e.g., directional speaker)

[0554] 4. Means of notification

[0555] 5. Supportive display means (e.g., smart glasses display)

[0556] Program processing

[0557] Video Acquisition

[0558] The monitoring tool will capture real-time footage of public transport vehicles, for example, through cameras built into smart glasses, allowing security staff to keep up to date with the latest developments.

[0559] Video transmission

[0560] The captured images are then sent to a server via Wi-Fi or mobile data networks, where high-resolution images are transmitted for improved analysis accuracy.

[0561] Video analysis

[0562] The server analyzes the received video data using an AI model (e.g., convolutional neural network), which extracts features such as suspicious movements, crowds of people, and abnormal sounds.

[0563] Trouble detection

[0564] The server then runs a problem detection algorithm based on the extracted features. For example, if unusually close movements of people or aggressive behavior are detected, this is recognized as a specific problem and a warning flag is raised.

[0565] Generate warning messages

[0566] The server generates a warning message based on the type of incident, location, and the target's behavior. For example, if a sexual assault is detected on a crowded train, a warning message saying "Please stop sexual assault" will be automatically generated.

[0567] Caution speech

[0568] The generated warning message is sent from the server to a speech device (directional speaker), which then speaks the message and delivers the warning directly to the target person. This allows the message to be delivered only to specific targets without disturbing surrounding passengers.

[0569] Notifications and Support Displays

[0570] At the same time, the server notifies security staff of any detected problems or warning messages via notification means, and visual support information is displayed on the smart glasses' display to help them respond quickly and appropriately.

[0571] Specific examples

[0572] As a concrete example, we will explain how to detect and respond to sexual harassment on a crowded train.

[0573] 1. Image acquisition: The built-in camera in the smart glasses captures images and sends them to the server. High-resolution images cover a wide area.

[0574] 2. Video analysis: Using an AI model, the server extracts features from the video received, such as suspicious movements and crowds of people, and performs a detailed analysis of people's movements and relative positions.

[0575] 3. Trouble detection: Abnormal behavior, such as sexual harassment, is detected and a warning flag is raised.

[0576] 4. Generation of warning message: The server generates a warning message saying, "Please stop harassing."

[0577] 5. Caution message generation and notification: The server sends the generated caution message to the directional speaker and issues a message, while simultaneously providing visual support information to security staff.

[0578] Prompt Sentence Examples

[0579] "Please detect abnormal movements and crowding of people on crowded trains and notify us so that appropriate measures can be taken."

[0580] This allows any trouble or violations of etiquette on public transport to be detected immediately, allowing security staff to respond quickly and appropriately.

[0581] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[0582] Step 1:

[0583] Users wear smart glasses while in public transport, and the built-in camera captures real-time images of passengers. The input is real-time images from inside the public transport vehicle, and the output is high-resolution video data.

[0584] Step 2:

[0585] The device (smart glasses) transmits the captured video to a server via Wi-Fi or a mobile data network. The input is high-resolution video data, and the output is the video data transmitted to the server.

[0586] Step 3:

[0587] The server analyzes the received video data using an AI model (for example, a convolutional neural network). The input is the transmitted video data, and the output is the extracted features (suspicious movements, crowds of people, abnormal sounds, etc.). Specific operations include inputting the video data into the AI ​​model and extracting the features.

[0588] Step 4:

[0589] The server runs a trouble detection algorithm based on the extracted features. The input is the extracted features, and the output is the trouble detection result (e.g., detection of sexual harassment). Specifically, it compares the extracted features with pre-set criteria to detect anomalies.

[0590] Step 5:

[0591] The server generates an appropriate warning message based on the results of the trouble detection. The input is the trouble detection result, and the output is the generated warning message (e.g., "Please stop harassing others"). Specifically, the server inputs the type of trouble and location information into a predefined warning message template to generate the appropriate warning message.

[0592] Step 6:

[0593] The server transmits the generated warning message to the speech generator (directional speaker). The input is the generated warning message, and the output is the data to be transmitted to the directional speaker. Specifically, the server transmits an appropriate message to the directional speaker.

[0594] Step 7:

[0595] The speech generator (directional speaker) speaks the transmitted warning message as a voice in a specific direction. The input is the transmitted data, and the output is the actual voice warning message. Specifically, it uses a voice synthesis function to speak a specific message to the target.

[0596] Step 8:

[0597] At the same time, the server uses a notification means to notify security staff of the detected trouble and the generated warning message. The input is the trouble detection result and the warning message, and the output is visual support information displayed on the smart glasses' display. Specifically, the server sends the notification data to the smart glasses, and appropriate support information is displayed on the display.

[0598] Prompt Sentence Examples

[0599] "Detect abnormal movements and crowding of people inside a moving train and provide appropriate countermeasures."

[0600] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[0601] The present invention is a system that uses AI technology to monitor troubles and violations of etiquette on public transportation and automatically urges appropriate cautions, and by combining it with an emotion engine that recognizes the user's emotions, it is intended to enhance the effectiveness of preventing troubles. It can be implemented in the following forms.

[0602] System configuration

[0603] The system includes the following elements:

[0604] 1. Terminal (surveillance camera)

[0605] 2. Server

[0606] 3. Terminal (directional speaker)

[0607] 4. Emotion Engine

[0608] Program processing and natural language explanation

[0609] Video monitoring and transmission

[0610] The devices (surveillance cameras) capture images in real time inside vehicles and public transportation. These images are sent to a server via a network. The surveillance cameras are high-resolution and are installed to cover a wide area.

[0611] Video analysis

[0612] The server analyzes the received video data and uses an AI model (such as a convolutional neural network) to extract features from the video, such as suspicious movements, crowds of people, and unusual sounds. This allows for the rapid and accurate detection of unusual behavior and abnormal situations in the video.

[0613] Trouble detection

[0614] The server then runs a problem detection algorithm based on the extracted features. For example, if abnormally close movements of people or aggressive behavior are detected, it will be recognized as a specific problem (e.g., sexual harassment, assault, etc.) and a warning flag will be raised.

[0615] Emotion recognition

[0616] When a problem is detected, the emotion engine analyzes the target's facial expressions and voice tone to recognize their emotional state (e.g., anger, anxiety, fear, etc.). The emotion engine uses an AI model to extract emotions from video and audio data.

[0617] Generate warning messages

[0618] Next, the server generates a warning message based on the emotional data of the target person recognized by the emotion engine. The conversational AI automatically creates an appropriate warning message based on the type of trouble, location, and the target person's emotional state. For example, if the target person is expressing anger, a message such as "Avoid trouble. Remain calm" will be generated.

[0619] Caution speech

[0620] The generated warning message is sent from the server to the device (directional speaker), which then directly warns the target person. The speaker also speaks in a tone and language that corresponds to the emotional state recognized by the emotion engine. This allows for more effective warning communication to the target person.

[0621] Specific examples

[0622] As a concrete example, we will explain how to detect and respond to problems on a crowded train.

[0623] 1. Video monitoring and transmission: The surveillance camera inside the vehicle captures the video and transmits it to the server. The video is high resolution and covers a wide area.

[0624] 2. Video analysis: Using an AI model, the server extracts features from the video received, such as suspicious movements and crowds of people, and analyzes people's movements and relative positions in detail.

[0625] 3. Trouble detection: AI analyzes these features to detect abnormal behavior, such as sexual harassment, which then raises a warning flag.

[0626] 4. Emotion Recognition: The emotion engine analyzes the facial expressions and voice of the target when a problem is detected and recognizes their emotional state. For example, if the target shows feelings of anxiety or anger.

[0627] 5. Generating warning messages: The server takes into account the recognized emotional data and generates warning messages such as "Please stay calm to avoid trouble" rather than "Please stop harassing others."

[0628] 6. Speech of warning: The server sends the generated warning message to the device (directional speaker), and the speaker speaks the warning to the target person. At the same time, the message is conveyed in a tone and language that matches the user's emotions.

[0629] These are the specific processing steps of the program of this system. In this way, by linking each step together, it is possible to effectively prevent trouble on public transportation and ensure the safety of passengers.

[0630] The processing flow will be explained below.

[0631] Step 1:

[0632] The device (surveillance camera) captures real-time images from inside a vehicle or public transport vehicle. The surveillance camera is set to capture high-resolution images and cover a wide area.

[0633] Step 2:

[0634] The video data captured by the terminal (surveillance camera) is compressed and sent over the network to the server, where it is converted into an appropriate format for efficient data transfer.

[0635] Step 3:

[0636] The server decodes the received video data and analyzes it in real time using dedicated video analysis software.

[0637] Step 4:

[0638] The server uses an AI model to extract features from the video data. These features include suspicious movements, crowds of people, and unusual sounds. Convolutional neural networks (CNNs) are used for extraction.

[0639] Step 5:

[0640] The server evaluates the extracted features to detect abnormal behavior and signs of trouble, where machine learning algorithms identify abnormal patterns based on past data.

[0641] Step 6:

[0642] The server sets a warning flag for any abnormal behavior it detects. For example, if assault or sexual harassment is detected, a warning flag for that behavior will be set.

[0643] Step 7:

[0644] When a problem is detected, the emotion engine analyzes the target's facial expressions and voice tone to recognize their emotional state (e.g., anger, anxiety, fear, etc.). The emotion engine uses an AI model to extract emotions from video and audio data.

[0645] Step 8:

[0646] The server organizes the details of the trouble (type of behavior, location, characteristics of the target, etc.) and the emotional state recognized by the emotion engine, and generates a warning message using conversational AI. For example, if the target shows anger, a message such as "Avoid trouble. Remain calm" will be generated.

[0647] Step 9:

[0648] The server converts the generated warning message into an appropriate data format and transmits it to the terminal (directional speaker) via the network.

[0649] Step 10:

[0650] The device (directional speaker) analyzes the received warning message and speaks it out to the target passenger. The directional speaker focuses the sound in a specific direction, warning the passenger without disturbing other passengers. The emotion engine also recognizes the passenger's emotional state and uses a tone and wording that matches that state, providing more effective warnings.

[0651] Step 11:

[0652] The server continues to analyze the surveillance camera footage and monitors the subject's reaction to the attention, checking for changes in the subject's movements and behavior.

[0653] Step 12:

[0654] If the user does not follow the warning, the server generates an additional warning message and sends it to the device (directional speaker) again. If necessary, the warning message is changed to a stronger one.

[0655] Step 13:

[0656] If the server cannot resolve the problem, it automatically sends an emergency notification to the police and the driver, which includes detailed information about the problem, enabling a prompt response.

[0657] Example 2

[0658] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0659] Trouble and bad manners on public transportation are serious issues that threaten passenger safety and comfort. Current monitoring systems have limitations in detecting suspicious movements and abnormal behavior, and are unable to respond appropriately while taking into account the emotional state of passengers. This makes it difficult to take prompt and effective measures when trouble occurs. There is also a lack of technology that can recognize passengers' emotions and psychological state and respond accordingly.

[0660] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[0661] In this invention, the server includes a monitoring means, a video analysis means, a trouble detection means, an emotion recognition means, a warning message generation means, and a speech utterance means. This makes it possible to monitor video in public transportation in real time, detect suspicious movements and crowding of people, and recognize the emotional state of passengers, and generate and speak appropriate warning messages based on those emotions.

[0662] "Monitoring means" refers to devices for capturing images inside public transportation, and primarily involves the use of surveillance cameras.

[0663] "Video analysis means" refers to technology or equipment that analyzes acquired video data and extracts features such as suspicious movements, crowds of people, and abnormal sounds.

[0664] "Trouble detection means" refers to algorithms and technologies for detecting abnormal behavior or trouble based on extracted features.

[0665] "Emotion recognition means" refers to technology or equipment that analyzes the facial expressions and tone of voice of a person when a problem is detected, and recognizes the emotional state of the person.

[0666] The "warning message generation means" refers to a technique or device for generating appropriate warning messages based on the recognized emotional state.

[0667] The "utterance means" is a device for uttering the generated warning message in a specific direction to convey it to the target person.

[0668] This invention is a system that uses AI technology to monitor troubles and bad manners on public transportation and automatically provides appropriate warnings. The system includes the following elements:

[0669] monitoring means

[0670] The terminal (surveillance camera) captures images of the inside of public transportation in real time. The surveillance camera has high resolution and is installed to cover a wide area. The captured images are sent to a server via a network.

[0671] Video analysis methods

[0672] The server analyzes the received video data, using AI models such as convolutional neural networks to extract features from the video, such as suspicious movements, crowds of people, and unusual sounds.

[0673] Trouble detection methods

[0674] The server runs a problem detection algorithm based on the extracted features. If abnormally close movements of people or aggressive behavior are detected, it will recognize it as a specific problem (e.g., assault) and raise a warning flag.

[0675] emotion recognition means

[0676] When the emotion engine detects a problem, it analyzes the target's facial expressions and vocal tone, using an AI model that extracts emotions from video and audio data to recognize the target's emotional state (e.g., anger, anxiety, fear, etc.).

[0677] Caution statement generation means

[0678] The server generates warning messages based on the emotional data of the target person recognized by the emotion engine. The conversational AI automatically creates appropriate warning messages based on the type of trouble, location, and the target person's emotional state. For example, if the target person is expressing anger, a message such as "Avoid trouble. Remain calm" will be generated.

[0679] Means of speech

[0680] The generated warning message is sent from the server to the device (directional speaker). The directional speaker then directly warns the target person. The speaker also speaks in a tone and language that corresponds to the emotional state recognized by the emotion engine. This operation allows for more effective warning communication to the target person.

[0681] Specific examples

[0682] As a concrete example, we will explain how to detect and respond to problems on a crowded train.

[0683] 1. Video monitoring and transmission: The terminal (surveillance camera) captures images from inside the train and transmits them to the server in real time. The camera has high resolution and covers a wide area.

[0684] 2. Video analysis: The server uses an AI model to analyze the video data it receives and extracts suspicious movements and crowds of people as features.

[0685] 3. Trouble detection: The server uses AI to analyze features and detect abnormal approaches or unnatural movements. This allows it to detect troubles such as sexual harassment and raise a warning flag.

[0686] 4. Emotion recognition: The emotion engine analyzes the subject's facial expressions and voice to recognize emotions such as anxiety or anger.

[0687] 5. Generating warning messages: Based on the emotional data, the server generates warning messages such as "Please stay calm to avoid trouble" rather than "Please stop harassing others."

[0688] 6. Speech warning: The server generates a warning message and sends it to the device (directional speaker), which then speaks it to the target person. By conveying the warning in a tone and using words that match the target person's emotions, a more effective response can be achieved.

[0689] Examples of prompt statements

[0690] Below is an example of a prompt sentence to input to the generative AI model.

[0691] "Please explain the system for detecting suspicious behavior on public transportation and taking appropriate action. For example, please explain in detail how to respond to sexual harassment on a crowded train."

[0692] The system serves as an effective means of improving safety within public transport.

[0693] The flow of the identification process in the second embodiment will be described with reference to FIG.

[0694] Step 1:

[0695] Video monitoring and transmission

[0696] explanation

[0697] The device (surveillance camera) captures images of the inside of public transportation in real time, and the captured image data is sent to a server via a network.

[0698] Specific actions

[0699] Input: Public transport footage

[0700] Data processing: The surveillance camera sensor captures the video and obtains detailed video data in high resolution.

[0701] Output: Video data sent to the server in real time

[0702] Step 2:

[0703] Video analysis

[0704] explanation

[0705] The server analyzes the received video data and uses an AI model (such as a convolutional neural network) to extract features such as suspicious movements, crowds of people, and unusual sounds from the video.

[0706] Specific actions

[0707] Input: Received video data

[0708] Data processing: Decoding video data, analyzing each frame, and using AI models to identify suspicious movements, crowds of people, and unusual sounds

[0709] Output: Extracted feature data

[0710] Step 3:

[0711] Trouble detection

[0712] explanation

[0713] The server runs a trouble-detection algorithm based on the extracted features. For example, if abnormally close movements of people or aggressive behavior are detected, it will recognize the problem and raise a warning flag.

[0714] Specific actions

[0715] Input: Extracted feature data

[0716] Data processing: Input feature data into the trouble detection algorithm to detect abnormal patterns

[0717] Output: Trouble detection flag

[0718] Step 4:

[0719] Emotion recognition

[0720] explanation

[0721] The emotion engine recognizes the emotional state of the person who detected the problem by analyzing facial expressions and tone of voice (e.g., anger, anxiety, fear, etc.).

[0722] Specific actions

[0723] Input: Subject's facial expression data and voice data

[0724] Data processing: The emotion engine uses AI models to analyze facial expressions and voice to identify emotional states.

[0725] Output: Recognized emotion data

[0726] Step 5:

[0727] Generate warning messages

[0728] explanation

[0729] The server generates warning messages based on the target person's emotional data recognized by the emotion engine. The conversational AI creates warning messages based on the type of trouble, location, and the target person's emotional state.

[0730] Specific actions

[0731] Input: Recognized emotion data, type of trouble, location information

[0732] Data processing: Input prompts into the generative AI model to generate appropriate warning messages

[0733] Output: Generated warning message

[0734] Step 6:

[0735] Caution speech

[0736] explanation

[0737] The generated warning message is sent from the server to the device (directional speaker), which then speaks the message to the target person.

[0738] Specific actions

[0739] Input: Generated warning text data

[0740] Data processing: A directional speaker uses a speech synthesis engine to play back warning messages and speak them in a specific direction.

[0741] Output: Audio warning to the target

[0742] In this way, each step works together to effectively prevent trouble on public transport and ensure passenger safety.

[0743] (Application example 2)

[0744] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0745] In commercial facilities, customers often violate etiquette and cause trouble, which often has a negative impact on store operations and other customers. The purpose of this invention is to maintain order and safety within commercial facilities by quickly and effectively detecting such trouble and violations of etiquette and automatically issuing appropriate warnings.

[0746] The identification process by the identification processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes a monitoring means, a video analysis means, a trouble detection means, an emotion recognition means, a warning message generation means, a speech means, and a notification means. This makes it possible to automatically detect troubles and etiquette violations within a commercial facility and issue appropriate warnings.

[0747] The "monitoring means" is a device that acquires images of the inside of a commercial facility in real time.

[0748] The "video analysis means" is a device or software for processing acquired video data and extracting specific features.

[0749] The "trouble detection means" is a device or algorithm that detects abnormal movements or behaviors based on features extracted from the video analysis means.

[0750] An "emotion recognition means" is a device or algorithm that analyzes the facial expressions and voice of a detected subject and recognizes their emotional state.

[0751] The "warning message generation means" is a device or software that generates appropriate warning messages for the target person based on the data obtained from the emotion recognition means.

[0752] The "utterance means" is a device for uttering the generated warning message as voice.

[0753] The "notification means" is a device or software that transmits the generated warning message to a smartphone or other device.

[0754] This invention is a system that detects troubles and bad manners in commercial facilities and automatically issues appropriate warnings. This system includes seven main means: "monitoring means," "video analysis means," "trouble detection means," "emotion recognition means," "warning message generation means," "speech means," and "notification means."

[0755] System configuration

[0756] The system includes the following elements:

[0757] 1. Monitoring measures

[0758] Camera devices and smartphone cameras installed within commercial facilities are used to capture images of the facility in real time.

[0759] 2. Video analysis methods

[0760] The captured video data is received and specific features are extracted using an AI model (e.g., convolutional neural network).

[0761] The features include abnormal movements, crowding of people, and abnormal behavior.

[0762] 3. Trouble detection methods

[0763] The server detects abnormal behavior based on the features obtained from the video analysis means.

[0764] In this case, a trouble detection algorithm is used, for example, to analyze theft or aggressive behavior and recognize it as trouble.

[0765] 4. Emotion recognition means

[0766] The server analyzes the subject's facial expressions and voice to recognize their emotional state.

[0767] To do this, it uses an AI model that extracts emotions from video and audio data.

[0768] 5. Caution statement generation means

[0769] The server generates appropriate warning messages based on the data obtained from the emotion recognition means.

[0770] For example, if the subject is feeling anxious, the system will automatically generate the phrase "Please remain calm."

[0771] 6. Means of speech

[0772] The warning messages sent from the server are spoken through speakers installed within the commercial facility or via a smartphone app.

[0773] Speech is delivered in a tone and language that corresponds to the emotion.

[0774] 7. Means of notification

[0775] The generated warning message is sent to store employees and the target person via a smartphone notification system.

[0776] For example, if any theft is detected, the store security team is immediately notified.

[0777] Specific examples

[0778] As a concrete example, we will explain the detection and response to theft behavior in a large shopping mall.

[0779] 1. Video monitoring and transmission: Surveillance cameras in the shopping mall capture video and transmit the data to a server. The cameras have high resolution and wide coverage.

[0780] 2. Video analysis: The server uses an AI model to extract suspicious movements and crowding of people from the received video data as features.

[0781] 3. Trouble detection: AI analyzes these features and detects, for example, theft. It raises a warning flag depending on the type of trouble.

[0782] 4. Emotion recognition: The server utilizes emotion recognition means to analyze the target person's facial expressions and voice to recognize their emotional state (e.g., anxiety, anger).

[0783] 5. Generating warning messages: The server generates appropriate warning messages (e.g., "Please stay calm") based on the emotion data.

[0784] 6. Caution notification: The generated caution message is sent directly to the target person via speech, and is also sent to employees and security teams via their smartphone notification system.

[0785] Prompt Sentence Examples

[0786] An example of a prompt to input to a generative AI model is as follows:

[0787] Prompt: If you detect any unusual behavior in the mall, create a message to encourage the person you spot to remain calm.

[0788] Response: "Remain calm and try to resolve the issue. If you need assistance, please speak to a member of staff."

[0789] The above configuration enables quick and effective detection of troubles and violations of etiquette within commercial facilities, and makes it possible to take appropriate measures when they occur.

[0790] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[0791] Step 1:

[0792] The server acquires real-time video footage of the commercial facility from the monitoring means (surveillance cameras). This video footage is high resolution and contains a wide range of data. The input is the video data acquired from the surveillance cameras, and the output is the real-time video data sent to the server.

[0793] Step 2:

[0794] The server analyzes the received video data using a video analysis method. Specifically, it uses a convolutional neural network (CNN) to extract features such as abnormal movements and crowding of people in the video. The input is the video data acquired in step 1, and the output is the extracted feature data.

[0795] Step 3:

[0796] The server uses the trouble detection means to detect troublesome behavior based on the features obtained from the video analysis means. For example, theft or aggressive behavior is analyzed using an algorithm and recognized as trouble. The input is the feature data extracted in step 2, and the output is the type of trouble detected and its location information.

[0797] Step 4:

[0798] The server uses emotion recognition means to analyze the facial expressions and voices of the people involved in the trouble and recognize their emotional state. Specifically, it identifies the emotional state using an AI model that extracts emotions. The input is a list of people involved in the trouble detected in step 3 and their video data, and the output is the emotional state of each person.

[0799] Step 5:

[0800] The server uses the warning message generation means to generate appropriate warning messages based on the emotion data obtained from the emotion recognition means. For example, it generates a message such as "Please stay calm to avoid problems." The input is the emotion data and trouble information obtained in step 4, and the output is the generated warning message.

[0801] Step 6:

[0802] The server sends the generated warning message to the speech means (directional speaker) and notification means (smartphone notification system). The warning message is spoken directly within the commercial facility through the speech means, and simultaneously notifies the target person or employee via their smartphone. The input is the warning message generated in step 5, and the output is the speech within the commercial facility and the smartphone notification.

[0803] These processing steps enable quick and effective detection of troubles and violations of etiquette within commercial facilities, and appropriate measures to be taken.

[0804] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0805] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0806] In the above embodiment, an example in which the specific processing is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the smart glasses 214.

[0807] [Third embodiment]

[0808] FIG. 5 shows an example of the configuration of a data processing system 310 according to the third embodiment.

[0809] 5, the data processing system 310 includes the data processing device 12 and a headset type terminal 314. An example of the data processing device 12 is a server.

[0810] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0811] The headset type terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a display 343. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the display 343 are also connected to the bus 52.

[0812] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[0813] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[0814] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[0815] Fig. 6 shows an example of the main functions of the data processing device 12 and the headset type terminal 314. As shown in Fig. 6, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[0816] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0817] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0818] In the headset type terminal 314, a reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0819] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the headset type terminal 314 will be referred to as the "terminal."

[0820] The present invention is a system that uses AI technology to monitor troubles and violations of etiquette on public transportation and automatically urges appropriate caution, and can be implemented in the following forms.

[0821] System configuration

[0822] The system includes the following elements:

[0823] 1. Terminal (surveillance camera)

[0824] 2. Server

[0825] 3. Terminal (directional speaker)

[0826] Program processing and natural language explanation

[0827] Video monitoring and transmission

[0828] The devices (surveillance cameras) capture images in real time inside vehicles and public transportation. These images are sent to a server via a network. The surveillance cameras are high-resolution and are installed to cover a wide area.

[0829] Video analysis

[0830] The server analyzes the received video data and uses an AI model (such as a convolutional neural network) to extract features from the video, such as suspicious movements, crowds of people, and unusual sounds. This allows for the rapid and accurate detection of unusual behavior and abnormal situations in the video.

[0831] Trouble detection

[0832] The server then runs a problem detection algorithm based on the extracted features. For example, if abnormally close movements of people or aggressive behavior are detected, it will be recognized as a specific problem (e.g., sexual harassment, assault, etc.) and a warning flag will be raised.

[0833] Generate warning messages

[0834] The server then generates a warning message based on the type of trouble, location, and behavior of the target. For example, if harassment is detected on a crowded train, the server generates a warning message saying, "Please stop harassing."

[0835] Caution speech

[0836] The generated warning message is sent from the server to a terminal (directional speaker), which then speaks the message and delivers the warning directly to the target person. Because the directional speaker can focus sound in a specific direction, it can deliver the message only to the target person without causing unnecessary inconvenience to surrounding passengers.

[0837] Specific examples

[0838] As a concrete example, we will explain how to detect and respond to sexual harassment on a crowded train.

[0839] 1. Video monitoring and transmission: The surveillance camera inside the vehicle captures the video and transmits it to the server. The video is high resolution and covers a wide area.

[0840] 2. Video analysis: Using an AI model, the server extracts features from the video received, such as suspicious movements and crowds of people, and analyzes people's movements and relative positions in detail.

[0841] 3. Trouble detection: AI analyzes these features to detect abnormal behavior, such as sexual harassment, which then raises a warning flag.

[0842] 4. Generation of warning message: When a problem is detected, the server generates a warning message such as "Please stop harassing the user."

[0843] 5. Caution speech: The server sends the generated caution text to the terminal (directional speaker), and the speaker speaks the caution to the target person.

[0844] In this way, this system effectively manages troubles and violations of etiquette on public transport, providing a safe and comfortable riding environment.

[0845] The processing flow will be explained below.

[0846] Step 1:

[0847] The device (surveillance camera) captures real-time images from inside a vehicle or public transport vehicle. The surveillance camera is set to capture high-resolution images and cover a wide area.

[0848] Step 2:

[0849] The video data captured by the terminal (surveillance camera) is compressed and sent over the network to the server, where it is converted into an appropriate format for efficient data transfer.

[0850] Step 3:

[0851] The server decodes the received video data and analyzes it in real time using dedicated video analysis software.

[0852] Step 4:

[0853] The server uses an AI model to extract features from the video data. These features include suspicious movements, crowds of people, and unusual sounds. Convolutional neural networks (CNNs) are used for extraction.

[0854] Step 5:

[0855] The server evaluates the extracted features to detect abnormal behavior and signs of trouble, where machine learning algorithms identify abnormal patterns based on past data.

[0856] Step 6:

[0857] The server sets a warning flag for any abnormal behavior it detects. For example, if assault or sexual harassment is detected, a warning flag for that behavior will be set.

[0858] Step 7:

[0859] The server collects details of the incident (type of behavior, location, characteristics of the victim, etc.) and uses conversational AI to generate a warning message, such as "Mobbing behavior has been detected. Please stop immediately."

[0860] Step 8:

[0861] The server converts the generated warning message into an appropriate data format and transmits it to the terminal (directional speaker) via the network.

[0862] Step 9:

[0863] The terminal (directional speaker) analyzes the received warning message and speaks it out to the target person. The directional speaker focuses the sound in a specific direction, warning the person without disturbing surrounding passengers.

[0864] Step 10:

[0865] The server continues to analyze the surveillance camera footage and monitors the subject's reaction to the attention, checking for changes in the subject's movements and behavior.

[0866] Step 11:

[0867] If the user does not follow the warning, the server generates an additional warning message and sends it to the device (directional speaker) again. If necessary, the warning message is changed to a stronger one.

[0868] Step 12:

[0869] If the server cannot resolve the problem, it automatically sends an emergency notification to the police and the driver, which includes detailed information about the problem, enabling a prompt response.

[0870] These are the specific processing steps of the program for this system. In this way, by linking each step together, troubles within public transportation systems can be effectively prevented and passenger safety can be ensured.

[0871] Example 1

[0872] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[0873] There is a need to monitor troubles and violations of etiquette on public transportation and respond quickly and appropriately. However, current systems have difficulty detecting troubles in real time and automatically issuing appropriate warnings. Rapid detection and response to troubles is particularly important on packed transportation, and an effective system for this purpose is required.

[0874] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0875] In this invention, the server includes a monitoring means, a video analysis means, a trouble detection means, a warning message generation means, and a speech means, which makes it possible to detect troubles and manner violations on public transportation in real time and to promptly and appropriately warn people.

[0876] "Monitoring means" refers to devices for capturing images inside public transportation, and mainly includes cameras.

[0877] The "video analysis means" is a device that has the function of analyzing received video data and extracting features such as suspicious movements, crowds of people, and abnormal sounds.

[0878] The "trouble detection means" refers to an algorithm or device for detecting trouble based on the feature amount extracted by the video analysis means.

[0879] The "warning message generation means" is a device that has the function of using a generative AI model to generate appropriate warning messages in response to detected problems.

[0880] The "utterance means" is a device for uttering the generated warning message to a specific target person, and particularly includes a directional speaker.

[0881] This invention is a system that uses AI technology to monitor troubles and bad manners on public transportation and automatically warns passengers of appropriate precautions. This system is implemented using the following hardware and software.

[0882] System configuration

[0883] The system includes the following elements:

[0884] 1. Terminal (surveillance camera)

[0885] 2. Server

[0886] 3. Terminal (directional speaker)

[0887] Video acquisition and transmission

[0888] Terminals (surveillance cameras) are installed inside public transportation facilities and capture high-resolution video in real time. The surveillance cameras are installed to cover a wide area, making them particularly useful for detecting trouble on crowded trains. The captured video data is sent to a server via a network. The transmitted data is encrypted to ensure security.

[0889] Video analysis

[0890] After the video data is sent to the server, the server receives the data. The server is equipped with a convolutional neural network (CNN) built using TensorFlow and PyTorch. The server uses this AI model to analyze the received video data and extract features such as suspicious movements, crowds of people, and unusual sounds. The analysis is performed in real time, dividing each video frame and extracting features using a specific algorithm.

[0891] Trouble detection

[0892] The server runs a problem detection algorithm based on the features obtained from the video analysis. For example, if abnormally close movements of people or aggressive behavior are detected, it will recognize this as a specific problem (e.g., sexual harassment or assault) and raise a warning flag. For example, if the server recognizes that a specific passenger is making inappropriate contact with another passenger on a crowded train, a warning flag will be raised.

[0893] Generate warning messages

[0894] When a warning flag is raised on the server, the server generates a warning message based on the problem. The message is generated using a generative AI model (e.g., GPT-3) based on a specific prompt. The warning message is generated using the following prompt:

[0895] "Trouble type: sexual harassment

[0896] Location: A crowded train

[0897] Situation: Passenger A makes inappropriate contact with another passenger B."

[0898] Based on this prompt, an appropriate warning message such as "Please stop harassing others" is generated.

[0899] Caution speech

[0900] The warning message generated by the server is sent to a terminal (directional speaker). The terminal (directional speaker) can focus sound in a specific direction, so the message is delivered to the target person only, without causing unnecessary inconvenience to surrounding passengers. For example, the speaker can issue a voice message saying, "Please stop molesting me" to a passenger sitting in a specific seat on a crowded train.

[0901] This system quickly and accurately detects any trouble or violations of etiquette on public transport in real time and provides appropriate warnings, resulting in a safe and comfortable riding environment.

[0902] The embodiment of the invention is configured as described above, which makes it possible to effectively monitor and quickly respond to problems within public transportation.

[0903] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0904] Program processing flow

[0905] Step 1: Acquire and transmit video

[0906] Users capture video using terminals (surveillance cameras) installed within public transportation. The cameras capture high-resolution video in real time and transmit the video data to a server via a network. The video captured by the camera is continuously divided into frames and transmitted to the server. During this process, the video data is encrypted to ensure security.

[0907] Input: Real-time camera footage

[0908] Output: Video data sent to the server

[0909] Step 2: Receiving the video and preparing for analysis

[0910] The server receives video data sent from the terminal (surveillance camera) via the network. The received video data is divided into frames and saved.

[0911] Input: Video data transmitted over the network

[0912] Output: Frame-by-frame video data stored on the server

[0913] Step 3: Analyzing the footage

[0914] The server analyzes the stored video data using a convolutional neural network (CNN). Specifically, it uses AI frameworks such as TensorFlow and PyTorch to perform real-time data analysis. As a result of the analysis, features such as suspicious movements, crowds of people, and unusual sounds are extracted. These features are used as input for predicting trouble.

[0915] Input: Frame-by-frame video data stored on the server

[0916] Output: Features such as suspicious movements, crowds of people, and abnormal sounds

[0917] Step 4: Detecting the problem

[0918] The server runs a problem detection algorithm based on the features obtained through the analysis. For example, if abnormal proximity or aggressive behavior is detected, it recognizes this as a specific problem (e.g., sexual harassment or assault) and raises a warning flag. This allows specific abnormal behavior to be identified in real time.

[0919] Input: Features extracted by analysis

[0920] Output: Trouble detection flag

[0921] Step 5: Generate warning text

[0922] The server generates a warning message when a warning flag is raised. The server uses a generative AI model (e.g., GPT-3) to generate an appropriate warning message based on a specific prompt. For example, given the prompt "Trouble type: sexual harassment, location: crowded train, situation: passenger A making inappropriate contact with another passenger B," the server generates a warning message saying, "Please stop sexual harassment."

[0923] Input: Trouble detection flag, prompt text

[0924] Output: Generated warning message

[0925] Step 6: Send and speak the warning message

[0926] The server generates a warning message and sends it to the device (directional speaker). The directional speaker can focus sound in a specific direction, delivering the message only to the target person without disturbing surrounding passengers. The speaker then speaks a warning message to the target person, saying, "Please stop harassing."

[0927] Input: Generated warning text

[0928] Output: Warning message spoken from the device (directional speaker)

[0929] In this way, this system can monitor troubles and violations of etiquette on public transportation in real time and automatically provide appropriate warnings.

[0930] (Application example 1)

[0931] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[0932] Systems for effectively monitoring and quickly responding to incidents and violations of etiquette on public transportation are limited, making it difficult for security staff to quickly grasp the situation and take appropriate action. Effective methods for focusing sound in a specific direction to draw attention are also lacking. Furthermore, existing systems struggle to provide real-time notifications and visual assistance, hindering security improvements.

[0933] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0934] In this invention, the server includes a monitoring means, an image analysis means, a trouble detection means, a warning message generation means, an output means, a notification means, and a support display means, which makes it possible to monitor and analyze troubles and manner violations on public transportation in real time, quickly detect troubles, generate and output appropriate warning messages, and immediately notify security staff and provide visual support information.

[0935] "Surveillance" refers to the equipment and technology used to capture footage inside public transport.

[0936] "Video analysis means" refers to devices and techniques for extracting and analyzing features from acquired video data.

[0937] "Trouble detection means" refers to devices and technologies for detecting trouble by discovering suspicious movements or abnormalities from features extracted by video analysis means.

[0938] "Warning message generation means" refers to a device and technology for generating appropriate warning messages in response to a detected problem.

[0939] "Speech means" refers to a device and technology for outputting the generated warning message as voice.

[0940] "Notification means" refers to devices and technologies used to notify security staff and other relevant parties of detected problems and generated warning messages.

[0941] "Support display means" refers to devices and technologies that visually provide necessary information to security staff when they respond.

[0942] The present invention is a system that uses AI technology to monitor troubles and bad manners on public transportation and automatically urges appropriate cautions. An embodiment of the present invention will be described.

[0943] System configuration

[0944] The system includes the following elements:

[0945] 1. Surveillance methods (e.g., built-in cameras in smart glasses)

[0946] 2. Server

[0947] 3. Speaking means (e.g., directional speaker)

[0948] 4. Means of notification

[0949] 5. Supportive display means (e.g., smart glasses display)

[0950] Program processing

[0951] Video Acquisition

[0952] The monitoring tool will capture real-time footage of public transport vehicles, for example, through cameras built into smart glasses, allowing security staff to keep up to date with the latest developments.

[0953] Video transmission

[0954] The captured images are then sent to a server via Wi-Fi or mobile data networks, where high-resolution images are transmitted for improved analysis accuracy.

[0955] Video analysis

[0956] The server analyzes the received video data using an AI model (e.g., convolutional neural network), which extracts features such as suspicious movements, crowds of people, and abnormal sounds.

[0957] Trouble detection

[0958] The server then runs a problem detection algorithm based on the extracted features. For example, if unusually close movements of people or aggressive behavior are detected, this is recognized as a specific problem and a warning flag is raised.

[0959] Generate warning messages

[0960] The server generates a warning message based on the type of incident, location, and the target's behavior. For example, if a sexual assault is detected on a crowded train, a warning message saying "Please stop sexual assault" will be automatically generated.

[0961] Caution speech

[0962] The generated warning message is sent from the server to a speech device (directional speaker), which then speaks the message and delivers the warning directly to the target person. This allows the message to be delivered only to specific targets without disturbing surrounding passengers.

[0963] Notifications and Support Displays

[0964] At the same time, the server notifies security staff of any detected problems or warning messages via notification means, and visual support information is displayed on the smart glasses' display to help them respond quickly and appropriately.

[0965] Specific examples

[0966] As a concrete example, we will explain how to detect and respond to sexual harassment on a crowded train.

[0967] 1. Image acquisition: The built-in camera in the smart glasses captures images and sends them to the server. High-resolution images cover a wide area.

[0968] 2. Video analysis: Using an AI model, the server extracts features from the video received, such as suspicious movements and crowds of people, and performs a detailed analysis of people's movements and relative positions.

[0969] 3. Trouble detection: Abnormal behavior, such as sexual harassment, is detected and a warning flag is raised.

[0970] 4. Generation of warning message: The server generates a warning message saying, "Please stop harassing."

[0971] 5. Caution message generation and notification: The server sends the generated caution message to the directional speaker and issues a message, while simultaneously providing visual support information to security staff.

[0972] Prompt Sentence Examples

[0973] "Please detect abnormal movements and crowding of people on crowded trains and notify us so that appropriate measures can be taken."

[0974] This allows any trouble or violations of etiquette on public transport to be detected immediately, allowing security staff to respond quickly and appropriately.

[0975] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[0976] Step 1:

[0977] Users wear smart glasses while in public transport, and the built-in camera captures real-time images of passengers. The input is real-time images from inside the public transport vehicle, and the output is high-resolution video data.

[0978] Step 2:

[0979] The device (smart glasses) transmits the captured video to a server via Wi-Fi or a mobile data network. The input is high-resolution video data, and the output is the video data transmitted to the server.

[0980] Step 3:

[0981] The server analyzes the received video data using an AI model (for example, a convolutional neural network). The input is the transmitted video data, and the output is the extracted features (suspicious movements, crowds of people, abnormal sounds, etc.). Specific operations include inputting the video data into the AI ​​model and extracting the features.

[0982] Step 4:

[0983] The server runs a trouble detection algorithm based on the extracted features. The input is the extracted features, and the output is the trouble detection result (e.g., detection of sexual harassment). Specifically, it compares the extracted features with pre-set criteria to detect anomalies.

[0984] Step 5:

[0985] The server generates an appropriate warning message based on the results of the trouble detection. The input is the trouble detection result, and the output is the generated warning message (e.g., "Please stop harassing others"). Specifically, the server inputs the type of trouble and location information into a predefined warning message template to generate the appropriate warning message.

[0986] Step 6:

[0987] The server transmits the generated warning message to the speech generator (directional speaker). The input is the generated warning message, and the output is the data to be transmitted to the directional speaker. Specifically, the server transmits an appropriate message to the directional speaker.

[0988] Step 7:

[0989] The speech generator (directional speaker) speaks the transmitted warning message as a voice in a specific direction. The input is the transmitted data, and the output is the actual voice warning message. Specifically, it uses a voice synthesis function to speak a specific message to the target.

[0990] Step 8:

[0991] At the same time, the server uses a notification means to notify security staff of the detected trouble and the generated warning message. The input is the trouble detection result and the warning message, and the output is visual support information displayed on the smart glasses' display. Specifically, the server sends the notification data to the smart glasses, and appropriate support information is displayed on the display.

[0992] Prompt Sentence Examples

[0993] "Detect abnormal movements and crowding of people inside a moving train and provide appropriate countermeasures."

[0994] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[0995] The present invention is a system that uses AI technology to monitor troubles and violations of etiquette on public transportation and automatically urges appropriate cautions, and by combining it with an emotion engine that recognizes the user's emotions, it is intended to enhance the effectiveness of preventing troubles. It can be implemented in the following forms.

[0996] System configuration

[0997] The system includes the following elements:

[0998] 1. Terminal (surveillance camera)

[0999] 2. Server

[1000] 3. Terminal (directional speaker)

[1001] 4. Emotion Engine

[1002] Program processing and natural language explanation

[1003] Video monitoring and transmission

[1004] The devices (surveillance cameras) capture images in real time inside vehicles and public transportation. These images are sent to a server via a network. The surveillance cameras are high-resolution and are installed to cover a wide area.

[1005] Video analysis

[1006] The server analyzes the received video data and uses an AI model (such as a convolutional neural network) to extract features from the video, such as suspicious movements, crowds of people, and unusual sounds. This allows for the rapid and accurate detection of unusual behavior and abnormal situations in the video.

[1007] Trouble detection

[1008] The server then runs a problem detection algorithm based on the extracted features. For example, if abnormally close movements of people or aggressive behavior are detected, it will be recognized as a specific problem (e.g., sexual harassment, assault, etc.) and a warning flag will be raised.

[1009] Emotion recognition

[1010] When a problem is detected, the emotion engine analyzes the target's facial expressions and voice tone to recognize their emotional state (e.g., anger, anxiety, fear, etc.). The emotion engine uses an AI model to extract emotions from video and audio data.

[1011] Generate warning messages

[1012] Next, the server generates a warning message based on the emotional data of the target person recognized by the emotion engine. The conversational AI automatically creates an appropriate warning message based on the type of trouble, location, and the target person's emotional state. For example, if the target person is expressing anger, a message such as "Avoid trouble. Remain calm" will be generated.

[1013] Caution speech

[1014] The generated warning message is sent from the server to the device (directional speaker), which then directly warns the target person. The speaker also speaks in a tone and language that corresponds to the emotional state recognized by the emotion engine. This allows for more effective warning communication to the target person.

[1015] Specific examples

[1016] As a concrete example, we will explain how to detect and respond to problems on a crowded train.

[1017] 1. Video monitoring and transmission: The surveillance camera inside the vehicle captures the video and transmits it to the server. The video is high resolution and covers a wide area.

[1018] 2. Video analysis: Using an AI model, the server extracts features from the video received, such as suspicious movements and crowds of people, and analyzes people's movements and relative positions in detail.

[1019] 3. Trouble detection: AI analyzes these features to detect abnormal behavior, such as sexual harassment, which then raises a warning flag.

[1020] 4. Emotion Recognition: The emotion engine analyzes the facial expressions and voice of the target when a problem is detected and recognizes their emotional state. For example, if the target shows feelings of anxiety or anger.

[1021] 5. Generating warning messages: The server takes into account the recognized emotional data and generates warning messages such as "Please stay calm to avoid trouble" rather than "Please stop harassing others."

[1022] 6. Speech of warning: The server sends the generated warning message to the device (directional speaker), and the speaker speaks the warning to the target person. At the same time, the message is conveyed in a tone and language that matches the user's emotions.

[1023] These are the specific processing steps of the program of this system. In this way, by linking each step together, it is possible to effectively prevent trouble on public transportation and ensure the safety of passengers.

[1024] The processing flow will be explained below.

[1025] Step 1:

[1026] The device (surveillance camera) captures real-time images from inside a vehicle or public transport vehicle. The surveillance camera is set to capture high-resolution images and cover a wide area.

[1027] Step 2:

[1028] The video data captured by the terminal (surveillance camera) is compressed and sent over the network to the server, where it is converted into an appropriate format for efficient data transfer.

[1029] Step 3:

[1030] The server decodes the received video data and analyzes it in real time using dedicated video analysis software.

[1031] Step 4:

[1032] The server uses an AI model to extract features from the video data. These features include suspicious movements, crowds of people, and unusual sounds. Convolutional neural networks (CNNs) are used for extraction.

[1033] Step 5:

[1034] The server evaluates the extracted features to detect abnormal behavior and signs of trouble, where machine learning algorithms identify abnormal patterns based on past data.

[1035] Step 6:

[1036] The server sets a warning flag for any abnormal behavior it detects. For example, if assault or sexual harassment is detected, a warning flag for that behavior will be set.

[1037] Step 7:

[1038] When a problem is detected, the emotion engine analyzes the target's facial expressions and voice tone to recognize their emotional state (e.g., anger, anxiety, fear, etc.). The emotion engine uses an AI model to extract emotions from video and audio data.

[1039] Step 8:

[1040] The server organizes the details of the trouble (type of behavior, location, characteristics of the target, etc.) and the emotional state recognized by the emotion engine, and generates a warning message using conversational AI. For example, if the target shows anger, a message such as "Avoid trouble. Remain calm" will be generated.

[1041] Step 9:

[1042] The server converts the generated warning message into an appropriate data format and transmits it to the terminal (directional speaker) via the network.

[1043] Step 10:

[1044] The device (directional speaker) analyzes the received warning message and speaks it out to the target passenger. The directional speaker focuses the sound in a specific direction, warning the passenger without disturbing other passengers. The emotion engine also recognizes the passenger's emotional state and uses a tone and wording that matches that state, providing more effective warnings.

[1045] Step 11:

[1046] The server continues to analyze the surveillance camera footage and monitors the subject's reaction to the attention, checking for changes in the subject's movements and behavior.

[1047] Step 12:

[1048] If the user does not follow the warning, the server generates an additional warning message and sends it to the device (directional speaker) again. If necessary, the warning message is changed to a stronger one.

[1049] Step 13:

[1050] If the server cannot resolve the problem, it automatically sends an emergency notification to the police and the driver, which includes detailed information about the problem, enabling a prompt response.

[1051] Example 2

[1052] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1053] Trouble and bad manners on public transportation are serious issues that threaten passenger safety and comfort. Current monitoring systems have limitations in detecting suspicious movements and abnormal behavior, and are unable to respond appropriately while taking into account the emotional state of passengers. This makes it difficult to take prompt and effective measures when trouble occurs. There is also a lack of technology that can recognize passengers' emotions and psychological state and respond accordingly.

[1054] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[1055] In this invention, the server includes a monitoring means, a video analysis means, a trouble detection means, an emotion recognition means, a warning message generation means, and a speech utterance means. This makes it possible to monitor video in public transportation in real time, detect suspicious movements and crowding of people, and recognize the emotional state of passengers, and generate and speak appropriate warning messages based on those emotions.

[1056] "Monitoring means" refers to devices for capturing images inside public transportation, and primarily involves the use of surveillance cameras.

[1057] "Video analysis means" refers to technology or equipment that analyzes acquired video data and extracts features such as suspicious movements, crowds of people, and abnormal sounds.

[1058] "Trouble detection means" refers to algorithms and technologies for detecting abnormal behavior or trouble based on extracted features.

[1059] "Emotion recognition means" refers to technology or equipment that analyzes the facial expressions and tone of voice of a person when a problem is detected, and recognizes the emotional state of the person.

[1060] The "warning message generation means" refers to a technique or device for generating appropriate warning messages based on the recognized emotional state.

[1061] The "utterance means" is a device for uttering the generated warning message in a specific direction to convey it to the target person.

[1062] This invention is a system that uses AI technology to monitor troubles and bad manners on public transportation and automatically provides appropriate warnings. The system includes the following elements:

[1063] monitoring means

[1064] The terminal (surveillance camera) captures images of the inside of public transportation in real time. The surveillance camera has high resolution and is installed to cover a wide area. The captured images are sent to a server via a network.

[1065] Video analysis methods

[1066] The server analyzes the received video data, using AI models such as convolutional neural networks to extract features from the video, such as suspicious movements, crowds of people, and unusual sounds.

[1067] Trouble detection methods

[1068] The server runs a problem detection algorithm based on the extracted features. If abnormally close movements of people or aggressive behavior are detected, it will recognize it as a specific problem (e.g., assault) and raise a warning flag.

[1069] emotion recognition means

[1070] When the emotion engine detects a problem, it analyzes the target's facial expressions and vocal tone, using an AI model that extracts emotions from video and audio data to recognize the target's emotional state (e.g., anger, anxiety, fear, etc.).

[1071] Caution statement generation means

[1072] The server generates warning messages based on the emotional data of the target person recognized by the emotion engine. The conversational AI automatically creates appropriate warning messages based on the type of trouble, location, and the target person's emotional state. For example, if the target person is expressing anger, a message such as "Avoid trouble. Remain calm" will be generated.

[1073] Means of speech

[1074] The generated warning message is sent from the server to the device (directional speaker). The directional speaker then directly warns the target person. The speaker also speaks in a tone and language that corresponds to the emotional state recognized by the emotion engine. This operation allows for more effective warning communication to the target person.

[1075] Specific examples

[1076] As a concrete example, we will explain how to detect and respond to problems on a crowded train.

[1077] 1. Video monitoring and transmission: The terminal (surveillance camera) captures images from inside the train and transmits them to the server in real time. The camera has high resolution and covers a wide area.

[1078] 2. Video analysis: The server uses an AI model to analyze the video data it receives and extracts suspicious movements and crowds of people as features.

[1079] 3. Trouble detection: The server uses AI to analyze features and detect abnormal approaches or unnatural movements. This allows it to detect troubles such as sexual harassment and raise a warning flag.

[1080] 4. Emotion recognition: The emotion engine analyzes the subject's facial expressions and voice to recognize emotions such as anxiety or anger.

[1081] 5. Generating warning messages: Based on the emotional data, the server generates warning messages such as "Please stay calm to avoid trouble" rather than "Please stop harassing others."

[1082] 6. Speech warning: The server generates a warning message and sends it to the device (directional speaker), which then speaks it to the target person. By conveying the warning in a tone and using words that match the target person's emotions, a more effective response can be achieved.

[1083] Examples of prompt statements

[1084] Below is an example of a prompt sentence to input to the generative AI model.

[1085] "Please explain the system for detecting suspicious behavior on public transportation and taking appropriate action. For example, please explain in detail how to respond to sexual harassment on a crowded train."

[1086] The system serves as an effective means of improving safety within public transport.

[1087] The flow of the identification process in the second embodiment will be described with reference to FIG.

[1088] Step 1:

[1089] Video monitoring and transmission

[1090] explanation

[1091] The device (surveillance camera) captures images of the inside of public transportation in real time, and the captured image data is sent to a server via a network.

[1092] Specific actions

[1093] Input: Public transport footage

[1094] Data processing: The surveillance camera sensor captures the video and obtains detailed video data in high resolution.

[1095] Output: Video data sent to the server in real time

[1096] Step 2:

[1097] Video analysis

[1098] explanation

[1099] The server analyzes the received video data and uses an AI model (such as a convolutional neural network) to extract features such as suspicious movements, crowds of people, and unusual sounds from the video.

[1100] Specific actions

[1101] Input: Received video data

[1102] Data processing: Decoding video data, analyzing each frame, and using AI models to identify suspicious movements, crowds of people, and unusual sounds

[1103] Output: Extracted feature data

[1104] Step 3:

[1105] Trouble detection

[1106] explanation

[1107] The server runs a trouble-detection algorithm based on the extracted features. For example, if abnormally close movements of people or aggressive behavior are detected, it will recognize the problem and raise a warning flag.

[1108] Specific actions

[1109] Input: Extracted feature data

[1110] Data processing: Input feature data into the trouble detection algorithm to detect abnormal patterns

[1111] Output: Trouble detection flag

[1112] Step 4:

[1113] Emotion recognition

[1114] explanation

[1115] The emotion engine recognizes the emotional state of the person who detected the problem by analyzing facial expressions and tone of voice (e.g., anger, anxiety, fear, etc.).

[1116] Specific actions

[1117] Input: Subject's facial expression data and voice data

[1118] Data processing: The emotion engine uses AI models to analyze facial expressions and voice to identify emotional states.

[1119] Output: Recognized emotion data

[1120] Step 5:

[1121] Generate warning messages

[1122] explanation

[1123] The server generates warning messages based on the target person's emotional data recognized by the emotion engine. The conversational AI creates warning messages based on the type of trouble, location, and the target person's emotional state.

[1124] Specific actions

[1125] Input: Recognized emotion data, type of trouble, location information

[1126] Data processing: Input prompts into the generative AI model to generate appropriate warning messages

[1127] Output: Generated warning message

[1128] Step 6:

[1129] Caution speech

[1130] explanation

[1131] The generated warning message is sent from the server to the device (directional speaker), which then speaks the message to the target person.

[1132] Specific actions

[1133] Input: Generated warning text data

[1134] Data processing: A directional speaker uses a speech synthesis engine to play back warning messages and speak them in a specific direction.

[1135] Output: Audio warning to the target

[1136] In this way, each step works together to effectively prevent trouble on public transport and ensure passenger safety.

[1137] (Application example 2)

[1138] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1139] In commercial facilities, customers often violate etiquette and cause trouble, which often has a negative impact on store operations and other customers. The purpose of this invention is to maintain order and safety within commercial facilities by quickly and effectively detecting such trouble and violations of etiquette and automatically issuing appropriate warnings.

[1140] The identification process by the identification processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes a monitoring means, a video analysis means, a trouble detection means, an emotion recognition means, a warning message generation means, a speech means, and a notification means. This makes it possible to automatically detect troubles and etiquette violations within a commercial facility and issue appropriate warnings.

[1141] The "monitoring means" is a device that acquires images of the inside of a commercial facility in real time.

[1142] The "video analysis means" is a device or software for processing acquired video data and extracting specific features.

[1143] The "trouble detection means" is a device or algorithm that detects abnormal movements or behaviors based on features extracted from the video analysis means.

[1144] An "emotion recognition means" is a device or algorithm that analyzes the facial expressions and voice of a detected subject and recognizes their emotional state.

[1145] The "warning message generation means" is a device or software that generates appropriate warning messages for the target person based on the data obtained from the emotion recognition means.

[1146] The "utterance means" is a device for uttering the generated warning message as voice.

[1147] The "notification means" is a device or software that transmits the generated warning message to a smartphone or other device.

[1148] This invention is a system that detects troubles and bad manners in commercial facilities and automatically issues appropriate warnings. This system includes seven main means: "monitoring means," "video analysis means," "trouble detection means," "emotion recognition means," "warning message generation means," "speech means," and "notification means."

[1149] System configuration

[1150] The system includes the following elements:

[1151] 1. Monitoring measures

[1152] Camera devices and smartphone cameras installed within commercial facilities are used to capture images of the facility in real time.

[1153] 2. Video analysis methods

[1154] The captured video data is received and specific features are extracted using an AI model (e.g., convolutional neural network).

[1155] The features include abnormal movements, crowding of people, and abnormal behavior.

[1156] 3. Trouble detection methods

[1157] The server detects abnormal behavior based on the features obtained from the video analysis means.

[1158] In this case, a trouble detection algorithm is used, for example, to analyze theft or aggressive behavior and recognize it as trouble.

[1159] 4. Emotion recognition means

[1160] The server analyzes the subject's facial expressions and voice to recognize their emotional state.

[1161] To do this, it uses an AI model that extracts emotions from video and audio data.

[1162] 5. Caution statement generation means

[1163] The server generates appropriate warning messages based on the data obtained from the emotion recognition means.

[1164] For example, if the subject is feeling anxious, the system will automatically generate the phrase "Please remain calm."

[1165] 6. Means of speech

[1166] The warning messages sent from the server are spoken through speakers installed within the commercial facility or via a smartphone app.

[1167] Speech is delivered in a tone and language that corresponds to the emotion.

[1168] 7. Means of notification

[1169] The generated warning message is sent to store employees and the target person via a smartphone notification system.

[1170] For example, if any theft is detected, the store security team is immediately notified.

[1171] Specific examples

[1172] As a concrete example, we will explain the detection and response to theft behavior in a large shopping mall.

[1173] 1. Video monitoring and transmission: Surveillance cameras in the shopping mall capture video and transmit the data to a server. The cameras have high resolution and wide coverage.

[1174] 2. Video analysis: The server uses an AI model to extract suspicious movements and crowding of people from the received video data as features.

[1175] 3. Trouble detection: AI analyzes these features and detects, for example, theft. It raises a warning flag depending on the type of trouble.

[1176] 4. Emotion recognition: The server utilizes emotion recognition means to analyze the target person's facial expressions and voice to recognize their emotional state (e.g., anxiety, anger).

[1177] 5. Generating warning messages: The server generates appropriate warning messages (e.g., "Please stay calm") based on the emotion data.

[1178] 6. Caution notification: The generated caution message is sent directly to the target person via speech, and is also sent to employees and security teams via their smartphone notification system.

[1179] Prompt Sentence Examples

[1180] An example of a prompt to input to a generative AI model is as follows:

[1181] Prompt: If you detect any unusual behavior in the mall, create a message to encourage the person you spot to remain calm.

[1182] Response: "Remain calm and try to resolve the issue. If you need assistance, please speak to a member of staff."

[1183] The above configuration enables quick and effective detection of troubles and violations of etiquette within commercial facilities, and makes it possible to take appropriate measures when they occur.

[1184] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[1185] Step 1:

[1186] The server acquires real-time video footage of the commercial facility from the monitoring means (surveillance cameras). This video footage is high resolution and contains a wide range of data. The input is the video data acquired from the surveillance cameras, and the output is the real-time video data sent to the server.

[1187] Step 2:

[1188] The server analyzes the received video data using a video analysis method. Specifically, it uses a convolutional neural network (CNN) to extract features such as abnormal movements and crowding of people in the video. The input is the video data acquired in step 1, and the output is the extracted feature data.

[1189] Step 3:

[1190] The server uses the trouble detection means to detect troublesome behavior based on the features obtained from the video analysis means. For example, theft or aggressive behavior is analyzed using an algorithm and recognized as trouble. The input is the feature data extracted in step 2, and the output is the type of trouble detected and its location information.

[1191] Step 4:

[1192] The server uses emotion recognition means to analyze the facial expressions and voices of the people involved in the trouble and recognize their emotional state. Specifically, it identifies the emotional state using an AI model that extracts emotions. The input is a list of people involved in the trouble detected in step 3 and their video data, and the output is the emotional state of each person.

[1193] Step 5:

[1194] The server uses the warning message generation means to generate appropriate warning messages based on the emotion data obtained from the emotion recognition means. For example, it generates a message such as "Please stay calm to avoid problems." The input is the emotion data and trouble information obtained in step 4, and the output is the generated warning message.

[1195] Step 6:

[1196] The server sends the generated warning message to the speech means (directional speaker) and notification means (smartphone notification system). The warning message is spoken directly within the commercial facility through the speech means, and simultaneously notifies the target person or employee via their smartphone. The input is the warning message generated in step 5, and the output is the speech within the commercial facility and the smartphone notification.

[1197] These processing steps enable quick and effective detection of troubles and violations of etiquette within commercial facilities, and appropriate measures to be taken.

[1198] The specific processing unit 290 transmits the result of the specific processing to the headset type terminal 314. In the headset type terminal 314, the control unit 46A causes the speaker 240 and the display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[1199] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[1200] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the headset type terminal 314.

[1201] [Fourth embodiment]

[1202] FIG. 7 shows an example of the configuration of a data processing system 410 according to the fourth embodiment.

[1203] 7, a data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.

[1204] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[1205] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a control target 443. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the control target 443 are also connected to the bus 52.

[1206] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[1207] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[1208] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[1209] The control object 443 includes a display device, LEDs in the eyes, and motors for driving the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the emotions of the robot 414 can be expressed by controlling these motors. In addition, the facial expressions of the robot 414 can also be expressed by controlling the light emission state of the LEDs in the eyes of the robot 414.

[1210] Fig. 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Fig. 8, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[1211] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[1212] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[1213] In the robot 414, the processor 46 performs the reception output process. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[1214] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1215] The present invention is a system that uses AI technology to monitor troubles and violations of etiquette on public transportation and automatically urges appropriate caution, and can be implemented in the following forms.

[1216] System configuration

[1217] The system includes the following elements:

[1218] 1. Terminal (surveillance camera)

[1219] 2. Server

[1220] 3. Terminal (directional speaker)

[1221] Program processing and natural language explanation

[1222] Video monitoring and transmission

[1223] The devices (surveillance cameras) capture images in real time inside vehicles and public transportation. These images are sent to a server via a network. The surveillance cameras are high-resolution and are installed to cover a wide area.

[1224] Video analysis

[1225] The server analyzes the received video data and uses an AI model (such as a convolutional neural network) to extract features from the video, such as suspicious movements, crowds of people, and unusual sounds. This allows for the rapid and accurate detection of unusual behavior and abnormal situations in the video.

[1226] Trouble detection

[1227] The server then runs a problem detection algorithm based on the extracted features. For example, if abnormally close movements of people or aggressive behavior are detected, it will be recognized as a specific problem (e.g., sexual harassment, assault, etc.) and a warning flag will be raised.

[1228] Generate warning messages

[1229] The server then generates a warning message based on the type of trouble, location, and behavior of the target. For example, if harassment is detected on a crowded train, the server generates a warning message saying, "Please stop harassing."

[1230] Caution speech

[1231] The generated warning message is sent from the server to a terminal (directional speaker), which then speaks the message and delivers the warning directly to the target person. Because the directional speaker can focus sound in a specific direction, it can deliver the message only to the target person without causing unnecessary inconvenience to surrounding passengers.

[1232] Specific examples

[1233] As a concrete example, we will explain how to detect and respond to sexual harassment on a crowded train.

[1234] 1. Video monitoring and transmission: The surveillance camera inside the vehicle captures the video and transmits it to the server. The video is high resolution and covers a wide area.

[1235] 2. Video analysis: Using an AI model, the server extracts features from the video received, such as suspicious movements and crowds of people, and analyzes people's movements and relative positions in detail.

[1236] 3. Trouble detection: AI analyzes these features to detect abnormal behavior, such as sexual harassment, which then raises a warning flag.

[1237] 4. Generation of warning message: When a problem is detected, the server generates a warning message such as "Please stop harassing the user."

[1238] 5. Caution speech: The server sends the generated caution text to the terminal (directional speaker), and the speaker speaks the caution to the target person.

[1239] In this way, this system effectively manages troubles and violations of etiquette on public transport, providing a safe and comfortable riding environment.

[1240] The processing flow will be explained below.

[1241] Step 1:

[1242] The device (surveillance camera) captures real-time images from inside a vehicle or public transport vehicle. The surveillance camera is set to capture high-resolution images and cover a wide area.

[1243] Step 2:

[1244] The video data captured by the terminal (surveillance camera) is compressed and sent over the network to the server, where it is converted into an appropriate format for efficient data transfer.

[1245] Step 3:

[1246] The server decodes the received video data and analyzes it in real time using dedicated video analysis software.

[1247] Step 4:

[1248] The server uses an AI model to extract features from the video data. These features include suspicious movements, crowds of people, and unusual sounds. Convolutional neural networks (CNNs) are used for extraction.

[1249] Step 5:

[1250] The server evaluates the extracted features to detect abnormal behavior and signs of trouble, where machine learning algorithms identify abnormal patterns based on past data.

[1251] Step 6:

[1252] The server sets a warning flag for any abnormal behavior it detects. For example, if assault or sexual harassment is detected, a warning flag for that behavior will be set.

[1253] Step 7:

[1254] The server collects details of the incident (type of behavior, location, characteristics of the victim, etc.) and uses conversational AI to generate a warning message, such as "Mobbing behavior has been detected. Please stop immediately."

[1255] Step 8:

[1256] The server converts the generated warning message into an appropriate data format and transmits it to the terminal (directional speaker) via the network.

[1257] Step 9:

[1258] The terminal (directional speaker) analyzes the received warning message and speaks it out to the target person. The directional speaker focuses the sound in a specific direction, warning the person without disturbing surrounding passengers.

[1259] Step 10:

[1260] The server continues to analyze the surveillance camera footage and monitors the subject's reaction to the attention, checking for changes in the subject's movements and behavior.

[1261] Step 11:

[1262] If the user does not follow the warning, the server generates an additional warning message and sends it to the device (directional speaker) again. If necessary, the warning message is changed to a stronger one.

[1263] Step 12:

[1264] If the server cannot resolve the problem, it automatically sends an emergency notification to the police and the driver, which includes detailed information about the problem, enabling a prompt response.

[1265] These are the specific processing steps of the program for this system. In this way, by linking each step together, troubles within public transportation systems can be effectively prevented and passenger safety can be ensured.

[1266] Example 1

[1267] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1268] There is a need to monitor troubles and violations of etiquette on public transportation and respond quickly and appropriately. However, current systems have difficulty detecting troubles in real time and automatically issuing appropriate warnings. Rapid detection and response to troubles is particularly important on packed transportation, and an effective system for this purpose is required.

[1269] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[1270] In this invention, the server includes a monitoring means, a video analysis means, a trouble detection means, a warning message generation means, and a speech means, which makes it possible to detect troubles and manner violations on public transportation in real time and to promptly and appropriately warn people.

[1271] "Monitoring means" refers to devices for capturing images inside public transportation, and mainly includes cameras.

[1272] The "video analysis means" is a device that has the function of analyzing received video data and extracting features such as suspicious movements, crowds of people, and abnormal sounds.

[1273] The "trouble detection means" refers to an algorithm or device for detecting trouble based on the feature amount extracted by the video analysis means.

[1274] The "warning message generation means" is a device that has the function of using a generative AI model to generate appropriate warning messages in response to detected problems.

[1275] The "utterance means" is a device for uttering the generated warning message to a specific target person, and particularly includes a directional speaker.

[1276] This invention is a system that uses AI technology to monitor troubles and bad manners on public transportation and automatically warns passengers of appropriate precautions. This system is implemented using the following hardware and software.

[1277] System configuration

[1278] The system includes the following elements:

[1279] 1. Terminal (surveillance camera)

[1280] 2. Server

[1281] 3. Terminal (directional speaker)

[1282] Video acquisition and transmission

[1283] Terminals (surveillance cameras) are installed inside public transportation facilities and capture high-resolution video in real time. The surveillance cameras are installed to cover a wide area, making them particularly useful for detecting trouble on crowded trains. The captured video data is sent to a server via a network. The transmitted data is encrypted to ensure security.

[1284] Video analysis

[1285] After the video data is sent to the server, the server receives the data. The server is equipped with a convolutional neural network (CNN) built using TensorFlow and PyTorch. The server uses this AI model to analyze the received video data and extract features such as suspicious movements, crowds of people, and unusual sounds. The analysis is performed in real time, dividing each video frame and extracting features using a specific algorithm.

[1286] Trouble detection

[1287] The server runs a problem detection algorithm based on the features obtained from the video analysis. For example, if abnormally close movements of people or aggressive behavior are detected, it will recognize this as a specific problem (e.g., sexual harassment or assault) and raise a warning flag. For example, if the server recognizes that a specific passenger is making inappropriate contact with another passenger on a crowded train, a warning flag will be raised.

[1288] Generate warning messages

[1289] When a warning flag is raised on the server, the server generates a warning message based on the problem. The message is generated using a generative AI model (e.g., GPT-3) based on a specific prompt. The warning message is generated using the following prompt:

[1290] "Trouble type: sexual harassment

[1291] Location: A crowded train

[1292] Situation: Passenger A makes inappropriate contact with another passenger B."

[1293] Based on this prompt, an appropriate warning message such as "Please stop harassing others" is generated.

[1294] Caution speech

[1295] The warning message generated by the server is sent to a terminal (directional speaker). The terminal (directional speaker) can focus sound in a specific direction, so the message is delivered to the target person only, without causing unnecessary inconvenience to surrounding passengers. For example, the speaker can issue a voice message saying, "Please stop molesting me" to a passenger sitting in a specific seat on a crowded train.

[1296] This system quickly and accurately detects any trouble or violations of etiquette on public transport in real time and provides appropriate warnings, resulting in a safe and comfortable riding environment.

[1297] The embodiment of the invention is configured as described above, which makes it possible to effectively monitor and quickly respond to problems within public transportation.

[1298] The flow of the identification process in the first embodiment will be described with reference to FIG.

[1299] Program processing flow

[1300] Step 1: Acquire and transmit video

[1301] Users capture video using terminals (surveillance cameras) installed within public transportation. The cameras capture high-resolution video in real time and transmit the video data to a server via a network. The video captured by the camera is continuously divided into frames and transmitted to the server. During this process, the video data is encrypted to ensure security.

[1302] Input: Real-time camera footage

[1303] Output: Video data sent to the server

[1304] Step 2: Receiving the video and preparing for analysis

[1305] The server receives video data sent from the terminal (surveillance camera) via the network. The received video data is divided into frames and saved.

[1306] Input: Video data transmitted over the network

[1307] Output: Frame-by-frame video data stored on the server

[1308] Step 3: Analyzing the footage

[1309] The server analyzes the stored video data using a convolutional neural network (CNN). Specifically, it uses AI frameworks such as TensorFlow and PyTorch to perform real-time data analysis. As a result of the analysis, features such as suspicious movements, crowds of people, and unusual sounds are extracted. These features are used as input for predicting trouble.

[1310] Input: Frame-by-frame video data stored on the server

[1311] Output: Features such as suspicious movements, crowds of people, and abnormal sounds

[1312] Step 4: Detecting the problem

[1313] The server runs a problem detection algorithm based on the features obtained through the analysis. For example, if abnormal proximity or aggressive behavior is detected, it recognizes this as a specific problem (e.g., sexual harassment or assault) and raises a warning flag. This allows specific abnormal behavior to be identified in real time.

[1314] Input: Features extracted by analysis

[1315] Output: Trouble detection flag

[1316] Step 5: Generate warning text

[1317] The server generates a warning message when a warning flag is raised. The server uses a generative AI model (e.g., GPT-3) to generate an appropriate warning message based on a specific prompt. For example, given the prompt "Trouble type: sexual harassment, location: crowded train, situation: passenger A making inappropriate contact with another passenger B," the server generates a warning message saying, "Please stop sexual harassment."

[1318] Input: Trouble detection flag, prompt text

[1319] Output: Generated warning message

[1320] Step 6: Send and speak the warning message

[1321] The server generates a warning message and sends it to the device (directional speaker). The directional speaker can focus sound in a specific direction, delivering the message only to the target person without disturbing surrounding passengers. The speaker then speaks a warning message to the target person, saying, "Please stop harassing."

[1322] Input: Generated warning text

[1323] Output: Warning message spoken from the device (directional speaker)

[1324] In this way, this system can monitor troubles and violations of etiquette on public transportation in real time and automatically provide appropriate warnings.

[1325] (Application example 1)

[1326] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1327] Systems for effectively monitoring and quickly responding to incidents and violations of etiquette on public transportation are limited, making it difficult for security staff to quickly grasp the situation and take appropriate action. Effective methods for focusing sound in a specific direction to draw attention are also lacking. Furthermore, existing systems struggle to provide real-time notifications and visual assistance, hindering security improvements.

[1328] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[1329] In this invention, the server includes a monitoring means, an image analysis means, a trouble detection means, a warning message generation means, an output means, a notification means, and a support display means, which makes it possible to monitor and analyze troubles and manner violations on public transportation in real time, quickly detect troubles, generate and output appropriate warning messages, and immediately notify security staff and provide visual support information.

[1330] "Surveillance" refers to the equipment and technology used to capture footage inside public transport.

[1331] "Video analysis means" refers to devices and techniques for extracting and analyzing features from acquired video data.

[1332] "Trouble detection means" refers to devices and technologies for detecting trouble by discovering suspicious movements or abnormalities from features extracted by video analysis means.

[1333] "Warning message generation means" refers to a device and technology for generating appropriate warning messages in response to a detected problem.

[1334] "Speech means" refers to a device and technology for outputting the generated warning message as voice.

[1335] "Notification means" refers to devices and technologies used to notify security staff and other relevant parties of detected problems and generated warning messages.

[1336] "Support display means" refers to devices and technologies that visually provide necessary information to security staff when they respond.

[1337] The present invention is a system that uses AI technology to monitor troubles and bad manners on public transportation and automatically urges appropriate cautions. An embodiment of the present invention will be described.

[1338] System configuration

[1339] The system includes the following elements:

[1340] 1. Surveillance methods (e.g., built-in cameras in smart glasses)

[1341] 2. Server

[1342] 3. Speaking means (e.g., directional speaker)

[1343] 4. Means of notification

[1344] 5. Supportive display means (e.g., smart glasses display)

[1345] Program processing

[1346] Video Acquisition

[1347] The monitoring tool will capture real-time footage of public transport vehicles, for example, through cameras built into smart glasses, allowing security staff to keep up to date with the latest developments.

[1348] Video transmission

[1349] The captured images are then sent to a server via Wi-Fi or mobile data networks, where high-resolution images are transmitted for improved analysis accuracy.

[1350] Video analysis

[1351] The server analyzes the received video data using an AI model (e.g., convolutional neural network), which extracts features such as suspicious movements, crowds of people, and abnormal sounds.

[1352] Trouble detection

[1353] The server then runs a problem detection algorithm based on the extracted features. For example, if unusually close movements of people or aggressive behavior are detected, this is recognized as a specific problem and a warning flag is raised.

[1354] Generate warning messages

[1355] The server generates a warning message based on the type of incident, location, and the target's behavior. For example, if a sexual assault is detected on a crowded train, a warning message saying "Please stop sexual assault" will be automatically generated.

[1356] Caution speech

[1357] The generated warning message is sent from the server to a speech device (directional speaker), which then speaks the message and delivers the warning directly to the target person. This allows the message to be delivered only to specific targets without disturbing surrounding passengers.

[1358] Notifications and Support Displays

[1359] At the same time, the server notifies security staff of any detected problems or warning messages via notification means, and visual support information is displayed on the smart glasses' display to help them respond quickly and appropriately.

[1360] Specific examples

[1361] As a concrete example, we will explain how to detect and respond to sexual harassment on a crowded train.

[1362] 1. Image acquisition: The built-in camera in the smart glasses captures images and sends them to the server. High-resolution images cover a wide area.

[1363] 2. Video analysis: Using an AI model, the server extracts features from the video received, such as suspicious movements and crowds of people, and performs a detailed analysis of people's movements and relative positions.

[1364] 3. Trouble detection: Abnormal behavior, such as sexual harassment, is detected and a warning flag is raised.

[1365] 4. Generation of warning message: The server generates a warning message saying, "Please stop harassing."

[1366] 5. Caution message generation and notification: The server sends the generated caution message to the directional speaker and issues a message, while simultaneously providing visual support information to security staff.

[1367] Prompt Sentence Examples

[1368] "Please detect abnormal movements and crowding of people on crowded trains and notify us so that appropriate measures can be taken."

[1369] This allows any trouble or violations of etiquette on public transport to be detected immediately, allowing security staff to respond quickly and appropriately.

[1370] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[1371] Step 1:

[1372] Users wear smart glasses while in public transport, and the built-in camera captures real-time images of passengers. The input is real-time images from inside the public transport vehicle, and the output is high-resolution video data.

[1373] Step 2:

[1374] The device (smart glasses) transmits the captured video to a server via Wi-Fi or a mobile data network. The input is high-resolution video data, and the output is the video data transmitted to the server.

[1375] Step 3:

[1376] The server analyzes the received video data using an AI model (for example, a convolutional neural network). The input is the transmitted video data, and the output is the extracted features (suspicious movements, crowds of people, abnormal sounds, etc.). Specific operations include inputting the video data into the AI ​​model and extracting the features.

[1377] Step 4:

[1378] The server runs a trouble detection algorithm based on the extracted features. The input is the extracted features, and the output is the trouble detection result (e.g., detection of sexual harassment). Specifically, it compares the extracted features with pre-set criteria to detect anomalies.

[1379] Step 5:

[1380] The server generates an appropriate warning message based on the results of the trouble detection. The input is the trouble detection result, and the output is the generated warning message (e.g., "Please stop harassing others"). Specifically, the server inputs the type of trouble and location information into a predefined warning message template to generate the appropriate warning message.

[1381] Step 6:

[1382] The server transmits the generated warning message to the speech generator (directional speaker). The input is the generated warning message, and the output is the data to be transmitted to the directional speaker. Specifically, the server transmits an appropriate message to the directional speaker.

[1383] Step 7:

[1384] The speech generator (directional speaker) speaks the transmitted warning message as a voice in a specific direction. The input is the transmitted data, and the output is the actual voice warning message. Specifically, it uses a voice synthesis function to speak a specific message to the target.

[1385] Step 8:

[1386] At the same time, the server uses a notification means to notify security staff of the detected trouble and the generated warning message. The input is the trouble detection result and the warning message, and the output is visual support information displayed on the smart glasses' display. Specifically, the server sends the notification data to the smart glasses, and appropriate support information is displayed on the display.

[1387] Prompt Sentence Examples

[1388] "Detect abnormal movements and crowding of people inside a moving train and provide appropriate countermeasures."

[1389] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[1390] The present invention is a system that uses AI technology to monitor troubles and violations of etiquette on public transportation and automatically urges appropriate cautions, and by combining it with an emotion engine that recognizes the user's emotions, it is intended to enhance the effectiveness of preventing troubles. It can be implemented in the following forms.

[1391] System configuration

[1392] The system includes the following elements:

[1393] 1. Terminal (surveillance camera)

[1394] 2. Server

[1395] 3. Terminal (directional speaker)

[1396] 4. Emotion Engine

[1397] Program processing and natural language explanation

[1398] Video monitoring and transmission

[1399] The devices (surveillance cameras) capture images in real time inside vehicles and public transportation. These images are sent to a server via a network. The surveillance cameras are high-resolution and are installed to cover a wide area.

[1400] Video analysis

[1401] The server analyzes the received video data and uses an AI model (such as a convolutional neural network) to extract features from the video, such as suspicious movements, crowds of people, and unusual sounds. This allows for the rapid and accurate detection of unusual behavior and abnormal situations in the video.

[1402] Trouble detection

[1403] The server then runs a problem detection algorithm based on the extracted features. For example, if abnormally close movements of people or aggressive behavior are detected, it will be recognized as a specific problem (e.g., sexual harassment, assault, etc.) and a warning flag will be raised.

[1404] Emotion recognition

[1405] When a problem is detected, the emotion engine analyzes the target's facial expressions and voice tone to recognize their emotional state (e.g., anger, anxiety, fear, etc.). The emotion engine uses an AI model to extract emotions from video and audio data.

[1406] Generate warning messages

[1407] Next, the server generates a warning message based on the emotional data of the target person recognized by the emotion engine. The conversational AI automatically creates an appropriate warning message based on the type of trouble, location, and the target person's emotional state. For example, if the target person is expressing anger, a message such as "Avoid trouble. Remain calm" will be generated.

[1408] Caution speech

[1409] The generated warning message is sent from the server to the device (directional speaker), which then directly warns the target person. The speaker also speaks in a tone and language that corresponds to the emotional state recognized by the emotion engine. This allows for more effective warning communication to the target person.

[1410] Specific examples

[1411] As a concrete example, we will explain how to detect and respond to problems on a crowded train.

[1412] 1. Video monitoring and transmission: The surveillance camera inside the vehicle captures the video and transmits it to the server. The video is high resolution and covers a wide area.

[1413] 2. Video analysis: Using an AI model, the server extracts features from the video received, such as suspicious movements and crowds of people, and analyzes people's movements and relative positions in detail.

[1414] 3. Trouble detection: AI analyzes these features to detect abnormal behavior, such as sexual harassment, which then raises a warning flag.

[1415] 4. Emotion Recognition: The emotion engine analyzes the facial expressions and voice of the target when a problem is detected and recognizes their emotional state. For example, if the target shows feelings of anxiety or anger.

[1416] 5. Generating warning messages: The server takes into account the recognized emotional data and generates warning messages such as "Please stay calm to avoid trouble" rather than "Please stop harassing others."

[1417] 6. Speech of warning: The server sends the generated warning message to the device (directional speaker), and the speaker speaks the warning to the target person. At the same time, the message is conveyed in a tone and language that matches the user's emotions.

[1418] These are the specific processing steps of the program of this system. In this way, by linking each step together, it is possible to effectively prevent trouble on public transportation and ensure the safety of passengers.

[1419] The processing flow will be explained below.

[1420] Step 1:

[1421] The device (surveillance camera) captures real-time images from inside a vehicle or public transport vehicle. The surveillance camera is set to capture high-resolution images and cover a wide area.

[1422] Step 2:

[1423] The video data captured by the terminal (surveillance camera) is compressed and sent over the network to the server, where it is converted into an appropriate format for efficient data transfer.

[1424] Step 3:

[1425] The server decodes the received video data and analyzes it in real time using dedicated video analysis software.

[1426] Step 4:

[1427] The server uses an AI model to extract features from the video data. These features include suspicious movements, crowds of people, and unusual sounds. Convolutional neural networks (CNNs) are used for extraction.

[1428] Step 5:

[1429] The server evaluates the extracted features to detect abnormal behavior and signs of trouble, where machine learning algorithms identify abnormal patterns based on past data.

[1430] Step 6:

[1431] The server sets a warning flag for any abnormal behavior it detects. For example, if assault or sexual harassment is detected, a warning flag for that behavior will be set.

[1432] Step 7:

[1433] When a problem is detected, the emotion engine analyzes the target's facial expressions and voice tone to recognize their emotional state (e.g., anger, anxiety, fear, etc.). The emotion engine uses an AI model to extract emotions from video and audio data.

[1434] Step 8:

[1435] The server organizes the details of the trouble (type of behavior, location, characteristics of the target, etc.) and the emotional state recognized by the emotion engine, and generates a warning message using conversational AI. For example, if the target shows anger, a message such as "Avoid trouble. Remain calm" will be generated.

[1436] Step 9:

[1437] The server converts the generated warning message into an appropriate data format and transmits it to the terminal (directional speaker) via the network.

[1438] Step 10:

[1439] The device (directional speaker) analyzes the received warning message and speaks it out to the target passenger. The directional speaker focuses the sound in a specific direction, warning the passenger without disturbing other passengers. The emotion engine also recognizes the passenger's emotional state and uses a tone and wording that matches that state, providing more effective warnings.

[1440] Step 11:

[1441] The server continues to analyze the surveillance camera footage and monitors the subject's reaction to the attention, checking for changes in the subject's movements and behavior.

[1442] Step 12:

[1443] If the user does not follow the warning, the server generates an additional warning message and sends it to the device (directional speaker) again. If necessary, the warning message is changed to a stronger one.

[1444] Step 13:

[1445] If the server cannot resolve the problem, it automatically sends an emergency notification to the police and the driver, which includes detailed information about the problem, enabling a prompt response.

[1446] Example 2

[1447] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1448] Trouble and bad manners on public transportation are serious issues that threaten passenger safety and comfort. Current monitoring systems have limitations in detecting suspicious movements and abnormal behavior, and are unable to respond appropriately while taking into account the emotional state of passengers. This makes it difficult to take prompt and effective measures when trouble occurs. There is also a lack of technology that can recognize passengers' emotions and psychological state and respond accordingly.

[1449] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[1450] In this invention, the server includes a monitoring means, a video analysis means, a trouble detection means, an emotion recognition means, a warning message generation means, and a speech utterance means. This makes it possible to monitor video in public transportation in real time, detect suspicious movements and crowding of people, and recognize the emotional state of passengers, and generate and speak appropriate warning messages based on those emotions.

[1451] "Monitoring means" refers to devices for capturing images inside public transportation, and primarily involves the use of surveillance cameras.

[1452] "Video analysis means" refers to technology or equipment that analyzes acquired video data and extracts features such as suspicious movements, crowds of people, and abnormal sounds.

[1453] "Trouble detection means" refers to algorithms and technologies for detecting abnormal behavior or trouble based on extracted features.

[1454] "Emotion recognition means" refers to technology or equipment that analyzes the facial expressions and tone of voice of a person when a problem is detected, and recognizes the emotional state of the person.

[1455] The "warning message generation means" refers to a technique or device for generating appropriate warning messages based on the recognized emotional state.

[1456] The "utterance means" is a device for uttering the generated warning message in a specific direction to convey it to the target person.

[1457] This invention is a system that uses AI technology to monitor troubles and bad manners on public transportation and automatically provides appropriate warnings. The system includes the following elements:

[1458] monitoring means

[1459] The terminal (surveillance camera) captures images of the inside of public transportation in real time. The surveillance camera has high resolution and is installed to cover a wide area. The captured images are sent to a server via a network.

[1460] Video analysis methods

[1461] The server analyzes the received video data, using AI models such as convolutional neural networks to extract features from the video, such as suspicious movements, crowds of people, and unusual sounds.

[1462] Trouble detection methods

[1463] The server runs a problem detection algorithm based on the extracted features. If abnormally close movements of people or aggressive behavior are detected, it will recognize it as a specific problem (e.g., assault) and raise a warning flag.

[1464] emotion recognition means

[1465] When the emotion engine detects a problem, it analyzes the target's facial expressions and vocal tone, using an AI model that extracts emotions from video and audio data to recognize the target's emotional state (e.g., anger, anxiety, fear, etc.).

[1466] Caution statement generation means

[1467] The server generates warning messages based on the emotional data of the target person recognized by the emotion engine. The conversational AI automatically creates appropriate warning messages based on the type of trouble, location, and the target person's emotional state. For example, if the target person is expressing anger, a message such as "Avoid trouble. Remain calm" will be generated.

[1468] Means of speech

[1469] The generated warning message is sent from the server to the device (directional speaker). The directional speaker then directly warns the target person. The speaker also speaks in a tone and language that corresponds to the emotional state recognized by the emotion engine. This operation allows for more effective warning communication to the target person.

[1470] Specific examples

[1471] As a concrete example, we will explain how to detect and respond to problems on a crowded train.

[1472] 1. Video monitoring and transmission: The terminal (surveillance camera) captures images from inside the train and transmits them to the server in real time. The camera has high resolution and covers a wide area.

[1473] 2. Video analysis: The server uses an AI model to analyze the video data it receives and extracts suspicious movements and crowds of people as features.

[1474] 3. Trouble detection: The server uses AI to analyze features and detect abnormal approaches or unnatural movements. This allows it to detect troubles such as sexual harassment and raise a warning flag.

[1475] 4. Emotion recognition: The emotion engine analyzes the subject's facial expressions and voice to recognize emotions such as anxiety or anger.

[1476] 5. Generating warning messages: Based on the emotional data, the server generates warning messages such as "Please stay calm to avoid trouble" rather than "Please stop harassing others."

[1477] 6. Speech warning: The server generates a warning message and sends it to the device (directional speaker), which then speaks it to the target person. By conveying the warning in a tone and using words that match the target person's emotions, a more effective response can be achieved.

[1478] Examples of prompt statements

[1479] Below is an example of a prompt sentence to input to the generative AI model.

[1480] "Please explain the system for detecting suspicious behavior on public transportation and taking appropriate action. For example, please explain in detail how to respond to sexual harassment on a crowded train."

[1481] The system serves as an effective means of improving safety within public transport.

[1482] The flow of the identification process in the second embodiment will be described with reference to FIG.

[1483] Step 1:

[1484] Video monitoring and transmission

[1485] explanation

[1486] The device (surveillance camera) captures images of the inside of public transportation in real time, and the captured image data is sent to a server via a network.

[1487] Specific actions

[1488] Input: Public transport footage

[1489] Data processing: The surveillance camera sensor captures the video and obtains detailed video data in high resolution.

[1490] Output: Video data sent to the server in real time

[1491] Step 2:

[1492] Video analysis

[1493] explanation

[1494] The server analyzes the received video data and uses an AI model (such as a convolutional neural network) to extract features such as suspicious movements, crowds of people, and unusual sounds from the video.

[1495] Specific actions

[1496] Input: Received video data

[1497] Data processing: Decoding video data, analyzing each frame, and using AI models to identify suspicious movements, crowds of people, and unusual sounds

[1498] Output: Extracted feature data

[1499] Step 3:

[1500] Trouble detection

[1501] explanation

[1502] The server runs a trouble-detection algorithm based on the extracted features. For example, if abnormally close movements of people or aggressive behavior are detected, it will recognize the problem and raise a warning flag.

[1503] Specific actions

[1504] Input: Extracted feature data

[1505] Data processing: Input feature data into the trouble detection algorithm to detect abnormal patterns

[1506] Output: Trouble detection flag

[1507] Step 4:

[1508] Emotion recognition

[1509] explanation

[1510] The emotion engine recognizes the emotional state of the person who detected the problem by analyzing facial expressions and tone of voice (e.g., anger, anxiety, fear, etc.).

[1511] Specific actions

[1512] Input: Subject's facial expression data and voice data

[1513] Data processing: The emotion engine uses AI models to analyze facial expressions and voice to identify emotional states.

[1514] Output: Recognized emotion data

[1515] Step 5:

[1516] Generate warning messages

[1517] explanation

[1518] The server generates warning messages based on the target person's emotional data recognized by the emotion engine. The conversational AI creates warning messages based on the type of trouble, location, and the target person's emotional state.

[1519] Specific actions

[1520] Input: Recognized emotion data, type of trouble, location information

[1521] Data processing: Input prompts into the generative AI model to generate appropriate warning messages

[1522] Output: Generated warning message

[1523] Step 6:

[1524] Caution speech

[1525] explanation

[1526] The generated warning message is sent from the server to the device (directional speaker), which then speaks the message to the target person.

[1527] Specific actions

[1528] Input: Generated warning text data

[1529] Data processing: A directional speaker uses a speech synthesis engine to play back warning messages and speak them in a specific direction.

[1530] Output: Audio warning to the target

[1531] In this way, each step works together to effectively prevent trouble on public transport and ensure passenger safety.

[1532] (Application example 2)

[1533] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1534] In commercial facilities, customers often violate etiquette and cause trouble, which often has a negative impact on store operations and other customers. The purpose of this invention is to maintain order and safety within commercial facilities by quickly and effectively detecting such trouble and violations of etiquette and automatically issuing appropriate warnings.

[1535] The identification process by the identification processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes a monitoring means, a video analysis means, a trouble detection means, an emotion recognition means, a warning message generation means, a speech means, and a notification means. This makes it possible to automatically detect troubles and etiquette violations within a commercial facility and issue appropriate warnings.

[1536] The "monitoring means" is a device that acquires images of the inside of a commercial facility in real time.

[1537] The "video analysis means" is a device or software for processing acquired video data and extracting specific features.

[1538] The "trouble detection means" is a device or algorithm that detects abnormal movements or behaviors based on features extracted from the video analysis means.

[1539] An "emotion recognition means" is a device or algorithm that analyzes the facial expressions and voice of a detected subject and recognizes their emotional state.

[1540] The "warning message generation means" is a device or software that generates appropriate warning messages for the target person based on the data obtained from the emotion recognition means.

[1541] The "utterance means" is a device for uttering the generated warning message as voice.

[1542] The "notification means" is a device or software that transmits the generated warning message to a smartphone or other device.

[1543] This invention is a system that detects troubles and bad manners in commercial facilities and automatically issues appropriate warnings. This system includes seven main means: "monitoring means," "video analysis means," "trouble detection means," "emotion recognition means," "warning message generation means," "speech means," and "notification means."

[1544] System configuration

[1545] The system includes the following elements:

[1546] 1. Monitoring measures

[1547] Camera devices and smartphone cameras installed within commercial facilities are used to capture images of the facility in real time.

[1548] 2. Video analysis methods

[1549] The captured video data is received and specific features are extracted using an AI model (e.g., convolutional neural network).

[1550] The features include abnormal movements, crowding of people, and abnormal behavior.

[1551] 3. Trouble detection methods

[1552] The server detects abnormal behavior based on the features obtained from the video analysis means.

[1553] In this case, a trouble detection algorithm is used, for example, to analyze theft or aggressive behavior and recognize it as trouble.

[1554] 4. Emotion recognition means

[1555] The server analyzes the subject's facial expressions and voice to recognize their emotional state.

[1556] To do this, it uses an AI model that extracts emotions from video and audio data.

[1557] 5. Caution statement generation means

[1558] The server generates appropriate warning messages based on the data obtained from the emotion recognition means.

[1559] For example, if the subject is feeling anxious, the system will automatically generate the phrase "Please remain calm."

[1560] 6. Means of speech

[1561] The warning messages sent from the server are spoken through speakers installed within the commercial facility or via a smartphone app.

[1562] Speech is delivered in a tone and language that corresponds to the emotion.

[1563] 7. Means of notification

[1564] The generated warning message is sent to store employees and the target person via a smartphone notification system.

[1565] For example, if any theft is detected, the store security team is immediately notified.

[1566] Specific examples

[1567] As a concrete example, we will explain the detection and response to theft behavior in a large shopping mall.

[1568] 1. Video monitoring and transmission: Surveillance cameras in the shopping mall capture video and transmit the data to a server. The cameras have high resolution and wide coverage.

[1569] 2. Video analysis: The server uses an AI model to extract suspicious movements and crowding of people from the received video data as features.

[1570] 3. Trouble detection: AI analyzes these features and detects, for example, theft. It raises a warning flag depending on the type of trouble.

[1571] 4. Emotion recognition: The server utilizes emotion recognition means to analyze the target person's facial expressions and voice to recognize their emotional state (e.g., anxiety, anger).

[1572] 5. Generating warning messages: The server generates appropriate warning messages (e.g., "Please stay calm") based on the emotion data.

[1573] 6. Caution notification: The generated caution message is sent directly to the target person via speech, and is also sent to employees and security teams via their smartphone notification system.

[1574] Prompt Sentence Examples

[1575] An example of a prompt to input to a generative AI model is as follows:

[1576] Prompt: If you detect any unusual behavior in the mall, create a message to encourage the person you spot to remain calm.

[1577] Response: "Remain calm and try to resolve the issue. If you need assistance, please speak to a member of staff."

[1578] The above configuration enables quick and effective detection of troubles and violations of etiquette within commercial facilities, and makes it possible to take appropriate measures when they occur.

[1579] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[1580] Step 1:

[1581] The server acquires real-time video footage of the commercial facility from the monitoring means (surveillance cameras). This video footage is high resolution and contains a wide range of data. The input is the video data acquired from the surveillance cameras, and the output is the real-time video data sent to the server.

[1582] Step 2:

[1583] The server analyzes the received video data using a video analysis method. Specifically, it uses a convolutional neural network (CNN) to extract features such as abnormal movements and crowding of people in the video. The input is the video data acquired in step 1, and the output is the extracted feature data.

[1584] Step 3:

[1585] The server uses the trouble detection means to detect troublesome behavior based on the features obtained from the video analysis means. For example, theft or aggressive behavior is analyzed using an algorithm and recognized as trouble. The input is the feature data extracted in step 2, and the output is the type of trouble detected and its location information.

[1586] Step 4:

[1587] The server uses emotion recognition means to analyze the facial expressions and voices of the people involved in the trouble and recognize their emotional state. Specifically, it identifies the emotional state using an AI model that extracts emotions. The input is a list of people involved in the trouble detected in step 3 and their video data, and the output is the emotional state of each person.

[1588] Step 5:

[1589] The server uses the warning message generation means to generate appropriate warning messages based on the emotion data obtained from the emotion recognition means. For example, it generates a message such as "Please stay calm to avoid problems." The input is the emotion data and trouble information obtained in step 4, and the output is the generated warning message.

[1590] Step 6:

[1591] The server sends the generated warning message to the speech means (directional speaker) and notification means (smartphone notification system). The warning message is spoken directly within the commercial facility through the speech means, and simultaneously notifies the target person or employee via their smartphone. The input is the warning message generated in step 5, and the output is the speech within the commercial facility and the smartphone notification.

[1592] These processing steps enable quick and effective detection of troubles and violations of etiquette within commercial facilities, and appropriate measures to be taken.

[1593] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the control target 443 to output the result of the specific processing. The microphone 238 acquires voice indicating a user input regarding the result of the specific processing. The control unit 46A transmits voice data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the voice data.

[1594] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[1595] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the robot 414.

[1596] The emotion identification model 59 as an emotion engine may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to an emotion map (see FIG. 9), which is a specific mapping. Similarly, the emotion identification model 59 may determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.

[1597] FIG. 9 is a diagram illustrating an emotion map 400 on which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. Emotions closer to the center of the concentric circles are more primitive. Emotions representing states and actions arising from a state of mind are arranged on the outer edges of the concentric circles. The concept of emotion includes both affect and mental states. Emotions generally generated from reactions occurring in the brain are arranged on the left side of the concentric circles. Emotions generally induced by situational judgment are arranged on the right side of the concentric circles. Emotions generally generated from reactions occurring in the brain and induced by situational judgment are arranged on the upper and lower sides of the concentric circles. Furthermore, the emotion of "pleasure" is arranged on the upper side of the concentric circles, and the emotion of "discomfort" is arranged on the lower side. In this way, in the emotion map 400, multiple emotions are mapped based on the structure by which emotions are generated, and emotions that tend to occur simultaneously are mapped close to each other.

[1598] These emotions are distributed in the 3 o'clock direction on emotion map 400, and typically fluctuate between relief and anxiety. In the right half of emotion map 400, situational awareness dominates over internal sensations, resulting in a sense of calm.

[1599] The inside of emotion map 400 represents what is going on in the mind, and the outside of emotion map 400 represents behavior, so the further you go outside emotion map 400, the more visible the emotions become (the more they are expressed in behavior).

[1600] Human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. Emotions can also be created for robots, automobiles, and motorcycles, based on various balances, such as posture and remaining battery life. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. An emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on Voice Emotion Recognition and Emotional Brain Physiological Signal Analysis Systems, Tokushima University, Doctoral Dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map lists emotions belonging to the "reaction" domain, where sensation is dominant. The right half of the emotion map lists emotions belonging to the "situation" domain, where situational awareness is dominant.

[1601] The emotion map defines two emotions that promote learning. One is a negative emotion on the situation side, around the middle of "repentance" or "reflection." In other words, this occurs when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is a positive emotion on the response side, around "desire." In other words, this occurs when the robot experiences positive feelings such as "I want more" or "I want to know more."

[1602] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values ​​indicating each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple pieces of training data that are combinations of user input and emotion values ​​indicating each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions that are located close to each other have similar values, as in the emotion map 900 shown in FIG. 10. FIG. 10 shows an example in which multiple emotions, "relieved," "calm," and "reassuring," have similar emotion values.

[1603] The system according to the present disclosure has been described above mainly with respect to the functions of the data processing device 12, but the system according to the present disclosure is not necessarily implemented on a server. The system according to the present disclosure may be implemented as a general information processing system. The present disclosure may be implemented, for example, as a software program running on a personal computer or an application running on a smartphone, etc. The method according to the present disclosure may be provided to users in the form of SaaS (Software as a Service).

[1604] In the above embodiment, an example was given in which the specific processing is performed by one computer 22, but the technology of the present disclosure is not limited to this, and the specific processing may be distributed and performed by a plurality of computers including the computer 22. For example, the data generation model 58 may be provided in an external device of the data processing device 12, and data may be generated in the external device in accordance with input data.

[1605] In the above embodiment, an example in which the specific processing program 56 is stored in the storage 32 has been described, but the technology of the present disclosure is not limited to this. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-transitory storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-transitory storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes the specific processing in accordance with the specific processing program 56.

[1606] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.

[1607] It is not necessary to store all of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store all of the specific processing program 56 in the storage 32; only a portion of the specific processing program 56 may be stored.

[1608] The hardware resource for executing a specific process can be any of the following processors: An example of a processor is a CPU, which is a general-purpose processor that functions as a hardware resource for executing a specific process by executing software, i.e., a program. Another example of a processor is a dedicated electrical circuit, such as an FPGA (Field-Programmable Gate Array), a PLD (Programmable Logic Device), or an ASIC (Application Specific Integrated Circuit), which is a processor with a circuit configuration designed specifically for executing a specific process. Each processor has built-in or connected memory, and each processor uses the memory to execute the specific process.

[1609] The hardware resource that executes the specific processing may be configured with one of these various processors, or may be configured with a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Also, the hardware resource that executes the specific processing may be a single processor.

[1610] As an example of a system configured with a single processor, first, one processor is configured by combining one or more CPUs and software, and this processor functions as a hardware resource that executes a specific process. Second, there is a system that uses a processor that realizes the functions of an entire system including multiple hardware resources that execute a specific process on a single IC chip, as typified by SoC (System-on-a-chip). In this way, a specific process is realized using one or more of the above-mentioned various processors as hardware resources.

[1611] Furthermore, the hardware structure of these various processors can be, more specifically, an electric circuit that combines circuit elements such as semiconductor devices. The specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps may be deleted, new steps may be added, or the processing order may be rearranged, without departing from the spirit of the invention.

[1612] The above-described description and illustrations are a detailed explanation of the parts related to the technology of the present disclosure and are merely an example of the technology of the present disclosure. For example, the above description of the configuration, functions, actions, and effects is an explanation of an example of the configuration, functions, actions, and effects of the parts related to the technology of the present disclosure. Therefore, it goes without saying that unnecessary parts may be deleted, new elements may be added, or replacements may be made to the above-described description and illustrations within the scope of the gist of the technology of the present disclosure. Furthermore, to avoid confusion and facilitate understanding of the parts related to the technology of the present disclosure, the above-described description and illustrations omit explanations of common technical knowledge that do not require particular explanation to enable the implementation of the technology of the present disclosure.

[1613] All publications, patent applications, and technical standards mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent application, or technical standard was specifically and individually indicated to be incorporated by reference.

[1614] The following is further disclosed regarding the above embodiment.

[1615] (Claim 1)

[1616] monitoring means;

[1617] Video analysis means;

[1618] A trouble detection means;

[1619] a warning message generating means;

[1620] A speech means;

[1621] A system including:

[1622] (Claim 2)

[1623] 2. The system according to claim 1, wherein the monitoring means acquires images of the inside of the public transportation facility using a surveillance camera.

[1624] (Claim 3)

[1625] 2. The system according to claim 1, wherein the trouble detection means extracts suspicious movements, crowds of people, and abnormal sounds as features from the received video data.

[1626] "Example 1"

[1627] (Claim 1)

[1628] monitoring means;

[1629] Video analysis means;

[1630] A trouble detection means;

[1631] a warning message generating means;

[1632] A speech means;

[1633] A system including:

[1634] (Claim 2)

[1635] 2. The system according to claim 1, wherein the monitoring means uses a camera to capture video images inside the public transportation facility.

[1636] (Claim 3)

[1637] The system according to claim 1, wherein the video analysis means analyzes the received video data using a convolutional neural network and extracts suspicious movements, crowds of people, and abnormal sounds as features.

[1638] (Claim 4)

[1639] The system according to claim 1, wherein the warning message generation means uses a generative AI model based on the extracted features to generate appropriate warning messages.

[1640] (Claim 5)

[1641] 2. The system according to claim 1, wherein the speech generating means uses a directional speaker to generate and generate a warning message in a specific direction.

[1642] "Application Example 1"

[1643] (Claim 1)

[1644] monitoring means;

[1645] Video analysis means;

[1646] A trouble detection means;

[1647] a warning message generating means;

[1648] A speech means;

[1649] A notification means;

[1650] Support display means;

[1651] A system including:

[1652] (Claim 2)

[1653] 2. The system according to claim 1, wherein the monitoring means acquires images of the inside of the public transportation facility using a surveillance camera.

[1654] (Claim 3)

[1655] 2. The system according to claim 1, wherein the trouble detection means extracts suspicious movements, crowds of people, and abnormal sounds as features from the received video data.

[1656] "Example 2: Combining Emotion Engines"

[1657] (Claim 1)

[1658] monitoring means;

[1659] Video analysis means;

[1660] A trouble detection means;

[1661] An emotion recognition means;

[1662] a warning message generating means;

[1663] A speech means;

[1664] A system including:

[1665] (Claim 2)

[1666] 2. The system according to claim 1, wherein the monitoring means acquires images of the inside of the public transportation facility using a surveillance camera.

[1667] (Claim 3)

[1668] 2. The system according to claim 1, wherein the trouble detection means extracts suspicious movements, crowds of people, and abnormal sounds as features from the received video data.

[1669] (Claim 4)

[1670] 2. The system according to claim 1, wherein the emotion recognition means analyzes the facial expression and tone of voice of the subject when a problem is detected, and recognizes the subject's emotional state.

[1671] (Claim 5)

[1672] 2. The system according to claim 1, wherein the warning message generating means generates appropriate warning messages based on the recognized emotional state.

[1673] (Claim 6)

[1674] 2. The system according to claim 1, wherein the speech means speaks the generated warning message in a specific direction.

[1675] "Application example 2 when combining emotion engines"

[1676] (Claim 1)

[1677] monitoring means;

[1678] Video analysis means;

[1679] A trouble detection means;

[1680] An emotion recognition means;

[1681] a warning message generating means;

[1682] A speech means;

[1683] A notification means;

[1684] A system including:

[1685] (Claim 2)

[1686] 2. The system according to claim 1, wherein the monitoring means acquires video images of the inside of the commercial facility using a video acquisition device.

[1687] (Claim 3)

[1688] 2. The system according to claim 1, wherein the trouble detection means extracts suspicious movements, crowds of people, and abnormal behavior as features from the received video data. [Explanation of symbols]

[1689] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Device 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robot< / url:> < / url:> < / url:> < / url:>

Claims

1. monitoring means; Video analysis means; A trouble detection means; a warning message generating means; A speech means; A system including:

2. 2. The system according to claim 1, wherein the monitoring means acquires images of the inside of the public transportation facility using a monitoring camera.

3. 2. The system according to claim 1, wherein the trouble detection means extracts suspicious movements, crowds of people, and abnormal sounds as features from the received video data.

Citation Information

Patent Citations

  • Persona chatbot control method and system

    JP2022180282A