System

The system uses real-time video analysis and a generative AI model to optimize elevator operation by avoiding unnecessary stops, improving passenger convenience and efficiency.

JP2026022479APending Publication Date: 2026-02-12SOFTBANK GROUP CORP
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2024123996
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-07-30
Publication Date
2026-02-12

AI Technical Summary

Technical Problem

Elevators in crowded facilities often make unnecessary stops, causing passenger inconvenience and delayed arrival times due to inefficient operation.

Method used

A system that uses a camera to capture real-time video data, preprocesses it, and inputs it into a generative AI model to quantify congestion, determining optimal stopping floors based on a threshold to avoid unnecessary stops.

Benefits of technology

Optimizes elevator operation by reducing passenger waiting times and energy waste, ensuring efficient and comfortable travel.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026022479000001_ABST
    Figure 2026022479000001_ABST
Patent Text Reader

Abstract

A system is provided.SOLUTION: A system comprising: means for acquiring an image in an elevator with a camera in real time; means for inputting the acquired image to a generative AI model and quantifying a degree of congestion in the elevator; and means for determining a stopping floor of the elevator when a numerical value of the degree of congestion exceeds a predetermined value.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The technology of the present disclosure relates to a system. [Background technology]

[0002] Patent document 1 discloses a persona chatbot control method performed by at least one processor, the method including the steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to a description of the chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance. [Prior art documents] [Patent documents]

[0003] [Patent Document 1] Japanese Patent Publication No. 2022-180282 Summary of the Invention [Problem to be solved by the invention]

[0004] In facilities used by many people, such as shopping malls and office buildings, elevators often stop at floors where new passengers are waiting, even when the elevator is crowded. In such cases, not only are passengers forced to wait without being able to board the elevator, but existing passengers also face the problem of delayed arrival times due to unnecessary stops. This results in wasted time and inconvenience for passengers, hindering efficient elevator operation. The present invention aims to provide a system that increases passenger convenience by avoiding unnecessary stops when elevators are crowded and operating efficiently. [Means for solving the problem]

[0005] This invention provides a system that captures video in real time using a camera installed inside an elevator, inputs the video data into a generative AI model, and quantifies the degree of congestion inside the elevator. Specifically, the video data captured by the camera is preprocessed and input into a generative AI model to analyze the number and placement of passengers and calculate the degree of congestion. It also includes a means for determining the floor at which the elevator should stop if the congestion level exceeds a preset threshold. As a result, the system controls operation when the elevator is crowded so that it does not stop at floors where passengers are not getting off, allowing passengers to efficiently reach their destination floor.

[0006] The "camera" is a device that captures images inside the elevator in real time.

[0007] "Video data" refers to video information captured by a camera inside an elevator.

[0008] A "generative AI model" is an artificial intelligence model that inputs video data and quantifies the degree of congestion inside an elevator.

[0009] The "crowding level" is a numerical value that indicates the degree of congestion calculated based on the number and arrangement of passengers in the elevator.

[0010] The "threshold" refers to a preset reference value that serves as a criterion for determining the degree of congestion.

[0011] "Stop floor" means a floor at which an elevator stops to let passengers on or off.

[0012] "Preprocessing" refers to the processing of data before inputting it into a generative AI model, such as removing noise from video data or adjusting resolution.

[0013] "Operation instructions" means instructions regarding elevator operation to the elevator control system.

[0014] An "elevator control system" is a system for controlling elevator operation.

[0015] "Passenger" refers to a person using an elevator.

[0016] The "stop list" is a list of floors where the elevator is scheduled to stop next.

[0017] "Excess passengers" means the excess number of passengers in an elevator when the standard value is exceeded, based on the total number of passengers in the elevator. [Brief explanation of the drawings]

[0018] [Figure 1] 1 is a conceptual diagram showing an example of the configuration of a data processing system according to a first embodiment. [Figure 2] 1 is a conceptual diagram showing an example of main functions of a data processing device and a smart device according to a first embodiment. [Figure 3] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a second embodiment. [Figure 4] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and smart glasses according to a second embodiment. [Figure 5] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a third embodiment. [Figure 6] FIG. 11 is a conceptual diagram showing an example of main functions of a data processing device and a headset-type terminal according to a third embodiment. [Figure 7] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a fourth embodiment. [Figure 8] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and a robot according to a fourth embodiment. [Figure 9] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 10] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 11]FIG. 3 is a sequence diagram showing a processing flow of the data processing system according to the first embodiment. [Figure 12] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 1. [Figure 13] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system according to the second embodiment when an emotion engine is combined. [Figure 14] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 2 when an emotion engine is combined. DETAILED DESCRIPTION OF THE INVENTION

[0019] An example of an embodiment of a system according to the technology of the present disclosure will be described below with reference to the accompanying drawings.

[0020] First, the terms used in the following description will be explained.

[0021] In the following embodiments, a coded processor (hereinafter simply referred to as a "processor") may be a single arithmetic device or a combination of multiple arithmetic devices. Furthermore, a processor may be a single type of arithmetic device or a combination of multiple types of arithmetic devices. Examples of arithmetic devices include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), and an APU (Accelerated Processing Unit).

[0022] In the following embodiments, a coded RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a working memory by a processor.

[0023] In the following embodiments, the coded storage is one or more non-volatile storage devices that store various programs, various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), and magnetic tapes.

[0024] In the following embodiments, a communication I / F (Interface) with a symbol is an interface including a communication processor, an antenna, etc. The communication I / F controls communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), Bluetooth (registered trademark), etc.

[0025] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." In other words, "A and / or B" means that it may be only A, only B, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" is also applied when three or more things are expressed connected by "and / or."

[0026] [First embodiment]

[0027] FIG. 1 shows an example of the configuration of a data processing system 10 according to the first embodiment.

[0028] 1, a data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.

[0029] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0030] The smart device 14 includes a computer 36, a reception device 38, an output device 40, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The reception device 38, the output device 40, and the camera 42 are also connected to the bus 52.

[0031] The reception device 38 includes a touch panel 38A, a microphone 38B, and the like, and receives user input. The touch panel 38A detects contact with an indicator (for example, a pen or a finger) to receive user input by the touch of the indicator. The microphone 38B detects the user's voice to receive user input by voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.

[0032] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form of expression that the user 20 can perceive (for example, audio and / or text). The display 40A displays visible information such as text and images in accordance with instructions from the processor 46. The speaker 40B outputs audio in accordance with instructions from the processor 46. The camera 42 is a compact digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.

[0033] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54.

[0034] FIG. 2 shows an example of the main functions of the data processing device 12 and the smart device 14.

[0035] 2, in the data processing device 12, a specific process is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific process is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0036] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0037] In the smart device 14, the processor 46 performs the reception output process. The storage 50 stores a reception output program 60. The reception output program 60 is used in conjunction with the specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0038] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0039] System Configuration

[0040] The present invention is implemented in the following configuration. A camera is installed inside the elevator, and a server collects real-time video data captured by the camera. The server preprocesses the video data and inputs it into a generative AI model to quantify the elevator congestion level. If the congestion level exceeds a set threshold, the server determines the floor at which the elevator will stop, and omits stops if necessary.

[0041] Overview of program processing flow

[0042] 1. Camera image acquisition

[0043] The server receives real-time video data from cameras installed inside the elevator, which is used to understand the situation of passengers inside the elevator.

[0044] 2. Preprocessing of video data

[0045] The server performs preprocessing such as noise removal and resolution adjustment on the acquired video data before inputting it into the generative AI model.

[0046] 3. Analysis of video data

[0047] The server inputs the preprocessed video data into a generative AI model to analyze the number and location of passengers, which then quantifies the degree of congestion inside the elevator.

[0048] 4. Determining congestion

[0049] The server verifies the congestion level value obtained from the generative AI model and determines whether the congestion level exceeds a set threshold.

[0050] 5. Determining the stopping floor

[0051] The server obtains the current floor stop list based on the buttons inside the elevator, and then updates the floor stop list so that the elevator does not stop at floors where passengers inside the elevator do not plan to get off.

[0052] 6. Elevator operation control

[0053] The server controls elevator operation based on the updated floor stop list to avoid unnecessary stops during peak hours, and notifies passengers of the next floor stop via a display and voice guidance inside the elevator.

[0054] Specific examples

[0055] 1. Example in a shopping mall

[0056] In an elevator in a shopping mall, the server acquires video data from the camera every second. For example, suppose the elevator is very crowded at a certain time.

[0057] The server preprocesses this video data and inputs it into a generative AI model, which detects that there are 30 passengers in the elevator and quantifies the congestion level as "85."

[0058] Since the congestion level exceeds the set threshold "70", the server determines that the elevator is crowded.

[0059] Next, the server checks the button status inside the elevator and confirms that although the elevator was scheduled to stop on the 3rd, 5th, and 7th floors, there are no passengers getting off on the 3rd and 7th floors.

[0060] The server updates the floor stop list and adjusts it to skip floors 3 and 7.

[0061] Ultimately, the elevator will stop only on the fifth floor, and a display and voice guidance inside the elevator will inform passengers that the next floor is floor 5. In this way, unnecessary stops can be avoided and the elevator can be operated efficiently.

[0062] This invention optimizes the operation of crowded elevators, reducing wasted time for passengers and reducing the amount of time that passengers waste waiting when they cannot board, thereby improving the overall efficiency of elevator use.

[0063] The processing flow will be explained below.

[0064] Step 1:

[0065] The server receives real-time video data from a camera installed inside the elevator, capturing the entire elevator interior and allowing the system to understand the situation of passengers inside the elevator.

[0066] Step 2:

[0067] The server temporarily stores the acquired video data and begins preprocessing, which includes removing noise and adjusting the resolution of the video, preparing it for input to the generative AI model.

[0068] Step 3:

[0069] The server inputs the preprocessed video data into a generative AI model, which then detects the number of passengers from the video data and analyzes their placement and movements.

[0070] Step 4:

[0071] The generative AI model quantifies the congestion level based on the number and location of passengers. The server receives this quantified congestion level and proceeds to the next step.

[0072] Step 5:

[0073] The server compares the quantified congestion level with a preset threshold. If the congestion level exceeds the threshold, the elevator is determined to be congested, but if it does not exceed the threshold, it continues to operate normally.

[0074] Step 6:

[0075] If the elevator is determined to be busy, the server retrieves the current floor stop list, which is based on floor button presses inside the elevator.

[0076] Step 7:

[0077] The server analyzes the stop floor list to identify floors where passengers in the elevator plan to get off, extracts floors where passengers do not plan to get off, and updates the stop floor list so that the elevator does not stop at those floors.

[0078] Step 8:

[0079] The server generates new operation instructions based on the updated floor list, including the floor where the elevator will next stop.

[0080] Step 9:

[0081] The server sends the generated operation instructions to the elevator control system, which operates the elevator based on these instructions and avoids unnecessary stops.

[0082] Step 10:

[0083] The server sends the next floor information to the display device and voice guidance device inside the elevator, explicitly notifying passengers of the next floor, thus ensuring that passengers inside the elevator are properly informed.

[0084] As a concrete example, elevators in a shopping mall tend to be crowded during the morning peak hours. For example, the server inputs video data acquired from a camera into a generative AI model to calculate a congestion level of "85," which it determines is above the threshold of "70." As a result, it determines that there are no passengers on the third and seventh floors, and generates operation instructions to skip stopping at those floors. The elevator then stops only on the fifth floor, allowing for efficient operation.

[0085] Example 1

[0086] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0087] Conventional elevator systems lacked the means to properly manage passenger congestion and optimize operational efficiency. This resulted in unnecessary elevator stops and congestion that prevented passengers from moving smoothly. This also led to problems such as wasted energy and increased passenger stress.

[0088] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0089] In this invention, the server includes means for capturing images of the inside of the elevator with a camera in real time, means for preprocessing the captured image data, means for inputting the preprocessed image data into a generative AI model to quantify the congestion level inside the elevator, means for determining the elevator's stopping floor when the congestion level exceeds a preset threshold, and means for notifying the next stopping floor via a display device or voice guidance inside the elevator. This improves elevator operation efficiency, reduces unnecessary waiting time for passengers, and contributes to energy savings.

[0090] The "camera" is a device for capturing images inside the elevator in real time.

[0091] "Video data" refers to real-time image information of the inside of an elevator captured by a camera.

[0092] "Preprocessing" refers to the process of removing noise, adjusting resolution, normalizing, etc. from video data, and converting it into a format suitable for the generative AI model.

[0093] The "generative AI model" is a machine learning model that uses a deep learning algorithm to analyze and quantify the number of passengers and congestion levels from input video data.

[0094] The "crowding level" is a numerical value that indicates the degree of congestion calculated based on the number and arrangement of passengers in the elevator.

[0095] A "threshold" is a reference value set to trigger a specific action when the congestion level exceeds this threshold.

[0096] "Stop floor" means a floor at which an elevator stops during a particular operation.

[0097] A "display device" is a device installed inside an elevator that visually provides passengers with information such as the next floor.

[0098] "Voice guidance" is a system that notifies passengers of the next floor and other information by voice inside the elevator.

[0099] The present invention is a system that analyzes the congestion level in an elevator in real time and optimizes operation efficiency. This system is implemented with the following configuration.

[0100] System Configuration

[0101] The system consists of a camera installed inside the elevator, a server, a generative AI model, and an elevator control system.

[0102] Hardware and Software

[0103] camera:

[0104] Cameras will be installed inside the elevators to capture high-resolution video in real time, providing images at 30 frames per second.

[0105] server:

[0106] The server is a central processing unit that receives video data sent from the camera and performs preprocessing and analysis. It is equipped with a high-performance CPU and GPU and has sufficient computing resources to run generative AI models.

[0107] Generative AI models:

[0108] This is a machine learning model that analyzes video data and quantifies the degree of congestion in an elevator. It uses a model that uses a deep learning algorithm (such as YOLOv5).

[0109] Elevator Control System:

[0110] The system controls elevator operation based on control signals sent from the server, and notifies passengers of the next floor via a display device and voice guidance system inside the elevator.

[0111] Example of operation

[0112] 1. Elevators in shopping malls:

[0113] The server acquires video data from the camera inside the elevator every second.

[0114] The server first pre-processes the acquired video data, which includes noise removal, resolution adjustment, and normalization.

[0115] The preprocessed video data is input into a generative AI model (e.g., YOLOv5) to analyze the number and location of passengers in the elevator. The generative AI model detects that there are 30 passengers in the elevator and quantifies the congestion level as "85."

[0116] The server checks whether the congestion level exceeds the set threshold value "70".

[0117] The server checks the button status inside the elevator and confirms that although the elevator was scheduled to stop on the 3rd, 5th, and 7th floors, there are no passengers getting off on the 3rd and 7th floors.

[0118] The server updates the stop floor list and adjusts it to skip floors 3 and 7.

[0119] The elevator will finally stop on the fifth floor, and a display and voice guidance inside the elevator will inform passengers that the next floor is the fifth floor.

[0120] Example prompts for generative AI models

[0121] "Please use the camera footage from inside the elevator to quantify the current congestion level. The video data shows 30 passengers. The congestion level is 85."

[0122] The system of the present invention optimizes elevator operation efficiency, reduces passenger waiting times, cuts energy waste, and ensures that passengers are transported to their destinations quickly and efficiently, even during peak hours.

[0123] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0124] Step 1: Acquire camera footage

[0125] The server acquires video data in real time from cameras installed inside the elevator.

[0126] Input: Real-time video data sent from the camera.

[0127] Specific operation: Video data is sent to the server as stream data at 30 frames per second.

[0128] Output: Captured real-time video data.

[0129] Step 2: Preprocessing the video data

[0130] The server preprocesses the acquired video data before inputting it into the generative AI model.

[0131] Input: Video data acquired from the camera.

[0132] Specific operations: noise removal, resolution adjustment (e.g. convert to 640x480 pixels), pixel value normalization (scaling to the range 0-1).

[0133] Output: Preprocessed video data.

[0134] Step 3: Analyzing the video data

[0135] The server inputs the preprocessed video data into a generative AI model to analyze passenger numbers and placements.

[0136] Input: Preprocessed video data.

[0137] Specific operation: Uses a generative AI model (e.g., YOLOv5) to detect the number of passengers in an elevator from video data and analyze their placement.

[0138] Output: Number of passengers, location information and congestion level (e.g. "85").

[0139] Step 4: Determine crowding level

[0140] The server verifies the congestion level values ​​obtained from the generative AI model.

[0141] Input: The congestion level output from the generative AI model.

[0142] Specific behavior: Compare the congestion level value with a pre-set threshold (e.g. "70").

[0143] Output: Boolean value (e.g. true) indicating whether the congestion threshold is exceeded.

[0144] Step 5: Determine the stops

[0145] The server obtains the button status inside the elevator and checks and updates the current floor stop list.

[0146] Input: Button status inside the elevator, congestion level judgment result.

[0147] Specific behavior: Get the current floor stop list and remove floors where passengers in the elevator do not plan to get off (e.g., floors 3 and 7) from the floor stop list.

[0148] Output: Updated floor list (e.g. only floor 5).

[0149] Step 6: Elevator operation control

[0150] The server controls elevator operation based on the updated stop list and notifies passengers of the next stop.

[0151] Input: Updated floor stop list.

[0152] What it does: Sends instructions to the elevator control system to stop only at the specified floor, and notifies passengers of the next stop via a visual display and voice announcement (e.g., "Next stop is floor 5").

[0153] Output: Elevator control signal, next floor notification information.

[0154] Through the above process, elevator operation efficiency can be optimized, congestion can be reduced, and energy can be saved.

[0155] (Application example 1)

[0156] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0157] In conventional factory robot movement, it has been difficult to grasp congestion conditions in real time and select efficient movement routes. This has resulted in reduced productivity and hindered efficient work. In particular, in areas of the factory where congestion frequently occurs, robot movement is hindered, resulting in a significant drop in production speed and work efficiency.

[0158] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0159] In this invention, the server includes a means for capturing images of the area in real time using a camera, a means for inputting the captured image data into a generative AI model to quantify the degree of congestion in the area, and a means for optimizing the robot's movement route when the numerical value of the degree of congestion exceeds a preset threshold. This makes it possible to analyze the congestion situation in real time and instruct the robot on the optimal movement route.

[0160] A "camera" is a device that captures video within a specified area in real time.

[0161] A "generative AI model" is an artificial intelligence model used to quantify the degree of congestion based on acquired video data.

[0162] The "degree of congestion" is a numerical value that indicates the degree of congestion calculated from the number of people and their locations within a specific area.

[0163] A "threshold" is a numerical value that serves as a reference for taking a specific action when the congestion level exceeds this threshold.

[0164] "Stop floor" means a floor where an elevator stops to allow passengers to board or alight.

[0165] A "movement route" is a path chosen by a robot or other moving object when it moves within an area.

[0166] "Notification" is the act of informing a robot or worker of the next action to be taken or the selected route.

[0167] "Preprocessing" refers to the process of processing the video data, such as removing noise and adjusting resolution, before inputting it into the generative AI model.

[0168] System Configuration

[0169] This invention is implemented with the following configuration: Cameras are installed in a factory area, and a server collects real-time video data captured by the cameras. The server is a system that preprocesses this video data and inputs it into a generative AI model to quantify the congestion level of the area. If the congestion level exceeds a set threshold, the server optimizes the robot's movement route and changes the route as necessary. It then notifies the robot or worker of this information.

[0170] Overview of program processing flow

[0171] 1. Camera image acquisition

[0172] The server collects real-time video data from cameras installed in the area, which is used to understand the congestion situation within the area.

[0173] 2. Preprocessing of video data

[0174] The server performs preprocessing such as noise removal and resolution adjustment on the acquired video data before inputting it into the generative AI model.

[0175] 3. Analysis of video data

[0176] The server inputs the preprocessed video data into a generative AI model to analyze the number of people and their locations within the area, which then quantifies the level of congestion.

[0177] 4. Determining congestion

[0178] The server verifies the congestion level value obtained from the generative AI model and determines whether the congestion level exceeds a set threshold.

[0179] 5. Deciding on a travel route

[0180] The server calculates the optimal route for the robot based on the congestion information, and selects a route that avoids congestion.

[0181] 6. Notification

[0182] The server notifies the robots and workers of the selected movement route, enabling them to act quickly and efficiently.

[0183] Hardware and software used

[0184] Hardware:

[0185] Camera: A device installed in a factory area that captures images within the area in real time.

[0186] Server: A computer for processing and analyzing video data.

[0187] software:

[0188] OpenCV: A library for acquiring and preprocessing camera footage.

[0189] Keras: A framework for loading and analyzing generative AI models.

[0190] Specific examples

[0191] When a specific area in the factory is extremely congested, the server receives real-time video data from the camera, preprocesses it, and then inputs it into the generative AI model. For example, if the area is highly congested, the server instructs the robot to select a different route. If the area is less congested, the robot will select the normal route to maximize movement efficiency.

[0192] Prompt Sentence Examples

[0193] "Please obtain video data from camera ID: {camera_id}, analyze the congestion level, and calculate the optimal travel route."

[0194] In this way, the operating efficiency of robots in the factory can be optimized, congestion can be avoided, and productivity can be improved.

[0195] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[0196] Step 1:

[0197] The server acquires video data in real time from cameras installed in the area. Specifically, it captures the video from the cameras and sends the video data to the server. The input data is the video stream from the cameras, and the output data is the video frames stored on the server.

[0198] Step 2:

[0199] The server preprocesses the captured video data, specifically using OpenCV to remove noise and adjust resolution. The input data is raw video frames, and the output data is video data converted into a format suitable for the generative AI model.

[0200] Step 3:

[0201] The server inputs the preprocessed video data into a generative AI model to analyze the number of people and their locations within the area. Specifically, it uses Keras to analyze the video data and quantify the degree of congestion. The input data is the preprocessed video data, and the output data is a numerical value indicating the degree of congestion.

[0202] Step 4:

[0203] The server verifies the congestion level value obtained from the generative AI model and determines whether the congestion level exceeds a set threshold. The input data is the congestion level value, and the output data is a boolean value indicating whether the congestion level exceeds the threshold.

[0204] Step 5:

[0205] The server calculates the optimal route for the robot based on the congestion information. Specifically, it executes an algorithm to select the least congested route. The input data is the congestion level value and default route information, and the output data is the optimized route.

[0206] Step 6:

[0207] The server notifies the robot and the worker of the selected movement route. Specifically, it sends the notification via the network to the robot's operation system or the worker's terminal. The input data is the optimized movement route, and the output data is the notification information.

[0208] Through the above processing steps, it is expected that the operating efficiency of robots in the factory will be optimized and productivity will be improved by avoiding congestion.

[0209] Furthermore, an emotion engine that estimates the user's emotion may be combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59 and perform identification processing using the user's emotion.

[0210] System Configuration

[0211] The present invention is a system that combines a camera installed inside an elevator with an emotion engine. The camera captures video footage inside the elevator in real time, preprocesses this video data, and inputs it into a generative AI model. The generative AI model analyzes the number and placement of passengers and quantifies the level of congestion. If this level of congestion exceeds a set threshold, the server uses the data from the emotion engine to determine the elevator's stopping floor. The emotion engine uses the camera's video data to recognize passenger emotions and, in conjunction with the level of congestion, determines the optimal stopping floor.

[0212] Overview of program processing flow

[0213] 1. Camera image acquisition

[0214] The server collects real-time video data from cameras installed inside the elevator, which serves as the basis for understanding the situation and emotions of passengers inside the elevator.

[0215] 2. Preprocessing of video data

[0216] The server pre-processes the captured video data, which includes noise reduction and resolution adjustment.

[0217] 3. Analysis of video data

[0218] The server inputs the preprocessed video data into a generative AI model, which analyzes the number and placement of passengers and quantifies the level of congestion.

[0219] 4. Emotional Recognition

[0220] The server also inputs the same video data into an emotion engine to recognize each passenger's emotions. The emotion engine uses facial recognition and facial expression analysis technology to detect passengers' feelings such as discomfort or stress.

[0221] 5. Integrating crowding and emotion data

[0222] The server integrates congestion data from the generative AI model with emotion data from the emotion engine, and optimizes elevator operation based on the integrated data.

[0223] 6. Determining the stopping floor

[0224] The server updates the elevator's floor stop list based on the integrated data, and adjusts the elevator to avoid stopping at floors where passengers feel uncomfortable, especially when the congestion level exceeds a threshold and discomfort is detected by the emotion engine.

[0225] 7. Elevator operation control

[0226] The server controls elevator operation based on the new stop list to avoid unnecessary stops, and uses displays and audio announcements inside the elevator to notify passengers of the next stop.

[0227] Specific examples

[0228] 1. Example in a shopping mall

[0229] In an elevator in a shopping mall, a server captures video data from a camera every second, including not only the number and location of passengers but also their facial expressions.

[0230] The server preprocesses this video data and then inputs it into the generative AI model and emotion engine. The generative AI model detects that there are 30 passengers and calculates the congestion level as 85. Meanwhile, the emotion engine recognizes that 10 passengers are feeling uncomfortable.

[0231] The server determines that the elevator is crowded because the congestion level of "85" exceeds the threshold of "70." Furthermore, taking into account the discomfort data, the server assesses that continuing normal operation may increase passenger discomfort.

[0232] An elevator that was scheduled to stop on floors 3, 5, and 7 will check to make sure there are no passengers on those floors, and update its operation instructions to stop only on floor 5, reflecting congestion and sentiment data, to most effectively reduce passenger discomfort.

[0233] The server sends the generated new operation instructions to the elevator control system, and passengers are notified that the next stop is floor 5 via a display and voice guidance inside the elevator.

[0234] This embodiment not only improves the efficiency of elevator operation but also enhances passenger convenience. It not only avoids unnecessary stops in crowded situations but also provides a comfortable user environment by taking passengers' emotions into consideration.

[0235] The processing flow will be explained below.

[0236] Step 1:

[0237] The server acquires real-time video data from a camera installed inside the elevator, which includes information on the number and location of passengers in the elevator, as well as emotional information such as facial expressions.

[0238] Step 2:

[0239] The server stores the acquired video data in temporary storage and begins preprocessing, which includes noise reduction and resolution adjustment, to prepare it for input to the generative AI model and emotion engine.

[0240] Step 3:

[0241] The server inputs the preprocessed video data into a generative AI model, which analyzes the number and location of passengers from the video data to quantify the level of congestion.

[0242] Step 4:

[0243] At the same time, the server inputs the pre-processed video data into the emotion engine, which uses facial recognition and facial expression analysis technologies to recognize each passenger's emotions (discomfort, stress, etc.).

[0244] Step 5:

[0245] As a result of the analysis, the generative AI model detects that there are 30 passengers in the elevator and calculates the congestion level as "85." This data is received by the server.

[0246] Step 6:

[0247] As a result of the analysis, the emotion engine recognizes that 10 passengers are feeling uncomfortable. This data is received by the server.

[0248] Step 7:

[0249] The server combines the congestion data obtained from the generative AI model with the emotion data obtained from the emotion engine. Since the congestion level is "85" and there are many passengers feeling uncomfortable, it determines that the elevator is crowded.

[0250] Step 8:

[0251] The server gets the elevator's current list of stops, say it's scheduled to stop at floors 3, 5, and 7.

[0252] Step 9:

[0253] The server analyzes the list of stops and identifies floors where the passengers in the elevator do not plan to get off. It confirms that the passengers do not plan to get off on the third and seventh floors.

[0254] Step 10:

[0255] Taking into account the congestion level and emotion data, the server updates the stop floor list to skip stops on floors 3 and 7 and stop only on floor 5 in order to most effectively reduce passenger discomfort.

[0256] Step 11:

[0257] The server generates new operation instructions based on the updated floor stop list and sends the instructions to the elevator control system to control elevator operation.

[0258] Step 12:

[0259] When the elevator is in operation, the server sends information about the next floor to the display device and voice guidance system inside the elevator, notifying passengers of the next floor. Specifically, the server informs passengers that the next floor is floor 5.

[0260] As a concrete example, suppose an elevator in a shopping mall is used by several passengers during the morning peak, with the camera detecting 30 passengers and the emotion engine recognizing the discomfort of 10 of them. Based on this data, the server calculates a congestion level of 85, exceeding the preset threshold of 70, and controls the elevator so that it stops only on the fifth floor, since no passengers are getting off on the third or seventh floor. Passengers are notified by display and audio guidance that the elevator will next stop on the fifth floor, ensuring efficient and comfortable operation.

[0261] Example 2

[0262] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0263] Conventional elevator operation systems lacked the means to achieve efficient operation, and were unable to determine the optimal stopping floor, especially during crowded times or based on passenger emotions. Furthermore, when passengers felt uncomfortable inside the elevator, conventional systems had difficulty responding appropriately. This sometimes led to lower satisfaction among elevator users.

[0264] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[0265] In this invention, the server includes means for capturing images of the inside of the elevator with a camera in real time, means for preprocessing the captured image data, means for inputting the preprocessed image data into a generative AI model and quantifying the degree of congestion inside the elevator, means for inputting the same image data into an emotion engine and recognizing passenger emotions, and means for integrating the congestion data obtained from the generative AI model and the emotion data obtained from the emotion engine and determining the floors at which the elevator should stop. This enables optimal operation that simultaneously takes into account the congestion status inside the elevator and the emotions of passengers.

[0266] The "camera" is a photographic device that captures images inside the elevator in real time.

[0267] "Video data" is digital data captured by a camera showing the situation inside the elevator.

[0268] "Preprocessing" refers to processes such as noise removal and resolution adjustment that are performed to improve the quality of acquired video data.

[0269] The "generative AI model" is an artificial intelligence model that analyzes video data and quantifies the degree of congestion inside an elevator.

[0270] "Crowding level" is a number that indicates the degree of congestion calculated by the generative AI model based on the number and placement of passengers in the elevator.

[0271] The "emotion engine" is an engine that analyzes video data and recognizes emotions from passengers' facial expressions.

[0272] "Integrated data" refers to data that integrates congestion data obtained from the generative AI model and emotion data obtained from the emotion engine.

[0273] "Stop floor" refers to the floor at which the elevator is scheduled to stop.

[0274] This invention is a system that combines a camera installed inside the elevator, a generative AI model, and an emotion engine, and simultaneously analyzes the congestion situation inside the elevator and the emotions of passengers to achieve optimal operation. Below, we will explain the hardware and software, as well as the specific processing and calculations.

[0275] Hardware

[0276] 1. Camera

[0277] These are IP cameras or surveillance cameras installed inside elevators. These cameras capture images from inside the elevator in real time and send them to a streaming server.

[0278] 2. Server

[0279] This is the central processing unit of the system, and performs preprocessing and analysis of video data. The server is equipped with a high-performance CPU and GPU, and runs deep learning frameworks such as Python, TensorFlow, and PyTorch.

[0280] software

[0281] 1. Preprocessing software

[0282] This software tool removes noise from video data, adjusts resolution, and adjusts contrast, etc. This improves the accuracy of analysis.

[0283] 2. Generative AI Models

[0284] This is an artificial intelligence model that analyzes the number and placement of passengers in an elevator from video data and quantifies the degree of congestion. It is usually implemented using Python, TensorFlow, and PyTorch.

[0285] 3. Emotion Engine

[0286] This is a software engine for recognizing passenger emotions from video data. It uses face recognition and facial expression analysis technologies and uses Python, OpenCV, and the Dlib library.

[0287] Example of a system

[0288] Elevator operation in shopping malls

[0289] The server collects video data every second from cameras installed in elevators in the shopping mall, including the number of passengers, their locations, and their facial expressions.

[0290] 1. Pretreatment

[0291] The server pre-processes this video data, removing noise and adjusting the resolution, for example, if the initial resolution is 720p, it will change it to 1080p to improve the quality of the data.

[0292] 2. Input to the generative AI model and emotion engine

[0293] The preprocessed data is input into the generative AI model and the emotion engine. The generative AI model detects that there are 30 passengers in the elevator and quantifies the congestion level as "85." Meanwhile, the emotion engine recognizes that 10 of the passengers are feeling uncomfortable.

[0294] 3. Integration and Operation Decisions

[0295] The server combines the congestion data obtained from the generative AI model with the emotion data obtained from the emotion engine. Based on this combined data, it determines the floors where the elevator should stop. For example, an elevator that is normally scheduled to stop on the third, fifth, and seventh floors will update its operation instructions to take into account the combined data and avoid stopping on the third and seventh floors, where there are no passengers, and stop only on the fifth floor.

[0296] Example prompt

[0297] For example, the following prompts might be used by a generative AI model:

[0298] "Analyze the number and location of passengers from video data inside the elevator, and quantify the congestion level on a scale of 0 to 100."

[0299] This invention not only improves the efficiency of elevator operation but also enhances passenger convenience. It not only avoids unnecessary stops in crowded situations but also provides a comfortable user environment by taking passengers' emotions into consideration.

[0300] The flow of the identification process in the second embodiment will be described with reference to FIG.

[0301] Program processing flow

[0302] The subject should be either the server, the terminal, or the user, and the explanation should be written in the plain voice (da / dearu style).

[0303] Step 1: Acquire camera footage

[0304] The server receives real-time video data from the camera. The camera captures video footage inside the elevator and streams it to the server. This video data includes the number and placement of passengers, facial expressions, etc.

[0305] Input: Real-time video from inside the elevator

[0306] Output: Captured video data

[0307] Specific behavior:

[0308] 1. The camera captures a frame every second and sends it to the streaming server.

[0309] 2. The server receives this stream and stores it in a database.

[0310] Step 2: Preprocessing the video data

[0311] The server pre-processes the captured video data, which includes noise reduction, resolution adjustment, and contrast enhancement.

[0312] Input: Captured video data

[0313] Output: Pre-processed video data

[0314] Specific behavior:

[0315] 1. A noise reduction filter is applied to the video data.

[0316] 2. Adjust the video resolution (e.g. convert from 720p to 1080p).

[0317] 3. Adjust contrast and brightness and clean abnormal pixels.

[0318] Step 3: Analyzing the video data

[0319] The server inputs the preprocessed video data into a generative AI model, which analyzes the number and location of passengers from the video data and quantifies the level of congestion.

[0320] Input: Preprocessed video data

[0321] Output: Numerical data of congestion

[0322] Specific behavior:

[0323] 1. Preprocessed video data is passed to a generative AI model.

[0324] 2. The model extracts facial and body features of each passenger and identifies their number and location.

[0325] 3. Based on the number and location of passengers, the congestion level is quantified on a scale of 0 to 100.

[0326] Step 4: Recognize emotions

[0327] The server then inputs the same video data into an emotion engine to recognize passenger emotions, which uses facial recognition and facial expression analysis to classify emotions.

[0328] Input: Preprocessed video data

[0329] Output: Passenger emotion data

[0330] Specific behavior:

[0331] 1. Detect the faces of each passenger from the preprocessed video data.

[0332] 2. Use a facial recognition algorithm to analyze each passenger's facial expression.

[0333] 3. Classify emotions such as discomfort, stress, and joy, and save the results as numerical data.

[0334] Step 5: Integrating crowding and emotion data

[0335] The server combines the congestion data obtained from the generative AI model with the emotion data obtained from the emotion engine, and this data is used to control elevator operation.

[0336] Input: Numerical data on congestion levels, passenger sentiment data

[0337] Output: Integrated traffic control data

[0338] Specific behavior:

[0339] 1. Obtain numerical data on congestion levels and emotional data.

[0340] 2. Combine both sets of data to create a detailed picture of elevator operation.

[0341] Step 6: Determine the stops

[0342] The server determines the elevator's stopping floor based on the integrated data, taking into account the level of congestion and passenger sentiment to select the optimal stopping floor.

[0343] Input: Integrated traffic control data

[0344] Output: Updated floor stop list

[0345] Specific behavior:

[0346] 1. Analyze the integrated data to assess congestion and emotions.

[0347] 2. If the threshold is exceeded, floors with fewer passengers or less discomfort will be selected preferentially.

[0348] 3. Create an updated list of stops.

[0349] Step 7: Elevator operation control

[0350] The server controls elevator operation based on the new stop list, avoiding unnecessary stops and notifying passengers of the operation status.

[0351] Input: Updated floor list

[0352] Output: Operation control instructions and passenger notifications

[0353] Specific behavior:

[0354] 1. Coordinate elevator operation based on a list of stops.

[0355] 2. Use displays and voice guidance systems inside the elevator to inform passengers of the next floor (e.g., announce "Next floor is floor 5").

[0356] (Application example 2)

[0357] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0358] Conventional elevator systems only analyzed congestion levels and did not take passengers' emotions into consideration. As a result, even when congestion levels were high, it was difficult to alleviate passenger discomfort, potentially resulting in a poor user experience. Furthermore, it was difficult to properly control elevator operation to avoid unnecessary elevator stops, hindering efficient operation.

[0359] The identification processing by the identification processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for capturing video images inside the elevator with a camera in real time, means for inputting the captured video data into a generative AI model and quantifying the congestion level inside the elevator, means for determining the elevator's stopping floor when the congestion level numerical value exceeds a preset threshold, means for inputting the video data into an emotion engine to recognize passenger emotions, means for integrating the emotion score recognized by the emotion engine with the congestion level data, and means for optimizing the elevator's stopping floors based on the integrated data. This enables efficient and comfortable elevator operation that takes into account not only the number of passengers but also the emotions of the passengers.

[0360] The "camera" is a device that captures images inside the elevator in real time.

[0361] "Video data" refers to video information inside the elevator captured by a camera.

[0362] The "generative AI model" is an artificial intelligence algorithm that analyzes the acquired video data and quantifies the degree of congestion inside the elevator.

[0363] The "crowding level" is a numerical value that indicates the degree of congestion calculated based on the number and arrangement of passengers in the elevator.

[0364] A "threshold" is a set value that, when exceeded, triggers a specific action (for example, determining the elevator's stopping floor).

[0365] The "emotion engine" is an algorithm that analyzes video data and recognizes passenger emotions.

[0366] The "emotion score" is a numerical representation of the passenger's emotional state as recognized by the emotion engine.

[0367] "Integrated data" is a dataset that integrates congestion data and emotion scores.

[0368] "Means for determining elevator stopping floors" refers to algorithms or devices that determine the optimal elevator stopping floor based on congestion and emotion score data.

[0369] This invention is a system for optimizing elevator operations in brick-and-mortar stores and shopping malls. The system consists of a camera, a server, a generative AI model, an emotion engine, and an elevator control system.

[0370] System Configuration

[0371] 1. Camera:

[0372] The cameras are installed inside the elevator and capture video data in real time, including the number, location, and facial expressions of passengers.

[0373] 2. Server:

[0374] The server preprocesses the acquired video data, which includes noise reduction and resolution adjustment. The preprocessed video data is then input into the generative AI model and emotion engine.

[0375] 3. Generative AI Model:

[0376] The generative AI model analyzes the pre-processed video data and quantifies the congestion level based on the number and location of passengers. For example, if there are 30 passengers, the congestion level is calculated as 85.

[0377] 4. Emotion Engine:

[0378] The emotion engine analyzes pre-processed video data to recognize passenger emotions. It uses facial recognition and facial expression analysis technology to detect emotions such as discomfort and stress. For example, it can recognize that 10 passengers are feeling discomfort.

[0379] 5. Elevator control system:

[0380] The server combines the congestion data from the generative AI model with the emotion data from the emotion engine. It then optimizes the elevator's floor stops based on the combined data. For example, if the congestion level of 85 exceeds the threshold of 70, it takes into account the discomfort data and determines that continuing normal operation will increase passenger discomfort.

[0381] Elevator stop lists will be updated and operational instructions will be updated to stop only on the first and fourth floors to minimize congestion.

[0382] Specific examples

[0383] In an elevator in a shopping mall, a server collects video data from a camera every second. This data includes not only the number and location of passengers, but also their facial expressions. The server preprocesses this video data before inputting it into a generative AI model and an emotion engine. For example, the generative AI model detects that there are 30 passengers and calculates the congestion level as "85." Meanwhile, the emotion engine recognizes that 10 passengers are feeling uncomfortable.

[0384] The server determines that the elevator is crowded because the congestion level of 85 exceeds the threshold of 70. Furthermore, taking into account discomfort data, the server assesses that continuing normal operation may increase passenger discomfort. The elevator, which was scheduled to stop on the third, fifth, and seventh floors, updates its operation instructions to stop only on the first and fourth floors, reflecting the congestion level and emotional data to most effectively reduce passenger discomfort.

[0385] Prompt Sentence Examples

[0386] Number of passengers in the elevator: 30

[0387] Passenger Sentiment Score:

[0388] 10 people found it "unpleasant"

[0389] 15 people answered "normal",

[0390] 5 people said "comfortable",

[0391] Crowding level: 85

[0392] Threshold: 70

[0393] In this way, the server can realize efficient and comfortable elevator operation that takes into account not only the number of passengers but also their emotions.

[0394] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[0395] Step 1:

[0396] The server acquires video data in real time from a camera installed inside the elevator. The input is the raw video from the camera, and the output is the acquired raw video data. This video data includes the number, location, and facial expressions of passengers inside the elevator.

[0397] Step 2:

[0398] The server preprocesses the acquired video data. The input is the raw video data acquired in step 1, and the output is preprocessed video data with noise removal and resolution adjustment. Specific operations include applying a noise removal filter and resizing the resolution to a fixed size.

[0399] Step 3:

[0400] The server inputs the preprocessed video data into the generative AI model. The input is the preprocessed video data, and the output is data that quantifies the degree of congestion inside the elevator. Specifically, the generative AI model analyzes the number and placement of passengers and calculates the degree of congestion as a number (e.g., "85").

[0401] Step 4:

[0402] The server also inputs the preprocessed video data into the emotion engine. The input is the preprocessed video data generated in step 2, and the output is the emotion score for each passenger. Specifically, the emotion engine uses face recognition technology and facial expression analysis technology to calculate each passenger's emotion (e.g., "uncomfortable," "neutral," or "comfortable") as a score.

[0403] Step 5:

[0404] The server integrates the congestion data from the generative AI model and the emotion data from the emotion engine. The input is the congestion data and emotion score, and the output is an integrated dataset. Specifically, it combines each data into a single dataset to comprehensively evaluate the elevator's condition.

[0405] Step 6:

[0406] The server optimizes elevator stops based on the integrated data. The input is the integrated dataset, and the output is an optimized elevator stop list. Specifically, if the congestion level exceeds a threshold and discomfort is high, the server adjusts the elevator stops to reduce discomfort (e.g., stopping only on the first and fourth floors).

[0407] Step 7:

[0408] The server sends the optimized elevator stop list to the elevator control system. The input is the optimized stop list, and the output is the operation instructions to be executed by the elevator control system. Specifically, the server sends new stop information to the control system to control the actual operation of the elevator.

[0409] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0410] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0411] In the above embodiment, an example in which the specific process is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific process may be performed by the smart device 14.

[0412] [Second embodiment]

[0413] FIG. 3 shows an example of the configuration of a data processing system 210 according to the second embodiment.

[0414] 3, the data processing system 210 includes the data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.

[0415] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0416] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, and the camera 42 are also connected to the bus 52.

[0417] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[0418] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[0419] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[0420] Fig. 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Fig. 4, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[0421] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0422] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0423] In the smart glasses 214, the reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0424] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal."

[0425] System Configuration

[0426] The present invention is implemented in the following configuration. A camera is installed inside the elevator, and a server collects real-time video data captured by the camera. The server preprocesses the video data and inputs it into a generative AI model to quantify the elevator congestion level. If the congestion level exceeds a set threshold, the server determines the floor at which the elevator will stop, and omits stops if necessary.

[0427] Overview of program processing flow

[0428] 1. Camera image acquisition

[0429] The server receives real-time video data from cameras installed inside the elevator, which is used to understand the situation of passengers inside the elevator.

[0430] 2. Preprocessing of video data

[0431] The server performs preprocessing such as noise removal and resolution adjustment on the acquired video data before inputting it into the generative AI model.

[0432] 3. Analysis of video data

[0433] The server inputs the preprocessed video data into a generative AI model to analyze the number and location of passengers, which then quantifies the degree of congestion inside the elevator.

[0434] 4. Determining congestion

[0435] The server verifies the congestion level value obtained from the generative AI model and determines whether the congestion level exceeds a set threshold.

[0436] 5. Determining the stopping floor

[0437] The server obtains the current floor stop list based on the buttons inside the elevator, and then updates the floor stop list so that the elevator does not stop at floors where passengers inside the elevator do not plan to get off.

[0438] 6. Elevator operation control

[0439] The server controls elevator operation based on the updated floor stop list to avoid unnecessary stops during peak hours, and notifies passengers of the next floor stop via a display and voice guidance inside the elevator.

[0440] Specific examples

[0441] 1. Example in a shopping mall

[0442] In an elevator in a shopping mall, the server acquires video data from the camera every second. For example, suppose the elevator is very crowded at a certain time.

[0443] The server preprocesses this video data and inputs it into a generative AI model, which detects that there are 30 passengers in the elevator and quantifies the congestion level as "85."

[0444] Since the congestion level exceeds the set threshold "70", the server determines that the elevator is crowded.

[0445] Next, the server checks the button status inside the elevator and confirms that although the elevator was scheduled to stop on the 3rd, 5th, and 7th floors, there are no passengers getting off on the 3rd and 7th floors.

[0446] The server updates the floor stop list and adjusts it to skip floors 3 and 7.

[0447] Ultimately, the elevator will stop only on the fifth floor, and a display and voice guidance inside the elevator will inform passengers that the next floor is floor 5. In this way, unnecessary stops can be avoided and the elevator can be operated efficiently.

[0448] This invention optimizes the operation of crowded elevators, reducing wasted time for passengers and reducing the amount of time that passengers waste waiting when they cannot board, thereby improving the overall efficiency of elevator use.

[0449] The processing flow will be explained below.

[0450] Step 1:

[0451] The server receives real-time video data from a camera installed inside the elevator, capturing the entire elevator interior and allowing the system to understand the situation of passengers inside the elevator.

[0452] Step 2:

[0453] The server temporarily stores the acquired video data and begins preprocessing, which includes removing noise and adjusting the resolution of the video, preparing it for input to the generative AI model.

[0454] Step 3:

[0455] The server inputs the preprocessed video data into a generative AI model, which then detects the number of passengers from the video data and analyzes their placement and movements.

[0456] Step 4:

[0457] The generative AI model quantifies the congestion level based on the number and location of passengers. The server receives this quantified congestion level and proceeds to the next step.

[0458] Step 5:

[0459] The server compares the quantified congestion level with a preset threshold. If the congestion level exceeds the threshold, the elevator is determined to be congested, but if it does not exceed the threshold, it continues to operate normally.

[0460] Step 6:

[0461] If the elevator is determined to be busy, the server retrieves the current floor stop list, which is based on floor button presses inside the elevator.

[0462] Step 7:

[0463] The server analyzes the stop floor list to identify floors where passengers in the elevator plan to get off, extracts floors where passengers do not plan to get off, and updates the stop floor list so that the elevator does not stop at those floors.

[0464] Step 8:

[0465] The server generates new operation instructions based on the updated floor list, including the floor where the elevator will next stop.

[0466] Step 9:

[0467] The server sends the generated operation instructions to the elevator control system, which operates the elevator based on these instructions and avoids unnecessary stops.

[0468] Step 10:

[0469] The server sends the next floor information to the display device and voice guidance device inside the elevator, explicitly notifying passengers of the next floor, thus ensuring that passengers inside the elevator are properly informed.

[0470] As a concrete example, elevators in a shopping mall tend to be crowded during the morning peak hours. For example, the server inputs video data acquired from a camera into a generative AI model to calculate a congestion level of "85," which it determines is above the threshold of "70." As a result, it determines that there are no passengers on the third and seventh floors, and generates operation instructions to skip stopping at those floors. The elevator then stops only on the fifth floor, allowing for efficient operation.

[0471] Example 1

[0472] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0473] Conventional elevator systems lacked the means to properly manage passenger congestion and optimize operational efficiency. This resulted in unnecessary elevator stops and congestion that prevented passengers from moving smoothly. This also led to problems such as wasted energy and increased passenger stress.

[0474] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0475] In this invention, the server includes means for capturing images of the inside of the elevator with a camera in real time, means for preprocessing the captured image data, means for inputting the preprocessed image data into a generative AI model to quantify the congestion level inside the elevator, means for determining the elevator's stopping floor when the congestion level exceeds a preset threshold, and means for notifying the next stopping floor via a display device or voice guidance inside the elevator. This improves elevator operation efficiency, reduces unnecessary waiting time for passengers, and contributes to energy savings.

[0476] The "camera" is a device for capturing images inside the elevator in real time.

[0477] "Video data" refers to real-time image information of the inside of an elevator captured by a camera.

[0478] "Preprocessing" refers to the process of removing noise, adjusting resolution, normalizing, etc. from video data, and converting it into a format suitable for the generative AI model.

[0479] The "generative AI model" is a machine learning model that uses a deep learning algorithm to analyze and quantify the number of passengers and congestion levels from input video data.

[0480] The "crowding level" is a numerical value that indicates the degree of congestion calculated based on the number and arrangement of passengers in the elevator.

[0481] A "threshold" is a reference value set to trigger a specific action when the congestion level exceeds this threshold.

[0482] "Stop floor" means a floor at which an elevator stops during a particular operation.

[0483] A "display device" is a device installed inside an elevator that visually provides passengers with information such as the next floor.

[0484] "Voice guidance" is a system that notifies passengers of the next floor and other information by voice inside the elevator.

[0485] The present invention is a system that analyzes the congestion level in an elevator in real time and optimizes operation efficiency. This system is implemented with the following configuration.

[0486] System Configuration

[0487] The system consists of a camera installed inside the elevator, a server, a generative AI model, and an elevator control system.

[0488] Hardware and Software

[0489] camera:

[0490] Cameras will be installed inside the elevators to capture high-resolution video in real time, providing images at 30 frames per second.

[0491] server:

[0492] The server is a central processing unit that receives video data sent from the camera and performs preprocessing and analysis. It is equipped with a high-performance CPU and GPU and has sufficient computing resources to run generative AI models.

[0493] Generative AI models:

[0494] This is a machine learning model that analyzes video data and quantifies the degree of congestion in an elevator. It uses a model that uses a deep learning algorithm (such as YOLOv5).

[0495] Elevator Control System:

[0496] The system controls elevator operation based on control signals sent from the server, and notifies passengers of the next floor via a display device and voice guidance system inside the elevator.

[0497] Example of operation

[0498] 1. Elevators in shopping malls:

[0499] The server acquires video data from the camera inside the elevator every second.

[0500] The server first pre-processes the acquired video data, which includes noise removal, resolution adjustment, and normalization.

[0501] The preprocessed video data is input into a generative AI model (e.g., YOLOv5) to analyze the number and location of passengers in the elevator. The generative AI model detects that there are 30 passengers in the elevator and quantifies the congestion level as "85."

[0502] The server checks whether the congestion level exceeds the set threshold value "70".

[0503] The server checks the button status inside the elevator and confirms that although the elevator was scheduled to stop on the 3rd, 5th, and 7th floors, there are no passengers getting off on the 3rd and 7th floors.

[0504] The server updates the stop floor list and adjusts it to skip floors 3 and 7.

[0505] The elevator will finally stop on the fifth floor, and a display and voice guidance inside the elevator will inform passengers that the next floor is the fifth floor.

[0506] Example prompts for generative AI models

[0507] "Please use the camera footage from inside the elevator to quantify the current congestion level. The video data shows 30 passengers. The congestion level is 85."

[0508] The system of the present invention optimizes elevator operation efficiency, reduces passenger waiting times, cuts energy waste, and ensures that passengers are transported to their destinations quickly and efficiently, even during peak hours.

[0509] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0510] Step 1: Acquire camera footage

[0511] The server acquires video data in real time from cameras installed inside the elevator.

[0512] Input: Real-time video data sent from the camera.

[0513] Specific operation: Video data is sent to the server as stream data at 30 frames per second.

[0514] Output: Captured real-time video data.

[0515] Step 2: Preprocessing the video data

[0516] The server preprocesses the acquired video data before inputting it into the generative AI model.

[0517] Input: Video data acquired from the camera.

[0518] Specific operations: noise removal, resolution adjustment (e.g. convert to 640x480 pixels), pixel value normalization (scaling to the range 0-1).

[0519] Output: Preprocessed video data.

[0520] Step 3: Analyzing the video data

[0521] The server inputs the preprocessed video data into a generative AI model to analyze passenger numbers and placements.

[0522] Input: Preprocessed video data.

[0523] Specific operation: Uses a generative AI model (e.g., YOLOv5) to detect the number of passengers in an elevator from video data and analyze their placement.

[0524] Output: Number of passengers, location information and congestion level (e.g. "85").

[0525] Step 4: Determine crowding level

[0526] The server verifies the congestion level values ​​obtained from the generative AI model.

[0527] Input: The congestion level output from the generative AI model.

[0528] Specific behavior: Compare the congestion level value with a pre-set threshold (e.g. "70").

[0529] Output: Boolean value (e.g. true) indicating whether the congestion threshold is exceeded.

[0530] Step 5: Determine the stops

[0531] The server obtains the button status inside the elevator and checks and updates the current floor stop list.

[0532] Input: Button status inside the elevator, congestion level judgment result.

[0533] Specific behavior: Get the current floor stop list and remove floors where passengers in the elevator do not plan to get off (e.g., floors 3 and 7) from the floor stop list.

[0534] Output: Updated floor list (e.g. only floor 5).

[0535] Step 6: Elevator operation control

[0536] The server controls elevator operation based on the updated stop list and notifies passengers of the next stop.

[0537] Input: Updated floor stop list.

[0538] What it does: Sends instructions to the elevator control system to stop only at the specified floor, and notifies passengers of the next stop via a visual display and voice announcement (e.g., "Next stop is floor 5").

[0539] Output: Elevator control signal, next floor notification information.

[0540] Through the above process, elevator operation efficiency can be optimized, congestion can be reduced, and energy can be saved.

[0541] (Application example 1)

[0542] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0543] In conventional factory robot movement, it has been difficult to grasp congestion conditions in real time and select efficient movement routes. This has resulted in reduced productivity and hindered efficient work. In particular, in areas of the factory where congestion frequently occurs, robot movement is hindered, resulting in a significant drop in production speed and work efficiency.

[0544] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0545] In this invention, the server includes a means for capturing images of the area in real time using a camera, a means for inputting the captured image data into a generative AI model to quantify the degree of congestion in the area, and a means for optimizing the robot's movement route when the numerical value of the degree of congestion exceeds a preset threshold. This makes it possible to analyze the congestion situation in real time and instruct the robot on the optimal movement route.

[0546] A "camera" is a device that captures video within a specified area in real time.

[0547] A "generative AI model" is an artificial intelligence model used to quantify the degree of congestion based on acquired video data.

[0548] The "degree of congestion" is a numerical value that indicates the degree of congestion calculated from the number of people and their locations within a specific area.

[0549] A "threshold" is a numerical value that serves as a reference for taking a specific action when the congestion level exceeds this threshold.

[0550] "Stop floor" means a floor where an elevator stops to allow passengers to board or alight.

[0551] A "movement route" is a path chosen by a robot or other moving object when it moves within an area.

[0552] "Notification" is the act of informing a robot or worker of the next action to be taken or the selected route.

[0553] "Preprocessing" refers to the process of processing the video data, such as removing noise and adjusting resolution, before inputting it into the generative AI model.

[0554] System Configuration

[0555] This invention is implemented with the following configuration: Cameras are installed in a factory area, and a server collects real-time video data captured by the cameras. The server is a system that preprocesses this video data and inputs it into a generative AI model to quantify the congestion level of the area. If the congestion level exceeds a set threshold, the server optimizes the robot's movement route and changes the route as necessary. It then notifies the robot or worker of this information.

[0556] Overview of program processing flow

[0557] 1. Camera image acquisition

[0558] The server collects real-time video data from cameras installed in the area, which is used to understand the congestion situation within the area.

[0559] 2. Preprocessing of video data

[0560] The server performs preprocessing such as noise removal and resolution adjustment on the acquired video data before inputting it into the generative AI model.

[0561] 3. Analysis of video data

[0562] The server inputs the preprocessed video data into a generative AI model to analyze the number of people and their locations within the area, which then quantifies the level of congestion.

[0563] 4. Determining congestion

[0564] The server verifies the congestion level value obtained from the generative AI model and determines whether the congestion level exceeds a set threshold.

[0565] 5. Deciding on a travel route

[0566] The server calculates the optimal route for the robot based on the congestion information, and selects a route that avoids congestion.

[0567] 6. Notification

[0568] The server notifies the robots and workers of the selected movement route, enabling them to act quickly and efficiently.

[0569] Hardware and software used

[0570] Hardware:

[0571] Camera: A device installed in a factory area that captures images within the area in real time.

[0572] Server: A computer for processing and analyzing video data.

[0573] software:

[0574] OpenCV: A library for acquiring and preprocessing camera footage.

[0575] Keras: A framework for loading and analyzing generative AI models.

[0576] Specific examples

[0577] When a specific area in the factory is extremely congested, the server receives real-time video data from the camera, preprocesses it, and then inputs it into the generative AI model. For example, if the area is highly congested, the server instructs the robot to select a different route. If the area is less congested, the robot will select the normal route to maximize movement efficiency.

[0578] Prompt Sentence Examples

[0579] "Please obtain video data from camera ID: {camera_id}, analyze the congestion level, and calculate the optimal travel route."

[0580] In this way, the operating efficiency of robots in the factory can be optimized, congestion can be avoided, and productivity can be improved.

[0581] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[0582] Step 1:

[0583] The server acquires video data in real time from cameras installed in the area. Specifically, it captures the video from the cameras and sends the video data to the server. The input data is the video stream from the cameras, and the output data is the video frames stored on the server.

[0584] Step 2:

[0585] The server preprocesses the captured video data, specifically using OpenCV to remove noise and adjust resolution. The input data is raw video frames, and the output data is video data converted into a format suitable for the generative AI model.

[0586] Step 3:

[0587] The server inputs the preprocessed video data into a generative AI model to analyze the number of people and their locations within the area. Specifically, it uses Keras to analyze the video data and quantify the degree of congestion. The input data is the preprocessed video data, and the output data is a numerical value indicating the degree of congestion.

[0588] Step 4:

[0589] The server verifies the congestion level value obtained from the generative AI model and determines whether the congestion level exceeds a set threshold. The input data is the congestion level value, and the output data is a boolean value indicating whether the congestion level exceeds the threshold.

[0590] Step 5:

[0591] The server calculates the optimal route for the robot based on the congestion information. Specifically, it executes an algorithm to select the least congested route. The input data is the congestion level value and default route information, and the output data is the optimized route.

[0592] Step 6:

[0593] The server notifies the robot and the worker of the selected movement route. Specifically, it sends the notification via the network to the robot's operation system or the worker's terminal. The input data is the optimized movement route, and the output data is the notification information.

[0594] Through the above processing steps, it is expected that the operating efficiency of robots in the factory will be optimized and productivity will be improved by avoiding congestion.

[0595] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[0596] System Configuration

[0597] The present invention is a system that combines a camera installed inside an elevator with an emotion engine. The camera captures video footage inside the elevator in real time, preprocesses this video data, and inputs it into a generative AI model. The generative AI model analyzes the number and placement of passengers and quantifies the level of congestion. If this level of congestion exceeds a set threshold, the server uses the data from the emotion engine to determine the elevator's stopping floor. The emotion engine uses the camera's video data to recognize passenger emotions and, in conjunction with the level of congestion, determines the optimal stopping floor.

[0598] Overview of program processing flow

[0599] 1. Camera image acquisition

[0600] The server collects real-time video data from cameras installed inside the elevator, which serves as the basis for understanding the situation and emotions of passengers inside the elevator.

[0601] 2. Preprocessing of video data

[0602] The server pre-processes the captured video data, which includes noise reduction and resolution adjustment.

[0603] 3. Analysis of video data

[0604] The server inputs the preprocessed video data into a generative AI model, which analyzes the number and placement of passengers and quantifies the level of congestion.

[0605] 4. Emotional Recognition

[0606] The server also inputs the same video data into an emotion engine to recognize each passenger's emotions. The emotion engine uses facial recognition and facial expression analysis technology to detect passengers' feelings such as discomfort or stress.

[0607] 5. Integrating crowding and emotion data

[0608] The server integrates congestion data from the generative AI model with emotion data from the emotion engine, and optimizes elevator operation based on the integrated data.

[0609] 6. Determining the stopping floor

[0610] The server updates the elevator's floor stop list based on the integrated data, and adjusts the elevator to avoid stopping at floors where passengers feel uncomfortable, especially when the congestion level exceeds a threshold and discomfort is detected by the emotion engine.

[0611] 7. Elevator operation control

[0612] The server controls elevator operation based on the new stop list to avoid unnecessary stops, and uses displays and audio announcements inside the elevator to notify passengers of the next stop.

[0613] Specific examples

[0614] 1. Example in a shopping mall

[0615] In an elevator in a shopping mall, a server captures video data from a camera every second, including not only the number and location of passengers but also their facial expressions.

[0616] The server preprocesses this video data and then inputs it into the generative AI model and emotion engine. The generative AI model detects that there are 30 passengers and calculates the congestion level as 85. Meanwhile, the emotion engine recognizes that 10 passengers are feeling uncomfortable.

[0617] The server determines that the elevator is crowded because the congestion level of "85" exceeds the threshold of "70." Furthermore, taking into account the discomfort data, the server assesses that continuing normal operation may increase passenger discomfort.

[0618] An elevator that was scheduled to stop on floors 3, 5, and 7 will check to make sure there are no passengers on those floors, and update its operation instructions to stop only on floor 5, reflecting congestion and sentiment data, to most effectively reduce passenger discomfort.

[0619] The server sends the generated new operation instructions to the elevator control system, and passengers are notified that the next stop is floor 5 via a display and voice guidance inside the elevator.

[0620] This embodiment not only improves the efficiency of elevator operation but also enhances passenger convenience. It not only avoids unnecessary stops in crowded situations but also provides a comfortable user environment by taking passengers' emotions into consideration.

[0621] The processing flow will be explained below.

[0622] Step 1:

[0623] The server acquires real-time video data from a camera installed inside the elevator, which includes information on the number and location of passengers in the elevator, as well as emotional information such as facial expressions.

[0624] Step 2:

[0625] The server stores the acquired video data in temporary storage and begins preprocessing, which includes noise reduction and resolution adjustment, to prepare it for input to the generative AI model and emotion engine.

[0626] Step 3:

[0627] The server inputs the preprocessed video data into a generative AI model, which analyzes the number and location of passengers from the video data to quantify the level of congestion.

[0628] Step 4:

[0629] At the same time, the server inputs the pre-processed video data into the emotion engine, which uses facial recognition and facial expression analysis technologies to recognize each passenger's emotions (discomfort, stress, etc.).

[0630] Step 5:

[0631] As a result of the analysis, the generative AI model detects that there are 30 passengers in the elevator and calculates the congestion level as "85." This data is received by the server.

[0632] Step 6:

[0633] As a result of the analysis, the emotion engine recognizes that 10 passengers are feeling uncomfortable. This data is received by the server.

[0634] Step 7:

[0635] The server combines the congestion data obtained from the generative AI model with the emotion data obtained from the emotion engine. Since the congestion level is "85" and there are many passengers feeling uncomfortable, it determines that the elevator is crowded.

[0636] Step 8:

[0637] The server gets the elevator's current list of stops, say it's scheduled to stop at floors 3, 5, and 7.

[0638] Step 9:

[0639] The server analyzes the list of stops and identifies floors where the passengers in the elevator do not plan to get off. It confirms that the passengers do not plan to get off on the third and seventh floors.

[0640] Step 10:

[0641] Taking into account the congestion level and emotion data, the server updates the stop floor list to skip stops on floors 3 and 7 and stop only on floor 5 in order to most effectively reduce passenger discomfort.

[0642] Step 11:

[0643] The server generates new operation instructions based on the updated floor stop list and sends the instructions to the elevator control system to control elevator operation.

[0644] Step 12:

[0645] When the elevator is in operation, the server sends information about the next floor to the display device and voice guidance system inside the elevator, notifying passengers of the next floor. Specifically, the server informs passengers that the next floor is floor 5.

[0646] As a concrete example, suppose an elevator in a shopping mall is used by several passengers during the morning peak, with the camera detecting 30 passengers and the emotion engine recognizing the discomfort of 10 of them. Based on this data, the server calculates a congestion level of 85, exceeding the preset threshold of 70, and controls the elevator so that it stops only on the fifth floor, since no passengers are getting off on the third or seventh floor. Passengers are notified by display and audio guidance that the elevator will next stop on the fifth floor, ensuring efficient and comfortable operation.

[0647] Example 2

[0648] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0649] Conventional elevator operation systems lacked the means to achieve efficient operation, and were unable to determine the optimal stopping floor, especially during crowded times or based on passenger emotions. Furthermore, when passengers felt uncomfortable inside the elevator, conventional systems had difficulty responding appropriately. This sometimes led to lower satisfaction among elevator users.

[0650] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[0651] In this invention, the server includes means for capturing images of the inside of the elevator with a camera in real time, means for preprocessing the captured image data, means for inputting the preprocessed image data into a generative AI model and quantifying the degree of congestion inside the elevator, means for inputting the same image data into an emotion engine and recognizing passenger emotions, and means for integrating the congestion data obtained from the generative AI model and the emotion data obtained from the emotion engine and determining the floors at which the elevator should stop. This enables optimal operation that simultaneously takes into account the congestion status inside the elevator and the emotions of passengers.

[0652] The "camera" is a photographic device that captures images inside the elevator in real time.

[0653] "Video data" is digital data captured by a camera showing the situation inside the elevator.

[0654] "Preprocessing" refers to processes such as noise removal and resolution adjustment that are performed to improve the quality of acquired video data.

[0655] The "generative AI model" is an artificial intelligence model that analyzes video data and quantifies the degree of congestion inside an elevator.

[0656] "Crowding level" is a number that indicates the degree of congestion calculated by the generative AI model based on the number and placement of passengers in the elevator.

[0657] The "emotion engine" is an engine that analyzes video data and recognizes emotions from passengers' facial expressions.

[0658] "Integrated data" refers to data that integrates congestion data obtained from the generative AI model and emotion data obtained from the emotion engine.

[0659] "Stop floor" refers to the floor at which the elevator is scheduled to stop.

[0660] This invention is a system that combines a camera installed inside the elevator, a generative AI model, and an emotion engine, and simultaneously analyzes the congestion situation inside the elevator and the emotions of passengers to achieve optimal operation. Below, we will explain the hardware and software, as well as the specific processing and calculations.

[0661] Hardware

[0662] 1. Camera

[0663] These are IP cameras or surveillance cameras installed inside elevators. These cameras capture images from inside the elevator in real time and send them to a streaming server.

[0664] 2. Server

[0665] This is the central processing unit of the system, and performs preprocessing and analysis of video data. The server is equipped with a high-performance CPU and GPU, and runs deep learning frameworks such as Python, TensorFlow, and PyTorch.

[0666] software

[0667] 1. Preprocessing software

[0668] This software tool removes noise from video data, adjusts resolution, and adjusts contrast, etc. This improves the accuracy of analysis.

[0669] 2. Generative AI Models

[0670] This is an artificial intelligence model that analyzes the number and placement of passengers in an elevator from video data and quantifies the degree of congestion. It is usually implemented using Python, TensorFlow, and PyTorch.

[0671] 3. Emotion Engine

[0672] This is a software engine for recognizing passenger emotions from video data. It uses face recognition and facial expression analysis technologies and uses Python, OpenCV, and the Dlib library.

[0673] Example of a system

[0674] Elevator operation in shopping malls

[0675] The server collects video data every second from cameras installed in elevators in the shopping mall, including the number of passengers, their locations, and their facial expressions.

[0676] 1. Pretreatment

[0677] The server pre-processes this video data, removing noise and adjusting the resolution, for example, if the initial resolution is 720p, it will change it to 1080p to improve the quality of the data.

[0678] 2. Input to the generative AI model and emotion engine

[0679] The preprocessed data is input into the generative AI model and the emotion engine. The generative AI model detects that there are 30 passengers in the elevator and quantifies the congestion level as "85." Meanwhile, the emotion engine recognizes that 10 of the passengers are feeling uncomfortable.

[0680] 3. Integration and Operation Decisions

[0681] The server combines the congestion data obtained from the generative AI model with the emotion data obtained from the emotion engine. Based on this combined data, it determines the floors where the elevator should stop. For example, an elevator that is normally scheduled to stop on the third, fifth, and seventh floors will update its operation instructions to take into account the combined data and avoid stopping on the third and seventh floors, where there are no passengers, and stop only on the fifth floor.

[0682] Example prompt

[0683] For example, the following prompts might be used by a generative AI model:

[0684] "Analyze the number and location of passengers from video data inside the elevator, and quantify the congestion level on a scale of 0 to 100."

[0685] This invention not only improves the efficiency of elevator operation but also enhances passenger convenience. It not only avoids unnecessary stops in crowded situations but also provides a comfortable user environment by taking passengers' emotions into consideration.

[0686] The flow of the identification process in the second embodiment will be described with reference to FIG.

[0687] Program processing flow

[0688] The subject should be either the server, the terminal, or the user, and the explanation should be written in the plain voice (da / dearu style).

[0689] Step 1: Acquire camera footage

[0690] The server receives real-time video data from the camera. The camera captures video footage inside the elevator and streams it to the server. This video data includes the number and placement of passengers, facial expressions, etc.

[0691] Input: Real-time video from inside the elevator

[0692] Output: Captured video data

[0693] Specific behavior:

[0694] 1. The camera captures a frame every second and sends it to the streaming server.

[0695] 2. The server receives this stream and stores it in a database.

[0696] Step 2: Preprocessing the video data

[0697] The server pre-processes the captured video data, which includes noise reduction, resolution adjustment, and contrast enhancement.

[0698] Input: Captured video data

[0699] Output: Pre-processed video data

[0700] Specific behavior:

[0701] 1. A noise reduction filter is applied to the video data.

[0702] 2. Adjust the video resolution (e.g. convert from 720p to 1080p).

[0703] 3. Adjust contrast and brightness and clean abnormal pixels.

[0704] Step 3: Analyzing the video data

[0705] The server inputs the preprocessed video data into a generative AI model, which analyzes the number and location of passengers from the video data and quantifies the level of congestion.

[0706] Input: Preprocessed video data

[0707] Output: Numerical data of congestion

[0708] Specific behavior:

[0709] 1. Preprocessed video data is passed to a generative AI model.

[0710] 2. The model extracts facial and body features of each passenger and identifies their number and location.

[0711] 3. Based on the number and location of passengers, the congestion level is quantified on a scale of 0 to 100.

[0712] Step 4: Recognize emotions

[0713] The server then inputs the same video data into an emotion engine to recognize passenger emotions, which uses facial recognition and facial expression analysis to classify emotions.

[0714] Input: Preprocessed video data

[0715] Output: Passenger emotion data

[0716] Specific behavior:

[0717] 1. Detect the faces of each passenger from the preprocessed video data.

[0718] 2. Use a facial recognition algorithm to analyze each passenger's facial expression.

[0719] 3. Classify emotions such as discomfort, stress, and joy, and save the results as numerical data.

[0720] Step 5: Integrating crowding and emotion data

[0721] The server combines the congestion data obtained from the generative AI model with the emotion data obtained from the emotion engine, and this data is used to control elevator operation.

[0722] Input: Numerical data on congestion levels, passenger sentiment data

[0723] Output: Integrated traffic control data

[0724] Specific behavior:

[0725] 1. Obtain numerical data on congestion levels and emotional data.

[0726] 2. Combine both sets of data to create a detailed picture of elevator operation.

[0727] Step 6: Determine the stops

[0728] The server determines the elevator's stopping floor based on the integrated data, taking into account the level of congestion and passenger sentiment to select the optimal stopping floor.

[0729] Input: Integrated traffic control data

[0730] Output: Updated floor stop list

[0731] Specific behavior:

[0732] 1. Analyze the integrated data to assess congestion and emotions.

[0733] 2. If the threshold is exceeded, floors with fewer passengers or less discomfort will be selected preferentially.

[0734] 3. Create an updated list of stops.

[0735] Step 7: Elevator operation control

[0736] The server controls elevator operation based on the new stop list, avoiding unnecessary stops and notifying passengers of the operation status.

[0737] Input: Updated floor list

[0738] Output: Operation control instructions and passenger notifications

[0739] Specific behavior:

[0740] 1. Coordinate elevator operation based on a list of stops.

[0741] 2. Use displays and voice guidance systems inside the elevator to inform passengers of the next floor (e.g., announce "Next floor is floor 5").

[0742] (Application example 2)

[0743] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0744] Conventional elevator systems only analyzed congestion levels and did not take passengers' emotions into consideration. As a result, even when congestion levels were high, it was difficult to alleviate passenger discomfort, potentially resulting in a poor user experience. Furthermore, it was difficult to properly control elevator operation to avoid unnecessary elevator stops, hindering efficient operation.

[0745] The identification processing by the identification processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for capturing video images inside the elevator with a camera in real time, means for inputting the captured video data into a generative AI model and quantifying the congestion level inside the elevator, means for determining the elevator's stopping floor when the congestion level numerical value exceeds a preset threshold, means for inputting the video data into an emotion engine to recognize passenger emotions, means for integrating the emotion score recognized by the emotion engine with the congestion level data, and means for optimizing the elevator's stopping floors based on the integrated data. This enables efficient and comfortable elevator operation that takes into account not only the number of passengers but also the emotions of the passengers.

[0746] The "camera" is a device that captures images inside the elevator in real time.

[0747] "Video data" refers to video information inside the elevator captured by a camera.

[0748] The "generative AI model" is an artificial intelligence algorithm that analyzes the acquired video data and quantifies the degree of congestion inside the elevator.

[0749] The "crowding level" is a numerical value that indicates the degree of congestion calculated based on the number and arrangement of passengers in the elevator.

[0750] A "threshold" is a set value that, when exceeded, triggers a specific action (for example, determining the elevator's stopping floor).

[0751] The "emotion engine" is an algorithm that analyzes video data and recognizes passenger emotions.

[0752] The "emotion score" is a numerical representation of the passenger's emotional state as recognized by the emotion engine.

[0753] "Integrated data" is a dataset that integrates congestion data and emotion scores.

[0754] "Means for determining elevator stopping floors" refers to algorithms or devices that determine the optimal elevator stopping floor based on congestion and emotion score data.

[0755] This invention is a system for optimizing elevator operations in brick-and-mortar stores and shopping malls. The system consists of a camera, a server, a generative AI model, an emotion engine, and an elevator control system.

[0756] System Configuration

[0757] 1. Camera:

[0758] The cameras are installed inside the elevator and capture video data in real time, including the number, location, and facial expressions of passengers.

[0759] 2. Server:

[0760] The server preprocesses the acquired video data, which includes noise reduction and resolution adjustment. The preprocessed video data is then input into the generative AI model and emotion engine.

[0761] 3. Generative AI Model:

[0762] The generative AI model analyzes the pre-processed video data and quantifies the congestion level based on the number and location of passengers. For example, if there are 30 passengers, the congestion level is calculated as 85.

[0763] 4. Emotion Engine:

[0764] The emotion engine analyzes pre-processed video data to recognize passenger emotions. It uses facial recognition and facial expression analysis technology to detect emotions such as discomfort and stress. For example, it can recognize that 10 passengers are feeling discomfort.

[0765] 5. Elevator control system:

[0766] The server combines the congestion data from the generative AI model with the emotion data from the emotion engine. It then optimizes the elevator's floor stops based on the combined data. For example, if the congestion level of 85 exceeds the threshold of 70, it takes into account the discomfort data and determines that continuing normal operation will increase passenger discomfort.

[0767] Elevator stop lists will be updated and operational instructions will be updated to stop only on the first and fourth floors to minimize congestion.

[0768] Specific examples

[0769] In an elevator in a shopping mall, a server collects video data from a camera every second. This data includes not only the number and location of passengers, but also their facial expressions. The server preprocesses this video data before inputting it into a generative AI model and an emotion engine. For example, the generative AI model detects that there are 30 passengers and calculates the congestion level as "85." Meanwhile, the emotion engine recognizes that 10 passengers are feeling uncomfortable.

[0770] The server determines that the elevator is crowded because the congestion level of 85 exceeds the threshold of 70. Furthermore, taking into account discomfort data, the server assesses that continuing normal operation may increase passenger discomfort. The elevator, which was scheduled to stop on the third, fifth, and seventh floors, updates its operation instructions to stop only on the first and fourth floors, reflecting the congestion level and emotional data to most effectively reduce passenger discomfort.

[0771] Prompt Sentence Examples

[0772] Number of passengers in the elevator: 30

[0773] Passenger Sentiment Score:

[0774] 10 people found it "unpleasant"

[0775] 15 people answered "normal",

[0776] 5 people said "comfortable",

[0777] Crowding level: 85

[0778] Threshold: 70

[0779] In this way, the server can realize efficient and comfortable elevator operation that takes into account not only the number of passengers but also their emotions.

[0780] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[0781] Step 1:

[0782] The server acquires video data in real time from a camera installed inside the elevator. The input is the raw video from the camera, and the output is the acquired raw video data. This video data includes the number, location, and facial expressions of passengers inside the elevator.

[0783] Step 2:

[0784] The server preprocesses the acquired video data. The input is the raw video data acquired in step 1, and the output is preprocessed video data with noise removal and resolution adjustment. Specific operations include applying a noise removal filter and resizing the resolution to a fixed size.

[0785] Step 3:

[0786] The server inputs the preprocessed video data into the generative AI model. The input is the preprocessed video data, and the output is data that quantifies the degree of congestion inside the elevator. Specifically, the generative AI model analyzes the number and placement of passengers and calculates the degree of congestion as a number (e.g., "85").

[0787] Step 4:

[0788] The server also inputs the preprocessed video data into the emotion engine. The input is the preprocessed video data generated in step 2, and the output is the emotion score for each passenger. Specifically, the emotion engine uses face recognition technology and facial expression analysis technology to calculate each passenger's emotion (e.g., "uncomfortable," "neutral," or "comfortable") as a score.

[0789] Step 5:

[0790] The server integrates the congestion data from the generative AI model and the emotion data from the emotion engine. The input is the congestion data and emotion score, and the output is an integrated dataset. Specifically, it combines each data into a single dataset to comprehensively evaluate the elevator's condition.

[0791] Step 6:

[0792] The server optimizes elevator stops based on the integrated data. The input is the integrated dataset, and the output is an optimized elevator stop list. Specifically, if the congestion level exceeds a threshold and discomfort is high, the server adjusts the elevator stops to reduce discomfort (e.g., stopping only on the first and fourth floors).

[0793] Step 7:

[0794] The server sends the optimized elevator stop list to the elevator control system. The input is the optimized stop list, and the output is the operation instructions to be executed by the elevator control system. Specifically, the server sends new stop information to the control system to control the actual operation of the elevator.

[0795] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0796] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0797] In the above embodiment, an example in which the specific processing is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the smart glasses 214.

[0798] [Third embodiment]

[0799] FIG. 5 shows an example of the configuration of a data processing system 310 according to the third embodiment.

[0800] 5, the data processing system 310 includes the data processing device 12 and a headset type terminal 314. An example of the data processing device 12 is a server.

[0801] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0802] The headset type terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a display 343. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the display 343 are also connected to the bus 52.

[0803] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[0804] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[0805] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[0806] Fig. 6 shows an example of the main functions of the data processing device 12 and the headset type terminal 314. As shown in Fig. 6, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[0807] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0808] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0809] In the headset type terminal 314, a reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0810] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the headset type terminal 314 will be referred to as the "terminal."

[0811] System Configuration

[0812] The present invention is implemented in the following configuration. A camera is installed inside the elevator, and a server collects real-time video data captured by the camera. The server preprocesses the video data and inputs it into a generative AI model to quantify the elevator congestion level. If the congestion level exceeds a set threshold, the server determines the floor at which the elevator will stop, and omits stops if necessary.

[0813] Overview of program processing flow

[0814] 1. Camera image acquisition

[0815] The server receives real-time video data from cameras installed inside the elevator, which is used to understand the situation of passengers inside the elevator.

[0816] 2. Preprocessing of video data

[0817] The server performs preprocessing such as noise removal and resolution adjustment on the acquired video data before inputting it into the generative AI model.

[0818] 3. Analysis of video data

[0819] The server inputs the preprocessed video data into a generative AI model to analyze the number and location of passengers, which then quantifies the degree of congestion inside the elevator.

[0820] 4. Determining congestion

[0821] The server verifies the congestion level value obtained from the generative AI model and determines whether the congestion level exceeds a set threshold.

[0822] 5. Determining the stopping floor

[0823] The server obtains the current floor stop list based on the buttons inside the elevator, and then updates the floor stop list so that the elevator does not stop at floors where passengers inside the elevator do not plan to get off.

[0824] 6. Elevator operation control

[0825] The server controls elevator operation based on the updated floor stop list to avoid unnecessary stops during peak hours, and notifies passengers of the next floor stop via a display and voice guidance inside the elevator.

[0826] Specific examples

[0827] 1. Example in a shopping mall

[0828] In an elevator in a shopping mall, the server acquires video data from the camera every second. For example, suppose the elevator is very crowded at a certain time.

[0829] The server preprocesses this video data and inputs it into a generative AI model, which detects that there are 30 passengers in the elevator and quantifies the congestion level as "85."

[0830] Since the congestion level exceeds the set threshold "70", the server determines that the elevator is crowded.

[0831] Next, the server checks the button status inside the elevator and confirms that although the elevator was scheduled to stop on the 3rd, 5th, and 7th floors, there are no passengers getting off on the 3rd and 7th floors.

[0832] The server updates the floor stop list and adjusts it to skip floors 3 and 7.

[0833] Ultimately, the elevator will stop only on the fifth floor, and a display and voice guidance inside the elevator will inform passengers that the next floor is floor 5. In this way, unnecessary stops can be avoided and the elevator can be operated efficiently.

[0834] This invention optimizes the operation of crowded elevators, reducing wasted time for passengers and reducing the amount of time that passengers waste waiting when they cannot board, thereby improving the overall efficiency of elevator use.

[0835] The processing flow will be explained below.

[0836] Step 1:

[0837] The server receives real-time video data from a camera installed inside the elevator, capturing the entire elevator interior and allowing the system to understand the situation of passengers inside the elevator.

[0838] Step 2:

[0839] The server temporarily stores the acquired video data and begins preprocessing, which includes removing noise and adjusting the resolution of the video, preparing it for input to the generative AI model.

[0840] Step 3:

[0841] The server inputs the preprocessed video data into a generative AI model, which then detects the number of passengers from the video data and analyzes their placement and movements.

[0842] Step 4:

[0843] The generative AI model quantifies the congestion level based on the number and location of passengers. The server receives this quantified congestion level and proceeds to the next step.

[0844] Step 5:

[0845] The server compares the quantified congestion level with a preset threshold. If the congestion level exceeds the threshold, the elevator is determined to be congested, but if it does not exceed the threshold, it continues to operate normally.

[0846] Step 6:

[0847] If the elevator is determined to be busy, the server retrieves the current floor stop list, which is based on floor button presses inside the elevator.

[0848] Step 7:

[0849] The server analyzes the stop floor list to identify floors where passengers in the elevator plan to get off, extracts floors where passengers do not plan to get off, and updates the stop floor list so that the elevator does not stop at those floors.

[0850] Step 8:

[0851] The server generates new operation instructions based on the updated floor list, including the floor where the elevator will next stop.

[0852] Step 9:

[0853] The server sends the generated operation instructions to the elevator control system, which operates the elevator based on these instructions and avoids unnecessary stops.

[0854] Step 10:

[0855] The server sends the next floor information to the display device and voice guidance device inside the elevator, explicitly notifying passengers of the next floor, thus ensuring that passengers inside the elevator are properly informed.

[0856] As a concrete example, elevators in a shopping mall tend to be crowded during the morning peak hours. For example, the server inputs video data acquired from a camera into a generative AI model to calculate a congestion level of "85," which it determines is above the threshold of "70." As a result, it determines that there are no passengers on the third and seventh floors, and generates operation instructions to skip stopping at those floors. The elevator then stops only on the fifth floor, allowing for efficient operation.

[0857] Example 1

[0858] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[0859] Conventional elevator systems lacked the means to properly manage passenger congestion and optimize operational efficiency. This resulted in unnecessary elevator stops and congestion that prevented passengers from moving smoothly. This also led to problems such as wasted energy and increased passenger stress.

[0860] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0861] In this invention, the server includes means for capturing images of the inside of the elevator with a camera in real time, means for preprocessing the captured image data, means for inputting the preprocessed image data into a generative AI model to quantify the congestion level inside the elevator, means for determining the elevator's stopping floor when the congestion level exceeds a preset threshold, and means for notifying the next stopping floor via a display device or voice guidance inside the elevator. This improves elevator operation efficiency, reduces unnecessary waiting time for passengers, and contributes to energy savings.

[0862] The "camera" is a device for capturing images inside the elevator in real time.

[0863] "Video data" refers to real-time image information of the inside of an elevator captured by a camera.

[0864] "Preprocessing" refers to the process of removing noise, adjusting resolution, normalizing, etc. from video data, and converting it into a format suitable for the generative AI model.

[0865] The "generative AI model" is a machine learning model that uses a deep learning algorithm to analyze and quantify the number of passengers and congestion levels from input video data.

[0866] The "crowding level" is a numerical value that indicates the degree of congestion calculated based on the number and arrangement of passengers in the elevator.

[0867] A "threshold" is a reference value set to trigger a specific action when the congestion level exceeds this threshold.

[0868] "Stop floor" means a floor at which an elevator stops during a particular operation.

[0869] A "display device" is a device installed inside an elevator that visually provides passengers with information such as the next floor.

[0870] "Voice guidance" is a system that notifies passengers of the next floor and other information by voice inside the elevator.

[0871] The present invention is a system that analyzes the congestion level in an elevator in real time and optimizes operation efficiency. This system is implemented with the following configuration.

[0872] System Configuration

[0873] The system consists of a camera installed inside the elevator, a server, a generative AI model, and an elevator control system.

[0874] Hardware and Software

[0875] camera:

[0876] Cameras will be installed inside the elevators to capture high-resolution video in real time, providing images at 30 frames per second.

[0877] server:

[0878] The server is a central processing unit that receives video data sent from the camera and performs preprocessing and analysis. It is equipped with a high-performance CPU and GPU and has sufficient computing resources to run generative AI models.

[0879] Generative AI models:

[0880] This is a machine learning model that analyzes video data and quantifies the degree of congestion in an elevator. It uses a model that uses a deep learning algorithm (such as YOLOv5).

[0881] Elevator Control System:

[0882] The system controls elevator operation based on control signals sent from the server, and notifies passengers of the next floor via a display device and voice guidance system inside the elevator.

[0883] Example of operation

[0884] 1. Elevators in shopping malls:

[0885] The server acquires video data from the camera inside the elevator every second.

[0886] The server first pre-processes the acquired video data, which includes noise removal, resolution adjustment, and normalization.

[0887] The preprocessed video data is input into a generative AI model (e.g., YOLOv5) to analyze the number and location of passengers in the elevator. The generative AI model detects that there are 30 passengers in the elevator and quantifies the congestion level as "85."

[0888] The server checks whether the congestion level exceeds the set threshold value "70".

[0889] The server checks the button status inside the elevator and confirms that although the elevator was scheduled to stop on the 3rd, 5th, and 7th floors, there are no passengers getting off on the 3rd and 7th floors.

[0890] The server updates the stop floor list and adjusts it to skip floors 3 and 7.

[0891] The elevator will finally stop on the fifth floor, and a display and voice guidance inside the elevator will inform passengers that the next floor is the fifth floor.

[0892] Example prompts for generative AI models

[0893] "Please use the camera footage from inside the elevator to quantify the current congestion level. The video data shows 30 passengers. The congestion level is 85."

[0894] The system of the present invention optimizes elevator operation efficiency, reduces passenger waiting times, cuts energy waste, and ensures that passengers are transported to their destinations quickly and efficiently, even during peak hours.

[0895] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0896] Step 1: Acquire camera footage

[0897] The server acquires video data in real time from cameras installed inside the elevator.

[0898] Input: Real-time video data sent from the camera.

[0899] Specific operation: Video data is sent to the server as stream data at 30 frames per second.

[0900] Output: Captured real-time video data.

[0901] Step 2: Preprocessing the video data

[0902] The server preprocesses the acquired video data before inputting it into the generative AI model.

[0903] Input: Video data acquired from the camera.

[0904] Specific operations: noise removal, resolution adjustment (e.g. convert to 640x480 pixels), pixel value normalization (scaling to the range 0-1).

[0905] Output: Preprocessed video data.

[0906] Step 3: Analyzing the video data

[0907] The server inputs the preprocessed video data into a generative AI model to analyze passenger numbers and placements.

[0908] Input: Preprocessed video data.

[0909] Specific operation: Uses a generative AI model (e.g., YOLOv5) to detect the number of passengers in an elevator from video data and analyze their placement.

[0910] Output: Number of passengers, location information and congestion level (e.g. "85").

[0911] Step 4: Determine crowding level

[0912] The server verifies the congestion level values ​​obtained from the generative AI model.

[0913] Input: The congestion level output from the generative AI model.

[0914] Specific behavior: Compare the congestion level value with a pre-set threshold (e.g. "70").

[0915] Output: Boolean value (e.g. true) indicating whether the congestion threshold is exceeded.

[0916] Step 5: Determine the stops

[0917] The server obtains the button status inside the elevator and checks and updates the current floor stop list.

[0918] Input: Button status inside the elevator, congestion level judgment result.

[0919] Specific behavior: Get the current floor stop list and remove floors where passengers in the elevator do not plan to get off (e.g., floors 3 and 7) from the floor stop list.

[0920] Output: Updated floor list (e.g. only floor 5).

[0921] Step 6: Elevator operation control

[0922] The server controls elevator operation based on the updated stop list and notifies passengers of the next stop.

[0923] Input: Updated floor stop list.

[0924] What it does: Sends instructions to the elevator control system to stop only at the specified floor, and notifies passengers of the next stop via a visual display and voice announcement (e.g., "Next stop is floor 5").

[0925] Output: Elevator control signal, next floor notification information.

[0926] Through the above process, elevator operation efficiency can be optimized, congestion can be reduced, and energy can be saved.

[0927] (Application example 1)

[0928] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[0929] In conventional factory robot movement, it has been difficult to grasp congestion conditions in real time and select efficient movement routes. This has resulted in reduced productivity and hindered efficient work. In particular, in areas of the factory where congestion frequently occurs, robot movement is hindered, resulting in a significant drop in production speed and work efficiency.

[0930] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0931] In this invention, the server includes a means for capturing images of the area in real time using a camera, a means for inputting the captured image data into a generative AI model to quantify the degree of congestion in the area, and a means for optimizing the robot's movement route when the numerical value of the degree of congestion exceeds a preset threshold. This makes it possible to analyze the congestion situation in real time and instruct the robot on the optimal movement route.

[0932] A "camera" is a device that captures video within a specified area in real time.

[0933] A "generative AI model" is an artificial intelligence model used to quantify the degree of congestion based on acquired video data.

[0934] The "degree of congestion" is a numerical value that indicates the degree of congestion calculated from the number of people and their locations within a specific area.

[0935] A "threshold" is a numerical value that serves as a reference for taking a specific action when the congestion level exceeds this threshold.

[0936] "Stop floor" means a floor where an elevator stops to allow passengers to board or alight.

[0937] A "movement route" is a path chosen by a robot or other moving object when it moves within an area.

[0938] "Notification" is the act of informing a robot or worker of the next action to be taken or the selected route.

[0939] "Preprocessing" refers to the process of processing the video data, such as removing noise and adjusting resolution, before inputting it into the generative AI model.

[0940] System Configuration

[0941] This invention is implemented with the following configuration: Cameras are installed in a factory area, and a server collects real-time video data captured by the cameras. The server is a system that preprocesses this video data and inputs it into a generative AI model to quantify the congestion level of the area. If the congestion level exceeds a set threshold, the server optimizes the robot's movement route and changes the route as necessary. It then notifies the robot or worker of this information.

[0942] Overview of program processing flow

[0943] 1. Camera image acquisition

[0944] The server collects real-time video data from cameras installed in the area, which is used to understand the congestion situation within the area.

[0945] 2. Preprocessing of video data

[0946] The server performs preprocessing such as noise removal and resolution adjustment on the acquired video data before inputting it into the generative AI model.

[0947] 3. Analysis of video data

[0948] The server inputs the preprocessed video data into a generative AI model to analyze the number of people and their locations within the area, which then quantifies the level of congestion.

[0949] 4. Determining congestion

[0950] The server verifies the congestion level value obtained from the generative AI model and determines whether the congestion level exceeds a set threshold.

[0951] 5. Deciding on a travel route

[0952] The server calculates the optimal route for the robot based on the congestion information, and selects a route that avoids congestion.

[0953] 6. Notification

[0954] The server notifies the robots and workers of the selected movement route, enabling them to act quickly and efficiently.

[0955] Hardware and software used

[0956] Hardware:

[0957] Camera: A device installed in a factory area that captures images within the area in real time.

[0958] Server: A computer for processing and analyzing video data.

[0959] software:

[0960] OpenCV: A library for acquiring and preprocessing camera footage.

[0961] Keras: A framework for loading and analyzing generative AI models.

[0962] Specific examples

[0963] When a specific area in the factory is extremely congested, the server receives real-time video data from the camera, preprocesses it, and then inputs it into the generative AI model. For example, if the area is highly congested, the server instructs the robot to select a different route. If the area is less congested, the robot will select the normal route to maximize movement efficiency.

[0964] Prompt Sentence Examples

[0965] "Please obtain video data from camera ID: {camera_id}, analyze the congestion level, and calculate the optimal travel route."

[0966] In this way, the operating efficiency of robots in the factory can be optimized, congestion can be avoided, and productivity can be improved.

[0967] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[0968] Step 1:

[0969] The server acquires video data in real time from cameras installed in the area. Specifically, it captures the video from the cameras and sends the video data to the server. The input data is the video stream from the cameras, and the output data is the video frames stored on the server.

[0970] Step 2:

[0971] The server preprocesses the captured video data, specifically using OpenCV to remove noise and adjust resolution. The input data is raw video frames, and the output data is video data converted into a format suitable for the generative AI model.

[0972] Step 3:

[0973] The server inputs the preprocessed video data into a generative AI model to analyze the number of people and their locations within the area. Specifically, it uses Keras to analyze the video data and quantify the degree of congestion. The input data is the preprocessed video data, and the output data is a numerical value indicating the degree of congestion.

[0974] Step 4:

[0975] The server verifies the congestion level value obtained from the generative AI model and determines whether the congestion level exceeds a set threshold. The input data is the congestion level value, and the output data is a boolean value indicating whether the congestion level exceeds the threshold.

[0976] Step 5:

[0977] The server calculates the optimal route for the robot based on the congestion information. Specifically, it executes an algorithm to select the least congested route. The input data is the congestion level value and default route information, and the output data is the optimized route.

[0978] Step 6:

[0979] The server notifies the robot and the worker of the selected movement route. Specifically, it sends the notification via the network to the robot's operation system or the worker's terminal. The input data is the optimized movement route, and the output data is the notification information.

[0980] Through the above processing steps, it is expected that the operating efficiency of robots in the factory will be optimized and productivity will be improved by avoiding congestion.

[0981] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[0982] System Configuration

[0983] The present invention is a system that combines a camera installed inside an elevator with an emotion engine. The camera captures video footage inside the elevator in real time, preprocesses this video data, and inputs it into a generative AI model. The generative AI model analyzes the number and placement of passengers and quantifies the level of congestion. If this level of congestion exceeds a set threshold, the server uses the data from the emotion engine to determine the elevator's stopping floor. The emotion engine uses the camera's video data to recognize passenger emotions and, in conjunction with the level of congestion, determines the optimal stopping floor.

[0984] Overview of program processing flow

[0985] 1. Camera image acquisition

[0986] The server collects real-time video data from cameras installed inside the elevator, which serves as the basis for understanding the situation and emotions of passengers inside the elevator.

[0987] 2. Preprocessing of video data

[0988] The server pre-processes the captured video data, which includes noise reduction and resolution adjustment.

[0989] 3. Analysis of video data

[0990] The server inputs the preprocessed video data into a generative AI model, which analyzes the number and placement of passengers and quantifies the level of congestion.

[0991] 4. Emotional Recognition

[0992] The server also inputs the same video data into an emotion engine to recognize each passenger's emotions. The emotion engine uses facial recognition and facial expression analysis technology to detect passengers' feelings such as discomfort or stress.

[0993] 5. Integrating crowding and emotion data

[0994] The server integrates congestion data from the generative AI model with emotion data from the emotion engine, and optimizes elevator operation based on the integrated data.

[0995] 6. Determining the stopping floor

[0996] The server updates the elevator's floor stop list based on the integrated data, and adjusts the elevator to avoid stopping at floors where passengers feel uncomfortable, especially when the congestion level exceeds a threshold and discomfort is detected by the emotion engine.

[0997] 7. Elevator operation control

[0998] The server controls elevator operation based on the new stop list to avoid unnecessary stops, and uses displays and audio announcements inside the elevator to notify passengers of the next stop.

[0999] Specific examples

[1000] 1. Example in a shopping mall

[1001] In an elevator in a shopping mall, a server captures video data from a camera every second, including not only the number and location of passengers but also their facial expressions.

[1002] The server preprocesses this video data and then inputs it into the generative AI model and emotion engine. The generative AI model detects that there are 30 passengers and calculates the congestion level as 85. Meanwhile, the emotion engine recognizes that 10 passengers are feeling uncomfortable.

[1003] The server determines that the elevator is crowded because the congestion level of "85" exceeds the threshold of "70." Furthermore, taking into account the discomfort data, the server assesses that continuing normal operation may increase passenger discomfort.

[1004] An elevator that was scheduled to stop on floors 3, 5, and 7 will check to make sure there are no passengers on those floors, and update its operation instructions to stop only on floor 5, reflecting congestion and sentiment data, to most effectively reduce passenger discomfort.

[1005] The server sends the generated new operation instructions to the elevator control system, and passengers are notified that the next stop is floor 5 via a display and voice guidance inside the elevator.

[1006] This embodiment not only improves the efficiency of elevator operation but also enhances passenger convenience. It not only avoids unnecessary stops in crowded situations but also provides a comfortable user environment by taking passengers' emotions into consideration.

[1007] The processing flow will be explained below.

[1008] Step 1:

[1009] The server acquires real-time video data from a camera installed inside the elevator, which includes information on the number and location of passengers in the elevator, as well as emotional information such as facial expressions.

[1010] Step 2:

[1011] The server stores the acquired video data in temporary storage and begins preprocessing, which includes noise reduction and resolution adjustment, to prepare it for input to the generative AI model and emotion engine.

[1012] Step 3:

[1013] The server inputs the preprocessed video data into a generative AI model, which analyzes the number and location of passengers from the video data to quantify the level of congestion.

[1014] Step 4:

[1015] At the same time, the server inputs the pre-processed video data into the emotion engine, which uses facial recognition and facial expression analysis technologies to recognize each passenger's emotions (discomfort, stress, etc.).

[1016] Step 5:

[1017] As a result of the analysis, the generative AI model detects that there are 30 passengers in the elevator and calculates the congestion level as "85." This data is received by the server.

[1018] Step 6:

[1019] As a result of the analysis, the emotion engine recognizes that 10 passengers are feeling uncomfortable. This data is received by the server.

[1020] Step 7:

[1021] The server combines the congestion data obtained from the generative AI model with the emotion data obtained from the emotion engine. Since the congestion level is "85" and there are many passengers feeling uncomfortable, it determines that the elevator is crowded.

[1022] Step 8:

[1023] The server gets the elevator's current list of stops, say it's scheduled to stop at floors 3, 5, and 7.

[1024] Step 9:

[1025] The server analyzes the list of stops and identifies floors where the passengers in the elevator do not plan to get off. It confirms that the passengers do not plan to get off on the third and seventh floors.

[1026] Step 10:

[1027] Taking into account the congestion level and emotion data, the server updates the stop floor list to skip stops on floors 3 and 7 and stop only on floor 5 in order to most effectively reduce passenger discomfort.

[1028] Step 11:

[1029] The server generates new operation instructions based on the updated floor stop list and sends the instructions to the elevator control system to control elevator operation.

[1030] Step 12:

[1031] When the elevator is in operation, the server sends information about the next floor to the display device and voice guidance system inside the elevator, notifying passengers of the next floor. Specifically, the server informs passengers that the next floor is floor 5.

[1032] As a concrete example, suppose an elevator in a shopping mall is used by several passengers during the morning peak, with the camera detecting 30 passengers and the emotion engine recognizing the discomfort of 10 of them. Based on this data, the server calculates a congestion level of 85, exceeding the preset threshold of 70, and controls the elevator so that it stops only on the fifth floor, since no passengers are getting off on the third or seventh floor. Passengers are notified by display and audio guidance that the elevator will next stop on the fifth floor, ensuring efficient and comfortable operation.

[1033] Example 2

[1034] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1035] Conventional elevator operation systems lacked the means to achieve efficient operation, and were unable to determine the optimal stopping floor, especially during crowded times or based on passenger emotions. Furthermore, when passengers felt uncomfortable inside the elevator, conventional systems had difficulty responding appropriately. This sometimes led to lower satisfaction among elevator users.

[1036] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[1037] In this invention, the server includes means for capturing images of the inside of the elevator with a camera in real time, means for preprocessing the captured image data, means for inputting the preprocessed image data into a generative AI model and quantifying the degree of congestion inside the elevator, means for inputting the same image data into an emotion engine and recognizing passenger emotions, and means for integrating the congestion data obtained from the generative AI model and the emotion data obtained from the emotion engine and determining the floors at which the elevator should stop. This enables optimal operation that simultaneously takes into account the congestion status inside the elevator and the emotions of passengers.

[1038] The "camera" is a photographic device that captures images inside the elevator in real time.

[1039] "Video data" is digital data captured by a camera showing the situation inside the elevator.

[1040] "Preprocessing" refers to processes such as noise removal and resolution adjustment that are performed to improve the quality of acquired video data.

[1041] The "generative AI model" is an artificial intelligence model that analyzes video data and quantifies the degree of congestion inside an elevator.

[1042] "Crowding level" is a number that indicates the degree of congestion calculated by the generative AI model based on the number and placement of passengers in the elevator.

[1043] The "emotion engine" is an engine that analyzes video data and recognizes emotions from passengers' facial expressions.

[1044] "Integrated data" refers to data that integrates congestion data obtained from the generative AI model and emotion data obtained from the emotion engine.

[1045] "Stop floor" refers to the floor at which the elevator is scheduled to stop.

[1046] This invention is a system that combines a camera installed inside the elevator, a generative AI model, and an emotion engine, and simultaneously analyzes the congestion situation inside the elevator and the emotions of passengers to achieve optimal operation. Below, we will explain the hardware and software, as well as the specific processing and calculations.

[1047] Hardware

[1048] 1. Camera

[1049] These are IP cameras or surveillance cameras installed inside elevators. These cameras capture images from inside the elevator in real time and send them to a streaming server.

[1050] 2. Server

[1051] This is the central processing unit of the system, and performs preprocessing and analysis of video data. The server is equipped with a high-performance CPU and GPU, and runs deep learning frameworks such as Python, TensorFlow, and PyTorch.

[1052] software

[1053] 1. Preprocessing software

[1054] This software tool removes noise from video data, adjusts resolution, and adjusts contrast, etc. This improves the accuracy of analysis.

[1055] 2. Generative AI Models

[1056] This is an artificial intelligence model that analyzes the number and placement of passengers in an elevator from video data and quantifies the degree of congestion. It is usually implemented using Python, TensorFlow, and PyTorch.

[1057] 3. Emotion Engine

[1058] This is a software engine for recognizing passenger emotions from video data. It uses face recognition and facial expression analysis technologies and uses Python, OpenCV, and the Dlib library.

[1059] Example of a system

[1060] Elevator operation in shopping malls

[1061] The server collects video data every second from cameras installed in elevators in the shopping mall, including the number of passengers, their locations, and their facial expressions.

[1062] 1. Pretreatment

[1063] The server pre-processes this video data, removing noise and adjusting the resolution, for example, if the initial resolution is 720p, it will change it to 1080p to improve the quality of the data.

[1064] 2. Input to the generative AI model and emotion engine

[1065] The preprocessed data is input into the generative AI model and the emotion engine. The generative AI model detects that there are 30 passengers in the elevator and quantifies the congestion level as "85." Meanwhile, the emotion engine recognizes that 10 of the passengers are feeling uncomfortable.

[1066] 3. Integration and Operation Decisions

[1067] The server combines the congestion data obtained from the generative AI model with the emotion data obtained from the emotion engine. Based on this combined data, it determines the floors where the elevator should stop. For example, an elevator that is normally scheduled to stop on the third, fifth, and seventh floors will update its operation instructions to take into account the combined data and avoid stopping on the third and seventh floors, where there are no passengers, and stop only on the fifth floor.

[1068] Example prompt

[1069] For example, the following prompts might be used by a generative AI model:

[1070] "Analyze the number and location of passengers from video data inside the elevator, and quantify the congestion level on a scale of 0 to 100."

[1071] This invention not only improves the efficiency of elevator operation but also enhances passenger convenience. It not only avoids unnecessary stops in crowded situations but also provides a comfortable user environment by taking passengers' emotions into consideration.

[1072] The flow of the identification process in the second embodiment will be described with reference to FIG.

[1073] Program processing flow

[1074] The subject should be either the server, the terminal, or the user, and the explanation should be written in the plain voice (da / dearu style).

[1075] Step 1: Acquire camera footage

[1076] The server receives real-time video data from the camera. The camera captures video footage inside the elevator and streams it to the server. This video data includes the number and placement of passengers, facial expressions, etc.

[1077] Input: Real-time video from inside the elevator

[1078] Output: Captured video data

[1079] Specific behavior:

[1080] 1. The camera captures a frame every second and sends it to the streaming server.

[1081] 2. The server receives this stream and stores it in a database.

[1082] Step 2: Preprocessing the video data

[1083] The server pre-processes the captured video data, which includes noise reduction, resolution adjustment, and contrast enhancement.

[1084] Input: Captured video data

[1085] Output: Pre-processed video data

[1086] Specific behavior:

[1087] 1. A noise reduction filter is applied to the video data.

[1088] 2. Adjust the video resolution (e.g. convert from 720p to 1080p).

[1089] 3. Adjust contrast and brightness and clean abnormal pixels.

[1090] Step 3: Analyzing the video data

[1091] The server inputs the preprocessed video data into a generative AI model, which analyzes the number and location of passengers from the video data and quantifies the level of congestion.

[1092] Input: Preprocessed video data

[1093] Output: Numerical data of congestion

[1094] Specific behavior:

[1095] 1. Preprocessed video data is passed to a generative AI model.

[1096] 2. The model extracts facial and body features of each passenger and identifies their number and location.

[1097] 3. Based on the number and location of passengers, the congestion level is quantified on a scale of 0 to 100.

[1098] Step 4: Recognize emotions

[1099] The server then inputs the same video data into an emotion engine to recognize passenger emotions, which uses facial recognition and facial expression analysis to classify emotions.

[1100] Input: Preprocessed video data

[1101] Output: Passenger emotion data

[1102] Specific behavior:

[1103] 1. Detect the faces of each passenger from the preprocessed video data.

[1104] 2. Use a facial recognition algorithm to analyze each passenger's facial expression.

[1105] 3. Classify emotions such as discomfort, stress, and joy, and save the results as numerical data.

[1106] Step 5: Integrating crowding and emotion data

[1107] The server combines the congestion data obtained from the generative AI model with the emotion data obtained from the emotion engine, and this data is used to control elevator operation.

[1108] Input: Numerical data on congestion levels, passenger sentiment data

[1109] Output: Integrated traffic control data

[1110] Specific behavior:

[1111] 1. Obtain numerical data on congestion levels and emotional data.

[1112] 2. Combine both sets of data to create a detailed picture of elevator operation.

[1113] Step 6: Determine the stops

[1114] The server determines the elevator's stopping floor based on the integrated data, taking into account the level of congestion and passenger sentiment to select the optimal stopping floor.

[1115] Input: Integrated traffic control data

[1116] Output: Updated floor stop list

[1117] Specific behavior:

[1118] 1. Analyze the integrated data to assess congestion and emotions.

[1119] 2. If the threshold is exceeded, floors with fewer passengers or less discomfort will be selected preferentially.

[1120] 3. Create an updated list of stops.

[1121] Step 7: Elevator operation control

[1122] The server controls elevator operation based on the new stop list, avoiding unnecessary stops and notifying passengers of the operation status.

[1123] Input: Updated floor list

[1124] Output: Operation control instructions and passenger notifications

[1125] Specific behavior:

[1126] 1. Coordinate elevator operation based on a list of stops.

[1127] 2. Use displays and voice guidance systems inside the elevator to inform passengers of the next floor (e.g., announce "Next floor is floor 5").

[1128] (Application example 2)

[1129] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1130] Conventional elevator systems only analyzed congestion levels and did not take passengers' emotions into consideration. As a result, even when congestion levels were high, it was difficult to alleviate passenger discomfort, potentially resulting in a poor user experience. Furthermore, it was difficult to properly control elevator operation to avoid unnecessary elevator stops, hindering efficient operation.

[1131] The identification processing by the identification processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for capturing video images inside the elevator with a camera in real time, means for inputting the captured video data into a generative AI model and quantifying the congestion level inside the elevator, means for determining the elevator's stopping floor when the congestion level numerical value exceeds a preset threshold, means for inputting the video data into an emotion engine to recognize passenger emotions, means for integrating the emotion score recognized by the emotion engine with the congestion level data, and means for optimizing the elevator's stopping floors based on the integrated data. This enables efficient and comfortable elevator operation that takes into account not only the number of passengers but also the emotions of the passengers.

[1132] The "camera" is a device that captures images inside the elevator in real time.

[1133] "Video data" refers to video information inside the elevator captured by a camera.

[1134] The "generative AI model" is an artificial intelligence algorithm that analyzes the acquired video data and quantifies the degree of congestion inside the elevator.

[1135] The "crowding level" is a numerical value that indicates the degree of congestion calculated based on the number and arrangement of passengers in the elevator.

[1136] A "threshold" is a set value that, when exceeded, triggers a specific action (for example, determining the elevator's stopping floor).

[1137] The "emotion engine" is an algorithm that analyzes video data and recognizes passenger emotions.

[1138] The "emotion score" is a numerical representation of the passenger's emotional state as recognized by the emotion engine.

[1139] "Integrated data" is a dataset that integrates congestion data and emotion scores.

[1140] "Means for determining elevator stopping floors" refers to algorithms or devices that determine the optimal elevator stopping floor based on congestion and emotion score data.

[1141] This invention is a system for optimizing elevator operations in brick-and-mortar stores and shopping malls. The system consists of a camera, a server, a generative AI model, an emotion engine, and an elevator control system.

[1142] System Configuration

[1143] 1. Camera:

[1144] The cameras are installed inside the elevator and capture video data in real time, including the number, location, and facial expressions of passengers.

[1145] 2. Server:

[1146] The server preprocesses the acquired video data, which includes noise reduction and resolution adjustment. The preprocessed video data is then input into the generative AI model and emotion engine.

[1147] 3. Generative AI Model:

[1148] The generative AI model analyzes the pre-processed video data and quantifies the congestion level based on the number and location of passengers. For example, if there are 30 passengers, the congestion level is calculated as 85.

[1149] 4. Emotion Engine:

[1150] The emotion engine analyzes pre-processed video data to recognize passenger emotions. It uses facial recognition and facial expression analysis technology to detect emotions such as discomfort and stress. For example, it can recognize that 10 passengers are feeling discomfort.

[1151] 5. Elevator control system:

[1152] The server combines the congestion data from the generative AI model with the emotion data from the emotion engine. It then optimizes the elevator's floor stops based on the combined data. For example, if the congestion level of 85 exceeds the threshold of 70, it takes into account the discomfort data and determines that continuing normal operation will increase passenger discomfort.

[1153] Elevator stop lists will be updated and operational instructions will be updated to stop only on the first and fourth floors to minimize congestion.

[1154] Specific examples

[1155] In an elevator in a shopping mall, a server collects video data from a camera every second. This data includes not only the number and location of passengers, but also their facial expressions. The server preprocesses this video data before inputting it into a generative AI model and an emotion engine. For example, the generative AI model detects that there are 30 passengers and calculates the congestion level as "85." Meanwhile, the emotion engine recognizes that 10 passengers are feeling uncomfortable.

[1156] The server determines that the elevator is crowded because the congestion level of 85 exceeds the threshold of 70. Furthermore, taking into account discomfort data, the server assesses that continuing normal operation may increase passenger discomfort. The elevator, which was scheduled to stop on the third, fifth, and seventh floors, updates its operation instructions to stop only on the first and fourth floors, reflecting the congestion level and emotional data to most effectively reduce passenger discomfort.

[1157] Prompt Sentence Examples

[1158] Number of passengers in the elevator: 30

[1159] Passenger Sentiment Score:

[1160] 10 people found it "unpleasant"

[1161] 15 people answered "normal",

[1162] 5 people said "comfortable",

[1163] Crowding level: 85

[1164] Threshold: 70

[1165] In this way, the server can realize efficient and comfortable elevator operation that takes into account not only the number of passengers but also their emotions.

[1166] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[1167] Step 1:

[1168] The server acquires video data in real time from a camera installed inside the elevator. The input is the raw video from the camera, and the output is the acquired raw video data. This video data includes the number, location, and facial expressions of passengers inside the elevator.

[1169] Step 2:

[1170] The server preprocesses the acquired video data. The input is the raw video data acquired in step 1, and the output is preprocessed video data with noise removal and resolution adjustment. Specific operations include applying a noise removal filter and resizing the resolution to a fixed size.

[1171] Step 3:

[1172] The server inputs the preprocessed video data into the generative AI model. The input is the preprocessed video data, and the output is data that quantifies the degree of congestion inside the elevator. Specifically, the generative AI model analyzes the number and placement of passengers and calculates the degree of congestion as a number (e.g., "85").

[1173] Step 4:

[1174] The server also inputs the preprocessed video data into the emotion engine. The input is the preprocessed video data generated in step 2, and the output is the emotion score for each passenger. Specifically, the emotion engine uses face recognition technology and facial expression analysis technology to calculate each passenger's emotion (e.g., "uncomfortable," "neutral," or "comfortable") as a score.

[1175] Step 5:

[1176] The server integrates the congestion data from the generative AI model and the emotion data from the emotion engine. The input is the congestion data and emotion score, and the output is an integrated dataset. Specifically, it combines each data into a single dataset to comprehensively evaluate the elevator's condition.

[1177] Step 6:

[1178] The server optimizes elevator stops based on the integrated data. The input is the integrated dataset, and the output is an optimized elevator stop list. Specifically, if the congestion level exceeds a threshold and discomfort is high, the server adjusts the elevator stops to reduce discomfort (e.g., stopping only on the first and fourth floors).

[1179] Step 7:

[1180] The server sends the optimized elevator stop list to the elevator control system. The input is the optimized stop list, and the output is the operation instructions to be executed by the elevator control system. Specifically, the server sends new stop information to the control system to control the actual operation of the elevator.

[1181] The specific processing unit 290 transmits the result of the specific processing to the headset type terminal 314. In the headset type terminal 314, the control unit 46A causes the speaker 240 and the display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[1182] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[1183] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the headset type terminal 314.

[1184] [Fourth embodiment]

[1185] FIG. 7 shows an example of the configuration of a data processing system 410 according to the fourth embodiment.

[1186] 7, a data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.

[1187] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[1188] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a control target 443. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the control target 443 are also connected to the bus 52.

[1189] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[1190] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[1191] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[1192] The control object 443 includes a display device, LEDs in the eyes, and motors for driving the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the emotions of the robot 414 can be expressed by controlling these motors. In addition, the facial expressions of the robot 414 can also be expressed by controlling the light emission state of the LEDs in the eyes of the robot 414.

[1193] Fig. 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Fig. 8, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[1194] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[1195] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[1196] In the robot 414, the processor 46 performs the reception output process. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[1197] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1198] System Configuration

[1199] The present invention is implemented in the following configuration. A camera is installed inside the elevator, and a server collects real-time video data captured by the camera. The server preprocesses the video data and inputs it into a generative AI model to quantify the elevator congestion level. If the congestion level exceeds a set threshold, the server determines the floor at which the elevator will stop, and omits stops if necessary.

[1200] Overview of program processing flow

[1201] 1. Camera image acquisition

[1202] The server receives real-time video data from cameras installed inside the elevator, which is used to understand the situation of passengers inside the elevator.

[1203] 2. Preprocessing of video data

[1204] The server performs preprocessing such as noise removal and resolution adjustment on the acquired video data before inputting it into the generative AI model.

[1205] 3. Analysis of video data

[1206] The server inputs the preprocessed video data into a generative AI model to analyze the number and location of passengers, which then quantifies the degree of congestion inside the elevator.

[1207] 4. Determining congestion

[1208] The server verifies the congestion level value obtained from the generative AI model and determines whether the congestion level exceeds a set threshold.

[1209] 5. Determining the stopping floor

[1210] The server obtains the current floor stop list based on the buttons inside the elevator, and then updates the floor stop list so that the elevator does not stop at floors where passengers inside the elevator do not plan to get off.

[1211] 6. Elevator operation control

[1212] The server controls elevator operation based on the updated floor stop list to avoid unnecessary stops during peak hours, and notifies passengers of the next floor stop via a display and voice guidance inside the elevator.

[1213] Specific examples

[1214] 1. Example in a shopping mall

[1215] In an elevator in a shopping mall, the server acquires video data from the camera every second. For example, suppose the elevator is very crowded at a certain time.

[1216] The server preprocesses this video data and inputs it into a generative AI model, which detects that there are 30 passengers in the elevator and quantifies the congestion level as "85."

[1217] Since the congestion level exceeds the set threshold "70", the server determines that the elevator is crowded.

[1218] Next, the server checks the button status inside the elevator and confirms that although the elevator was scheduled to stop on the 3rd, 5th, and 7th floors, there are no passengers getting off on the 3rd and 7th floors.

[1219] The server updates the floor stop list and adjusts it to skip floors 3 and 7.

[1220] Ultimately, the elevator will stop only on the fifth floor, and a display and voice guidance inside the elevator will inform passengers that the next floor is floor 5. In this way, unnecessary stops can be avoided and the elevator can be operated efficiently.

[1221] This invention optimizes the operation of crowded elevators, reducing wasted time for passengers and reducing the amount of time that passengers waste waiting when they cannot board, thereby improving the overall efficiency of elevator use.

[1222] The processing flow will be explained below.

[1223] Step 1:

[1224] The server receives real-time video data from a camera installed inside the elevator, capturing the entire elevator interior and allowing the system to understand the situation of passengers inside the elevator.

[1225] Step 2:

[1226] The server temporarily stores the acquired video data and begins preprocessing, which includes removing noise and adjusting the resolution of the video, preparing it for input to the generative AI model.

[1227] Step 3:

[1228] The server inputs the preprocessed video data into a generative AI model, which then detects the number of passengers from the video data and analyzes their placement and movements.

[1229] Step 4:

[1230] The generative AI model quantifies the congestion level based on the number and location of passengers. The server receives this quantified congestion level and proceeds to the next step.

[1231] Step 5:

[1232] The server compares the quantified congestion level with a preset threshold. If the congestion level exceeds the threshold, the elevator is determined to be congested, but if it does not exceed the threshold, it continues to operate normally.

[1233] Step 6:

[1234] If the elevator is determined to be busy, the server retrieves the current floor stop list, which is based on floor button presses inside the elevator.

[1235] Step 7:

[1236] The server analyzes the stop floor list to identify floors where passengers in the elevator plan to get off, extracts floors where passengers do not plan to get off, and updates the stop floor list so that the elevator does not stop at those floors.

[1237] Step 8:

[1238] The server generates new operation instructions based on the updated floor list, including the floor where the elevator will next stop.

[1239] Step 9:

[1240] The server sends the generated operation instructions to the elevator control system, which operates the elevator based on these instructions and avoids unnecessary stops.

[1241] Step 10:

[1242] The server sends the next floor information to the display device and voice guidance device inside the elevator, explicitly notifying passengers of the next floor, thus ensuring that passengers inside the elevator are properly informed.

[1243] As a concrete example, elevators in a shopping mall tend to be crowded during the morning peak hours. For example, the server inputs video data acquired from a camera into a generative AI model to calculate a congestion level of "85," which it determines is above the threshold of "70." As a result, it determines that there are no passengers on the third and seventh floors, and generates operation instructions to skip stopping at those floors. The elevator then stops only on the fifth floor, allowing for efficient operation.

[1244] Example 1

[1245] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1246] Conventional elevator systems lacked the means to properly manage passenger congestion and optimize operational efficiency. This resulted in unnecessary elevator stops and congestion that prevented passengers from moving smoothly. This also led to problems such as wasted energy and increased passenger stress.

[1247] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[1248] In this invention, the server includes means for capturing images of the inside of the elevator with a camera in real time, means for preprocessing the captured image data, means for inputting the preprocessed image data into a generative AI model to quantify the congestion level inside the elevator, means for determining the elevator's stopping floor when the congestion level exceeds a preset threshold, and means for notifying the next stopping floor via a display device or voice guidance inside the elevator. This improves elevator operation efficiency, reduces unnecessary waiting time for passengers, and contributes to energy savings.

[1249] The "camera" is a device for capturing images inside the elevator in real time.

[1250] "Video data" refers to real-time image information of the inside of an elevator captured by a camera.

[1251] "Preprocessing" refers to the process of removing noise, adjusting resolution, normalizing, etc. from video data, and converting it into a format suitable for the generative AI model.

[1252] The "generative AI model" is a machine learning model that uses a deep learning algorithm to analyze and quantify the number of passengers and congestion levels from input video data.

[1253] The "crowding level" is a numerical value that indicates the degree of congestion calculated based on the number and arrangement of passengers in the elevator.

[1254] A "threshold" is a reference value set to trigger a specific action when the congestion level exceeds this threshold.

[1255] "Stop floor" means a floor at which an elevator stops during a particular operation.

[1256] A "display device" is a device installed inside an elevator that visually provides passengers with information such as the next floor.

[1257] "Voice guidance" is a system that notifies passengers of the next floor and other information by voice inside the elevator.

[1258] The present invention is a system that analyzes the congestion level in an elevator in real time and optimizes operation efficiency. This system is implemented with the following configuration.

[1259] System Configuration

[1260] The system consists of a camera installed inside the elevator, a server, a generative AI model, and an elevator control system.

[1261] Hardware and Software

[1262] camera:

[1263] Cameras will be installed inside the elevators to capture high-resolution video in real time, providing images at 30 frames per second.

[1264] server:

[1265] The server is a central processing unit that receives video data sent from the camera and performs preprocessing and analysis. It is equipped with a high-performance CPU and GPU and has sufficient computing resources to run generative AI models.

[1266] Generative AI models:

[1267] This is a machine learning model that analyzes video data and quantifies the degree of congestion in an elevator. It uses a model that uses a deep learning algorithm (such as YOLOv5).

[1268] Elevator Control System:

[1269] The system controls elevator operation based on control signals sent from the server, and notifies passengers of the next floor via a display device and voice guidance system inside the elevator.

[1270] Example of operation

[1271] 1. Elevators in shopping malls:

[1272] The server acquires video data from the camera inside the elevator every second.

[1273] The server first pre-processes the acquired video data, which includes noise removal, resolution adjustment, and normalization.

[1274] The preprocessed video data is input into a generative AI model (e.g., YOLOv5) to analyze the number and location of passengers in the elevator. The generative AI model detects that there are 30 passengers in the elevator and quantifies the congestion level as "85."

[1275] The server checks whether the congestion level exceeds the set threshold value "70".

[1276] The server checks the button status inside the elevator and confirms that although the elevator was scheduled to stop on the 3rd, 5th, and 7th floors, there are no passengers getting off on the 3rd and 7th floors.

[1277] The server updates the stop floor list and adjusts it to skip floors 3 and 7.

[1278] The elevator will finally stop on the fifth floor, and a display and voice guidance inside the elevator will inform passengers that the next floor is the fifth floor.

[1279] Example prompts for generative AI models

[1280] "Please use the camera footage from inside the elevator to quantify the current congestion level. The video data shows 30 passengers. The congestion level is 85."

[1281] The system of the present invention optimizes elevator operation efficiency, reduces passenger waiting times, cuts energy waste, and ensures that passengers are transported to their destinations quickly and efficiently, even during peak hours.

[1282] The flow of the identification process in the first embodiment will be described with reference to FIG.

[1283] Step 1: Acquire camera footage

[1284] The server acquires video data in real time from cameras installed inside the elevator.

[1285] Input: Real-time video data sent from the camera.

[1286] Specific operation: Video data is sent to the server as stream data at 30 frames per second.

[1287] Output: Captured real-time video data.

[1288] Step 2: Preprocessing the video data

[1289] The server preprocesses the acquired video data before inputting it into the generative AI model.

[1290] Input: Video data acquired from the camera.

[1291] Specific operations: noise removal, resolution adjustment (e.g. convert to 640x480 pixels), pixel value normalization (scaling to the range 0-1).

[1292] Output: Preprocessed video data.

[1293] Step 3: Analyzing the video data

[1294] The server inputs the preprocessed video data into a generative AI model to analyze passenger numbers and placements.

[1295] Input: Preprocessed video data.

[1296] Specific operation: Uses a generative AI model (e.g., YOLOv5) to detect the number of passengers in an elevator from video data and analyze their placement.

[1297] Output: Number of passengers, location information and congestion level (e.g. "85").

[1298] Step 4: Determine crowding level

[1299] The server verifies the congestion level values ​​obtained from the generative AI model.

[1300] Input: The congestion level output from the generative AI model.

[1301] Specific behavior: Compare the congestion level value with a pre-set threshold (e.g. "70").

[1302] Output: Boolean value (e.g. true) indicating whether the congestion threshold is exceeded.

[1303] Step 5: Determine the stops

[1304] The server obtains the button status inside the elevator and checks and updates the current floor stop list.

[1305] Input: Button status inside the elevator, congestion level judgment result.

[1306] Specific behavior: Get the current floor stop list and remove floors where passengers in the elevator do not plan to get off (e.g., floors 3 and 7) from the floor stop list.

[1307] Output: Updated floor list (e.g. only floor 5).

[1308] Step 6: Elevator operation control

[1309] The server controls elevator operation based on the updated stop list and notifies passengers of the next stop.

[1310] Input: Updated floor stop list.

[1311] What it does: Sends instructions to the elevator control system to stop only at the specified floor, and notifies passengers of the next stop via a visual display and voice announcement (e.g., "Next stop is floor 5").

[1312] Output: Elevator control signal, next floor notification information.

[1313] Through the above process, elevator operation efficiency can be optimized, congestion can be reduced, and energy can be saved.

[1314] (Application example 1)

[1315] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1316] In conventional factory robot movement, it has been difficult to grasp congestion conditions in real time and select efficient movement routes. This has resulted in reduced productivity and hindered efficient work. In particular, in areas of the factory where congestion frequently occurs, robot movement is hindered, resulting in a significant drop in production speed and work efficiency.

[1317] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[1318] In this invention, the server includes a means for capturing images of the area in real time using a camera, a means for inputting the captured image data into a generative AI model to quantify the degree of congestion in the area, and a means for optimizing the robot's movement route when the numerical value of the degree of congestion exceeds a preset threshold. This makes it possible to analyze the congestion situation in real time and instruct the robot on the optimal movement route.

[1319] A "camera" is a device that captures video within a specified area in real time.

[1320] A "generative AI model" is an artificial intelligence model used to quantify the degree of congestion based on acquired video data.

[1321] The "degree of congestion" is a numerical value that indicates the degree of congestion calculated from the number of people and their locations within a specific area.

[1322] A "threshold" is a numerical value that serves as a reference for taking a specific action when the congestion level exceeds this threshold.

[1323] "Stop floor" means a floor where an elevator stops to allow passengers to board or alight.

[1324] A "movement route" is a path chosen by a robot or other moving object when it moves within an area.

[1325] "Notification" is the act of informing a robot or worker of the next action to be taken or the selected route.

[1326] "Preprocessing" refers to the process of processing the video data, such as removing noise and adjusting resolution, before inputting it into the generative AI model.

[1327] System Configuration

[1328] This invention is implemented with the following configuration: Cameras are installed in a factory area, and a server collects real-time video data captured by the cameras. The server is a system that preprocesses this video data and inputs it into a generative AI model to quantify the congestion level of the area. If the congestion level exceeds a set threshold, the server optimizes the robot's movement route and changes the route as necessary. It then notifies the robot or worker of this information.

[1329] Overview of program processing flow

[1330] 1. Camera image acquisition

[1331] The server collects real-time video data from cameras installed in the area, which is used to understand the congestion situation within the area.

[1332] 2. Preprocessing of video data

[1333] The server performs preprocessing such as noise removal and resolution adjustment on the acquired video data before inputting it into the generative AI model.

[1334] 3. Analysis of video data

[1335] The server inputs the preprocessed video data into a generative AI model to analyze the number of people and their locations within the area, which then quantifies the level of congestion.

[1336] 4. Determining congestion

[1337] The server verifies the congestion level value obtained from the generative AI model and determines whether the congestion level exceeds a set threshold.

[1338] 5. Deciding on a travel route

[1339] The server calculates the optimal route for the robot based on the congestion information, and selects a route that avoids congestion.

[1340] 6. Notification

[1341] The server notifies the robots and workers of the selected movement route, enabling them to act quickly and efficiently.

[1342] Hardware and software used

[1343] Hardware:

[1344] Camera: A device installed in a factory area that captures images within the area in real time.

[1345] Server: A computer for processing and analyzing video data.

[1346] software:

[1347] OpenCV: A library for acquiring and preprocessing camera footage.

[1348] Keras: A framework for loading and analyzing generative AI models.

[1349] Specific examples

[1350] When a specific area in the factory is extremely congested, the server receives real-time video data from the camera, preprocesses it, and then inputs it into the generative AI model. For example, if the area is highly congested, the server instructs the robot to select a different route. If the area is less congested, the robot will select the normal route to maximize movement efficiency.

[1351] Prompt Sentence Examples

[1352] "Please obtain video data from camera ID: {camera_id}, analyze the congestion level, and calculate the optimal travel route."

[1353] In this way, the operating efficiency of robots in the factory can be optimized, congestion can be avoided, and productivity can be improved.

[1354] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[1355] Step 1:

[1356] The server acquires video data in real time from cameras installed in the area. Specifically, it captures the video from the cameras and sends the video data to the server. The input data is the video stream from the cameras, and the output data is the video frames stored on the server.

[1357] Step 2:

[1358] The server preprocesses the captured video data, specifically using OpenCV to remove noise and adjust resolution. The input data is raw video frames, and the output data is video data converted into a format suitable for the generative AI model.

[1359] Step 3:

[1360] The server inputs the preprocessed video data into a generative AI model to analyze the number of people and their locations within the area. Specifically, it uses Keras to analyze the video data and quantify the degree of congestion. The input data is the preprocessed video data, and the output data is a numerical value indicating the degree of congestion.

[1361] Step 4:

[1362] The server verifies the congestion level value obtained from the generative AI model and determines whether the congestion level exceeds a set threshold. The input data is the congestion level value, and the output data is a boolean value indicating whether the congestion level exceeds the threshold.

[1363] Step 5:

[1364] The server calculates the optimal route for the robot based on the congestion information. Specifically, it executes an algorithm to select the least congested route. The input data is the congestion level value and default route information, and the output data is the optimized route.

[1365] Step 6:

[1366] The server notifies the robot and the worker of the selected movement route. Specifically, it sends the notification via the network to the robot's operation system or the worker's terminal. The input data is the optimized movement route, and the output data is the notification information.

[1367] Through the above processing steps, it is expected that the operating efficiency of robots in the factory will be optimized and productivity will be improved by avoiding congestion.

[1368] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[1369] System Configuration

[1370] The present invention is a system that combines a camera installed inside an elevator with an emotion engine. The camera captures video footage inside the elevator in real time, preprocesses this video data, and inputs it into a generative AI model. The generative AI model analyzes the number and placement of passengers and quantifies the level of congestion. If this level of congestion exceeds a set threshold, the server uses the data from the emotion engine to determine the elevator's stopping floor. The emotion engine uses the camera's video data to recognize passenger emotions and, in conjunction with the level of congestion, determines the optimal stopping floor.

[1371] Overview of program processing flow

[1372] 1. Camera image acquisition

[1373] The server collects real-time video data from cameras installed inside the elevator, which serves as the basis for understanding the situation and emotions of passengers inside the elevator.

[1374] 2. Preprocessing of video data

[1375] The server pre-processes the captured video data, which includes noise reduction and resolution adjustment.

[1376] 3. Analysis of video data

[1377] The server inputs the preprocessed video data into a generative AI model, which analyzes the number and placement of passengers and quantifies the level of congestion.

[1378] 4. Emotional Recognition

[1379] The server also inputs the same video data into an emotion engine to recognize each passenger's emotions. The emotion engine uses facial recognition and facial expression analysis technology to detect passengers' feelings such as discomfort or stress.

[1380] 5. Integrating crowding and emotion data

[1381] The server integrates congestion data from the generative AI model with emotion data from the emotion engine, and optimizes elevator operation based on the integrated data.

[1382] 6. Determining the stopping floor

[1383] The server updates the elevator's floor stop list based on the integrated data, and adjusts the elevator to avoid stopping at floors where passengers feel uncomfortable, especially when the congestion level exceeds a threshold and discomfort is detected by the emotion engine.

[1384] 7. Elevator operation control

[1385] The server controls elevator operation based on the new stop list to avoid unnecessary stops, and uses displays and audio announcements inside the elevator to notify passengers of the next stop.

[1386] Specific examples

[1387] 1. Example in a shopping mall

[1388] In an elevator in a shopping mall, a server captures video data from a camera every second, including not only the number and location of passengers but also their facial expressions.

[1389] The server preprocesses this video data and then inputs it into the generative AI model and emotion engine. The generative AI model detects that there are 30 passengers and calculates the congestion level as 85. Meanwhile, the emotion engine recognizes that 10 passengers are feeling uncomfortable.

[1390] The server determines that the elevator is crowded because the congestion level of "85" exceeds the threshold of "70." Furthermore, taking into account the discomfort data, the server assesses that continuing normal operation may increase passenger discomfort.

[1391] An elevator that was scheduled to stop on floors 3, 5, and 7 will check to make sure there are no passengers on those floors, and update its operation instructions to stop only on floor 5, reflecting congestion and sentiment data, to most effectively reduce passenger discomfort.

[1392] The server sends the generated new operation instructions to the elevator control system, and passengers are notified that the next stop is floor 5 via a display and voice guidance inside the elevator.

[1393] This embodiment not only improves the efficiency of elevator operation but also enhances passenger convenience. It not only avoids unnecessary stops in crowded situations but also provides a comfortable user environment by taking passengers' emotions into consideration.

[1394] The processing flow will be explained below.

[1395] Step 1:

[1396] The server acquires real-time video data from a camera installed inside the elevator, which includes information on the number and location of passengers in the elevator, as well as emotional information such as facial expressions.

[1397] Step 2:

[1398] The server stores the acquired video data in temporary storage and begins preprocessing, which includes noise reduction and resolution adjustment, to prepare it for input to the generative AI model and emotion engine.

[1399] Step 3:

[1400] The server inputs the preprocessed video data into a generative AI model, which analyzes the number and location of passengers from the video data to quantify the level of congestion.

[1401] Step 4:

[1402] At the same time, the server inputs the pre-processed video data into the emotion engine, which uses facial recognition and facial expression analysis technologies to recognize each passenger's emotions (discomfort, stress, etc.).

[1403] Step 5:

[1404] As a result of the analysis, the generative AI model detects that there are 30 passengers in the elevator and calculates the congestion level as "85." This data is received by the server.

[1405] Step 6:

[1406] As a result of the analysis, the emotion engine recognizes that 10 passengers are feeling uncomfortable. This data is received by the server.

[1407] Step 7:

[1408] The server combines the congestion data obtained from the generative AI model with the emotion data obtained from the emotion engine. Since the congestion level is "85" and there are many passengers feeling uncomfortable, it determines that the elevator is crowded.

[1409] Step 8:

[1410] The server gets the elevator's current list of stops, say it's scheduled to stop at floors 3, 5, and 7.

[1411] Step 9:

[1412] The server analyzes the list of stops and identifies floors where the passengers in the elevator do not plan to get off. It confirms that the passengers do not plan to get off on the third and seventh floors.

[1413] Step 10:

[1414] Taking into account the congestion level and emotion data, the server updates the stop floor list to skip stops on floors 3 and 7 and stop only on floor 5 in order to most effectively reduce passenger discomfort.

[1415] Step 11:

[1416] The server generates new operation instructions based on the updated floor stop list and sends the instructions to the elevator control system to control elevator operation.

[1417] Step 12:

[1418] When the elevator is in operation, the server sends information about the next floor to the display device and voice guidance system inside the elevator, notifying passengers of the next floor. Specifically, the server informs passengers that the next floor is floor 5.

[1419] As a concrete example, suppose an elevator in a shopping mall is used by several passengers during the morning peak, with the camera detecting 30 passengers and the emotion engine recognizing the discomfort of 10 of them. Based on this data, the server calculates a congestion level of 85, exceeding the preset threshold of 70, and controls the elevator so that it stops only on the fifth floor, since no passengers are getting off on the third or seventh floor. Passengers are notified by display and audio guidance that the elevator will next stop on the fifth floor, ensuring efficient and comfortable operation.

[1420] Example 2

[1421] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1422] Conventional elevator operation systems lacked the means to achieve efficient operation, and were unable to determine the optimal stopping floor, especially during crowded times or based on passenger emotions. Furthermore, when passengers felt uncomfortable inside the elevator, conventional systems had difficulty responding appropriately. This sometimes led to lower satisfaction among elevator users.

[1423] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[1424] In this invention, the server includes means for capturing images of the inside of the elevator with a camera in real time, means for preprocessing the captured image data, means for inputting the preprocessed image data into a generative AI model and quantifying the degree of congestion inside the elevator, means for inputting the same image data into an emotion engine and recognizing passenger emotions, and means for integrating the congestion data obtained from the generative AI model and the emotion data obtained from the emotion engine and determining the floors at which the elevator should stop. This enables optimal operation that simultaneously takes into account the congestion status inside the elevator and the emotions of passengers.

[1425] The "camera" is a photographic device that captures images inside the elevator in real time.

[1426] "Video data" is digital data captured by a camera showing the situation inside the elevator.

[1427] "Preprocessing" refers to processes such as noise removal and resolution adjustment that are performed to improve the quality of acquired video data.

[1428] The "generative AI model" is an artificial intelligence model that analyzes video data and quantifies the degree of congestion inside an elevator.

[1429] "Crowding level" is a number that indicates the degree of congestion calculated by the generative AI model based on the number and placement of passengers in the elevator.

[1430] The "emotion engine" is an engine that analyzes video data and recognizes emotions from passengers' facial expressions.

[1431] "Integrated data" refers to data that integrates congestion data obtained from the generative AI model and emotion data obtained from the emotion engine.

[1432] "Stop floor" refers to the floor at which the elevator is scheduled to stop.

[1433] This invention is a system that combines a camera installed inside the elevator, a generative AI model, and an emotion engine, and simultaneously analyzes the congestion situation inside the elevator and the emotions of passengers to achieve optimal operation. Below, we will explain the hardware and software, as well as the specific processing and calculations.

[1434] Hardware

[1435] 1. Camera

[1436] These are IP cameras or surveillance cameras installed inside elevators. These cameras capture images from inside the elevator in real time and send them to a streaming server.

[1437] 2. Server

[1438] This is the central processing unit of the system, and performs preprocessing and analysis of video data. The server is equipped with a high-performance CPU and GPU, and runs deep learning frameworks such as Python, TensorFlow, and PyTorch.

[1439] software

[1440] 1. Preprocessing software

[1441] This software tool removes noise from video data, adjusts resolution, and adjusts contrast, etc. This improves the accuracy of analysis.

[1442] 2. Generative AI Models

[1443] This is an artificial intelligence model that analyzes the number and placement of passengers in an elevator from video data and quantifies the degree of congestion. It is usually implemented using Python, TensorFlow, and PyTorch.

[1444] 3. Emotion Engine

[1445] This is a software engine for recognizing passenger emotions from video data. It uses face recognition and facial expression analysis technologies and uses Python, OpenCV, and the Dlib library.

[1446] Example of a system

[1447] Elevator operation in shopping malls

[1448] The server collects video data every second from cameras installed in elevators in the shopping mall, including the number of passengers, their locations, and their facial expressions.

[1449] 1. Pretreatment

[1450] The server pre-processes this video data, removing noise and adjusting the resolution, for example, if the initial resolution is 720p, it will change it to 1080p to improve the quality of the data.

[1451] 2. Input to the generative AI model and emotion engine

[1452] The preprocessed data is input into the generative AI model and the emotion engine. The generative AI model detects that there are 30 passengers in the elevator and quantifies the congestion level as "85." Meanwhile, the emotion engine recognizes that 10 of the passengers are feeling uncomfortable.

[1453] 3. Integration and Operation Decisions

[1454] The server combines the congestion data obtained from the generative AI model with the emotion data obtained from the emotion engine. Based on this combined data, it determines the floors where the elevator should stop. For example, an elevator that is normally scheduled to stop on the third, fifth, and seventh floors will update its operation instructions to take into account the combined data and avoid stopping on the third and seventh floors, where there are no passengers, and stop only on the fifth floor.

[1455] Example prompt

[1456] For example, the following prompts might be used by a generative AI model:

[1457] "Analyze the number and location of passengers from video data inside the elevator, and quantify the congestion level on a scale of 0 to 100."

[1458] This invention not only improves the efficiency of elevator operation but also enhances passenger convenience. It not only avoids unnecessary stops in crowded situations but also provides a comfortable user environment by taking passengers' emotions into consideration.

[1459] The flow of the identification process in the second embodiment will be described with reference to FIG.

[1460] Program processing flow

[1461] The subject should be either the server, the terminal, or the user, and the explanation should be written in the plain voice (da / dearu style).

[1462] Step 1: Acquire camera footage

[1463] The server receives real-time video data from the camera. The camera captures video footage inside the elevator and streams it to the server. This video data includes the number and placement of passengers, facial expressions, etc.

[1464] Input: Real-time video from inside the elevator

[1465] Output: Captured video data

[1466] Specific behavior:

[1467] 1. The camera captures a frame every second and sends it to the streaming server.

[1468] 2. The server receives this stream and stores it in a database.

[1469] Step 2: Preprocessing the video data

[1470] The server pre-processes the captured video data, which includes noise reduction, resolution adjustment, and contrast enhancement.

[1471] Input: Captured video data

[1472] Output: Pre-processed video data

[1473] Specific behavior:

[1474] 1. A noise reduction filter is applied to the video data.

[1475] 2. Adjust the video resolution (e.g. convert from 720p to 1080p).

[1476] 3. Adjust contrast and brightness and clean abnormal pixels.

[1477] Step 3: Analyzing the video data

[1478] The server inputs the preprocessed video data into a generative AI model, which analyzes the number and location of passengers from the video data and quantifies the level of congestion.

[1479] Input: Preprocessed video data

[1480] Output: Numerical data of congestion

[1481] Specific behavior:

[1482] 1. Preprocessed video data is passed to a generative AI model.

[1483] 2. The model extracts facial and body features of each passenger and identifies their number and location.

[1484] 3. Based on the number and location of passengers, the congestion level is quantified on a scale of 0 to 100.

[1485] Step 4: Recognize emotions

[1486] The server then inputs the same video data into an emotion engine to recognize passenger emotions, which uses facial recognition and facial expression analysis to classify emotions.

[1487] Input: Preprocessed video data

[1488] Output: Passenger emotion data

[1489] Specific behavior:

[1490] 1. Detect the faces of each passenger from the preprocessed video data.

[1491] 2. Use a facial recognition algorithm to analyze each passenger's facial expression.

[1492] 3. Classify emotions such as discomfort, stress, and joy, and save the results as numerical data.

[1493] Step 5: Integrating crowding and emotion data

[1494] The server combines the congestion data obtained from the generative AI model with the emotion data obtained from the emotion engine, and this data is used to control elevator operation.

[1495] Input: Numerical data on congestion levels, passenger sentiment data

[1496] Output: Integrated traffic control data

[1497] Specific behavior:

[1498] 1. Obtain numerical data on congestion levels and emotional data.

[1499] 2. Combine both sets of data to create a detailed picture of elevator operation.

[1500] Step 6: Determine the stops

[1501] The server determines the elevator's stopping floor based on the integrated data, taking into account the level of congestion and passenger sentiment to select the optimal stopping floor.

[1502] Input: Integrated traffic control data

[1503] Output: Updated floor stop list

[1504] Specific behavior:

[1505] 1. Analyze the integrated data to assess congestion and emotions.

[1506] 2. If the threshold is exceeded, floors with fewer passengers or less discomfort will be selected preferentially.

[1507] 3. Create an updated list of stops.

[1508] Step 7: Elevator operation control

[1509] The server controls elevator operation based on the new stop list, avoiding unnecessary stops and notifying passengers of the operation status.

[1510] Input: Updated floor list

[1511] Output: Operation control instructions and passenger notifications

[1512] Specific behavior:

[1513] 1. Coordinate elevator operation based on a list of stops.

[1514] 2. Use displays and voice guidance systems inside the elevator to inform passengers of the next floor (e.g., announce "Next floor is floor 5").

[1515] (Application example 2)

[1516] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1517] Conventional elevator systems only analyzed congestion levels and did not take passengers' emotions into consideration. As a result, even when congestion levels were high, it was difficult to alleviate passenger discomfort, potentially resulting in a poor user experience. Furthermore, it was difficult to properly control elevator operation to avoid unnecessary elevator stops, hindering efficient operation.

[1518] The identification processing by the identification processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for capturing video images inside the elevator with a camera in real time, means for inputting the captured video data into a generative AI model and quantifying the congestion level inside the elevator, means for determining the elevator's stopping floor when the congestion level numerical value exceeds a preset threshold, means for inputting the video data into an emotion engine to recognize passenger emotions, means for integrating the emotion score recognized by the emotion engine with the congestion level data, and means for optimizing the elevator's stopping floors based on the integrated data. This enables efficient and comfortable elevator operation that takes into account not only the number of passengers but also the emotions of the passengers.

[1519] The "camera" is a device that captures images inside the elevator in real time.

[1520] "Video data" refers to video information inside the elevator captured by a camera.

[1521] The "generative AI model" is an artificial intelligence algorithm that analyzes the acquired video data and quantifies the degree of congestion inside the elevator.

[1522] The "crowding level" is a numerical value that indicates the degree of congestion calculated based on the number and arrangement of passengers in the elevator.

[1523] A "threshold" is a set value that, when exceeded, triggers a specific action (for example, determining the elevator's stopping floor).

[1524] The "emotion engine" is an algorithm that analyzes video data and recognizes passenger emotions.

[1525] The "emotion score" is a numerical representation of the passenger's emotional state as recognized by the emotion engine.

[1526] "Integrated data" is a dataset that integrates congestion data and emotion scores.

[1527] "Means for determining elevator stopping floors" refers to algorithms or devices that determine the optimal elevator stopping floor based on congestion and emotion score data.

[1528] This invention is a system for optimizing elevator operations in brick-and-mortar stores and shopping malls. The system consists of a camera, a server, a generative AI model, an emotion engine, and an elevator control system.

[1529] System Configuration

[1530] 1. Camera:

[1531] The cameras are installed inside the elevator and capture video data in real time, including the number, location, and facial expressions of passengers.

[1532] 2. Server:

[1533] The server preprocesses the acquired video data, which includes noise reduction and resolution adjustment. The preprocessed video data is then input into the generative AI model and emotion engine.

[1534] 3. Generative AI Model:

[1535] The generative AI model analyzes the pre-processed video data and quantifies the congestion level based on the number and location of passengers. For example, if there are 30 passengers, the congestion level is calculated as 85.

[1536] 4. Emotion Engine:

[1537] The emotion engine analyzes pre-processed video data to recognize passenger emotions. It uses facial recognition and facial expression analysis technology to detect emotions such as discomfort and stress. For example, it can recognize that 10 passengers are feeling discomfort.

[1538] 5. Elevator control system:

[1539] The server combines the congestion data from the generative AI model with the emotion data from the emotion engine. It then optimizes the elevator's floor stops based on the combined data. For example, if the congestion level of 85 exceeds the threshold of 70, it takes into account the discomfort data and determines that continuing normal operation will increase passenger discomfort.

[1540] Elevator stop lists will be updated and operational instructions will be updated to stop only on the first and fourth floors to minimize congestion.

[1541] Specific examples

[1542] In an elevator in a shopping mall, a server collects video data from a camera every second. This data includes not only the number and location of passengers, but also their facial expressions. The server preprocesses this video data before inputting it into a generative AI model and an emotion engine. For example, the generative AI model detects that there are 30 passengers and calculates the congestion level as "85." Meanwhile, the emotion engine recognizes that 10 passengers are feeling uncomfortable.

[1543] The server determines that the elevator is crowded because the congestion level of 85 exceeds the threshold of 70. Furthermore, taking into account discomfort data, the server assesses that continuing normal operation may increase passenger discomfort. The elevator, which was scheduled to stop on the third, fifth, and seventh floors, updates its operation instructions to stop only on the first and fourth floors, reflecting the congestion level and emotional data to most effectively reduce passenger discomfort.

[1544] Prompt Sentence Examples

[1545] Number of passengers in the elevator: 30

[1546] Passenger Sentiment Score:

[1547] 10 people found it "unpleasant,"

[1548] 15 people answered "normal",

[1549] 5 people said "comfortable",

[1550] Crowding level: 85

[1551] Threshold: 70

[1552] In this way, the server can realize efficient and comfortable elevator operation that takes into account not only the number of passengers but also their emotions.

[1553] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[1554] Step 1:

[1555] The server acquires video data in real time from a camera installed inside the elevator. The input is the raw video from the camera, and the output is the acquired raw video data. This video data includes the number, location, and facial expressions of passengers inside the elevator.

[1556] Step 2:

[1557] The server preprocesses the acquired video data. The input is the raw video data acquired in step 1, and the output is preprocessed video data with noise removal and resolution adjustment. Specific operations include applying a noise removal filter and resizing the resolution to a fixed size.

[1558] Step 3:

[1559] The server inputs the preprocessed video data into the generative AI model. The input is the preprocessed video data, and the output is data that quantifies the degree of congestion inside the elevator. Specifically, the generative AI model analyzes the number and placement of passengers and calculates the degree of congestion as a number (e.g., "85").

[1560] Step 4:

[1561] The server also inputs the preprocessed video data into the emotion engine. The input is the preprocessed video data generated in step 2, and the output is the emotion score for each passenger. Specifically, the emotion engine uses face recognition technology and facial expression analysis technology to calculate each passenger's emotion (e.g., "uncomfortable," "neutral," or "comfortable") as a score.

[1562] Step 5:

[1563] The server integrates the congestion data from the generative AI model and the emotion data from the emotion engine. The input is the congestion data and emotion score, and the output is an integrated dataset. Specifically, it combines each data into a single dataset to comprehensively evaluate the elevator's condition.

[1564] Step 6:

[1565] The server optimizes elevator stops based on the integrated data. The input is the integrated dataset, and the output is an optimized elevator stop list. Specifically, if the congestion level exceeds a threshold and discomfort is high, the server adjusts the elevator stops to reduce discomfort (e.g., stopping only on the first and fourth floors).

[1566] Step 7:

[1567] The server sends the optimized elevator stop list to the elevator control system. The input is the optimized stop list, and the output is the operation instructions to be executed by the elevator control system. Specifically, the server sends new stop information to the control system to control the actual operation of the elevator.

[1568] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the control target 443 to output the result of the specific processing. The microphone 238 acquires voice indicating a user input regarding the result of the specific processing. The control unit 46A transmits voice data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the voice data.

[1569] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[1570] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the robot 414.

[1571] The emotion identification model 59 as an emotion engine may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to an emotion map (see FIG. 9), which is a specific mapping. Similarly, the emotion identification model 59 may determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.

[1572] FIG. 9 is a diagram illustrating an emotion map 400 on which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. Emotions closer to the center of the concentric circles are more primitive. Emotions representing states and actions arising from a state of mind are arranged on the outer edges of the concentric circles. The concept of emotion includes both affect and mental states. Emotions generally generated from reactions occurring in the brain are arranged on the left side of the concentric circles. Emotions generally induced by situational judgment are arranged on the right side of the concentric circles. Emotions generally generated from reactions occurring in the brain and induced by situational judgment are arranged on the upper and lower sides of the concentric circles. Furthermore, the emotion of "pleasure" is arranged on the upper side of the concentric circles, and the emotion of "discomfort" is arranged on the lower side. In this way, in the emotion map 400, multiple emotions are mapped based on the structure by which emotions are generated, and emotions that tend to occur simultaneously are mapped close to each other.

[1573] These emotions are distributed in the 3 o'clock direction on emotion map 400, and typically fluctuate between relief and anxiety. In the right half of emotion map 400, situational awareness dominates over internal sensations, resulting in a sense of calm.

[1574] The inside of emotion map 400 represents what is going on in the mind, and the outside of emotion map 400 represents behavior, so the further you go outside emotion map 400, the more visible the emotions become (the more they are expressed in behavior).

[1575] Human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. Emotions can also be created for robots, automobiles, and motorcycles, based on various balances, such as posture and remaining battery life. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. An emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on Voice Emotion Recognition and Emotional Brain Physiological Signal Analysis Systems, Tokushima University, Doctoral Dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map lists emotions belonging to the "reaction" domain, where sensation is dominant. The right half of the emotion map lists emotions belonging to the "situation" domain, where situational awareness is dominant.

[1576] The emotion map defines two emotions that promote learning. One is a negative emotion on the situation side, around the middle of "repentance" or "reflection." In other words, this occurs when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is a positive emotion on the response side, around "desire." In other words, this occurs when the robot experiences positive feelings such as "I want more" or "I want to know more."

[1577] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values ​​indicating each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple pieces of training data that are combinations of user input and emotion values ​​indicating each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions that are located close to each other have similar values, as in the emotion map 900 shown in FIG. 10. FIG. 10 shows an example in which multiple emotions, "relieved," "calm," and "reassuring," have similar emotion values.

[1578] The system according to the present disclosure has been described above mainly with respect to the functions of the data processing device 12, but the system according to the present disclosure is not necessarily implemented on a server. The system according to the present disclosure may be implemented as a general information processing system. The present disclosure may be implemented, for example, as a software program running on a personal computer or an application running on a smartphone, etc. The method according to the present disclosure may be provided to users in the form of SaaS (Software as a Service).

[1579] In the above embodiment, an example was given in which the specific processing is performed by one computer 22, but the technology of the present disclosure is not limited to this, and the specific processing may be distributed and performed by a plurality of computers including the computer 22. For example, the data generation model 58 may be provided in an external device of the data processing device 12, and data may be generated in the external device in accordance with input data.

[1580] In the above embodiment, an example in which the specific processing program 56 is stored in the storage 32 has been described, but the technology of the present disclosure is not limited to this. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-transitory storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-transitory storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes the specific processing in accordance with the specific processing program 56.

[1581] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.

[1582] It is not necessary to store all of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store all of the specific processing program 56 in the storage 32; only a portion of the specific processing program 56 may be stored.

[1583] The hardware resource for executing a specific process can be any of the following processors: An example of a processor is a CPU, which is a general-purpose processor that functions as a hardware resource for executing a specific process by executing software, i.e., a program. Another example of a processor is a dedicated electrical circuit, such as an FPGA (Field-Programmable Gate Array), a PLD (Programmable Logic Device), or an ASIC (Application Specific Integrated Circuit), which is a processor with a circuit configuration designed specifically for executing a specific process. Each processor has built-in or connected memory, and each processor uses the memory to execute the specific process.

[1584] The hardware resource that executes the specific processing may be configured with one of these various processors, or may be configured with a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Also, the hardware resource that executes the specific processing may be a single processor.

[1585] As an example of a system configured with a single processor, first, one processor is configured by combining one or more CPUs and software, and this processor functions as a hardware resource that executes a specific process. Second, there is a system that uses a processor that realizes the functions of an entire system including multiple hardware resources that execute a specific process on a single IC chip, as typified by SoC (System-on-a-chip). In this way, a specific process is realized using one or more of the above-mentioned various processors as hardware resources.

[1586] Furthermore, the hardware structure of these various processors can be, more specifically, an electric circuit that combines circuit elements such as semiconductor devices. The specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps may be deleted, new steps may be added, or the processing order may be rearranged, without departing from the spirit of the invention.

[1587] The above-described description and illustrations are a detailed explanation of the parts related to the technology of the present disclosure and are merely an example of the technology of the present disclosure. For example, the above description of the configuration, functions, actions, and effects is an explanation of an example of the configuration, functions, actions, and effects of the parts related to the technology of the present disclosure. Therefore, it goes without saying that unnecessary parts may be deleted, new elements may be added, or replacements may be made to the above-described description and illustrations within the scope of the gist of the technology of the present disclosure. Furthermore, to avoid confusion and facilitate understanding of the parts related to the technology of the present disclosure, the above-described description and illustrations omit explanations of common technical knowledge that do not require particular explanation to enable the implementation of the technology of the present disclosure.

[1588] All publications, patent applications, and technical standards mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent application, or technical standard was specifically and individually indicated to be incorporated by reference.

[1589] The following is further disclosed regarding the above embodiment.

[1590] (Claim 1)

[1591] A means of capturing images of the inside of the elevator in real time using a camera,

[1592] The acquired video data is input into a generative AI model to quantify the degree of congestion inside the elevator.

[1593] a means for determining a floor at which the elevator should stop when the congestion degree value exceeds a preset threshold;

[1594] A system including:

[1595] (Claim 2)

[1596] 10. The system of claim 1, further comprising means for pre-processing video data acquired from the camera.

[1597] (Claim 3)

[1598] 10. The system of claim 1, further comprising means for calculating a congestion level based on the number and placement of passengers detected from the video data using a generative AI model.

[1599] "Example 1"

[1600] (Claim 1)

[1601] A means of capturing images of the inside of the elevator in real time using a camera,

[1602] means for pre-processing the acquired video data;

[1603] A means to input the preprocessed video data into a generative AI model and quantify the degree of congestion inside the elevator,

[1604] a means for determining a floor at which the elevator should stop when the congestion degree value exceeds a preset threshold;

[1605] A means for informing the next floor via a display device or voice guidance in the elevator;

[1606] A system including:

[1607] (Claim 2)

[1608] 10. The system of claim 1, further comprising means for normalizing the video data obtained from the camera.

[1609] (Claim 3)

[1610] 10. The system of claim 1, further comprising means for calculating a congestion level based on the number and placement of passengers detected from the preprocessed video data by the generative AI model.

[1611] "Application Example 1"

[1612] (Claim 1)

[1613] A means of capturing images of the inside of the elevator in real time using a camera,

[1614] The acquired video data is input into a generative AI model to quantify the degree of congestion inside the elevator.

[1615] a means for determining a floor at which the elevator should stop when the congestion degree value exceeds a preset threshold;

[1616] A means for optimizing a travel route based on congestion information of an area where the image is acquired;

[1617] A means of notifying optimal travel routes based on congestion information;

[1618] A system including:

[1619] (Claim 2)

[1620] 10. The system of claim 1, further comprising means for pre-processing video data acquired from the camera.

[1621] (Claim 3)

[1622] The system of claim 1, further comprising means for calculating a congestion level based on the number of people and their locations detected from the video data using a generative AI model.

[1623] "Example 2: Combining Emotion Engines"

[1624] (Claim 1)

[1625] A means of capturing images of the inside of the elevator in real time using a camera,

[1626] means for pre-processing the acquired video data;

[1627] A means to input the preprocessed video data into a generative AI model and quantify the degree of congestion inside the elevator,

[1628] The same video data is input into an emotion engine to recognize passenger emotions.

[1629] A means for integrating the congestion data obtained from the generative AI model and the emotion data obtained from the emotion engine to determine the elevator's stopping floor;

[1630] A system including:

[1631] (Claim 2)

[1632] 10. The system of claim 1, further comprising means for calculating a congestion level based on the number and placement of passengers detected from the video data using a generative AI model.

[1633] (Claim 3)

[1634] 10. The system of claim 1, further comprising: means for determining, based on the integrated data, an optimal elevator stop when congestion exceeds a threshold and discomfort is detected.

[1635] "Application example 2 when combining emotion engines"

[1636] (Claim 1)

[1637] A means of capturing images of the inside of the elevator in real time using a camera,

[1638] The acquired video data is input into a generative AI model to quantify the degree of congestion inside the elevator.

[1639] a means for determining a floor at which the elevator should stop when the congestion degree value exceeds a preset threshold;

[1640] means for inputting video data into an emotion engine for recognizing passenger emotions;

[1641] a means for integrating the emotion scores recognized by the emotion engine with the congestion data;

[1642] a means for optimizing elevator stops based on the integrated data;

[1643] A system including:

[1644] (Claim 2)

[1645] 10. The system of claim 1, further comprising means for pre-processing video data acquired from the camera.

[1646] (Claim 3)

[1647] 10. The system of claim 1, further comprising means for calculating a congestion level based on the number and placement of passengers detected from the video data using a generative AI model. [Explanation of symbols]

[1648] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Device 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robot< / url:> < / url:> < / url:> < / url:>

Claims

1. A means of capturing images of the inside of the elevator in real time using a camera, The acquired video data is input into a generative AI model to quantify the degree of congestion inside the elevator. a means for determining a floor at which the elevator should stop when the congestion degree value exceeds a preset threshold; A system including:

2. The system of claim 1 further comprising means for pre-processing video data acquired from the camera.

3. The system of claim 1 , further comprising means for calculating a congestion level based on the number and placement of passengers detected from the video data using a generative AI model.

Citation Information

Patent Citations

  • Persona chatbot control method and system

    JP2022180282A