system

A system using a generative model to analyze and tag logistics operations in real-time video footage addresses the inefficiencies of traditional methods, enabling rapid identification and correction of logistics issues.

JP2026070241APending Publication Date: 2026-04-27SOFTBANK GROUP CORP
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
SOFTBANK GROUP CORP
Filing Date
2024-10-15
Publication Date
2026-04-27

AI Technical Summary

Technical Problem

Existing methods for identifying and extracting relevant information from vast amounts of video footage of logistics operations are time-consuming and labor-intensive, making it difficult to quickly address issues such as packaging errors and shipping delays.

Method used

A system that utilizes a generative model to analyze video footage in real-time, automatically tag logistics operations with order identification information, and provide relevant video fragments upon user search requests, enhancing the efficiency of logistics operations by allowing quick identification and resolution of problems.

Benefits of technology

The system significantly improves quality control and operational efficiency by enabling rapid extraction and review of logistics operations, allowing users to identify and correct issues promptly.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026070241000001_ABST
    Figure 2026070241000001_ABST
Patent Text Reader

Abstract

We provide the system. [Solution] Video footage of logistics operations activities obtained based on order identification information, A means for automatically extracting relevant video fragments by analyzing them using a generative model, Means for storing video fragments related to the corresponding order identification information in a storage device, A means for providing video fragments related to order identification information based on a search request from a user, A system that includes this.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0005] , ,

[0001] The technology of the present disclosure relates to a system.

Background Art

[0002] Patent Document 1 discloses a method for controlling a persona chatbot, which is performed by at least one processor, the method including the steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to an explanation of a character of the chatbot, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance that responds to the user utterance.

Prior Art Documents

Patent Documents

[0006] "Order identification information" is a unique ID associated with a specific logistics or warehouse management instruction, and is information used to track the associated transaction or process.

[0007] "Logistics operations" refers to a series of activities in warehouses and distribution centers, including handling, moving, packaging, and shipping goods.

[0008] "Video footage" refers to visual data recorded by cameras, specifically digital video footage used to monitor and record logistics operations.

[0009] A "generative model" is a collection of machine learning algorithms that use AI technology to analyze data and identify specific patterns or objects.

[0010] A "video clip" is a portion of a video document that corresponds to a specific scene or time period.

[0011] A "storage device" is a hardware device for storing and managing digital data, and is a device that has the function of accumulating information in electronic form.

[0012] A "search request" is a query or inquiry made by a user to a system in order to retrieve specific information.

[0013] "To provide" means to deliver necessary information or data to the user in an appropriate format, and to present the requested results. [Brief explanation of the drawing]

[0014] [Figure 1] This is a conceptual diagram showing an example of the configuration of a data processing system according to the first embodiment. [Figure 2] This is a conceptual diagram showing an example of the essential functions of a data processing device and a smart device according to the first embodiment. [Figure 3] This is a conceptual diagram showing an example of the configuration of a data processing system according to the second embodiment. [Figure 4] This is a conceptual diagram showing an example of the main functions of a data processing device and smart glasses according to the second embodiment. [Figure 5] This is a conceptual diagram showing an example of the configuration of a data processing system according to the third embodiment. [Figure 6] This is a conceptual diagram showing an example of the main functions of a data processing device and a headset-type terminal according to the third embodiment. [Figure 7] This is a conceptual diagram showing an example of the configuration of a data processing system according to the fourth embodiment. [Figure 8] This is a conceptual diagram showing an example of the main functions of a data processing device and a robot according to the fourth embodiment. [Figure 9] This shows an emotion map where multiple emotions are mapped. [Figure 10] This shows an emotion map where multiple emotions are mapped. [Figure 11] This is a sequence diagram showing the processing flow of the data processing system in Example 1. [Figure 12] This is a sequence diagram showing the processing flow of the data processing system in Application Example 1. [Figure 13] This is a sequence diagram showing the processing flow of the data processing system in Example 2, which incorporates an emotion engine. [Figure 14] This is a sequence diagram showing the processing flow of the data processing system in Application Example 2, which combines an emotion engine. [Modes for carrying out the invention]

[0015] Hereinafter, an example of an embodiment of a system according to the technology of the present disclosure will be described with reference to the accompanying drawings.

[0016] First, the terms used in the following description will be explained.

[0017] In the following embodiments, a numbered processor (hereinafter simply referred to as "processor") may be a single arithmetic unit or a combination of multiple arithmetic units. Also, the processor may be a single type of arithmetic unit or a combination of multiple types of arithmetic units. Examples of arithmetic units include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), an APU (Accelerated Processing Unit), and the like.

[0018] In the following embodiments, a numbered RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a work memory by the processor.

[0019] In the following embodiments, a numbered storage is one or more non-volatile storage devices that store various programs and various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), or magnetic tapes, and the like.

[0020] In the following embodiments, the signed communication interface (I / F) is an interface that includes a communication processor and an antenna, etc. The communication interface manages communication between multiple computers. Examples of communication standards applicable to the communication interface include wireless communication standards such as 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), or Bluetooth (registered trademark).

[0021] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." That is, "A and / or B" means that it may be A alone, or B alone, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" applies when expressing three or more things linked by "and / or."

[0022] [First Embodiment]

[0023] Figure 1 shows an example of the configuration of the data processing system 10 according to the first embodiment.

[0024] As shown in Figure 1, the data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.

[0025] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0026] The smart device 14 comprises a computer 36, a reception device 38, an output device 40, a camera 42, and a communication interface 44. The computer 36 comprises a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The reception device 38, output device 40, and camera 42 are also connected to the bus 52.

[0027] The reception device 38 is equipped with a touch panel 38A and a microphone 38B, etc., and receives user input. The touch panel 38A receives user input by detecting contact with an object (e.g., a pen or finger). The microphone 38B receives user input by detecting the user's voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.

[0028] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form perceptible to the user 20 (e.g., audio and / or text). The display 40A displays visible information such as text and images according to instructions from the processor 46. The speaker 40B outputs audio according to instructions from the processor 46. The camera 42 is a small digital camera equipped with an optical system such as a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.

[0029] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various types of information between processor 46 and processor 28 via network 54.

[0030] Figure 2 shows an example of the main functions of the data processing device 12 and the smart device 14.

[0031] As shown in Figure 2, in the data processing device 12, a specific processing is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" related to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.

[0032] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0033] In the smart device 14, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The reception output program 60 is used in conjunction with a specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0034] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".

[0035] As an embodiment of this invention, an information processing system consisting mainly of a server, a terminal, and a user is constructed. This system has the function of receiving video material obtained from surveillance cameras used in logistics or warehouses, and extracting and managing the corresponding video fragments based on order identification information.

[0036] Server Role

[0037] The server receives video footage in real time from multiple cameras installed at the logistics site, efficiently compresses it, and stores it in storage. Equipped with a generative model, it analyzes the stored video footage to identify logistics operations within the video and assigns related order identification information. This automatically organizes and manages video segments for each order. Furthermore, in response to user search requests, it quickly identifies the necessary video segments and provides them to the user's terminal.

[0038] User roles

[0039] Users access the system using a client terminal and submit search requests specifying order numbers and specific time periods. When a request reaches the server, the server searches its database for video clips based on the request and extracts the relevant sections. By reviewing the video provided in this response and verifying the details of logistics operations, users can quickly identify the cause of problems such as packaging errors.

[0040] Terminal role

[0041] The terminal has the role of forwarding search requests sent by the user to the server, and also has the function of playing video clips received from the server. This interface allows users to obtain the necessary information in a user-friendly environment, which helps in the rapid resolution of problems.

[0042] As a specific example, if a user searches for order number "12345," the server identifies video clips of logistics operations linked to that order from previously collected video footage and delivers them to the terminal in real time. This allows the user to immediately review the video and take necessary actions. This invention will significantly contribute to quality control and efficiency improvement in logistics operations.

[0043] The following describes the processing flow.

[0044] Step 1:

[0045] The server receives video footage in real time from cameras installed within the logistics facility. These cameras are located in each packing area, and the captured video is compressed and stored in a storage device.

[0046] Step 2:

[0047] The server analyzes the stored video footage in batch processing. During this process, it utilizes a generative model to sequentially analyze each frame within the video, recognizing logistics operations and automatically tagging each frame with corresponding order identification information.

[0048] Step 3:

[0049] The server organizes and stores tagged video frames in a database. This makes it easy to search for and extract video fragments associated with each order identification.

[0050] Step 4:

[0051] The user uses the client terminal interface to specify the order number and a specific time, and sends a search request to the server. This request is to retrieve video clips that meet specific criteria.

[0052] Step 5:

[0053] The server receives a search request from the user and quickly performs a search on the video frames stored in the database. It identifies the relevant video fragments and prepares the necessary data.

[0054] Step 6:

[0055] The server sends the identified video clips to the user's device in a streaming or downloadable format, allowing the user to quickly access the information they need.

[0056] Step 7:

[0057] Users can play back video clips received on their devices to review logistics operations. This allows them to identify, for example, the cause of a packing error and take necessary corrective actions.

[0058] (Example 1)

[0059] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."

[0060] In logistics operations, efficiently monitoring work status and preventing misdeliveries and delays requires the management of visual information using monitoring devices and the identification of appropriate operational activities. However, conventional methods have the challenge of not being able to quickly extract necessary information from vast amounts of data.

[0061] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0062] In this invention, the server includes means for receiving visual information acquired from multiple monitoring devices in real time, effectively compressing it, and storing it in a storage device; means for analyzing the received visual information using a generation model to identify logistics operation activities and automatically assign related identification information; and means for quickly searching for and providing visual information fragments organized based on the identification information in response to usage requests. This makes it possible to accurately grasp the work situation at the logistics site and quickly extract relevant information.

[0063] A "monitoring device" is a device installed in logistics sites or related locations to acquire visual information.

[0064] "Visual information" refers to all video data acquired by monitoring devices and is used for monitoring and analyzing logistics operations.

[0065] A "generative model" refers to an algorithm or learning model used to analyze received visual information and recognize and classify specific patterns or activities.

[0066] "Identification information" refers to labels or codes assigned to associate specific logistics operations or orders, and is used to organize and manage information.

[0067] "Storage device" refers to the storage medium or storage system that a server uses to save and store data.

[0068] A "request for use" refers to a request from a user for searching or retrieving data, which acts as a trigger for the server to provide information based on that request.

[0069] The following describes embodiments for carrying out the invention.

[0070] This invention is a system that efficiently manages and analyzes various logistics operations based on visual information from monitoring devices installed at logistics sites. The server receives visual information in real time from multiple monitoring devices and stores it in a storage device using image compression technologies such as H.264 and HEVC while maintaining high quality. This enables the operation of a massive amount of data.

[0071] The server is equipped with a generative AI model that uses deep learning algorithms to analyze visual information. Through this analysis, the server identifies logistics operations and automatically assigns relevant identification information, thereby properly organizing and managing the visual information fragments.

[0072] Users access the server via a client terminal and make search requests based on visual information using prompt messages. For example, when a user enters a prompt message such as "Search for video with order number 12345," the server immediately retrieves the corresponding visual information fragment and provides it to the user terminal.

[0073] The terminal receives and plays back visual information fragments provided by the server, creating an environment where users can smoothly verify logistics operations. This allows users to quickly identify problems such as packing errors or shipping delays based on the received video and take necessary corrective measures. This system contributes to improving quality control and operational efficiency in logistics operations.

[0074] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0075] Step 1:

[0076] The server receives visual information from the monitoring device in real time. At this time, the received data (input) is a raw video stream. The server compresses this video stream using compression technologies such as H.264 or HEVC to reduce the data size (data processing). After that, it saves the compressed visual information (output) to the storage device.

[0077] Step 2:

[0078] The server inputs the stored visual information into a generating AI model and performs data analysis. This model analyzes the received compressed video data (input) and identifies logistics operation activities (data computation). Specifically, it uses a deep learning algorithm to identify objects and actions within the video. As a result of the analysis, identification information corresponding to specific activities within the video (output) is generated and attached to the visual information fragment.

[0079] Step 3:

[0080] The user sends a search request to the server through their client terminal. This request consists of a prompt statement (input), for example, "Search for video with order number 12345." Based on this request, the server refers to the identification information and searches for and retrieves relevant visual information fragments from the database (data processing). The retrieved visual information fragments (output) are then sent to the user's terminal.

[0081] Step 4:

[0082] The user plays back the visual information fragments received on the terminal to confirm logistics operations. Through this playback, the user analyzes the specific activity details obtained from the video, identifying, for example, packaging defects or shipping errors (data calculation and situation confirmation). Based on this, the user can immediately take necessary corrective measures.

[0083] (Application Example 1)

[0084] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."

[0085] To improve operational efficiency in logistics centers and to quickly verify orders and troubleshoot problems, a system is needed that allows on-site staff to obtain necessary information in a short amount of time. However, traditional methods for identifying specific operational activities linked to particular orders from vast amounts of video data are time-consuming and labor-intensive, making immediate countermeasures difficult.

[0086] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0087] In this invention, the server includes means for analyzing video footage of logistics operations acquired based on order identification information and automatically extracting the corresponding video fragments; means for storing the video fragments related to the corresponding order identification information on an electronic storage medium; and means for acquiring order identification information using a voice input device and immediately displaying the related video. This makes it possible for on-site staff to acquire necessary order information by voice using a smart device and immediately confirm its contents.

[0088] "Order identification information" refers to a unique identifier used to identify each order in logistics.

[0089] "Logistics operations" refers to a series of business processes within a logistics center, including the movement, picking, packing, and shipping of goods and packages.

[0090] "Video footage" refers to video data recorded by surveillance cameras, etc., which visually captures the details of logistics operations.

[0091] A "generative model" is a system that uses machine learning algorithms to analyze video materials and recognize and classify specific objects.

[0092] "Electronic storage media" refers to devices or materials used to store digital data, and includes hard disks, SSDs, and cloud storage.

[0093] A "voice input device" is a device that captures the user's voice and sends it to the system as a voice command.

[0094] A "video clip" is a portion of a video recording of a logistics operation related to specific order identification information.

[0095] "On-site staff" refers to workers and personnel who perform their duties on a daily basis at a logistics center.

[0096] As an embodiment of this invention, a video management system for a logistics center is configured. The server acquires video data in real time from surveillance cameras installed in the logistics center, analyzes it using a generating AI model to identify logistics operation activities, and extracts video fragments based on the respective order identification information. These video fragments are stored on an electronic storage medium and provided to the user as needed.

[0097] The terminal is designed to allow staff to search for specific order numbers by voice using voice input devices such as smart glasses. The voice input is converted into a digital signal and sent to a server. The server analyzes this voice signal, quickly identifies the relevant video clips, and provides them to the terminal in streaming or download format.

[0098] As a concrete example, when a field staff member wears smart glasses and issues the voice command "Search for order 12345," the server identifies the video footage associated with order identification information "12345" and immediately displays it in the staff member's field of view. This function allows staff members to confirm the details of logistics operations without delay and take appropriate action.

[0099] By using a generative AI model, it is possible to efficiently extract relevant logistics activities from video footage and provide videos based on specific orders. This system significantly improves operational efficiency within logistics centers.

[0100] An example of a prompt message would be, "Please display the latest logistics operation video for order 12345." This operation is performed using voice recognition technology via smart glasses.

[0101] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0102] Step 1:

[0103] The server receives real-time video data from surveillance cameras in the logistics center. It receives continuous video streams from each camera as input and analyzes them using a generative AI model. As output, it identifies logistics operations within the video and generates video data labeled based on corresponding order identification information. Specifically, it executes video analysis algorithms to perform object recognition and motion detection.

[0104] Step 2:

[0105] The server stores the analyzed video data on an electronic storage medium. It receives labeled video data as input and converts it into a format for storage in the database. As output, the video data is categorized by order identification information and stored in storage in a searchable format. Specifically, it creates an index to organize and store video segments and their associated metadata.

[0106] Step 3:

[0107] The user searches for a specific order number by issuing a voice command using smart glasses. The input is voice data collected by a voice input device. The output is the order number obtained by converting the voice signal into text. Specifically, voice recognition software is used to interpret the voice command and identify the order number.

[0108] Step 4:

[0109] The server searches the database for relevant video data based on the order number received from the user. The order number, in text format, is used as input, and a query is generated to search for the corresponding video data. The output extracts data for the relevant video segments. Specifically, the system executes a database query to quickly retrieve the relevant video data.

[0110] Step 5:

[0111] The terminal receives video fragments from the server and transmits them to the user's smart glasses in streaming format. The input is video data from the server. The output is the video being visually displayed to the staff. Specifically, the operation involves transferring video data via a network connection and displaying the video using the smart glasses' video playback function.

[0112] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0113] To implement this invention, a system comprising a server, a terminal, and a user is required. This system includes a function to analyze and provide video of logistics operations based on order identification information, as well as an emotion engine that recognizes the user's emotions.

[0114] Server Role

[0115] The server receives real-time video footage from cameras within the logistics facility and analyzes it using a generative model. The generative model identifies logistics operations in each frame and tags them with relevant order identification information. This process automatically organizes and stores the video fragments in a database.

[0116] In addition, the server analyzes the user's emotions through an emotion engine. This enables the provision of optimal video segments tailored to the user's state and the application of a customized user interface.

[0117] User roles

[0118] When a user accesses the system from a client terminal and makes a search request specifying an order number and time, the emotion engine analyzes their emotions and adjusts the user interface and presentation methods accordingly. This provides the user with a more comfortable and efficient operating experience.

[0119] Terminal role

[0120] The device sends the user's search request to the server and displays the video clips received from the server in a playable state. Based on information obtained from the emotion engine, the video being played and the user interface are customized to match the user's emotions.

[0121] For example, if a user expresses impatience while searching for the order "12345," the emotion engine analyzes that emotion, and the server provides the terminal with faster display speeds and additional guidelines. This allows the user to smoothly identify the problem and consider solutions, improving overall work efficiency. In this way, customization using an emotion engine can improve the user experience.

[0122] The following describes the processing flow.

[0123] Step 1:

[0124] The server receives video footage from surveillance cameras within the logistics facility in real time and efficiently stores it in its storage device. This ensures that logistics operations are continuously recorded.

[0125] Step 2:

[0126] The server analyzes the stored video using a generative model. The analysis is performed frame by frame, identifying logistics operations and related order identification information, and tagging the frame accordingly. This information is then organized in a database.

[0127] Step 3:

[0128] The user executes a search request via the system interface through a client terminal, specifying the order number and a specific time. The search request is then sent to the server.

[0129] Step 4:

[0130] The server receives a search request from the user and quickly searches the database for the relevant video clip. Furthermore, the emotion engine recognizes the user's emotions from their facial expressions and actions, and generates analysis results.

[0131] Step 5:

[0132] Based on the analysis results of the emotion engine, the server determines how to provide information according to the user's emotional state. For example, if the user is showing signs of stress, the server prepares a simplified explanation.

[0133] Step 6:

[0134] The device receives video clips transmitted from the server via streaming or download and begins playback. Simultaneously, the user interface adjusts to the user's emotions.

[0135] Step 7:

[0136] Users review the played-back footage and evaluate the logistics operations in detail. User feedback is further analyzed by an emotion engine to improve the system's responsiveness and usability.

[0137] (Example 2)

[0138] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".

[0139] There is a need to efficiently extract necessary video clips from a vast amount of video footage related to logistics operations based on specific order identification information, and present them to the user quickly and appropriately. Furthermore, there is a lack of technology to adjust search results and the user interface according to the user's emotional state, thus necessitating an improvement in the user experience.

[0140] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0141] In this invention, the server includes means for analyzing video footage of logistics operations acquired based on order identification information using a generating AI model to automatically extract relevant video fragments, means for storing the relevant video fragments related to the order identification information in a storage device, and means for adjusting the user interface during search requests and information provision using an emotion engine that analyzes the user's emotions. As a result, the user can efficiently acquire video related to a specific order and comfortably use the system with an interface customized to the user's emotions.

[0142] "Order identification information" refers to information used to identify a specific order related to a logistics operation.

[0143] "Logistics operations" refers to a series of tasks and processes related to the movement, storage, and management of goods.

[0144] A "video clip" is a portion of video footage extracted from video material based on specific information.

[0145] A "generative AI model" is a computational model that uses artificial intelligence technology to analyze given data and generate or identify specific patterns or information.

[0146] An "emotion engine" is a program that analyzes user input and behavior to estimate the user's emotional state.

[0147] A "user terminal" is an electronic device, such as a computer or smartphone, that a user uses to perform operations.

[0148] A "storage device" refers to hardware or software used to store data or information.

[0149] "Streaming" is a technology or method that transmits data instantly and plays it back in real time on the receiving end.

[0150] This invention is implemented primarily as a system consisting of a server, a terminal, and a user. In this system, the server receives video data acquired in real time from cameras installed within the logistics facility and performs analysis using a generative AI model. The generative AI model identifies logistics operation activities in each video frame and tags the video fragments with appropriate order identification information. Following this video analysis, the video fragments are automatically saved to a storage device.

[0151] The server also plays a role in analyzing the user's emotions through its emotion engine. This allows the server to provide video clips and customize the user interface based on the user's emotions when they make a search request using a client terminal, specifying an order number and time. For example, if a user searches for the order "12345" and expresses anxiety, the server will use the emotion engine's analysis to provide faster video clips and display additional guidelines.

[0152] The terminal displays video clips transmitted from the server in a playable format and adjusts the interface based on the user's emotions using an emotion engine. This adjustment allows the user to have an efficient and comfortable operating experience.

[0153] An example of a prompt message that might be input into the generating AI model is, "Search for videos of logistics activities related to order number 12345, adjust the display speed to be faster, and provide guidelines to the user." This allows the system to immediately process the appropriate information and provide customized feedback to the user.

[0154] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0155] Step 1:

[0156] The server receives video data in real time from cameras within the logistics facility. High-resolution video frames acquired from the cameras are provided to the server as input for the received video. The server temporarily stores this frame data for the next analysis step.

[0157] Step 2:

[0158] The server analyzes the received video frames using a generating AI model. In this analysis step, the AI ​​model identifies the movement of items and people within the frame and identifies logistics operations. As a result, order identification information corresponding to each frame is generated. This allows the system to understand which order a particular logistics activity is related to.

[0159] Step 3:

[0160] The server tags video frames with order identification information based on the analysis results. Identification information obtained from the AI ​​model is used as input, and tagged video frames are obtained as output. These are organized and stored in a database to enable efficient searching.

[0161] Step 4:

[0162] The server uses an emotion engine to analyze user input and actions to infer their emotional state. Here, user input data is provided to the emotion engine, and information about the user's emotions is generated as output. This information is then used in the next step to customize the user interface.

[0163] Step 5:

[0164] The user uses a client terminal to send a search request specifying a particular order number and time. This request is passed to the server in the form of prompt statements. The server processes this request and prepares to send the corresponding video clips to the terminal.

[0165] Step 6:

[0166] Based on the emotion analysis results, the server determines interface settings and video display methods appropriate for the user's emotional state. At this stage, the adjusted settings are sent to the terminal and applied to improve the user's experience.

[0167] Step 7:

[0168] The device displays video clips received from the server in a playable format. Furthermore, it can adjust the video playback speed and add guidelines based on information from the emotion engine. This allows users to view the video in a way that resonates with their own emotions.

[0169] (Application Example 2)

[0170] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as a "server" and the smart device 14 as a "terminal".

[0171] To improve operational efficiency and reduce errors in logistics centers, there is a need for methods to make operational information easier to understand. Furthermore, a system is needed that supports smooth work execution by providing support tailored to the emotions of the workers. The challenge is to improve the efficiency and safety of the entire logistics process.

[0172] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0173] In this invention, the server includes means for analyzing video information of logistics operations acquired based on order identification information using a generative model to automatically identify the corresponding video elements; means for storing the corresponding video elements related to the order identification information in a storage device; means for providing video data related to the order identification information based on a search request from the user; means for identifying the user's emotions using an emotion analysis engine and customizing the user interface; and means for presenting emotion-appropriate instruction information along with video data to a display device used in the logistics center. This makes it possible for workers to receive appropriate video information and emotion-appropriate instructions in real time, improving the efficiency and safety of logistics operations.

[0174] "Order identification information" refers to information used in logistics operations to identify a specific order.

[0175] "Logistics operations" refers to a series of tasks performed within a logistics facility, such as receiving, storing, picking, and shipping goods.

[0176] A "generative model" is an algorithm that learns patterns and features from input data and uses that data for analysis and classification.

[0177] "Video elements" refer to segments or frames related to a specific order within video materials, including logistics operations.

[0178] "Storage device" refers to hardware and software used to hold digital data, and includes databases and storage devices.

[0179] A "search request" is a request for inquiry or retrieval made by a user to obtain a specific order or information from the system.

[0180] An "emotion analysis engine" is a software engine that analyzes a user's facial expressions and behavior to determine their emotional state.

[0181] A "user interface" refers to the screen display and input devices used for interaction between a user and a computer system.

[0182] A "display device" is hardware used to provide information visually, and includes screens and displays.

[0183] "Instructional information" refers to information provided to guide operations or actions, and includes commands and guidelines.

[0184] A specific embodiment of this invention involves a system consisting of a server, a terminal, and a user. The server receives real-time video footage of delivery activities from cameras within the logistics facility. The received video is analyzed using a generative AI model. The generative model analyzes each video frame and tags relevant video elements based on order identification information. These video elements are stored in a memory device and organized for quick access when a search is needed.

[0185] The device receives relevant video elements from the server based on the user's exploration request and plays them back via streaming or download. The user's emotions are evaluated in real time using an emotion analysis engine, and the user interface is dynamically adjusted accordingly. This adjustment allows the user to use the system efficiently; for example, if the user is feeling stressed, the system automatically displays instructions in a visually easy-to-understand format.

[0186] Users can obtain real-time instruction information through a display device such as smart glasses. This helps prevent judgment errors during logistics operations and supports efficient work progress.

[0187] As a concrete example, consider a scenario where a logistics center worker processes a specific order number, "12345." If the worker experiences stress, an emotion analysis engine analyzes their state, and the server suggests improved delivery routes and alerts to ensure smoother work. This process is visually presented through the display of smart glasses.

[0188] An example of a prompt message describing the operation of this system is: "Please tell me about real-time support while working at the logistics center. How do you handle users who are experiencing stress?"

[0189] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0190] Step 1:

[0191] The server receives real-time video data from cameras placed within the logistics facility. The input is the raw video stream from the cameras, and the output is video data stored in temporary storage for analysis. The server prepares this data to be input into a generating AI model.

[0192] Step 2:

[0193] The server analyzes the received video data using a generative AI model. The input is the video data saved in step 1, and the output is order identification information and the type of logistics operation activity tagged to each frame. Based on this tagged data, the server organizes and stores the information in a database.

[0194] Step 3:

[0195] The user makes a search request through their terminal, specifying a particular order identifier. The input is the order number from the user, and the output is a request for the associated video fragment. This request is sent to the server.

[0196] Step 4:

[0197] The server retrieves the relevant video fragment from the database based on the user's search request. The input is the user's order number, and the output is the corresponding video fragment. The server then prepares to send this data to the terminal in streaming or download format.

[0198] Step 5:

[0199] The terminal plays video clips received from the server to the user. The input is the video clips from the server, and the output is the video displayed on the terminal. This video provides the user with necessary information for their work and helps improve efficiency.

[0200] Step 6:

[0201] The server uses an emotion analysis engine to evaluate the user's emotions in real time. Input is the user's facial image and biometric data, and output is the analyzed emotion information. Based on this information, the user interface is automatically adjusted to provide optimal support tailored to the user's situation.

[0202] Step 7:

[0203] Users receive improved delivery routes and alerts as needed through the display device. Input is instructional information from the server, and output is visual instructional information presented on the display device. This allows users to obtain information that helps improve their work in real time.

[0204] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0205] Data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of data generation model 58 is ChatGPT (registered trademark) (Internet search).<URL: https: / / openai.com / blog / chatgpt> ), Gemini (registered trademark) (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0206] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart device 14.

[0207] [Second Embodiment]

[0208] Figure 3 shows an example of the configuration of the data processing system 210 according to the second embodiment.

[0209] As shown in Figure 3, the data processing system 210 includes a data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.

[0210] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0211] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication interface 44. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, and camera 42 are also connected to the bus 52.

[0212] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0213] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[0214] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0215] Figure 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Figure 4, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0216] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0217] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0218] In the smart glasses 214, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0219] Next, the identification processing performed by the identification processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".

[0220] As an embodiment of this invention, an information processing system consisting mainly of a server, a terminal, and a user is constructed. This system has the function of receiving video material obtained from surveillance cameras used in logistics or warehouses, and extracting and managing the corresponding video fragments based on order identification information.

[0221] Server Role

[0222] The server receives video footage in real time from multiple cameras installed at the logistics site, efficiently compresses it, and stores it in storage. Equipped with a generative model, it analyzes the stored video footage to identify logistics operations within the video and assigns related order identification information. This automatically organizes and manages video segments for each order. Furthermore, in response to user search requests, it quickly identifies the necessary video segments and provides them to the user's terminal.

[0223] User roles

[0224] Users access the system using a client terminal and submit search requests specifying order numbers and specific time periods. When a request reaches the server, the server searches its database for video clips based on the request and extracts the relevant sections. By reviewing the video provided in this response and verifying the details of logistics operations, users can quickly identify the cause of problems such as packaging errors.

[0225] Terminal role

[0226] The terminal has the role of forwarding search requests sent by the user to the server, and also has the function of playing video clips received from the server. This interface allows users to obtain the necessary information in a user-friendly environment, which helps in the rapid resolution of problems.

[0227] As a specific example, if a user searches for order number "12345," the server identifies video clips of logistics operations linked to that order from previously collected video footage and delivers them to the terminal in real time. This allows the user to immediately review the video and take necessary actions. This invention will significantly contribute to quality control and efficiency improvement in logistics operations.

[0228] The following describes the processing flow.

[0229] Step 1:

[0230] The server receives video footage in real time from cameras installed within the logistics facility. These cameras are located in each packing area, and the captured video is compressed and stored in a storage device.

[0231] Step 2:

[0232] The server analyzes the stored video footage in batch processing. During this process, it utilizes a generative model to sequentially analyze each frame within the video, recognizing logistics operations and automatically tagging each frame with corresponding order identification information.

[0233] Step 3:

[0234] The server organizes and stores tagged video frames in a database. This makes it easy to search for and extract video fragments associated with each order identification.

[0235] Step 4:

[0236] The user uses the client terminal interface to specify the order number and a specific time, and sends a search request to the server. This request is to retrieve video clips that meet specific criteria.

[0237] Step 5:

[0238] The server receives a search request from the user and quickly performs a search on the video frames stored in the database. It identifies the relevant video fragments and prepares the necessary data.

[0239] Step 6:

[0240] The server sends the identified video clips to the user's device in a streaming or downloadable format, allowing the user to quickly access the information they need.

[0241] Step 7:

[0242] Users can play back video clips received on their devices to review logistics operations. This allows them to identify, for example, the cause of a packing error and take necessary corrective actions.

[0243] (Example 1)

[0244] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."

[0245] In logistics operations, efficiently monitoring work status and preventing misdeliveries and delays requires the management of visual information using monitoring devices and the identification of appropriate operational activities. However, conventional methods have the challenge of not being able to quickly extract necessary information from vast amounts of data.

[0246] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0247] In this invention, the server includes means for receiving visual information acquired from multiple monitoring devices in real time, effectively compressing it, and storing it in a storage device; means for analyzing the received visual information using a generation model to identify logistics operation activities and automatically assign related identification information; and means for quickly searching for and providing visual information fragments organized based on the identification information in response to usage requests. This makes it possible to accurately grasp the work situation at the logistics site and quickly extract relevant information.

[0248] A "monitoring device" is a device installed in logistics sites or related locations to acquire visual information.

[0249] "Visual information" refers to all video data acquired by monitoring devices and is used for monitoring and analyzing logistics operations.

[0250] A "generative model" refers to an algorithm or learning model used to analyze received visual information and recognize and classify specific patterns or activities.

[0251] "Identification information" refers to labels or codes assigned to associate specific logistics operations or orders, and is used to organize and manage information.

[0252] "Storage device" refers to the storage medium or storage system that a server uses to save and store data.

[0253] A "request for use" refers to a request from a user for searching or retrieving data, which acts as a trigger for the server to provide information based on that request.

[0254] The following describes embodiments for carrying out the invention.

[0255] This invention is a system that efficiently manages and analyzes various logistics operations based on visual information from monitoring devices installed at logistics sites. The server receives visual information in real time from multiple monitoring devices and stores it in a storage device using image compression technologies such as H.264 and HEVC while maintaining high quality. This enables the operation of a massive amount of data.

[0256] The server is equipped with a generative AI model that uses deep learning algorithms to analyze visual information. Through this analysis, the server identifies logistics operations and automatically assigns relevant identification information, thereby properly organizing and managing the visual information fragments.

[0257] Users access the server via a client terminal and make search requests based on visual information using prompt messages. For example, when a user enters a prompt message such as "Search for video with order number 12345," the server immediately retrieves the corresponding visual information fragment and provides it to the user terminal.

[0258] The terminal receives and plays back visual information fragments provided by the server, creating an environment where users can smoothly verify logistics operations. This allows users to quickly identify problems such as packing errors or shipping delays based on the received video and take necessary corrective measures. This system contributes to improving quality control and operational efficiency in logistics operations.

[0259] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0260] Step 1:

[0261] The server receives visual information from the monitoring device in real time. At this time, the received data (input) is a raw video stream. The server compresses this video stream using compression technologies such as H.264 or HEVC to reduce the data size (data processing). After that, it saves the compressed visual information (output) to the storage device.

[0262] Step 2:

[0263] The server inputs the stored visual information into a generating AI model and performs data analysis. This model analyzes the received compressed video data (input) and identifies logistics operation activities (data computation). Specifically, it uses a deep learning algorithm to identify objects and actions within the video. As a result of the analysis, identification information corresponding to specific activities within the video (output) is generated and attached to the visual information fragment.

[0264] Step 3:

[0265] The user sends a search request to the server through their client terminal. This request consists of a prompt statement (input), for example, "Search for video with order number 12345." Based on this request, the server refers to the identification information and searches for and retrieves relevant visual information fragments from the database (data processing). The retrieved visual information fragments (output) are then sent to the user's terminal.

[0266] Step 4:

[0267] The user plays back the visual information fragments received on the terminal to confirm logistics operations. Through this playback, the user analyzes the specific activity details obtained from the video, identifying, for example, packaging defects or shipping errors (data calculation and situation confirmation). Based on this, the user can immediately take necessary corrective measures.

[0268] (Application Example 1)

[0269] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."

[0270] To improve operational efficiency in logistics centers and to quickly verify orders and troubleshoot problems, a system is needed that allows on-site staff to obtain necessary information in a short amount of time. However, traditional methods for identifying specific operational activities linked to particular orders from vast amounts of video data are time-consuming and labor-intensive, making immediate countermeasures difficult.

[0271] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0272] In this invention, the server includes means for analyzing video footage of logistics operations acquired based on order identification information and automatically extracting the corresponding video fragments; means for storing the video fragments related to the corresponding order identification information on an electronic storage medium; and means for acquiring order identification information using a voice input device and immediately displaying the related video. This makes it possible for on-site staff to acquire necessary order information by voice using a smart device and immediately confirm its contents.

[0273] "Order identification information" refers to a unique identifier used to identify each order in logistics.

[0274] "Logistics operations" refers to a series of business processes within a logistics center, including the movement, picking, packing, and shipping of goods and packages.

[0275] "Video footage" refers to video data recorded by surveillance cameras, etc., which visually captures the details of logistics operations.

[0276] A "generative model" is a system that uses machine learning algorithms to analyze video materials and recognize and classify specific objects.

[0277] "Electronic storage media" refers to devices or materials used to store digital data, and includes hard disks, SSDs, and cloud storage.

[0278] A "voice input device" is a device that captures the user's voice and sends it to the system as a voice command.

[0279] A "video clip" is a portion of a video recording of a logistics operation related to specific order identification information.

[0280] "On-site staff" refers to workers and personnel who perform their duties on a daily basis at a logistics center.

[0281] As an embodiment of this invention, a video management system for a logistics center is configured. The server acquires video data in real time from surveillance cameras installed in the logistics center, analyzes it using a generating AI model to identify logistics operation activities, and extracts video fragments based on the respective order identification information. These video fragments are stored on an electronic storage medium and provided to the user as needed.

[0282] The terminal is designed so that staff can search for a specific order number by voice using a voice input device such as smart glasses. The voice input is converted into a digital signal and sent to the server. The server analyzes this voice signal, quickly identifies the relevant video clip, and provides it to the terminal in the form of streaming or downloading.

[0283] As a specific example, when on-site staff wears smart glasses and issues a voice command "Search for order 12345", the server identifies the video clip related to the order identification information "12345" and immediately displays the video in the staff's field of vision. With this function, staff can check the content of logistics operations without delay and take appropriate actions.

[0284] By using a generative AI model, relevant logistics activities can be efficiently extracted from video materials, and video can be provided based on a specific order. This system significantly improves the business efficiency within the logistics center.

[0285] As an example of the prompt text, a form such as "Please display the latest logistics operation video related to order 12345" can be considered. This operation is performed from smart glasses using voice recognition technology.

[0286] The flow of the specific process in Application Example 1 will be described using FIG. 12.

[0287] Step 1:

[0288] The server receives video data in real time from the surveillance cameras in the logistics center. It receives the video streams continuously sent from each camera as input and analyzes it using a generative AI model. As output, it generates video data with labels based on the corresponding order identification information, identifying the logistics operation activities within the video. As a specific operation, it executes a video analysis algorithm to perform object recognition and motion detection.

[0289] Step 2:

[0290] The server stores the analyzed video data on an electronic storage medium. It receives labeled video data as input and converts it into a format for storage in the database. As output, the video data is categorized by order identification information and stored in storage in a searchable format. Specifically, it creates an index to organize and store video segments and their associated metadata.

[0291] Step 3:

[0292] The user searches for a specific order number by issuing a voice command using smart glasses. The input is voice data collected by a voice input device. The output is the order number obtained by converting the voice signal into text. Specifically, voice recognition software is used to interpret the voice command and identify the order number.

[0293] Step 4:

[0294] The server searches the database for relevant video data based on the order number received from the user. The order number, in text format, is used as input, and a query is generated to search for the corresponding video data. The output extracts data for the relevant video segments. Specifically, the system executes a database query to quickly retrieve the relevant video data.

[0295] Step 5:

[0296] The terminal receives video fragments from the server and transmits them to the user's smart glasses in streaming format. The input is video data from the server. The output is the video being visually displayed to the staff. Specifically, the operation involves transferring video data via a network connection and displaying the video using the smart glasses' video playback function.

[0297] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0298] To implement this invention, a system comprising a server, a terminal, and a user is required. This system includes a function to analyze and provide video of logistics operations based on order identification information, as well as an emotion engine that recognizes the user's emotions.

[0299] Server Role

[0300] The server receives real-time video footage from cameras within the logistics facility and analyzes it using a generative model. The generative model identifies logistics operations in each frame and tags them with relevant order identification information. This process automatically organizes and stores the video fragments in a database.

[0301] In addition, the server analyzes the user's emotions through an emotion engine. This enables the provision of optimal video segments tailored to the user's state and the application of a customized user interface.

[0302] User roles

[0303] When a user accesses the system from a client terminal and makes a search request specifying an order number and time, the emotion engine analyzes their emotions and adjusts the user interface and presentation methods accordingly. This provides the user with a more comfortable and efficient operating experience.

[0304] Terminal role

[0305] The device sends the user's search request to the server and displays the video clips received from the server in a playable state. Based on information obtained from the emotion engine, the video being played and the user interface are customized to match the user's emotions.

[0306] As a specific example, when the user shows impatience while searching for an order of "12345", the emotion engine analyzes the emotion, and the server provides a faster display speed and additional guidelines to the terminal. As a result, the user can smoothly identify the problem and consider countermeasures, improving the overall work efficiency. Thus, customization using the emotion engine can improve the user experience.

[0307] The following describes the processing flow.

[0308] Step 1:

[0309] The server receives video materials from surveillance cameras in the logistics facility in real time and efficiently stores them in the storage device. As a result, the state of the logistics operation activities is continuously recorded.

[0310] Step 2:

[0311] The server analyzes the stored video using a generation model. The analysis is performed for each frame, and the logistics operation activities and related order identification information are identified and tagged to the frame. This information is organized in the database.

[0312] Step 3:

[0313] The user executes a search request by specifying an order number or a specific time through the interface of the system via the client terminal. The search request is sent to the server.

[0314] Step 4:

[0315] The server receives the search request from the user and quickly searches the database for the corresponding video clip. Furthermore, the emotion engine recognizes the emotion from the user's expression and movement, and an analysis result is generated.

[0316] Step 5:

[0317] Based on the analysis results of the emotion engine, the server determines how to provide information according to the user's emotional state. For example, if the user is showing signs of stress, the server prepares a simplified explanation.

[0318] Step 6:

[0319] The device receives video clips transmitted from the server via streaming or download and begins playback. Simultaneously, the user interface adjusts to the user's emotions.

[0320] Step 7:

[0321] Users review the played-back footage and evaluate the logistics operations in detail. User feedback is further analyzed by an emotion engine to improve the system's responsiveness and usability.

[0322] (Example 2)

[0323] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".

[0324] There is a need to efficiently extract necessary video clips from a vast amount of video footage related to logistics operations based on specific order identification information, and present them to the user quickly and appropriately. Furthermore, there is a lack of technology to adjust search results and the user interface according to the user's emotional state, thus necessitating an improvement in the user experience.

[0325] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0326] In this invention, the server includes means for analyzing video footage of logistics operations acquired based on order identification information using a generating AI model to automatically extract relevant video fragments, means for storing the relevant video fragments related to the order identification information in a storage device, and means for adjusting the user interface during search requests and information provision using an emotion engine that analyzes the user's emotions. As a result, the user can efficiently acquire video related to a specific order and comfortably use the system with an interface customized to the user's emotions.

[0327] "Order identification information" refers to information used to identify a specific order related to a logistics operation.

[0328] "Logistics operations" refers to a series of tasks and processes related to the movement, storage, and management of goods.

[0329] A "video clip" is a portion of video footage extracted from video material based on specific information.

[0330] A "generative AI model" is a computational model that uses artificial intelligence technology to analyze given data and generate or identify specific patterns or information.

[0331] An "emotion engine" is a program that analyzes user input and behavior to estimate the user's emotional state.

[0332] A "user terminal" is an electronic device, such as a computer or smartphone, that a user uses to perform operations.

[0333] A "storage device" refers to hardware or software used to store data or information.

[0334] "Streaming" is a technology or method that transmits data instantly and plays it back in real time on the receiving end.

[0335] This invention is implemented primarily as a system consisting of a server, a terminal, and a user. In this system, the server receives video data acquired in real time from cameras installed within the logistics facility and performs analysis using a generative AI model. The generative AI model identifies logistics operation activities in each video frame and tags the video fragments with appropriate order identification information. Following this video analysis, the video fragments are automatically saved to a storage device.

[0336] The server also plays a role in analyzing the user's emotions through its emotion engine. This allows the server to provide video clips and customize the user interface based on the user's emotions when they make a search request using a client terminal, specifying an order number and time. For example, if a user searches for the order "12345" and expresses anxiety, the server will use the emotion engine's analysis to provide faster video clips and display additional guidelines.

[0337] The terminal displays video clips transmitted from the server in a playable format and adjusts the interface based on the user's emotions using an emotion engine. This adjustment allows the user to have an efficient and comfortable operating experience.

[0338] An example of a prompt message that might be input into the generating AI model is, "Search for videos of logistics activities related to order number 12345, adjust the display speed to be faster, and provide guidelines to the user." This allows the system to immediately process the appropriate information and provide customized feedback to the user.

[0339] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0340] Step 1:

[0341] The server receives video data in real time from cameras within the logistics facility. High-resolution video frames acquired from the cameras are provided to the server as input for the received video. The server temporarily stores this frame data for the next analysis step.

[0342] Step 2:

[0343] The server analyzes the received video frames using a generating AI model. In this analysis step, the AI ​​model identifies the movement of items and people within the frame and identifies logistics operations. As a result, order identification information corresponding to each frame is generated. This allows the system to understand which order a particular logistics activity is related to.

[0344] Step 3:

[0345] The server tags video frames with order identification information based on the analysis results. Identification information obtained from the AI ​​model is used as input, and tagged video frames are obtained as output. These are organized and stored in a database to enable efficient searching.

[0346] Step 4:

[0347] The server uses an emotion engine to analyze user input and actions to infer their emotional state. Here, user input data is provided to the emotion engine, and information about the user's emotions is generated as output. This information is then used in the next step to customize the user interface.

[0348] Step 5:

[0349] The user uses a client terminal to send a search request specifying a particular order number and time. This request is passed to the server in the form of prompt statements. The server processes this request and prepares to send the corresponding video clips to the terminal.

[0350] Step 6:

[0351] Based on the emotion analysis results, the server determines interface settings and video display methods appropriate for the user's emotional state. At this stage, the adjusted settings are sent to the terminal and applied to improve the user's experience.

[0352] Step 7:

[0353] The device displays video clips received from the server in a playable format. Furthermore, it can adjust the video playback speed and add guidelines based on information from the emotion engine. This allows users to view the video in a way that resonates with their own emotions.

[0354] (Application Example 2)

[0355] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."

[0356] To improve operational efficiency and reduce errors in logistics centers, there is a need for methods to make operational information easier to understand. Furthermore, a system is needed that supports smooth work execution by providing support tailored to the emotions of the workers. The challenge is to improve the efficiency and safety of the entire logistics process.

[0357] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0358] In this invention, the server includes means for analyzing video information of logistics operations acquired based on order identification information using a generative model to automatically identify the corresponding video elements; means for storing the corresponding video elements related to the order identification information in a storage device; means for providing video data related to the order identification information based on a search request from the user; means for identifying the user's emotions using an emotion analysis engine and customizing the user interface; and means for presenting emotion-appropriate instruction information along with video data to a display device used in the logistics center. This makes it possible for workers to receive appropriate video information and emotion-appropriate instructions in real time, improving the efficiency and safety of logistics operations.

[0359] "Order identification information" refers to information used in logistics operations to identify a specific order.

[0360] "Logistics operations" refers to a series of tasks performed within a logistics facility, such as receiving, storing, picking, and shipping goods.

[0361] A "generative model" is an algorithm that learns patterns and features from input data and uses that data for analysis and classification.

[0362] "Video elements" refer to segments or frames related to a specific order within video materials, including logistics operations.

[0363] "Storage device" refers to hardware and software used to hold digital data, and includes databases and storage devices.

[0364] A "search request" is a request for inquiry or retrieval made by a user to obtain a specific order or information from the system.

[0365] An "emotion analysis engine" is a software engine that analyzes a user's facial expressions and behavior to determine their emotional state.

[0366] A "user interface" refers to the screen display and input devices used for interaction between a user and a computer system.

[0367] A "display device" is hardware used to provide information visually, and includes screens and displays.

[0368] "Instructional information" refers to information provided to guide operations or actions, and includes commands and guidelines.

[0369] A specific embodiment of this invention involves a system consisting of a server, a terminal, and a user. The server receives real-time video footage of delivery activities from cameras within the logistics facility. The received video is analyzed using a generative AI model. The generative model analyzes each video frame and tags relevant video elements based on order identification information. These video elements are stored in a memory device and organized for quick access when a search is needed.

[0370] The device receives relevant video elements from the server based on the user's exploration request and plays them back via streaming or download. The user's emotions are evaluated in real time using an emotion analysis engine, and the user interface is dynamically adjusted accordingly. This adjustment allows the user to use the system efficiently; for example, if the user is feeling stressed, the system automatically displays instructions in a visually easy-to-understand format.

[0371] Users can obtain real-time instruction information through a display device such as smart glasses. This helps prevent judgment errors during logistics operations and supports efficient work progress.

[0372] As a concrete example, consider a scenario where a logistics center worker processes a specific order number, "12345." If the worker experiences stress, an emotion analysis engine analyzes their state, and the server suggests improved delivery routes and alerts to ensure smoother work. This process is visually presented through the display of smart glasses.

[0373] An example of a prompt message describing the operation of this system is: "Please tell me about real-time support while working at the logistics center. How do you handle users who are experiencing stress?"

[0374] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0375] Step 1:

[0376] The server receives real-time video data from cameras placed within the logistics facility. The input is the raw video stream from the cameras, and the output is video data stored in temporary storage for analysis. The server prepares this data to be input into a generating AI model.

[0377] Step 2:

[0378] The server analyzes the received video data using a generative AI model. The input is the video data saved in step 1, and the output is order identification information and the type of logistics operation activity tagged to each frame. Based on this tagged data, the server organizes and stores the information in a database.

[0379] Step 3:

[0380] The user makes a search request through their terminal, specifying a particular order identifier. The input is the order number from the user, and the output is a request for the associated video fragment. This request is sent to the server.

[0381] Step 4:

[0382] The server retrieves the relevant video fragment from the database based on the user's search request. The input is the user's order number, and the output is the corresponding video fragment. The server then prepares to send this data to the terminal in streaming or download format.

[0383] Step 5:

[0384] The terminal plays video clips received from the server to the user. The input is the video clips from the server, and the output is the video displayed on the terminal. This video provides the user with necessary information for their work and helps improve efficiency.

[0385] Step 6:

[0386] The server uses an emotion analysis engine to evaluate the user's emotions in real time. Input is the user's facial image and biometric data, and output is the analyzed emotion information. Based on this information, the user interface is automatically adjusted to provide optimal support tailored to the user's situation.

[0387] Step 7:

[0388] Users receive improved delivery routes and alerts as needed through the display device. Input is instructional information from the server, and output is visual instructional information presented on the display device. This allows users to obtain information that helps improve their work in real time.

[0389] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0390] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0391] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart glasses 214.

[0392] [Third Embodiment]

[0393] Figure 5 shows an example of the configuration of the data processing system 310 according to the third embodiment.

[0394] As shown in Figure 5, the data processing system 310 includes a data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.

[0395] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0396] The headset terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a display 343. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and display 343 are also connected to the bus 52.

[0397] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0398] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[0399] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0400] Figure 6 shows an example of the main functions of the data processing device 12 and the headset terminal 314. As shown in Figure 6, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0401] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0402] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0403] In the headset terminal 314, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0404] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the headset terminal 314 will be referred to as the "terminal".

[0405] As an embodiment of this invention, an information processing system consisting mainly of a server, a terminal, and a user is constructed. This system has the function of receiving video material obtained from surveillance cameras used in logistics or warehouses, and extracting and managing the corresponding video fragments based on order identification information.

[0406] Server Role

[0407] The server receives video footage in real time from multiple cameras installed at the logistics site, efficiently compresses it, and stores it in storage. Equipped with a generative model, it analyzes the stored video footage to identify logistics operations within the video and assigns related order identification information. This automatically organizes and manages video segments for each order. Furthermore, in response to user search requests, it quickly identifies the necessary video segments and provides them to the user's terminal.

[0408] User roles

[0409] Users access the system using a client terminal and submit search requests specifying order numbers and specific time periods. When a request reaches the server, the server searches its database for video clips based on the request and extracts the relevant sections. By reviewing the video provided in this response and verifying the details of logistics operations, users can quickly identify the cause of problems such as packaging errors.

[0410] Terminal role

[0411] The terminal has the role of forwarding search requests sent by the user to the server, and also has the function of playing video clips received from the server. This interface allows users to obtain the necessary information in a user-friendly environment, which helps in the rapid resolution of problems.

[0412] As a specific example, if a user searches for order number "12345," the server identifies video clips of logistics operations linked to that order from previously collected video footage and delivers them to the terminal in real time. This allows the user to immediately review the video and take necessary actions. This invention will significantly contribute to quality control and efficiency improvement in logistics operations.

[0413] The following describes the processing flow.

[0414] Step 1:

[0415] The server receives video footage in real time from cameras installed within the logistics facility. These cameras are located in each packing area, and the captured video is compressed and stored in a storage device.

[0416] Step 2:

[0417] The server analyzes the stored video footage in batch processing. During this process, it utilizes a generative model to sequentially analyze each frame within the video, recognizing logistics operations and automatically tagging each frame with corresponding order identification information.

[0418] Step 3:

[0419] The server organizes and stores tagged video frames in a database. This makes it easy to search for and extract video fragments associated with each order identification.

[0420] Step 4:

[0421] The user uses the client terminal interface to specify the order number and a specific time, and sends a search request to the server. This request is to retrieve video clips that meet specific criteria.

[0422] Step 5:

[0423] The server receives a search request from the user and quickly performs a search on the video frames stored in the database. It identifies the relevant video fragments and prepares the necessary data.

[0424] Step 6:

[0425] The server sends the identified video clips to the user's device in a streaming or downloadable format, allowing the user to quickly access the information they need.

[0426] Step 7:

[0427] Users can play back video clips received on their devices to review logistics operations. This allows them to identify, for example, the cause of a packing error and take necessary corrective actions.

[0428] (Example 1)

[0429] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0430] In logistics operations, efficiently monitoring work status and preventing misdeliveries and delays requires the management of visual information using monitoring devices and the identification of appropriate operational activities. However, conventional methods have the challenge of not being able to quickly extract necessary information from vast amounts of data.

[0431] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0432] In this invention, the server includes means for receiving visual information acquired from multiple monitoring devices in real time, effectively compressing it, and storing it in a storage device; means for analyzing the received visual information using a generation model to identify logistics operation activities and automatically assign related identification information; and means for quickly searching for and providing visual information fragments organized based on the identification information in response to usage requests. This makes it possible to accurately grasp the work situation at the logistics site and quickly extract relevant information.

[0433] A "monitoring device" is a device installed in logistics sites or related locations to acquire visual information.

[0434] "Visual information" refers to all video data acquired by monitoring devices and is used for monitoring and analyzing logistics operations.

[0435] A "generative model" refers to an algorithm or learning model used to analyze received visual information and recognize and classify specific patterns or activities.

[0436] "Identification information" refers to labels or codes assigned to associate specific logistics operations or orders, and is used to organize and manage information.

[0437] "Storage device" refers to the storage medium or storage system that a server uses to save and store data.

[0438] A "request for use" refers to a request from a user for searching or retrieving data, which acts as a trigger for the server to provide information based on that request.

[0439] The following describes embodiments for carrying out the invention.

[0440] This invention is a system that efficiently manages and analyzes various logistics operations based on visual information from monitoring devices installed at logistics sites. The server receives visual information in real time from multiple monitoring devices and stores it in a storage device using image compression technologies such as H.264 and HEVC while maintaining high quality. This enables the operation of a massive amount of data.

[0441] The server is equipped with a generative AI model that uses deep learning algorithms to analyze visual information. Through this analysis, the server identifies logistics operations and automatically assigns relevant identification information, thereby properly organizing and managing the visual information fragments.

[0442] Users access the server via a client terminal and make search requests based on visual information using prompt messages. For example, when a user enters a prompt message such as "Search for video with order number 12345," the server immediately retrieves the corresponding visual information fragment and provides it to the user terminal.

[0443] The terminal receives and plays back visual information fragments provided by the server, creating an environment where users can smoothly verify logistics operations. This allows users to quickly identify problems such as packing errors or shipping delays based on the received video and take necessary corrective measures. This system contributes to improving quality control and operational efficiency in logistics operations.

[0444] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0445] Step 1:

[0446] The server receives visual information from the monitoring device in real time. At this time, the received data (input) is a raw video stream. The server compresses this video stream using compression technologies such as H.264 or HEVC to reduce the data size (data processing). After that, it saves the compressed visual information (output) to the storage device.

[0447] Step 2:

[0448] The server inputs the stored visual information into a generating AI model and performs data analysis. This model analyzes the received compressed video data (input) and identifies logistics operation activities (data computation). Specifically, it uses a deep learning algorithm to identify objects and actions within the video. As a result of the analysis, identification information corresponding to specific activities within the video (output) is generated and attached to the visual information fragment.

[0449] Step 3:

[0450] The user sends a search request to the server through their client terminal. This request consists of a prompt statement (input), for example, "Search for video with order number 12345." Based on this request, the server refers to the identification information and searches for and retrieves relevant visual information fragments from the database (data processing). The retrieved visual information fragments (output) are then sent to the user's terminal.

[0451] Step 4:

[0452] The user plays back the visual information fragments received on the terminal to confirm logistics operations. Through this playback, the user analyzes the specific activity details obtained from the video, identifying, for example, packaging defects or shipping errors (data calculation and situation confirmation). Based on this, the user can immediately take necessary corrective measures.

[0453] (Application Example 1)

[0454] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0455] To improve operational efficiency in logistics centers and to quickly verify orders and troubleshoot problems, a system is needed that allows on-site staff to obtain necessary information in a short amount of time. However, traditional methods for identifying specific operational activities linked to particular orders from vast amounts of video data are time-consuming and labor-intensive, making immediate countermeasures difficult.

[0456] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0457] In this invention, the server includes means for analyzing video footage of logistics operations acquired based on order identification information and automatically extracting the corresponding video fragments; means for storing the video fragments related to the corresponding order identification information on an electronic storage medium; and means for acquiring order identification information using a voice input device and immediately displaying the related video. This makes it possible for on-site staff to acquire necessary order information by voice using a smart device and immediately confirm its contents.

[0458] "Order identification information" refers to a unique identifier used to identify each order in logistics.

[0459] "Logistics operations" refers to a series of business processes within a logistics center, including the movement, picking, packing, and shipping of goods and packages.

[0460] "Video footage" refers to video data recorded by surveillance cameras, etc., which visually captures the details of logistics operations.

[0461] A "generative model" is a system that uses machine learning algorithms to analyze video materials and recognize and classify specific objects.

[0462] "Electronic storage media" refers to devices or materials used to store digital data, and includes hard disks, SSDs, and cloud storage.

[0463] A "voice input device" is a device that captures the user's voice and sends it to the system as a voice command.

[0464] A "video clip" is a portion of a video recording of a logistics operation related to specific order identification information.

[0465] "On-site staff" refers to workers and personnel who perform their duties on a daily basis at a logistics center.

[0466] As an embodiment of this invention, a video management system for a logistics center is configured. The server acquires video data in real time from surveillance cameras installed in the logistics center, analyzes it using a generating AI model to identify logistics operation activities, and extracts video fragments based on the respective order identification information. These video fragments are stored on an electronic storage medium and provided to the user as needed.

[0467] The terminal is designed to allow staff to search for specific order numbers by voice using voice input devices such as smart glasses. The voice input is converted into a digital signal and sent to a server. The server analyzes this voice signal, quickly identifies the relevant video clips, and provides them to the terminal in streaming or download format.

[0468] As a concrete example, when a field staff member wears smart glasses and issues the voice command "Search for order 12345," the server identifies the video footage associated with order identification information "12345" and immediately displays it in the staff member's field of view. This function allows staff members to confirm the details of logistics operations without delay and take appropriate action.

[0469] By using a generative AI model, it is possible to efficiently extract relevant logistics activities from video footage and provide videos based on specific orders. This system significantly improves operational efficiency within logistics centers.

[0470] An example of a prompt message would be, "Please display the latest logistics operation video for order 12345." This operation is performed using voice recognition technology via smart glasses.

[0471] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0472] Step 1:

[0473] The server receives real-time video data from surveillance cameras in the logistics center. It receives continuous video streams from each camera as input and analyzes them using a generative AI model. As output, it identifies logistics operations within the video and generates video data labeled based on corresponding order identification information. Specifically, it executes video analysis algorithms to perform object recognition and motion detection.

[0474] Step 2:

[0475] The server stores the analyzed video data on an electronic storage medium. It receives labeled video data as input and converts it into a format for storage in the database. As output, the video data is categorized by order identification information and stored in storage in a searchable format. Specifically, it creates an index to organize and store video segments and their associated metadata.

[0476] Step 3:

[0477] The user searches for a specific order number by issuing a voice command using smart glasses. The input is voice data collected by a voice input device. The output is the order number obtained by converting the voice signal into text. Specifically, voice recognition software is used to interpret the voice command and identify the order number.

[0478] Step 4:

[0479] The server searches the database for relevant video data based on the order number received from the user. The order number, in text format, is used as input, and a query is generated to search for the corresponding video data. The output extracts data for the relevant video segments. Specifically, the system executes a database query to quickly retrieve the relevant video data.

[0480] Step 5:

[0481] The terminal receives video fragments from the server and transmits them to the user's smart glasses in streaming format. The input is video data from the server. The output is the video being visually displayed to the staff. Specifically, the operation involves transferring video data via a network connection and displaying the video using the smart glasses' video playback function.

[0482] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0483] To implement this invention, a system comprising a server, a terminal, and a user is required. This system includes a function to analyze and provide video of logistics operations based on order identification information, as well as an emotion engine that recognizes the user's emotions.

[0484] Server Role

[0485] The server receives real-time video footage from cameras within the logistics facility and analyzes it using a generative model. The generative model identifies logistics operations in each frame and tags them with relevant order identification information. This process automatically organizes and stores the video fragments in a database.

[0486] In addition, the server analyzes the user's emotions through an emotion engine. This enables the provision of optimal video segments tailored to the user's state and the application of a customized user interface.

[0487] User roles

[0488] When a user accesses the system from a client terminal and makes a search request specifying an order number and time, the emotion engine analyzes their emotions and adjusts the user interface and presentation methods accordingly. This provides the user with a more comfortable and efficient operating experience.

[0489] Terminal role

[0490] The device sends the user's search request to the server and displays the video clips received from the server in a playable state. Based on information obtained from the emotion engine, the video being played and the user interface are customized to match the user's emotions.

[0491] For example, if a user expresses impatience while searching for the order "12345," the emotion engine analyzes that emotion, and the server provides the terminal with faster display speeds and additional guidelines. This allows the user to smoothly identify the problem and consider solutions, improving overall work efficiency. In this way, customization using an emotion engine can improve the user experience.

[0492] The following describes the processing flow.

[0493] Step 1:

[0494] The server receives video footage from surveillance cameras within the logistics facility in real time and efficiently stores it in its storage device. This ensures that logistics operations are continuously recorded.

[0495] Step 2:

[0496] The server analyzes the stored video using a generative model. The analysis is performed frame by frame, identifying logistics operations and related order identification information, and tagging the frame accordingly. This information is then organized in a database.

[0497] Step 3:

[0498] The user executes a search request via the system interface through a client terminal, specifying the order number and a specific time. The search request is then sent to the server.

[0499] Step 4:

[0500] The server receives a search request from the user and quickly searches the database for the relevant video clip. Furthermore, the emotion engine recognizes the user's emotions from their facial expressions and actions, and generates analysis results.

[0501] Step 5:

[0502] Based on the analysis results of the emotion engine, the server determines how to provide information according to the user's emotional state. For example, if the user is showing signs of stress, the server prepares a simplified explanation.

[0503] Step 6:

[0504] The device receives video clips transmitted from the server via streaming or download and begins playback. Simultaneously, the user interface adjusts to the user's emotions.

[0505] Step 7:

[0506] Users review the played-back footage and evaluate the logistics operations in detail. User feedback is further analyzed by an emotion engine to improve the system's responsiveness and usability.

[0507] (Example 2)

[0508] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0509] There is a need to efficiently extract necessary video clips from a vast amount of video footage related to logistics operations based on specific order identification information, and present them to the user quickly and appropriately. Furthermore, there is a lack of technology to adjust search results and the user interface according to the user's emotional state, thus necessitating an improvement in the user experience.

[0510] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0511] In this invention, the server includes means for analyzing video footage of logistics operations acquired based on order identification information using a generating AI model to automatically extract relevant video fragments, means for storing the relevant video fragments related to the order identification information in a storage device, and means for adjusting the user interface during search requests and information provision using an emotion engine that analyzes the user's emotions. As a result, the user can efficiently acquire video related to a specific order and comfortably use the system with an interface customized to the user's emotions.

[0512] "Order identification information" refers to information used to identify a specific order related to a logistics operation.

[0513] "Logistics operations" refers to a series of tasks and processes related to the movement, storage, and management of goods.

[0514] A "video clip" is a portion of video footage extracted from video material based on specific information.

[0515] A "generative AI model" is a computational model that uses artificial intelligence technology to analyze given data and generate or identify specific patterns or information.

[0516] An "emotion engine" is a program that analyzes user input and behavior to estimate the user's emotional state.

[0517] A "user terminal" is an electronic device, such as a computer or smartphone, that a user uses to perform operations.

[0518] A "storage device" refers to hardware or software used to store data or information.

[0519] "Streaming" is a technology or method that transmits data instantly and plays it back in real time on the receiving end.

[0520] This invention is implemented primarily as a system consisting of a server, a terminal, and a user. In this system, the server receives video data acquired in real time from cameras installed within the logistics facility and performs analysis using a generative AI model. The generative AI model identifies logistics operation activities in each video frame and tags the video fragments with appropriate order identification information. Following this video analysis, the video fragments are automatically saved to a storage device.

[0521] The server also plays a role in analyzing the user's emotions through its emotion engine. This allows the server to provide video clips and customize the user interface based on the user's emotions when they make a search request using a client terminal, specifying an order number and time. For example, if a user searches for the order "12345" and expresses anxiety, the server will use the emotion engine's analysis to provide faster video clips and display additional guidelines.

[0522] The terminal displays video clips transmitted from the server in a playable format and adjusts the interface based on the user's emotions using an emotion engine. This adjustment allows the user to have an efficient and comfortable operating experience.

[0523] An example of a prompt message that might be input into the generating AI model is, "Search for videos of logistics activities related to order number 12345, adjust the display speed to be faster, and provide guidelines to the user." This allows the system to immediately process the appropriate information and provide customized feedback to the user.

[0524] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0525] Step 1:

[0526] The server receives video data in real time from cameras within the logistics facility. High-resolution video frames acquired from the cameras are provided to the server as input for the received video. The server temporarily stores this frame data for the next analysis step.

[0527] Step 2:

[0528] The server analyzes the received video frames using a generating AI model. In this analysis step, the AI ​​model identifies the movement of items and people within the frame and identifies logistics operations. As a result, order identification information corresponding to each frame is generated. This allows the system to understand which order a particular logistics activity is related to.

[0529] Step 3:

[0530] The server tags video frames with order identification information based on the analysis results. Identification information obtained from the AI ​​model is used as input, and tagged video frames are obtained as output. These are organized and stored in a database to enable efficient searching.

[0531] Step 4:

[0532] The server uses an emotion engine to analyze user input and actions to infer their emotional state. Here, user input data is provided to the emotion engine, and information about the user's emotions is generated as output. This information is then used in the next step to customize the user interface.

[0533] Step 5:

[0534] The user uses a client terminal to send a search request specifying a particular order number and time. This request is passed to the server in the form of prompt statements. The server processes this request and prepares to send the corresponding video clips to the terminal.

[0535] Step 6:

[0536] Based on the emotion analysis results, the server determines interface settings and video display methods appropriate for the user's emotional state. At this stage, the adjusted settings are sent to the terminal and applied to improve the user's experience.

[0537] Step 7:

[0538] The device displays video clips received from the server in a playable format. Furthermore, it can adjust the video playback speed and add guidelines based on information from the emotion engine. This allows users to view the video in a way that resonates with their own emotions.

[0539] (Application Example 2)

[0540] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0541] To improve operational efficiency and reduce errors in logistics centers, there is a need for methods to make operational information easier to understand. Furthermore, a system is needed that supports smooth work execution by providing support tailored to the emotions of the workers. The challenge is to improve the efficiency and safety of the entire logistics process.

[0542] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0543] In this invention, the server includes means for analyzing video information of logistics operations acquired based on order identification information using a generative model to automatically identify the corresponding video elements; means for storing the corresponding video elements related to the order identification information in a storage device; means for providing video data related to the order identification information based on a search request from the user; means for identifying the user's emotions using an emotion analysis engine and customizing the user interface; and means for presenting emotion-appropriate instruction information along with video data to a display device used in the logistics center. This makes it possible for workers to receive appropriate video information and emotion-appropriate instructions in real time, improving the efficiency and safety of logistics operations.

[0544] "Order identification information" refers to information used in logistics operations to identify a specific order.

[0545] "Logistics operations" refers to a series of tasks performed within a logistics facility, such as receiving, storing, picking, and shipping goods.

[0546] A "generative model" is an algorithm that learns patterns and features from input data and uses that data for analysis and classification.

[0547] "Video elements" refer to segments or frames related to a specific order within video materials, including logistics operations.

[0548] "Storage device" refers to hardware and software used to hold digital data, and includes databases and storage devices.

[0549] A "search request" is a request for inquiry or retrieval made by a user to obtain a specific order or information from the system.

[0550] An "emotion analysis engine" is a software engine that analyzes a user's facial expressions and behavior to determine their emotional state.

[0551] A "user interface" refers to the screen display and input devices used for interaction between a user and a computer system.

[0552] A "display device" is hardware used to provide information visually, and includes screens and displays.

[0553] "Instructional information" refers to information provided to guide operations or actions, and includes commands and guidelines.

[0554] A specific embodiment of this invention involves a system consisting of a server, a terminal, and a user. The server receives real-time video footage of delivery activities from cameras within the logistics facility. The received video is analyzed using a generative AI model. The generative model analyzes each video frame and tags relevant video elements based on order identification information. These video elements are stored in a memory device and organized for quick access when a search is needed.

[0555] The device receives relevant video elements from the server based on the user's exploration request and plays them back via streaming or download. The user's emotions are evaluated in real time using an emotion analysis engine, and the user interface is dynamically adjusted accordingly. This adjustment allows the user to use the system efficiently; for example, if the user is feeling stressed, the system automatically displays instructions in a visually easy-to-understand format.

[0556] Users can obtain real-time instruction information through a display device such as smart glasses. This helps prevent judgment errors during logistics operations and supports efficient work progress.

[0557] As a concrete example, consider a scenario where a logistics center worker processes a specific order number, "12345." If the worker experiences stress, an emotion analysis engine analyzes their state, and the server suggests improved delivery routes and alerts to ensure smoother work. This process is visually presented through the display of smart glasses.

[0558] An example of a prompt message describing the operation of this system is: "Please tell me about real-time support while working at the logistics center. How do you handle users who are experiencing stress?"

[0559] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0560] Step 1:

[0561] The server receives real-time video data from cameras placed within the logistics facility. The input is the raw video stream from the cameras, and the output is video data stored in temporary storage for analysis. The server prepares this data to be input into a generating AI model.

[0562] Step 2:

[0563] The server analyzes the received video data using a generative AI model. The input is the video data saved in step 1, and the output is order identification information and the type of logistics operation activity tagged to each frame. Based on this tagged data, the server organizes and stores the information in a database.

[0564] Step 3:

[0565] The user makes a search request through their terminal, specifying a particular order identifier. The input is the order number from the user, and the output is a request for the associated video fragment. This request is sent to the server.

[0566] Step 4:

[0567] The server retrieves the relevant video fragment from the database based on the user's search request. The input is the user's order number, and the output is the corresponding video fragment. The server then prepares to send this data to the terminal in streaming or download format.

[0568] Step 5:

[0569] The terminal plays video clips received from the server to the user. The input is the video clips from the server, and the output is the video displayed on the terminal. This video provides the user with necessary information for their work and helps improve efficiency.

[0570] Step 6:

[0571] The server uses an emotion analysis engine to evaluate the user's emotions in real time. Input is the user's facial image and biometric data, and output is the analyzed emotion information. Based on this information, the user interface is automatically adjusted to provide optimal support tailored to the user's situation.

[0572] Step 7:

[0573] Users receive improved delivery routes and alerts as needed through the display device. Input is instructional information from the server, and output is visual instructional information presented on the display device. This allows users to obtain information that helps improve their work in real time.

[0574] The specific processing unit 290 transmits the result of the specific processing to the headset terminal 314. In the headset terminal 314, the control unit 46A causes the speaker 240 and display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0575] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet Search).<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0576] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and specific processing may also be performed by the headset terminal 314.

[0577] [Fourth Embodiment]

[0578] Figure 7 shows an example of the configuration of the data processing system 410 according to the fourth embodiment.

[0579] As shown in Figure 7, the data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.

[0580] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0581] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a controlled object 443. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and controlled object 443 are also connected to the bus 52.

[0582] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0583] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[0584] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0585] The controlled object 443 includes a display device, LEDs in the eyes, and motors that drive the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors in the arms, hands, and feet. Some of the robot 414's emotions can be expressed by controlling these motors. Furthermore, the robot 414's facial expressions can also be expressed by controlling the illumination state of the LEDs in its eyes.

[0586] Figure 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Figure 8, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0587] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0588] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0589] In robot 414, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0590] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0591] As an embodiment of this invention, an information processing system consisting mainly of a server, a terminal, and a user is constructed. This system has the function of receiving video material obtained from surveillance cameras used in logistics or warehouses, and extracting and managing the corresponding video fragments based on order identification information.

[0592] Server Role

[0593] The server receives video footage in real time from multiple cameras installed at the logistics site, efficiently compresses it, and stores it in storage. Equipped with a generative model, it analyzes the stored video footage to identify logistics operations within the video and assigns related order identification information. This automatically organizes and manages video segments for each order. Furthermore, in response to user search requests, it quickly identifies the necessary video segments and provides them to the user's terminal.

[0594] User roles

[0595] Users access the system using a client terminal and submit search requests specifying order numbers and specific time periods. When a request reaches the server, the server searches its database for video clips based on the request and extracts the relevant sections. By reviewing the video provided in this response and verifying the details of logistics operations, users can quickly identify the cause of problems such as packaging errors.

[0596] Terminal role

[0597] The terminal has the role of forwarding search requests sent by the user to the server, and also has the function of playing video clips received from the server. This interface allows users to obtain the necessary information in a user-friendly environment, which helps in the rapid resolution of problems.

[0598] As a specific example, if a user searches for order number "12345," the server identifies video clips of logistics operations linked to that order from previously collected video footage and delivers them to the terminal in real time. This allows the user to immediately review the video and take necessary actions. This invention will significantly contribute to quality control and efficiency improvement in logistics operations.

[0599] The following describes the processing flow.

[0600] Step 1:

[0601] The server receives video footage in real time from cameras installed within the logistics facility. These cameras are located in each packing area, and the captured video is compressed and stored in a storage device.

[0602] Step 2:

[0603] The server analyzes the stored video footage in batch processing. During this process, it utilizes a generative model to sequentially analyze each frame within the video, recognizing logistics operations and automatically tagging each frame with corresponding order identification information.

[0604] Step 3:

[0605] The server organizes and stores tagged video frames in a database. This makes it easy to search for and extract video fragments associated with each order identification.

[0606] Step 4:

[0607] The user uses the client terminal interface to specify the order number and a specific time, and sends a search request to the server. This request is to retrieve video clips that meet specific criteria.

[0608] Step 5:

[0609] The server receives a search request from the user and quickly performs a search on the video frames stored in the database. It identifies the relevant video fragments and prepares the necessary data.

[0610] Step 6:

[0611] The server sends the identified video clips to the user's device in a streaming or downloadable format, allowing the user to quickly access the information they need.

[0612] Step 7:

[0613] Users can play back video clips received on their devices to review logistics operations. This allows them to identify, for example, the cause of a packing error and take necessary corrective actions.

[0614] (Example 1)

[0615] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0616] In logistics operations, efficiently monitoring work status and preventing misdeliveries and delays requires the management of visual information using monitoring devices and the identification of appropriate operational activities. However, conventional methods have the challenge of not being able to quickly extract necessary information from vast amounts of data.

[0617] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0618] In this invention, the server includes means for receiving visual information acquired from multiple monitoring devices in real time, effectively compressing it, and storing it in a storage device; means for analyzing the received visual information using a generation model to identify logistics operation activities and automatically assign related identification information; and means for quickly searching for and providing visual information fragments organized based on the identification information in response to usage requests. This makes it possible to accurately grasp the work situation at the logistics site and quickly extract relevant information.

[0619] A "monitoring device" is a device installed in logistics sites or related locations to acquire visual information.

[0620] "Visual information" refers to all video data acquired by monitoring devices and is used for monitoring and analyzing logistics operations.

[0621] A "generative model" refers to an algorithm or learning model used to analyze received visual information and recognize and classify specific patterns or activities.

[0622] "Identification information" refers to labels or codes assigned to associate specific logistics operations or orders, and is used to organize and manage information.

[0623] "Storage device" refers to the storage medium or storage system that a server uses to save and store data.

[0624] A "request for use" refers to a request from a user for searching or retrieving data, which acts as a trigger for the server to provide information based on that request.

[0625] The following describes embodiments for carrying out the invention.

[0626] This invention is a system that efficiently manages and analyzes various logistics operations based on visual information from monitoring devices installed at logistics sites. The server receives visual information in real time from multiple monitoring devices and stores it in a storage device using image compression technologies such as H.264 and HEVC while maintaining high quality. This enables the operation of a massive amount of data.

[0627] The server is equipped with a generative AI model that uses deep learning algorithms to analyze visual information. Through this analysis, the server identifies logistics operations and automatically assigns relevant identification information, thereby properly organizing and managing the visual information fragments.

[0628] Users access the server via a client terminal and make search requests based on visual information using prompt messages. For example, when a user enters a prompt message such as "Search for video with order number 12345," the server immediately retrieves the corresponding visual information fragment and provides it to the user terminal.

[0629] The terminal receives and plays back visual information fragments provided by the server, creating an environment where users can smoothly verify logistics operations. This allows users to quickly identify problems such as packing errors or shipping delays based on the received video and take necessary corrective measures. This system contributes to improving quality control and operational efficiency in logistics operations.

[0630] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0631] Step 1:

[0632] The server receives visual information from the monitoring device in real time. At this time, the received data (input) is a raw video stream. The server compresses this video stream using compression technologies such as H.264 or HEVC to reduce the data size (data processing). After that, it saves the compressed visual information (output) to the storage device.

[0633] Step 2:

[0634] The server inputs the stored visual information into a generating AI model and performs data analysis. This model analyzes the received compressed video data (input) and identifies logistics operation activities (data computation). Specifically, it uses a deep learning algorithm to identify objects and actions within the video. As a result of the analysis, identification information corresponding to specific activities within the video (output) is generated and attached to the visual information fragment.

[0635] Step 3:

[0636] The user sends a search request to the server through their client terminal. This request consists of a prompt statement (input), for example, "Search for video with order number 12345." Based on this request, the server refers to the identification information and searches for and retrieves relevant visual information fragments from the database (data processing). The retrieved visual information fragments (output) are then sent to the user's terminal.

[0637] Step 4:

[0638] The user plays back the visual information fragments received on the terminal to confirm logistics operations. Through this playback, the user analyzes the specific activity details obtained from the video, identifying, for example, packaging defects or shipping errors (data calculation and situation confirmation). Based on this, the user can immediately take necessary corrective measures.

[0639] (Application Example 1)

[0640] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0641] To improve operational efficiency in logistics centers and to quickly verify orders and troubleshoot problems, a system is needed that allows on-site staff to obtain necessary information in a short amount of time. However, traditional methods for identifying specific operational activities linked to particular orders from vast amounts of video data are time-consuming and labor-intensive, making immediate countermeasures difficult.

[0642] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0643] In this invention, the server includes means for analyzing video footage of logistics operations acquired based on order identification information and automatically extracting the corresponding video fragments; means for storing the video fragments related to the corresponding order identification information on an electronic storage medium; and means for acquiring order identification information using a voice input device and immediately displaying the related video. This makes it possible for on-site staff to acquire necessary order information by voice using a smart device and immediately confirm its contents.

[0644] "Order identification information" refers to a unique identifier used to identify each order in logistics.

[0645] "Logistics operations" refers to a series of business processes within a logistics center, including the movement, picking, packing, and shipping of goods and packages.

[0646] "Video footage" refers to video data recorded by surveillance cameras, etc., which visually captures the details of logistics operations.

[0647] A "generative model" is a system that uses machine learning algorithms to analyze video materials and recognize and classify specific objects.

[0648] "Electronic storage media" refers to devices or materials used to store digital data, and includes hard disks, SSDs, and cloud storage.

[0649] A "voice input device" is a device that captures the user's voice and sends it to the system as a voice command.

[0650] A "video clip" is a portion of a video recording of a logistics operation related to specific order identification information.

[0651] "On-site staff" refers to workers and personnel who perform their duties on a daily basis at a logistics center.

[0652] As an embodiment of this invention, a video management system for a logistics center is configured. The server acquires video data in real time from surveillance cameras installed in the logistics center, analyzes it using a generating AI model to identify logistics operation activities, and extracts video fragments based on the respective order identification information. These video fragments are stored on an electronic storage medium and provided to the user as needed.

[0653] The terminal is designed to allow staff to search for specific order numbers by voice using voice input devices such as smart glasses. The voice input is converted into a digital signal and sent to a server. The server analyzes this voice signal, quickly identifies the relevant video clips, and provides them to the terminal in streaming or download format.

[0654] As a concrete example, when a field staff member wears smart glasses and issues the voice command "Search for order 12345," the server identifies the video footage associated with order identification information "12345" and immediately displays it in the staff member's field of view. This function allows staff members to confirm the details of logistics operations without delay and take appropriate action.

[0655] By using a generative AI model, it is possible to efficiently extract relevant logistics activities from video footage and provide videos based on specific orders. This system significantly improves operational efficiency within logistics centers.

[0656] An example of a prompt message would be, "Please display the latest logistics operation video for order 12345." This operation is performed using voice recognition technology via smart glasses.

[0657] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0658] Step 1:

[0659] The server receives real-time video data from surveillance cameras in the logistics center. It receives continuous video streams from each camera as input and analyzes them using a generative AI model. As output, it identifies logistics operations within the video and generates video data labeled based on corresponding order identification information. Specifically, it executes video analysis algorithms to perform object recognition and motion detection.

[0660] Step 2:

[0661] The server stores the analyzed video data on an electronic storage medium. It receives labeled video data as input and converts it into a format for storage in the database. As output, the video data is categorized by order identification information and stored in storage in a searchable format. Specifically, it creates an index to organize and store video segments and their associated metadata.

[0662] Step 3:

[0663] The user searches for a specific order number by issuing a voice command using smart glasses. The input is voice data collected by a voice input device. The output is the order number obtained by converting the voice signal into text. Specifically, voice recognition software is used to interpret the voice command and identify the order number.

[0664] Step 4:

[0665] The server searches the database for relevant video data based on the order number received from the user. The order number, in text format, is used as input, and a query is generated to search for the corresponding video data. The output extracts data for the relevant video segments. Specifically, the system executes a database query to quickly retrieve the relevant video data.

[0666] Step 5:

[0667] The terminal receives video fragments from the server and transmits them to the user's smart glasses in streaming format. The input is video data from the server. The output is the video being visually displayed to the staff. Specifically, the operation involves transferring video data via a network connection and displaying the video using the smart glasses' video playback function.

[0668] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0669] To implement this invention, a system comprising a server, a terminal, and a user is required. This system includes a function to analyze and provide video of logistics operations based on order identification information, as well as an emotion engine that recognizes the user's emotions.

[0670] Server Role

[0671] The server receives real-time video footage from cameras within the logistics facility and analyzes it using a generative model. The generative model identifies logistics operations in each frame and tags them with relevant order identification information. This process automatically organizes and stores the video fragments in a database.

[0672] In addition, the server analyzes the user's emotions through an emotion engine. This enables the provision of optimal video segments tailored to the user's state and the application of a customized user interface.

[0673] User roles

[0674] When a user accesses the system from a client terminal and makes a search request specifying an order number and time, the emotion engine analyzes their emotions and adjusts the user interface and presentation methods accordingly. This provides the user with a more comfortable and efficient operating experience.

[0675] Terminal role

[0676] The device sends the user's search request to the server and displays the video clips received from the server in a playable state. Based on information obtained from the emotion engine, the video being played and the user interface are customized to match the user's emotions.

[0677] For example, if a user expresses impatience while searching for the order "12345," the emotion engine analyzes that emotion, and the server provides the terminal with faster display speeds and additional guidelines. This allows the user to smoothly identify the problem and consider solutions, improving overall work efficiency. In this way, customization using an emotion engine can improve the user experience.

[0678] The following describes the processing flow.

[0679] Step 1:

[0680] The server receives video footage from surveillance cameras within the logistics facility in real time and efficiently stores it in its storage device. This ensures that logistics operations are continuously recorded.

[0681] Step 2:

[0682] The server analyzes the stored video using a generative model. The analysis is performed frame by frame, identifying logistics operations and related order identification information, and tagging the frame accordingly. This information is then organized in a database.

[0683] Step 3:

[0684] The user executes a search request via the system interface through a client terminal, specifying the order number and a specific time. The search request is then sent to the server.

[0685] Step 4:

[0686] The server receives a search request from the user and quickly searches the database for the relevant video clip. Furthermore, the emotion engine recognizes the user's emotions from their facial expressions and actions, and generates analysis results.

[0687] Step 5:

[0688] Based on the analysis results of the emotion engine, the server determines how to provide information according to the user's emotional state. For example, if the user is showing signs of stress, the server prepares a simplified explanation.

[0689] Step 6:

[0690] The device receives video clips transmitted from the server via streaming or download and begins playback. Simultaneously, the user interface adjusts to the user's emotions.

[0691] Step 7:

[0692] Users review the played-back footage and evaluate the logistics operations in detail. User feedback is further analyzed by an emotion engine to improve the system's responsiveness and usability.

[0693] (Example 2)

[0694] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0695] There is a need to efficiently extract necessary video clips from a vast amount of video footage related to logistics operations based on specific order identification information, and present them to the user quickly and appropriately. Furthermore, there is a lack of technology to adjust search results and the user interface according to the user's emotional state, thus necessitating an improvement in the user experience.

[0696] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0697] In this invention, the server includes means for analyzing video footage of logistics operations acquired based on order identification information using a generating AI model to automatically extract relevant video fragments, means for storing the relevant video fragments related to the order identification information in a storage device, and means for adjusting the user interface during search requests and information provision using an emotion engine that analyzes the user's emotions. As a result, the user can efficiently acquire video related to a specific order and comfortably use the system with an interface customized to the user's emotions.

[0698] "Order identification information" refers to information used to identify a specific order related to a logistics operation.

[0699] "Logistics operations" refers to a series of tasks and processes related to the movement, storage, and management of goods.

[0700] A "video clip" is a portion of video footage extracted from video material based on specific information.

[0701] A "generative AI model" is a computational model that uses artificial intelligence technology to analyze given data and generate or identify specific patterns or information.

[0702] An "emotion engine" is a program that analyzes user input and behavior to estimate the user's emotional state.

[0703] A "user terminal" is an electronic device, such as a computer or smartphone, that a user uses to perform operations.

[0704] A "storage device" refers to hardware or software used to store data or information.

[0705] "Streaming" is a technology or method that transmits data instantly and plays it back in real time on the receiving end.

[0706] This invention is implemented primarily as a system consisting of a server, a terminal, and a user. In this system, the server receives video data acquired in real time from cameras installed within the logistics facility and performs analysis using a generative AI model. The generative AI model identifies logistics operation activities in each video frame and tags the video fragments with appropriate order identification information. Following this video analysis, the video fragments are automatically saved to a storage device.

[0707] The server also plays a role in analyzing the user's emotions through its emotion engine. This allows the server to provide video clips and customize the user interface based on the user's emotions when they make a search request using a client terminal, specifying an order number and time. For example, if a user searches for the order "12345" and expresses anxiety, the server will use the emotion engine's analysis to provide faster video clips and display additional guidelines.

[0708] The terminal displays video clips transmitted from the server in a playable format and adjusts the interface based on the user's emotions using an emotion engine. This adjustment allows the user to have an efficient and comfortable operating experience.

[0709] An example of a prompt message that might be input into the generating AI model is, "Search for videos of logistics activities related to order number 12345, adjust the display speed to be faster, and provide guidelines to the user." This allows the system to immediately process the appropriate information and provide customized feedback to the user.

[0710] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0711] Step 1:

[0712] The server receives video data in real time from cameras within the logistics facility. High-resolution video frames acquired from the cameras are provided to the server as input for the received video. The server temporarily stores this frame data for the next analysis step.

[0713] Step 2:

[0714] The server analyzes the received video frames using a generating AI model. In this analysis step, the AI ​​model identifies the movement of items and people within the frame and identifies logistics operations. As a result, order identification information corresponding to each frame is generated. This allows the system to understand which order a particular logistics activity is related to.

[0715] Step 3:

[0716] The server tags video frames with order identification information based on the analysis results. Identification information obtained from the AI ​​model is used as input, and tagged video frames are obtained as output. These are organized and stored in a database to enable efficient searching.

[0717] Step 4:

[0718] The server uses an emotion engine to analyze user input and actions to infer their emotional state. Here, user input data is provided to the emotion engine, and information about the user's emotions is generated as output. This information is then used in the next step to customize the user interface.

[0719] Step 5:

[0720] The user uses a client terminal to send a search request specifying a particular order number and time. This request is passed to the server in the form of prompt statements. The server processes this request and prepares to send the corresponding video clips to the terminal.

[0721] Step 6:

[0722] Based on the emotion analysis results, the server determines interface settings and video display methods appropriate for the user's emotional state. At this stage, the adjusted settings are sent to the terminal and applied to improve the user's experience.

[0723] Step 7:

[0724] The device displays video clips received from the server in a playable format. Furthermore, it can adjust the video playback speed and add guidelines based on information from the emotion engine. This allows users to view the video in a way that resonates with their own emotions.

[0725] (Application Example 2)

[0726] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0727] To improve operational efficiency and reduce errors in logistics centers, there is a need for methods to make operational information easier to understand. Furthermore, a system is needed that supports smooth work execution by providing support tailored to the emotions of the workers. The challenge is to improve the efficiency and safety of the entire logistics process.

[0728] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0729] In this invention, the server includes means for analyzing video information of logistics operations acquired based on order identification information using a generative model to automatically identify the corresponding video elements; means for storing the corresponding video elements related to the order identification information in a storage device; means for providing video data related to the order identification information based on a search request from the user; means for identifying the user's emotions using an emotion analysis engine and customizing the user interface; and means for presenting emotion-appropriate instruction information along with video data to a display device used in the logistics center. This makes it possible for workers to receive appropriate video information and emotion-appropriate instructions in real time, improving the efficiency and safety of logistics operations.

[0730] "Order identification information" refers to information used in logistics operations to identify a specific order.

[0731] "Logistics operations" refers to a series of tasks performed within a logistics facility, such as receiving, storing, picking, and shipping goods.

[0732] A "generative model" is an algorithm that learns patterns and features from input data and uses that data for analysis and classification.

[0733] "Video elements" refer to segments or frames related to a specific order within video materials, including logistics operations.

[0734] "Storage device" refers to hardware and software used to hold digital data, and includes databases and storage devices.

[0735] A "search request" is a request for inquiry or retrieval made by a user to obtain a specific order or information from the system.

[0736] An "emotion analysis engine" is a software engine that analyzes a user's facial expressions and behavior to determine their emotional state.

[0737] A "user interface" refers to the screen display and input devices used for interaction between a user and a computer system.

[0738] A "display device" is hardware used to provide information visually, and includes screens and displays.

[0739] "Instructional information" refers to information provided to guide operations or actions, and includes commands and guidelines.

[0740] A specific embodiment of this invention involves a system consisting of a server, a terminal, and a user. The server receives real-time video footage of delivery activities from cameras within the logistics facility. The received video is analyzed using a generative AI model. The generative model analyzes each video frame and tags relevant video elements based on order identification information. These video elements are stored in a memory device and organized for quick access when a search is needed.

[0741] The device receives relevant video elements from the server based on the user's exploration request and plays them back via streaming or download. The user's emotions are evaluated in real time using an emotion analysis engine, and the user interface is dynamically adjusted accordingly. This adjustment allows the user to use the system efficiently; for example, if the user is feeling stressed, the system automatically displays instructions in a visually easy-to-understand format.

[0742] Users can obtain real-time instruction information through a display device such as smart glasses. This helps prevent judgment errors during logistics operations and supports efficient work progress.

[0743] As a concrete example, consider a scenario where a logistics center worker processes a specific order number, "12345." If the worker experiences stress, an emotion analysis engine analyzes their state, and the server suggests improved delivery routes and alerts to ensure smoother work. This process is visually presented through the display of smart glasses.

[0744] An example of a prompt message describing the operation of this system is: "Please tell me about real-time support while working at the logistics center. How do you handle users who are experiencing stress?"

[0745] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0746] Step 1:

[0747] The server receives real-time video data from cameras placed within the logistics facility. The input is the raw video stream from the cameras, and the output is video data stored in temporary storage for analysis. The server prepares this data to be input into a generating AI model.

[0748] Step 2:

[0749] The server analyzes the received video data using a generative AI model. The input is the video data saved in step 1, and the output is order identification information and the type of logistics operation activity tagged to each frame. Based on this tagged data, the server organizes and stores the information in a database.

[0750] Step 3:

[0751] The user makes a search request through their terminal, specifying a particular order identifier. The input is the order number from the user, and the output is a request for the associated video fragment. This request is sent to the server.

[0752] Step 4:

[0753] The server retrieves the relevant video fragment from the database based on the user's search request. The input is the user's order number, and the output is the corresponding video fragment. The server then prepares to send this data to the terminal in streaming or download format.

[0754] Step 5:

[0755] The terminal plays video clips received from the server to the user. The input is the video clips from the server, and the output is the video displayed on the terminal. This video provides the user with necessary information for their work and helps improve efficiency.

[0756] Step 6:

[0757] The server uses an emotion analysis engine to evaluate the user's emotions in real time. Input is the user's facial image and biometric data, and output is the analyzed emotion information. Based on this information, the user interface is automatically adjusted to provide optimal support tailored to the user's situation.

[0758] Step 7:

[0759] Users receive improved delivery routes and alerts as needed through the display device. Input is instructional information from the server, and output is visual instructional information presented on the display device. This allows users to obtain information that helps improve their work in real time.

[0760] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the controlled object 443 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0761] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet Search).<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0762] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the robot 414.

[0763] Furthermore, the emotion identification model 59, acting as an emotion engine, may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to a specific mapping, which is an emotion map (see Figure 9). Similarly, the emotion identification model 59 may also determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.

[0764] Figure 9 shows an emotion map 400 in which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. The closer to the center of the concentric circles, the more primitive the emotions are located. Further out of the concentric circles, emotions representing states and actions arising from mental states are located. Emotion is a concept that includes feelings and mental states. On the left side of the concentric circles, emotions that are generally generated from reactions occurring in the brain are located. On the right side of the concentric circles, emotions that are generally induced by situational judgment are located. Above and below the concentric circles, emotions that are generally generated from reactions occurring in the brain and induced by situational judgment are located. In addition, the emotion of "pleasure" is located on the upper side of the concentric circles, and the emotion of "displeasure" is located on the lower side. Thus, in the emotion map 400, multiple emotions are mapped based on the structure in which emotions arise, and emotions that are likely to occur simultaneously are mapped close together.

[0765] These emotions are distributed at the 3 o'clock position on the Emotion Map 400, and usually fluctuate between feelings of security and anxiety. In the right half of the Emotion Map 400, situational awareness takes precedence over internal feelings, resulting in a calm impression.

[0766] The inside of the Emotion Map 400 represents inner thoughts, while the outside represents actions. Therefore, the further you go from the outside of the Emotion Map 400, the more visible (expressed in actions) your emotions become.

[0767] Here, human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. Similarly, in robots, cars, motorcycles, etc., emotions can be created based on various balances, such as posture and battery level. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. The emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on a system for analyzing brain physiological signals of speech emotion recognition and emotion, Tokushima University, doctoral dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map contains emotions belonging to a region called "response," where sensation is dominant. The right half of the emotion map contains emotions belonging to a region called "situation," where situational awareness is dominant.

[0768] The emotion map defines two emotions that promote learning. One is the emotion around the middle of the negative "repentance" and "reflection" on the situation side. In other words, it is when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is the emotion around the positive "desire" on the reaction side. In other words, it is when the robot has positive feelings such as "I want more" or "I want to know more."

[0769] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values ​​representing each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple training data sets, which are combinations of user input and emotion values ​​representing each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions located close together have similar values, as shown in the emotion map 900 in Figure 10. Figure 10 shows an example where multiple emotions such as "reassured," "calm," and "confident" have similar emotion values.

[0770] The above description primarily focuses on the functions of the data processing device 12 in relation to this disclosure. However, the system related to this disclosure is not necessarily implemented on a server. The system related to this disclosure may be implemented as a general information processing system. This disclosure may be implemented, for example, as a software program that runs on a personal computer or as an application that runs on a smartphone. The method related to this disclosure may be provided to users in SaaS (Software as a Service) format.

[0771] In the above embodiment, an example was given in which a specific process is performed by a single computer 22. However, the technology of this disclosure is not limited thereto, and a distributed processing of the specific process may be performed by multiple computers, including computer 22. For example, a data generation model 58 may be provided in an external device of the data processing device 12, and the external device may generate data according to the input data.

[0772] In the above embodiment, an example was given in which the specific processing program 56 is stored in the storage 32, but the technology of this disclosure is not limited thereto. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-temporary storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-temporary storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes specific processing according to the specific processing program 56.

[0773] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.

[0774] Furthermore, it is not necessary to store the entirety of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store the entirety of the specific processing program 56 in the storage 32; it is acceptable to store only a portion of the specific processing program 56.

[0775] The following types of processors can be used as hardware resources to perform specific processing. Examples of processors include a CPU, a general-purpose processor that functions as a hardware resource to perform specific processing by executing software, i.e., a program. Other examples of processors include dedicated electrical circuits, such as FPGAs (Field-Programmable Gate Arrays), PLDs (Programmable Logic Devices), or ASICs (Application Specific Integrated Circuits), which have circuit configurations specifically designed to perform specific processing. All of these processors have built-in or connected memory, and all of them perform specific processing by using memory.

[0776] The hardware resource that performs a specific process may consist of one of these various processors, or it may consist of a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Alternatively, the hardware resource that performs a specific process may consist of a single processor.

[0777] Examples of configurations using a single processor include, firstly, a configuration in which one or more CPUs and software are combined to form a single processor, and this processor functions as a hardware resource that performs a specific process. Secondly, there is a configuration using a processor that realizes the functions of the entire system, including multiple hardware resources that perform a specific process, on a single IC chip, as exemplified by SoCs (System-on-a-chip). In this way, a specific process is realized using one or more of the above types of processors as hardware resources.

[0778] Furthermore, the hardware structure of these various processors can more specifically utilize electrical circuits that combine circuit elements such as semiconductor devices. Also, the specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps can be deleted, new steps added, or the processing order rearranged, as long as it does not deviate from the main purpose.

[0779] The descriptions and illustrations presented above are detailed explanations of the technical aspects of this disclosure and are merely examples of the technical aspects. For example, the above descriptions of the structure, function, operation, and effect are examples of the structure, function, operation, and effect of the technical aspects of this disclosure. Therefore, it goes without saying that you may delete unnecessary parts, add new elements, or replace elements in the descriptions and illustrations presented above, as long as you do not deviate from the essence of the technical aspects of this disclosure. Furthermore, in order to avoid confusion and facilitate understanding of the technical aspects of this disclosure, explanations of common technical knowledge and the like that do not require special explanation to enable the implementation of the technical aspects of this disclosure have been omitted from the descriptions and illustrations presented above.

[0780] All documents, patent applications, and technical standards described herein are incorporated by reference to the same extent as if each individual document, patent application, and technical standard were specifically and individually noted to be incorporated by reference.

[0781] The following is further disclosed regarding the embodiments described above.

[0782] (Claim 1)

[0783] Video footage of logistics operations activities obtained based on order identification information,

[0784] A means for automatically extracting relevant video fragments by analyzing them using a generative model,

[0785] Means for storing video fragments related to the corresponding order identification information in a storage device,

[0786] A means for providing video fragments related to order identification information based on a search request from a user,

[0787] A system that includes this.

[0788] (Claim 2)

[0789] The system according to claim 1, wherein the generation model has means for recognizing logistics operation activities for each frame of video fragments and assigning corresponding order identification information.

[0790] (Claim 3)

[0791] The system according to claim 1, comprising means for transmitting video fragments to a user terminal in streaming or download format.

[0792] "Example 1"

[0793] (Claim 1)

[0794] A means for receiving visual information acquired from multiple monitoring devices in real time, effectively compressing it, and storing it in a storage device,

[0795] A means for analyzing received visual information using a generative model, identifying logistics operation activities, and automatically assigning related identification information,

[0796] A means for quickly searching and providing visual information fragments organized based on identification information in response to usage requests,

[0797] A system that includes this.

[0798] (Claim 2)

[0799] The system according to claim 1, wherein the generative model has means for analyzing each frame of visual information fragments to recognize logistics operation activities and assigning corresponding identification information.

[0800] (Claim 3)

[0801] The system according to claim 1, comprising means for transmitting visual information fragments corresponding to user information processing devices in real time.

[0802] "Application Example 1"

[0803] (Claim 1)

[0804] Video footage of logistics operations activities obtained based on order identification information,

[0805] A means for automatically extracting relevant video fragments by analyzing them using a generative model,

[0806] A means for storing video fragments related to the corresponding order identification information on an electronic storage medium,

[0807] A means for providing video fragments related to order identification information based on a search request from a user,

[0808] A means for acquiring order identification information using an audio input device and immediately displaying related video,

[0809] A system that includes this.

[0810] (Claim 2)

[0811] The system according to claim 1, wherein the generation model has means for recognizing logistics operation activities for each frame of video fragments and assigning corresponding order identification information.

[0812] (Claim 3)

[0813] The system according to claim 1, comprising means for transmitting video fragments to a user terminal in streaming or download format.

[0814] "Example 2 of combining an emotion engine"

[0815] (Claim 1)

[0816] Video footage of logistics operations activities obtained based on order identification information,

[0817] A means for automatically extracting relevant video fragments by analyzing them using a generative AI model,

[0818] Means for storing video fragments related to the corresponding order identification information in a storage device,

[0819] A means for providing video fragments related to order identification information based on a search request from a user,

[0820] A means of adjusting the user interface for search requests and information provision using an emotion engine that analyzes user emotions,

[0821] A system that includes this.

[0822] (Claim 2)

[0823] The system according to claim 1, wherein the generating AI model has means for recognizing logistics operation activities for each frame of a video fragment and assigning corresponding order identification information.

[0824] (Claim 3)

[0825] The system according to claim 1, comprising means for transmitting video fragments to a user terminal in streaming or download format and for customizing the interface based on the output of an emotion engine.

[0826] "Application example 2 when combining with an emotional engine"

[0827] (Claim 1)

[0828] Video information of logistics operations obtained based on order identification information,

[0829] A means of automatically identifying the relevant video elements by analyzing them using a generative model,

[0830] Means for storing video elements related to the corresponding order identification information in a storage device,

[0831] A means for providing video data related to order identification information based on a search request from a user,

[0832] A means of identifying a user's emotions using an emotion analysis engine and customizing the user interface,

[0833] A means of displaying emotion-responsive instruction information along with video data to a display device used in a logistics center,

[0834] A system that includes this.

[0835] (Claim 2)

[0836] The system according to claim 1, wherein the generation model has means for identifying logistics operation activities for each frame of video elements and attaching corresponding order identification information.

[0837] (Claim 3)

[0838] The system according to claim 1, comprising means for transmitting video elements corresponding to a user terminal in streaming or download format. [Explanation of Symbols]

[0839] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Devices 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robots< / url:> < / url:> < / url:> < / url:>

Claims

1. Video footage of logistics operations activities obtained based on order identification information, A means for automatically extracting relevant video fragments by analyzing them using a generative model, Means for storing video fragments related to the corresponding order identification information in a storage device, A means for providing video fragments related to order identification information based on a search request from a user, A system that includes this.

2. The system according to claim 1, wherein the generation model has means for recognizing logistics operation activities for each frame of video fragments and assigning corresponding order identification information.

3. The system according to claim 1, further comprising means for transmitting video fragments corresponding to a user terminal in streaming or download format.

Citation Information

Patent Citations

  • Persona chatbot control method and system

    JP2022180282A