system
An automated baggage inspection system using generative AI and user feedback improves detection accuracy and efficiency by preprocessing video data from inspection devices, addressing human reliance and security vulnerabilities.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- SOFTBANK GROUP CORP
- Filing Date
- 2024-10-15
- Publication Date
- 2026-04-27
AI Technical Summary
Current baggage inspection systems at airports and transportation hubs rely heavily on human visual inspection, leading to inefficiencies, errors, and security vulnerabilities.
An automated system that collects and preprocesses video data from baggage inspection devices, uses a generative AI model to detect dangerous goods, and incorporates user verification feedback to improve the AI model's accuracy.
Enhances security and efficiency by automating baggage inspection, optimizing human resources, and continuously improving detection accuracy through user feedback.
Smart Images

Figure 2026070107000001_ABST
Abstract
Description
Technical Field
[0001] The technology of the present disclosure relates to a system.
Background Art
[0002] Patent Document 1 discloses a persona chatbot control method performed by at least one processor, including steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to an explanation of a chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance.
Prior Art Documents
Patent Documents
[0003]
Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0004] In the past, baggage inspection at airports and transportation hubs has relied on visual inspection by humans, making it a challenge to secure human resources. Also, problems such as work errors, delays during congestion, and security defects have emerged. Therefore, there is a need for an automated system that combines accuracy and safety while improving inspection efficiency.
Means for Solving the Problems
[0005] This invention comprises means for collecting video data from a baggage inspection device and preprocessing it into a format suitable for analysis, and means for automatically detecting dangerous goods or prohibited items using a generating AI model. Furthermore, the system configuration determines the degree of danger based on the detection results and notifies the user of the results, thereby streamlining human verification. In addition, the system utilizes user verification results as feedback to retrain the generating AI model, continuously improving the accuracy and efficiency of the analysis. This automates the baggage inspection process, thereby improving security and service.
[0006] A "baggage inspection device" is a device used to check the contents of baggage, and typically uses X-rays or other sensors to visualize the inside.
[0007] "Video data" refers to visual information acquired by baggage inspection devices, and serves as basic data for analyzing the condition of the inside of baggage.
[0008] "Preprocessing" refers to the process of converting video data into a format suitable for analysis, and includes steps such as noise reduction and resolution adjustment.
[0009] A "generative AI model" is an artificial intelligence model trained using machine learning algorithms, and is a technology used to recognize objects in video and detect dangerous or prohibited items.
[0010] "Dangerous goods or prohibited items" refer to items that are restricted from being brought into airports or on transportation, and include items considered to pose a safety risk, such as knives and explosives.
[0011] "Detection results" refer to information obtained after analysis by a generative AI model, and are output containing information about the identified object.
[0012] "Feedback" refers to the confirmation results provided by users, and this data is used to retrain the generated AI model to improve its accuracy.
[0013] "Retraining" is the process of improving the performance of an existing generative AI model by adding new data and training it. [Brief explanation of the drawing]
[0014] [Figure 1] This is a conceptual diagram showing an example of the configuration of a data processing system according to the first embodiment. [Figure 2] This is a conceptual diagram showing an example of the essential functions of a data processing device and a smart device according to the first embodiment. [Figure 3] This is a conceptual diagram showing an example of the configuration of a data processing system according to the second embodiment. [Figure 4] This is a conceptual diagram showing an example of the main functions of a data processing device and smart glasses according to the second embodiment. [Figure 5] This is a conceptual diagram showing an example of the configuration of a data processing system according to the third embodiment. [Figure 6] This is a conceptual diagram showing an example of the main functions of a data processing device and a headset-type terminal according to the third embodiment. [Figure 7] This is a conceptual diagram showing an example of the configuration of a data processing system according to the fourth embodiment. [Figure 8] This is a conceptual diagram showing an example of the main functions of a data processing device and a robot according to the fourth embodiment. [Figure 9] This shows an emotion map where multiple emotions are mapped. [Figure 10] This shows an emotion map where multiple emotions are mapped. [Figure 11] This is a sequence diagram showing the processing flow of the data processing system in Example 1. [Figure 12] This is a sequence diagram showing the processing flow of the data processing system in Application Example 1. [Figure 13] This is a sequence diagram showing the processing flow of the data processing system in Example 2 when an emotion engine is combined. [Figure 14]It is a sequence diagram showing the processing flow of a data processing system in Application Example 2 when a sentiment engine is combined.
Embodiments for Carrying Out the Invention
[0015] Hereinafter, an example of an embodiment of a system according to the technology of the present disclosure will be described with reference to the accompanying drawings.
[0016] First, the terms used in the following description will be explained.
[0017] In the following embodiments, a numbered processor (hereinafter simply referred to as "processor") may be a single arithmetic unit or a combination of multiple arithmetic units. Also, the processor may be a single type of arithmetic unit or a combination of multiple types of arithmetic units. Examples of arithmetic units include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), an APU (Accelerated Processing Unit), and the like.
[0018] In the following embodiments, a numbered RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a work memory by the processor.
[0019] In the following embodiments, a numbered storage is one or more non-volatile storage devices that store various programs and various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), or magnetic tapes, and the like.
[0020] In the following embodiments, the signed communication interface (I / F) is an interface that includes a communication processor and an antenna, etc. The communication interface manages communication between multiple computers. Examples of communication standards applicable to the communication interface include wireless communication standards such as 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), or Bluetooth (registered trademark).
[0021] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." That is, "A and / or B" means that it may be A alone, or B alone, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" applies when expressing three or more things linked by "and / or."
[0022] [First Embodiment]
[0023] Figure 1 shows an example of the configuration of the data processing system 10 according to the first embodiment.
[0024] As shown in Figure 1, the data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.
[0025] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0026] The smart device 14 comprises a computer 36, a reception device 38, an output device 40, a camera 42, and a communication interface 44. The computer 36 comprises a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The reception device 38, output device 40, and camera 42 are also connected to the bus 52.
[0027] The reception device 38 is equipped with a touch panel 38A and a microphone 38B, etc., and receives user input. The touch panel 38A receives user input by detecting contact with an object (e.g., a pen or finger). The microphone 38B receives user input by detecting the user's voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.
[0028] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form perceptible to the user 20 (e.g., audio and / or text). The display 40A displays visible information such as text and images according to instructions from the processor 46. The speaker 40B outputs audio according to instructions from the processor 46. The camera 42 is a small digital camera equipped with an optical system such as a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.
[0029] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various types of information between processor 46 and processor 28 via network 54.
[0030] Figure 2 shows an example of the main functions of the data processing device 12 and the smart device 14.
[0031] As shown in Figure 2, in the data processing device 12, a specific processing is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" related to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.
[0032] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0033] In the smart device 14, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The reception output program 60 is used in conjunction with a specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0034] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".
[0035] This invention relates to a system for automating baggage screening at airports and transportation hubs. This system primarily consists of a server, terminals, and users. The program processing and specific examples are described below.
[0036] First, the server collects video data from baggage inspection devices in real time. This video data is based on X-ray and 3D images and provides a detailed representation of the inside of the baggage. This data is first pre-processed and converted into a format suitable for analysis. Pre-processing includes noise reduction, image normalization, and resolution adjustment.
[0037] Next, the pre-processed video data is input into a generating AI model. This AI model is pre-trained using object recognition technology and analyzes objects in the video to detect hazardous materials and prohibited items. The server receives the analysis results from the generating AI and makes a determination based on whether the detected items are dangerous. This determination is then generated as an alarm according to the level of risk.
[0038] The server notifies the terminal of the detection results, making them available for viewing by security staff users. The results are presented in text and visual form, showing the type and location of the hazardous material. Users can use this information to perform manual verification as needed.
[0039] Furthermore, the results manually verified by the user are sent to the server as feedback. This feedback information is used to retrain the generating AI model, contributing to improved accuracy in subsequent analyses.
[0040] As a concrete example, consider a scene from a baggage inspection at an airport. When a traveler's luggage passes through the inspection machine, the server immediately acquires video data, which is then analyzed by a generative AI that detects the presence of a knife. The server sends an alarm to the terminal indicating the presence of a knife, and security staff ask the person being inspected to check their luggage and actually remove the knife for verification. This verification result is sent to the server as feedback, and the generative AI learns to perform more accurate analyses.
[0041] This system enables efficient and accurate baggage screening, optimizing human resources and improving passenger security. Furthermore, the system is highly scalable, making it easy to implement at different airports and transportation hubs.
[0042] The following describes the processing flow.
[0043] Step 1:
[0044] The server acquires video data from baggage inspection devices in real time. This data visualizes the inside of the baggage using X-ray sensors and other imaging technologies.
[0045] Step 2:
[0046] The server preprocesses the acquired video data. This preprocessing involves noise reduction, image normalization, and resolution adjustment, and converting the data into the optimal format for analysis.
[0047] Step 3:
[0048] The server inputs pre-processed data into a generative AI model. The generative AI model is pre-trained and uses object recognition technology to analyze objects in the video and detect hazardous materials and prohibited items.
[0049] Step 4:
[0050] The server determines the level of danger of the detected items based on the analysis results generated by the AI. Based on the results, it generates an alarm corresponding to the degree of risk.
[0051] Step 5:
[0052] The server notifies the terminal of the judgment result. The notification is provided as text and visual information so that security staff can review the content.
[0053] Step 6:
[0054] Users (security staff) can check the detection results on their terminals and, if necessary, manually inspect baggage. Actions taken after the inspection will be determined according to the situation at hand.
[0055] Step 7:
[0056] The results of manual verification performed by the user are fed back to the server via the terminal. This feedback is used as data to improve the accuracy of the generated AI model.
[0057] Step 8:
[0058] The server utilizes the collected feedback to retrain the generated AI model. This improves the model's analytical accuracy and efficiency, strengthening the reliability of future inspections.
[0059] (Example 1)
[0060] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."
[0061] Traditional baggage inspections required significant human resources to detect dangerous and prohibited items, resulting in limitations in efficiency and accuracy. Furthermore, continuous system improvements to enhance the accuracy of dangerous item detection were difficult.
[0062] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0063] In this invention, the server includes means for collecting image information obtained from a video acquisition device, means for preprocessing the image information into a format suitable for analysis, and means for inputting the preprocessed image information into a generating AI model to identify items that pose a safety problem. This makes it possible to detect hazardous materials efficiently and accurately.
[0064] An "image acquisition device" is a device installed to acquire image information of luggage, and mainly refers to equipment that visualizes the internal structure using X-rays or 3D images.
[0065] "Image information" refers to digital data collected by video acquisition devices to represent the inside of luggage, and this is the subject of analysis.
[0066] "Preprocessing" refers to the process of converting acquired image information into a format suitable for analysis, such as by removing noise or adjusting the resolution.
[0067] A "generative AI model" refers to an artificial intelligence model that is trained using machine learning techniques and has the ability to identify items with safety issues from image information.
[0068] "Items with safety issues" refer to items that are prohibited from being brought in or items that pose a risk, and which require action after detection.
[0069] "Feedback" refers to the flow of information sent back to the server by operators for manual verification and to improve detection accuracy, which is then used to retrain the generated AI model.
[0070] This invention is a system for streamlining and improving the accuracy of baggage screening at airports and transportation hubs. It mainly consists of a server, terminals, and users.
[0071] The server first collects image information obtained from the video acquisition device. This image information is used to visualize the inside of the luggage using X-rays and 3D images. The server then preprocesses the image information, applying noise reduction and resolution adjustments to convert it into a format suitable for analysis. This preprocessing prepares the generated AI model for effective analysis of the image information.
[0072] The pre-processed image information is input by the server into a generating AI model. This AI model is trained using machine learning techniques and has the ability to identify items that pose a safety risk. For example, it can detect prohibited items such as knives and explosives. Based on these detections, the server evaluates the level of safety and notifies the terminal of the evaluation result.
[0073] Users of the terminal review the notified evaluation results and manually check their luggage if necessary. The user's verification results are sent to the server as feedback. This feedback information is used to retrain the generating AI model and improve the accuracy of the analysis.
[0074] As a concrete example, during baggage screening at an airport, when a passenger's luggage passes through the screening device, the server immediately acquires image information and analyzes it using a generative AI model. If the analysis determines that the luggage may contain dangerous materials, that information is sent to a terminal. Security staff receive the warning from the terminal and take steps to manually check the luggage. Furthermore, as an example of a prompt message to improve detection accuracy, instructions such as "Identify suspicious items in the luggage" can be provided to the generative AI model.
[0075] This invention improves the accuracy and efficiency of baggage inspection, enhancing security while optimizing human resources.
[0076] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0077] Step 1:
[0078] The server acquires image information about the luggage from the image acquisition device. The input at this time is raw data of X-rays and 3D images. This data is used to visualize the internal structure of the luggage. The server receives this image data and prepares it for preprocessing in the next processing step.
[0079] Step 2:
[0080] The server preprocesses the acquired image information. The input is the raw image data received in the previous step. The server applies a noise reduction filter and adjusts the resolution to a predetermined setting to convert the image into a format suitable for analysis. This preprocessing removes unnecessary information from the image, making it suitable for recognition by the generated AI model. The output is clean, preprocessed image data.
[0081] Step 3:
[0082] The server inputs pre-processed image data into the generative AI model. The input is pre-processed image data. The server passes this data to the generative AI model for machine learning analysis. The generative AI model uses its pre-trained object recognition capabilities to identify items that pose a safety risk, such as knives or explosives. The output is the analysis results, including the detected items and their location information.
[0083] Step 4:
[0084] The server evaluates the level of safety based on the analysis results from the generated AI model and notifies the terminal of the evaluation result. The input is the analysis result obtained in the previous step. The server analyzes the analysis result and determines the level of risk. Based on the determination result, it sets an alarm level and sends that information to the terminal. The output is alarm information that reflects the risk evaluation.
[0085] Step 5:
[0086] The terminal displays alarm information to the security staff, who are the users. The input is alarm information sent from the server. The terminal displays this information to the user as text and visuals, clearly indicating the type and location of the detected item. The user then performs further manual verification based on this information.
[0087] Step 6:
[0088] The user performs a manual check and sends the results to the server as feedback. The input is the details of the check performed by the user. The user checks their luggage according to the terminal information and checks for the presence of actual dangerous goods. The check results are recorded as feedback and sent to the server. The output is the feedback information resulting from the check.
[0089] Step 7:
[0090] The server receives feedback information and uses it to retrain the generative AI model. The input is the feedback information. Based on this feedback, the server retrains the model to improve the accuracy of the next analysis. The output is the improved generative AI model.
[0091] (Application Example 1)
[0092] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."
[0093] Baggage screening is essential for ensuring the safety of public and commercial facilities, but current methods are time-consuming and consume a large amount of human resources. Furthermore, the accuracy of detecting dangerous goods varies depending on the situation, and there is a need for efficient and highly accurate automation. In addition, the insufficient real-time situation assessment and rapid response also pose a safety challenge.
[0094] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0095] In this invention, the server includes means for collecting video data acquired from a baggage inspection device, means for preprocessing the video data into a format suitable for analysis, means for inputting the preprocessed video data into a generating AI model to detect dangerous or prohibited items, means for photographing items at the entrance of a public facility using a camera on an individual device and transmitting the data to a cloud server, means for performing data analysis on the cloud server and determining the danger of the items, and means for presenting the analysis results to a user device in real time and issuing an alarm. This makes it possible to efficiently and accurately automate the verification of baggage safety.
[0096] A "baggage inspection device" is a device used to check the internal condition of baggage and has the function of acquiring video data.
[0097] "Video data" refers to data including X-ray images and 3D images of baggage acquired by baggage inspection equipment.
[0098] A "generative AI model" is a model that is pre-trained using artificial intelligence technology to analyze objects from video data and detect hazardous materials or prohibited items.
[0099] A "cloud server" is an external computing resource that can be accessed via a network and is used for data analysis and information storage.
[0100] A "user device" is a terminal device that allows users to view analysis results and provide feedback as needed, and typically refers to a personal digital assistant (PDI).
[0101] "Preprocessing" refers to a series of processes performed to convert video data into a format suitable for analysis.
[0102] "Feedback" refers to the confirmation results provided by users, and this information is used to retrain the generated AI model.
[0103] "Dangerous goods" are items that may threaten public safety and are subject to detection.
[0104] "Notification" refers to the act of informing the user of detected information, and is the result provided as an alarm.
[0105] This invention is a system designed to streamline baggage inspection. The server collects video data from baggage inspection devices in real time and preprocesses it into a format suitable for analysis. Preprocessing includes noise reduction, image normalization, and resolution adjustment.
[0106] The pre-processed data is input into a generating AI model on a cloud server. This AI model is trained using object recognition technology and has the ability to analyze and detect hazardous materials and prohibited items from video data. The analysis results determine the level of danger of the items and are sent from the server in a format that notifies the user's device. The notification is displayed in real time on the user's device and can be accessed via a portable terminal.
[0107] Based on this notification, the user reviews the baggage inspection results and performs additional manual inspections if necessary. After the review is complete, the results are sent back to the server as feedback. This feedback data is used to continuously train the generative AI model, enabling improvements in analysis accuracy.
[0108] A concrete example is a baggage inspection scene at the entrance of a public facility. When a customer enters the facility, they open a dedicated smartphone app and take a picture of their baggage with the camera. This application connects to a cloud server and detects any dangerous items hidden in the baggage. An example of a prompt message used would be, "Please analyze whether this baggage contains any dangerous items." This enables quick and appropriate safety checks.
[0109] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0110] Step 1:
[0111] The server collects video data from the baggage inspection device. The input consists of X-ray images and 3D images from the baggage inspection device. The server receives these images and temporarily stores them in preparation for the next processing step. The output is image data for the pre-processing stage.
[0112] Step 2:
[0113] The server preprocesses the received video data. It performs noise reduction, image normalization, and resolution adjustment, and converts it into a format suitable for analysis. The input is the original image data, which is the output of step 1, and the output is preprocessed data suitable for the generative AI model.
[0114] Step 3:
[0115] The server inputs pre-processed video data into a generating AI model. The generating AI model analyzes the data and detects hazardous materials and prohibited items. The input is the pre-processed data which is the output of step 2, and the output is a list of detected items and their attribute data.
[0116] Step 4:
[0117] The server determines the level of risk of an item based on the analysis results from the generated AI model. The determination is expressed as a score or alarm corresponding to the degree of risk. The input is the detection result from step 3, and the output is alarm data based on the level of risk.
[0118] Step 5:
[0119] The terminal receives alarm data from the server and notifies the user in real time. Notifications are made in the form of voice, pop-up messages, warning lights, etc. The input is the alarm data which is the output of step 4, and the output on the terminal is the alarm display for the user.
[0120] Step 6:
[0121] The user reviews the notification and performs manual inspections if necessary. The results of the review are sent to the server as feedback. The input is the output from step 5, and the output is text data or image data of the reviewed content.
[0122] Step 7:
[0123] The server receives feedback from the user and uses it to retrain the generative AI model. The feedback data is added to the training dataset of the generative AI model, improving the accuracy of the analysis. The input is the feedback data, which is the output of step 6, and the output is the updated parameters of the AI model.
[0124] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0125] This invention relates to an automated baggage inspection system that incorporates an emotion engine. This system consists of collaborative work between a baggage inspection device, a server, a terminal, and a user. The program processing and specific examples are described below.
[0126] The server first acquires video data from the baggage inspection device. This data includes X-ray and visible light images, showing the contents of the baggage in detail. The server then preprocesses this data into a format that is easy to analyze. Specifically, this involves processes such as noise reduction and image quality improvement.
[0127] Next, the server inputs the pre-processed video data into a generating AI model. This model analyzes and detects hazardous materials and prohibited items based on object recognition technology. Based on the results, the server evaluates the level of danger and generates a result. The generated result is then sent to the terminal.
[0128] This system incorporates an emotion engine that monitors the emotions of the security staff users in real time. It analyzes the user's voice and facial expressions to understand their emotions based on the results displayed on the terminal, and then weights the feedback accordingly. For example, if a user expresses concern, that feedback becomes more important in retraining the generative AI model.
[0129] Furthermore, if the emotion engine indicates that the user's emotions are anxious or alarming, the server automatically triggers additional testing procedures. This process enhances the response to potential risks.
[0130] As a concrete example, when a passenger's baggage passes through the inspection device, the server acquires the video data and begins analysis. As the detected results are displayed on the terminal, the emotion engine analyzes the user's reaction. For example, if the user shows tension or anxiety, the server issues an additional alert and prompts further manual verification. This entire process ensures that inspections are precise and flexible.
[0131] This system configuration further enhances the security of airports and transportation systems, streamlines security processes, and offers flexibility to adapt to diverse environments.
[0132] The following describes the processing flow.
[0133] Step 1:
[0134] The server acquires video data from baggage inspection devices in real time. This data is acquired using X-ray technology to visualize the inside of the baggage.
[0135] Step 2:
[0136] The server preprocesses the acquired video data and converts it into a format suitable for analysis. Specifically, it removes noise and adjusts the image contrast and resolution.
[0137] Step 3:
[0138] The server inputs pre-processed data into a generating AI model. This model uses object recognition technology to analyze objects in the video and detect hazardous materials and prohibited items.
[0139] Step 4:
[0140] The server determines the level of risk of an item based on the analysis results of the generated AI model and sends the result to the terminal. This information includes the type and location of the item, as well as the risk level.
[0141] Step 5:
[0142] The terminal displays the received inspection results to the user, who is a security staff member. Here, the user can visually confirm the results.
[0143] Step 6:
[0144] The emotion engine analyzes the user's voice and facial expressions in real time to determine their emotional state. Emotional data is recorded for feedback.
[0145] Step 7:
[0146] The user reviews the results and provides feedback based on their reaction. This feedback, along with the analysis results from the emotion engine, is sent to the server.
[0147] Step 8:
[0148] The server receives user feedback and uses it as retraining data for the generated AI model. This process aims to improve the model's accuracy.
[0149] Step 9:
[0150] The server automatically triggers additional checks if the emotion engine detects any emotions such as anxiety or tension in the user. This strengthens the system's ability to address potential security risks.
[0151] (Example 2)
[0152] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".
[0153] In baggage inspections, there is a need for reliable detection of dangerous and prohibited items, as well as improvements in inspection speed and efficiency. However, current systems have limitations in improving detection accuracy, and in particular, they cannot consider the influence of user emotions on inspection results, resulting in insufficient risk management.
[0154] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0155] In this invention, the server includes means for collecting video data acquired from a baggage inspection device, means for preprocessing the video data into a format suitable for analysis, means for inputting the preprocessed video data into a generating AI model to detect dangerous goods or prohibited items, and means for analyzing the user's voice and facial expressions using an emotion analysis device to evaluate emotions. This improves the accuracy of dangerous goods detection and enables flexible adjustment of inspection procedures that take into account the user's emotions.
[0156] A "baggage inspection device" is a technological device used at airports, train stations, and other locations to examine passengers' baggage and check for the presence of dangerous goods or prohibited items.
[0157] "Visual data" refers to digital information generated by X-rays or other methods to visually show the detailed condition of items in baggage.
[0158] A "generative AI model" is a form of artificial intelligence trained using a deep learning framework, which utilizes object recognition technology to identify and classify items within luggage.
[0159] An "emotion analysis device" is a system component that detects and analyzes a user's voice tone, facial expressions, etc., to evaluate their emotions.
[0160] "Risk assessment" is the process of evaluating the hazardous nature of detected items and determining subsequent actions based on that assessment.
[0161] An embodiment of this invention consists of a baggage inspection device, a server, a terminal, and collaborative work by a user.
[0162] The server first acquires X-ray and visible light image data from the baggage inspection device. This data is used to visually show the contents of the baggage in detail and to check for the presence of dangerous or prohibited items. The server then performs preprocessing on this data, such as noise reduction and image quality enhancement, to prepare it for analysis. Image processing libraries are among the software used at this stage.
[0163] Next, the server inputs the pre-processed video data into a generating AI model. This model is built on deep learning frameworks such as TENSORFLOW® and PyTorch, and uses object recognition technology to identify dangerous or prohibited items in the luggage. It then evaluates the degree of danger based on the identified items and generates an evaluation result. The generated result is transmitted to the terminal in real time and presented to the user.
[0164] The terminal is also equipped with an emotion analyzer that analyzes the user's voice and facial expressions in real time to understand their emotional state. For example, if the user shows anxiety or apprehension, that information is fed back to the server and used to adjust the examination procedure. This enables a more flexible and precise response.
[0165] As a concrete example, when a passenger's baggage passes through an inspection device, the server acquires the video data and initiates a process to detect dangerous items. The detection results are displayed on the terminal, and simultaneously, sentiment analysis is performed to analyze the user's reaction. If the user shows signs of anxiety, the server recommends further verification steps, enabling dynamic risk management.
[0166] Examples of prompt messages include phrases like, "Identify dangerous items in passengers' baggage and adjust the inspection procedure based on the user's emotional response." In this way, security processes at airports and transportation facilities are conducted efficiently and safely.
[0167] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0168] Step 1:
[0169] The server acquires X-ray and visible light video data from the baggage inspection device. This data provides detailed information about the items inside the baggage and is transmitted to the server as a multidimensional array. The input video data is denoised, processed to improve image quality, and converted into a format suitable for analysis. The output is pre-processed, clear video data.
[0170] Step 2:
[0171] The server inputs pre-processed video data into a generating AI model. The AI model, built using TensorFlow or PyTorch, utilizes object recognition technology to analyze dangerous or prohibited items in luggage. As a data computation, the model identifies patterns through numerous neuron layers. The output is the identification result regarding dangerous items and its detection accuracy.
[0172] Step 3:
[0173] The server evaluates the risk level of the target based on the identification results obtained from the AI model and calculates the risk level. The risk assessment is calculated according to pre-set criteria based on the content of the identification results. A detailed report summarizing this assessment result is generated and sent to the terminal. The output is a report including the risk level.
[0174] Step 4:
[0175] The terminal receives a report sent from the server and displays it to the user. At this time, an emotion analysis device is activated, capturing the user's voice tone and facial expressions in real time. Based on the captured emotion data, the user's response is analyzed, and their emotional state is estimated. The output is the current emotion evaluation result.
[0176] Step 5:
[0177] The server receives the sentiment assessment results sent as feedback from the terminal and dynamically adjusts the testing procedure. Specifically, if the user's emotions indicate anxiety or caution, it triggers additional testing or a more detailed review. This adjustment enables flexible and effective risk management. The output is the adjusted testing procedure.
[0178] (Application Example 2)
[0179] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as a "server" and the smart device 14 as a "terminal".
[0180] In modern baggage screening systems, the detection of dangerous or prohibited items still largely relies on human judgment, and the accuracy and efficiency of detection have not improved sufficiently. Furthermore, the emotions and cognitive state of security staff conducting the inspections can influence their ability to respond to potential risks. However, there is a lack of means to effectively incorporate these factors and improve accuracy through real-time analysis and feedback.
[0181] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[0182] In this invention, the server includes means for collecting video information acquired from a baggage inspection device, means for pre-processing the video information into a format suitable for analysis, means for inputting the pre-processed video information into a generating AI algorithm to detect dangerous or prohibited items, and means for presenting the results in real time using a visual display worn by the user and analyzing the user's reaction. This enables advanced detection of dangerous items that takes into account the emotional state of security staff and allows for faster countermeasures.
[0183] A "baggage inspection device" is a device that scans passenger and cargo luggage and makes visible the items contained inside.
[0184] "Visual information" refers to digital video data acquired by baggage inspection equipment, including X-ray images and visible light images.
[0185] A "generative AI algorithm" is a program that uses artificial intelligence technology to extract features from input data and identify dangerous or prohibited items.
[0186] "Object recognition techniques" are technologies that use algorithms to analyze and identify specific objects within images or videos.
[0187] A "visual display" is a device incorporated into devices such as smart glasses that visually presents information to the user.
[0188] "User emotional data" refers to data on the emotional state of the user, evaluated from their facial expressions and voice, acquired through visual displays and other sensors.
[0189] "Feedback" refers to information that is used to improve system operation and algorithms based on user feedback.
[0190] This system is designed to provide baggage inspection with integrated anomaly detection and user assistance. The server first collects video information acquired by the baggage inspection device. This information is comprehensive image data, including X-ray and visible light images in particular.
[0191] The server first preprocesses the received video information. This processing includes noise reduction and image quality improvement. The goal is to prepare the data so that the generative AI algorithm can perform analysis with high accuracy. Next, the server sends the preprocessed video information to the generative AI algorithm for analysis to identify dangerous or prohibited items. This generative AI algorithm utilizes object recognition techniques and has the ability to detect targets efficiently and effectively.
[0192] The analysis results are presented in real time to the user wearing a visual display. This visual display is typically implemented as a device like smart glasses and is responsible for visualizing the information. The system also analyzes the user's facial expressions and voice tone to collect emotional data. This emotional data reflects the user's level of alertness and anxiety and is incorporated into the system as feedback.
[0193] As a concrete example, when airport security staff use glasses-type visual displays to inspect passengers' baggage, any abnormalities are immediately detected by the images displayed on the screen. If the staff member feels uneasy, the system triggers additional inspection procedures based on emotional data. This entire process ensures safer and more efficient baggage inspection.
[0194] An example of a prompt to the generating AI model is as follows: "Analyze the X-ray images of the luggage to detect dangerous items. Also, analyze the user's facial expression data to assess their emotions and evaluate their importance." This prompt allows the system to perform a more precise analysis using the AI algorithm.
[0195] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0196] Step 1:
[0197] The server acquires video information from the baggage inspection device. This information includes X-ray images and visible light images. The acquired video information is received as input data and passed to the next preprocessing step.
[0198] Step 2:
[0199] The server performs preprocessing on the acquired video information. This preprocessing includes noise reduction and image quality correction to improve image quality. Finally, the data is converted into a format that is easy for the generating AI algorithm to analyze. The output of this process is the preprocessed video information.
[0200] Step 3:
[0201] The server inputs pre-processed video information into a generating AI algorithm. The generating AI algorithm analyzes the data using object recognition techniques to detect dangerous or prohibited items. The results of the analysis are output as a list of detected items.
[0202] Step 4:
[0203] The server assesses the risk level of the detected items based on the list and prepares to notify the user. This assessment sorts the items in descending order of risk based on their characteristics. The assessment results are output as data for notification.
[0204] Step 5:
[0205] On the device, evaluation results are presented to the user via a visual display. The visual display shows information in real time on smart glasses. The user reviews the presented information and provides feedback in response. This feedback is recorded as input data to the system.
[0206] Step 6:
[0207] To collect user emotion data, the device analyzes the user's voice and facial expressions. The analysis uses an emotion analysis engine to evaluate the emotional state. The output of this analysis is the user's emotion parameters.
[0208] Step 7:
[0209] The server retrains the generating AI algorithm based on user feedback and emotion parameters. This retraining is aimed at improving analysis accuracy and efficiency, optimizing the performance of future baggage screening.
[0210] Step 8:
[0211] The server automatically triggers additional testing procedures based on the user's emotional state. Signs of anxiety or alertness identified through emotion analysis are treated as high risk, and immediate additional action is taken. The output of this step is an instruction to execute the automated additional procedure.
[0212] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0213] Data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of data generation model 58 is ChatGPT (registered trademark) (Internet search).<URL: https: / / openai.com / blog / chatgpt> ), Gemini (registered trademark) (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0214] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart device 14.
[0215] [Second Embodiment]
[0216] Figure 3 shows an example of the configuration of the data processing system 210 according to the second embodiment.
[0217] As shown in Figure 3, the data processing system 210 includes a data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.
[0218] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0219] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication interface 44. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, and camera 42 are also connected to the bus 52.
[0220] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[0221] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).
[0222] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[0223] Figure 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Figure 4, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[0224] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0225] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0226] In the smart glasses 214, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0227] Next, the identification processing performed by the identification processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".
[0228] This invention relates to a system for automating baggage screening at airports and transportation hubs. This system primarily consists of a server, terminals, and users. The program processing and specific examples are described below.
[0229] First, the server collects video data from baggage inspection devices in real time. This video data is based on X-ray and 3D images and provides a detailed representation of the inside of the baggage. This data is first pre-processed and converted into a format suitable for analysis. Pre-processing includes noise reduction, image normalization, and resolution adjustment.
[0230] Next, the pre-processed video data is input into a generating AI model. This AI model is pre-trained using object recognition technology and analyzes objects in the video to detect hazardous materials and prohibited items. The server receives the analysis results from the generating AI and makes a determination based on whether the detected items are dangerous. This determination is then generated as an alarm according to the level of risk.
[0231] The server notifies the terminal of the detection results, making them available for viewing by security staff users. The results are presented in text and visual form, showing the type and location of the hazardous material. Users can use this information to perform manual verification as needed.
[0232] Furthermore, the results manually verified by the user are sent to the server as feedback. This feedback information is used to retrain the generating AI model, contributing to improved accuracy in subsequent analyses.
[0233] As a concrete example, consider a scene from a baggage inspection at an airport. When a traveler's luggage passes through the inspection machine, the server immediately acquires video data, which is then analyzed by a generative AI that detects the presence of a knife. The server sends an alarm to the terminal indicating the presence of a knife, and security staff ask the person being inspected to check their luggage and actually remove the knife for verification. This verification result is sent to the server as feedback, and the generative AI learns to perform more accurate analyses.
[0234] This system enables efficient and accurate baggage screening, optimizing human resources and improving passenger security. Furthermore, the system is highly scalable, making it easy to implement at different airports and transportation hubs.
[0235] The following describes the processing flow.
[0236] Step 1:
[0237] The server acquires video data from baggage inspection devices in real time. This data visualizes the inside of the baggage using X-ray sensors and other imaging technologies.
[0238] Step 2:
[0239] The server preprocesses the acquired video data. This preprocessing involves noise reduction, image normalization, and resolution adjustment, and converting the data into the optimal format for analysis.
[0240] Step 3:
[0241] The server inputs pre-processed data into a generative AI model. The generative AI model is pre-trained and uses object recognition technology to analyze objects in the video and detect hazardous materials and prohibited items.
[0242] Step 4:
[0243] The server determines the level of danger of the detected items based on the analysis results generated by the AI. Based on the results, it generates an alarm corresponding to the degree of risk.
[0244] Step 5:
[0245] The server notifies the terminal of the judgment result. The notification is provided as text and visual information so that security staff can review the content.
[0246] Step 6:
[0247] Users (security staff) can check the detection results on their terminals and, if necessary, manually inspect baggage. Actions taken after the inspection will be determined according to the situation at hand.
[0248] Step 7:
[0249] The results of manual verification performed by the user are fed back to the server via the terminal. This feedback is used as data to improve the accuracy of the generated AI model.
[0250] Step 8:
[0251] The server utilizes the collected feedback to retrain the generated AI model. This improves the model's analytical accuracy and efficiency, strengthening the reliability of future inspections.
[0252] (Example 1)
[0253] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."
[0254] Traditional baggage inspections required significant human resources to detect dangerous and prohibited items, resulting in limitations in efficiency and accuracy. Furthermore, continuous system improvements to enhance the accuracy of dangerous item detection were difficult.
[0255] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0256] In this invention, the server includes means for collecting image information obtained from a video acquisition device, means for preprocessing the image information into a format suitable for analysis, and means for inputting the preprocessed image information into a generating AI model to identify items that pose a safety problem. This makes it possible to detect hazardous materials efficiently and accurately.
[0257] An "image acquisition device" is a device installed to acquire image information of luggage, and mainly refers to equipment that visualizes the internal structure using X-rays or 3D images.
[0258] "Image information" refers to digital data collected by video acquisition devices to represent the inside of luggage, and this is the subject of analysis.
[0259] "Preprocessing" refers to the process of converting acquired image information into a format suitable for analysis, such as by removing noise or adjusting the resolution.
[0260] A "generative AI model" refers to an artificial intelligence model that is trained using machine learning techniques and has the ability to identify items with safety issues from image information.
[0261] "Items with safety issues" refer to items that are prohibited from being brought in or items that pose a risk, and which require action after detection.
[0262] "Feedback" refers to the flow of information sent back to the server by operators for manual verification and to improve detection accuracy, which is then used to retrain the generated AI model.
[0263] This invention is a system for streamlining and improving the accuracy of baggage screening at airports and transportation hubs. It mainly consists of a server, terminals, and users.
[0264] The server first collects image information obtained from the video acquisition device. This image information is used to visualize the inside of the luggage using X-rays and 3D images. The server then preprocesses the image information, applying noise reduction and resolution adjustments to convert it into a format suitable for analysis. This preprocessing prepares the generated AI model for effective analysis of the image information.
[0265] The pre-processed image information is input by the server into a generating AI model. This AI model is trained using machine learning techniques and has the ability to identify items that pose a safety risk. For example, it can detect prohibited items such as knives and explosives. Based on these detections, the server evaluates the level of safety and notifies the terminal of the evaluation result.
[0266] Users of the terminal review the notified evaluation results and manually check their luggage if necessary. The user's verification results are sent to the server as feedback. This feedback information is used to retrain the generating AI model and improve the accuracy of the analysis.
[0267] As a concrete example, during baggage screening at an airport, when a passenger's luggage passes through the screening device, the server immediately acquires image information and analyzes it using a generative AI model. If the analysis determines that the luggage may contain dangerous materials, that information is sent to a terminal. Security staff receive the warning from the terminal and take steps to manually check the luggage. Furthermore, as an example of a prompt message to improve detection accuracy, instructions such as "Identify suspicious items in the luggage" can be provided to the generative AI model.
[0268] This invention improves the accuracy and efficiency of baggage inspection, enhancing security while optimizing human resources.
[0269] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0270] Step 1:
[0271] The server acquires image information about the luggage from the image acquisition device. The input at this time is raw data of X-rays and 3D images. This data is used to visualize the internal structure of the luggage. The server receives this image data and prepares it for preprocessing in the next processing step.
[0272] Step 2:
[0273] The server preprocesses the acquired image information. The input is the raw image data received in the previous step. The server applies a noise reduction filter and adjusts the resolution to a predetermined setting to convert the image into a format suitable for analysis. This preprocessing removes unnecessary information from the image, making it suitable for recognition by the generated AI model. The output is clean, preprocessed image data.
[0274] Step 3:
[0275] The server inputs pre-processed image data into the generative AI model. The input is pre-processed image data. The server passes this data to the generative AI model for machine learning analysis. The generative AI model uses its pre-trained object recognition capabilities to identify items that pose a safety risk, such as knives or explosives. The output is the analysis results, including the detected items and their location information.
[0276] Step 4:
[0277] The server evaluates the level of safety based on the analysis results from the generated AI model and notifies the terminal of the evaluation result. The input is the analysis result obtained in the previous step. The server analyzes the analysis result and determines the level of risk. Based on the determination result, it sets an alarm level and sends that information to the terminal. The output is alarm information that reflects the risk evaluation.
[0278] Step 5:
[0279] The terminal displays alarm information to the security staff, who are the users. The input is alarm information sent from the server. The terminal displays this information to the user as text and visuals, clearly indicating the type and location of the detected item. The user then performs further manual verification based on this information.
[0280] Step 6:
[0281] The user performs a manual check and sends the results to the server as feedback. The input is the details of the check performed by the user. The user checks their luggage according to the terminal information and checks for the presence of actual dangerous goods. The check results are recorded as feedback and sent to the server. The output is the feedback information resulting from the check.
[0282] Step 7:
[0283] The server receives feedback information and uses it for retraining the generative AI model. The input is the feedback information. Based on this feedback, the server retrains the model to improve the next analysis accuracy. The output is the improved generative AI model.
[0284] (Application Example 1)
[0285] Next, Application Example 1 will be described. In the following description, the data processing device 12 is referred to as the "server", and the smart glasses 214 are referred to as the "terminal".
[0286] Luggage inspection is essential for ensuring the safety of public facilities and commercial facilities. However, the current method has problems such as taking a long time and consuming a large amount of human resources. In addition, the detection accuracy of dangerous goods varies depending on the situation, and efficient and highly accurate automation is required. Furthermore, the lack of sufficient real-time situation awareness and prompt response is also a safety issue. )
[0287] The specific processing by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0288] In this invention, the server includes means for collecting video data acquired from a luggage inspection device, means for preprocessing the video data into a form suitable for analysis, means for inputting the preprocessed video data into a generative AI model to detect dangerous goods or prohibited items, means for using the camera of an individual device to photograph items at the entrance of a public facility and transmit the data to a cloud server, means for performing data analysis on the cloud server to determine the danger level of the items, and means for presenting the analysis result to a user device in real time and issuing an alarm. Thereby, it becomes possible to automatically and accurately confirm the safety of luggage efficiently.
[0289] The "luggage inspection device" is a device for checking the internal state of luggage and has a function of acquiring video data.
[0290] "Video data" refers to data including X-ray images and 3D images of baggage acquired by baggage inspection equipment.
[0291] A "generative AI model" is a model that is pre-trained using artificial intelligence technology to analyze objects from video data and detect hazardous materials or prohibited items.
[0292] A "cloud server" is an external computing resource accessible via a network, used for data analysis and information storage.
[0293] A "user device" is a terminal device that allows users to view analysis results and provide feedback as needed, and typically refers to a personal digital assistant (PDI).
[0294] "Preprocessing" refers to a series of processes performed to convert video data into a format suitable for analysis.
[0295] "Feedback" refers to the confirmation results provided by users, and this information is used to retrain the generated AI model.
[0296] "Dangerous goods" are items that may threaten public safety and are subject to detection.
[0297] "Notification" refers to the act of informing the user of detected information, and is the result provided as an alarm.
[0298] This invention is a system designed to streamline baggage inspection. The server collects video data from baggage inspection devices in real time and preprocesses it into a format suitable for analysis. Preprocessing includes noise reduction, image normalization, and resolution adjustment.
[0299] The pre-processed data is input into a generating AI model on a cloud server. This AI model is trained using object recognition technology and has the ability to analyze and detect hazardous materials and prohibited items from video data. The analysis results determine the level of danger of the items and are sent from the server in a format that notifies the user's device. The notification is displayed in real time on the user's device and can be accessed via a portable terminal.
[0300] Based on this notification, the user reviews the baggage inspection results and performs additional manual inspections if necessary. After the review is complete, the results are sent back to the server as feedback. This feedback data is used to continuously train the generative AI model, enabling improvements in analysis accuracy.
[0301] A concrete example is a baggage inspection scene at the entrance of a public facility. When a customer enters the facility, they open a dedicated smartphone app and take a picture of their baggage with the camera. This application connects to a cloud server and detects any dangerous items hidden in the baggage. An example of a prompt message used would be, "Please analyze whether this baggage contains any dangerous items." This enables quick and appropriate safety checks.
[0302] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0303] Step 1:
[0304] The server collects video data from the baggage inspection device. The input consists of X-ray images and 3D images from the baggage inspection device. The server receives these images and temporarily stores them in preparation for the next processing step. The output is image data for the pre-processing stage.
[0305] Step 2:
[0306] The server preprocesses the received video data. It performs noise removal, image normalization, resolution adjustment, and converts it into a format suitable for analysis. The input is the original image data which is the output of Step 1, and the output is the preprocessed data suitable for the generative AI model.
[0307] Step 3:
[0308] The server inputs the preprocessed video data into the generative AI model. The generative AI model analyzes the data and detects dangerous goods and prohibited items. The input is the preprocessed data which is the output of Step 2, and the output is a list of items as detection results and their attribute data.
[0309] Step 4:
[0310] Based on the analysis results from the generative AI model, the server determines the risk level of the items. The determination is expressed as a score or an alert according to the degree of risk. The input is the detection result of Step 3, and the output is the alert data based on the risk level. [[ID=十七]]
[0311] Step 5:
[0312] The terminal receives the alert data from the server and notifies the user in real time. The notification is made in the form of voice, pop-up messages, warning lamps, etc. The input is the alert data which is the output of Step 4, and the output on the terminal is the alert display for the user.
[0313] Step 6:
[0314] The user checks the notified content and performs a manual inspection if necessary. The check result is sent to the server as feedback. The input is the output of Step 5, and the output is the text data or image data of the confirmed content.
[0315] Step 7:
[0316] The server receives feedback from the user and uses it to retrain the generative AI model. The feedback data is added to the training dataset of the generative AI model, improving the accuracy of the analysis. The input is the feedback data, which is the output of step 6, and the output is the updated parameters of the AI model.
[0317] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0318] This invention relates to an automated baggage inspection system that incorporates an emotion engine. This system consists of collaborative work between a baggage inspection device, a server, a terminal, and a user. The program processing and specific examples are described below.
[0319] The server first acquires video data from the baggage inspection device. This data includes X-ray and visible light images, showing the contents of the baggage in detail. The server then preprocesses this data into a format that is easy to analyze. Specifically, this involves processes such as noise reduction and image quality improvement.
[0320] Next, the server inputs the pre-processed video data into a generating AI model. This model analyzes and detects hazardous materials and prohibited items based on object recognition technology. Based on the results, the server evaluates the level of danger and generates a result. The generated result is then sent to the terminal.
[0321] This system incorporates an emotion engine that monitors the emotions of the security staff users in real time. It analyzes the user's voice and facial expressions to understand their emotions based on the results displayed on the terminal, and then weights the feedback accordingly. For example, if a user expresses concern, that feedback becomes more important in retraining the generative AI model.
[0322] Furthermore, if the emotion engine indicates that the user's emotions are anxious or alarming, the server automatically triggers additional testing procedures. This process enhances the response to potential risks.
[0323] As a concrete example, when a passenger's baggage passes through the inspection device, the server acquires the video data and begins analysis. As the detected results are displayed on the terminal, the emotion engine analyzes the user's reaction. For example, if the user shows tension or anxiety, the server issues an additional alert and prompts further manual verification. This entire process ensures that inspections are precise and flexible.
[0324] This system configuration further enhances the security of airports and transportation systems, streamlines security processes, and offers flexibility to adapt to diverse environments.
[0325] The following describes the processing flow.
[0326] Step 1:
[0327] The server acquires video data from baggage inspection devices in real time. This data is acquired using X-ray technology to visualize the inside of the baggage.
[0328] Step 2:
[0329] The server preprocesses the acquired video data and converts it into a format suitable for analysis. Specifically, it removes noise and adjusts the image contrast and resolution.
[0330] Step 3:
[0331] The server inputs pre-processed data into a generating AI model. This model uses object recognition technology to analyze objects in the video and detect hazardous materials and prohibited items.
[0332] Step 4:
[0333] The server determines the level of risk of an item based on the analysis results of the generated AI model and sends the result to the terminal. This information includes the type and location of the item, as well as the risk level.
[0334] Step 5:
[0335] The terminal displays the received inspection results to the user, who is a security staff member. Here, the user can visually confirm the results.
[0336] Step 6:
[0337] The emotion engine analyzes the user's voice and facial expressions in real time to determine their emotional state. Emotional data is recorded for feedback.
[0338] Step 7:
[0339] The user reviews the results and provides feedback based on their reaction. This feedback, along with the analysis results from the emotion engine, is sent to the server.
[0340] Step 8:
[0341] The server receives user feedback and uses it as retraining data for the generated AI model. This process aims to improve the model's accuracy.
[0342] Step 9:
[0343] The server automatically triggers additional checks if the emotion engine detects any emotions such as anxiety or tension in the user. This strengthens the system's ability to address potential security risks.
[0344] (Example 2)
[0345] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".
[0346] In baggage inspections, there is a need for reliable detection of dangerous and prohibited items, as well as improvements in inspection speed and efficiency. However, current systems have limitations in improving detection accuracy, and in particular, they cannot consider the influence of user emotions on inspection results, resulting in insufficient risk management.
[0347] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0348] In this invention, the server includes means for collecting video data acquired from a baggage inspection device, means for preprocessing the video data into a format suitable for analysis, means for inputting the preprocessed video data into a generating AI model to detect dangerous goods or prohibited items, and means for analyzing the user's voice and facial expressions using an emotion analysis device to evaluate emotions. This improves the accuracy of dangerous goods detection and enables flexible adjustment of inspection procedures that take into account the user's emotions.
[0349] A "baggage inspection device" is a technological device used at airports, train stations, and other locations to examine passengers' baggage and check for the presence of dangerous goods or prohibited items.
[0350] "Visual data" refers to digital information generated by X-rays or other methods to visually show the detailed condition of items in baggage.
[0351] A "generative AI model" is a form of artificial intelligence trained using a deep learning framework, which utilizes object recognition technology to identify and classify items within luggage.
[0352] An "emotion analysis device" is a system component that detects and analyzes a user's voice tone, facial expressions, etc., to evaluate their emotions.
[0353] "Risk assessment" is the process of evaluating the hazardous nature of detected items and determining subsequent actions based on that assessment.
[0354] An embodiment of this invention consists of a baggage inspection device, a server, a terminal, and collaborative work by a user.
[0355] The server first acquires X-ray and visible light image data from the baggage inspection device. This data is used to visually show the contents of the baggage in detail and to check for the presence of dangerous or prohibited items. The server then performs preprocessing on this data, such as noise reduction and image quality enhancement, to prepare it for analysis. Image processing libraries are among the software used at this stage.
[0356] Next, the server inputs the pre-processed video data into a generating AI model. This model is built on deep learning frameworks such as TensorFlow and PyTorch, and uses object recognition technology to identify dangerous or prohibited items in the luggage. It then evaluates the level of danger based on the identified items and generates an evaluation result. The generated result is transmitted to the terminal in real time and presented to the user.
[0357] The terminal is also equipped with an emotion analyzer that analyzes the user's voice and facial expressions in real time to understand their emotional state. For example, if the user shows anxiety or apprehension, that information is fed back to the server and used to adjust the examination procedure. This enables a more flexible and precise response.
[0358] As a concrete example, when a passenger's baggage passes through an inspection device, the server acquires the video data and initiates a process to detect dangerous items. The detection results are displayed on the terminal, and simultaneously, sentiment analysis is performed to analyze the user's reaction. If the user shows signs of anxiety, the server recommends further verification steps, enabling dynamic risk management.
[0359] Examples of prompt messages include phrases like, "Identify dangerous items in passengers' baggage and adjust the inspection procedure based on the user's emotional response." In this way, security processes at airports and transportation facilities are conducted efficiently and safely.
[0360] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0361] Step 1:
[0362] The server acquires X-ray and visible light video data from the baggage inspection device. This data provides detailed information about the items inside the baggage and is transmitted to the server as a multidimensional array. The input video data is denoised, processed to improve image quality, and converted into a format suitable for analysis. The output is pre-processed, clear video data.
[0363] Step 2:
[0364] The server inputs pre-processed video data into a generating AI model. The AI model, built using TensorFlow or PyTorch, utilizes object recognition technology to analyze dangerous or prohibited items in luggage. As a data computation, the model identifies patterns through numerous neuron layers. The output is the identification result regarding dangerous items and its detection accuracy.
[0365] Step 3:
[0366] The server evaluates the risk level of the target based on the identification results obtained from the AI model and calculates the risk level. The risk assessment is calculated according to pre-set criteria based on the content of the identification results. A detailed report summarizing this assessment result is generated and sent to the terminal. The output is a report including the risk level.
[0367] Step 4:
[0368] The terminal receives a report sent from the server and displays it to the user. At this time, an emotion analysis device is activated, capturing the user's voice tone and facial expressions in real time. Based on the captured emotion data, the user's response is analyzed, and their emotional state is estimated. The output is the current emotion evaluation result.
[0369] Step 5:
[0370] The server receives the sentiment assessment results sent as feedback from the terminal and dynamically adjusts the testing procedure. Specifically, if the user's emotions indicate anxiety or caution, it triggers additional testing or a more detailed review. This adjustment enables flexible and effective risk management. The output is the adjusted testing procedure.
[0371] (Application Example 2)
[0372] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."
[0373] In modern baggage screening systems, the detection of dangerous or prohibited items still largely relies on human judgment, and the accuracy and efficiency of detection have not improved sufficiently. Furthermore, the emotions and cognitive state of security staff conducting the inspections can influence their ability to respond to potential risks. However, there is a lack of means to effectively incorporate these factors and improve accuracy through real-time analysis and feedback.
[0374] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[0375] In this invention, the server includes means for collecting video information acquired from a baggage inspection device, means for pre-processing the video information into a format suitable for analysis, means for inputting the pre-processed video information into a generating AI algorithm to detect dangerous or prohibited items, and means for presenting the results in real time using a visual display worn by the user and analyzing the user's reaction. This enables advanced detection of dangerous items that takes into account the emotional state of security staff and allows for faster countermeasures.
[0376] A "baggage inspection device" is a device that scans passenger and cargo luggage and makes visible the items contained inside.
[0377] "Visual information" refers to digital video data acquired by baggage inspection equipment, including X-ray images and visible light images.
[0378] A "generative AI algorithm" is a program that uses artificial intelligence technology to extract features from input data and identify dangerous or prohibited items.
[0379] "Object recognition techniques" are technologies that use algorithms to analyze and identify specific objects within images or videos.
[0380] A "visual display" is a device incorporated into devices such as smart glasses that visually presents information to the user.
[0381] "User emotional data" refers to data on the emotional state of the user, evaluated from their facial expressions and voice, acquired through visual displays and other sensors.
[0382] "Feedback" refers to information that is used to improve system operation and algorithms based on user feedback.
[0383] This system is designed to provide baggage inspection with integrated anomaly detection and user assistance. The server first collects video information acquired by the baggage inspection device. This information is comprehensive image data, including X-ray and visible light images in particular.
[0384] The server first preprocesses the received video information. This processing includes noise reduction and image quality improvement. The goal is to prepare the data so that the generative AI algorithm can perform analysis with high accuracy. Next, the server sends the preprocessed video information to the generative AI algorithm for analysis to identify dangerous or prohibited items. This generative AI algorithm utilizes object recognition techniques and has the ability to detect targets efficiently and effectively.
[0385] The analysis results are presented in real time to the user wearing a visual display. This visual display is typically implemented as a device like smart glasses and is responsible for visualizing the information. The system also analyzes the user's facial expressions and voice tone to collect emotional data. This emotional data reflects the user's level of alertness and anxiety and is incorporated into the system as feedback.
[0386] As a concrete example, when airport security staff use glasses-type visual displays to inspect passengers' baggage, any abnormalities are immediately detected by the images displayed on the screen. If the staff member feels uneasy, the system triggers additional inspection procedures based on emotional data. This entire process ensures safer and more efficient baggage inspection.
[0387] An example of a prompt to the generating AI model is as follows: "Analyze the X-ray images of the luggage to detect dangerous items. Also, analyze the user's facial expression data to assess their emotions and evaluate their importance." This prompt allows the system to perform a more precise analysis using the AI algorithm.
[0388] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0389] Step 1:
[0390] The server acquires video information from the baggage inspection device. This information includes X-ray images and visible light images. The acquired video information is received as input data and passed to the next preprocessing step.
[0391] Step 2:
[0392] The server performs preprocessing on the acquired video information. This preprocessing includes noise reduction and image quality correction to improve image quality. Finally, the data is converted into a format that is easy for the generating AI algorithm to analyze. The output of this process is the preprocessed video information.
[0393] Step 3:
[0394] The server inputs pre-processed video information into a generating AI algorithm. The generating AI algorithm analyzes the data using object recognition techniques to detect dangerous or prohibited items. The results of the analysis are output as a list of detected items.
[0395] Step 4:
[0396] The server assesses the risk level of the detected items based on the list and prepares to notify the user. This assessment sorts the items in descending order of risk based on their characteristics. The assessment results are output as data for notification.
[0397] Step 5:
[0398] On the device, evaluation results are presented to the user via a visual display. The visual display shows information in real time on smart glasses. The user reviews the presented information and provides feedback in response. This feedback is recorded as input data to the system.
[0399] Step 6:
[0400] To collect user emotion data, the device analyzes the user's voice and facial expressions. The analysis uses an emotion analysis engine to evaluate the emotional state. The output of this analysis is the user's emotion parameters.
[0401] Step 7:
[0402] The server retrains the generating AI algorithm based on user feedback and emotion parameters. This retraining is aimed at improving analysis accuracy and efficiency, optimizing the performance of future baggage screening.
[0403] Step 8:
[0404] The server automatically triggers additional testing procedures based on the user's emotional state. Signs of anxiety or alertness identified through emotion analysis are treated as high risk, and immediate additional action is taken. The output of this step is an instruction to execute the automated additional procedure.
[0405] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[0406] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0407] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart glasses 214.
[0408] [Third Embodiment]
[0409] Figure 5 shows an example of the configuration of the data processing system 310 according to the third embodiment.
[0410] As shown in Figure 5, the data processing system 310 includes a data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.
[0411] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0412] The headset terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a display 343. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and display 343 are also connected to the bus 52.
[0413] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[0414] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).
[0415] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[0416] Figure 6 shows an example of the main functions of the data processing device 12 and the headset terminal 314. As shown in Figure 6, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[0417] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0418] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0419] In the headset terminal 314, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0420] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the headset terminal 314 will be referred to as the "terminal".
[0421] This invention relates to a system for automating baggage screening at airports and transportation hubs. This system primarily consists of a server, terminals, and users. The program processing and specific examples are described below.
[0422] First, the server collects video data from baggage inspection devices in real time. This video data is based on X-ray and 3D images and provides a detailed representation of the inside of the baggage. This data is first pre-processed and converted into a format suitable for analysis. Pre-processing includes noise reduction, image normalization, and resolution adjustment.
[0423] Next, the pre-processed video data is input into a generating AI model. This AI model is pre-trained using object recognition technology and analyzes objects in the video to detect hazardous materials and prohibited items. The server receives the analysis results from the generating AI and makes a determination based on whether the detected items are dangerous. This determination is then generated as an alarm according to the level of risk.
[0424] The server notifies the terminal of the detection results, making them available for viewing by security staff users. The results are presented in text and visual form, showing the type and location of the hazardous material. Users can use this information to perform manual verification as needed.
[0425] Furthermore, the results manually verified by the user are sent to the server as feedback. This feedback information is used to retrain the generating AI model, contributing to improved accuracy in subsequent analyses.
[0426] As a concrete example, consider a scene from a baggage inspection at an airport. When a traveler's luggage passes through the inspection machine, the server immediately acquires video data, which is then analyzed by a generative AI that detects the presence of a knife. The server sends an alarm to the terminal indicating the presence of a knife, and security staff ask the person being inspected to check their luggage and actually remove the knife for verification. This verification result is sent to the server as feedback, and the generative AI learns to perform more accurate analyses.
[0427] This system enables efficient and accurate baggage screening, optimizing human resources and improving passenger security. Furthermore, the system is highly scalable, making it easy to implement at different airports and transportation hubs.
[0428] The following describes the processing flow.
[0429] Step 1:
[0430] The server acquires video data from baggage inspection devices in real time. This data visualizes the inside of the baggage using X-ray sensors and other imaging technologies.
[0431] Step 2:
[0432] The server preprocesses the acquired video data. This preprocessing involves noise reduction, image normalization, and resolution adjustment, and converting the data into the optimal format for analysis.
[0433] Step 3:
[0434] The server inputs pre-processed data into a generative AI model. The generative AI model is pre-trained and uses object recognition technology to analyze objects in the video and detect hazardous materials and prohibited items.
[0435] Step 4:
[0436] The server determines the level of danger of the detected items based on the analysis results generated by the AI. Based on the results, it generates an alarm corresponding to the degree of risk.
[0437] Step 5:
[0438] The server notifies the terminal of the judgment result. The notification is provided as text and visual information so that security staff can review the content.
[0439] Step 6:
[0440] Users (security staff) can check the detection results on their terminals and, if necessary, manually inspect baggage. Actions taken after the inspection will be determined according to the situation at hand.
[0441] Step 7:
[0442] The results of manual verification performed by the user are fed back to the server via the terminal. This feedback is used as data to improve the accuracy of the generated AI model.
[0443] Step 8:
[0444] The server utilizes the collected feedback to retrain the generated AI model. This improves the model's analytical accuracy and efficiency, strengthening the reliability of future inspections.
[0445] (Example 1)
[0446] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0447] Traditional baggage inspections required significant human resources to detect dangerous and prohibited items, resulting in limitations in efficiency and accuracy. Furthermore, continuous system improvements to enhance the accuracy of dangerous item detection were difficult.
[0448] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0449] In this invention, the server includes means for collecting image information obtained from a video acquisition device, means for preprocessing the image information into a format suitable for analysis, and means for inputting the preprocessed image information into a generating AI model to identify items that pose a safety problem. This makes it possible to detect hazardous materials efficiently and accurately.
[0450] An "image acquisition device" is a device installed to acquire image information of luggage, and mainly refers to equipment that visualizes the internal structure using X-rays or 3D images.
[0451] "Image information" refers to digital data collected by video acquisition devices to represent the inside of luggage, and this is the subject of analysis.
[0452] "Preprocessing" refers to the process of converting acquired image information into a format suitable for analysis, such as by removing noise or adjusting the resolution.
[0453] A "generative AI model" refers to an artificial intelligence model that is trained using machine learning techniques and has the ability to identify items with safety issues from image information.
[0454] "Items with safety issues" refer to items that are prohibited from being brought in or items that pose a risk, and which require action after detection.
[0455] "Feedback" refers to the flow of information sent back to the server by operators for manual verification and to improve detection accuracy, which is then used to retrain the generated AI model.
[0456] This invention is a system for streamlining and improving the accuracy of baggage screening at airports and transportation hubs. It mainly consists of a server, terminals, and users.
[0457] The server first collects image information obtained from the video acquisition device. This image information is used to visualize the inside of the luggage using X-rays and 3D images. The server then preprocesses the image information, applying noise reduction and resolution adjustments to convert it into a format suitable for analysis. This preprocessing prepares the generated AI model for effective analysis of the image information.
[0458] The pre-processed image information is input by the server into a generating AI model. This AI model is trained using machine learning techniques and has the ability to identify items that pose a safety risk. For example, it can detect prohibited items such as knives and explosives. Based on these detections, the server evaluates the level of safety and notifies the terminal of the evaluation result.
[0459] Users of the terminal review the notified evaluation results and manually check their luggage if necessary. The user's verification results are sent to the server as feedback. This feedback information is used to retrain the generating AI model and improve the accuracy of the analysis.
[0460] As a concrete example, during baggage screening at an airport, when a passenger's luggage passes through the screening device, the server immediately acquires image information and analyzes it using a generative AI model. If the analysis determines that the luggage may contain dangerous materials, that information is sent to a terminal. Security staff receive the warning from the terminal and take steps to manually check the luggage. Furthermore, as an example of a prompt message to improve detection accuracy, instructions such as "Identify suspicious items in the luggage" can be provided to the generative AI model.
[0461] This invention improves the accuracy and efficiency of baggage inspection, enhancing security while optimizing human resources.
[0462] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0463] Step 1:
[0464] The server acquires image information about the luggage from the image acquisition device. The input at this time is raw data of X-rays and 3D images. This data is used to visualize the internal structure of the luggage. The server receives this image data and prepares it for preprocessing in the next processing step.
[0465] Step 2:
[0466] The server preprocesses the acquired image information. The input is the raw image data received in the previous step. The server applies a noise reduction filter and adjusts the resolution to a predetermined setting to convert the image into a format suitable for analysis. This preprocessing removes unnecessary information from the image, making it suitable for recognition by the generated AI model. The output is clean, preprocessed image data.
[0467] Step 3:
[0468] The server inputs pre-processed image data into the generative AI model. The input is pre-processed image data. The server passes this data to the generative AI model for machine learning analysis. The generative AI model uses its pre-trained object recognition capabilities to identify items that pose a safety risk, such as knives or explosives. The output is the analysis results, including the detected items and their location information.
[0469] Step 4:
[0470] The server evaluates the level of safety based on the analysis results from the generated AI model and notifies the terminal of the evaluation result. The input is the analysis result obtained in the previous step. The server analyzes the analysis result and determines the level of risk. Based on the determination result, it sets an alarm level and sends that information to the terminal. The output is alarm information that reflects the risk evaluation.
[0471] Step 5:
[0472] The terminal displays alarm information to the security staff, who are the users. The input is alarm information sent from the server. The terminal displays this information to the user as text and visuals, clearly indicating the type and location of the detected item. The user then performs further manual verification based on this information.
[0473] Step 6:
[0474] The user performs a manual check and sends the results to the server as feedback. The input is the details of the check performed by the user. The user checks their luggage according to the terminal information and checks for the presence of actual dangerous goods. The check results are recorded as feedback and sent to the server. The output is the feedback information resulting from the check.
[0475] Step 7:
[0476] The server receives feedback information and uses it to retrain the generative AI model. The input is the feedback information. Based on this feedback, the server retrains the model to improve the accuracy of the next analysis. The output is the improved generative AI model.
[0477] (Application Example 1)
[0478] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0479] Baggage screening is essential for ensuring the safety of public and commercial facilities, but current methods are time-consuming and consume a large amount of human resources. Furthermore, the accuracy of detecting dangerous goods varies depending on the situation, and there is a need for efficient and highly accurate automation. In addition, the insufficient real-time situation assessment and rapid response also pose a safety challenge.
[0480] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0481] In this invention, the server includes means for collecting video data acquired from a baggage inspection device, means for preprocessing the video data into a format suitable for analysis, means for inputting the preprocessed video data into a generating AI model to detect dangerous or prohibited items, means for photographing items at the entrance of a public facility using a camera on an individual device and transmitting the data to a cloud server, means for performing data analysis on the cloud server and determining the danger of the items, and means for presenting the analysis results to a user device in real time and issuing an alarm. This makes it possible to efficiently and accurately automate the verification of baggage safety.
[0482] A "baggage inspection device" is a device used to check the internal condition of baggage and has the function of acquiring video data.
[0483] "Video data" refers to data including X-ray images and 3D images of baggage acquired by baggage inspection equipment.
[0484] A "generative AI model" is a model that is pre-trained using artificial intelligence technology to analyze objects from video data and detect hazardous materials or prohibited items.
[0485] A "cloud server" is an external computing resource that can be accessed via a network and is used for data analysis and information storage.
[0486] A "user device" is a terminal device that allows users to view analysis results and provide feedback as needed, and typically refers to a personal digital assistant (PDI).
[0487] "Preprocessing" refers to a series of processes performed to convert video data into a format suitable for analysis.
[0488] "Feedback" refers to the confirmation results provided by users, and this information is used to retrain the generated AI model.
[0489] "Dangerous goods" are items that may threaten public safety and are subject to detection.
[0490] "Notification" refers to the act of informing the user of detected information, and is the result provided as an alarm.
[0491] This invention is a system designed to streamline baggage inspection. The server collects video data from baggage inspection devices in real time and preprocesses it into a format suitable for analysis. Preprocessing includes noise reduction, image normalization, and resolution adjustment.
[0492] The pre-processed data is input into a generating AI model on a cloud server. This AI model is trained using object recognition technology and has the ability to analyze and detect hazardous materials and prohibited items from video data. The analysis results determine the level of danger of the items and are sent from the server in a format that notifies the user's device. The notification is displayed in real time on the user's device and can be accessed via a portable terminal.
[0493] Based on this notification, the user reviews the baggage inspection results and performs additional manual inspections if necessary. After the review is complete, the results are sent back to the server as feedback. This feedback data is used to continuously train the generative AI model, enabling improvements in analysis accuracy.
[0494] A concrete example is a baggage inspection scene at the entrance of a public facility. When a customer enters the facility, they open a dedicated smartphone app and take a picture of their baggage with the camera. This application connects to a cloud server and detects any dangerous items hidden in the baggage. An example of a prompt message used would be, "Please analyze whether this baggage contains any dangerous items." This enables quick and appropriate safety checks.
[0495] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0496] Step 1:
[0497] The server collects video data from the baggage inspection device. The input consists of X-ray images and 3D images from the baggage inspection device. The server receives these images and temporarily stores them in preparation for the next processing step. The output is image data for the pre-processing stage.
[0498] Step 2:
[0499] The server preprocesses the received video data. It performs noise reduction, image normalization, and resolution adjustment, and converts it into a format suitable for analysis. The input is the original image data, which is the output of step 1, and the output is preprocessed data suitable for the generative AI model.
[0500] Step 3:
[0501] The server inputs pre-processed video data into a generating AI model. The generating AI model analyzes the data and detects hazardous materials and prohibited items. The input is the pre-processed data which is the output of step 2, and the output is a list of detected items and their attribute data.
[0502] Step 4:
[0503] The server determines the level of risk of an item based on the analysis results from the generated AI model. The determination is expressed as a score or alarm corresponding to the degree of risk. The input is the detection result from step 3, and the output is alarm data based on the level of risk.
[0504] Step 5:
[0505] The terminal receives alarm data from the server and notifies the user in real time. Notifications are made in the form of voice, pop-up messages, warning lights, etc. The input is the alarm data which is the output of step 4, and the output on the terminal is the alarm display for the user.
[0506] Step 6:
[0507] The user reviews the notification and performs manual inspections if necessary. The results of the review are sent to the server as feedback. The input is the output from step 5, and the output is text data or image data of the reviewed content.
[0508] Step 7:
[0509] The server receives feedback from the user and uses it to retrain the generative AI model. The feedback data is added to the training dataset of the generative AI model, improving the accuracy of the analysis. The input is the feedback data, which is the output of step 6, and the output is the updated parameters of the AI model.
[0510] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0511] This invention relates to an automated baggage inspection system that incorporates an emotion engine. This system consists of collaborative work between a baggage inspection device, a server, a terminal, and a user. The program processing and specific examples are described below.
[0512] The server first acquires video data from the baggage inspection device. This data includes X-ray and visible light images, showing the contents of the baggage in detail. The server then preprocesses this data into a format that is easy to analyze. Specifically, this involves processes such as noise reduction and image quality improvement.
[0513] Next, the server inputs the pre-processed video data into a generating AI model. This model analyzes and detects hazardous materials and prohibited items based on object recognition technology. Based on the results, the server evaluates the level of danger and generates a result. The generated result is then sent to the terminal.
[0514] This system incorporates an emotion engine that monitors the emotions of the security staff users in real time. It analyzes the user's voice and facial expressions to understand their emotions based on the results displayed on the terminal, and then weights the feedback accordingly. For example, if a user expresses concern, that feedback becomes more important in retraining the generative AI model.
[0515] Furthermore, if the emotion engine indicates that the user's emotions are anxious or alarming, the server automatically triggers additional testing procedures. This process enhances the response to potential risks.
[0516] As a concrete example, when a passenger's baggage passes through the inspection device, the server acquires the video data and begins analysis. As the detected results are displayed on the terminal, the emotion engine analyzes the user's reaction. For example, if the user shows tension or anxiety, the server issues an additional alert and prompts further manual verification. This entire process ensures that inspections are precise and flexible.
[0517] This system configuration further enhances the security of airports and transportation systems, streamlines security processes, and offers flexibility to adapt to diverse environments.
[0518] The following describes the processing flow.
[0519] Step 1:
[0520] The server acquires video data from baggage inspection devices in real time. This data is acquired using X-ray technology to visualize the inside of the baggage.
[0521] Step 2:
[0522] The server preprocesses the acquired video data and converts it into a format suitable for analysis. Specifically, it removes noise and adjusts the image contrast and resolution.
[0523] Step 3:
[0524] The server inputs pre-processed data into a generating AI model. This model uses object recognition technology to analyze objects in the video and detect hazardous materials and prohibited items.
[0525] Step 4:
[0526] The server determines the level of risk of an item based on the analysis results of the generated AI model and sends the result to the terminal. This information includes the type and location of the item, as well as the risk level.
[0527] Step 5:
[0528] The terminal displays the received inspection results to the user, who is a security staff member. Here, the user can visually confirm the results.
[0529] Step 6:
[0530] The emotion engine analyzes the user's voice and facial expressions in real time to determine their emotional state. Emotional data is recorded for feedback.
[0531] Step 7:
[0532] The user reviews the results and provides feedback based on their reaction. This feedback, along with the analysis results from the emotion engine, is sent to the server.
[0533] Step 8:
[0534] The server receives user feedback and uses it as retraining data for the generated AI model. This process aims to improve the model's accuracy.
[0535] Step 9:
[0536] The server automatically triggers additional checks if the emotion engine detects any emotions such as anxiety or tension in the user. This strengthens the system's ability to address potential security risks.
[0537] (Example 2)
[0538] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0539] In baggage inspections, there is a need for reliable detection of dangerous and prohibited items, as well as improvements in inspection speed and efficiency. However, current systems have limitations in improving detection accuracy, and in particular, they cannot consider the influence of user emotions on inspection results, resulting in insufficient risk management.
[0540] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0541] In this invention, the server includes means for collecting video data acquired from a baggage inspection device, means for preprocessing the video data into a format suitable for analysis, means for inputting the preprocessed video data into a generating AI model to detect dangerous goods or prohibited items, and means for analyzing the user's voice and facial expressions using an emotion analysis device to evaluate emotions. This improves the accuracy of dangerous goods detection and enables flexible adjustment of inspection procedures that take into account the user's emotions.
[0542] A "baggage inspection device" is a technological device used at airports, train stations, and other locations to examine passengers' baggage and check for the presence of dangerous goods or prohibited items.
[0543] "Visual data" refers to digital information generated by X-rays or other methods to visually show the detailed condition of items in baggage.
[0544] A "generative AI model" is a form of artificial intelligence trained using a deep learning framework, which utilizes object recognition technology to identify and classify items within luggage.
[0545] An "emotion analysis device" is a system component that detects and analyzes a user's voice tone, facial expressions, etc., to evaluate their emotions.
[0546] "Risk assessment" is the process of evaluating the hazardous nature of detected items and determining subsequent actions based on that assessment.
[0547] An embodiment of this invention consists of a baggage inspection device, a server, a terminal, and collaborative work by a user.
[0548] The server first acquires X-ray and visible light image data from the baggage inspection device. This data is used to visually show the contents of the baggage in detail and to check for the presence of dangerous or prohibited items. The server then performs preprocessing on this data, such as noise reduction and image quality enhancement, to prepare it for analysis. Image processing libraries are among the software used at this stage.
[0549] Next, the server inputs the pre-processed video data into a generating AI model. This model is built on deep learning frameworks such as TensorFlow and PyTorch, and uses object recognition technology to identify dangerous or prohibited items in the luggage. It then evaluates the level of danger based on the identified items and generates an evaluation result. The generated result is transmitted to the terminal in real time and presented to the user.
[0550] The terminal is also equipped with an emotion analyzer that analyzes the user's voice and facial expressions in real time to understand their emotional state. For example, if the user shows anxiety or apprehension, that information is fed back to the server and used to adjust the examination procedure. This enables a more flexible and precise response.
[0551] As a concrete example, when a passenger's baggage passes through an inspection device, the server acquires the video data and initiates a process to detect dangerous items. The detection results are displayed on the terminal, and simultaneously, sentiment analysis is performed to analyze the user's reaction. If the user shows signs of anxiety, the server recommends further verification steps, enabling dynamic risk management.
[0552] Examples of prompt messages include phrases like, "Identify dangerous items in passengers' baggage and adjust the inspection procedure based on the user's emotional response." In this way, security processes at airports and transportation facilities are conducted efficiently and safely.
[0553] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0554] Step 1:
[0555] The server acquires X-ray and visible light video data from the baggage inspection device. This data provides detailed information about the items inside the baggage and is transmitted to the server as a multidimensional array. The input video data is denoised, processed to improve image quality, and converted into a format suitable for analysis. The output is pre-processed, clear video data.
[0556] Step 2:
[0557] The server inputs pre-processed video data into a generating AI model. The AI model, built using TensorFlow or PyTorch, utilizes object recognition technology to analyze dangerous or prohibited items in luggage. As a data computation, the model identifies patterns through numerous neuron layers. The output is the identification result regarding dangerous items and its detection accuracy.
[0558] Step 3:
[0559] The server evaluates the risk level of the target based on the identification results obtained from the AI model and calculates the risk level. The risk assessment is calculated according to pre-set criteria based on the content of the identification results. A detailed report summarizing this assessment result is generated and sent to the terminal. The output is a report including the risk level.
[0560] Step 4:
[0561] The terminal receives a report sent from the server and displays it to the user. At this time, an emotion analysis device is activated, capturing the user's voice tone and facial expressions in real time. Based on the captured emotion data, the user's response is analyzed, and their emotional state is estimated. The output is the current emotion evaluation result.
[0562] Step 5:
[0563] The server receives the sentiment assessment results sent as feedback from the terminal and dynamically adjusts the testing procedure. Specifically, if the user's emotions indicate anxiety or caution, it triggers additional testing or a more detailed review. This adjustment enables flexible and effective risk management. The output is the adjusted testing procedure.
[0564] (Application Example 2)
[0565] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0566] In modern baggage screening systems, the detection of dangerous or prohibited items still largely relies on human judgment, and the accuracy and efficiency of detection have not improved sufficiently. Furthermore, the emotions and cognitive state of security staff conducting the inspections can influence their ability to respond to potential risks. However, there is a lack of means to effectively incorporate these factors and improve accuracy through real-time analysis and feedback.
[0567] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[0568] In this invention, the server includes means for collecting video information acquired from a baggage inspection device, means for pre-processing the video information into a format suitable for analysis, means for inputting the pre-processed video information into a generating AI algorithm to detect dangerous or prohibited items, and means for presenting the results in real time using a visual display worn by the user and analyzing the user's reaction. This enables advanced detection of dangerous items that takes into account the emotional state of security staff and allows for faster countermeasures.
[0569] A "baggage inspection device" is a device that scans passenger and cargo luggage and makes visible the items contained inside.
[0570] "Visual information" refers to digital video data acquired by baggage inspection equipment, including X-ray images and visible light images.
[0571] A "generative AI algorithm" is a program that uses artificial intelligence technology to extract features from input data and identify dangerous or prohibited items.
[0572] "Object recognition techniques" are technologies that use algorithms to analyze and identify specific objects within images or videos.
[0573] A "visual display" is a device incorporated into devices such as smart glasses that visually presents information to the user.
[0574] "User emotional data" refers to data on the emotional state of the user, evaluated from their facial expressions and voice, acquired through visual displays and other sensors.
[0575] "Feedback" refers to information that is used to improve system operation and algorithms based on user feedback.
[0576] This system is designed to provide baggage inspection with integrated anomaly detection and user assistance. The server first collects video information acquired by the baggage inspection device. This information is comprehensive image data, including X-ray and visible light images in particular.
[0577] The server first preprocesses the received video information. This processing includes noise reduction and image quality improvement. The goal is to prepare the data so that the generative AI algorithm can perform analysis with high accuracy. Next, the server sends the preprocessed video information to the generative AI algorithm for analysis to identify dangerous or prohibited items. This generative AI algorithm utilizes object recognition techniques and has the ability to detect targets efficiently and effectively.
[0578] The analysis results are presented in real time to the user wearing a visual display. This visual display is typically implemented as a device like smart glasses and is responsible for visualizing the information. The system also analyzes the user's facial expressions and voice tone to collect emotional data. This emotional data reflects the user's level of alertness and anxiety and is incorporated into the system as feedback.
[0579] As a concrete example, when airport security staff use glasses-type visual displays to inspect passengers' baggage, any abnormalities are immediately detected by the images displayed on the screen. If the staff member feels uneasy, the system triggers additional inspection procedures based on emotional data. This entire process ensures safer and more efficient baggage inspection.
[0580] An example of a prompt to the generating AI model is as follows: "Analyze the X-ray images of the luggage to detect dangerous items. Also, analyze the user's facial expression data to assess their emotions and evaluate their importance." This prompt allows the system to perform a more precise analysis using the AI algorithm.
[0581] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0582] Step 1:
[0583] The server acquires video information from the baggage inspection device. This information includes X-ray images and visible light images. The acquired video information is received as input data and passed to the next preprocessing step.
[0584] Step 2:
[0585] The server performs preprocessing on the acquired video information. This preprocessing includes noise reduction and image quality correction to improve image quality. Finally, the data is converted into a format that is easy for the generating AI algorithm to analyze. The output of this process is the preprocessed video information.
[0586] Step 3:
[0587] The server inputs pre-processed video information into a generating AI algorithm. The generating AI algorithm analyzes the data using object recognition techniques to detect dangerous or prohibited items. The results of the analysis are output as a list of detected items.
[0588] Step 4:
[0589] The server assesses the risk level of the detected items based on the list and prepares to notify the user. This assessment sorts the items in descending order of risk based on their characteristics. The assessment results are output as data for notification.
[0590] Step 5:
[0591] On the device, evaluation results are presented to the user via a visual display. The visual display shows information in real time on smart glasses. The user reviews the presented information and provides feedback in response. This feedback is recorded as input data to the system.
[0592] Step 6:
[0593] To collect user emotion data, the device analyzes the user's voice and facial expressions. The analysis uses an emotion analysis engine to evaluate the emotional state. The output of this analysis is the user's emotion parameters.
[0594] Step 7:
[0595] The server retrains the generating AI algorithm based on user feedback and emotion parameters. This retraining is aimed at improving analysis accuracy and efficiency, optimizing the performance of future baggage screening.
[0596] Step 8:
[0597] The server automatically triggers additional testing procedures based on the user's emotional state. Signs of anxiety or alertness identified through emotion analysis are treated as high risk, and immediate additional action is taken. The output of this step is an instruction to execute the automated additional procedure.
[0598] The specific processing unit 290 transmits the result of the specific processing to the headset terminal 314. In the headset terminal 314, the control unit 46A causes the speaker 240 and display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[0599] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0600] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and specific processing may also be performed by the headset terminal 314.
[0601] [Fourth Embodiment]
[0602] Figure 7 shows an example of the configuration of the data processing system 410 according to the fourth embodiment.
[0603] As shown in Figure 7, the data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.
[0604] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0605] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a controlled object 443. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and controlled object 443 are also connected to the bus 52.
[0606] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[0607] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).
[0608] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[0609] The controlled object 443 includes a display device, LEDs in the eyes, and motors that drive the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the robot 414's emotions can be expressed by controlling these motors. Furthermore, the robot 414's facial expressions can also be expressed by controlling the illumination state of the LEDs in its eyes.
[0610] Figure 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Figure 8, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[0611] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0612] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0613] In robot 414, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0614] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0615] This invention relates to a system for automating baggage screening at airports and transportation hubs. This system primarily consists of a server, terminals, and users. The program processing and specific examples are described below.
[0616] First, the server collects video data from baggage inspection devices in real time. This video data is based on X-ray and 3D images and provides a detailed representation of the inside of the baggage. This data is first pre-processed and converted into a format suitable for analysis. Pre-processing includes noise reduction, image normalization, and resolution adjustment.
[0617] Next, the pre-processed video data is input into a generating AI model. This AI model is pre-trained using object recognition technology and analyzes objects in the video to detect hazardous materials and prohibited items. The server receives the analysis results from the generating AI and makes a determination based on whether the detected items are dangerous. This determination is then generated as an alarm according to the level of risk.
[0618] The server notifies the terminal of the detection results, making them available for viewing by security staff users. The results are presented in text and visual form, showing the type and location of the hazardous material. Users can use this information to perform manual verification as needed.
[0619] Furthermore, the results manually verified by the user are sent to the server as feedback. This feedback information is used to retrain the generating AI model, contributing to improved accuracy in subsequent analyses.
[0620] As a concrete example, consider a scene from a baggage inspection at an airport. When a traveler's luggage passes through the inspection machine, the server immediately acquires video data, which is then analyzed by a generative AI that detects the presence of a knife. The server sends an alarm to the terminal indicating the presence of a knife, and security staff ask the person being inspected to check their luggage and actually remove the knife for verification. This verification result is sent to the server as feedback, and the generative AI learns to perform more accurate analyses.
[0621] This system enables efficient and accurate baggage screening, optimizing human resources and improving passenger security. Furthermore, the system is highly scalable, making it easy to implement at different airports and transportation hubs.
[0622] The following describes the processing flow.
[0623] Step 1:
[0624] The server acquires video data from baggage inspection devices in real time. This data visualizes the inside of the baggage using X-ray sensors and other imaging technologies.
[0625] Step 2:
[0626] The server preprocesses the acquired video data. This preprocessing involves noise reduction, image normalization, and resolution adjustment, and converting the data into the optimal format for analysis.
[0627] Step 3:
[0628] The server inputs pre-processed data into a generative AI model. The generative AI model is pre-trained and uses object recognition technology to analyze objects in the video and detect hazardous materials and prohibited items.
[0629] Step 4:
[0630] The server determines the level of danger of the detected items based on the analysis results generated by the AI. Based on the results, it generates an alarm corresponding to the degree of risk.
[0631] Step 5:
[0632] The server notifies the terminal of the judgment result. The notification is provided as text and visual information so that security staff can review the content.
[0633] Step 6:
[0634] Users (security staff) can check the detection results on their terminals and, if necessary, manually inspect baggage. Actions taken after the inspection will be determined according to the situation at hand.
[0635] Step 7:
[0636] The results of manual verification performed by the user are fed back to the server via the terminal. This feedback is used as data to improve the accuracy of the generated AI model.
[0637] Step 8:
[0638] The server utilizes the collected feedback to retrain the generated AI model. This improves the model's analytical accuracy and efficiency, strengthening the reliability of future inspections.
[0639] (Example 1)
[0640] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0641] Traditional baggage inspections required significant human resources to detect dangerous and prohibited items, resulting in limitations in efficiency and accuracy. Furthermore, continuous system improvements to enhance the accuracy of dangerous item detection were difficult.
[0642] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0643] In this invention, the server includes means for collecting image information obtained from a video acquisition device, means for preprocessing the image information into a format suitable for analysis, and means for inputting the preprocessed image information into a generating AI model to identify items that pose a safety problem. This makes it possible to detect hazardous materials efficiently and accurately.
[0644] An "image acquisition device" is a device installed to acquire image information of luggage, and mainly refers to equipment that visualizes the internal structure using X-rays or 3D images.
[0645] "Image information" refers to digital data collected by video acquisition devices to represent the inside of luggage, and this is the subject of analysis.
[0646] "Preprocessing" refers to the process of converting acquired image information into a format suitable for analysis, such as by removing noise or adjusting the resolution.
[0647] A "generative AI model" refers to an artificial intelligence model that is trained using machine learning techniques and has the ability to identify items with safety issues from image information.
[0648] "Items with safety issues" refer to items that are prohibited from being brought in or items that pose a risk, and which require action after detection.
[0649] "Feedback" refers to the flow of information sent back to the server by operators for manual verification and to improve detection accuracy, which is then used to retrain the generated AI model.
[0650] This invention is a system for streamlining and improving the accuracy of baggage screening at airports and transportation hubs. It mainly consists of a server, terminals, and users.
[0651] The server first collects image information obtained from the video acquisition device. This image information is used to visualize the inside of the luggage using X-rays and 3D images. The server then preprocesses the image information, applying noise reduction and resolution adjustments to convert it into a format suitable for analysis. This preprocessing prepares the generated AI model for effective analysis of the image information.
[0652] The pre-processed image information is input by the server into a generating AI model. This AI model is trained using machine learning techniques and has the ability to identify items that pose a safety risk. For example, it can detect prohibited items such as knives and explosives. Based on these detections, the server evaluates the level of safety and notifies the terminal of the evaluation result.
[0653] Users of the terminal review the notified evaluation results and manually check their luggage if necessary. The user's verification results are sent to the server as feedback. This feedback information is used to retrain the generating AI model and improve the accuracy of the analysis.
[0654] As a concrete example, during baggage screening at an airport, when a passenger's luggage passes through the screening device, the server immediately acquires image information and analyzes it using a generative AI model. If the analysis determines that the luggage may contain dangerous materials, that information is sent to a terminal. Security staff receive the warning from the terminal and take steps to manually check the luggage. Furthermore, as an example of a prompt message to improve detection accuracy, instructions such as "Identify suspicious items in the luggage" can be provided to the generative AI model.
[0655] This invention improves the accuracy and efficiency of baggage inspection, enhancing security while optimizing human resources.
[0656] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0657] Step 1:
[0658] The server acquires image information about the luggage from the image acquisition device. The input at this time is raw data of X-rays and 3D images. This data is used to visualize the internal structure of the luggage. The server receives this image data and prepares it for preprocessing in the next processing step.
[0659] Step 2:
[0660] The server preprocesses the acquired image information. The input is the raw image data received in the previous step. The server applies a noise reduction filter and adjusts the resolution to a predetermined setting to convert the image into a format suitable for analysis. This preprocessing removes unnecessary information from the image, making it suitable for recognition by the generated AI model. The output is clean, preprocessed image data.
[0661] Step 3:
[0662] The server inputs pre-processed image data into the generative AI model. The input is pre-processed image data. The server passes this data to the generative AI model for machine learning analysis. The generative AI model uses its pre-trained object recognition capabilities to identify items that pose a safety risk, such as knives or explosives. The output is the analysis results, including the detected items and their location information.
[0663] Step 4:
[0664] The server evaluates the level of safety based on the analysis results from the generated AI model and notifies the terminal of the evaluation result. The input is the analysis result obtained in the previous step. The server analyzes the analysis result and determines the level of risk. Based on the determination result, it sets an alarm level and sends that information to the terminal. The output is alarm information that reflects the risk evaluation.
[0665] Step 5:
[0666] The terminal displays alarm information to the security staff, who are the users. The input is alarm information sent from the server. The terminal displays this information to the user as text and visuals, clearly indicating the type and location of the detected item. The user then performs further manual verification based on this information.
[0667] Step 6:
[0668] The user performs a manual check and sends the results to the server as feedback. The input is the details of the check performed by the user. The user checks their luggage according to the terminal information and checks for the presence of actual dangerous goods. The check results are recorded as feedback and sent to the server. The output is the feedback information resulting from the check.
[0669] Step 7:
[0670] The server receives feedback information and uses it to retrain the generative AI model. The input is the feedback information. Based on this feedback, the server retrains the model to improve the accuracy of the next analysis. The output is the improved generative AI model.
[0671] (Application Example 1)
[0672] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0673] Baggage screening is essential for ensuring the safety of public and commercial facilities, but current methods are time-consuming and consume a large amount of human resources. Furthermore, the accuracy of detecting dangerous goods varies depending on the situation, and there is a need for efficient and highly accurate automation. In addition, the insufficient real-time situation assessment and rapid response also pose a safety challenge.
[0674] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0675] In this invention, the server includes means for collecting video data acquired from a baggage inspection device, means for preprocessing the video data into a format suitable for analysis, means for inputting the preprocessed video data into a generating AI model to detect dangerous or prohibited items, means for photographing items at the entrance of a public facility using a camera on an individual device and transmitting the data to a cloud server, means for performing data analysis on the cloud server and determining the danger of the items, and means for presenting the analysis results to a user device in real time and issuing an alarm. This makes it possible to efficiently and accurately automate the verification of baggage safety.
[0676] A "baggage inspection device" is a device used to check the internal condition of baggage and has the function of acquiring video data.
[0677] "Video data" refers to data including X-ray images and 3D images of baggage acquired by baggage inspection equipment.
[0678] A "generative AI model" is a model that is pre-trained using artificial intelligence technology to analyze objects from video data and detect hazardous materials or prohibited items.
[0679] A "cloud server" is an external computing resource that can be accessed via a network and is used for data analysis and information storage.
[0680] A "user device" is a terminal device that allows users to view analysis results and provide feedback as needed, and typically refers to a personal digital assistant (PDI).
[0681] "Preprocessing" refers to a series of processes performed to convert video data into a format suitable for analysis.
[0682] "Feedback" refers to the confirmation results provided by users, and this information is used to retrain the generated AI model.
[0683] "Dangerous goods" are items that may threaten public safety and are subject to detection.
[0684] "Notification" refers to the act of informing the user of detected information, and is the result provided as an alarm.
[0685] This invention is a system designed to streamline baggage inspection. The server collects video data from baggage inspection devices in real time and preprocesses it into a format suitable for analysis. Preprocessing includes noise reduction, image normalization, and resolution adjustment.
[0686] The pre-processed data is input into a generating AI model on a cloud server. This AI model is trained using object recognition technology and has the ability to analyze and detect hazardous materials and prohibited items from video data. The analysis results determine the level of danger of the items and are sent from the server in a format that notifies the user's device. The notification is displayed in real time on the user's device and can be accessed via a portable terminal.
[0687] Based on this notification, the user reviews the baggage inspection results and performs additional manual inspections if necessary. After the review is complete, the results are sent back to the server as feedback. This feedback data is used to continuously train the generative AI model, enabling improvements in analysis accuracy.
[0688] A concrete example is a baggage inspection scene at the entrance of a public facility. When a customer enters the facility, they open a dedicated smartphone app and take a picture of their baggage with the camera. This application connects to a cloud server and detects any dangerous items hidden in the baggage. An example of a prompt message used would be, "Please analyze whether this baggage contains any dangerous items." This enables quick and appropriate safety checks.
[0689] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0690] Step 1:
[0691] The server collects video data from the baggage inspection device. The input consists of X-ray images and 3D images from the baggage inspection device. The server receives these images and temporarily stores them in preparation for the next processing step. The output is image data for the pre-processing stage.
[0692] Step 2:
[0693] The server preprocesses the received video data. It performs noise reduction, image normalization, and resolution adjustment, and converts it into a format suitable for analysis. The input is the original image data, which is the output of step 1, and the output is preprocessed data suitable for the generative AI model.
[0694] Step 3:
[0695] The server inputs pre-processed video data into a generating AI model. The generating AI model analyzes the data and detects hazardous materials and prohibited items. The input is the pre-processed data which is the output of step 2, and the output is a list of detected items and their attribute data.
[0696] Step 4:
[0697] The server determines the level of risk of an item based on the analysis results from the generated AI model. The determination is expressed as a score or alarm corresponding to the degree of risk. The input is the detection result from step 3, and the output is alarm data based on the level of risk.
[0698] Step 5:
[0699] The terminal receives alarm data from the server and notifies the user in real time. Notifications are made in the form of voice, pop-up messages, warning lights, etc. The input is the alarm data which is the output of step 4, and the output on the terminal is the alarm display for the user.
[0700] Step 6:
[0701] The user reviews the notification and performs manual inspections if necessary. The results of the review are sent to the server as feedback. The input is the output from step 5, and the output is text data or image data of the reviewed content.
[0702] Step 7:
[0703] The server receives feedback from the user and uses it to retrain the generative AI model. The feedback data is added to the training dataset of the generative AI model, improving the accuracy of the analysis. The input is the feedback data, which is the output of step 6, and the output is the updated parameters of the AI model.
[0704] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0705] This invention relates to an automated baggage inspection system that incorporates an emotion engine. This system consists of collaborative work between a baggage inspection device, a server, a terminal, and a user. The program processing and specific examples are described below.
[0706] The server first acquires video data from the baggage inspection device. This data includes X-ray and visible light images, showing the contents of the baggage in detail. The server then preprocesses this data into a format that is easy to analyze. Specifically, this involves processes such as noise reduction and image quality improvement.
[0707] Next, the server inputs the pre-processed video data into a generating AI model. This model analyzes and detects hazardous materials and prohibited items based on object recognition technology. Based on the results, the server evaluates the level of danger and generates a result. The generated result is then sent to the terminal.
[0708] This system incorporates an emotion engine that monitors the emotions of the security staff users in real time. It analyzes the user's voice and facial expressions to understand their emotions based on the results displayed on the terminal, and then weights the feedback accordingly. For example, if a user expresses concern, that feedback becomes more important in retraining the generative AI model.
[0709] Furthermore, if the emotion engine indicates that the user's emotions are anxious or alarming, the server automatically triggers additional testing procedures. This process enhances the response to potential risks.
[0710] As a concrete example, when a passenger's baggage passes through the inspection device, the server acquires the video data and begins analysis. As the detected results are displayed on the terminal, the emotion engine analyzes the user's reaction. For example, if the user shows tension or anxiety, the server issues an additional alert and prompts further manual verification. This entire process ensures that inspections are precise and flexible.
[0711] This system configuration further enhances the security of airports and transportation systems, streamlines security processes, and offers flexibility to adapt to diverse environments.
[0712] The following describes the processing flow.
[0713] Step 1:
[0714] The server acquires video data from baggage inspection devices in real time. This data is acquired using X-ray technology to visualize the inside of the baggage.
[0715] Step 2:
[0716] The server preprocesses the acquired video data and converts it into a format suitable for analysis. Specifically, it removes noise and adjusts the image contrast and resolution.
[0717] Step 3:
[0718] The server inputs pre-processed data into a generating AI model. This model uses object recognition technology to analyze objects in the video and detect hazardous materials and prohibited items.
[0719] Step 4:
[0720] The server determines the level of risk of an item based on the analysis results of the generated AI model and sends the result to the terminal. This information includes the type and location of the item, as well as the risk level.
[0721] Step 5:
[0722] The terminal displays the received inspection results to the user, who is a security staff member. Here, the user can visually confirm the results.
[0723] Step 6:
[0724] The emotion engine analyzes the user's voice and facial expressions in real time to determine their emotional state. Emotional data is recorded for feedback.
[0725] Step 7:
[0726] The user reviews the results and provides feedback based on their reaction. This feedback, along with the analysis results from the emotion engine, is sent to the server.
[0727] Step 8:
[0728] The server receives user feedback and uses it as retraining data for the generated AI model. This process aims to improve the model's accuracy.
[0729] Step 9:
[0730] The server automatically triggers additional checks if the emotion engine detects any emotions such as anxiety or tension in the user. This strengthens the system's ability to address potential security risks.
[0731] (Example 2)
[0732] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0733] In baggage inspections, there is a need for reliable detection of dangerous and prohibited items, as well as improvements in inspection speed and efficiency. However, current systems have limitations in improving detection accuracy, and in particular, they cannot consider the influence of user emotions on inspection results, resulting in insufficient risk management.
[0734] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0735] In this invention, the server includes means for collecting video data acquired from a baggage inspection device, means for preprocessing the video data into a format suitable for analysis, means for inputting the preprocessed video data into a generating AI model to detect dangerous goods or prohibited items, and means for analyzing the user's voice and facial expressions using an emotion analysis device to evaluate emotions. This improves the accuracy of dangerous goods detection and enables flexible adjustment of inspection procedures that take into account the user's emotions.
[0736] A "baggage inspection device" is a technological device used at airports, train stations, and other locations to examine passengers' baggage and check for the presence of dangerous goods or prohibited items.
[0737] "Visual data" refers to digital information generated by X-rays or other methods to visually show the detailed condition of items in baggage.
[0738] A "generative AI model" is a form of artificial intelligence trained using a deep learning framework, which utilizes object recognition technology to identify and classify items within luggage.
[0739] An "emotion analysis device" is a system component that detects and analyzes a user's voice tone, facial expressions, etc., to evaluate their emotions.
[0740] "Risk assessment" is the process of evaluating the hazardous nature of detected items and determining subsequent actions based on that assessment.
[0741] An embodiment of this invention consists of a baggage inspection device, a server, a terminal, and collaborative work by a user.
[0742] The server first acquires X-ray and visible light image data from the baggage inspection device. This data is used to visually show the contents of the baggage in detail and to check for the presence of dangerous or prohibited items. The server then performs preprocessing on this data, such as noise reduction and image quality enhancement, to prepare it for analysis. Image processing libraries are among the software used at this stage.
[0743] Next, the server inputs the pre-processed video data into a generating AI model. This model is built on deep learning frameworks such as TensorFlow and PyTorch, and uses object recognition technology to identify dangerous or prohibited items in the luggage. It then evaluates the level of danger based on the identified items and generates an evaluation result. The generated result is transmitted to the terminal in real time and presented to the user.
[0744] The terminal is also equipped with an emotion analyzer that analyzes the user's voice and facial expressions in real time to understand their emotional state. For example, if the user shows anxiety or apprehension, that information is fed back to the server and used to adjust the examination procedure. This enables a more flexible and precise response.
[0745] As a concrete example, when a passenger's baggage passes through an inspection device, the server acquires the video data and initiates a process to detect dangerous items. The detection results are displayed on the terminal, and simultaneously, sentiment analysis is performed to analyze the user's reaction. If the user shows signs of anxiety, the server recommends further verification steps, enabling dynamic risk management.
[0746] Examples of prompt messages include phrases like, "Identify dangerous items in passengers' baggage and adjust the inspection procedure based on the user's emotional response." In this way, security processes at airports and transportation facilities are conducted efficiently and safely.
[0747] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0748] Step 1:
[0749] The server acquires X-ray and visible light video data from the baggage inspection device. This data provides detailed information about the items inside the baggage and is transmitted to the server as a multidimensional array. The input video data is denoised, processed to improve image quality, and converted into a format suitable for analysis. The output is pre-processed, clear video data.
[0750] Step 2:
[0751] The server inputs pre-processed video data into a generating AI model. The AI model, built using TensorFlow or PyTorch, utilizes object recognition technology to analyze dangerous or prohibited items in luggage. As a data computation, the model identifies patterns through numerous neuron layers. The output is the identification result regarding dangerous items and its detection accuracy.
[0752] Step 3:
[0753] The server evaluates the risk level of the target based on the identification results obtained from the AI model and calculates the risk level. The risk assessment is calculated according to pre-set criteria based on the content of the identification results. A detailed report summarizing this assessment result is generated and sent to the terminal. The output is a report including the risk level.
[0754] Step 4:
[0755] The terminal receives a report sent from the server and displays it to the user. At this time, an emotion analysis device is activated, capturing the user's voice tone and facial expressions in real time. Based on the captured emotion data, the user's response is analyzed, and their emotional state is estimated. The output is the current emotion evaluation result.
[0756] Step 5:
[0757] The server receives the sentiment assessment results sent as feedback from the terminal and dynamically adjusts the testing procedure. Specifically, if the user's emotions indicate anxiety or caution, it triggers additional testing or a more detailed review. This adjustment enables flexible and effective risk management. The output is the adjusted testing procedure.
[0758] (Application Example 2)
[0759] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0760] In modern baggage screening systems, the detection of dangerous or prohibited items still largely relies on human judgment, and the accuracy and efficiency of detection have not improved sufficiently. Furthermore, the emotions and cognitive state of security staff conducting the inspections can influence their ability to respond to potential risks. However, there is a lack of means to effectively incorporate these factors and improve accuracy through real-time analysis and feedback.
[0761] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[0762] In this invention, the server includes means for collecting video information acquired from a baggage inspection device, means for pre-processing the video information into a format suitable for analysis, means for inputting the pre-processed video information into a generating AI algorithm to detect dangerous or prohibited items, and means for presenting the results in real time using a visual display worn by the user and analyzing the user's reaction. This enables advanced detection of dangerous items that takes into account the emotional state of security staff and allows for faster countermeasures.
[0763] A "baggage inspection device" is a device that scans passenger and cargo luggage and makes visible the items contained inside.
[0764] "Visual information" refers to digital video data acquired by baggage inspection equipment, including X-ray images and visible light images.
[0765] A "generative AI algorithm" is a program that uses artificial intelligence technology to extract features from input data and identify dangerous or prohibited items.
[0766] "Object recognition techniques" are technologies that use algorithms to analyze and identify specific objects within images or videos.
[0767] A "visual display" is a device incorporated into devices such as smart glasses that visually presents information to the user.
[0768] "User emotional data" refers to data on the emotional state of the user, evaluated from their facial expressions and voice, acquired through visual displays and other sensors.
[0769] "Feedback" refers to information that is used to improve system operation and algorithms based on user feedback.
[0770] This system is designed to provide baggage inspection with integrated anomaly detection and user assistance. The server first collects video information acquired by the baggage inspection device. This information is comprehensive image data, including X-ray and visible light images in particular.
[0771] The server first preprocesses the received video information. This processing includes noise reduction and image quality improvement. The goal is to prepare the data so that the generative AI algorithm can perform analysis with high accuracy. Next, the server sends the preprocessed video information to the generative AI algorithm for analysis to identify dangerous or prohibited items. This generative AI algorithm utilizes object recognition techniques and has the ability to detect targets efficiently and effectively.
[0772] The analysis results are presented in real time to the user wearing a visual display. This visual display is typically implemented as a device like smart glasses and is responsible for visualizing the information. The system also analyzes the user's facial expressions and voice tone to collect emotional data. This emotional data reflects the user's level of alertness and anxiety and is incorporated into the system as feedback.
[0773] As a concrete example, when airport security staff use glasses-type visual displays to inspect passengers' baggage, any abnormalities are immediately detected by the images displayed on the screen. If the staff member feels uneasy, the system triggers additional inspection procedures based on emotional data. This entire process ensures safer and more efficient baggage inspection.
[0774] An example of a prompt to the generating AI model is as follows: "Analyze the X-ray images of the luggage to detect dangerous items. Also, analyze the user's facial expression data to assess their emotions and evaluate their importance." This prompt allows the system to perform a more precise analysis using the AI algorithm.
[0775] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0776] Step 1:
[0777] The server acquires video information from the baggage inspection device. This information includes X-ray images and visible light images. The acquired video information is received as input data and passed to the next preprocessing step.
[0778] Step 2:
[0779] The server performs preprocessing on the acquired video information. This preprocessing includes noise reduction and image quality correction to improve image quality. Finally, the data is converted into a format that is easy for the generating AI algorithm to analyze. The output of this process is the preprocessed video information.
[0780] Step 3:
[0781] The server inputs pre-processed video information into a generating AI algorithm. The generating AI algorithm analyzes the data using object recognition techniques to detect dangerous or prohibited items. The results of the analysis are output as a list of detected items.
[0782] Step 4:
[0783] The server assesses the risk level of the detected items based on the list and prepares to notify the user. This assessment sorts the items in descending order of risk based on their characteristics. The assessment results are output as data for notification.
[0784] Step 5:
[0785] On the device, evaluation results are presented to the user via a visual display. The visual display shows information in real time on smart glasses. The user reviews the presented information and provides feedback in response. This feedback is recorded as input data to the system.
[0786] Step 6:
[0787] To collect user emotion data, the device analyzes the user's voice and facial expressions. The analysis uses an emotion analysis engine to evaluate the emotional state. The output of this analysis is the user's emotion parameters.
[0788] Step 7:
[0789] The server retrains the generating AI algorithm based on user feedback and emotion parameters. This retraining is aimed at improving analysis accuracy and efficiency, optimizing the performance of future baggage screening.
[0790] Step 8:
[0791] The server automatically triggers additional testing procedures based on the user's emotional state. Signs of anxiety or alertness identified through emotion analysis are treated as high risk, and immediate additional action is taken. The output of this step is an instruction to execute the automated additional procedure.
[0792] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the controlled object 443 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[0793] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0794] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the robot 414.
[0795] Furthermore, the emotion identification model 59, acting as an emotion engine, may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to a specific mapping, which is an emotion map (see Figure 9). Similarly, the emotion identification model 59 may also determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.
[0796] Figure 9 shows an emotion map 400 in which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. The closer to the center of the concentric circles, the more primitive the emotions are located. Further out of the concentric circles, emotions representing states and actions arising from mental states are located. Emotion is a concept that includes feelings and mental states. On the left side of the concentric circles, emotions that are generally generated from reactions occurring in the brain are located. On the right side of the concentric circles, emotions that are generally induced by situational judgment are located. Above and below the concentric circles, emotions that are generally generated from reactions occurring in the brain and induced by situational judgment are located. In addition, the emotion of "pleasure" is located on the upper side of the concentric circles, and the emotion of "displeasure" is located on the lower side. Thus, in the emotion map 400, multiple emotions are mapped based on the structure in which emotions arise, and emotions that are likely to occur simultaneously are mapped close together.
[0797] These emotions are distributed at the 3 o'clock position on the Emotion Map 400, and usually fluctuate between feelings of security and anxiety. In the right half of the Emotion Map 400, situational awareness takes precedence over internal feelings, resulting in a calm impression.
[0798] The inside of the Emotion Map 400 represents inner thoughts, while the outside represents actions. Therefore, the further you go from the outside of the Emotion Map 400, the more visible (expressed in actions) your emotions become.
[0799] Here, human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. Similarly, in robots, cars, motorcycles, etc., emotions can be created based on various balances, such as posture and battery level. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. The emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on a system for analyzing brain physiological signals of speech emotion recognition and emotion, Tokushima University, doctoral dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map contains emotions belonging to a region called "response," where sensation is dominant. The right half of the emotion map contains emotions belonging to a region called "situation," where situational awareness is dominant.
[0800] The emotion map defines two emotions that promote learning. One is the emotion around the middle of the negative "repentance" and "reflection" on the situation side. In other words, it is when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is the emotion around the positive "desire" on the reaction side. In other words, it is when the robot has positive feelings such as "I want more" or "I want to know more."
[0801] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values representing each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple training data sets, which are combinations of user input and emotion values representing each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions located close together have similar values, as shown in the emotion map 900 in Figure 10. Figure 10 shows an example where multiple emotions such as "reassured," "calm," and "confident" have similar emotion values.
[0802] The above description primarily focuses on the functions of the data processing device 12 in relation to this disclosure. However, the system related to this disclosure is not necessarily implemented on a server. The system related to this disclosure may be implemented as a general information processing system. This disclosure may be implemented, for example, as a software program that runs on a personal computer or as an application that runs on a smartphone. The method related to this disclosure may be provided to users in SaaS (Software as a Service) format.
[0803] In the above embodiment, an example was given in which a specific process is performed by a single computer 22. However, the technology of this disclosure is not limited thereto, and a distributed processing of the specific process may be performed by multiple computers, including computer 22. For example, a data generation model 58 may be provided in an external device of the data processing device 12, and the external device may generate data according to the input data.
[0804] In the above embodiment, an example was given in which the specific processing program 56 is stored in the storage 32, but the technology of this disclosure is not limited thereto. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-temporary storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-temporary storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes specific processing according to the specific processing program 56.
[0805] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.
[0806] Furthermore, it is not necessary to store the entirety of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store the entirety of the specific processing program 56 in the storage 32; it is acceptable to store only a portion of the specific processing program 56.
[0807] The following types of processors can be used as hardware resources to perform specific processing. Examples of processors include a CPU, a general-purpose processor that functions as a hardware resource to perform specific processing by executing software, i.e., a program. Other examples of processors include dedicated electrical circuits, such as FPGAs (Field-Programmable Gate Arrays), PLDs (Programmable Logic Devices), or ASICs (Application Specific Integrated Circuits), which have circuit configurations specifically designed to perform specific processing. All of these processors have built-in or connected memory, and all of them perform specific processing by using memory.
[0808] The hardware resource that performs a specific process may consist of one of these various processors, or it may consist of a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Alternatively, the hardware resource that performs a specific process may consist of a single processor.
[0809] Examples of configurations using a single processor include, firstly, a configuration in which one or more CPUs and software are combined to form a single processor, and this processor functions as a hardware resource that performs a specific process. Secondly, there is a configuration using a processor that realizes the functions of the entire system, including multiple hardware resources that perform a specific process, on a single IC chip, as exemplified by SoCs (System-on-a-chip). In this way, a specific process is realized using one or more of the above types of processors as hardware resources.
[0810] Furthermore, the hardware structure of these various processors can more specifically utilize electrical circuits that combine circuit elements such as semiconductor devices. Also, the specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps can be deleted, new steps added, or the processing order rearranged, as long as it does not deviate from the main purpose.
[0811] The descriptions and illustrations presented above are detailed explanations of the technical aspects of this disclosure and are merely examples of the technical aspects. For example, the above descriptions of the structure, function, operation, and effect are examples of the structure, function, operation, and effect of the technical aspects of this disclosure. Therefore, it goes without saying that you may delete unnecessary parts, add new elements, or replace elements in the descriptions and illustrations presented above, as long as you do not deviate from the essence of the technical aspects of this disclosure. Furthermore, in order to avoid confusion and facilitate understanding of the technical aspects of this disclosure, explanations of common technical knowledge and the like that do not require special explanation to enable the implementation of the technical aspects of this disclosure have been omitted from the descriptions and illustrations presented above.
[0812] All documents, patent applications, and technical standards described herein are incorporated by reference to the same extent as if each individual document, patent application, and technical standard were specifically and individually noted to be incorporated by reference.
[0813] The following is further disclosed regarding the embodiments described above.
[0814] (Claim 1)
[0815] A means of collecting video data obtained from baggage inspection devices,
[0816] A means for preprocessing the aforementioned video data into a format suitable for analysis,
[0817] The aforementioned pre-processed video data is input into a generating AI model, and means for detecting hazardous materials or prohibited items,
[0818] A means for determining the degree of risk based on the detection results and notifying the result,
[0819] A system that includes a means for users to provide feedback on their findings based on the notified results.
[0820] (Claim 2)
[0821] The system according to claim 1, characterized in that the generated AI model uses object recognition technology to analyze X-ray images and improve detection accuracy.
[0822] (Claim 3)
[0823] The system according to claim 1, characterized in that the generated AI model is retrained using the aforementioned feedback to improve the accuracy and efficiency of future analyses.
[0824] "Example 1"
[0825] (Claim 1)
[0826] A means for collecting image information obtained from a video acquisition device,
[0827] A preprocessing means for converting the aforementioned image information into a format suitable for analysis,
[0828] The means of inputting the pre-processed image information into a generating AI model to identify items with safety issues,
[0829] A means for evaluating the safety level based on identified items and notifying the evaluation results,
[0830] A system that includes means for operators to provide feedback on the results based on the notified evaluation results.
[0831] (Claim 2)
[0832] The system according to claim 1, characterized in that the generating AI model uses object recognition technology to analyze image information and improve identification accuracy.
[0833] (Claim 3)
[0834] The system according to claim 1, characterized in that it reconstructs the generated AI model using the aforementioned feedback to improve future analysis accuracy and work efficiency.
[0835] "Application Example 1"
[0836] (Claim 1)
[0837] A means of collecting video data obtained from baggage inspection devices,
[0838] A means for preprocessing the aforementioned video data into a format suitable for analysis,
[0839] The aforementioned pre-processed video data is input into a generating AI model, and means for detecting hazardous materials or prohibited items,
[0840] A means for determining the degree of risk based on the detection results and notifying the result,
[0841] Based on the notified results, a means for the user to provide feedback on the verification results,
[0842] A means of using a camera on an individual device to photograph items at the entrance of a public facility and transmitting the data to a cloud server,
[0843] A means of performing data analysis on a cloud server to determine the hazards of goods,
[0844] A system that includes means for displaying analysis results to the user's device in real time and issuing alarms.
[0845] (Claim 2)
[0846] The system according to claim 1, characterized in that the generated AI model uses object recognition technology to analyze X-ray images and improve detection accuracy.
[0847] (Claim 3)
[0848] The system according to claim 1, characterized in that the generated AI model is retrained using the aforementioned feedback to improve the accuracy and efficiency of future analyses.
[0849] "Example 2 of combining an emotion engine"
[0850] (Claim 1)
[0851] A means of collecting video data obtained from baggage inspection devices,
[0852] A means for preprocessing the aforementioned video data into a format suitable for analysis,
[0853] The aforementioned pre-processed video data is input into a generating AI model, and means for detecting hazardous materials or prohibited items,
[0854] A means for determining the degree of risk based on the detection results and notifying the result,
[0855] A means of analyzing a user's voice and facial expressions using an emotion analysis device to evaluate their emotions,
[0856] A system including means for adjusting the testing procedure in accordance with the results of the evaluation of the aforementioned emotions.
[0857] (Claim 2)
[0858] The system according to claim 1, characterized in that the generating AI model uses object recognition technology to analyze X-ray images, improves detection accuracy, and flexibly adjusts the examination procedure using the emotion evaluation results.
[0859] (Claim 3)
[0860] The system according to claim 1, characterized in that it retrains the generative AI model using the aforementioned feedback and emotion analysis results to improve future analysis accuracy and efficiency.
[0861] "Application example 2 when combining with an emotional engine"
[0862] (Claim 1)
[0863] A means of collecting video information obtained from baggage inspection devices,
[0864] means for preprocessing the aforementioned video information into a format suitable for analysis,
[0865] The means for inputting the pre-processed video information into a generating AI algorithm to detect dangerous or prohibited items,
[0866] A means for determining the degree of risk based on the detection results and notifying the result,
[0867] Based on the notified results, a means for users to provide feedback on the verification results,
[0868] A means for presenting the results in real time using a visual display worn by the user and analyzing the user's reaction,
[0869] A system including means for automatically triggering additional testing procedures based on the aforementioned user's emotional data.
[0870] (Claim 2)
[0871] The system according to claim 1, characterized in that the generation AI algorithm uses an object recognition method to analyze X-ray images and improve detection accuracy.
[0872] (Claim 3)
[0873] The system according to claim 1, characterized in that it retrains the generative AI algorithm using the aforementioned feedback and emotion data to improve the accuracy and efficiency of future analyses. [Explanation of Symbols]
[0874] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Devices 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robots< / url:> < / url:> < / url:> < / url:>
Claims
1. A means of collecting video data obtained from baggage inspection devices, A means for preprocessing the aforementioned video data into a format suitable for analysis, The aforementioned pre-processed video data is input into a generating AI model, and means for detecting hazardous materials or prohibited items, A means for determining the degree of risk based on the detection results and notifying the result, A system that includes a means for users to provide feedback on their findings based on the notified results.
2. The system according to claim 1, characterized in that the generated AI model uses object recognition technology to analyze X-ray images and improve detection accuracy.
3. The system according to claim 1, characterized in that the generated AI model is retrained using the aforementioned feedback to improve the accuracy and efficiency of future analyses.
Citation Information
Patent Citations
Persona chatbot control method and system
JP2022180282A