system

A system using mobile device imaging and machine learning to provide accurate waste sorting instructions addresses regional rule variations, enhancing waste disposal accuracy and efficiency.

JP2026071565APending Publication Date: 2026-04-30SOFTBANK GROUP CORP
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
SOFTBANK GROUP CORP
Filing Date
2024-10-17
Publication Date
2026-04-30

AI Technical Summary

Technical Problem

Users face difficulties in accurately sorting waste due to varying regional rules, leading to improper disposal and environmental issues.

Method used

A system that utilizes a mobile device to capture images of waste, analyzes them using machine learning, and refers to local waste sorting regulations to provide users with accurate sorting instructions.

Benefits of technology

Enables users to easily and accurately sort waste, improving local waste management efficiency and adherence to environmental regulations.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026071565000001_ABST
    Figure 2026071565000001_ABST
Patent Text Reader

Abstract

We provide the system. [Solution] A means of receiving images of waste taken by the user with a mobile device, A means for classifying waste by analyzing the aforementioned images, Based on the above classification, means for identifying the correct sorting method by referring to the relevant local waste sorting regulations, Means for notifying the user of the specified sorting method, A system that includes this.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The technology of the present disclosure relates to a system.

Background Art

[0002] Patent Document 1 discloses a persona chatbot control method performed by at least one processor, including steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to an explanation of a chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance as a response to the user utterance.

Prior Art Documents

Patent Documents

[0003]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0004] Due to different waste sorting rules in each region, many users find it difficult to immediately determine what the correct sorting method is. As a result, inappropriate sorting is carried out, which may cause problems in waste reprocessing and adverse effects on the environment. The purpose of the present invention is to solve such problems of inaccurate sorting and provide a system that enables users to easily and accurately sort waste.

Means for Solving the Problems

[0005] The present invention includes means for receiving images of waste captured by a user with a mobile device, and means for analyzing these images to classify the waste. Furthermore, it includes means for identifying the correct sorting method based on this classification and by referring to relevant local waste sorting regulations. It also includes means for notifying the user of the identified sorting method, thereby providing a system that allows the user to easily and accurately sort waste.

[0006] "Users" refers to people who want to know the correct way to sort waste.

[0007] "Mobile device" refers to mobile devices and smartphones equipped with a camera function.

[0008] "Waste" refers to objects or materials that a user intends to dispose of.

[0009] "Image" refers to visual information of waste captured with a mobile device.

[0010] "Means of receiving" refers to the interface or process for collecting image data from a mobile device to a server.

[0011] "Means of analysis and classification" refers to the process of identifying the type of waste using image analysis technology.

[0012] "Relevant area" refers to the municipality or district where the user currently resides or lives.

[0013] "Waste sorting regulations" refer to guidelines established in each region for the proper sorting and disposal of waste.

[0014] "Means of identification" refers to the process of selecting an appropriate classification method based on the analysis results.

[0015] "Means of notification" refers to communication or display technologies used to transmit information on sorting methods to users. [Brief explanation of the drawing]

[0016] [Figure 1] It is a conceptual diagram showing an example of the configuration of a data processing system according to the first embodiment. [Figure 2] It is a conceptual diagram showing an example of the main functions of a data processing device and a smart device according to the first embodiment. [Figure 3] It is a conceptual diagram showing an example of the configuration of a data processing system according to the second embodiment. [Figure 4] It is a conceptual diagram showing an example of the main functions of a data processing device and smart glasses according to the second embodiment. [Figure 5] It is a conceptual diagram showing an example of the configuration of a data processing system according to the third embodiment. [Figure 6] It is a conceptual diagram showing an example of the main functions of a data processing device and a headset-type terminal according to the third embodiment. [Figure 7] It is a conceptual diagram showing an example of the configuration of a data processing system according to the fourth embodiment. [Figure 8] It is a conceptual diagram showing an example of the main functions of a data processing device and a robot according to the fourth embodiment. [Figure 9] Shows an emotion map to which multiple emotions are mapped. [Figure 10] Shows an emotion map to which multiple emotions are mapped. [Figure 11] It is a sequence diagram showing the processing flow of the data processing system in Example 1. [Figure 12] It is a sequence diagram showing the processing flow of the data processing system in Application Example 1. [Figure 13] It is a sequence diagram showing the processing flow of the data processing system in Example 2 when the emotion engine is combined. [Figure 14] It is a sequence diagram showing the processing flow of the data processing system in Application Example 2 when the emotion engine is combined.

Modes for Carrying Out the Invention

[0017] Hereinafter, an example of an embodiment of a system according to the technology of the present disclosure will be described with reference to the accompanying drawings.

[0018] First, the terms used in the following description will be explained.

[0019] In the following embodiments, a labeled processor (hereinafter simply referred to as "processor") may be a single arithmetic unit or a combination of multiple arithmetic units. Also, the processor may be a single type of arithmetic unit or a combination of multiple types of arithmetic units. Examples of arithmetic units include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), an APU (Accelerated Processing Unit), and the like.

[0020] In the following embodiments, a labeled RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a work memory by the processor.

[0021] In the following embodiments, a labeled storage is one or more non-volatile storage devices that store various programs and various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), or magnetic tapes, and the like.

[0022] In the following embodiments, the signed communication interface (I / F) is an interface that includes a communication processor and an antenna, etc. The communication interface manages communication between multiple computers. Examples of communication standards applicable to the communication interface include wireless communication standards such as 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), or Bluetooth (registered trademark).

[0023] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." That is, "A and / or B" means that it may be A alone, or B alone, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" applies when expressing three or more things linked by "and / or."

[0024] [First Embodiment]

[0025] Figure 1 shows an example of the configuration of the data processing system 10 according to the first embodiment.

[0026] As shown in Figure 1, the data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.

[0027] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0028] The smart device 14 comprises a computer 36, a reception device 38, an output device 40, a camera 42, and a communication interface 44. The computer 36 comprises a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The reception device 38, output device 40, and camera 42 are also connected to the bus 52.

[0029] The reception device 38 is equipped with a touch panel 38A and a microphone 38B, etc., and receives user input. The touch panel 38A receives user input by detecting contact with an object (e.g., a pen or finger). The microphone 38B receives user input by detecting the user's voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.

[0030] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form perceptible to the user 20 (e.g., audio and / or text). The display 40A displays visible information such as text and images according to instructions from the processor 46. The speaker 40B outputs audio according to instructions from the processor 46. The camera 42 is a small digital camera equipped with an optical system such as a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.

[0031] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various types of information between processor 46 and processor 28 via network 54.

[0032] Figure 2 shows an example of the main functions of the data processing device 12 and the smart device 14.

[0033] As shown in Figure 2, in the data processing device 12, a specific processing is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" related to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.

[0034] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0035] In the smart device 14, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The reception output program 60 is used in conjunction with a specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0036] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".

[0037] The system of the present invention assists users in sorting the waste they generate on a daily basis and consists of a mobile terminal, a server, and related software. The program processing of the system is described below in natural language.

[0038] The main components of this system include image capture, image analysis, information retrieval, and classification instruction functions.

[0039] First, the user uses their mobile device to take a picture of the waste using the deployed camera function. This image is then sent to the server via the application.

[0040] The server applies an image analysis algorithm to the received image to identify the type of waste. This analysis uses machine learning techniques to detect the characteristics of the waste with high accuracy based on past data.

[0041] Furthermore, the server uses the user's location information to refer to a database of relevant regions and retrieve sorting rules that are suitable for the characteristics of the waste. Subsequently, the retrieved rules are compared with the analysis results to identify the correct waste sorting method.

[0042] The identified sorting method is transmitted from the server to the terminal. The terminal receives this information and displays it to the user in an easy-to-understand format. This display includes detailed instructions such as the waste sorting category, collection date, and the appropriate type of garbage bag.

[0043] As a concrete example, if a user takes a picture of a plastic bottle from their home, the server will recognize it as plastic waste and provide instructions to separate it as a "recyclable resource" according to the rules of the relevant area.

[0044] Thus, this invention helps users easily and accurately sort waste, contributing to the efficiency of local waste management.

[0045] The following describes the processing flow.

[0046] Step 1:

[0047] Users use their mobile devices to take photos of waste using the application's camera function. After taking the photo, the image is uploaded to the server by tapping the send button within the app.

[0048] Step 2:

[0049] The server receives image data of the waste sent by the user. The received images are temporarily stored in storage and prepared to be sent to the analysis process.

[0050] Step 3:

[0051] The server uses an image analysis algorithm to analyze the image data. Here, a machine learning model extracts features of the waste and identifies its type. For example, it determines categories such as plastic bottles and cans based on shape, color, label information, etc.

[0052] Step 4:

[0053] The server uses the user's location information to access the waste sorting database for that area. It searches for sorting rules corresponding to the identified waste type and obtains the correct sorting method.

[0054] Step 5:

[0055] The server uses the acquired sorting information to generate an easy-to-understand instruction message for the user. This message includes the sorting category and handling precautions.

[0056] Step 6:

[0057] The server sends the generated instruction message to the terminal. The terminal displays the received message on the app screen, showing the user the proper way to sort waste.

[0058] Step 7:

[0059] Users follow instructions from the terminal to sort waste into designated categories and dispose of it appropriately in accordance with local regulations.

[0060] (Example 1)

[0061] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."

[0062] Conventional waste sorting methods require users to manually understand and apply classification criteria, which is time-consuming and cumbersome. Furthermore, adapting to varying waste sorting rules across different regions is difficult, increasing the risk of incorrect disposal. This invention aims to solve these problems and enable users to sort waste quickly and accurately.

[0063] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0064] In this invention, the server includes means for receiving image data generated by a user in a portable information terminal equipped with a camera function, means for analyzing the image data using machine learning technology to recognize the characteristics of the waste, and means for referring to location-dependent regional waste disposal procedure data based on the recognition to identify an appropriate classification method. This enables the user to efficiently sort waste in accordance with local waste sorting rules.

[0065] An "information processing device" is an electronic device programmed to process data and perform specific tasks.

[0066] A "personal information terminal" is an electronic device that a user can carry with them and that has data input and display functions.

[0067] "Image data" refers to visual information generated using a camera function, expressed in digital format.

[0068] "Machine learning technology" is a collection of algorithms and methodologies for learning from past data and making predictions and classifications on new data.

[0069] "Waste characteristics" refer to classification criteria based on the material, shape, and intended use of an object.

[0070] "Location information" is digital data used to identify a specific location, and is acquired using technologies such as GPS.

[0071] "Regional waste disposal procedure data" refers to information regarding waste sorting methods and regulations established in each region.

[0072] A "classification method" is a set of procedures and criteria for dividing waste into specific categories based on its characteristics.

[0073] To implement the present invention, the system is implemented with a configuration including an information processing device, a personal information terminal, and related software. Specifically, the personal information terminal is equipped with a camera function, which the user uses to generate image data of objects that are routinely discarded. This image data is then transmitted to a server through a dedicated application.

[0074] The server analyzes the received image data. Machine learning techniques are used for the analysis, utilizing popular machine learning frameworks such as TENSORFLOW® and PyTorch. This technology allows the server to recognize the characteristics of the waste with high accuracy and identify its material and shape.

[0075] Furthermore, the server uses location information obtained from the user's mobile device to refer to waste disposal procedure data for each region. Location information is typically obtained using GPS technology, helping the server identify sorting rules relevant to the user's current location. Based on this information, the server identifies the appropriate sorting method and notifies the user.

[0076] Notifications are sent in real time, and the device displays the information clearly and visually to the user. The display includes instructions such as the waste classification category, collection date, and the appropriate type of garbage bag.

[0077] For example, if a user takes a picture of a plastic bottle from their home, the system will recognize it as plastic waste and provide instructions to process it as a "recyclable resource" as defined by the local area. This makes it easy for users to dispose of their waste according to local regulations.

[0078] An example of a prompt to a generating AI model is, "Based on the image data taken by the user, output instructions regarding the appropriate sorting category and specific collection date." This prompt allows the AI ​​model to provide accurate and practical sorting instructions.

[0079] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0080] Step 1:

[0081] The user uses the camera function of their mobile device to take pictures of the waste. The captured image data is sent to the server via the application. The input is the image data captured by the camera, and the output is the digital image data received by the server.

[0082] Step 2:

[0083] The server analyzes the received image data using a machine learning algorithm. Specifically, a model built using TensorFlow or PyTorch classifies objects in the image based on their characteristics as waste. The input to this step is the image data received by the server, and the output is information about the characteristics of the waste based on that image.

[0084] Step 3:

[0085] The server uses location information obtained from the user's terminal to refer to a regional waste disposal procedure database. The input is location information obtained using GPS, and the output is a list of region-specific sorting methods based on that location information.

[0086] Step 4:

[0087] The server compares the image analysis results with the acquired local sorting rules to identify the appropriate waste classification method. This comparison process determines which sorting category the waste belongs to. The input is the waste's characteristics and a list of local sorting methods, and the output is the optimal sorting method.

[0088] Step 5:

[0089] The server transmits the identified sorting method to the user's mobile device. Communication takes place in real time, and the user can receive the results immediately. The input is the sorting method identified by the server, and the output is the sorting instruction information displayed on the user's device.

[0090] Step 6:

[0091] The terminal visually displays the received sorting information to the user. The display screen shows the waste classification category, collection date, and appropriate garbage bag type, allowing the user to process the waste accordingly. The input is the notified sorting instruction information, and the output is the detailed instruction content displayed on the user's screen.

[0092] (Application Example 1)

[0093] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."

[0094] Waste sorting is a crucial issue for environmental protection and efficient resource utilization, but conventional methods have limitations in terms of accuracy and efficiency. In particular, large amounts of waste are generated in industry and manufacturing, requiring proper disposal. On-site sorting often relies on manual labor, leading to problems such as the depletion of human resources and increased environmental burden due to errors. Therefore, there is a need for automated, high-precision, and highly efficient sorting systems.

[0095] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0096] In this invention, the server includes means for receiving visual information of waste captured by a user with a device having a camera function, means for classifying the waste by analyzing the visual information, means for identifying the correct disposal method by referring to the relevant local waste disposal regulations, and means for physically separating the waste using a powered machine. This enables automatic and accurate sorting of waste.

[0097] "User" refers to an individual or organization that utilizes the waste sorting system.

[0098] "Photography function" refers to the capability of a device to capture visual information about waste.

[0099] "Device" refers to hardware components, including cameras and sensors.

[0100] "Visual information" refers to image data acquired using a camera or other imaging function.

[0101] "Analysis" refers to the process of identifying the type and characteristics of waste from visual information using machine learning algorithms.

[0102] "Classification" refers to dividing waste into predetermined categories based on the results of analysis.

[0103] "Local waste disposal regulations" refer to laws and guidelines concerning the collection, treatment, and recycling of waste that apply in a specific area.

[0104] "Disposal method" refers to the procedures for properly processing or recycling waste.

[0105] A "powered machine" refers to a device that supplies energy for performing mechanical work.

[0106] "Physically separating" refers to the process of physically separating waste using mechanical equipment rather than manual methods.

[0107] In embodiments of the present invention, a waste sorting system is used. The system consists of a device with a camera function (hereinafter referred to as a terminal), a server that processes data, and related software. The user uses the camera function of the terminal to acquire visual information of the waste to be sorted. The acquired image data is transmitted to the server via the network.

[0108] The server first uses machine learning algorithms on the received visual information to identify the type of waste. This analysis utilizes previously accumulated waste characteristic information to enable highly accurate classification. Based on the identified information, the server considers the user's location and refers to the waste disposal regulations database for the relevant area to determine the appropriate disposal method.

[0109] The identified disposal method includes information on means of physically separating the waste using powered machinery, and the generated instructions are sent to the terminal. The terminal displays visual information, disposal methods, and precautions to the user in an easy-to-understand format.

[0110] For example, if a user takes a picture of a plastic part used in a factory, the server will identify it as recyclable material and generate instructions on how to transport it to the appropriate recycling bin.

[0111] Examples of prompts to input into a generative AI model:

[0112] "Please explain how a robot operating within a factory can photograph waste materials, perform image analysis, and generate sorting instructions based on the results."

[0113] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0114] Step 1:

[0115] The user takes a picture of the waste using their device. The input is capturing the actual waste with the camera. The output is digital image data generated as pixel information. Specifically, the user presses the capture button, which activates the camera and saves a photograph of the object.

[0116] Step 2:

[0117] The device transmits the captured image data to the server via the network. The input is the digital image obtained in step 1. The output is the image data received by the server. Specifically, the device establishes a network connection and uploads the image data to the server using the HTTP protocol.

[0118] Step 3:

[0119] The server analyzes the received image data to identify the type of waste. The input is image data sent from the terminal. The output is the identified waste type and characteristic information. Specifically, the server uses computer vision technology and machine learning algorithms to extract features from the image and classify them by comparing them with a database.

[0120] Step 4:

[0121] The server obtains the user's location information and refers to a database of waste disposal regulations for the relevant region. Inputs are the type of waste and the user's location. Outputs are instructions for the appropriate disposal method. Specifically, the server uses GPS data to determine the location and retrieves relevant information from the database based on that location.

[0122] Step 5:

[0123] The server determines the means of physically separating the waste using a powered machine and sends instructions to the terminal. The input is the instruction on the disposal method. The output is a sorting method guide displayed on the terminal. Specifically, the server converts the instruction content into text format and sends it to the terminal.

[0124] Step 6:

[0125] The terminal visually presents the received instructions to the user. The input is the sorting method instructions sent from the server. The output is the instruction screen that the user can view. Specifically, the terminal displays the instruction content on its screen and provides information using icons and text so that the user can easily understand it.

[0126] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0127] This invention is a system that assists users in processing waste they sort on a daily basis and incorporates an emotion engine to provide appropriate feedback according to the user's emotional state. The system's program processing is described below in natural language.

[0128] This system consists of functions related to image capture, emotion recognition, image analysis, information retrieval, and classification instructions.

[0129] First, the user uses their mobile device to take a picture of the waste using the camera. This image is sent to the server via the application. In addition, the device simultaneously collects data on the user's voice and facial expressions, which are then transmitted to the server for analysis of the user's emotional state by an emotion engine.

[0130] The server analyzes received images of waste using an image analysis algorithm to identify the type of waste. This analysis utilizes a machine learning model, achieving highly accurate classification based on accumulated data. In addition, an emotion engine analyzes the user's voice tone and facial expression data to determine their current emotional state (e.g., joy, surprise, dissatisfaction).

[0131] Next, the server uses the user's location information to refer to the relevant regional waste sorting database. Here, it searches for the appropriate sorting rules corresponding to the identified waste type and determines the correct disposal method for the waste.

[0132] Based on the identified classification method and the user's emotional state, the server generates encouraging messages and situation-specific warning messages. For example, if the user is dissatisfied, a positive message such as "You're doing well, let's move on!" might be generated.

[0133] This information is sent from the server to the terminal, which then displays sorting instructions and messages on the screen to provide the user with appropriate guidance and motivation.

[0134] As a concrete example, when a user photographs a plastic bottle that should be recycled, the server recognizes it as plastic waste and instructs it to separate it as "recyclable material" according to the relevant local rules. At the same time, if the emotion engine detects dissatisfaction from the user's facial expression, it also displays a message to boost their motivation.

[0135] Thus, the system of the present invention not only supports users in properly sorting waste, but also provides emotionally responsive feedback to realize a more comfortable and effective user experience.

[0136] The following describes the processing flow.

[0137] Step 1:

[0138] The user launches the application on their mobile device and takes a picture of the waste using the camera function. If audio or facial expressions are available, the app will also record this data simultaneously.

[0139] Step 2:

[0140] The device transmits captured image data, recorded audio, and facial expression data to a server via the app. This data also includes the user's location information.

[0141] Step 3:

[0142] The server receives the image data and first performs image analysis using a machine learning algorithm. This analysis extracts the characteristics of the waste and identifies its type.

[0143] Step 4:

[0144] The server analyzes voice and facial expression data using an emotion engine to determine the user's emotional state. For example, it analyzes voice tone and detects changes in facial expressions to identify whether the user is happy or dissatisfied.

[0145] Step 5:

[0146] The server accesses a waste sorting database for the relevant region based on the user's location information. Using the analysis results, it retrieves the appropriate sorting method for the waste from the database.

[0147] Step 6:

[0148] The server generates notification messages for the user based on the acquired sorting method. Furthermore, it adds positive messages and words of encouragement depending on the user's emotional state.

[0149] Step 7:

[0150] The terminal receives a message sent from the server and displays it on the screen. This message includes instructions on how to properly sort waste and emotionally responsive feedback.

[0151] Step 8:

[0152] Users review the displayed information and follow the instructions to properly sort their waste. The positive messages displayed are also expected to boost user motivation.

[0153] (Example 2)

[0154] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".

[0155] In today's environmental challenges, the sorting of waste that users perform daily is important, but accurate classification and adherence to local sorting rules are difficult, sometimes resulting in improper disposal. Furthermore, insufficient feedback on user behavior makes it difficult to maintain motivation for sorting.

[0156] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0157] In this invention, the server includes means for receiving images of objects captured by the user with a communication terminal, means for analyzing the images and classifying the objects, and means for receiving audio data and facial expression data and determining the user's emotional state. This makes it possible to support accurate object classification while providing appropriate feedback according to the user's emotions.

[0158] A "communication terminal" is an electronic device used by a user to send and receive information, and is equipped with camera and voice input functions.

[0159] An "image of an object" is digital data that shows the visible shape of an object, captured using the camera of a communication terminal.

[0160] "Analyzing an image" is the process of identifying the features of an object from digital data and classifying its attributes.

[0161] "Object classification" is the act of dividing objects into specific categories based on their analyzed characteristics.

[0162] "Audio data" refers to information that represents acoustic signals acquired through the microphone of a communication terminal in digital format.

[0163] "Facial expression data" refers to digital data containing information about facial expressions captured by the camera of a communication device.

[0164] "User's emotional state" refers to information indicating the psychological state obtained by analyzing voice data and facial expression data, and includes examples such as "joy," "dissatisfaction," and "surprise."

[0165] "Generating adaptive notification messages that respond to emotions" means creating feedback messages that are appropriate for individual situations, taking into account the user's emotional state.

[0166] This invention is a system that links a communication terminal and a server to assist users in sorting objects on a daily basis. Specifically, the communication terminal uses a camera and microphone to capture image data by photographing objects, and in addition, it collects the user's voice and facial expressions. This data is transmitted from the terminal to the server.

[0167] The server uses generative AI models, such as Convolutional Neural Networks (CNNs), to analyze the features of objects based on the received image data. The image analysis algorithm classifies objects into specific categories. The server also utilizes facial recognition and natural language processing (NLP) technologies to analyze audio and facial data and determine the user's emotional state (joy, surprise, dissatisfaction, etc.).

[0168] Based on the analyzed data, the server references the user's location information and retrieves region-specific object classification rules from a database. Following these rules, it identifies the optimal method for processing the object and generates feedback tailored to the identified emotional state. For example, if the user expresses dissatisfaction, it might create an encouraging message such as, "Very well done, keep up the good work next time!"

[0169] Ultimately, this information is transmitted to the terminal, and the results are displayed to the user on the screen. This allows the user to efficiently perform the appropriate processing of the object while maintaining motivation.

[0170] For example, if a user photographs a reusable plastic bottle, the server will recognize the object as "plastic" and instruct the user to separate it as "recyclable waste" according to local rules. At the same time, if sentiment analysis determines that the user is frustrated, it will display a positive message such as, "Great job, please recycle next time!"

[0171] An example of a prompt to input into the generating AI model is, "Please tell me the steps to determine the type of object, apply the sorting rules for that location, and generate the optimal action."

[0172] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0173] Step 1:

[0174] The user takes an image of an object using the camera on the communication terminal. The input image data must be high resolution so that the entire object is clearly visible. As output, the captured image is saved to the memory of the communication terminal.

[0175] Step 2:

[0176] The device prepares to send the captured image data to the server. It also collects user voice data using the microphone and acquires facial expression data using the camera. The inputs are image, audio, and facial expression data, and these are sent to the server as output.

[0177] Step 3:

[0178] The server processes the received image data through an image analysis system to extract object features. This analysis uses a Convolutional Neural Network (CNN) to analyze features such as object shape and color at the pixel level. The input is image data, and the output is the object classification result.

[0179] Step 4:

[0180] The server analyzes voice and facial expression data to determine the user's emotional state. This involves analyzing voice tone and using facial expression recognition algorithms. The input is voice and facial expression data, and the output is an emotional state (e.g., joy, anxiety, anger).

[0181] Step 5:

[0182] The server queries a database for local object sorting rules based on the object classification results and the user's location information, and selects the optimal processing method. The input is the object classification results and location information, and the output is the specific processing method based on the sorting rules.

[0183] Step 6:

[0184] The server generates notification messages that reflect how the object is processed and the user's emotional state. For example, if the user is in a "dissatisfied" state, it will create a message that includes positive language. The input is the sorting method and emotional state, and the output is a customized message sent to the user.

[0185] Step 7:

[0186] The terminal displays sorting instructions and messages received from the server on its screen. Based on the displayed information, the user can perform appropriate sorting actions. The input is the message from the server, and the output is visually displayed guidance information.

[0187] (Application Example 2)

[0188] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as a "server" and the smart device 14 as a "terminal".

[0189] Accurate waste sorting is crucial for environmental protection, but it's not easy to achieve in daily operations. Furthermore, workers' emotional states and motivation can significantly impact work efficiency. Therefore, there is a need for a system that supports proper waste classification while simultaneously providing feedback that takes into account the user's feelings.

[0190] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0191] In this invention, the server includes means for receiving images of waste captured by the user with a video acquisition device, means for analyzing the images to identify the type of waste, means for identifying an appropriate classification method by referring to waste sorting information in the relevant area based on the classification, means for analyzing emotional data to determine the user's emotional state, and means for generating an appropriate response message for the user based on the emotional state. This makes it possible to provide support for properly classifying waste and to provide feedback that takes the user's emotions into consideration.

[0192] A "video acquisition device" is a recording device that allows the user to take images of waste and transmit them to the system.

[0193] A "server" is a central control unit that analyzes and processes received data.

[0194] "Means for identifying the type of waste" refers to a system that analyzes images of the waste that have been photographed and performs a process to identify its type.

[0195] "Waste sorting information" refers to information regarding the correct classification and disposal methods of waste in a specific area.

[0196] "Emotional data" refers to data obtained from the user's voice and facial expressions, and is used to determine the user's emotional state.

[0197] "Means for generating response messages" refers to a function that generates and presents appropriate text based on the user's emotions and the results of waste classification.

[0198] The system for realizing this invention includes a video acquisition device for the user to capture images of waste and transmit them to a server. When the user captures images of waste using the video acquisition device, the data is transmitted to the server via a terminal. The server analyzes the image data using a machine learning algorithm to identify the type of waste. This analysis is typically performed using software such as TensorFlow.

[0199] Furthermore, the server utilizes emotion recognition software such as Azure Cognitive Services to analyze the user's voice and facial expression data. This analysis determines the user's current emotional state, and based on the results, individually tailored response messages are generated.

[0200] The server also refers to the database and uses location information to determine how to classify the waste. At this time, it uses waste sorting information for the relevant area to inform the user of the appropriate classification method.

[0201] For example, if a user photographs a plastic bottle with a video acquisition device, the system identifies it as plastic waste and instructs the user to classify it as a recyclable resource according to local rules. At the same time, if the system analyzes that the user is feeling fatigued, it provides a motivational message such as, "You've made good progress so far today, let's keep going a little longer!"

[0202] The following are examples of prompts used when utilizing generative AI models.

[0203] By asking a simple question such as, "Where should I dispose of this waste?", you can get appropriate instructions from the system.

[0204] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0205] Step 1:

[0206] The user captures images of the waste using a video acquisition device. The input is the actual appearance of the waste, and the output is digital image data. This image data is transmitted to the server via the terminal.

[0207] Step 2:

[0208] The server analyzes the image data it receives. The input is image data of waste sent by the user. The server uses TensorFlow to identify the type of waste based on a machine learning model. The output is the waste classification information.

[0209] Step 3:

[0210] The server collects user voice and facial expression data and performs emotion recognition. The input is facial expression data obtained from voice and video. Azure Cognitive Services is used to determine the user's emotional state (e.g., joy, surprise, dissatisfaction). This result is the output of the emotion analysis.

[0211] Step 4:

[0212] The server obtains the user's location information and refers to a waste sorting information database. The inputs are classified waste information and the user's location information. Based on this, it searches for waste sorting information in the relevant area and identifies the correct processing method. As output, specific sorting instructions are generated.

[0213] Step 5:

[0214] The server generates appropriate response messages based on the user's emotional state. Input consists of the results of the emotion analysis and classification instructions. The goal is to generate flexible messages tailored to the emotional state, thereby increasing user motivation. Output is a feedback message presented to the user.

[0215] Step 6:

[0216] The terminal notifies the user of sorting instructions and response messages received from the server. The terminal displays sorting methods and messages on the screen. This enables the user to dispose of waste properly and simultaneously provides feedback to maintain motivation for the task.

[0217] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0218] Data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of data generation model 58 is ChatGPT (registered trademark) (Internet search).<URL: https: / / openai.com / blog / chatgpt> ), Gemini (registered trademark) (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0219] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart device 14.

[0220] [Second Embodiment]

[0221] Figure 3 shows an example of the configuration of the data processing system 210 according to the second embodiment.

[0222] As shown in Figure 3, the data processing system 210 includes a data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.

[0223] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0224] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication interface 44. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, and camera 42 are also connected to the bus 52.

[0225] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0226] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[0227] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0228] Figure 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Figure 4, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0229] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0230] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0231] In the smart glasses 214, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0232] Next, the identification processing performed by the identification processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".

[0233] The system of the present invention assists users in sorting the waste they generate on a daily basis and consists of a mobile terminal, a server, and related software. The program processing of the system is described below in natural language.

[0234] The main components of this system include image capture, image analysis, information retrieval, and classification instruction functions.

[0235] First, the user uses their mobile device to take a picture of the waste using the deployed camera function. This image is then sent to the server via the application.

[0236] The server applies an image analysis algorithm to the received image to identify the type of waste. This analysis uses machine learning techniques to detect the characteristics of the waste with high accuracy based on past data.

[0237] Furthermore, the server uses the user's location information to refer to a database of relevant regions and retrieve sorting rules that are suitable for the characteristics of the waste. Subsequently, the retrieved rules are compared with the analysis results to identify the correct waste sorting method.

[0238] The identified sorting method is transmitted from the server to the terminal. The terminal receives this information and displays it to the user in an easy-to-understand format. This display includes detailed instructions such as the waste sorting category, collection date, and the appropriate type of garbage bag.

[0239] As a concrete example, if a user takes a picture of a plastic bottle from their home, the server will recognize it as plastic waste and provide instructions to separate it as a "recyclable resource" according to the rules of the relevant area.

[0240] Thus, this invention helps users easily and accurately sort waste, contributing to the efficiency of local waste management.

[0241] The following describes the processing flow.

[0242] Step 1:

[0243] Users use their mobile devices to take photos of waste using the application's camera function. After taking the photo, the image is uploaded to the server by tapping the send button within the app.

[0244] Step 2:

[0245] The server receives image data of the waste sent by the user. The received images are temporarily stored in storage and prepared to be sent to the analysis process.

[0246] Step 3:

[0247] The server uses an image analysis algorithm to analyze the image data. Here, a machine learning model extracts features of the waste and identifies its type. For example, it determines categories such as plastic bottles and cans based on shape, color, label information, etc.

[0248] Step 4:

[0249] The server uses the user's location information to access the waste sorting database for that area. It searches for sorting rules corresponding to the identified waste type and obtains the correct sorting method.

[0250] Step 5:

[0251] The server uses the acquired sorting information to generate an easy-to-understand instruction message for the user. This message includes the sorting category and handling precautions.

[0252] Step 6:

[0253] The server sends the generated instruction message to the terminal. The terminal displays the received message on the app screen, showing the user the proper way to sort waste.

[0254] Step 7:

[0255] Users follow instructions from the terminal to sort waste into designated categories and dispose of it appropriately in accordance with local regulations.

[0256] (Example 1)

[0257] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".

[0258] Conventional waste sorting methods require users to manually understand and apply classification criteria, which is time-consuming and cumbersome. Furthermore, adapting to varying waste sorting rules across different regions is difficult, increasing the risk of incorrect disposal. This invention aims to solve these problems and enable users to sort waste quickly and accurately.

[0259] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0260] In this invention, the server includes means for receiving image data generated by a user in a portable information terminal equipped with a camera function, means for analyzing the image data using machine learning technology to recognize the characteristics of the waste, and means for referring to location-dependent regional waste disposal procedure data based on the recognition to identify an appropriate classification method. This enables the user to efficiently sort waste in accordance with local waste sorting rules.

[0261] An "information processing device" is an electronic device programmed to process data and perform specific tasks.

[0262] A "personal information terminal" is an electronic device that a user can carry with them and that has data input and display functions.

[0263] "Image data" refers to visual information generated using a camera function, expressed in digital format.

[0264] "Machine learning technology" is a collection of algorithms and methodologies for learning from past data and making predictions and classifications on new data.

[0265] "Waste characteristics" refer to classification criteria based on the material, shape, and intended use of an object.

[0266] "Location information" is digital data used to identify a specific location, and is acquired using technologies such as GPS.

[0267] "Regional waste disposal procedure data" refers to information regarding waste sorting methods and regulations established in each region.

[0268] A "classification method" is a set of procedures and criteria for dividing waste into specific categories based on its characteristics.

[0269] To implement the present invention, the system is implemented with a configuration including an information processing device, a personal information terminal, and related software. Specifically, the personal information terminal is equipped with a camera function, which the user uses to generate image data of objects that are routinely discarded. This image data is then transmitted to a server through a dedicated application.

[0270] The server analyzes the received image data. Machine learning techniques are used for the analysis, utilizing popular machine learning frameworks such as TensorFlow and PyTorch. This technology allows the server to recognize the characteristics of the waste with high accuracy and identify its material and shape.

[0271] Furthermore, the server uses location information obtained from the user's mobile device to refer to waste disposal procedure data for each region. Location information is typically obtained using GPS technology, helping the server identify sorting rules relevant to the user's current location. Based on this information, the server identifies the appropriate sorting method and notifies the user.

[0272] Notifications are sent in real time, and the device displays the information clearly and visually to the user. The display includes instructions such as the waste classification category, collection date, and the appropriate type of garbage bag.

[0273] For example, if a user takes a picture of a plastic bottle from their home, the system will recognize it as plastic waste and provide instructions to process it as a "recyclable resource" as defined by the local area. This makes it easy for users to dispose of their waste according to local regulations.

[0274] An example of a prompt to a generating AI model is, "Based on the image data taken by the user, output instructions regarding the appropriate sorting category and specific collection date." This prompt allows the AI ​​model to provide accurate and practical sorting instructions.

[0275] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0276] Step 1:

[0277] The user uses the camera function of their mobile device to take pictures of the waste. The captured image data is sent to the server via the application. The input is the image data captured by the camera, and the output is the digital image data received by the server.

[0278] Step 2:

[0279] The server analyzes the received image data using a machine learning algorithm. Specifically, a model built using TensorFlow or PyTorch classifies objects in the image based on their characteristics as waste. The input to this step is the image data received by the server, and the output is information about the characteristics of the waste based on that image.

[0280] Step 3:

[0281] The server uses location information obtained from the user's terminal to refer to a regional waste disposal procedure database. The input is location information obtained using GPS, and the output is a list of region-specific sorting methods based on that location information.

[0282] Step 4:

[0283] The server compares the image analysis results with the classification rules of the obtained region to identify an appropriate waste classification method. Through this comparison operation, it is determined to which classification category the waste belongs. The input is the characteristic information of the waste and a list of regional classification methods, and the output is the optimal classification method.

[0284] Step 5:

[0285] The server transmits the identified classification method to the user's mobile information terminal. Communication is performed in real time, and the user can receive the result immediately. The input is the classification method identified by the server, and the output is the classification instruction information displayed on the user's terminal.

[0286] Step 6:

[0287] The terminal visually and clearly displays the received classification information to the user. The display screen shows the waste classification category, collection date, appropriate types of garbage bags, etc., and the user can process the waste according to the content. The input is the notified classification instruction information, and the output is the detailed instruction content displayed on the user's screen.

[0288] (Application Example 1)

[0289] Next, Application Example 1 will be described. In the following description, the data processing device 12 is referred to as the "server", and the smart glasses 214 are referred to as the "terminal".

[0290] Waste classification is an important issue in environmental protection and effective utilization of resources, but conventional methods have limitations in classification accuracy and efficiency. Especially in industries and manufacturing, a large amount of waste is generated, and its proper treatment is necessary. The classification work on-site often relies on manual labor, and the consumption of human resources and the increase in environmental burden due to mistakes have become problems. Therefore, the realization of an automated high-precision and high-efficiency classification system is required.

[0291] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0292] In this invention, the server includes means for receiving visual information of waste captured by a user with a device having a camera function, means for classifying the waste by analyzing the visual information, means for identifying the correct disposal method by referring to the relevant local waste disposal regulations, and means for physically separating the waste using a powered machine. This enables automatic and accurate sorting of waste.

[0293] "User" refers to an individual or organization that utilizes the waste sorting system.

[0294] "Photography function" refers to the capability of a device to capture visual information about waste.

[0295] "Device" refers to hardware components, including cameras and sensors.

[0296] "Visual information" refers to image data acquired using a camera or other imaging function.

[0297] "Analysis" refers to the process of identifying the type and characteristics of waste from visual information using machine learning algorithms.

[0298] "Classification" refers to dividing waste into predetermined categories based on the results of analysis.

[0299] "Local waste disposal regulations" refer to laws and guidelines concerning the collection, treatment, and recycling of waste that apply in a specific area.

[0300] "Disposal method" refers to the procedures for properly processing or recycling waste.

[0301] A "powered machine" refers to a device that supplies energy for performing mechanical work.

[0302] "Physically separate" refers to the operation of physically separating waste using mechanical devices rather than manually.

[0303] In an embodiment of the present invention, a waste separation system is used. The system is composed of a device with a photographing function (hereinafter referred to as a terminal), a server that performs data processing, and related software. The user uses the camera function of the terminal to obtain visual information of the waste to be separated. The acquired image data is transmitted to the server through the network.

[0304] First, the server uses a machine learning algorithm for the received visual information to identify the type of waste. For this analysis, the characteristic information of waste accumulated in the past is utilized to enable highly accurate classification. Based on the identified information, the server refers to the waste treatment rule database of the corresponding region considering the user's location information and specifies an appropriate disposal method.

[0305] The specified disposal method also includes information on means for physically separating waste using a power machine, and the generated instruction is transmitted to the terminal. The terminal displays the visual information, disposal method, and precautions to the user in an easy-to-understand manner.

[0306] As a specific example, when a user photographs plastic parts used in a factory, the server identifies it as a recyclable material and generates an instruction to transport it to an appropriate recycling box.

[0307] Examples of prompt sentences to input into the generated AI model:

[0308] "Please teach me the method of generating a separation instruction from the results of image analysis when a robot operating in a factory photographs waste."

[0309] The flow of the specific process in Application Example 1 will be described with reference to FIG. 12.

[0310] Step 1:

[0311] The user takes a picture of the waste using their device. The input is capturing the actual waste with the camera. The output is digital image data generated as pixel information. Specifically, the user presses the capture button, which activates the camera and saves a photograph of the object.

[0312] Step 2:

[0313] The device transmits the captured image data to the server via the network. The input is the digital image obtained in step 1. The output is the image data received by the server. Specifically, the device establishes a network connection and uploads the image data to the server using the HTTP protocol.

[0314] Step 3:

[0315] The server analyzes the received image data to identify the type of waste. The input is image data sent from the terminal. The output is the identified waste type and characteristic information. Specifically, the server uses computer vision technology and machine learning algorithms to extract features from the image and classify them by comparing them with a database.

[0316] Step 4:

[0317] The server obtains the user's location information and refers to a database of waste disposal regulations for the relevant region. Inputs are the type of waste and the user's location. Outputs are instructions for the appropriate disposal method. Specifically, the server uses GPS data to determine the location and retrieves relevant information from the database based on that location.

[0318] Step 5:

[0319] The server determines the means of physically separating the waste using a powered machine and sends instructions to the terminal. The input is the instruction on the disposal method. The output is a sorting method guide displayed on the terminal. Specifically, the server converts the instruction content into text format and sends it to the terminal.

[0320] Step 6:

[0321] The terminal visually presents the received instructions to the user. The input is the sorting method instructions sent from the server. The output is the instruction screen that the user can view. Specifically, the terminal displays the instruction content on its screen and provides information using icons and text so that the user can easily understand it.

[0322] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0323] This invention is a system that assists users in processing waste they sort on a daily basis and incorporates an emotion engine to provide appropriate feedback according to the user's emotional state. The system's program processing is described below in natural language.

[0324] This system consists of functions related to image capture, emotion recognition, image analysis, information retrieval, and classification instructions.

[0325] First, the user uses their mobile device to take a picture of the waste using the camera. This image is sent to the server via the application. In addition, the device simultaneously collects data on the user's voice and facial expressions, which are then transmitted to the server for analysis of the user's emotional state by an emotion engine.

[0326] The server analyzes received images of waste using an image analysis algorithm to identify the type of waste. This analysis utilizes a machine learning model, achieving highly accurate classification based on accumulated data. In addition, an emotion engine analyzes the user's voice tone and facial expression data to determine their current emotional state (e.g., joy, surprise, dissatisfaction).

[0327] Next, the server uses the user's location information to refer to the relevant regional waste sorting database. Here, it searches for the appropriate sorting rules corresponding to the identified waste type and determines the correct disposal method for the waste.

[0328] Based on the identified classification method and the user's emotional state, the server generates encouraging messages and situation-specific warning messages. For example, if the user is dissatisfied, a positive message such as "You're doing well, let's move on!" might be generated.

[0329] This information is sent from the server to the terminal, which then displays sorting instructions and messages on the screen to provide the user with appropriate guidance and motivation.

[0330] As a concrete example, when a user photographs a plastic bottle that should be recycled, the server recognizes it as plastic waste and instructs it to separate it as "recyclable material" according to the relevant local rules. At the same time, if the emotion engine detects dissatisfaction from the user's facial expression, it also displays a message to boost their motivation.

[0331] Thus, the system of the present invention not only supports users in properly sorting waste, but also provides emotionally responsive feedback to realize a more comfortable and effective user experience.

[0332] The following describes the processing flow.

[0333] Step 1:

[0334] The user launches the application on their mobile device and takes a picture of the waste using the camera function. If audio or facial expressions are available, the app will also record this data simultaneously.

[0335] Step 2:

[0336] The device transmits captured image data, recorded audio, and facial expression data to a server via the app. This data also includes the user's location information.

[0337] Step 3:

[0338] The server receives the image data and first performs image analysis using a machine learning algorithm. This analysis extracts the characteristics of the waste and identifies its type.

[0339] Step 4:

[0340] The server analyzes voice and facial expression data using an emotion engine to determine the user's emotional state. For example, it analyzes voice tone and detects changes in facial expressions to identify whether the user is happy or dissatisfied.

[0341] Step 5:

[0342] The server accesses a waste sorting database for the relevant region based on the user's location information. Using the analysis results, it retrieves the appropriate sorting method for the waste from the database.

[0343] Step 6:

[0344] The server generates notification messages for the user based on the acquired sorting method. Furthermore, it adds positive messages and words of encouragement depending on the user's emotional state.

[0345] Step 7:

[0346] The terminal receives a message sent from the server and displays it on the screen. This message includes instructions on how to properly sort waste and emotionally responsive feedback.

[0347] Step 8:

[0348] Users review the displayed information and follow the instructions to properly sort their waste. The positive messages displayed are also expected to boost user motivation.

[0349] (Example 2)

[0350] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".

[0351] In today's environmental challenges, the sorting of waste that users perform daily is important, but accurate classification and adherence to local sorting rules are difficult, sometimes resulting in improper disposal. Furthermore, insufficient feedback on user behavior makes it difficult to maintain motivation for sorting.

[0352] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0353] In this invention, the server includes means for receiving images of objects captured by the user with a communication terminal, means for analyzing the images and classifying the objects, and means for receiving audio data and facial expression data and determining the user's emotional state. This makes it possible to support accurate object classification while providing appropriate feedback according to the user's emotions.

[0354] A "communication terminal" is an electronic device used by a user to send and receive information, and is equipped with camera and voice input functions.

[0355] An "image of an object" is digital data that shows the visible shape of an object, captured using the camera of a communication terminal.

[0356] "Analyzing an image" is the process of identifying the features of an object from digital data and classifying its attributes.

[0357] "Object classification" is the act of dividing objects into specific categories based on their analyzed characteristics.

[0358] "Audio data" refers to information that represents acoustic signals acquired through the microphone of a communication terminal in digital format.

[0359] "Facial expression data" refers to digital data containing information about facial expressions captured by the camera of a communication device.

[0360] "User's emotional state" refers to information indicating the psychological state obtained by analyzing voice data and facial expression data, and includes examples such as "joy," "dissatisfaction," and "surprise."

[0361] "Generating adaptive notification messages that respond to emotions" means creating feedback messages that are appropriate for individual situations, taking into account the user's emotional state.

[0362] This invention is a system that links a communication terminal and a server to assist users in sorting objects on a daily basis. Specifically, the communication terminal uses a camera and microphone to capture image data by photographing objects, and in addition, it collects the user's voice and facial expressions. This data is transmitted from the terminal to the server.

[0363] The server uses generative AI models, such as Convolutional Neural Networks (CNNs), to analyze the features of objects based on the received image data. The image analysis algorithm classifies objects into specific categories. The server also utilizes facial recognition and natural language processing (NLP) technologies to analyze audio and facial data and determine the user's emotional state (joy, surprise, dissatisfaction, etc.).

[0364] Based on the analyzed data, the server references the user's location information and retrieves region-specific object classification rules from a database. Following these rules, it identifies the optimal method for processing the object and generates feedback tailored to the identified emotional state. For example, if the user expresses dissatisfaction, it might create an encouraging message such as, "Very well done, keep up the good work next time!"

[0365] Ultimately, this information is transmitted to the terminal, and the results are displayed to the user on the screen. This allows the user to efficiently perform the appropriate processing of the object while maintaining motivation.

[0366] For example, if a user photographs a reusable plastic bottle, the server will recognize the object as "plastic" and instruct the user to separate it as "recyclable waste" according to local rules. At the same time, if sentiment analysis determines that the user is frustrated, it will display a positive message such as, "Great job, please recycle next time!"

[0367] An example of a prompt to input into the generating AI model is, "Please tell me the steps to determine the type of object, apply the sorting rules for that location, and generate the optimal action."

[0368] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0369] Step 1:

[0370] The user takes an image of an object using the camera on the communication terminal. The input image data must be high resolution so that the entire object is clearly visible. As output, the captured image is saved to the memory of the communication terminal.

[0371] Step 2:

[0372] The device prepares to send the captured image data to the server. It also collects user voice data using the microphone and acquires facial expression data using the camera. The inputs are image, audio, and facial expression data, and these are sent to the server as output.

[0373] Step 3:

[0374] The server processes the received image data through an image analysis system to extract object features. This analysis uses a Convolutional Neural Network (CNN) to analyze features such as object shape and color at the pixel level. The input is image data, and the output is the object classification result.

[0375] Step 4:

[0376] The server analyzes voice and facial expression data to determine the user's emotional state. This involves analyzing voice tone and using facial expression recognition algorithms. The input is voice and facial expression data, and the output is an emotional state (e.g., joy, anxiety, anger).

[0377] Step 5:

[0378] The server queries a database for local object sorting rules based on the object classification results and the user's location information, and selects the optimal processing method. The input is the object classification results and location information, and the output is the specific processing method based on the sorting rules.

[0379] Step 6:

[0380] The server generates notification messages that reflect how the object is processed and the user's emotional state. For example, if the user is in a "dissatisfied" state, it will create a message that includes positive language. The input is the sorting method and emotional state, and the output is a customized message sent to the user.

[0381] Step 7:

[0382] The terminal displays sorting instructions and messages received from the server on its screen. Based on the displayed information, the user can perform appropriate sorting actions. The input is the message from the server, and the output is visually displayed guidance information.

[0383] (Application Example 2)

[0384] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."

[0385] Accurate waste sorting is crucial for environmental protection, but it's not easy to achieve in daily operations. Furthermore, workers' emotional states and motivation can significantly impact work efficiency. Therefore, there is a need for a system that supports proper waste classification while simultaneously providing feedback that takes into account the user's feelings.

[0386] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0387] In this invention, the server includes means for receiving images of waste captured by the user with a video acquisition device, means for analyzing the images to identify the type of waste, means for identifying an appropriate classification method by referring to waste sorting information in the relevant area based on the classification, means for analyzing emotional data to determine the user's emotional state, and means for generating an appropriate response message for the user based on the emotional state. This makes it possible to provide support for properly classifying waste and to provide feedback that takes the user's emotions into consideration.

[0388] A "video acquisition device" is a recording device that allows the user to take images of waste and transmit them to the system.

[0389] A "server" is a central control unit that analyzes and processes received data.

[0390] "Means for identifying the type of waste" refers to a system that analyzes images of the waste that have been photographed and performs a process to identify its type.

[0391] "Waste sorting information" refers to information regarding the correct classification and disposal methods of waste in a specific area.

[0392] "Emotional data" refers to data obtained from the user's voice and facial expressions, and is used to determine the user's emotional state.

[0393] "Means for generating response messages" refers to a function that generates and presents appropriate text based on the user's emotions and the results of waste classification.

[0394] The system for realizing this invention includes a video acquisition device for the user to capture images of waste and transmit them to a server. When the user captures images of waste using the video acquisition device, the data is transmitted to the server via a terminal. The server analyzes the image data using a machine learning algorithm to identify the type of waste. This analysis is typically performed using software such as TensorFlow.

[0395] Furthermore, the server utilizes emotion recognition software such as Azure Cognitive Services to analyze the user's voice and facial expression data. This analysis determines the user's current emotional state, and based on the results, individually tailored response messages are generated.

[0396] The server also refers to the database and uses location information to determine how to classify the waste. At this time, it uses waste sorting information for the relevant area to inform the user of the appropriate classification method.

[0397] For example, if a user photographs a plastic bottle with a video acquisition device, the system identifies it as plastic waste and instructs the user to classify it as a recyclable resource according to local rules. At the same time, if the system analyzes that the user is feeling fatigued, it provides a motivational message such as, "You've made good progress so far today, let's keep going a little longer!"

[0398] The following are examples of prompts used when utilizing generative AI models.

[0399] By asking a simple question such as, "Where should I dispose of this waste?", you can get appropriate instructions from the system.

[0400] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0401] Step 1:

[0402] The user captures images of the waste using a video acquisition device. The input is the actual appearance of the waste, and the output is digital image data. This image data is transmitted to the server via the terminal.

[0403] Step 2:

[0404] The server analyzes the image data it receives. The input is image data of waste sent by the user. The server uses TensorFlow to identify the type of waste based on a machine learning model. The output is the waste classification information.

[0405] Step 3:

[0406] The server collects user voice and facial expression data and performs emotion recognition. The input is facial expression data obtained from voice and video. Azure Cognitive Services is used to determine the user's emotional state (e.g., joy, surprise, dissatisfaction). This result is the output of the emotion analysis.

[0407] Step 4:

[0408] The server obtains the user's location information and refers to a waste sorting information database. The inputs are classified waste information and the user's location information. Based on this, it searches for waste sorting information in the relevant area and identifies the correct processing method. As output, specific sorting instructions are generated.

[0409] Step 5:

[0410] The server generates appropriate response messages based on the user's emotional state. Input consists of the results of the emotion analysis and classification instructions. The goal is to generate flexible messages tailored to the emotional state, thereby increasing user motivation. Output is a feedback message presented to the user.

[0411] Step 6:

[0412] The terminal notifies the user of sorting instructions and response messages received from the server. The terminal displays sorting methods and messages on the screen. This enables the user to dispose of waste properly and simultaneously provides feedback to maintain motivation for the task.

[0413] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0414] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet Search).<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0415] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart glasses 214.

[0416] [Third Embodiment]

[0417] Figure 5 shows an example of the configuration of the data processing system 310 according to the third embodiment.

[0418] As shown in Figure 5, the data processing system 310 includes a data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.

[0419] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0420] The headset terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a display 343. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and display 343 are also connected to the bus 52.

[0421] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0422] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[0423] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0424] Figure 6 shows an example of the main functions of the data processing device 12 and the headset terminal 314. As shown in Figure 6, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0425] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0426] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0427] In the headset terminal 314, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0428] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the headset terminal 314 will be referred to as the "terminal".

[0429] The system of the present invention assists users in sorting the waste they generate on a daily basis and consists of a mobile terminal, a server, and related software. The program processing of the system is described below in natural language.

[0430] The main components of this system include image capture, image analysis, information retrieval, and classification instruction functions.

[0431] First, the user uses their mobile device to take a picture of the waste using the deployed camera function. This image is then sent to the server via the application.

[0432] The server applies an image analysis algorithm to the received image to identify the type of waste. This analysis uses machine learning techniques to detect the characteristics of the waste with high accuracy based on past data.

[0433] Furthermore, the server uses the user's location information to refer to a database of relevant regions and retrieve sorting rules that are suitable for the characteristics of the waste. Subsequently, the retrieved rules are compared with the analysis results to identify the correct waste sorting method.

[0434] The identified sorting method is transmitted from the server to the terminal. The terminal receives this information and displays it to the user in an easy-to-understand format. This display includes detailed instructions such as the waste sorting category, collection date, and the appropriate type of garbage bag.

[0435] As a concrete example, if a user takes a picture of a plastic bottle from their home, the server will recognize it as plastic waste and provide instructions to separate it as a "recyclable resource" according to the rules of the relevant area.

[0436] Thus, this invention helps users easily and accurately sort waste, contributing to the efficiency of local waste management.

[0437] The following describes the processing flow.

[0438] Step 1:

[0439] Users use their mobile devices to take photos of waste using the application's camera function. After taking the photo, the image is uploaded to the server by tapping the send button within the app.

[0440] Step 2:

[0441] The server receives image data of the waste sent by the user. The received images are temporarily stored in storage and prepared to be sent to the analysis process.

[0442] Step 3:

[0443] The server uses an image analysis algorithm to analyze the image data. Here, a machine learning model extracts features of the waste and identifies its type. For example, it determines categories such as plastic bottles and cans based on shape, color, label information, etc.

[0444] Step 4:

[0445] The server uses the user's location information to access the waste sorting database for that area. It searches for sorting rules corresponding to the identified waste type and obtains the correct sorting method.

[0446] Step 5:

[0447] The server uses the acquired sorting information to generate an easy-to-understand instruction message for the user. This message includes the sorting category and handling precautions.

[0448] Step 6:

[0449] The server sends the generated instruction message to the terminal. The terminal displays the received message on the app screen, showing the user the proper way to sort waste.

[0450] Step 7:

[0451] Users follow instructions from the terminal to sort waste into designated categories and dispose of it appropriately in accordance with local regulations.

[0452] (Example 1)

[0453] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0454] Conventional waste sorting methods require users to manually understand and apply classification criteria, which is time-consuming and cumbersome. Furthermore, adapting to varying waste sorting rules across different regions is difficult, increasing the risk of incorrect disposal. This invention aims to solve these problems and enable users to sort waste quickly and accurately.

[0455] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0456] In this invention, the server includes means for receiving image data generated by a user in a portable information terminal equipped with a camera function, means for analyzing the image data using machine learning technology to recognize the characteristics of the waste, and means for referring to location-dependent regional waste disposal procedure data based on the recognition to identify an appropriate classification method. This enables the user to efficiently sort waste in accordance with local waste sorting rules.

[0457] An "information processing device" is an electronic device programmed to process data and perform specific tasks.

[0458] A "personal information terminal" is an electronic device that a user can carry with them and that has data input and display functions.

[0459] "Image data" refers to visual information generated using a camera function, expressed in digital format.

[0460] "Machine learning technology" is a collection of algorithms and methodologies for learning from past data and making predictions and classifications on new data.

[0461] "Waste characteristics" refer to classification criteria based on the material, shape, and intended use of an object.

[0462] "Location information" is digital data used to identify a specific location, and is acquired using technologies such as GPS.

[0463] "Regional waste disposal procedure data" refers to information regarding waste sorting methods and regulations established in each region.

[0464] A "classification method" is a set of procedures and criteria for dividing waste into specific categories based on its characteristics.

[0465] To implement the present invention, the system is implemented with a configuration including an information processing device, a personal information terminal, and related software. Specifically, the personal information terminal is equipped with a camera function, which the user uses to generate image data of objects that are routinely discarded. This image data is then transmitted to a server through a dedicated application.

[0466] The server analyzes the received image data. Machine learning techniques are used for the analysis, utilizing popular machine learning frameworks such as TensorFlow and PyTorch. This technology allows the server to recognize the characteristics of the waste with high accuracy and identify its material and shape.

[0467] Furthermore, the server uses location information obtained from the user's mobile device to refer to waste disposal procedure data for each region. Location information is typically obtained using GPS technology, helping the server identify sorting rules relevant to the user's current location. Based on this information, the server identifies the appropriate sorting method and notifies the user.

[0468] Notifications are sent in real time, and the device displays the information clearly and visually to the user. The display includes instructions such as the waste classification category, collection date, and the appropriate type of garbage bag.

[0469] For example, if a user takes a picture of a plastic bottle from their home, the system will recognize it as plastic waste and provide instructions to process it as a "recyclable resource" as defined by the local area. This makes it easy for users to dispose of their waste according to local regulations.

[0470] An example of a prompt to a generating AI model is, "Based on the image data taken by the user, output instructions regarding the appropriate sorting category and specific collection date." This prompt allows the AI ​​model to provide accurate and practical sorting instructions.

[0471] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0472] Step 1:

[0473] The user uses the camera function of their mobile device to take pictures of the waste. The captured image data is sent to the server via the application. The input is the image data captured by the camera, and the output is the digital image data received by the server.

[0474] Step 2:

[0475] The server analyzes the received image data using a machine learning algorithm. Specifically, a model built using TensorFlow or PyTorch classifies objects in the image based on their characteristics as waste. The input to this step is the image data received by the server, and the output is information about the characteristics of the waste based on that image.

[0476] Step 3:

[0477] The server uses location information obtained from the user's terminal to refer to a regional waste disposal procedure database. The input is location information obtained using GPS, and the output is a list of region-specific sorting methods based on that location information.

[0478] Step 4:

[0479] The server compares the image analysis results with the acquired local sorting rules to identify the appropriate waste classification method. This comparison process determines which sorting category the waste belongs to. The input is the waste's characteristics and a list of local sorting methods, and the output is the optimal sorting method.

[0480] Step 5:

[0481] The server transmits the identified sorting method to the user's mobile device. Communication takes place in real time, and the user can receive the results immediately. The input is the sorting method identified by the server, and the output is the sorting instruction information displayed on the user's device.

[0482] Step 6:

[0483] The terminal visually displays the received sorting information to the user. The display screen shows the waste classification category, collection date, and appropriate garbage bag type, allowing the user to process the waste accordingly. The input is the notified sorting instruction information, and the output is the detailed instruction content displayed on the user's screen.

[0484] (Application Example 1)

[0485] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0486] Waste sorting is a crucial issue for environmental protection and efficient resource utilization, but conventional methods have limitations in terms of accuracy and efficiency. In particular, large amounts of waste are generated in industry and manufacturing, requiring proper disposal. On-site sorting often relies on manual labor, leading to problems such as the depletion of human resources and increased environmental burden due to errors. Therefore, there is a need for automated, high-precision, and highly efficient sorting systems.

[0487] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0488] In this invention, the server includes means for receiving visual information of waste captured by a user with a device having a camera function, means for classifying the waste by analyzing the visual information, means for identifying the correct disposal method by referring to the relevant local waste disposal regulations, and means for physically separating the waste using a powered machine. This enables automatic and accurate sorting of waste.

[0489] "User" refers to an individual or organization that utilizes the waste sorting system.

[0490] "Photography function" refers to the capability of a device to capture visual information about waste.

[0491] "Device" refers to hardware components, including cameras and sensors.

[0492] "Visual information" refers to image data acquired using a camera or other imaging function.

[0493] "Analysis" refers to the process of identifying the type and characteristics of waste from visual information using machine learning algorithms.

[0494] "Classification" refers to dividing waste into predetermined categories based on the results of analysis.

[0495] "Local waste disposal regulations" refer to laws and guidelines concerning the collection, treatment, and recycling of waste that apply in a specific area.

[0496] "Disposal method" refers to the procedures for properly processing or recycling waste.

[0497] A "powered machine" refers to a device that supplies energy for performing mechanical work.

[0498] "Physically separating" refers to the process of physically separating waste using mechanical equipment rather than manual methods.

[0499] In embodiments of the present invention, a waste sorting system is used. The system consists of a device with a camera function (hereinafter referred to as a terminal), a server that processes data, and related software. The user uses the camera function of the terminal to acquire visual information of the waste to be sorted. The acquired image data is transmitted to the server via the network.

[0500] The server first uses machine learning algorithms on the received visual information to identify the type of waste. This analysis utilizes previously accumulated waste characteristic information to enable highly accurate classification. Based on the identified information, the server considers the user's location and refers to the waste disposal regulations database for the relevant area to determine the appropriate disposal method.

[0501] The identified disposal method includes information on means of physically separating the waste using powered machinery, and the generated instructions are sent to the terminal. The terminal displays visual information, disposal methods, and precautions to the user in an easy-to-understand format.

[0502] For example, if a user takes a picture of a plastic part used in a factory, the server will identify it as recyclable material and generate instructions on how to transport it to the appropriate recycling bin.

[0503] Examples of prompts to input into a generative AI model:

[0504] "Please explain how a robot operating within a factory can photograph waste materials, perform image analysis, and generate sorting instructions based on the results."

[0505] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0506] Step 1:

[0507] The user takes a picture of the waste using their device. The input is capturing the actual waste with the camera. The output is digital image data generated as pixel information. Specifically, the user presses the capture button, which activates the camera and saves a photograph of the object.

[0508] Step 2:

[0509] The device transmits the captured image data to the server via the network. The input is the digital image obtained in step 1. The output is the image data received by the server. Specifically, the device establishes a network connection and uploads the image data to the server using the HTTP protocol.

[0510] Step 3:

[0511] The server analyzes the received image data to identify the type of waste. The input is image data sent from the terminal. The output is the identified waste type and characteristic information. Specifically, the server uses computer vision technology and machine learning algorithms to extract features from the image and classify them by comparing them with a database.

[0512] Step 4:

[0513] The server obtains the user's location information and refers to a database of waste disposal regulations for the relevant region. Inputs are the type of waste and the user's location. Outputs are instructions for the appropriate disposal method. Specifically, the server uses GPS data to determine the location and retrieves relevant information from the database based on that location.

[0514] Step 5:

[0515] The server determines the means of physically separating the waste using a powered machine and sends instructions to the terminal. The input is the instruction on the disposal method. The output is a sorting method guide displayed on the terminal. Specifically, the server converts the instruction content into text format and sends it to the terminal.

[0516] Step 6:

[0517] The terminal visually presents the received instructions to the user. The input is the sorting method instructions sent from the server. The output is the instruction screen that the user can view. Specifically, the terminal displays the instruction content on its screen and provides information using icons and text so that the user can easily understand it.

[0518] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0519] This invention is a system that assists users in processing waste they sort on a daily basis and incorporates an emotion engine to provide appropriate feedback according to the user's emotional state. The system's program processing is described below in natural language.

[0520] This system consists of functions related to image capture, emotion recognition, image analysis, information retrieval, and classification instructions.

[0521] First, the user uses their mobile device to take a picture of the waste using the camera. This image is sent to the server via the application. In addition, the device simultaneously collects data on the user's voice and facial expressions, which are then transmitted to the server for analysis of the user's emotional state by an emotion engine.

[0522] The server analyzes received images of waste using an image analysis algorithm to identify the type of waste. This analysis utilizes a machine learning model, achieving highly accurate classification based on accumulated data. In addition, an emotion engine analyzes the user's voice tone and facial expression data to determine their current emotional state (e.g., joy, surprise, dissatisfaction).

[0523] Next, the server uses the user's location information to refer to the relevant regional waste sorting database. Here, it searches for the appropriate sorting rules corresponding to the identified waste type and determines the correct disposal method for the waste.

[0524] Based on the identified classification method and the user's emotional state, the server generates encouraging messages and situation-specific warning messages. For example, if the user is dissatisfied, a positive message such as "You're doing well, let's move on!" might be generated.

[0525] This information is sent from the server to the terminal, which then displays sorting instructions and messages on the screen to provide the user with appropriate guidance and motivation.

[0526] As a concrete example, when a user photographs a plastic bottle that should be recycled, the server recognizes it as plastic waste and instructs it to separate it as "recyclable material" according to the relevant local rules. At the same time, if the emotion engine detects dissatisfaction from the user's facial expression, it also displays a message to boost their motivation.

[0527] Thus, the system of the present invention not only supports users in properly sorting waste, but also provides emotionally responsive feedback to realize a more comfortable and effective user experience.

[0528] The following describes the processing flow.

[0529] Step 1:

[0530] The user launches the application on their mobile device and takes a picture of the waste using the camera function. If audio or facial expressions are available, the app will also record this data simultaneously.

[0531] Step 2:

[0532] The device transmits captured image data, recorded audio, and facial expression data to a server via the app. This data also includes the user's location information.

[0533] Step 3:

[0534] The server receives the image data and first performs image analysis using a machine learning algorithm. This analysis extracts the characteristics of the waste and identifies its type.

[0535] Step 4:

[0536] The server analyzes voice and facial expression data using an emotion engine to determine the user's emotional state. For example, it analyzes voice tone and detects changes in facial expressions to identify whether the user is happy or dissatisfied.

[0537] Step 5:

[0538] The server accesses a waste sorting database for the relevant region based on the user's location information. Using the analysis results, it retrieves the appropriate sorting method for the waste from the database.

[0539] Step 6:

[0540] The server generates notification messages for the user based on the acquired sorting method. Furthermore, it adds positive messages and words of encouragement depending on the user's emotional state.

[0541] Step 7:

[0542] The terminal receives a message sent from the server and displays it on the screen. This message includes instructions on how to properly sort waste and emotionally responsive feedback.

[0543] Step 8:

[0544] Users review the displayed information and follow the instructions to properly sort their waste. The positive messages displayed are also expected to boost user motivation.

[0545] (Example 2)

[0546] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0547] In today's environmental challenges, the sorting of waste that users perform daily is important, but accurate classification and adherence to local sorting rules are difficult, sometimes resulting in improper disposal. Furthermore, insufficient feedback on user behavior makes it difficult to maintain motivation for sorting.

[0548] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0549] In this invention, the server includes means for receiving images of objects captured by the user with a communication terminal, means for analyzing the images and classifying the objects, and means for receiving audio data and facial expression data and determining the user's emotional state. This makes it possible to support accurate object classification while providing appropriate feedback according to the user's emotions.

[0550] A "communication terminal" is an electronic device used by a user to send and receive information, and is equipped with camera and voice input functions.

[0551] An "image of an object" is digital data that shows the visible shape of an object, captured using the camera of a communication terminal.

[0552] "Analyzing an image" is the process of identifying the features of an object from digital data and classifying its attributes.

[0553] "Object classification" is the act of dividing objects into specific categories based on their analyzed characteristics.

[0554] "Audio data" refers to information that represents acoustic signals acquired through the microphone of a communication terminal in digital format.

[0555] "Facial expression data" refers to digital data containing information about facial expressions captured by the camera of a communication device.

[0556] "User's emotional state" refers to information indicating the psychological state obtained by analyzing voice data and facial expression data, and includes examples such as "joy," "dissatisfaction," and "surprise."

[0557] "Generating adaptive notification messages that respond to emotions" means creating feedback messages that are appropriate for individual situations, taking into account the user's emotional state.

[0558] This invention is a system that links a communication terminal and a server to assist users in sorting objects on a daily basis. Specifically, the communication terminal uses a camera and microphone to capture image data by photographing objects, and in addition, it collects the user's voice and facial expressions. This data is transmitted from the terminal to the server.

[0559] The server uses generative AI models, such as Convolutional Neural Networks (CNNs), to analyze the features of objects based on the received image data. The image analysis algorithm classifies objects into specific categories. The server also utilizes facial recognition and natural language processing (NLP) technologies to analyze audio and facial data and determine the user's emotional state (joy, surprise, dissatisfaction, etc.).

[0560] Based on the analyzed data, the server references the user's location information and retrieves region-specific object classification rules from a database. Following these rules, it identifies the optimal method for processing the object and generates feedback tailored to the identified emotional state. For example, if the user expresses dissatisfaction, it might create an encouraging message such as, "Very well done, keep up the good work next time!"

[0561] Ultimately, this information is transmitted to the terminal, and the results are displayed to the user on the screen. This allows the user to efficiently perform the appropriate processing of the object while maintaining motivation.

[0562] For example, if a user photographs a reusable plastic bottle, the server will recognize the object as "plastic" and instruct the user to separate it as "recyclable waste" according to local rules. At the same time, if sentiment analysis determines that the user is frustrated, it will display a positive message such as, "Great job, please recycle next time!"

[0563] An example of a prompt to input into the generating AI model is, "Please tell me the steps to determine the type of object, apply the sorting rules for that location, and generate the optimal action."

[0564] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0565] Step 1:

[0566] The user takes an image of an object using the camera on the communication terminal. The input image data must be high resolution so that the entire object is clearly visible. As output, the captured image is saved to the memory of the communication terminal.

[0567] Step 2:

[0568] The device prepares to send the captured image data to the server. It also collects user voice data using the microphone and acquires facial expression data using the camera. The inputs are image, audio, and facial expression data, and these are sent to the server as output.

[0569] Step 3:

[0570] The server processes the received image data through an image analysis system to extract object features. This analysis uses a Convolutional Neural Network (CNN) to analyze features such as object shape and color at the pixel level. The input is image data, and the output is the object classification result.

[0571] Step 4:

[0572] The server analyzes voice and facial expression data to determine the user's emotional state. This involves analyzing voice tone and using facial expression recognition algorithms. The input is voice and facial expression data, and the output is an emotional state (e.g., joy, anxiety, anger).

[0573] Step 5:

[0574] The server queries a database for local object sorting rules based on the object classification results and the user's location information, and selects the optimal processing method. The input is the object classification results and location information, and the output is the specific processing method based on the sorting rules.

[0575] Step 6:

[0576] The server generates notification messages that reflect how the object is processed and the user's emotional state. For example, if the user is in a "dissatisfied" state, it will create a message that includes positive language. The input is the sorting method and emotional state, and the output is a customized message sent to the user.

[0577] Step 7:

[0578] The terminal displays sorting instructions and messages received from the server on its screen. Based on the displayed information, the user can perform appropriate sorting actions. The input is the message from the server, and the output is visually displayed guidance information.

[0579] (Application Example 2)

[0580] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0581] Accurate waste sorting is crucial for environmental protection, but it's not easy to achieve in daily operations. Furthermore, workers' emotional states and motivation can significantly impact work efficiency. Therefore, there is a need for a system that supports proper waste classification while simultaneously providing feedback that takes into account the user's feelings.

[0582] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0583] In this invention, the server includes means for receiving images of waste captured by the user with a video acquisition device, means for analyzing the images to identify the type of waste, means for identifying an appropriate classification method by referring to waste sorting information in the relevant area based on the classification, means for analyzing emotional data to determine the user's emotional state, and means for generating an appropriate response message for the user based on the emotional state. This makes it possible to provide support for properly classifying waste and to provide feedback that takes the user's emotions into consideration.

[0584] A "video acquisition device" is a recording device that allows the user to take images of waste and transmit them to the system.

[0585] A "server" is a central control unit that analyzes and processes received data.

[0586] "Means for identifying the type of waste" refers to a system that analyzes images of the waste that have been photographed and performs a process to identify its type.

[0587] "Waste sorting information" refers to information regarding the correct classification and disposal methods of waste in a specific area.

[0588] "Emotional data" refers to data obtained from the user's voice and facial expressions, and is used to determine the user's emotional state.

[0589] "Means for generating response messages" refers to a function that generates and presents appropriate text based on the user's emotions and the results of waste classification.

[0590] The system for realizing this invention includes a video acquisition device for the user to capture images of waste and transmit them to a server. When the user captures images of waste using the video acquisition device, the data is transmitted to the server via a terminal. The server analyzes the image data using a machine learning algorithm to identify the type of waste. This analysis is typically performed using software such as TensorFlow.

[0591] Furthermore, the server utilizes emotion recognition software such as Azure Cognitive Services to analyze the user's voice and facial expression data. This analysis determines the user's current emotional state, and based on the results, individually tailored response messages are generated.

[0592] The server also refers to the database and uses location information to determine how to classify the waste. At this time, it uses waste sorting information for the relevant area to inform the user of the appropriate classification method.

[0593] For example, if a user photographs a plastic bottle with a video acquisition device, the system identifies it as plastic waste and instructs the user to classify it as a recyclable resource according to local rules. At the same time, if the system analyzes that the user is feeling fatigued, it provides a motivational message such as, "You've made good progress so far today, let's keep going a little longer!"

[0594] The following are examples of prompts used when utilizing generative AI models.

[0595] By asking a simple question such as, "Where should I dispose of this waste?", you can get appropriate instructions from the system.

[0596] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0597] Step 1:

[0598] The user captures images of the waste using a video acquisition device. The input is the actual appearance of the waste, and the output is digital image data. This image data is transmitted to the server via the terminal.

[0599] Step 2:

[0600] The server analyzes the image data it receives. The input is image data of waste sent by the user. The server uses TensorFlow to identify the type of waste based on a machine learning model. The output is the waste classification information.

[0601] Step 3:

[0602] The server collects user voice and facial expression data and performs emotion recognition. The input is facial expression data obtained from voice and video. Azure Cognitive Services is used to determine the user's emotional state (e.g., joy, surprise, dissatisfaction). This result is the output of the emotion analysis.

[0603] Step 4:

[0604] The server obtains the user's location information and refers to a waste sorting information database. The inputs are classified waste information and the user's location information. Based on this, it searches for waste sorting information in the relevant area and identifies the correct processing method. As output, specific sorting instructions are generated.

[0605] Step 5:

[0606] The server generates appropriate response messages based on the user's emotional state. Input consists of the results of the emotion analysis and classification instructions. The goal is to generate flexible messages tailored to the emotional state, thereby increasing user motivation. Output is a feedback message presented to the user.

[0607] Step 6:

[0608] The terminal notifies the user of sorting instructions and response messages received from the server. The terminal displays sorting methods and messages on the screen. This enables the user to dispose of waste properly and simultaneously provides feedback to maintain motivation for the task.

[0609] The specific processing unit 290 transmits the result of the specific processing to the headset terminal 314. In the headset terminal 314, the control unit 46A causes the speaker 240 and display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0610] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet Search).<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0611] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and specific processing may also be performed by the headset terminal 314.

[0612] [Fourth Embodiment]

[0613] Figure 7 shows an example of the configuration of the data processing system 410 according to the fourth embodiment.

[0614] As shown in Figure 7, the data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.

[0615] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0616] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a controlled object 443. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and controlled object 443 are also connected to the bus 52.

[0617] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0618] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[0619] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0620] The controlled object 443 includes a display device, LEDs in the eyes, and motors that drive the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the robot 414's emotions can be expressed by controlling these motors. Furthermore, the robot 414's facial expressions can also be expressed by controlling the illumination state of the LEDs in its eyes.

[0621] Figure 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Figure 8, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0622] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0623] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0624] In robot 414, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0625] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0626] The system of the present invention assists users in sorting the waste they generate on a daily basis and consists of a mobile terminal, a server, and related software. The program processing of the system is described below in natural language.

[0627] The main components of this system include image capture, image analysis, information retrieval, and classification instruction functions.

[0628] First, the user uses their mobile device to take a picture of the waste using the deployed camera function. This image is then sent to the server via the application.

[0629] The server applies an image analysis algorithm to the received image to identify the type of waste. This analysis uses machine learning techniques to detect the characteristics of the waste with high accuracy based on past data.

[0630] Furthermore, the server uses the user's location information to refer to a database of relevant regions and retrieve sorting rules that are suitable for the characteristics of the waste. Subsequently, the retrieved rules are compared with the analysis results to identify the correct waste sorting method.

[0631] The identified sorting method is transmitted from the server to the terminal. The terminal receives this information and displays it to the user in an easy-to-understand format. This display includes detailed instructions such as the waste sorting category, collection date, and the appropriate type of garbage bag.

[0632] As a concrete example, if a user takes a picture of a plastic bottle from their home, the server will recognize it as plastic waste and provide instructions to separate it as a "recyclable resource" according to the rules of the relevant area.

[0633] Thus, this invention helps users easily and accurately sort waste, contributing to the efficiency of local waste management.

[0634] The following describes the processing flow.

[0635] Step 1:

[0636] Users use their mobile devices to take photos of waste using the application's camera function. After taking the photo, the image is uploaded to the server by tapping the send button within the app.

[0637] Step 2:

[0638] The server receives image data of the waste sent by the user. The received images are temporarily stored in storage and prepared to be sent to the analysis process.

[0639] Step 3:

[0640] The server uses an image analysis algorithm to analyze the image data. Here, a machine learning model extracts features of the waste and identifies its type. For example, it determines categories such as plastic bottles and cans based on shape, color, label information, etc.

[0641] Step 4:

[0642] The server uses the user's location information to access the waste sorting database for that area. It searches for sorting rules corresponding to the identified waste type and obtains the correct sorting method.

[0643] Step 5:

[0644] The server uses the acquired sorting information to generate an easy-to-understand instruction message for the user. This message includes the sorting category and handling precautions.

[0645] Step 6:

[0646] The server sends the generated instruction message to the terminal. The terminal displays the received message on the app screen, showing the user the proper way to sort waste.

[0647] Step 7:

[0648] Users follow instructions from the terminal to sort waste into designated categories and dispose of it appropriately in accordance with local regulations.

[0649] (Example 1)

[0650] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0651] Conventional waste sorting methods require users to manually understand and apply classification criteria, which is time-consuming and cumbersome. Furthermore, adapting to varying waste sorting rules across different regions is difficult, increasing the risk of incorrect disposal. This invention aims to solve these problems and enable users to sort waste quickly and accurately.

[0652] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0653] In this invention, the server includes means for receiving image data generated by a user in a portable information terminal equipped with a camera function, means for analyzing the image data using machine learning technology to recognize the characteristics of the waste, and means for referring to location-dependent regional waste disposal procedure data based on the recognition to identify an appropriate classification method. This enables the user to efficiently sort waste in accordance with local waste sorting rules.

[0654] An "information processing device" is an electronic device programmed to process data and perform specific tasks.

[0655] A "personal information terminal" is an electronic device that a user can carry with them and that has data input and display functions.

[0656] "Image data" refers to visual information generated using a camera function, expressed in digital format.

[0657] "Machine learning technology" is a collection of algorithms and methodologies for learning from past data and making predictions and classifications on new data.

[0658] "Waste characteristics" refer to classification criteria based on the material, shape, and intended use of an object.

[0659] "Location information" is digital data used to identify a specific location, and is acquired using technologies such as GPS.

[0660] "Regional waste disposal procedure data" refers to information regarding waste sorting methods and regulations established in each region.

[0661] A "classification method" is a set of procedures and criteria for dividing waste into specific categories based on its characteristics.

[0662] To implement the present invention, the system is implemented with a configuration including an information processing device, a personal information terminal, and related software. Specifically, the personal information terminal is equipped with a camera function, which the user uses to generate image data of objects that are routinely discarded. This image data is then transmitted to a server through a dedicated application.

[0663] The server analyzes the received image data. Machine learning techniques are used for the analysis, utilizing popular machine learning frameworks such as TensorFlow and PyTorch. This technology allows the server to recognize the characteristics of the waste with high accuracy and identify its material and shape.

[0664] Furthermore, the server uses location information obtained from the user's mobile device to refer to waste disposal procedure data for each region. Location information is typically obtained using GPS technology, helping the server identify sorting rules relevant to the user's current location. Based on this information, the server identifies the appropriate sorting method and notifies the user.

[0665] Notifications are sent in real time, and the device displays the information clearly and visually to the user. The display includes instructions such as the waste classification category, collection date, and the appropriate type of garbage bag.

[0666] For example, if a user takes a picture of a plastic bottle from their home, the system will recognize it as plastic waste and provide instructions to process it as a "recyclable resource" as defined by the local area. This makes it easy for users to dispose of their waste according to local regulations.

[0667] An example of a prompt to a generating AI model is, "Based on the image data taken by the user, output instructions regarding the appropriate sorting category and specific collection date." This prompt allows the AI ​​model to provide accurate and practical sorting instructions.

[0668] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0669] Step 1:

[0670] The user uses the camera function of their mobile device to take pictures of the waste. The captured image data is sent to the server via the application. The input is the image data captured by the camera, and the output is the digital image data received by the server.

[0671] Step 2:

[0672] The server analyzes the received image data using a machine learning algorithm. Specifically, a model built using TensorFlow or PyTorch classifies objects in the image based on their characteristics as waste. The input to this step is the image data received by the server, and the output is information about the characteristics of the waste based on that image.

[0673] Step 3:

[0674] The server uses location information obtained from the user's terminal to refer to a regional waste disposal procedure database. The input is location information obtained using GPS, and the output is a list of region-specific sorting methods based on that location information.

[0675] Step 4:

[0676] The server compares the image analysis results with the acquired local sorting rules to identify the appropriate waste classification method. This comparison process determines which sorting category the waste belongs to. The input is the waste's characteristics and a list of local sorting methods, and the output is the optimal sorting method.

[0677] Step 5:

[0678] The server transmits the identified sorting method to the user's mobile device. Communication takes place in real time, and the user can receive the results immediately. The input is the sorting method identified by the server, and the output is the sorting instruction information displayed on the user's device.

[0679] Step 6:

[0680] The terminal visually displays the received sorting information to the user. The display screen shows the waste classification category, collection date, and appropriate garbage bag type, allowing the user to process the waste accordingly. The input is the notified sorting instruction information, and the output is the detailed instruction content displayed on the user's screen.

[0681] (Application Example 1)

[0682] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0683] Waste sorting is a crucial issue for environmental protection and efficient resource utilization, but conventional methods have limitations in terms of accuracy and efficiency. In particular, large amounts of waste are generated in industry and manufacturing, requiring proper disposal. On-site sorting often relies on manual labor, leading to problems such as the depletion of human resources and increased environmental burden due to errors. Therefore, there is a need for automated, high-precision, and highly efficient sorting systems.

[0684] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0685] In this invention, the server includes means for receiving visual information of waste captured by a user with a device having a camera function, means for classifying the waste by analyzing the visual information, means for identifying the correct disposal method by referring to the relevant local waste disposal regulations, and means for physically separating the waste using a powered machine. This enables automatic and accurate sorting of waste.

[0686] "User" refers to an individual or organization that utilizes the waste sorting system.

[0687] "Photography function" refers to the capability of a device to capture visual information about waste.

[0688] "Device" refers to hardware components, including cameras and sensors.

[0689] "Visual information" refers to image data acquired using a camera or other imaging function.

[0690] "Analysis" refers to the process of identifying the type and characteristics of waste from visual information using machine learning algorithms.

[0691] "Classification" refers to dividing waste into predetermined categories based on the results of analysis.

[0692] "Local waste disposal regulations" refer to laws and guidelines concerning the collection, treatment, and recycling of waste that apply in a specific area.

[0693] "Disposal method" refers to the procedures for properly processing or recycling waste.

[0694] A "powered machine" refers to a device that supplies energy for performing mechanical work.

[0695] "Physically separating" refers to the process of physically separating waste using mechanical equipment rather than manual methods.

[0696] In embodiments of the present invention, a waste sorting system is used. The system consists of a device with a camera function (hereinafter referred to as a terminal), a server that processes data, and related software. The user uses the camera function of the terminal to acquire visual information of the waste to be sorted. The acquired image data is transmitted to the server via the network.

[0697] The server first uses machine learning algorithms on the received visual information to identify the type of waste. This analysis utilizes previously accumulated waste characteristic information to enable highly accurate classification. Based on the identified information, the server considers the user's location and refers to the waste disposal regulations database for the relevant area to determine the appropriate disposal method.

[0698] The identified disposal method includes information on means of physically separating the waste using powered machinery, and the generated instructions are sent to the terminal. The terminal displays visual information, disposal methods, and precautions to the user in an easy-to-understand format.

[0699] For example, if a user takes a picture of a plastic part used in a factory, the server will identify it as recyclable material and generate instructions on how to transport it to the appropriate recycling bin.

[0700] Examples of prompts to input into a generative AI model:

[0701] "Please explain how a robot operating within a factory can photograph waste materials, perform image analysis, and generate sorting instructions based on the results."

[0702] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0703] Step 1:

[0704] The user takes a picture of the waste using their device. The input is capturing the actual waste with the camera. The output is digital image data generated as pixel information. Specifically, the user presses the capture button, which activates the camera and saves a photograph of the object.

[0705] Step 2:

[0706] The device transmits the captured image data to the server via the network. The input is the digital image obtained in step 1. The output is the image data received by the server. Specifically, the device establishes a network connection and uploads the image data to the server using the HTTP protocol.

[0707] Step 3:

[0708] The server analyzes the received image data to identify the type of waste. The input is image data sent from the terminal. The output is the identified waste type and characteristic information. Specifically, the server uses computer vision technology and machine learning algorithms to extract features from the image and classify them by comparing them with a database.

[0709] Step 4:

[0710] The server obtains the user's location information and refers to a database of waste disposal regulations for the relevant region. Inputs are the type of waste and the user's location. Outputs are instructions for the appropriate disposal method. Specifically, the server uses GPS data to determine the location and retrieves relevant information from the database based on that location.

[0711] Step 5:

[0712] The server determines the means of physically separating the waste using a powered machine and sends instructions to the terminal. The input is the instruction on the disposal method. The output is a sorting method guide displayed on the terminal. Specifically, the server converts the instruction content into text format and sends it to the terminal.

[0713] Step 6:

[0714] The terminal visually presents the received instructions to the user. The input is the sorting method instructions sent from the server. The output is the instruction screen that the user can view. Specifically, the terminal displays the instruction content on its screen and provides information using icons and text so that the user can easily understand it.

[0715] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0716] This invention is a system that assists users in processing waste they sort on a daily basis and incorporates an emotion engine to provide appropriate feedback according to the user's emotional state. The system's program processing is described below in natural language.

[0717] This system consists of functions related to image capture, emotion recognition, image analysis, information retrieval, and classification instructions.

[0718] First, the user uses their mobile device to take a picture of the waste using the camera. This image is sent to the server via the application. In addition, the device simultaneously collects data on the user's voice and facial expressions, which are then transmitted to the server for analysis of the user's emotional state by an emotion engine.

[0719] The server analyzes received images of waste using an image analysis algorithm to identify the type of waste. This analysis utilizes a machine learning model, achieving highly accurate classification based on accumulated data. In addition, an emotion engine analyzes the user's voice tone and facial expression data to determine their current emotional state (e.g., joy, surprise, dissatisfaction).

[0720] Next, the server uses the user's location information to refer to the relevant regional waste sorting database. Here, it searches for the appropriate sorting rules corresponding to the identified waste type and determines the correct disposal method for the waste.

[0721] Based on the identified classification method and the user's emotional state, the server generates encouraging messages and situation-specific warning messages. For example, if the user is dissatisfied, a positive message such as "You're doing well, let's move on!" might be generated.

[0722] This information is sent from the server to the terminal, which then displays sorting instructions and messages on the screen to provide the user with appropriate guidance and motivation.

[0723] As a concrete example, when a user photographs a plastic bottle that should be recycled, the server recognizes it as plastic waste and instructs it to separate it as "recyclable material" according to the relevant local rules. At the same time, if the emotion engine detects dissatisfaction from the user's facial expression, it also displays a message to boost their motivation.

[0724] Thus, the system of the present invention not only supports users in properly sorting waste, but also provides emotionally responsive feedback to realize a more comfortable and effective user experience.

[0725] The following describes the processing flow.

[0726] Step 1:

[0727] The user launches the application on their mobile device and takes a picture of the waste using the camera function. If audio or facial expressions are available, the app will also record this data simultaneously.

[0728] Step 2:

[0729] The device transmits captured image data, recorded audio, and facial expression data to a server via the app. This data also includes the user's location information.

[0730] Step 3:

[0731] The server receives the image data and first performs image analysis using a machine learning algorithm. This analysis extracts the characteristics of the waste and identifies its type.

[0732] Step 4:

[0733] The server analyzes voice and facial expression data using an emotion engine to determine the user's emotional state. For example, it analyzes voice tone and detects changes in facial expressions to identify whether the user is happy or dissatisfied.

[0734] Step 5:

[0735] The server accesses a waste sorting database for the relevant region based on the user's location information. Using the analysis results, it retrieves the appropriate sorting method for the waste from the database.

[0736] Step 6:

[0737] The server generates notification messages for the user based on the acquired sorting method. Furthermore, it adds positive messages and words of encouragement depending on the user's emotional state.

[0738] Step 7:

[0739] The terminal receives a message sent from the server and displays it on the screen. This message includes instructions on how to properly sort waste and emotionally responsive feedback.

[0740] Step 8:

[0741] Users review the displayed information and follow the instructions to properly sort their waste. The positive messages displayed are also expected to boost user motivation.

[0742] (Example 2)

[0743] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0744] In today's environmental challenges, the sorting of waste that users perform daily is important, but accurate classification and adherence to local sorting rules are difficult, sometimes resulting in improper disposal. Furthermore, insufficient feedback on user behavior makes it difficult to maintain motivation for sorting.

[0745] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0746] In this invention, the server includes means for receiving images of objects captured by the user with a communication terminal, means for analyzing the images and classifying the objects, and means for receiving audio data and facial expression data and determining the user's emotional state. This makes it possible to support accurate object classification while providing appropriate feedback according to the user's emotions.

[0747] A "communication terminal" is an electronic device used by a user to send and receive information, and is equipped with camera and voice input functions.

[0748] An "image of an object" is digital data that shows the visible shape of an object, captured using the camera of a communication terminal.

[0749] "Analyzing an image" is the process of identifying the features of an object from digital data and classifying its attributes.

[0750] "Object classification" is the act of dividing objects into specific categories based on their analyzed characteristics.

[0751] "Audio data" refers to information that represents acoustic signals acquired through the microphone of a communication terminal in digital format.

[0752] "Facial expression data" refers to digital data containing information about facial expressions captured by the camera of a communication device.

[0753] "User's emotional state" refers to information indicating the psychological state obtained by analyzing voice data and facial expression data, and includes examples such as "joy," "dissatisfaction," and "surprise."

[0754] "Generating adaptive notification messages that respond to emotions" means creating feedback messages that are appropriate for individual situations, taking into account the user's emotional state.

[0755] This invention is a system that links a communication terminal and a server to assist users in sorting objects on a daily basis. Specifically, the communication terminal uses a camera and microphone to capture image data by photographing objects, and in addition, it collects the user's voice and facial expressions. This data is transmitted from the terminal to the server.

[0756] The server uses generative AI models, such as Convolutional Neural Networks (CNNs), to analyze the features of objects based on the received image data. The image analysis algorithm classifies objects into specific categories. The server also utilizes facial recognition and natural language processing (NLP) technologies to analyze audio and facial data and determine the user's emotional state (joy, surprise, dissatisfaction, etc.).

[0757] Based on the analyzed data, the server references the user's location information and retrieves region-specific object classification rules from a database. Following these rules, it identifies the optimal method for processing the object and generates feedback tailored to the identified emotional state. For example, if the user expresses dissatisfaction, it might create an encouraging message such as, "Very well done, keep up the good work next time!"

[0758] Ultimately, this information is transmitted to the terminal, and the results are displayed to the user on the screen. This allows the user to efficiently perform the appropriate processing of the object while maintaining motivation.

[0759] For example, if a user photographs a reusable plastic bottle, the server will recognize the object as "plastic" and instruct the user to separate it as "recyclable waste" according to local rules. At the same time, if sentiment analysis determines that the user is frustrated, it will display a positive message such as, "Great job, please recycle next time!"

[0760] An example of a prompt to input into the generating AI model is, "Please tell me the steps to determine the type of object, apply the sorting rules for that location, and generate the optimal action."

[0761] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0762] Step 1:

[0763] The user takes an image of an object using the camera on the communication terminal. The input image data must be high resolution so that the entire object is clearly visible. As output, the captured image is saved to the memory of the communication terminal.

[0764] Step 2:

[0765] The device prepares to send the captured image data to the server. It also collects user voice data using the microphone and acquires facial expression data using the camera. The inputs are image, audio, and facial expression data, and these are sent to the server as output.

[0766] Step 3:

[0767] The server processes the received image data through an image analysis system to extract object features. This analysis uses a Convolutional Neural Network (CNN) to analyze features such as object shape and color at the pixel level. The input is image data, and the output is the object classification result.

[0768] Step 4:

[0769] The server analyzes voice and facial expression data to determine the user's emotional state. This involves analyzing voice tone and using facial expression recognition algorithms. The input is voice and facial expression data, and the output is an emotional state (e.g., joy, anxiety, anger).

[0770] Step 5:

[0771] The server queries a database for local object sorting rules based on the object classification results and the user's location information, and selects the optimal processing method. The input is the object classification results and location information, and the output is the specific processing method based on the sorting rules.

[0772] Step 6:

[0773] The server generates notification messages that reflect how the object is processed and the user's emotional state. For example, if the user is in a "dissatisfied" state, it will create a message that includes positive language. The input is the sorting method and emotional state, and the output is a customized message sent to the user.

[0774] Step 7:

[0775] The terminal displays sorting instructions and messages received from the server on its screen. Based on the displayed information, the user can perform appropriate sorting actions. The input is the message from the server, and the output is visually displayed guidance information.

[0776] (Application Example 2)

[0777] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0778] Accurate waste sorting is crucial for environmental protection, but it's not easy to achieve in daily operations. Furthermore, workers' emotional states and motivation can significantly impact work efficiency. Therefore, there is a need for a system that supports proper waste classification while simultaneously providing feedback that takes into account the user's feelings.

[0779] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0780] In this invention, the server includes means for receiving images of waste captured by the user with a video acquisition device, means for analyzing the images to identify the type of waste, means for identifying an appropriate classification method by referring to waste sorting information in the relevant area based on the classification, means for analyzing emotional data to determine the user's emotional state, and means for generating an appropriate response message for the user based on the emotional state. This makes it possible to provide support for properly classifying waste and to provide feedback that takes the user's emotions into consideration.

[0781] A "video acquisition device" is a recording device that allows the user to take images of waste and transmit them to the system.

[0782] A "server" is a central control unit that analyzes and processes received data.

[0783] "Means for identifying the type of waste" refers to a system that analyzes images of the waste that have been photographed and performs a process to identify its type.

[0784] "Waste sorting information" refers to information regarding the correct classification and disposal methods of waste in a specific area.

[0785] "Emotional data" refers to data obtained from the user's voice and facial expressions, and is used to determine the user's emotional state.

[0786] "Means for generating response messages" refers to a function that generates and presents appropriate text based on the user's emotions and the results of waste classification.

[0787] The system for realizing this invention includes a video acquisition device for the user to capture images of waste and transmit them to a server. When the user captures images of waste using the video acquisition device, the data is transmitted to the server via a terminal. The server analyzes the image data using a machine learning algorithm to identify the type of waste. This analysis is typically performed using software such as TensorFlow.

[0788] Furthermore, the server utilizes emotion recognition software such as Azure Cognitive Services to analyze the user's voice and facial expression data. This analysis determines the user's current emotional state, and based on the results, individually tailored response messages are generated.

[0789] The server also refers to the database and uses location information to determine how to classify the waste. At this time, it uses waste sorting information for the relevant area to inform the user of the appropriate classification method.

[0790] For example, if a user photographs a plastic bottle with a video acquisition device, the system identifies it as plastic waste and instructs the user to classify it as a recyclable resource according to local rules. At the same time, if the system analyzes that the user is feeling fatigued, it provides a motivational message such as, "You've made good progress so far today, let's keep going a little longer!"

[0791] The following are examples of prompts used when utilizing generative AI models.

[0792] By asking a simple question such as, "Where should I dispose of this waste?", you can get appropriate instructions from the system.

[0793] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0794] Step 1:

[0795] The user captures images of the waste using a video acquisition device. The input is the actual appearance of the waste, and the output is digital image data. This image data is transmitted to the server via the terminal.

[0796] Step 2:

[0797] The server analyzes the image data it receives. The input is image data of waste sent by the user. The server uses TensorFlow to identify the type of waste based on a machine learning model. The output is the waste classification information.

[0798] Step 3:

[0799] The server collects user voice and facial expression data and performs emotion recognition. The input is facial expression data obtained from voice and video. Azure Cognitive Services is used to determine the user's emotional state (e.g., joy, surprise, dissatisfaction). This result is the output of the emotion analysis.

[0800] Step 4:

[0801] The server obtains the user's location information and refers to a waste sorting information database. The inputs are classified waste information and the user's location information. Based on this, it searches for waste sorting information in the relevant area and identifies the correct processing method. As output, specific sorting instructions are generated.

[0802] Step 5:

[0803] The server generates appropriate response messages based on the user's emotional state. Input consists of the results of the emotion analysis and classification instructions. The goal is to generate flexible messages tailored to the emotional state, thereby increasing user motivation. Output is a feedback message presented to the user.

[0804] Step 6:

[0805] The terminal notifies the user of sorting instructions and response messages received from the server. The terminal displays sorting methods and messages on the screen. This enables the user to dispose of waste properly and simultaneously provides feedback to maintain motivation for the task.

[0806] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the controlled object 443 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0807] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet Search).<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0808] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the robot 414.

[0809] Furthermore, the emotion identification model 59, acting as an emotion engine, may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to a specific mapping, which is an emotion map (see Figure 9). Similarly, the emotion identification model 59 may also determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.

[0810] Figure 9 shows an emotion map 400 in which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. The closer to the center of the concentric circles, the more primitive the emotions are located. Further out of the concentric circles, emotions representing states and actions arising from mental states are located. Emotion is a concept that includes feelings and mental states. On the left side of the concentric circles, emotions that are generally generated from reactions occurring in the brain are located. On the right side of the concentric circles, emotions that are generally induced by situational judgment are located. Above and below the concentric circles, emotions that are generally generated from reactions occurring in the brain and induced by situational judgment are located. In addition, the emotion of "pleasure" is located on the upper side of the concentric circles, and the emotion of "displeasure" is located on the lower side. Thus, in the emotion map 400, multiple emotions are mapped based on the structure in which emotions arise, and emotions that are likely to occur simultaneously are mapped close together.

[0811] These emotions are distributed at the 3 o'clock position on the Emotion Map 400, and usually fluctuate between feelings of security and anxiety. In the right half of the Emotion Map 400, situational awareness takes precedence over internal feelings, resulting in a calm impression.

[0812] The inside of the Emotion Map 400 represents inner thoughts, while the outside represents actions. Therefore, the further you go from the outside of the Emotion Map 400, the more visible (expressed in actions) your emotions become.

[0813] Here, human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. Similarly, in robots, cars, motorcycles, etc., emotions can be created based on various balances, such as posture and battery level. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. The emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on a system for analyzing brain physiological signals of speech emotion recognition and emotion, Tokushima University, doctoral dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map contains emotions belonging to a region called "response," where sensation is dominant. The right half of the emotion map contains emotions belonging to a region called "situation," where situational awareness is dominant.

[0814] The emotion map defines two emotions that promote learning. One is the emotion around the middle of the negative "repentance" and "reflection" on the situation side. In other words, it is when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is the emotion around the positive "desire" on the reaction side. In other words, it is when the robot has positive feelings such as "I want more" or "I want to know more."

[0815] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values ​​representing each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple training data sets, which are combinations of user input and emotion values ​​representing each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions located close together have similar values, as shown in the emotion map 900 in Figure 10. Figure 10 shows an example where multiple emotions such as "reassured," "calm," and "confident" have similar emotion values.

[0816] The above description primarily focuses on the functions of the data processing device 12 in relation to this disclosure. However, the system related to this disclosure is not necessarily implemented on a server. The system related to this disclosure may be implemented as a general information processing system. This disclosure may be implemented, for example, as a software program that runs on a personal computer or as an application that runs on a smartphone. The method related to this disclosure may be provided to users in SaaS (Software as a Service) format.

[0817] In the above embodiment, an example was given in which a specific process is performed by a single computer 22. However, the technology of this disclosure is not limited thereto, and a distributed processing of the specific process may be performed by multiple computers, including computer 22. For example, a data generation model 58 may be provided in an external device of the data processing device 12, and the external device may generate data according to the input data.

[0818] In the above embodiment, an example was given in which the specific processing program 56 is stored in the storage 32, but the technology of this disclosure is not limited thereto. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-temporary storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-temporary storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes specific processing according to the specific processing program 56.

[0819] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.

[0820] Furthermore, it is not necessary to store the entirety of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store the entirety of the specific processing program 56 in the storage 32; it is acceptable to store only a portion of the specific processing program 56.

[0821] The following types of processors can be used as hardware resources to perform specific processing. Examples of processors include a CPU, a general-purpose processor that functions as a hardware resource to perform specific processing by executing software, i.e., a program. Other examples of processors include dedicated electrical circuits, such as FPGAs (Field-Programmable Gate Arrays), PLDs (Programmable Logic Devices), or ASICs (Application Specific Integrated Circuits), which have circuit configurations specifically designed to perform specific processing. All of these processors have built-in or connected memory, and all of them perform specific processing by using memory.

[0822] The hardware resource that performs a specific process may consist of one of these various processors, or it may consist of a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Alternatively, the hardware resource that performs a specific process may consist of a single processor.

[0823] Examples of configurations using a single processor include, firstly, a configuration in which one or more CPUs and software are combined to form a single processor, and this processor functions as a hardware resource that performs a specific process. Secondly, there is a configuration using a processor that realizes the functions of the entire system, including multiple hardware resources that perform a specific process, on a single IC chip, as exemplified by SoCs (System-on-a-chip). In this way, a specific process is realized using one or more of the above types of processors as hardware resources.

[0824] Furthermore, the hardware structure of these various processors can more specifically utilize electrical circuits that combine circuit elements such as semiconductor devices. Also, the specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps can be deleted, new steps added, or the processing order rearranged, as long as it does not deviate from the main purpose.

[0825] The descriptions and illustrations presented above are detailed explanations of the technical aspects of this disclosure and are merely examples of the technical aspects. For example, the above descriptions of the structure, function, operation, and effect are examples of the structure, function, operation, and effect of the technical aspects of this disclosure. Therefore, it goes without saying that you may delete unnecessary parts, add new elements, or replace elements in the descriptions and illustrations presented above, as long as you do not deviate from the essence of the technical aspects of this disclosure. Furthermore, in order to avoid confusion and facilitate understanding of the technical aspects of this disclosure, explanations of common technical knowledge and the like that do not require special explanation to enable the implementation of the technical aspects of this disclosure have been omitted from the descriptions and illustrations presented above.

[0826] All documents, patent applications, and technical standards described herein are incorporated by reference to the same extent as if each individual document, patent application, and technical standard were specifically and individually noted as being incorporated by reference.

[0827] The following is further disclosed regarding the embodiments described above.

[0828] (Claim 1)

[0829] A means of receiving images of waste taken by the user with a mobile device,

[0830] A means for classifying waste by analyzing the aforementioned images,

[0831] Based on the above classification, means for identifying the correct sorting method by referring to the relevant local waste sorting regulations,

[0832] Means for notifying the user of the specified sorting method,

[0833] A system that includes this.

[0834] (Claim 2)

[0835] The system according to claim 1, wherein the analysis means identifies the characteristics of waste using a machine learning algorithm.

[0836] (Claim 3)

[0837] The system according to claim 1, wherein the reference means acquires data from a waste sorting database based on location information.

[0838] "Example 1"

[0839] (Claim 1)

[0840] In a portable information terminal equipped with a camera function, the information processing device includes means for receiving image data generated by the user,

[0841] A means for analyzing the aforementioned image data using machine learning technology to recognize the characteristics of the waste,

[0842] Based on the aforementioned understanding, a means for referring to location-dependent regional waste disposal procedure data and identifying an appropriate classification method,

[0843] A means of presenting the identified classification method to the user,

[0844] A system that includes this.

[0845] (Claim 2)

[0846] The system according to claim 1, wherein the analysis means includes an information processing system that identifies the characteristics of waste with high accuracy by comparison with a prior database.

[0847] (Claim 3)

[0848] The system according to claim 1, wherein the aforementioned reference means acquires information from a waste classification database based on location information obtained by location identification technology.

[0849] "Application Example 1"

[0850] (Claim 1)

[0851] A means for receiving visual information of waste captured by a user with a device having a camera function,

[0852] A means for classifying waste by analyzing the aforementioned visual information,

[0853] Based on the above classification, a means of identifying the correct disposal method by referring to the relevant local waste disposal regulations,

[0854] Means for notifying the user of the specified disposal method,

[0855] A means of physically separating waste using powered machinery,

[0856] A system that includes this.

[0857] (Claim 2)

[0858] The system according to claim 1, wherein the analysis means identifies the characteristics of waste using a learning algorithm.

[0859] (Claim 3)

[0860] The system according to claim 1, wherein the aforementioned reference means obtains information from a waste disposal database based on location information.

[0861] "Example 2 of combining an emotion engine"

[0862] (Claim 1)

[0863] A means for receiving images of objects taken by the user with a communication terminal,

[0864] A means for classifying objects by analyzing the aforementioned image,

[0865] A means of receiving voice data and facial expression data to determine the user's emotional state,

[0866] Based on the aforementioned classification, means for identifying the precise sorting method by referring to the relevant local object sorting rules,

[0867] Means for generating adaptive notification messages based on identified classification methods and the user's emotional state,

[0868] Means for providing the aforementioned notification message to the user,

[0869] A system that includes this.

[0870] (Claim 2)

[0871] The system according to claim 1, wherein the analysis means identifies the features of an object using a data processing algorithm.

[0872] (Claim 3)

[0873] The system according to claim 1, wherein the reference means acquires data from an object classification database based on location information.

[0874] "Application example 2 when combining with an emotional engine"

[0875] (Claim 1)

[0876] A means for receiving images of waste captured by the user with a video acquisition device,

[0877] A means for analyzing the aforementioned image to identify the type of waste,

[0878] Based on the above classification, means for identifying an appropriate classification method by referring to waste sorting information in the relevant area,

[0879] Means for notifying the user of the identified classification method,

[0880] A means of analyzing emotional data to determine the user's emotional state,

[0881] A means for generating an appropriate response message for the user based on the aforementioned emotional state,

[0882] A system that includes this.

[0883] (Claim 2)

[0884] The system according to claim 1, wherein the analysis means identifies the characteristics of waste using a learning algorithm.

[0885] (Claim 3)

[0886] The system according to claim 1, wherein the reference means acquires data from a waste classification information base based on location information. [Explanation of Symbols]

[0887] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Devices 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robots< / url:> < / url:> < / url:> < / url:>

Claims

1. A means of receiving images of waste taken by the user with a mobile device, A means for classifying waste by analyzing the aforementioned images, Based on the above classification, means for identifying the correct sorting method by referring to the relevant local waste sorting regulations, Means for notifying the user of the specified sorting method, A system that includes this.

2. The system according to claim 1, wherein the analysis means identifies the characteristics of waste using a machine learning algorithm.

3. The system according to claim 1, wherein the aforementioned reference means acquires data from a waste sorting database based on location information.

Citation Information

Patent Citations

  • Persona chatbot control method and system

    JP2022180282A