Information processing system
By automatically detecting household stains or pests through image acquisition units and artificial intelligence models, and providing cleaning guidance in conjunction with chatbots, this technology solves the problems of low efficiency and heavy user workload in existing household cleaning management, and achieves efficient and intelligent household environment management.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-27
- Publication Date
- 2026-03-10
AI Technical Summary
Existing home cleaning management systems struggle to identify stains or pests accurately and promptly, resulting in low efficiency, heavy user workload, and a lack of scientific cleaning methods and tools, making it difficult to maintain a clean living environment in the long term.
The system regularly collects home image data through an image acquisition unit, uses an artificial intelligence model to automatically detect stains or pests, generates cleaning notifications, and provides cleaning guidance through a chatbot, thus achieving automated and personalized home cleaning management.
It enables automated, high-precision detection and early notification of the home environment, provides real-time cleaning method suggestions, reduces the burden on users, and improves the hygiene level of the living environment.
Smart Images

Figure CN121640355A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The technology of the present disclosure relates to an information processing system. BACKGROUND
[0002] In Japanese Patent Application Publication No. 2022-180282, a role chatbot control method executed by at least one processor is disclosed, which includes the steps of receiving a user utterance, adding the user utterance to a prompt containing an instruction sentence associated with an explanation about a role of a chatbot, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance.
[0003] The existing home cleaning management relies on manual inspection and subjective judgment, which is difficult to accurately identify cleaning targets such as stains or pests in the home in a timely manner, and has problems such as low efficiency, high risk of omission, and heavy user operation burden. In addition, some users lack scientific cleaning methods and tool knowledge, and cannot efficiently complete cleaning work, resulting in difficulty in maintaining the health of the living environment for a long time. SUMMARY
[0004] The present disclosure provides an information processing system, which includes an image acquisition unit, an image analysis unit, an analysis result recording unit, a notification generation unit, a notification sending unit, a user interface unit, and a chatbot unit. The system periodically collects and stream transfers image data through the camera in the home, automatically detects stains or pests using an artificial intelligence model, records and manages the detection results and related information. When the detection confidence is higher than the set threshold, the system automatically generates and pushes a cleaning or attention notification message to the user terminal, and displays it in real time through the user interface. At the same time, the system supports users to ask cleaning methods and other questions through the chatbot, and automatically generates practical replies using natural language processing technology to assist users to scientifically and efficiently complete home cleaning work, thereby improving the health level of the living environment and reducing the management and operation burden of the users.
[0005] The "image acquisition unit" refers to a device or module that can periodically receive and stream acquire image data from a camera device.
[0006] The "image analysis unit" refers to a device or module that pre-processes the received image data and identifies stains or pests using an artificial intelligence model.
[0007] The "analysis result recording unit" refers to a device or module that records and stores related data such as detection results, confidence scores, timestamps, and camera location information into a database.
[0008] The "notification generation unit" refers to a device or module that automatically extracts a position that needs to be cleaned or paid attention to when the confidence score of the detection result exceeds a set threshold, and generates a corresponding notification message.
[0009] The "notification sending unit" refers to a device or module responsible for sending the generated notification message to the user terminal device.
[0010] The "user interface unit" refers to a device or module that displays information to the user through the terminal device in the form of push notifications or in-app messages.
[0011] The "chatbot unit" refers to a device or module that can receive questions raised by the user, analyze them using natural language processing technology, and display specific answers to the user in the user interface unit.
[0012] The "camera device" refers to a camera installed at different positions in the room for collecting and sending image data.
[0013] The "artificial intelligence model" refers to a machine learning algorithm, such as a convolutional neural network (CNN), used to analyze image content and implement stain or pest detection and recognition.
[0014] The "database" refers to a data management system used to store analysis results, detection data, time information, and camera positions.
[0015] The "confidence score" refers to a numerical value output by the artificial intelligence model that represents the degree of confidence in the detection result.
[0016] The "threshold value" refers to a confidence score benchmark for determining whether a notification is needed. Only when the value is higher than this value will a notification be sent.
[0017] The "terminal device" refers to an electronic device used by the user to receive notifications and interact with the system, such as a smartphone or tablet. BRIEF DESCRIPTION OF DRAWINGS
[0018] Figure 1 is a conceptual diagram representing an example of the configuration of the data processing system of the first embodiment.
[0019] Figure 2 is a conceptual diagram representing an example of the main part functions of the data processing device and the smart device of the first embodiment.
[0020] Figure 3 is a conceptual diagram representing an example of the configuration of the data processing system of the second embodiment.
[0021] Figure 4 is a conceptual diagram representing an example of the main part functions of the data processing device and the smart glasses of the second embodiment.
[0022] Figure 5 is a conceptual diagram showing an example of the configuration of the data processing system of the third embodiment.
[0023] Figure 6 is a conceptual diagram showing an example of the main part functions of the data processing device and the head-mounted terminal of the third embodiment.
[0024] Figure 7 is a conceptual diagram showing an example of the configuration of the data processing system of the fourth embodiment.
[0025] Figure 8 is a conceptual diagram showing an example of the main part functions of the data processing device and the robot of the fourth embodiment.
[0026] Figure 9 represents an emotion map mapping a plurality of emotions.
[0027] Figure 10 represents an emotion map mapping a plurality of emotions.
[0028] Figure 11 is a sequence diagram showing a processing flow of the data processing system of the first embodiment.
[0029] Figure 12 is a sequence diagram showing a processing flow of the data processing system in Application Example 1.
[0030] Figure 13 is a sequence diagram showing a processing flow of the data processing system of the second embodiment.
[0031] Figure 14 is a sequence diagram showing a processing flow of the data processing system in Application Example 2. DETAILED DESCRIPTION
[0032] Hereinafter, an example of an embodiment of a system to which the technology of the present disclosure is applied will be described with reference to the drawings.
[0033] First, the terms used in the following description will be explained.
[0034] In the following embodiments, the processor (hereinafter, simply referred to as "processor") denoted by the reference sign can be one arithmetic device or a combination of a plurality of arithmetic devices. Further, the processor can be one arithmetic device or a combination of a plurality of arithmetic devices. As an example of the arithmetic device, a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), an APU (Accelerated Processing Unit), or the like can be cited.
[0035] In the following embodiments, the RAM (Random Access Memory) denoted by the reference sign is a memory that temporarily stores information, and is used as a work memory by the processor.
[0036] In the following embodiments, the memory denoted by the reference sign is one or a plurality of nonvolatile storage devices that store various programs, various parameters, and the like. As an example of the nonvolatile storage device, a flash memory (SSD (Solid State Drive)), a magnetic disk (for example, a hard disk), a magnetic tape, or the like can be cited.
[0037] In the following embodiments, the communication I / F (Interface) denoted by the reference sign is an interface that includes a communication processor and an antenna, and the like. The communication I / F is responsible for communication between a plurality of computers. As an example of the communication standard suitable for the communication I / F, a wireless communication standard including 5G (5th Generation Mobile Communication System), Wi-Fi (wireless fidelity) (registered trademark), or Bluetooth (registered trademark), or the like can be cited.
[0038] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B". That is, "A and / or B" means that it can be only A, or only B, or a combination of A and B. Further, in the present specification, when three or more items are connected by "and / or", the same understanding as "A and / or B" is applied.
[0039] First Embodiment
[0040] Figure 1 An example of the configuration of the data processing system 10 of the first embodiment is shown in FIG. 1.
[0041] As shown in Figure 1 Fig. 1, a data processing system 10 is provided with a data processing apparatus 12 and an intelligent device 14. As an example of the data processing apparatus 12, a server can be cited.
[0042] The data processing apparatus 12 is provided with a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of the "computer" involved in the technology of the present disclosure. The computer 22 is provided with a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. In addition, the database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. As an example of the network 54, a WAN (Wide Area Network), a LAN (Local Area Network), or the like can be cited.
[0043] The intelligent device 14 is provided with a computer 36, a reception apparatus 38, an output apparatus 40, a camera 42, and a communication I / F 44. The computer 36 is provided with a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. In addition, the reception apparatus 38, the output apparatus 40, the camera 42, and the communication I / F 44 are also connected to the bus 52.
[0044] The reception apparatus 38 is provided with a touch screen 38A and a microphone 38B or the like, and receives user input. The touch screen 38A receives user input by contact of a pointing body (for example, a pen or a finger or the like) by detecting the contact of the pointing body. The microphone 38B receives user input by voice by detecting the voice of the user. A control section 46A in the processor 46 transmits data indicating the user input received by the touch screen 38A and the microphone 38B to the data processing apparatus 12. In the data processing apparatus 12, a specific processing section 290 acquires the data indicating the user input.
[0045] The output apparatus 40 is provided with a display 40A and a speaker 40B or the like, and presents data to the user 20 by outputting data in a presentation form (for example, voice and / or text) that the user 20 can perceive. The display 40A displays visual information such as text and images in accordance with an instruction from the processor 46. The speaker 40B outputs voice in accordance with an instruction from the processor 46. The camera 42 is a small digital camera equipped with an optical system such as a lens, an aperture, and a shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.
[0046] Communication I / F44 is connected to network 54. Communication I / F44 and 26 are responsible for sending and receiving various information between processor 46 and processor 28 via network 54.
[0047] Figure 2 The diagram shows an example of the main functions of the data processing device 12 and the smart device 14.
[0048] like Figure 2 As shown, in the data processing apparatus 12, specific processing is performed by the processor 28. A specific processing program 56 is stored in the memory 32. The specific processing program 56 is an example of a "program" as understood in this disclosure. The processor 28 reads the specific processing program 56 from the memory 32 and executes the read specific processing program 56 on the RAM 30. Specific processing is implemented by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.
[0049] The memory 32 stores a data generation model 58 and an emotion-specific model 59. The data generation model 58 and the emotion-specific model 59 are used by the specific processing unit 290. The specific processing unit 290 can use the emotion-specific model 59 to infer the user's emotions and perform specific processing based on the user's emotions. In the emotion inference function (emotion-specific function) using the emotion-specific model 59, various inferences and predictions related to the user's emotions are performed, including inferences and predictions of the user's emotions, but this is not limited to this example. Furthermore, emotion inference and prediction may also include, for example, emotion analysis (parsing).
[0050] In the smart device 14, the processor 46 performs the acceptance output processing. The memory 50 stores the acceptance output program 60. The acceptance output program 60 is used in conjunction with the data processing system 10 and the specific processing program 56. The processor 46 reads the acceptance output program 60 from the memory 50 and executes the read acceptance output program 60 on the RAM 48. The acceptance output processing is implemented by the processor 46 acting as the control unit 46A according to the acceptance output program 60 executed on the RAM 48. Furthermore, the smart device 14 has the same data generation model and emotion-specific model as the data generation model 58 and the emotion-specific model 59, and these models can also be used to perform the same processing as the specific processing unit 290. The acceptance output processing is implemented by the processor 46 acting as the control unit 46A according to the acceptance output program 60 executed on the RAM 48.
[0051] In addition, a device other than the data processing device 12 can have the data generation model 58. For example, a server device (for example, a generation server) can have the data generation model 58. In this case, the data processing device 12 obtains a processing result (a prediction result, etc.) using the data generation model 58 by communicating with the server device having the data generation model 58. Furthermore, the data processing device 12 can be a server device, or can be a terminal device (for example, a mobile phone, a robot, a household appliance, etc.) held by a user. Next, an example of the processing of the data processing system 10 of the first embodiment will be described.
[0052] Embodiment 1
[0053] A flow of the specific processing in Embodiment 1 will be described. Each part of the system described below is implemented by the data processing device 12 and the intelligent device 14. Furthermore, the data processing device 12 is referred to as a “server”, and the intelligent device 14 is referred to as a “terminal”.
[0054] In the existing home environment, cleaning and maintenance often requires a lot of time and effort, especially the difficulty in finding dirt or pests leads to delayed cleaning response. In addition, due to the lack of targeted cleaning method and tool information, users often cannot efficiently and appropriately deal with environmental health problems in the home. Therefore, the prior art cannot realize early automatic detection, intelligent notification and personalized cleaning guidance of the abnormality of the home environment, resulting in low efficiency of the cleaning work in the home and difficulty in maintaining the health condition.
[0055] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Embodiment 1 is implemented by the following means.
[0056] In the present application, the server includes a shooting component for periodically acquiring image information, a preprocessing unit for denoising and image transformation of the acquired image, an image analysis unit for image analysis using a convolutional artificial intelligence model, an information recording unit for recording the detection result and its evaluation value to an information storage device, an information generation unit for generating a notification content when the abnormal situation evaluation value exceeds the threshold value, a communication unit for sending the notification content to an information display terminal through wireless communication, a display unit for displaying the notification content on the terminal, and a dialogue assistance unit for receiving user inquiries and automatically generating answers using a generative artificial intelligence model and displaying. Thus, the automatic, high-precision detection and early notification of abnormal situations (such as dirt, pests, etc.) in the home environment can be realized, and real-time cleaning method suggestions can be provided to users, achieving efficient and intelligent management of the home environment.
[0057] The “shooting component” refers to an image collector capable of acquiring real-time image information in a home or target environment, including but not limited to a camera or an image sensing device.
[0058] "Image acquisition unit" refers to a device or functional module for receiving image information from a shooting component in a periodic and continuous manner.
[0059] "Preprocessing unit" refers to a device or program module that performs preliminary processing such as noise reduction, resolution conversion, color space conversion, etc. on acquired image information, usually implemented through image processing algorithms.
[0060] "Image processing component" refers to a hardware or software component specifically designed to perform image data processing operations, such as CPU, GPU, FPGA, or corresponding image processing software.
[0061] "Convolutional artificial intelligence model" refers to an artificial intelligence model based on a convolutional neural network (CNN) structure, used to extract and analyze image features, and implement algorithms for target detection, recognition, etc.
[0062] "Machine learning algorithm" refers to a computer algorithm that can learn from data and automatically improve analysis capabilities, used for pattern recognition and inference processing of image data, etc.
[0063] "Image analysis unit" refers to a processing module that performs intelligent analysis and discrimination on preprocessed image information, efficiently detecting the state or abnormal condition of target objects.
[0064] "Information storage device" refers to a storage hardware or database system that can record and save image, detection results, time information, location, and other related data.
[0065] "Information recording unit" refers to a functional unit that sequentially writes detection results and related parameters into the information storage device.
[0066] "Evaluation value" refers to a quantitative score based on artificial intelligence model output results, indicating the confidence or severity of the abnormal state of the target object.
[0067] "Information generation unit" refers to a module that automatically generates notification or operation instruction content based on detection records and evaluation values.
[0068] "Communication unit" refers to a component that transmits information such as notification content to an information display terminal through wired or wireless communication methods.
[0069] "Information display terminal" refers to a terminal device such as a smartphone or tablet that can receive information from a server and present notifications and interactive interfaces to users.
[0070] "Display unit" refers to a module that displays notification content and other information on the information display terminal.
[0071] A "user" refers to a family member or user who uses the system to receive notifications, ask questions, and perform operations according to the guidance.
[0072] "Question text" refers to the natural language description input by the user to obtain answers related to cleaning methods, object information, etc.
[0073] "Generative artificial intelligence model" refers to an artificial intelligence algorithm model capable of automatically generating natural language answers or content based on input prompts, such as large language models.
[0074] "Prompt sentence" refers to the text content input as a generative artificial intelligence model to guide it to generate answers on a specific topic.
[0075] "Dialogue assistance unit" refers to a functional unit for receiving user questions, calling a generative artificial intelligence model to automatically generate answers, and displaying them.
[0076] To facilitate the accurate implementation of the present application by others, the embodiments of the present application are described in detail below in conjunction with specific hardware, software structure and usage.
[0077] The system described in the present application is mainly composed of a server, a terminal and a user.
[0078] The server can be deployed on a high-performance computing device, such as a cloud server or a local microcomputer. The server connects multiple cameras (such as general network cameras) deployed in various parts of the home. The cameras transmit real-time or timed image streams to the server through the home wireless local area network. The server uses image processing software (such as OpenCV) to preprocess the received image data, including noise removal (Gaussian filtering), resolution conversion (such as uniform to 1280x720), color space conversion (such as RGB to grayscale). These operations can effectively reduce the computational burden of subsequent analysis, improve detection accuracy and speed.
[0079] After preprocessing, the server locally or calls cloud services to intelligently analyze the image data based on convolutional neural networks (such as CNN models developed using TensorFlow or PyTorch). The server uses machine learning algorithms to automatically detect abnormal conditions such as dirt, garbage, and insect damage in the image, and generates a confidence score for each detection. The server writes the detected abnormal categories, locations, timestamps, confidence scores, and other information into a database system (such as MySQL) for subsequent tracking and management.
[0080] The server monitors the latest detection results in the database in real time or at regular intervals. For abnormal detections with high confidence and exceeding a set threshold, the server automatically generates prompt notification content. The information generation module composes the content in the form of a message text, including cleaning suggestions, danger warnings, etc. The server sends the notification content to the terminal, such as the user's smartphone or tablet, through a wireless communication module, such as through a mobile message push service.
[0081] The terminal is responsible for receiving and displaying the notification messages issued by the server. Through the mobile phone APP or special program, the user can obtain detailed warning content in real time in the notification bar. For example, the phone screen will display: "Stains found in kitchen drain, please clean up in time." The terminal also provides a chat robot interface for the user, who can directly input natural language questions in the interface.
[0082] After receiving the cleaning or abnormal prompt, the user can respond in a timely manner according to the notification content. For example, after receiving information about the dirt in the kitchen drain, the user can prepare a brush, cleaning agent, and other tools and immediately handle it. In addition, the user can input inquiry text (prompt statements) in the APP chat interface, such as "How to clean the drain?", "How to remove the kitchen odor?", "What to do if there are small insects at home?" etc. The terminal uploads the user's natural language questions to the server.
[0083] After receiving the user's inquiry, the server uses an integrated generative artificial intelligence model (such as based on GPT-3 or other large language models) to understand the semantics and generate content for the question. The server automatically builds detailed and personalized operation guides or answers, such as "The cleaning method for the drain is as follows: 1. Remove the drain cover; 2. Use a brush to remove dirt and debris; 3. Use disinfectant to clean and thoroughly flush water." The server sends the answer to the terminal through wireless communication. The user can receive detailed responses in the APP chat interface and follow the suggested operations.
[0084] The generative artificial intelligence model introduced in this invention can automatically generate high-quality personalized answers based on different "prompt statements", which greatly helps to improve home cleaning and environmental intelligent management.
[0085] Examples of prompt statements are as follows:
[0086] "What is the cleaning method for the drain?"
[0087] "How to clean the kitchen drain with simple tools?"
[0088] "What to do if there are small insects in the bathroom?"
[0089] "Please give me detailed instructions on how to disinfect the kitchen countertop."
[0090] Usage Figure 11 The process of handling is described.
[0091] Step 1:
[0092] The server receives real-time or scheduled image data from cameras deployed within the home.
[0093] The input is video stream data transmitted by each camera over the network, and the output is raw image data stored locally on the server. The specific actions include receiving one or more frames of pictures taken by multiple cameras at regular intervals through the network interface and saving them to the local or temporary cache.
[0094] Step 2:
[0095] The server pre-processes the received raw image data.
[0096] The input is raw image data, and the output is image data after noise reduction, resolution conversion, and color space conversion. The server uses image processing libraries (such as OpenCV) to perform Gaussian filtering to remove noise, unify the resolution to a predetermined size (such as 1280x720), and convert color images to grayscale images. The specific actions are to call the program to sequentially complete the above processing for each frame of picture.
[0097] Step 3:
[0098] The server uses artificial intelligence models such as convolutional neural networks to analyze the pre-processed image data.
[0099] The input is pre-processed image data, and the output is detection results, including abnormal types (such as dirt, pests), abnormal locations, and confidence scores. The server inputs the processed pictures into the AI model, which automatically identifies abnormal areas and outputs analysis reports. Specific actions include: the model detects pollution in the kitchen drain and gives a confidence value of 95%.
[0100] Step 4:
[0101] The server records the detection results and related information of the AI model into the database.
[0102] The input is detection result data (abnormal category, location, confidence, detection time, etc.), and the output is a new abnormal record stored in the database. The specific actions are to call the database interface to write records for all abnormal information obtained through analysis, including the geographical or camera location where the anomaly occurred, the timestamp, and the confidence score.
[0103] Step 5:
[0104] The server generates notification content based on the detection results in the database and the set threshold.
[0105] Input: latest high-confidence anomaly detection records in the database, output: generated notification message text. The server filters out anomaly events with confidence exceeding a certain threshold (such as 90%), and automatically writes notification information to remind users to clean up or pay attention in time. Specific actions such as generating a prompt text "stains found in kitchen drain, please clean up in time."
[0106] Step 6:
[0107] The server pushes the generated notification message to the terminal device through wireless communication.
[0108] Input: notification text to be pushed, output: reminder information received by the terminal device. The server calls the push service (such as APP notification channel), and transmits the message to the user's registered mobile device through the network. The specific action is that the server calls the message push API and sends the information in real time.
[0109] Step 7:
[0110] The terminal receives and displays the notification information on the user's terminal.
[0111] Input: push notification from the server, output: reminder display on the terminal screen. The terminal automatically pops up a message prompt through the APP or applet, such as displaying the corresponding content in the notification bar. Specific actions such as popping up a prompt "stains found in kitchen drain, please clean up in time." on the phone screen.
[0112] Step 8:
[0113] The user completes the actual operation according to the notification content, or inputs questions in the chat robot interface of the terminal.
[0114] Input: user's natural language question (such as "how to clean the drain?"), output: inquiry text sent to the server. After reading the notification, the user performs cleaning and other actions, and can input questions related to cleaning in the APP dialogue box. The specific action is that the user types his own questions on the mobile terminal.
[0115] Step 9:
[0116] The terminal sends the user's input inquiry text to the server in real time.
[0117] Input: user input question text, output: inquiry text received by the server. The terminal uploads the text to the server through the network interface. The specific action is that the APP background packages the text into a request and sends it to the server's dialogue processing interface.
[0118] Step 10:
[0119] The server uses generative artificial intelligence model to understand the inquiry text and generate answers.
[0120] The input is the original text of the user's inquiry (prompt sentence), and the output is the detailed answer text automatically generated by the system. The server calls a large language model (such as GPT-3) to analyze the input question and generate a detailed answer to the specific household problem. Specific actions such as generating "Drain cleaning method: 1. Remove the drain cover; 2. Use a brush to remove dirt; 3. Rinse with disinfectant."
[0121] Step 11:
[0122] The server sends the generated answer through wireless communication to the terminal.
[0123] The input is the generated answer text, and the output is the message content received by the terminal. The server calls the interface to return the answer data to the user device. Specific actions such as uploading the content of "Drain cleaning method:..." to the terminal APP.
[0124] Step 12:
[0125] The terminal displays the answer returned by the server in the chat interface.
[0126] The input is the answer text pushed by the server, and the output is the reply content displayed on the APP interface. The terminal automatically displays the answer in the chat window, and the user can read the detailed step-by-step instructions and operate accordingly. The specific action is to refresh the phone interface and display "Drain cleaning method: 1. Remove the drain cover...".
[0127] Application Example 1
[0128] The flow of specific processing in Application Example 1 is described below. Each part of the system described below is implemented by the data processing device 12 and the intelligent device 14. In addition, the data processing device 12 is referred to as "server", and the intelligent device 14 is referred to as "terminal".
[0129] Existing home monitoring and cleaning systems can usually only achieve basic video monitoring or local pollution detection, lack efficient identification and automatic response to abnormal behavior (such as suspicious activity), and cannot provide personalized and humanized safety and cleaning suggestions based on the user's real-time emotional state. In addition, existing systems often lack natural language-based intelligent interaction with users, and users cannot easily obtain exclusive recommendations for the current home situation. Therefore, how to realize a home intelligent management system that combines multi-target detection, abnormal behavior identification, emotion perception, and intelligent dialogue functions is a technical problem that needs to be solved in this field.
[0130] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is implemented through the following means.
[0131] In the present application, the server comprises a collection unit for receiving video data in streaming form periodically, a processing unit for image pre-processing of the video data, an analysis unit for implementing anomaly behavior, pollution or biological detection using a machine learning model, a recording unit for recording detection results and key information, a generation unit for automatically generating and pushing response instruction messages, a sending unit, and a support unit for inferring user emotions according to user state data and adjusting message content accordingly, and is equipped with a dialogue support unit based on a generative artificial intelligence model, which can perform natural language intelligent interaction. Thus, automatic monitoring, problem identification, data recording, intelligent notification, and emotional and personalized interactive response of the home environment can be realized, significantly improving the automation level and user experience of home intelligent management.
[0132] The "collection device" refers to an electronic device capable of collecting video data in a designated environment and outputting in streaming form, such as a camera or other image collection unit.
[0133] The "collection unit" refers to a functional module or component for periodically receiving video stream data from the collection device and managing data input.
[0134] The "processing unit" refers to a functional module or component for performing image denoising, resolution adjustment, color space transformation, and other pre-processing operations on the received video data.
[0135] The "machine learning model" refers to a data processing algorithm or artificial intelligence system that can automatically analyze and identify targets such as abnormal behavior, objects, and living beings in video images based on a large amount of training data.
[0136] The "analysis unit" refers to a functional module or component that uses a machine learning model to detect targets and identify abnormal behavior in pre-processed video data.
[0137] The "recording unit" refers to a functional module or component that stores information such as detection content, confidence, collection source, and time in an electronic storage medium.
[0138] The "generation unit" refers to a functional module or component that automatically generates message content such as preset notifications, response instructions, and processing suggestions based on detection results.
[0139] The "sending unit" refers to a functional module or component that pushes generated instructions, notifications, and other information to external information terminals through network communication.
[0140] The "information terminal" refers to a user device such as a smartphone or tablet computer that can receive and display messages and instructions sent by the server and interact with the server.
[0141] The "display unit" refers to a functional module or component that displays the received content to the user in the form of a pop-up window, a reminder, an interface message, etc. on the information terminal.
[0142] The "generative artificial intelligence model" refers to an artificial intelligence algorithm system that can automatically generate natural language replies or content based on deep learning technology, etc.
[0143] The "dialogue support unit" refers to a functional module or component that receives user query information, automatically generates personalized reply content through a generative artificial intelligence model, and displays it.
[0144] The "user state data" refers to data that can reflect the user's emotional state, behavior, or health status, which is collected in real time through sensors, terminals, or collection devices.
[0145] The "support unit" refers to a functional module or component that analyzes user state data, infers user emotional state, and adjusts message or reply content accordingly.
[0146] The present application relates to an artificial intelligence-based video monitoring and home management system, which cooperates with the server, terminal and user, realizes comprehensive monitoring and personalized interaction of pollution, pests, abnormal behavior, etc. in the home environment through multi-module integration.
[0147] The server receives video data periodically through a collection device (such as a network camera or other hardware) using a streaming media protocol such as RTSP. The server uses general-purpose computing hardware (such as an X86 server or an embedded controller) in combination with an image processing software library (such as OpenCV) to preprocess the received video data, including noise removal, resolution conversion, color space transformation, etc. After preprocessing, the server inputs the preprocessed video data into a machine learning platform (such as TensorFlow or PyTorch) to detect stains, pests, and human abnormal behavior in the picture using a pre-trained deep learning model (convolutional neural network CNN).
[0148] The detected content, along with associated information such as confidence, collection time, source location, etc., is stored in a database system (such as MongoDB or PostgreSQL) by the server. The server monitors the database content in real time, and once the confidence of the identified pollution or abnormal situation exceeds a certain threshold, it automatically generates a specific response message. The message content is dynamically formed by the server using a text template or a generative artificial intelligence model (such as GPT-3, BERT, etc.) in combination with the existing situation. The server sends the notification information to the user's terminal device (such as a smartphone or tablet) in real time through a push message using a communication module to call a push service (such as Firebase Cloud Messaging).
[0149] After the terminal device receives the push notification, it is automatically displayed to the user in the form of a pop-up window, a notification bar message, etc. through the application interface. The user can actively query the cleaning method, security suggestions, etc. through the chat interface provided by the terminal and input natural language questions. The terminal uploads the user input content to the server in real time, the server uses the generative artificial intelligence model to analyze and understand the user's intention, generates detailed and personalized response content based on the existing home state and specific prompts, and feeds back to the user through the terminal.
[0150] The server is also integrated with an emotion recognition module, which can use deep learning emotion analysis models such as emotion classification BERT and EmotionCNN to determine the user's current emotional state (such as fatigue, anxiety, fear, etc.) through user images or audio data recorded by the collection device, and dynamically adjust the push notification or AI dialogue reply wording according to the state, realizing humanized and personalized information interaction.
[0151] Specific examples are as follows:
[0152] For example, the user's kitchen camera detects a noticeable stain on the drain, the server completes image preprocessing using OpenCV, loads a deep learning model using TensorFlow to identify the location of the oil stain, and stores the detection details in MongoDB. After the server determines that it needs to be cleaned, it automatically generates a push notification to the user's phone: "There is an oil stain on the kitchen drain. Please clean it in time." After receiving the reminder, the user can ask "How to thoroughly clean the drain?" through the chat interface on the phone, and the server uses the GPT-3 model to generate "In three steps: first, open the drain cover; second, clean the residue with a brush; third, use disinfectant." as a response.
[0153] The server can also analyze the tone and content of the user's voice, and if it identifies that the user is in a state of fatigue, the server will push the content to become "You seem to be quite tired now. If you need to, you can handle the cleaning task later."
[0154] The prompt statement of the generative artificial intelligence model is, for example:
[0155] "Home kitchen camera detects suspicious person. Please provide detailed instructions on the best safety handling process."
[0156] "Please provide step-by-step instructions on how to efficiently clean the bathroom drain."
[0157] "If the user is in a low mood, how to encourage and care for them to complete the home cleaning?"
[0158] "If there are pests such as cockroaches and ants in the home, how to handle and prevent pest infestation?"
[0159] Through the above-mentioned hardware and software integration mode and data processing flow, the system can not only actively identify and remind the family cleaning and security problems, but also can provide intelligent and dynamic human-computer interaction experience according to the real-time state and needs of the user, greatly improving the convenience and safety of family management.
[0160] Using Figure 12 The flow of processing is described.
[0161] Step 1:
[0162] The server receives real-time video streams within the home from multiple acquisition devices (cameras) through network protocols (such as RTSP) at regular intervals. The input is the raw video data stream sent by each acquisition device in real time, and the output is the raw video frame sequence cached locally by the server. The server establishes network connections for cameras in different areas (such as the kitchen, living room, and bathroom) to achieve full-range data collection.
[0163] Step 2:
[0164] The server uses image processing libraries (such as OpenCV) to preprocess the received video frames. The input is the raw video frame, and the server performs noise removal, resolution scaling, color space conversion, and other data processing on the video frame. The output is the preprocessed video frame after format unification and noise removal, which is used for subsequent AI analysis. Specific operations include applying Gaussian filtering, image scaling, RGB to grayscale, etc.
[0165] Step 3:
[0166] The server inputs the preprocessed video frame into a deep learning model (such as a convolutional neural network on TensorFlow). The input is the processed video frame, and the server detects whether there are objects or events such as stains, pests, and abnormal behaviors in the picture. Data processing includes feature extraction and recognition, and the output is a list of detection results, including object type, detection confidence, frame position, etc.
[0167] Step 4:
[0168] The server writes the detailed results of each detected object or event (such as object category, confidence, camera position, and timestamp) into a database (such as MongoDB). The input is the detection result data output by the AI model, and the data operation includes persistently saving structured data. The output is the newly added record entry in the database.
[0169] Step 5:
[0170] The server periodically queries the latest detection data in the database to determine whether the confidence of any detection object exceeds the preset threshold. The input is the detection result record in the database, and the server performs conditional filtering and logical judgment on the data. The output is a list of events that need to be notified, such as "there is oil stain in the kitchen drain" or "unknown person appears".
[0171] Step 6:
[0172] The server generates a notification message. The input is the list of events that need to be notified, and the server forms specific notification content according to the event type and situation text template, or calls generative artificial intelligence models (such as GPT-3 / BERT). Data processing includes automatic generation of message text. The output is one or more message contents to be pushed.
[0173] Step 7:
[0174] The server sends the generated notification or instruction to the terminal through the push service (such as Firebase Cloud Messaging). The input is the message content to be pushed and the target terminal information, and the data calculation is to package the message and call the push API. The output is the push message received on the user terminal.
[0175] Step 8:
[0176] The terminal receives the push message from the server and immediately displays it to the user through interface pop-up, notification bar, etc. The input is the message package pushed by the server, and the terminal system parses the message and triggers local reminder operation. The output is the notification interface visible on the user device.
[0177] Step 9:
[0178] The user actively inputs queries through the dialogue interface of the terminal, such as "how to clean the drain" or "what to do if suspicious person is found". The input is the user's natural language input text, and the terminal packages the content and uploads it to the server. The output is the user query data package received by the server.
[0179] Step 10:
[0180] After receiving the user query, the server uses generative artificial intelligence models (such as GPT-3) to analyze and generate responses to the input text. The input is the user query text and the current home detection data, and the data processing is natural language understanding and generation. The output is a detailed and personalized reply text.
[0181] Step 11:
[0182] The server sends the generated reply text back to the terminal. The input is the AI-generated response message and the target terminal information, and the data calculation is the message communication protocol packaging and sending. The output is the new message received by the terminal.
[0183] Step 12:
[0184] The terminal displays the response content returned by the dialogue interface display server. The input is the text data returned by the server, and the terminal displays the interface after parsing the content. The output is an AI reply readable by the user, which provides guidance for the user's home management and cleaning behavior.
[0185] Step 13:
[0186] The server runs an emotion analysis model by analyzing information from the collection device or terminal (such as audio, video, sensor data), and infers the user's emotional state. The input is real-time user state data, and the data processing includes emotion feature extraction and classification. The output is the user's current emotional label, such as "fatigue", "anxiety", "joy", etc.
[0187] Step 14:
[0188] The server adjusts the notification language and intelligent question and answer content based on the user's emotional state obtained in step 13. The input is the emotional label and the original message text, and the data processing includes information rewriting or style adjustment. The output is a more personalized, caring, and personalized push message or AI reply.
[0189] In addition, a sentiment engine for inferring the user's emotions can also be combined. That is, the specific processing unit 290 can also use the emotion-specific model 59 to infer the user's emotions and perform specific processing using the user's emotions.
[0190] Embodiment 2
[0191] The flow of the specific processing in Embodiment 2 is described. Each part of the system described below is implemented by the data processing device 12 and the intelligent device 14. In addition, the data processing device 12 is referred to as "server", and the intelligent device 14 is referred to as "terminal".
[0192] The existing home cleaning assistance system has limitations in automatically detecting stains or pests, and cannot intelligently adjust the notification content according to the user's current emotional state and actual needs, enhancing the user experience. In addition, after receiving the reminder, the user often needs to manually query the cleaning method, lacking intelligent interaction and personalized suggestions, resulting in insufficient humanization and convenience of the system, and cannot effectively reduce the user's psychological burden and actual operation pressure.
[0193] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Embodiment 2 is implemented by the following means.
[0194] In the present application, the server comprises a video acquisition device for periodically receiving video data, an image acquisition unit, a preprocessing unit, an image analysis unit for detection using a generative artificial intelligence model, a result recording unit, a notification generation unit, a notification sending unit, a display unit, a dialogue processing unit, and an emotion analysis unit. Thereby, real-time and automatic detection of stains or organisms within the home can be achieved, and personalized and emotional prompt content can be generated and sent. Users can ask questions and obtain intelligent answers through the interactive terminal, and the system can automatically adjust the notification language based on the user's emotional state, achieving efficient, convenient, and personalized home cleaning assistance experience.
[0195] The "video acquisition device" refers to a hardware device for periodically acquiring or continuously receiving image data in a monitored environment, such as a camera installed in the home.
[0196] The "image acquisition unit" refers to a software or hardware module for receiving and storing image data from the video acquisition device in a streaming format.
[0197] The "preprocessing unit" refers to a software or hardware module that performs preliminary data processing on the received image data, such as noise reduction, resolution adjustment, color space conversion, etc.
[0198] The "generative artificial intelligence model" refers to a deep learning network that can automatically generate results or answers based on input images or data, typically including but not limited to convolutional neural networks.
[0199] The "image analysis unit" refers to a software and hardware component that uses a generative artificial intelligence model to perform target detection, classification, and recognition on preprocessed image data.
[0200] The "result recording unit" refers to a module that stores the detection results, confidence scores, time information, device location information, etc. of the image analysis unit in an information storage device.
[0201] The "notification generation unit" refers to a module that automatically generates message content or prompt statements based on detection results and data rules.
[0202] The "notification sending unit" refers to a module that pushes or sends generated notification messages to information communication devices through a network.
[0203] The "display unit" refers to an interface or component that displays notification content on an information communication terminal in a real-time manner to the user.
[0204] The "dialogue processing unit" refers to a module that receives user queries, understands them based on natural language processing models, and automatically generates answer content for output.
[0205] The "emotion analysis unit" refers to a software and hardware module that analyzes the user's voice data or image data to determine the user's emotional state and adjust the expression of the prompt content accordingly.
[0206] Embodiments of the present application are described as follows.
[0207] The present application relates to an automatic cleaning assistance system based on a generative artificial intelligence model. The system realizes intelligent and automatic stain or biological detection, notification generation, user interaction and emotional adaptation in a home environment through the cooperation of a server, a terminal and a user.
[0208] The server is installed in a local or cloud data center and serves as the main processing unit of the system. The server configuration includes a video acquisition device for regularly receiving image data collected by a home camera, an image acquisition unit for realizing stream media reception and storage, and a preprocessing unit for image denoising and format conversion. The server mainly runs the OpenCV software library to realize Gaussian filtering processing, resolution scaling (e.g., from 1080p to 720p), and color space conversion (e.g., from RGB to grayscale) of images.
[0209] The preprocessed image data is input to the generative artificial intelligence model on the server for analysis. The generative artificial intelligence model used in the present application has a convolutional neural network structure and is deployed in combination with the TensorFlow deep learning framework. This model can detect target objects such as stains or pests in each frame of image and output the detection category and confidence score. All detection results and related information (including timestamp, device location information, etc.) are recorded through the information storage device in the server, and general database management software such as MySQL database can be used.
[0210] When the server finds that the confidence score exceeds the preset threshold, the notification generation unit will automatically generate personalized prompt content according to the specific circumstances detected. To achieve efficient interaction with the user, the server sends messages to the user terminal device through a push service (such as the general push API Firebase Cloud Messaging). The terminal device can be an information communication device such as a smartphone or tablet computer. The application program on the terminal integrates message notification display, which can show cleaning suggestions and abnormal reminders to the user in the form of pop-up windows, message bars, etc.
[0211] After receiving the cleaning suggestion notification, the user can directly input questions in text or voice input mode through the interface on the terminal. For example, the user can input "What are the cleaning methods for the drain?" The conversation processing unit on the server side will input the user's natural language question into a natural language processing model (such as BERT, implemented through the Transformers open source library), and automatically generate a structured and step-by-step answer, which is returned to the terminal to be displayed to the user.
[0212] The emotion analysis unit of the present application can perform emotion recognition on the user's voice data or image data. The server analyzes the user's current emotional state using a speech-to-text API (such as a general speech recognition API) combined with an emotion judgment algorithm (such as Emotion API), and automatically adjusts the expression of the prompt content accordingly. For example, when the user is identified as "tired", the system will generate a prompt sentence expressing concern or delaying the suggestion, such as "Today you seem to be more tired, you can do cleaning later."
[0213] The system can also generate AI prompt sentences as follows:
[0214] Please generate a notification statement based on the following situation: the camera detects that the kitchen drain has stains, and the user's emotional state is tired. Please add a gentle suggestion and encouragement in the notification.
[0215] The hardware platform used above can adapt to general home network environment, and the software environment is based on common image processing and artificial intelligence framework, which is convenient for actual deployment and expansion. The whole system process takes the server as the core processing center, and the terminal as the information interaction bridge. Users can obtain cleaning suggestions and technical support in a natural and personalized way, improving the automation and intelligence level of home health management.
[0216] Usage Figure 13 The process is described.
[0217] Step 1:
[0218] The server regularly receives image data streams from video capture devices deployed inside the home. The input is multiple camera-captured H.264 encoded image data streams. The server captures raw video frames in real time through network ports and stores them in local cache. The output is raw streaming image data.
[0219] Step 2:
[0220] The server performs preprocessing operations on the received raw image data. The input is raw streaming image data, and the output is preprocessed image data. The server uses OpenCV library to perform Gaussian filter denoising on each frame of image, performs resolution scaling (such as reducing 1080p to 720p), and converts RGB color image to grayscale image. This processing improves the accuracy and speed of subsequent AI analysis.
[0221] Step 3:
[0222] The server inputs the pre-processed image data into a generative artificial intelligence model for analysis. The input is the pre-processed image data, and the output is the detected target and confidence score (e.g., stain type, pest type, and their confidence). The server uses a convolutional neural network under TensorFlow to perform target detection on each frame of image, and labels each detection result and calculates the confidence score.
[0223] Step 4:
[0224] The server records the AI detection results and related additional information into an information storage device (e.g., a database). The input is the detection label, confidence, image collection timestamp, camera location, and other metadata. The output is a structured record written to the database. The server calls the database interface to insert these data into the corresponding data table for subsequent retrieval and query.
[0225] Step 5:
[0226] The server retrieves the latest detection results from the database at regular intervals and determines whether their confidence exceeds a pre-set threshold. The input is the latest detection record in the database. The output is the detection information that needs to be reminded. The server filters high-confidence data and extracts the corresponding monitoring location and type, and prepares to generate a user reminder if necessary.
[0227] Step 6:
[0228] The server automatically generates targeted prompt sentences based on the detection results and set templates. The input is the detection information that needs to be reminded, such as "stains detected at kitchen drain." The output is a structured notification message. The server splices or generates personalized notification text based on the detection location and category, and can combine sentiment judgment for appropriate expression processing.
[0229] Step 7:
[0230] The server sends the generated reminder content to the user terminal using a push message service (e.g., a general push API). The input is the notification message, and the output is the message that has been pushed to the terminal. The server calls the push service API to distribute the message to the user's smartphone or tablet application, and confirms the message status.
[0231] Step 8:
[0232] The terminal receives the notification message and displays it in real time on the user interface. The input is the notification information pushed down, and the output is the message displayed on the user interface. The terminal triggers a pop-up window, notification bar content refresh, or application message list update to remind the user to pay attention.
[0233] Step 9:
[0234] The user inputs a query through the terminal interface, such as requesting a specific cleaning method. The input is the user's text or voice question on the terminal, and the output is the question data to be uploaded. The user can ask cleaning-related questions through text input or voice-to-text function.
[0235] Step 10:
[0236] The terminal sends the user's input question to the server in real time. The input is the user's question content, and the output is the request successfully uploaded to the server. The terminal calls the communication interface, packages the user's question as JSON or structured request, and uploads it to the server API.
[0237] Step 11:
[0238] After receiving the user's question, the server calls the natural language processing module (such as the BERT model) to analyze the question content and automatically generate detailed and step-by-step answers. The input is the user's question text and environmental context information, and the output is the step-by-step or detailed answer. The server uses intelligent question and answer models to analyze the problem intent, retrieve the knowledge base, and return targeted answers.
[0239] Step 12:
[0240] The server performs sentiment analysis on the user's voice or image data. If the user uploads voice information, the server uses voice recognition and sentiment determination algorithms to evaluate the state. The input is the user's voice or video segment, and the output is the user's emotional state label (such as fatigue, happiness, anxiety, etc.). The server adjusts the wording of the subsequent generated notification information based on the sentiment analysis results, such as adding appropriate care or encouragement content.
[0241] Step 13:
[0242] The server optimizes the final notification or suggestion by combining detection conditions, user interaction content, and emotional state, and pushes it to the terminal. The input is the detection result, user state, and push template, and the output is the final push prompt statement. The server ensures that each piece of information is personalized and emotional, improving the intelligent service level of the system.
[0243] Step 14:
[0244] The terminal displays the final notification or answer information pushed by the server to the user in real time. The input is the final prompt statement or answer, and the output is the message content displayed on the user interface, waiting for the user's next operation or feedback to achieve a closed-loop interaction.
[0245] Application Example 2
[0246] The flow of the specific processing in Application Example 2 will be described. Each part of the system described below is implemented by the data processing device 12 and the smart device 14. In addition, the data processing device 12 is referred to as a "server", and the smart device 14 is referred to as a "terminal".
[0247] In existing actual sites, it is difficult to provide personalized service responses according to the immediate emotional state of the user, resulting in a decrease in customer satisfaction. In addition, there is a lack of a system that can analyze the emotional state in real time based on multi-modal information such as user expressions, voices, etc., and push personalized notifications to staff, and the overall customer experience quality cannot be improved.
[0248] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is implemented by the following means.
[0249] In the present application, the server includes: an imaging device and a video stream information acquisition unit for periodically acquiring image information, an information analysis unit for individual recognition and emotion analysis using an artificial intelligence model, an information recording unit for recording analysis results together with confidence, time and space information, a generation processing unit for automatically generating personalized notifications according to emotional state, a notification sending unit for displaying notifications in real time on wearable terminals, a prompt sentence generation unit for automatically generating prompt sentences of the generative artificial intelligence model based on user behavior and historical analysis information, and a dialogue response unit for analyzing user queries through a natural language model and providing semantic responses. Thereby, the customer emotional state can be recognized in real time, and personalized service notifications can be automatically generated and pushed to staff according to the analysis results, effectively improving customer satisfaction and service quality.
[0250] The "imaging device" refers to an information acquisition hardware device capable of image acquisition of a target scene and output of image information or a video stream.
[0251] The "information acquisition unit" refers to a functional module for periodically acquiring image information and inputting to the system in the form of streaming media, etc.
[0252] The "audio information" refers to data related to user voice content and environmental sound acquired by a sound acquisition device.
[0253] The "information analysis unit" refers to a processing module that uses an artificial intelligence model to preprocess, recognize individuals, and analyze emotional states of acquired image information and audio information.
[0254] The "artificial intelligence model" refers to a software and hardware integrated system capable of performing pattern recognition, classification, reasoning, etc. algorithms on input multi-modal information, including but not limited to deep learning models.
[0255] "Individual identification" refers to a processing method for identifying and distinguishing specific individuals based on characteristics such as images, sounds, etc.
[0256] "Emotional state analysis" refers to the process of determining the internal emotional state of a user by analyzing information such as facial expressions, speech, etc.
[0257] "Information recording unit" refers to a device or functional module that stores the results of data processing together with confidence scores, time information, and spatial information in a database.
[0258] "Time information" refers to timestamp, date, and other identification information associated with data collection, processing, and other steps.
[0259] "Space information" refers to relevant location information about imaging devices or collection locations.
[0260] "Analysis history information" refers to a collection of all or part of the analysis results, confidence, time, and location-related information recorded in the system.
[0261] "Generation processing unit" refers to a functional module that automatically generates personalized notification text based on analysis results.
[0262] "Notification sending unit" refers to a functional module that distributes generated notification content to target terminals or display devices.
[0263] "Information prompting unit" refers to a module that can output notification information visually or aurally on a portable information terminal or augmented reality device.
[0264] "Conversational response unit" refers to a functional module that receives user queries, analyzes and generates response content through a natural language processing model, and returns the content to the user through the information prompting unit.
[0265] "Natural language processing model" refers to an artificial intelligence processing algorithm that can structurally understand, semantically analyze, and generate responses to natural language text.
[0266] "Generative artificial intelligence model" refers to an artificial intelligence model that can automatically generate text, suggestions, or notification content based on input prompts.
[0267] "Prompt sentence generation unit" refers to a module that generates prompt sentences for input to a generative artificial intelligence model based on user behavior, expressions, speech content, and analysis history.
[0268] The present invention relates to a user emotional state recognition and personalized service notification system based on multi-modal information of images and audio. To facilitate implementation, the present invention is described in detail below in terms of hardware configuration, software composition, data processing method, and specific application scenarios.
[0269] The system includes multiple terminal devices (such as smart glasses, mobile terminals), one or more servers, and databases and human-machine interfaces. The terminal can be equipped with a camera and a microphone for real-time collection of customer image and audio information (such as facial expressions and speech content). These raw data are transmitted to the server through a wireless network.
[0270] The server has high computing power (recommended configuration of GPU) and is pre-installed with various software resources, including but not limited to OpenCV (for image preprocessing), DeepFace (for face recognition and emotion analysis), natural language processing models (such as NLP services based on large language models), database systems (such as MySQL), etc.
[0271] When the terminal collects image data including the customer's face, the server first uses OpenCV to pre-process the image stream, including noise reduction, resolution adjustment, color space conversion, etc. Then, the server calls the DeepFace model to perform face detection, individual recognition, and emotion state inference on each frame of data, such as recognizing typical emotions such as "anger", "happiness", "surprise", etc., and outputting confidence scores. At the same time, the server can use speech recognition and analysis modules to further analyze the user's intent and emotion based on the recorded audio.
[0272] The analysis results, along with metadata such as time (such as timestamp), camera spatial position, etc., are stored in the server-side database and regularly archived as analysis history data. The server further calls generative artificial intelligence models (such as general large language models similar to GPT-4) based on relevant prompt statements to generate personalized notification text for the staff. For example, "The customer is currently in a bad mood, please actively listen to their demands."
[0273] The server sends these notification texts to the terminal device in real time through the network, and the terminal's display screen or AR display module immediately pops up a window to remind the staff. If the terminal supports speech synthesis, it can also voice the notification content. After receiving the personalized notification, the staff can adjust the service method to achieve the best service that meets the customer's current mood.
[0274] In addition to automatic push, users (such as staff or customers) can also submit free text questions through the terminal interface, such as "What is the customer's current mood?" The server uses natural language processing models to analyze user questions and generates and returns clear answers on the human-machine interface. For example, "The customer is currently in a happy mood, please continue to provide enthusiastic service."
[0275] In a specific implementation, the server can also automatically analyze the user's behavior, expression, speech content and historical data, construct structured or targeted prompt sentences for the generative artificial intelligence model, and use them as model inputs to automatically optimize and upgrade service recommendation strategies.
[0276] Specific examples
[0277] The terminal collects a video clip of the customer frowning during a peak period of human flow. The server determines that the customer is in an anxious state according to image and voice analysis, automatically generates a notification that "the customer's emotion is relatively tense, and it is recommended to speed up the checkout process", and pushes it to the smart glasses worn by the store clerk in real time. The store clerk adjusts the service process according to the prompt to improve customer satisfaction.
[0278] Prompt sentence examples
[0279] Input: Upload customer facial image
[0280] Prompt: Please analyze the expression of the person in the picture and determine their current emotional state, and give specific service recommendations.
[0281] Output: The customer is detected to be angry, and it is recommended to communicate in a gentle tone.
[0282] Input: The customer stops in front of the checkout counter and frowns, what is the current emotion?
[0283] Output: The system determines that the customer is in an anxious state, and it is recommended that the store clerk actively inquire about their needs.
[0284] Through the coordinated operation of the above software and hardware, the system can realize real-time recognition of customer multi-modal emotional state and automatic push of personalized service recommendations, effectively improving the quality of interactive service and customer satisfaction in offline scenarios.
[0285] Use Figure 14 The process of processing is described.
[0286] Step 1:
[0287] The terminal collects image and audio data of the customer in real time through the camera and microphone, and generates video stream and audio stream regularly. The input is the image and sound signal of the on-site customer. The terminal transmits the collected video stream and audio stream to the server through network encryption. The output is real-time video stream and audio stream data packets.
[0288] Step 2:
[0289] The server receives the video stream and audio stream from the terminal, and uses OpenCV to perform image preprocessing on the video frames, including denoising, resolution adjustment, and color format unification. The audio data is standardized. The input is the original video stream and audio stream, and the server performs format conversion, frame processing, and cleaning based on this. The output is preprocessed video frame data and standardized audio data.
[0290] Step 3:
[0291] The server submits the preprocessed video frames to the DeepFace model for face detection and individual recognition, and extracts features from the audio data. The input is the preprocessed video frames and audio data, and the server calls the artificial intelligence model to analyze the faces in the image, determine the identity and infer the expression and emotional state (such as anger, happiness, surprise, etc.), and perform comprehensive analysis of the emotional features of the audio. The output is the identification result, emotional category, and corresponding confidence score of the customer.
[0292] Step 4:
[0293] The server stores the analysis results, confidence scores, collection times, and camera spatial position information in the database. The input is the analysis results and meta information generated in the previous step, and the server performs data structuring and organization storage. The output is a database record containing emotional state, confidence, timestamp, and spatial information.
[0294] Step 5:
[0295] The server automatically generates personalized service suggestions for the staff based on the latest emotional analysis results and historical data. The server calls the generative artificial intelligence model, inputs the customer's current emotional category, confidence, and related background data, and generates natural service suggestion text that fits the scene (such as "Please speed up the service for anxious customers") based on the generated prompt sentences. The output is a personalized notification message suitable for the current scene.
[0296] Step 6:
[0297] The server pushes the generated notification message to the terminal through the network, and the terminal receives and displays it on the display screen (or AR display) in the form of a pop-up window, vibration, or voice prompt. The input is the notification text generated by the server, and the terminal presents it to the staff in various prompt modes. The output is a real-time service reminder or voice broadcast on the terminal.
[0298] Step 7:
[0299] Users (such as shop assistants or customers) can proactively ask free-form text or voice questions (e.g., "What is this customer's current mood?") through the terminal's human-computer interface. The server receives the questions and uses a natural language processing model for semantic analysis and response generation. The input is the user's text or voice question; the server parses it, combines it with current database analysis results to generate a corresponding answer, and then outputs it through the terminal. The output is the automated answer to the question.
[0300] Step 8:
[0301] The server continuously archives historical recognition and interaction data, and uses a prompt generation unit to automatically construct more effective generative AI models to input prompts, thereby optimizing the generation of subsequent personalized notifications and recommendations. Inputs include customer behavior, facial expressions, voice content, and historical database information. After data structuring and transformation, the server outputs highly targeted and context-sensitive prompts.
[0302] The specific processing unit 290 sends the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires sound representing user input regarding the result of the specific processing. The control unit 46A sends the sound data representing user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the sound data.
[0303] Data generation model 58 is a so-called generative AI (Artificial Intelligence). Examples of data generation models 58 include ChatGPT (registered trademark) (accessible via the internet (URL: https: / / openai.com / blog / chatgpt)). Data generation model 58 is obtained through deep learning on a neural network. Input to data generation model 58 are prompt words containing instructions, and inference data such as sound data representing sound, text data representing text, and image data representing images (e.g., still image data or animation data). Data generation model 58 infers from the input inference data based on the instructions represented by the prompt words and outputs the inference result in one or more data forms, such as sound data, text data, and image data. Data generation model 58 includes, for example, text generation AI, image generation AI, and multimodal generation AI. Here, inference refers to, for example, analysis, classification, prediction, and / or induction. The specific processing unit 290 performs the aforementioned specific processing while using data generation model 58. The data generation model 58 can also be a model finely tuned to output inference results from prompts that do not contain instructions. In this case, the data generation model 58 can output inference results based on prompts that do not contain instructions. The data processing apparatus 12, etc., includes various data generation models 58, including AI other than the generation AI. AI other than the generation AI can be, for example, linear regression, logistic regression, decision trees, random forests, support vector machines (SVM), k-means clustering, convolutional neural networks (CNN), recurrent neural networks (RNN), generative adversarial networks (GAN), or Naive Bayes, and can perform various processes, but is not limited to this example. Furthermore, the AI can also be an AI agent. Furthermore, when the processing of the above-mentioned parts is performed by AI, the processing can be partially or entirely performed by AI, but is not limited to this example. Furthermore, the processing performed by the AI including the generation AI can be replaced by processing in the rule base, and the processing in the rule base can also be replaced by processing performed by the AI including the generation AI.
[0304] Furthermore, the processing of the aforementioned data processing system 10 is performed by the specific processing unit 290 of the data processing device 12 or the control unit 46A of the smart device 14, but it can also be performed by both the specific processing unit 290 of the data processing device 12 and the control unit 46A of the smart device 14. Additionally, the specific processing unit 290 of the data processing device 12 acquires or collects information required for processing from the smart device 14 or external devices, and the smart device 14 acquires or collects information required for processing from the data processing device 12 or external devices.
[0305] For example, the collection unit is implemented by the control unit 46A of the smart device 14 or the specific processing unit 290 of the data processing device 12. For example, the acquisition unit uses the camera 42 or communication I / F 44 of the smart device 14 to acquire step data, which is then processed by the specific processing unit 290 of the data processing device 12. For example, the analysis unit is implemented by the specific processing unit 290 of the data processing device 12, which analyzes the data from the collection unit and the acquisition unit. For example, the generation unit is implemented by the specific processing unit 290 of the data processing device 12, which uses a generation AI to generate a menu. For example, the serving unit is implemented by the output device 40 of the smart device 14 or the specific processing unit 290 of the data processing device 12, which provides the generated menu to the user. The correspondence between each unit and the device or control unit is not limited to the above examples and various changes can be made.
[0306] In the above embodiments, examples of specific processing by the data processing device 12 are given, but the technology disclosed herein is not limited to this, and specific processing may also be performed by the smart device 14.
[0307] Second Implementation Method
[0308] Figure 3 An example of the configuration of the data processing system 210 according to the second embodiment is shown.
[0309] like Figure 3 As shown, the data processing system 210 includes a data processing device 12 and smart glasses 214. A server can be cited as an example of the data processing device 12.
[0310] The data processing apparatus 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" as understood in this disclosure. The computer 22 includes a processor 28, RAM 30, and memory 32. The processor 28, RAM 30, and memory 32 are connected to a bus 34. Furthermore, the database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0311] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, RAM 48, and memory 50. The processor 46, RAM 48, and memory 50 are connected to a bus 52. Furthermore, the microphone 238, speaker 240, camera 42, and communication I / F 44 are also connected to the bus 52.
[0312] Microphone 238 receives instructions from user 20 by receiving sounds emitted by user 20. Microphone 238 captures sounds emitted by user 20 and converts the captured sounds into sound data, which is then output to processor 46. Speaker 240 outputs sound according to instructions from processor 46.
[0313] Camera 42 is a small digital camera equipped with an optical system such as a lens, aperture and shutter, and imaging elements such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, to capture images of the user 20's surroundings (e.g., the field of view defined by an angle equivalent to the field of vision of an average healthy person).
[0314] Communication I / F44 is connected to network 54. Communication I / F44 and 26 are responsible for the transmission and reception of various information between processor 46 and processor 28 via network 54. The transmission and reception of various information between processor 46 and processor 28 using communication I / F44 and 26 is performed in a secure state.
[0315] Figure 4 This illustrates an example of the main functions of the data processing device 12 and the smart glasses 214. For example... Figure 4 As shown, in the data processing device 12, specific processing is performed by the processor 28. The specific processing program 56 is stored in the memory 32.
[0316] The specific processing program 56 is an example of a "program" involved in the technology of this disclosure. The processor 28 reads the specific processing program 56 from the memory 32 and executes the read specific processing program 56 on the RAM 30. Specific processing is implemented by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.
[0317] The memory 32 stores a data generation model 58 and an emotion-specific model 59. The data generation model 58 and the emotion-specific model 59 are used by the specific processing unit 290. The specific processing unit 290 can use the emotion-specific model 59 to infer the user's emotions and perform specific processing based on the user's emotions. In the emotion inference function (emotion-specific function) using the emotion-specific model 59, various inferences and predictions related to the user's emotions are performed, including inferences and predictions of the user's emotions, but this is not limited to this example. Furthermore, emotion inference and prediction may also include, for example, emotion analysis (parsing).
[0318] In the smart glasses 214, the processor 46 performs reception and output processing. The memory 50 stores the reception and output program 60. The processor 46 reads the reception and output program 60 from the memory 50 and executes the read reception and output program 60 on the RAM 48. The reception and output processing is implemented by the processor 46 operating as a control unit 46A according to the reception and output program 60 executed on the RAM 48. Furthermore, the smart glasses 214 has the same data generation model and emotion-specific model as the data generation model 58 and the emotion-specific model 59, and these models can also be used to perform the same processing as the specific processing unit 290.
[0319] Next, the specific processing of the specific processing unit 290 of the data processing device 12 will be described. Each part of the system described below is implemented by the data processing device 12 and the smart glasses 214. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".
[0320] Example 1
[0321] The process is the same as that of the specific process described in Embodiment 1 in the first embodiment above, so the description is omitted.
[0322] Application Example 1
[0323] The process is the same as that in the specific processing described in Application Example 1 of the first embodiment above, so the description is omitted.
[0324] Example 2
[0325] The process is the same as that of the specific process in Embodiment 2 described in the first embodiment above, so the description is omitted.
[0326] Application Example 2
[0327] The process is the same as that in the specific processing described in Application Example 2 of the first embodiment above, so the description is omitted.
[0328] The specific processing unit 290 sends the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A outputs the result of the specific processing to the speaker 240. The microphone 238 acquires sound input representing the user's input regarding the result of the specific processing. The control unit 46A sends the sound data representing the user's input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the sound data.
[0329] Data generation model 58 is a so-called generative AI (Artificial Intelligence). Examples of data generation models 58 include ChatGPT (registered trademark) (accessible via the internet (URL: https: / / openai.com / blog / chatgpt)). Data generation model 58 is obtained through deep learning on a neural network. Input to data generation model 58 are prompt words containing instructions, and inference data such as sound data representing sound, text data representing text, and image data representing images (e.g., still image data or animation data). Data generation model 58 infers from the input inference data based on the instructions represented by the prompt words and outputs the inference result in one or more data forms, such as sound data, text data, and image data. Data generation model 58 includes, for example, text generation AI, image generation AI, and multimodal generation AI. Here, inference refers to, for example, analysis, classification, prediction, and / or induction. The specific processing unit 290 performs the aforementioned specific processing while using data generation model 58. The data generation model 58 can also be a model finely tuned to output inference results from prompts that do not contain instructions. In this case, the data generation model 58 can output inference results based on prompts that do not contain instructions. The data processing apparatus 12, etc., includes various data generation models 58, including AI other than the generation AI. AI other than the generation AI can be, for example, linear regression, logistic regression, decision trees, random forests, support vector machines (SVM), k-means clustering, convolutional neural networks (CNN), recurrent neural networks (RNN), generative adversarial networks (GAN), or Naive Bayes, and can perform various processes, but is not limited to this example. Furthermore, the AI can also be an AI agent. Furthermore, when the processing of the above-mentioned parts is performed by AI, the processing can be partially or entirely performed by AI, but is not limited to this example. Furthermore, the processing performed by the AI including the generation AI can be replaced by processing in the rule base, and the processing in the rule base can also be replaced by processing performed by the AI including the generation AI.
[0330] Furthermore, the processing of the aforementioned data processing system 10 is performed by the specific processing unit 290 of the data processing device 12 or the control unit 46A of the smart glasses 214, but it can also be performed by both the specific processing unit 290 of the data processing device 12 and the control unit 46A of the smart glasses 214. Additionally, the specific processing unit 290 of the data processing device 12 acquires or collects information required for processing from the smart glasses 214 or external devices, and the smart glasses 214 acquires or collects information required for processing from the data processing device 12 or external devices.
[0331] For example, the collection unit is implemented by the control unit 46A of the smart glasses 214 or the specific processing unit 290 of the data processing device 12. For example, the acquisition unit uses the camera 42 or communication I / F 44 of the smart glasses 214 to acquire step data, which is then processed by the specific processing unit 290 of the data processing device 12. For example, the analysis unit is implemented by the specific processing unit 290 of the data processing device 12, which analyzes the data from the collection unit and the acquisition unit. For example, the generation unit is implemented by the specific processing unit 290 of the data processing device 12, which uses a generation AI to generate a menu. For example, the serving unit is implemented by the speaker 240 of the smart glasses 214 or the specific processing unit 290 of the data processing device 12, which provides the generated menu to the user. The correspondence between each unit and the device or control unit is not limited to the above examples and various changes can be made.
[0332] In the above embodiments, examples of specific processing by the data processing device 12 are given, but the technology disclosed herein is not limited to this, and specific processing may also be performed by the smart glasses 214.
[0333] Third Implementation Method
[0334] Figure 5 An example of the configuration of the data processing system 310 according to the third embodiment is shown.
[0335] like Figure 5 As shown, the data processing system 310 includes a data processing device 12 and a head-mounted terminal 314. A server can be cited as an example of the data processing device 12.
[0336] The data processing apparatus 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" as understood in this disclosure. The computer 22 includes a processor 28, RAM 30, and memory 32. The processor 28, RAM 30, and memory 32 are connected to a bus 34. Furthermore, the database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0337] The head-mounted terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a display 343. The computer 36 includes a processor 46, RAM 48, and memory 50. The processor 46, RAM 48, and memory 50 are connected to a bus 52. Furthermore, the microphone 238, speaker 240, camera 42, display 343, and communication I / F 44 are also connected to the bus 52.
[0338] Microphone 238 receives instructions from user 20 by receiving sounds emitted by user 20. Microphone 238 captures sounds emitted by user 20 and converts the captured sounds into sound data, which is then output to processor 46. Speaker 240 outputs sound according to instructions from processor 46.
[0339] Camera 42 is a small digital camera equipped with an optical system such as a lens, aperture and shutter, and imaging elements such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, to capture images of the user 20's surroundings (e.g., the field of view defined by an angle equivalent to the field of vision of an average healthy person).
[0340] Communication I / F44 is connected to network 54. Communication I / F44 and 26 are responsible for the transmission and reception of various information between processor 46 and processor 28 via network 54. The transmission and reception of various information between processor 46 and processor 28 using communication I / F44 and 26 is performed in a secure state.
[0341] Figure 6 This illustrates an example of the main functions of the data processing device 12 and the head-mounted terminal 314. For example... Figure 6 As shown, in the data processing device 12, specific processing is performed by the processor 28. The specific processing program 56 is stored in the memory 32.
[0342] The specific processing program 56 is an example of a "program" involved in the technology of this disclosure. The processor 28 reads the specific processing program 56 from the memory 32 and executes the read specific processing program 56 on the RAM 30. Specific processing is implemented by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.
[0343] The memory 32 stores the data generation model 58 and the emotion-specific model 59. The data generation model 58 and the emotion-specific model 59 are used by the specific processing unit 290.
[0344] In the head-mounted terminal 314, the processor 46 performs the acceptance / output processing. The memory 50 stores the acceptance / output program 60. The processor 46 reads the acceptance / output program 60 from the memory 50 and executes the read acceptance / output program 60 on the RAM 48. The acceptance / output processing is implemented by the processor 46 operating as a control unit 46A according to the acceptance / output program 60 executed on the RAM 48.
[0345] Next, the specific processing of the specific processing unit 290 of the data processing device 12 will be described. Each part of the system described below is implemented by the data processing device 12 and the head-mounted terminal 314. In the following description, the data processing device 12 will be referred to as the "server" and the head-mounted terminal 314 will be referred to as the "terminal".
[0346] Example 1
[0347] The process is the same as that of the specific process described in Embodiment 1 in the first embodiment above, so the description is omitted.
[0348] Application Example 1
[0349] The process is the same as that in the specific processing described in Application Example 1 of the first embodiment above, so the description is omitted.
[0350] Example 2
[0351] The process is the same as that of the specific process in Embodiment 2 described in the first embodiment above, so the description is omitted.
[0352] Application Example 2
[0353] The process is the same as that in the specific processing described in Application Example 2 of the first embodiment above, so the description is omitted.
[0354] The specific processing unit 290 sends the result of the specific processing to the head-mounted terminal 314. In the head-mounted terminal 314, the control unit 46A outputs the result of the specific processing to the speaker 240 and the display 343. The microphone 238 acquires sound input representing the result of the specific processing. The control unit 46A sends the sound data representing the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the sound data.
[0355] Data generation model 58 is a so-called generative AI (Artificial Intelligence). Examples of data generation models 58 include ChatGPT (registered trademark) (accessible via the internet (URL: https: / / openai.com / blog / chatgpt)). Data generation model 58 is obtained through deep learning on a neural network. Input to data generation model 58 includes prompt words containing instructions, and inference data such as sound data representing sound, text data representing text, and image data representing images (e.g., still image data or animation data). Data generation model 58 infers the input inference data based on the instructions represented by the prompt words and outputs the inference result in one or more data forms such as sound data, text data, and image data. Data generation model 58 includes, for example, text generation AI, image generation AI, and multimodal generation AI. Here, inference refers to, for example, analysis, classification, prediction, and / or induction. The specific processing unit 290 performs the aforementioned specific processing while using data generation model 58. The data generation model 58 can also be a model finely tuned to output inference results from prompts that do not contain instructions. In this case, the data generation model 58 can output inference results based on prompts that do not contain instructions. The data processing apparatus 12, etc., includes various data generation models 58, including AI other than the generation AI. AI other than the generation AI can be, for example, linear regression, logistic regression, decision trees, random forests, support vector machines (SVM), k-means clustering, convolutional neural networks (CNN), recurrent neural networks (RNN), generative adversarial networks (GAN), or Naive Bayes, and can perform various processes, but is not limited to this example. Furthermore, the AI can also be an AI agent. Furthermore, when the processing of the above-mentioned parts is performed by AI, the processing can be partially or entirely performed by AI, but is not limited to this example. Furthermore, the processing performed by the AI including the generation AI can be replaced by processing in the rule base, and the processing in the rule base can also be replaced by processing performed by the AI including the generation AI.
[0356] Furthermore, the processing of the aforementioned data processing system 10 is performed by the specific processing unit 290 of the data processing device 12 or the control unit 46A of the head-mounted terminal 314, but it can also be performed by both the specific processing unit 290 of the data processing device 12 and the control unit 46A of the head-mounted terminal 314. Additionally, the specific processing unit 290 of the data processing device 12 acquires or collects information required for processing from the head-mounted terminal 314 or external devices, and the head-mounted terminal 314 acquires or collects information required for processing from the data processing device 12 or external devices.
[0357] For example, the collection unit is implemented by the control unit 46A of the head-mounted terminal 314 or the specific processing unit 290 of the data processing device 12. For example, the acquisition unit uses the camera 42 or communication I / F 44 of the head-mounted terminal 314 to acquire step data, which is then processed by the specific processing unit 290 of the data processing device 12. For example, the analysis unit is implemented by the specific processing unit 290 of the data processing device 12 to analyze the data from the collection unit and the acquisition unit. For example, the generation unit is implemented by the specific processing unit 290 of the data processing device 12 to generate a menu using a generation AI. For example, the serving unit is implemented by the speaker 240 and display 343 of the head-mounted terminal 314 or the specific processing unit 290 of the data processing device 12 to provide the generated menu to the user. The correspondence between each unit and the device or control unit is not limited to the above examples and various changes can be made.
[0358] In the above embodiments, examples of specific processing by the data processing device 12 are given, but the technology disclosed herein is not limited to this, and specific processing may also be performed by the head-mounted terminal 314.
[0359] Fourth Implementation Method
[0360] Figure 7 An example of the configuration of the data processing system 410 according to the fourth embodiment is shown.
[0361] like Figure 7 As shown, the data processing system 410 includes a data processing device 12 and a robot 414. A server can be cited as an example of the data processing device 12.
[0362] The data processing apparatus 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" as understood in this disclosure. The computer 22 includes a processor 28, RAM 30, and memory 32. The processor 28, RAM 30, and memory 32 are connected to a bus 34. Furthermore, the database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0363] Robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a controlled object 443. The computer 36 includes a processor 46, RAM 48, and memory 50. The processor 46, RAM 48, and memory 50 are connected to a bus 52. Furthermore, the microphone 238, speaker 240, camera 42, controlled object 443, and communication I / F 44 are also connected to the bus 52.
[0364] Microphone 238 receives instructions from user 20 by receiving sounds emitted by user 20. Microphone 238 captures sounds emitted by user 20 and converts the captured sounds into sound data, which is then output to processor 46. Speaker 240 outputs sound according to instructions from processor 46.
[0365] Camera 42 is a small digital camera equipped with an optical system such as a lens, aperture and shutter, and imaging elements such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, to photograph the area around robot 414 (e.g., the field of view defined by a perspective equivalent to the field of vision of an average healthy person).
[0366] Communication I / F44 is connected to network 54. Communication I / F44 and 26 are responsible for the transmission and reception of various information between processor 46 and processor 28 via network 54. The transmission and reception of various information between processor 46 and processor 28 using communication I / F44 and 26 is performed in a secure state.
[0367] The controlled object 443 includes a display device, LEDs (light-emitting diodes) for the eyes, and motors for driving the arms, hands, and feet. The posture or movement of the robot 414 is controlled by controlling the motors in the arms, hands, and feet. Some of the robot 414's emotions can be expressed by controlling these motors. In addition, the robot 414's facial expressions can also be expressed by controlling the illumination state of the LEDs in its eyes.
[0368] Figure 8 This illustrates an example of the main functions of the data processing device 12 and the robot 414. For example... Figure 8 As shown, in the data processing device 12, specific processing is performed by the processor 28. The specific processing program 56 is stored in the memory 32.
[0369] The specific processing program 56 is an example of a "program" involved in the technology of this disclosure. The processor 28 reads the specific processing program 56 from the memory 32 and executes the read specific processing program 56 on the RAM 30. Specific processing is implemented by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.
[0370] The memory 32 stores the data generation model 58 and the emotion-specific model 59. The data generation model 58 and the emotion-specific model 59 are used by the specific processing unit 290.
[0371] In robot 414, the processor 46 performs the acceptance and output processing. The memory 50 stores the acceptance and output program 60. The processor 46 reads the acceptance and output program 60 from the memory 50 and executes the read acceptance and output program 60 on RAM 48. The acceptance and output processing is implemented by the processor 46 acting as the control unit 46A according to the acceptance and output program 60 executed on RAM 48.
[0372] Next, the specific processing of the specific processing unit 290 of the data processing device 12 will be described. Each part of the system described below is implemented by the data processing device 12 and the robot 414. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 will be referred to as the "terminal".
[0373] Example 1
[0374] The process is the same as that of the specific process described in Embodiment 1 in the first embodiment above, so the description is omitted.
[0375] Application Example 1
[0376] The process is the same as that in the specific processing described in Application Example 1 of the first embodiment above, so the description is omitted.
[0377] Example 2
[0378] The process is the same as that of the specific process described in Embodiment 2 in the first embodiment above, so the description is omitted.
[0379] Application Example 2
[0380] The process is the same as that in the specific processing described in Application Example 2 of the first embodiment above, so the description is omitted.
[0381] The specific processing unit 290 sends the result of the specific processing to the robot 414. In the robot 414, the control unit 46A outputs the result of the specific processing to the speaker 240 and the controlled object 443. The microphone 238 acquires sound input representing the result of the specific processing. The control unit 46A sends the sound data representing the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the sound data.
[0382] Data generation model 58 is a so-called generative AI (Artificial Intelligence). Examples of data generation models 58 include ChatGPT (registered trademark) (accessible via the internet (URL: https: / / openai.com / blog / chatgpt)). Data generation model 58 is obtained through deep learning on a neural network. Input to data generation model 58 are prompt words containing instructions, and inference data such as sound data representing sound, text data representing text, and image data representing images (e.g., still image data or animation data). Data generation model 58 infers from the input inference data based on the instructions represented by the prompt words and outputs the inference result in one or more data forms, such as sound data, text data, and image data. Data generation model 58 includes, for example, text generation AI, image generation AI, and multimodal generation AI. Here, inference refers to, for example, analysis, classification, prediction, and / or induction. The specific processing unit 290 performs the aforementioned specific processing while using data generation model 58. The data generation model 58 can also be a model finely tuned to output inference results from prompts that do not contain instructions. In this case, the data generation model 58 can output inference results based on prompts that do not contain instructions. The data processing apparatus 12, etc., includes various data generation models 58, including AI other than the generation AI. AI other than the generation AI can be, for example, linear regression, logistic regression, decision trees, random forests, support vector machines (SVM), k-means clustering, convolutional neural networks (CNN), recurrent neural networks (RNN), generative adversarial networks (GAN), or Naive Bayes, and can perform various processes, but is not limited to this example. Furthermore, the AI can also be an AI agent. Furthermore, when the processing of the above-mentioned parts is performed by AI, the processing can be partially or entirely performed by AI, but is not limited to this example. Furthermore, the processing performed by the AI including the generation AI can be replaced by processing in the rule base, and the processing in the rule base can also be replaced by processing performed by the AI including the generation AI.
[0383] Furthermore, the processing of the aforementioned data processing system 10 is performed by the specific processing unit 290 of the data processing device 12 or the control unit 46A of the robot 414, but it can also be performed by both the specific processing unit 290 of the data processing device 12 and the control unit 46A of the robot 414. Additionally, the specific processing unit 290 of the data processing device 12 acquires or collects information required for processing from the robot 414 or external devices, and the robot 414 acquires or collects information required for processing from the data processing device 12 or external devices.
[0384] For example, the collection unit is implemented by the control unit 46A of the robot 414 or the specific processing unit 290 of the data processing device 12. For example, the acquisition unit uses the camera 42 or communication I / F 44 of the robot 414 to acquire step data, which is then processed by the specific processing unit 290 of the data processing device 12. For example, the analysis unit is implemented by the specific processing unit 290 of the data processing device 12, which analyzes the data from the collection unit and the acquisition unit. For example, the generation unit is implemented by the specific processing unit 290 of the data processing device 12, which uses a generation AI to generate a menu. For example, the serving unit is implemented by the speaker 240 of the robot 414 and the control object 443 or the specific processing unit 290 of the data processing device 12, which provides the generated menu to the user. The correspondence between each unit and the device or control unit is not limited to the above examples and various changes can be made.
[0385] In the above embodiments, examples of specific processing by the data processing device 12 are given, but the technology disclosed herein is not limited to this, and specific processing may also be performed by the robot 414.
[0386] Furthermore, the emotion-specific model 59, acting as an emotion engine, can determine a user's emotion based on a specific mapping. Specifically, the emotion-specific model 59 can determine a user's emotion based on an emotion graph that serves as a specific mapping (see [reference]). Figure 9 The system determines the user's emotions. Furthermore, the emotion-specific model 59 can similarly determine the robot's emotions, and the specific processing unit 290 performs specific processing based on the robot's emotions.
[0387] Figure 9 This is a diagram representing an emotion map 400 that maps multiple emotions. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. The closer to the center of the concentric circles, the more primitive the emotion is. On the outer side of the concentric circles, emotions representing states or behaviors arising from mood are arranged. Emotions are concepts that include feelings and mental states. Emotions generated by reactions occurring in the brain are arranged roughly to the left of the concentric circles. Emotions derived from situational judgments are arranged roughly to the right of the concentric circles. Emotions generated by reactions occurring in the brain and derived from situational judgments are arranged roughly above and below the concentric circles. Furthermore, "pleasant" emotions are arranged above the concentric circles, and "unpleasant" emotions are arranged below them. Thus, in the emotion map 400, multiple emotions are mapped based on the structure that generates emotions, and emotions that are likely to occur simultaneously are mapped close to each other.
[0388] These emotions are distributed at the three o'clock position of the emotion map 400, typically fluctuating between peace and anxiety. In the right half of the emotion map 400, situational awareness dominates over internal sensation, thus resulting in an impression of calm.
[0389] The inner side of the emotion map 400 represents the inner state, while the outer side represents behavior. Therefore, the further outward you are from the emotion map 400, the more visible the emotion becomes (manifested in behavior).
[0390] Here, human emotions are based on various balances such as posture and blood sugar levels. When these balances deviate from an ideal state, it indicates an unpleasant state; when they approach the ideal state, it indicates a pleasant state. Emotions in robots, cars, motorcycles, etc., can also be created in the following way: based on various balances such as posture and remaining battery power, when these balances deviate from an ideal state, it indicates an unpleasant state; when they approach the ideal state, it indicates a pleasant state. Emotion maps can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on a Brain Physiological Signal Analysis System for Voice Emotion Recognition and Emotion, Tokushima University, Doctoral Dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). In the left half of the emotion map, emotions belonging to the sensory-dominated region, called "response," are arranged. Furthermore, in the right half of the emotion map, emotions belonging to the situational cognition-dominated region, called "situation," are arranged.
[0391] In the emotion map, two types of emotions that promote learning are defined. One is a negative emotion on the situational side, in the middle or peripheral region of "repentance" or "reflection." This occurs when the robot experiences negative emotions such as "I don't want to experience this feeling again" or "I don't want to be blamed again." The other is a positive emotion on the response side, near the "desire" region. This occurs when there are positive feelings such as "wanting more" or "wanting to know more."
[0392] The emotion-specific model 59 inputs user input into a pre-trained neural network to obtain emotion values representing each emotion shown in the emotion map 400, thereby determining the user's emotion. This neural network is pre-trained based on multiple training data sets, which are combinations of user input and emotion values representing each emotion shown in the emotion map 400. Furthermore, this neural network... Figure 10 As shown in the sentiment graph 900, it was trained in a way that sentiments that are configured close to each other have similar values. Figure 10 The text shows examples of emotions such as "peace of mind", "stability", and "reassurance" that have similar emotion values.
[0393] The above description focuses on the functions of the data processing device 12, but the system of this disclosure is not necessarily installed on a server. The system of this disclosure can also be installed as a general information processing system. This disclosure can also be installed, for example, as a software program running on a personal computer, an application running on a smartphone, etc. The method of this disclosure can also be provided to users in the form of SaaS (Software as a Service).
[0394] In the above embodiments, an example of a specific process being performed by a single computer 22 is given. However, the technology disclosed herein is not limited to this, and the specific process can also be distributed among multiple computers, including computer 22. For example, the data generation model 58 can be located on an external device of the data processing apparatus 12, where data is generated based on the input data.
[0395] In the above embodiments, examples of storing a specific processing program 56 in the memory 32 have been described, but the technology disclosed herein is not limited thereto. For example, the specific processing program 56 may also be stored in a portable computer-readable non-transitory storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-transitory storage medium is installed into the computer 22 of the data processing device 12. The processor 28 executes specific processing according to the specific processing program 56.
[0396] Alternatively, a specific processing program 56 may be pre-stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 according to the requirements of the data processing device 12.
[0397] In addition, it is not necessary to store all the specific processing program 56 in the storage device such as the server connected to the data processing device 12 via the network 54 or in the memory 32; a portion of the specific processing program 56 may be stored in advance.
[0398] As hardware resources for performing specific processes, various processors, as shown below, can be used. For example, a CPU can be listed as a processor, which functions as a general-purpose processor that performs specific processes by executing software, i.e., a program. Furthermore, processors can be listed as special-purpose circuits such as FPGAs (Field-Programmable Gate Arrays), PLDs (Programmable Logic Devices), or ASICs (Application-Specific Integrated Circuits), which are processors with circuitry specifically designed to perform specific processes. Each processor has built-in or connected memory, and each processor executes specific processes using that memory.
[0399] The hardware resources for performing a specific process can consist of one of these various processors, or a combination of two or more processors of the same or different types (e.g., a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Alternatively, the hardware resources for performing a specific process can be a single processor.
[0400] As an example of a single processor, there are two approaches: First, a processor is composed of a combination of one or more CPUs and software, which functions as a hardware resource to perform a specific process; second, as represented by a SoC (System-on-a-chip), a processor is used to implement the functionality of the entire system, which includes multiple hardware resources for performing a specific process, using a single IC (Integrated Circuit) chip. In this way, the specific process is implemented by using one or more of the aforementioned processors as hardware resources.
[0401] Furthermore, the hardware architecture of these various processors, more specifically, can utilize circuits that combine semiconductor elements and other circuit components. Moreover, the specific process described above is just one example. Therefore, without departing from the main point, unnecessary steps can certainly be deleted, new steps added, or the processing order changed.
[0402] The descriptions and illustrations above are detailed explanations of a portion of the technology disclosed herein, and are merely one example of the technology disclosed herein. For example, the above descriptions of the structure, function, effect, and results are just one example of the structure, function, effect, and results of a portion of the technology disclosed herein. Therefore, without departing from the spirit of the technology disclosed herein, unnecessary parts may be deleted, new elements added, or replacements may be made to the descriptions and illustrations above. Furthermore, to avoid confusion and facilitate understanding of a portion of the technology disclosed herein, explanations of common technical knowledge that do not require special explanation under the premise of being able to implement the technology disclosed herein have been omitted from the descriptions and illustrations above.
[0403] All documents, patent applications and technical specifications set forth in this specification are incorporated herein by reference to the same extent that each document, patent application and technical specification is specifically and individually described therein and referenced by reference.
[0404] In addition, the following notes are provided in response to the above explanation.
[0405] Example 1
[0406] (Note 1)
[0407] An information processing system includes: a shooting component for periodically acquiring image information, wherein the image information is acquired in a continuous transmission format; a preprocessing unit for preprocessing the acquired image information, performing noise reduction, resolution transformation, and color space transformation using an image processing component; an image analysis unit for analyzing the preprocessed image information using machine learning algorithms such as a convolutional artificial intelligence model running on computer resources, thereby extracting abnormal states of a target object with high confidence; an information recording unit for recording the extraction results and evaluation values of the abnormal states of the target object into an information storage device, including time information and shooting location; an information generation unit for determining the relevant location and generating a notification content indicating corresponding measures when the evaluation value exceeds a predetermined threshold based on the recorded evaluation value; a communication unit for sending the notification content to an information display terminal via wireless communication; a display unit for displaying the notification content on the information display terminal in the form of a screen notification or an in-application notification; and a dialogue assistance unit for receiving user-submitted query text, performing content analysis using a generative artificial intelligence model, automatically generating specific answers, and displaying them through the display unit.
[0408] (Note 2)
[0409] The information processing system according to Appendix 1 further includes: a functional unit for automatically generating prompt statements input to the generative artificial intelligence model and providing them to the model.
[0410] (Note 3)
[0411] According to the information processing system described in Appendix 1, the image processing component, computer resources, information storage device, and generative artificial intelligence model are implemented in the same computer system or distributed processing environment.
[0412] Application Example 1
[0413] (Note 1)
[0414] An information processing system includes: an acquisition unit for periodically receiving video data from multiple acquisition devices in streaming media format; a processing unit for preprocessing the received video data using image processing methods; an analysis unit for analyzing the preprocessed video data using a machine learning model to identify abnormal behavior, objects, or organisms; a recording unit for storing detection content, confidence level, acquisition source location, and time information to a storage medium based on the analysis results; a generation unit for automatically generating corresponding instruction messages when the analysis results exceed a benchmark value, filtering out the required items; a sending unit for sending the generated instruction messages to an information terminal via a communication function; a display unit for instantly notifying or displaying the messages on the information terminal; a dialogue support unit for receiving user query information, analyzing the query information using a generative artificial intelligence model, generating predetermined response content, and displaying it on the information terminal; and a support unit for analyzing user status data from the acquisition devices or the information terminal, inferring user emotions, and adjusting message content or response content based on the inference results.
[0415] (Note 2)
[0416] According to the information processing system described in Note 1, the parsing unit combines multiple machine learning algorithms to perform comprehensive analysis in abnormal behavior identification, and makes a judgment after integrating multiple results.
[0417] (Note 3)
[0418] According to the information processing system described in Appendix 1, the dialogue support unit generates personalized response content for the user by inputting multiple prompts into the generative artificial intelligence model.
[0419] Example 2
[0420] (Note 1)
[0421] An information processing system includes: a video acquisition device for periodically receiving image data, and an image acquisition unit for receiving the image data in streaming media format; a preprocessing unit for preprocessing the received image data; an image analysis unit for detecting stains or organisms in the image data processed by the preprocessing unit using a generative artificial intelligence model; a result recording unit for recording the detection results and confidence scores of the image analysis unit, along with time information and device location information, into an information storage device; a notification generation unit for extracting the location and automatically generating prompt content based on the detection results obtained from the information storage device when the confidence score exceeds a predetermined threshold; a notification sending unit for sending the generated prompt content to an information communication device; a display unit for displaying the prompt content in real time on the information communication device; a dialogue processing unit for receiving user queries, parsing them using a natural language processing model, generating corresponding user responses, and displaying them through the information communication device; and a sentiment analysis unit for performing sentiment analysis on voice data or image data obtained from the user and automatically adjusting the expression of the prompt content based on the analysis results.
[0422] (Note 2)
[0423] According to the information processing system described in Note 1, the generative artificial intelligence model is composed of a neural network that includes convolution operations.
[0424] (Note 3)
[0425] According to the information processing system described in Note 1, the emotion analysis unit determines the user's state through a voice analysis algorithm and an emotion judgment algorithm, and automatically generates prompts containing encouraging or comforting suggestions based on the detection results.
[0426] Application Example 2
[0427] (Note 1)
[0428] An information processing system includes: an imaging device for periodically acquiring image information and acquiring the image information in the form of a video stream; an information analysis unit for preprocessing the acquired image and audio information and using an artificial intelligence model for individual recognition and emotional state analysis; an information recording unit for recording the analysis results along with confidence scores and recording time and spatial information as historical analysis information; a generation processing unit for automatically generating personalized notification text based on the analyzed emotional state or detection state; a notification sending unit for sending the generated notification text to a display device for display; an information prompting unit for displaying the notification text live on the display screen of a portable information terminal or the visual interface of an augmented reality device; and a conversational response unit for receiving user query information, analyzing it using a natural language processing model, dynamically generating response text, and responding through the information prompting unit.
[0429] (Note 2)
[0430] The information processing system according to Appendix 1 further includes: a prompting statement generation unit for automatically generating prompting statements for input to a generative artificial intelligence model based on the user's behavioral tendencies, facial expressions, voice content, and analysis history information, and inputting these prompting statements into the generative artificial intelligence model.
[0431] (Note 3)
[0432] According to the information processing system described in Note 1, the artificial intelligence model combines a facial feature extraction algorithm for individual recognition and emotion inference, and combines the natural language processing model for semantic analysis.
Claims
1. An information processing system, characterized by comprising: Comprise: An image acquisition unit for periodically receiving image data captured by camera devices and receiving the image data in a streaming format; An image analysis unit for pre-processing the received image data and detecting stains or pests using an artificial intelligence model; An analysis result recording unit for recording the detection results together with a confidence score, and including a time stamp and camera location information; A notification generation unit for extracting locations that need cleaning or attention from the detection results when the confidence score exceeds a preset threshold, and generating corresponding notification messages; A notification sending unit for sending the generated notification messages to user terminals; A user interface unit for displaying the notification messages in the form of real-time push notifications or in-app messages; A chatbot unit for receiving questions from users, parsing them using a natural language processing model, generating specific replies, and displaying them through the user interface unit.
2. The information processing system according to claim 1, characterized by, The image acquisition unit can communicate with multiple camera devices within the home and acquire image data in real time from different locations.
3. The information processing system according to claim 1, characterized by, The analysis result recording unit can store the detection results in chronological order and support subsequent queries and retrievals.
Citation Information
Patent Citations
Persona chatbot control method and system
JP2022180282A