system
The system uses cameras and microphones to analyze pet behaviors and vocalizations, providing real-time emotional feedback through a speaker collar, addressing the challenge of accurately understanding pet emotions and enhancing pet-owner communication.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-09-30
- Publication Date
- 2026-04-09
AI Technical Summary
Conventional communication with pets relies on observing their cries and behaviors to infer emotions, which is inaccurate and difficult for ordinary pet owners to understand, leading to anxiety and stress in pet-owner relationships.
A system that uses cameras and microphones to capture animal behavior, gestures, and vocalizations, preprocesses the data, analyzes emotions on a server, and generates sound via a speaker attached to the animal's collar to convey emotions accurately and in real-time.
Enables pet owners to understand their pets' emotions more accurately and respond appropriately, improving communication and reducing anxiety through real-time emotional feedback.
Smart Images

Figure 2026062142000001_ABST
Abstract
Description
Technical Field
[0001] The technology of the present disclosure relates to a system.
Background Art
[0002] Patent Document 1 discloses a method for controlling a persona chatbot, which is performed by at least one processor, including steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to an explanation of a chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance.
Prior Art Documents
Patent Documents
[0003]
Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0004] Conventional communication with pets depends on observing their cries and behaviors to infer their emotions, and it is difficult to accurately understand the emotions of pets. Therefore, pet owners cannot immediately grasp the emotions of their pets and respond appropriately, and may feel anxiety and stress in the relationship with their pets. In particular, this problem is prominent for ordinary pet owners who do not have specialized knowledge of ethology. Therefore, an object of the present invention is to provide a system that accurately analyzes the behaviors, mannerisms, expressions, and cries of pets and expresses their emotions in voice, so that pet owners can understand the emotions of their pets in real time.
Means for Solving the Problems
[0005] The present invention provides a system comprising: means for acquiring an animal's behavior, gestures, facial expressions, and vocalizations from a camera and microphone; means for preprocessing the data acquired from the camera and microphone; means for analyzing the preprocessed data and estimating the animal's emotions; means for generating sound based on the estimated emotions; and means for emitting the generated sound from a speaker attached to the animal's collar. This system allows pet owners to understand their pets' emotions in real time and facilitate smooth communication with them. Furthermore, by providing means for transmitting the data to a server, the analysis processing is performed on a high-performance server, achieving highly accurate emotion estimation. Additionally, by providing means for estimating the animal's emotions via the server, advanced emotional analysis of pets can be performed.
[0006] A "camera" is a device used to capture video data of an animal's behavior, gestures, facial expressions, and other similar information.
[0007] A "microphone" is a sound recording device used to acquire animal sounds and other audio data.
[0008] "Animals" primarily refer to non-human living beings such as cats and dogs kept as pets.
[0009] "Behavior" refers to the various movements and activities exhibited by animals, as well as their state and emotions.
[0010] "Gestures" refer to subtle movements and gestures expressed through the body and facial movements of animals.
[0011] "Facial expressions" are signs of emotion expressed through patterns on an animal's face, as well as the movements of its eyes and mouth.
[0012] "A sound" refers to an animal's vocalization, which expresses its emotions or needs.
[0013] "Preprocessing" refers to data processing that converts data acquired from cameras and microphones into a format that is easy to analyze.
[0014] "Analysis" is the process of evaluating an animal's behavior, gestures, facial expressions, and vocalizations based on pre-processed data, and estimating its emotions.
[0015] "Emotion" refers to a psychological or emotional state that can be inferred from an animal's behavior, gestures, facial expressions, vocalizations, etc.
[0016] "Voice generation" is the process of creating synthesized speech based on emotions estimated through analysis.
[0017] A "speaker" is a sound output device attached to an animal's collar or similar object to emit generated sound.
[0018] A "collar" is an accessory worn around an animal's neck and is a device used to attach equipment such as speakers.
[0019] A "server" is a high-performance computing device that receives data transmitted from terminals and performs analysis processing on it. [Brief explanation of the drawing]
[0020] [Figure 1] This is a conceptual diagram showing an example of the configuration of a data processing system according to the first embodiment. [Figure 2] This is a conceptual diagram showing an example of the essential functions of a data processing device and a smart device according to the first embodiment. [Figure 3] This is a conceptual diagram showing an example of the configuration of a data processing system according to the second embodiment. [Figure 4] This is a conceptual diagram showing an example of the main functions of a data processing device and smart glasses according to the second embodiment. [Figure 5] This is a conceptual diagram showing an example of the configuration of a data processing system according to the third embodiment. [Figure 6] This is a conceptual diagram showing an example of the main functions of a data processing device and a headset-type terminal according to the third embodiment. [Figure 7] It is a conceptual diagram showing an example of the configuration of a data processing system according to the fourth embodiment. [Figure 8] It is a conceptual diagram showing an example of the main functions of a data processing device and a robot according to the fourth embodiment. [Figure 9] It shows an emotion map to which a plurality of emotions are mapped. [Figure 10] It shows an emotion map to which a plurality of emotions are mapped. [Figure 11] It is a sequence diagram showing the processing flow of the data processing system in Example 1. [Figure 12] It is a sequence diagram showing the processing flow of the data processing system in Application Example 1. [Figure 13] It is a sequence diagram showing the processing flow of the data processing system in Example 2 when an emotion engine is combined. [Figure 14] It is a sequence diagram showing the processing flow of the data processing system in Application Example 2 when an emotion engine is combined.
Mode for Carrying Out the Invention
[0021] Hereinafter, an example of an embodiment of a system according to the technology of the present disclosure will be described with reference to the accompanying drawings.
[0022] First, the terms used in the following description will be explained.
[0023] In the following embodiments, the signed processor (hereinafter simply referred to as "processor") may be a single arithmetic unit or a combination of multiple arithmetic units. Furthermore, the processor may be a single type of arithmetic unit or a combination of multiple types of arithmetic units. Examples of arithmetic units include CPU (Central Processing Unit), GPU (Graphics Processing Unit), GPGPU (General-Purpose computing on Graphics Processing Units), and APU (Accelerated Processing Unit).
[0024] In the following embodiments, signed RAM (Random Access Memory) is a memory that temporarily stores information and is used as work memory by the processor.
[0025] In the following embodiments, the signed storage is one or more non-volatile storage devices that store various programs and various parameters. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), or magnetic tapes.
[0026] In the following embodiments, the signed communication interface (I / F) is an interface that includes a communication processor and an antenna, etc. The communication interface manages communication between multiple computers. Examples of communication standards applicable to the communication interface include wireless communication standards such as 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), or Bluetooth (registered trademark).
[0027] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." That is, "A and / or B" means that it may be A alone, or B alone, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" applies when expressing three or more things linked by "and / or."
[0028] [First Embodiment]
[0029] Figure 1 shows an example of the configuration of the data processing system 10 according to the first embodiment.
[0030] As shown in Figure 1, the data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.
[0031] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0032] The smart device 14 comprises a computer 36, a reception device 38, an output device 40, a camera 42, and a communication interface 44. The computer 36 comprises a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The reception device 38, output device 40, and camera 42 are also connected to the bus 52.
[0033] The reception device 38 is equipped with a touch panel 38A and a microphone 38B, etc., and receives user input. The touch panel 38A receives user input by detecting contact with an object (e.g., a pen or finger). The microphone 38B receives user input by detecting the user's voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.
[0034] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form perceptible to the user 20 (e.g., audio and / or text). The display 40A displays visible information such as text and images according to instructions from the processor 46. The speaker 40B outputs audio according to instructions from the processor 46. The camera 42 is a small digital camera equipped with an optical system such as a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.
[0035] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various types of information between processor 46 and processor 28 via network 54.
[0036] Figure 2 shows an example of the main functions of the data processing device 12 and the smart device 14.
[0037] As shown in Figure 2, in the data processing device 12, a specific processing is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" related to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.
[0038] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0039] In the smart device 14, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The reception output program 60 is used in conjunction with a specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0040] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".
[0041] This invention is a system that uses a camera and microphone to acquire the behavior, gestures, facial expressions, and vocalizations of animals, analyzes this data to estimate the animals' emotions, and expresses the estimated emotions in sound. Specific embodiments for carrying out this invention are described below.
[0042] System Overview
[0043] This system uses multiple cameras and microphones placed in a room to capture animal behavior and vocalizations in real time, and sends the data to a server for analysis. Based on the analysis results, sounds are generated and emitted from speakers attached to the animals' collars, allowing users to understand the animals' emotions.
[0044] System Configuration
[0045] 1. Camera and microphone:
[0046] Multiple cameras and microphones installed in the room capture animal behavior and sounds in real time. The cameras continuously record animal movements, while the microphones record sounds such as vocalizations.
[0047] 2. Preprocessing at the terminal:
[0048] The device preprocesses the data acquired from the camera and microphone. Specifically, it removes background noise from the video and filters the audio data to reduce noise. This preprocessed data is then sent to the server for further analysis.
[0049] 3. Analysis on the server:
[0050] The server analyzes pre-processed data sent from the terminal. It identifies animal behavior, gestures, and facial expressions from image data, and analyzes animal sounds from audio data. This allows it to estimate the animal's emotions.
[0051] 4. Speech generation:
[0052] Based on the emotions of the animals estimated by the server, the appropriate voice text is selected and generated as speech. For example, if the emotion "happy" is estimated, the corresponding speech will be "Feels good!".
[0053] 5. Audio playback via speaker:
[0054] The generated audio is transmitted via a device to a speaker attached to the animal's collar. The audio emitted from the speaker allows the user to understand the animal's emotions in real time.
[0055] Specific examples of actions
[0056] Example 1: When a cat is happy to be petted
[0057] 1. Device: The camera records the user's actions as they pet the cat, and the microphone captures the cat's meows.
[0058] 2. Terminal: Preprocess this data to remove unwanted noise.
[0059] 3. Terminal: Sends pre-processed data to the server.
[0060] 4. Server: Analyzes data and estimates whether the cat is "happy" based on its gestures and meows.
[0061] 5. Server: Generates the voice "Feels good!" which corresponds to the "feeling happy" state.
[0062] 6. Terminal: Sends the generated audio to the speaker and plays it.
[0063] 7. User: Hearing the voice saying "Feels good!" from the cat's collar, the user understands that the cat is happy.
[0064] Example 2: When the cat looks anxious
[0065] 1. Device: The camera and microphone capture the cat's actions and meows when it is startled by something.
[0066] 2. Terminal: Preprocess this data.
[0067] 3. Terminal: Sends pre-processed data to the server.
[0068] 4. Server: Analyzes data and estimates whether the cat is "anxious" based on its behavior and meows.
[0069] 5. Server: Generates the voice message "What's wrong?" which corresponds to the state of "feeling anxious".
[0070] 6. Terminal: Sends the generated audio to the speaker and plays it.
[0071] 7. User: Hearing the voice saying "What's wrong?" from the cat's collar, the user understands that the cat is anxious and takes action to soothe it.
[0072] In this way, the system provides support for better communication between animals and their owners, and makes it possible to understand pets' emotions more accurately and in real time.
[0073] The following describes the processing flow.
[0074] Step 1:
[0075] The device initializes the camera and microphone installed in the room and captures the animal's behavior, gestures, facial expressions, and vocalizations in real time. Specifically, it continuously captures frame images from the camera and records audio data from the microphone.
[0076] Step 2:
[0077] The terminal preprocesses the acquired video and audio data. A noise reduction filter is applied to the video data to extract the animal's areas of interest. Background noise is removed from the audio data, converting it into a clear audio signal.
[0078] Step 3:
[0079] The terminal packages the pre-processed data and sends it to the server via the internet. The data package includes video frames and filtered audio data.
[0080] Step 4:
[0081] The server receives the data package sent from the terminal and prepares it for analysis. It performs a consistency check on the received data to confirm that there are no missing or corrupted files.
[0082] Step 5:
[0083] The server analyzes the video data and applies models to identify animal behavior, gestures, and facial expressions. For example, it can detect the posture a cat is in when being petted, or specific gestures.
[0084] Step 6:
[0085] The server analyzes the audio data and applies a speech recognition model to estimate emotions from animal sounds. It analyzes the patterns and phonemes of the sounds to identify emotions.
[0086] Step 7:
[0087] The server integrates the analysis results of video and audio data to estimate the animal's overall emotions. For example, if the animal shows signs of happiness, it will be estimated to be "happy."
[0088] Step 8:
[0089] The server generates an appropriate voice message based on the estimated emotion. For example, if the emotion is "happy," it will generate the voice message "Feels good!"
[0090] Step 9:
[0091] The server sends the generated audio data to the terminal. This includes the audio message and associated metadata.
[0092] Step 10:
[0093] The device receives audio data from the server and transmits it to a speaker attached to the animal's collar, playing the audio.
[0094] Step 11:
[0095] Users can understand animal emotions through voice messages emitted from the speaker. For example, by hearing the voice say "That feels good!", they can recognize that the cat is happy.
[0096] This allows users to understand the animals' emotions in real time and engage in appropriate dialogue and responses.
[0097] (Example 1)
[0098] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."
[0099] Conventional pet emotion understanding systems have difficulty accurately estimating emotions from animal behavior and vocalizations, potentially leading owners to misunderstand their pets' feelings. Furthermore, the time required for real-time emotion estimation and feedback hinders smooth communication between pets and their owners. To address these issues, the present invention provides a system that accurately and in real-time estimates animal emotions and expresses those emotions verbally.
[0100] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0101] In this invention, the server includes means for acquiring animal behavior, gestures, facial expressions, and vocalizations from a camera and a microphone; means for preprocessing the data acquired from the camera and microphone; means for transmitting the preprocessed data to the server; means for analyzing the preprocessed data and estimating the animal's emotions; means for generating sound based on the estimated emotions; and means for emitting the generated sound from a speaker attached to the animal's collar. This makes it possible to estimate the animal's emotions with high accuracy and in real time, and to express those emotions in sound.
[0102] A "camera" is a device used to capture images of an animal's behavior, gestures, and facial expressions as video data.
[0103] A "microphone" is a sound-collecting device used to capture animal sounds and ambient noises as audio data.
[0104] "Data preprocessing" is the process of removing noise from video and audio data acquired from cameras and microphones, and preparing it in a format that is easy to analyze.
[0105] A "server" is a central processing unit that receives data transmitted from terminals, performs analysis, and estimates emotions.
[0106] "Analysis" refers to the process of identifying and analyzing animal behavior and vocalizations based on pre-processed data.
[0107] "Emotion estimation" is the process of identifying an animal's emotions based on the results of an analysis.
[0108] "Voice generation" is the process of selecting appropriate voice text to express estimated emotions and generating it as voice data.
[0109] A "speaker" is an output device that plays back generated audio data as sound.
[0110] A "terminal" is a device that preprocesses data acquired from cameras and microphones and transmits it to a server.
[0111] This invention is a system that uses a camera and microphone to acquire animal behavior, gestures, facial expressions, and vocalizations in real time, analyzes this data to estimate the animal's emotions, and expresses those estimated emotions in sound. This system is designed to make it easier for users to understand animal emotions. Its specific form is described below.
[0112] Hardware and software to be used
[0113] Camera: Used to continuously record the behavior and facial expressions of animals. Examples include the Logitec HD Pro Webcam.
[0114] Microphone: Used to record animal sounds and ambient noise. Examples include the Blue Yeti USB Mic.
[0115] Terminal: A small PC used for data preprocessing and communication with the server. For example, a Raspberry Pi can be used.
[0116] Server: A central processing unit that performs data analysis and sentiment estimation. Examples include AWS® and Google® Cloud Platform.
[0117] Generative AI models: AI / ML frameworks used for data analysis and sentiment estimation. Specific examples include TENSORFLOW® and PyTorch.
[0118] Text-to-speech software: Generates speech based on estimated emotions. Examples include Amazon Polly and Google Text-to-Speech.
[0119] Speaker: An output device for playing back the generated sound, attached to the animal's collar.
[0120] Specific examples of actions
[0121] Example 1: When a cat is happy to be petted
[0122] 1. The device uses a camera to record the user's actions as they pet the cat, and a microphone to capture the cat's meows.
[0123] 2. The terminal preprocesses this data to remove unwanted noise.
[0124] 3. The terminal sends the pre-processed data to the server.
[0125] 4. The server analyzes the data and estimates whether the cat is "happy" based on its gestures and meows.
[0126] 5. The server generates the voice "Feels good!" which corresponds to the state of being "happy".
[0127] 6. The device sends the generated audio to the speaker and plays it.
[0128] 7. The user hears the cat say "Feels good!" from the cat's collar and understands that the cat is happy.
[0129] Example 2: When the cat looks anxious
[0130] 1. The device uses its camera and microphone to capture any actions or sounds the cat makes when it is startled by something.
[0131] 2. The terminal preprocesses this data.
[0132] 3. The terminal sends the pre-processed data to the server.
[0133] 4. The server analyzes the data and estimates whether the cat is "anxious" based on its behavior and meows.
[0134] 5. The server generates the voice message "What's wrong?" which corresponds to the state of being "anxious".
[0135] 6. The device sends the generated audio to the speaker and plays it.
[0136] 7. The user hears a voice saying "What's wrong?" from the cat's collar, understands that the cat is anxious, and takes action to soothe it.
[0137] Example of a prompt
[0138] "Analyze animal behavioral and auditory data to estimate the animal's emotions and generate appropriate voice-to-text. Specifically, if the animal is happy, generate the voice 'Feels good!', and if it is anxious, generate the voice 'What's wrong?'."
[0139] In this way, the present invention provides support for better communication between animals and their owners, and makes it possible to understand pets' emotions more accurately and in real time.
[0140] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0141] Step 1:
[0142] Data acquisition
[0143] The device uses a camera and microphone installed in the room to capture animal behavior and sounds in real time. The camera captures video of the animals, and the microphone records their audio. Specifically, the camera captures video data frame by frame, and the microphone continuously records audio.
[0144] Input: Animal behavior and sounds
[0145] Output: Raw video data and raw audio data
[0146] Step 2:
[0147] Data preprocessing
[0148] The device preprocesses the acquired video and audio data. Specifically, it removes background noise from the video data and filters the audio data to reduce unwanted noise. For video data, it masks parts other than animals, and for audio data, it applies filtering techniques to remove ambient noise.
[0149] Input: Raw video data and raw audio data
[0150] Output: Pre-processed video data and pre-processed audio data
[0151] Step 3:
[0152] Data transmission
[0153] The terminal sends pre-processed data to the server. This pre-processed video and audio data are transmitted to the server in real time using a high-speed network connection.
[0154] Input: Pre-processed video data and pre-processed audio data
[0155] Output: Data sent to the server
[0156] Step 4:
[0157] Data Analysis
[0158] The server analyzes the received pre-processed data. Specifically, it uses a generative AI model (e.g., TensorFlow or PyTorch) to identify animal behavior, gestures, and facial expressions from video data, and analyze animal sounds from audio data. Based on prompts, it obtains the analysis results.
[0159] Input: Pre-processed video data and pre-processed audio data
[0160] Output: Behavioral analysis results and vocalization analysis results
[0161] Step 5:
[0162] Emotion estimation
[0163] The server estimates the animal's emotions based on the results of data analysis. Specifically, it combines behavioral analysis results and vocalization analysis results to estimate the animal's current emotional state (e.g., "happy," "anxious").
[0164] Input: Behavioral analysis results and vocalization analysis results
[0165] Output: Estimated emotional state
[0166] Step 6:
[0167] Speech generation
[0168] The server generates speech based on estimated emotions. Specifically, it selects appropriate speech text corresponding to the emotional state and generates speech data using speech synthesis software (such as Amazon Polly or Google Text-to-Speech).
[0169] Input: Estimated emotional state
[0170] Output: Generated audio data
[0171] Step 7:
[0172] Audio transmission and playback
[0173] The device receives audio data from the server and transmits it to a speaker attached to the animal's collar, which then plays the audio. Specifically, the device transmits the audio data to the playback speaker via Bluetooth or Wi-Fi, and the speaker outputs the audio.
[0174] Input: Generated audio data
[0175] Output: Audio played through the speaker
[0176] Through these steps, users can understand the animals' emotions in real time, enabling better communication with them.
[0177] (Application Example 1)
[0178] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."
[0179] Conventional animal emotion estimation systems can estimate an animal's emotions, but they lack the ability to detect abnormal behavior and notify the user in real time. As a result, it is difficult for pet owners and caregivers to respond quickly to abnormal animal behavior. Furthermore, safety issues remain unresolved. This invention aims to improve animal safety and user peace of mind by detecting not only animal emotions but also abnormal behavior and notifying the user.
[0180] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0181] In this invention, the server includes means for acquiring animal behavior, gestures, facial expressions, and vocalizations from a camera and a microphone; means for preprocessing the data acquired from the camera and microphone; means for analyzing the preprocessed data and estimating the animal's emotions; means for generating sound based on the estimated emotions; means for emitting the generated sound from a speaker attached to the animal's collar; means for analyzing the data to detect abnormal animal behavior; and means for generating a warning based on the detected abnormal behavior and notifying the user. This makes it possible to detect both animal emotions and abnormal behavior in real time and notify the user.
[0182] A "camera" is a device used to capture images of an animal's behavior, gestures, and facial expressions.
[0183] A "microphone" is a sound acquisition device used to record animal sounds.
[0184] "Data preprocessing" refers to the process of removing noise and converting the data format of video and audio data acquired from cameras and microphones.
[0185] "Data analysis" refers to algorithmic processing that identifies animal behavior, gestures, and vocalizations from pre-processed video and audio data, and estimates the animal's emotions and abnormal behavior.
[0186] "Emotion estimation" is the process of inferring what emotional state an animal is in based on its behavior, gestures, and vocalizations.
[0187] "Voice generation" is the process of generating voice messages that convey emotions to the user based on estimated animal emotions.
[0188] A "speaker attached to an animal's collar" is an audio playback device installed on the collar worn by an animal, used to play generated audio messages.
[0189] "Detection of abnormal behavior" is the process of determining whether an animal is behaving in an unusual way based on data analysis.
[0190] "Notification to users" refers to the process of sending warning messages or notifications to users in real time based on detected abnormal behavior.
[0191] "Real-time" is a time standard that means processing or notifications are performed almost instantaneously or with very little delay.
[0192] The present invention is a system that uses a camera and microphone to acquire animal behavior, gestures, facial expressions, and vocalizations, analyzes this data to estimate the animal's emotions and abnormal behavior, and notifies the user as necessary. Specific embodiments for carrying out the present invention are described below.
[0193] System Overview
[0194] This system uses multiple cameras and microphones placed in a room or outdoors to capture animal behavior and vocalizations in real time, and transmits the data to a server for analysis. Based on the analysis results, it emits sounds from a speaker attached to the animal's collar, allowing the user to understand the animal's emotions. Furthermore, it has a function to notify the user of any abnormal behavior if it is detected.
[0195] System Configuration
[0196] 1. Camera and microphone
[0197] Multiple cameras and microphones, installed indoors or outdoors, capture animal behavior and sounds in real time. The cameras continuously record animal movements, while the microphones record sounds such as vocalizations.
[0198] 2. Preprocessing at the terminal
[0199] The device preprocesses the data acquired from the camera and microphone. Specifically, it removes background noise from the video and filters the audio data to reduce noise. This preprocessed data is then sent to the server for further analysis.
[0200] 3. Analysis on the server
[0201] The server analyzes pre-processed data sent from the terminal. It identifies animal behavior, gestures, and facial expressions from image data, and analyzes vocalizations from audio data. This allows it to estimate the animal's emotions and abnormal behavior. The server performs the analysis using machine learning algorithms such as TensorFlow.
[0202] 4. Speech generation
[0203] Based on the emotions of the animal estimated by the server, the appropriate voice text is selected and generated as speech. For example, if the emotion "happy" is estimated, the corresponding speech will be "Feels good!". In addition, if abnormal behavior is detected, a warning voice such as "Be careful!" will be generated.
[0204] 5. Audio playback via speaker
[0205] The generated audio is transmitted via a device to a speaker attached to the animal's collar. The audio emitted from the speaker allows the user to understand the animal's emotions and any abnormal behavior in real time.
[0206] Specific examples of actions
[0207] Example 1: When a dog senses a suspicious person and becomes alert.
[0208] 1. Device: The camera and microphone capture the dog's alert behavior and barking.
[0209] 2. Terminal: Preprocess this data to remove unwanted noise.
[0210] 3. Terminal: Sends pre-processed data to the server.
[0211] 4. Server: Analyzes data, estimates a "vigilant" state from the dog's movements and barks, and further detects abnormal behavior.
[0212] 5. Server: Generates the voice message "Be careful!" corresponding to the "Alert" state.
[0213] 6. Terminal: Sends the generated audio to the speaker and plays it.
[0214] 7. User: Hearing the "Be careful!" voice message from the dog's collar, the user understands that the dog is on alert and takes appropriate action.
[0215] Example 2: When the dog is playing safely
[0216] 1. Device: The camera and microphone capture the dog's playful behavior and barking.
[0217] 2. Terminal: Preprocess this data.
[0218] 3. Terminal: Sends pre-processed data to the server.
[0219] 4. Server: Analyzes data and estimates whether the dog is "enjoying" itself based on its movements and barks.
[0220] 5. Server: Generates the voice message "This is fun!" which corresponds to the "enjoying" state.
[0221] 6. Terminal: Sends the generated audio to the speaker and plays it.
[0222] 7. User: Confirm that the dog is happy by hearing the voice message "Happy!" coming from the dog's collar.
[0223] Examples of use and prompts for generative AI models
[0224] Usage example
[0225] If a security dog detects an intruder, it records the situation with a camera, analyzes any unusual barking sounds with a microphone, and emits a warning message saying, "An intruder has been detected!"
[0226] Example of a prompt:
[0227] basic information:
[0228] Animals used: Dogs
[0229] Usage scenario: Security guarding
[0230] detail:
[0231] Capture a real-time video stream with the camera.
[0232] The sounds are recorded in real time using a microphone.
[0233] Analyze the data using a TensorFlow model to detect anomalies.
[0234] If an anomaly is detected, the user will be notified with an audio alert.
[0235] In this way, the system detects not only the animals' emotions but also abnormal behavior in real time and notifies the user, thereby improving the safety of the animals and the user's sense of security.
[0236] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0237] Step 1:
[0238] Capture animal behavior and sounds from cameras and microphones.
[0239] Subject: terminal
[0240] Input: Animal movements and sounds
[0241] Output: Unprocessed video and audio data
[0242] Specific operation: The device acquires animal behavior and sounds in real time from multiple cameras and microphones installed indoors or outdoors. This includes capturing video and audio streams.
[0243] Step 2:
[0244] Preprocess the data acquired from the camera and microphone.
[0245] Subject: terminal
[0246] Input: Unprocessed video and audio data
[0247] Output: Pre-processed video and audio data
[0248] Specific operation: The terminal removes background noise from video data and filters audio data to reduce noise. This process uses OpenCV and audio filtering algorithms.
[0249] Step 3:
[0250] The pre-processed data is sent to the server.
[0251] Subject: terminal
[0252] Input: Pre-processed video and audio data
[0253] Output: Data sent to the server
[0254] Specific operation: The terminal sends pre-processed data to the server via the internet or local network. The data may be encrypted for security purposes during this process.
[0255] Step 4:
[0256] The server analyzes the data to estimate the animals' emotions and abnormal behaviors.
[0257] Subject: Server
[0258] Input: Pre-processed video and audio data
[0259] Output: Emotion and abnormal behavior estimation results
[0260] Specific operation: The server uses machine learning algorithms such as TensorFlow to analyze video and audio data and estimate emotions and abnormal behavior from animal behavior, gestures, and vocalizations. Specifically, it uses a trained model to classify emotional states such as "happy," "alert," and "anxious" and detect abnormal behavior.
[0261] Step 5:
[0262] It generates voices based on estimated emotions and abnormal behaviors.
[0263] Subject: Server
[0264] Input: Emotion and abnormal behavior estimation results
[0265] Output: Generated audio data
[0266] Specific operation: The server selects speech text corresponding to the estimated emotions and abnormal behavior, and converts the text into speech. This is done using speech synthesis software (such as the Google Text-to-Speech API).
[0267] Step 6:
[0268] The generated sound is emitted from a speaker attached to the animal's collar.
[0269] Subject: terminal
[0270] Input: Generated audio data
[0271] Output: Sound emitted from an animal collar
[0272] Specific operation: Audio data is generated and sent from the server to the terminal. The terminal then sends the audio data to a speaker attached to the animal's collar. The speaker plays the received audio data in real time.
[0273] Step 7:
[0274] The system notifies the user of abnormal behavior.
[0275] Subject: Server
[0276] Input: Abnormal behavior detection results
[0277] Output: Warning notification to the user
[0278] Specific operation: If abnormal behavior is detected, the server will send a warning message to the user's smartphone or other devices. This will be done using methods such as push notifications, email, or SMS. The user will receive the notification and be able to take prompt action.
[0279] This processing flow allows for the detection of not only animal emotions but also abnormal behavior in real time, and notifies the user, thereby improving animal safety and user peace of mind.
[0280] Furthermore, an emotion engine for estimating the user's emotion may be combined. That is, the specific processing unit 290 may estimate the user's emotion using the emotion recognition model 59 and perform specific processing using the user's emotion.
[0281] In addition to a system for estimating the emotion of an animal, the present invention is a system combined with an emotion engine that also recognizes the emotion of a user. The specific embodiments of the present invention will be described below.
[0282] Overview of the System
[0283] This system uses cameras and microphones installed in multiple locations in the room to acquire the actions, gestures, expressions, and vocalizations of animals, analyzes the data, and estimates the emotions of the animals. In addition, the gestures, expressions, and tone of the user's voice are also acquired using cameras and microphones to recognize the user's emotion. This recognized emotion generates appropriate response actions for the animals and provides them to the user as voice or other feedback, thereby promoting interaction.
[0284] Configuration of the System
[0285] 1. Cameras and Microphones:
[0286] Multiple cameras and microphones installed in the room acquire the actions, vocalizations, gestures, and expressions of animals and users in real time. The cameras continuously capture high-resolution video, and the microphones record high-sensitivity audio data.
[0287] 2. Preprocessing on the Terminal:
[0288] The terminal preprocesses the data acquired from the cameras and microphones. The video data is applied with a noise removal filter to extract the regions of interest of animals and users. The audio data removes background noise and converts it into a clear audio signal. Furthermore, preprocessing for analyzing the tone and language pattern of the user's voice is also performed.
[0289] 3. Analysis on the Server:
[0290] The server analyzes the pre-processed data sent from the terminal. First, it applies a model to the video data to identify the actions, gestures, and facial expressions of the animals and the user. This allows it to detect things like whether the cat is being petted or whether the user is smiling.
[0291] 4. Estimating animal emotions:
[0292] The server applies a speech recognition model to estimate emotions from animal sounds and gestures. It analyzes the patterns and phonemes of the sounds to identify emotions such as "happy" or "anxious."
[0293] 5. User emotion recognition:
[0294] The server analyzes the user's gestures, facial expressions, and tone of voice to recognize their emotions. For example, if the user is smiling, it recognizes that they are "having fun," and if they are frowning, it recognizes that they are "anxious."
[0295] 6. Speech generation and interaction proposals:
[0296] The server generates appropriate voice messages based on the animal's emotions and the user's emotions. Furthermore, it suggests appropriate interactions between the user and the animal. For example, if the animal appears anxious and the user is also anxious, the system will make a voice suggestion such as, "Please gently stroke the cat to reassure it."
[0297] 7. Audio playback via speaker:
[0298] The generated audio is transmitted via the device to a speaker attached to the animal's collar. Feedback audio is also transmitted from the device to the user's speaker.
[0299] Specific examples of actions
[0300] Example 1: When the cat is being petted and is happy
[0301] 1. Terminal: The camera records the user's action of petting the cat, and the microphone acquires the cat's cries.
[0302] 2. Terminal: Preprocess this data to remove unnecessary noise.
[0303] 3. Terminal: Send the preprocessed data to the server.
[0304] 4. Server: Analyze the data and estimate the "happy" state from the cat's gestures and cries.
[0305] 5. Server: On the other hand, analyze the user's expression and tone of voice to recognize that the user is in a "happy" state.
[0306] 6. Server: Generate the voice "Feeling good!" corresponding to the "happy" state, and at the same time, generate a voice to feedback to the user "The cat is happy".
[0307] 7. Terminal: Send the generated voice to the cat's collar and the user's speaker and play it.
[0308] 8. User: Hear the voice "Feeling good!" from the cat's collar, understand that the cat is happy, and confirm that they themselves are also happy.
[0309] Example 2: When the cat looks uneasy
[0310] 1. Terminal: The camera and microphone acquire the cat's startled actions and cries.
[0311] 2. Terminal: Preprocess this data.
[0312] 3. Terminal: Send the preprocessed data to the server.
[0313] 4. Server: Analyzes data and estimates whether the cat is "anxious" based on its behavior and meows.
[0314] 5. Server: On the other hand, it analyzes the user's facial expressions and tone of voice to recognize that the user is also "anxious."
[0315] 6. Server: Generates the voice message "What's wrong?" to indicate an anxious state, and simultaneously generates a voice message to the user saying, "Please gently stroke the cat to reassure it."
[0316] 7. Terminal: The generated audio is sent to the cat's collar and the user's speaker for playback.
[0317] 8. User: Hearing the sound from the speaker, the user understands that the cat is anxious and takes action to reassure the cat by gently stroking it.
[0318] In this way, the system can provide support for users and animals to understand each other's emotions and communicate smoothly.
[0319] The following describes the processing flow.
[0320] Step 1:
[0321] The device initializes the camera and microphone installed in the room and captures animal behavior, gestures, facial expressions, and vocalizations, as well as the user's gestures, facial expressions, and voice tone in real time. The camera continuously captures frame images, and the microphone records audio data.
[0322] Step 2:
[0323] The device preprocesses the acquired video and audio data. First, it removes noise from the video data and extracts the areas of interest for the animals and the user. Next, it filters the audio data to remove background noise and converts it into a clear audio signal.
[0324] Step 3:
[0325] The terminal packages pre-processed animal and user data and transmits it to a server via the internet. The data package includes video frames and filtered audio data.
[0326] Step 4:
[0327] The server receives the data package sent from the terminal and prepares it for analysis. It performs a consistency check on the received data to confirm that there are no missing or corrupted files.
[0328] Step 5:
[0329] The server analyzes the video data and applies models to identify animal behavior, gestures, and facial expressions, as well as user gestures and facial expressions. For example, it can detect whether an animal is being petted or whether the user is smiling.
[0330] Step 6:
[0331] The server analyzes audio data and applies a speech recognition model to estimate emotions from animal sounds. It also analyzes the tone and phonemes of the user's voice to identify their emotions.
[0332] Step 7:
[0333] The server integrates the analysis results of video and audio data to estimate the overall emotions of the animal and the user. For example, if the animal is showing signs of happiness, it is estimated to be "happy," and if the user is enjoying themselves, it is recognized as "enjoying themselves."
[0334] Step 8:
[0335] The server generates an appropriate voice message based on the estimated emotion. For example, if the server estimates that the animal is "happy," it will generate a voice message saying "Feels good!" and provide the user with feedback such as "The cat is happy."
[0336] Step 9:
[0337] The server sends the generated audio data to the terminal. This includes the audio message and associated metadata.
[0338] Step 10:
[0339] The device receives audio data from the server and transmits it to a speaker attached to the animal's collar, while simultaneously transmitting it to the user's speaker for playback.
[0340] Step 11:
[0341] Through voice messages emitted from the speaker, users can understand the emotions of animals and recognize how their own actions affect them. For example, they might hear the voice saying "Feels good!" from the animal's collar, confirm that the cat is happy, and become aware of their own enjoyment through the feedback "The cat is happy."
[0342] This allows users to understand the animals' emotions in real time and improve their relationship with them by interacting with them appropriately.
[0343] (Example 2)
[0344] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".
[0345] Conventional animal emotion estimation systems only estimate emotions based on animal behavior and vocalizations, and merely provide feedback to the animals. However, since the emotions of users living with animals directly influence the animals' behavior and state, a comprehensive analysis that includes the user's emotions is desirable. Furthermore, providing appropriate interaction suggestions based on the emotions of both the animal and the user can improve the relationship between the animal and the user. Against this backdrop, there is a need for a system that simultaneously analyzes the emotions of both animals and users and provides appropriate feedback and interaction suggestions.
[0346] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0347] In this invention, the server includes means for preprocessing data acquired from a camera and a microphone, means for analyzing the preprocessed data and estimating the emotions of the animal and the user, and means for generating speech based on the estimated emotions. This enables high-precision analysis of the emotions of the animal and the user, and allows for appropriate feedback and interaction suggestions based on the results.
[0348] A "camera" is an optical device used to capture the behavior, gestures, and facial expressions of animals and users.
[0349] A "microphone" is an audio device used to capture the voices of animals and users.
[0350] "Data preprocessing" refers to operations such as noise reduction and extraction of regions of interest on data acquired from cameras and microphones.
[0351] A "server" is an information processing device that analyzes pre-processed data, estimates the emotions of animals and users, and generates necessary feedback.
[0352] "Emotion estimation" is the process of inferring the emotional state of animals and users from acquired and pre-processed data.
[0353] "Voice generation" is the process of creating appropriate voice messages based on estimated emotions.
[0354] An "output device" is a device that emits the generated sound from an animal's collar or a speaker on the user's side.
[0355] "Interaction suggestion" is a process that proposes appropriate ways for animals and users to interact based on emotion estimation results.
[0356] This invention is a system that simultaneously analyzes the emotions of both animals and users, and provides appropriate feedback and interaction suggestions. Specific embodiments are described below.
[0357] First, multiple high-resolution cameras and high-sensitivity microphones installed in the room are used to capture the behavior, gestures, facial expressions, and voices of the animals and the user in real time. The cameras capture video at 30 frames per second, while the microphones collect audio data at a sampling rate of 20 kHz.
[0358] The terminal performs preprocessing on the acquired data. Specifically, it uses the open-source OpenCV library to denoise video data and extract regions of interest, and the librosa library to reduce background noise in audio data and convert it into a clear audio signal.
[0359] The pre-processed data is sent from the terminal to the server. The server uses the YOLOv3 model for image analysis and the Google Speech-to-Text API for speech analysis to analyze the received data. This analysis identifies the behavior, gestures, facial expressions, sounds, and voice tone of animals and users.
[0360] Next, the server estimates the emotions of both the animals and the users. For animal emotion estimation, a TensorFlow-based speech emotion recognition model is used, while for user emotion estimation, a combination of an image recognition model (e.g., OpenPose) and a speech emotion recognition model is used.
[0361] Based on the estimated emotions, the server generates an appropriate voice message. The Google Cloud Text-to-Speech API is used for voice generation. The generated voice message is sent to an output device attached to the animal's collar and to the user's output device (speaker) for playback.
[0362] As a concrete example, let's consider the case where a cat is enjoying being petted. In this case, the camera records the user's actions as they pet the cat, and the microphone captures the cat's meows. The device preprocesses this data and sends it to the server. The server analyzes the data, estimating the cat's "happy" state from its gestures and meows, while also analyzing the user's facial expressions and tone of voice to recognize that they are "enjoying" the experience. Based on this, the server generates the audio "That feels good!" and provides feedback to the user saying "The cat is happy." This audio is played from the animal's collar and the user's speaker.
[0363] As another example, consider a case where a cat appears anxious. In this case, the camera and microphone capture the cat's startled movements and meows, which the device preprocesses and sends to the server. The server estimates that the cat is "anxious" and similarly recognizes this from the user's facial expressions and tone of voice. The server generates a voice message saying "What's wrong?" and provides feedback to the user saying "Please gently stroke the cat to reassure it." This voice message is also played back from the output device.
[0364] Specific examples of prompt statements are as follows:
[0365] 1. "Explain how a system that uses a camera and microphone to analyze a cat's emotions provides feedback when the cat is happy while being petted."
[0366] 2. "Please describe the procedure for analyzing data from cats that appear surprised and anxious, and then suggesting appropriate interactions for the user."
[0367] As described above, this system analyzes the emotions of animals and users with high accuracy and improves the relationship between animals and users by providing appropriate feedback and interaction suggestions based on the results.
[0368] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0369] Step 1:
[0370] Data acquisition using camera and microphone
[0371] The device uses multiple high-resolution cameras and high-sensitivity microphones installed in the room to capture the behavior, gestures, facial expressions, and voices of animals and users in real time. Specifically, the cameras capture video at a rate of 30 frames per second, and the microphones collect audio data at a sampling rate of 20 kHz. This inputs data such as animal movements and sounds, and the user's facial expressions and voice tone. The output of this step is raw video and audio data.
[0372] Step 2:
[0373] Data preprocessing
[0374] The terminal performs preprocessing on the acquired raw video and audio data. Using the OpenCV library, it removes background noise from the video data and extracts the areas of interest for the animals and the user. It also uses the librosa library to reduce background noise from the audio data, converting it into a clear audio signal. Since the input video data contains noise, it undergoes filtering, and noise reduction is applied to the audio data. The output of this step is the preprocessed video and audio data.
[0375] Step 3:
[0376] Pre-processed data transmission to the server
[0377] The terminal sends pre-processed video and audio data to the server. The data is sent asynchronously using an HTTP POST request. As a specific example, compressed pre-processed data is sent to the server via a REST API. The input to this step is the pre-processed data, and the output is the data transfer to the server.
[0378] Step 4:
[0379] Server-based data analysis
[0380] The server analyzes the received preprocessed data. For image analysis, the YOLOv3 model is used to identify animal and user behavior, gestures, and facial expressions. For audio data analysis, the Google Speech-to-Text API is used to convert vocalizations and tone into text data for further analysis. The input for this step is preprocessed video and audio data, and the output is behavioral and emotion labels as analysis results. A concrete example of its operation is analyzing a cat's posture and vocalizations while being petted to identify that "the cat is happy."
[0381] Step 5:
[0382] Estimating animal emotions
[0383] The server estimates emotions from animal sounds and gestures. Using a TensorFlow-based speech emotion recognition model, it analyzes vocal patterns and phonemes to identify emotions such as "happy" or "anxious." The input for this step is the analyzed audio and video data, and the output is the animal's emotion label. A concrete example of its operation is analyzing a cat's high-pitched meow to estimate the emotion of "happiness."
[0384] Step 6:
[0385] User emotion recognition
[0386] The server analyzes the user's gestures, facial expressions, and voice tone to recognize their emotions. It uses a combination of an image recognition model (e.g., OpenPose) and a voice emotion recognition model. The input for this step is the analyzed video and audio data, and the output is the user's emotion label. A concrete example of its operation is detecting the user's smile and cheerful voice tone to identify a state of "enjoyment."
[0387] Step 7:
[0388] Speech generation and interaction proposals
[0389] The server generates an appropriate voice message based on the animal's emotions and the user's emotions. Using the Google Cloud Text-to-Speech API, the generated voice message is sent to an output device attached to the animal's collar and the user's output device (speaker) for playback. The input for this step is the emotion label, and the output is the generated voice message. A concrete example of its operation is the generation of a message such as "The cat is happy."
[0390] Step 8:
[0391] Audio playback and user feedback
[0392] The terminal transmits and plays the generated audio to an output device attached to the animal's collar and to the user's output device. The user listens to the audio from the speaker and takes appropriate action according to the animal's emotional state. The input for this step is the generated audio message, and the output is the playback of the audio. A concrete example of this operation would be the audio "Feels good!" being played from the cat's collar, and the message "The cat is happy" being played from the user's speaker.
[0393] (Application Example 2)
[0394] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."
[0395] In traditional pet shops, it was difficult for staff to accurately understand the emotions of customers and their pets, resulting in difficulties in recommending appropriate products and services. Understanding pets' emotions is particularly crucial for providing stress-free service, but effective methods for achieving this were lacking. Therefore, there is a growing need for a system that recognizes the emotions of customers and their pets in real time and provides optimal recommendations and feedback based on that understanding, in order to improve customer satisfaction.
[0396] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[0397] In this invention, the server includes means for acquiring animal behavior, gestures, facial expressions, and vocalizations from a camera and microphone; means for analyzing customer gestures, facial expressions, tone of voice, and content from the data to recognize customer emotions; and means for suggesting optimal products and services based on the emotions of the animals and customers. This makes it possible to recognize the emotions of customers and pets in real time within a pet shop and to provide optimal products and services based on that recognition.
[0398] A "camera" is a device that acquires visual information and converts it into digital images or video data.
[0399] A "microphone" is a device that captures sound and converts it into digital audio data.
[0400] "Data preprocessing" is the process of removing noise from acquired data and converting it into a format suitable for analysis.
[0401] "Means for estimating animal emotions" refers to algorithms or devices that identify and estimate an animal's emotional state based on its behavior, gestures, facial expressions, and vocalizations.
[0402] "Means for recognizing customer emotions" refers to an algorithm or device that identifies and recognizes a customer's emotional state based on their gestures, facial expressions, tone of voice, and content.
[0403] "Means of proposing optimal products and services" refers to algorithms or devices that select and propose appropriate products and services based on the emotional state of animals and customers.
[0404] "Means for generating speech" refers to algorithms or devices that create natural-sounding speech from text.
[0405] A "speaker" is a device that reproduces digital audio data as sound.
[0406] A "server" is a computing device or a system on a network that provides data reception, storage, and analysis functions.
[0407] The embodiments for carrying out this invention will be described in detail below. As an example of the embodiment of the invention, we will particularly illustrate its application to a system aimed at improving customer service in a pet shop.
[0408] System Configuration
[0409] This system recognizes the emotions of customers and their pets in real time and suggests the most suitable products and services based on that. The system includes a camera, microphone, terminal, server, and speaker.
[0410] Hardware configuration
[0411] 1. Camera
[0412] High-resolution cameras (e.g., Logitech C920) capture the behavior, gestures, and facial expressions of customers and pets inside the store.
[0413] 2. Mike
[0414] High-sensitivity microphones (e.g., Blue Yeti) can capture the voices of customers or pets.
[0415] 3. Terminal
[0416] Use a laptop or mobile device (e.g., iPhone®) to preprocess data from the camera and microphone.
[0417] 4. Server
[0418] High-performance servers (e.g., AWS EC2) analyze pre-processed data to estimate the emotions of animals and customers.
[0419] 5. Speakers
[0420] Speakers installed in the store and speakers attached to pet collars play sounds generated by the system.
[0421] Software Configuration
[0422] 1. Preprocessing
[0423] The device performs noise reduction on data acquired from the camera and microphone, and detects faces and voices. Specifically, it uses OpenCV and the SpeechRecognition library.
[0424] 2. Data Analysis
[0425] The server uses the DeepFace library to perform facial recognition and estimate customer emotions in order to analyze pre-processed data sent from the terminal. It also uses Google's Speech-to-Text API for speech analysis.
[0426] 3. Emotion estimation
[0427] The server estimates the emotions of animals and customers based on the analyzed data. It identifies emotions such as "happy" or "anxious" from the animals' vocalizations and gestures, and identifies emotions such as "enjoying" or "anxious" from the customers' tone of voice and facial expressions.
[0428] 4. Suggestions and Feedback
[0429] Based on the emotions of animals and customers, the server suggests the most suitable products and services and generates voice messages. Voice generation uses gTTS (Google Text-to-Speech) and is played back using the Playsound library.
[0430] Specific example of processing
[0431] For example, if a customer visits a pet shop with their pet and is looking at products, the system will perform the following actions.
[0432] Example of a prompt
[0433] Camera image analysis prompt:
[0434] "Analyze the dominant emotion from this frame. Frame: {frame_data}"
[0435] Voice analysis prompt:
[0436] "Estimate the customer's emotions from their voice. Voice data: {audio_data}"
[0437] Based on camera footage and audio data, the system analyzes the emotions of the customer and their pet, understanding situations in real time, such as "the pet is interested in a toy" or "the customer is looking for new food." Based on this, the system generates voice messages such as "Your pet is interested in this toy. How about this toy?" and provides feedback to the customer through the speaker.
[0438] This improves the purchasing experience because customers can more easily understand their pet's condition and be offered appropriate products and services.
[0439] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0440] Step 1:
[0441] The device uses a high-resolution camera and a high-sensitivity microphone to acquire video and audio data of customers and pets in the store. The input consists of video data from the camera and audio data from the microphone, and this data is sent to the next step.
[0442] Step 2:
[0443] The video and audio data acquired by the device are preprocessed. Specifically, for video data, noise is removed using OpenCV, and faces and animal gestures are detected. For audio data, noise is removed using the SpeechRecognition library, and it is converted into clear audio data. The input is the video and audio data acquired in step 1, and the output is the preprocessed data.
[0444] Step 3:
[0445] The terminal sends pre-processed data to the server. The input is pre-processed data, which is sent from the terminal to the server via the network.
[0446] Step 4:
[0447] The server receives pre-processed data and analyzes the video data using the DeepFace library to estimate emotions from the customer's facial expressions. The input is video data sent from the terminal, and the output is the result of the customer's emotion recognition. The prompt message "Customer's face frame: {frame_data}" is used.
[0448] Step 5:
[0449] The server converts audio data into text using Google's Speech-to-Text API, and then analyzes the content and tone of voice to recognize the customer's emotions. The input is pre-processed audio data, and the output is the emotion recognition result based on speech recognition. The prompt message used is "Audio input data: {audio_data}".
[0450] Step 6:
[0451] The server analyzes animal video and audio data to determine their behavior, gestures, and vocalizations, and estimates their emotions. The input is video and audio data transmitted from a terminal, and the output is the estimated animal emotions. DeepFace and an audio analysis model are used.
[0452] Step 7:
[0453] The server generates messages suggesting optimal products and services based on animal emotion estimation results and customer emotion recognition results. gTTS is used to convert text to speech during this process. The input is the emotion recognition and emotion estimation results, and the output is the generated speech message.
[0454] Step 8:
[0455] The device sends the generated voice message to a speaker, which plays it back through the store's speakers and a speaker attached to the pet's collar. The input is the generated voice message, and the output is voice feedback to the customer and their pet.
[0456] Step 9:
[0457] The user listens to feedback from the speaker and selects appropriate products or services based on the system's suggestions. The input is voice feedback, and the output is the user's actions.
[0458] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0459] Data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of data generation model 58 is ChatGPT (registered trademark) (Internet search).<URL: https: / / openai.com / blog / chatgpt> ), Gemini (registered trademark) (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0460] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart device 14.
[0461] [Second Embodiment]
[0462] Figure 3 shows an example of the configuration of the data processing system 210 according to the second embodiment.
[0463] As shown in Figure 3, the data processing system 210 includes a data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.
[0464] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0465] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication interface 44. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, and camera 42 are also connected to the bus 52.
[0466] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[0467] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).
[0468] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[0469] Figure 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Figure 4, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[0470] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0471] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0472] In the smart glasses 214, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0473] Next, the identification processing performed by the identification processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".
[0474] This invention is a system that uses a camera and microphone to acquire the behavior, gestures, facial expressions, and vocalizations of animals, analyzes this data to estimate the animals' emotions, and expresses the estimated emotions in sound. Specific embodiments for carrying out this invention are described below.
[0475] System Overview
[0476] This system uses multiple cameras and microphones placed in a room to capture animal behavior and vocalizations in real time, and sends the data to a server for analysis. Based on the analysis results, sounds are generated and emitted from speakers attached to the animals' collars, allowing users to understand the animals' emotions.
[0477] System Configuration
[0478] 1. Camera and microphone:
[0479] Multiple cameras and microphones installed in the room capture animal behavior and sounds in real time. The cameras continuously record animal movements, while the microphones record sounds such as vocalizations.
[0480] 2. Preprocessing at the terminal:
[0481] The device preprocesses the data acquired from the camera and microphone. Specifically, it removes background noise from the video and filters the audio data to reduce noise. This preprocessed data is then sent to the server for further analysis.
[0482] 3. Analysis on the server:
[0483] The server analyzes pre-processed data sent from the terminal. It identifies animal behavior, gestures, and facial expressions from image data, and analyzes animal sounds from audio data. This allows it to estimate the animal's emotions.
[0484] 4. Speech generation:
[0485] Based on the emotions of the animals estimated by the server, the appropriate voice text is selected and generated as speech. For example, if the emotion "happy" is estimated, the corresponding speech will be "Feels good!".
[0486] 5. Audio playback via speaker:
[0487] The generated audio is transmitted via a device to a speaker attached to the animal's collar. The audio emitted from the speaker allows the user to understand the animal's emotions in real time.
[0488] Specific examples of actions
[0489] Example 1: When a cat is happy to be petted
[0490] 1. Device: The camera records the user's actions as they pet the cat, and the microphone captures the cat's meows.
[0491] 2. Terminal: Preprocess this data to remove unwanted noise.
[0492] 3. Terminal: Sends pre-processed data to the server.
[0493] 4. Server: Analyzes data and estimates whether the cat is "happy" based on its gestures and meows.
[0494] 5. Server: Generates the voice "Feels good!" which corresponds to the "feeling happy" state.
[0495] 6. Terminal: Sends the generated audio to the speaker and plays it.
[0496] 7. User: Hearing the voice saying "Feels good!" from the cat's collar, the user understands that the cat is happy.
[0497] Example 2: When the cat looks anxious
[0498] 1. Device: The camera and microphone capture the cat's actions and meows when it is startled by something.
[0499] 2. Terminal: Preprocess this data.
[0500] 3. Terminal: Sends pre-processed data to the server.
[0501] 4. Server: Analyzes data and estimates whether the cat is "anxious" based on its behavior and meows.
[0502] 5. Server: Generates the voice message "What's wrong?" which corresponds to the state of "feeling anxious".
[0503] 6. Terminal: Sends the generated audio to the speaker and plays it.
[0504] 7. User: Hearing the voice saying "What's wrong?" from the cat's collar, the user understands that the cat is anxious and takes action to soothe it.
[0505] In this way, the system provides support for better communication between animals and their owners, and makes it possible to understand pets' emotions more accurately and in real time.
[0506] The following describes the processing flow.
[0507] Step 1:
[0508] The device initializes the camera and microphone installed in the room and captures the animal's behavior, gestures, facial expressions, and vocalizations in real time. Specifically, it continuously captures frame images from the camera and records audio data from the microphone.
[0509] Step 2:
[0510] The terminal preprocesses the acquired video and audio data. The video data is filtered to remove noise and extract the animal's areas of interest. The audio data is processed to remove background noise and convert it into a clear audio signal.
[0511] Step 3:
[0512] The terminal packages the pre-processed data and sends it to the server via the internet. The data package includes video frames and filtered audio data.
[0513] Step 4:
[0514] The server receives the data package sent from the terminal and prepares it for analysis. It performs a consistency check on the received data to confirm that there are no missing or corrupted files.
[0515] Step 5:
[0516] The server analyzes the video data and applies models to identify animal behavior, gestures, and facial expressions. For example, it can detect the posture a cat is in when being petted, or specific gestures.
[0517] Step 6:
[0518] The server analyzes the audio data and applies a speech recognition model to estimate emotions from animal sounds. It analyzes the patterns and phonemes of the sounds to identify emotions.
[0519] Step 7:
[0520] The server integrates the analysis results of video and audio data to estimate the animal's overall emotions. For example, if the animal shows signs of happiness, it will be estimated to be "happy."
[0521] Step 8:
[0522] The server generates an appropriate voice message based on the estimated emotion. For example, if the emotion is "happy," it will generate the voice message "Feels good!"
[0523] Step 9:
[0524] The server sends the generated audio data to the terminal. This includes the audio message and associated metadata.
[0525] Step 10:
[0526] The device receives audio data from the server and transmits it to a speaker attached to the animal's collar, playing the audio.
[0527] Step 11:
[0528] Users can understand animal emotions through voice messages emitted from the speaker. For example, by hearing the voice say "That feels good!", they can recognize that the cat is happy.
[0529] This allows users to understand the animals' emotions in real time and engage in appropriate dialogue and responses.
[0530] (Example 1)
[0531] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."
[0532] Conventional pet emotion understanding systems have difficulty accurately estimating emotions from animal behavior and vocalizations, potentially leading owners to misunderstand their pets' feelings. Furthermore, the time required for real-time emotion estimation and feedback hinders smooth communication between pets and their owners. To address these issues, the present invention provides a system that accurately and in real-time estimates animal emotions and expresses those emotions verbally.
[0533] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0534] In this invention, the server includes means for acquiring animal behavior, gestures, facial expressions, and vocalizations from a camera and a microphone; means for preprocessing the data acquired from the camera and microphone; means for transmitting the preprocessed data to the server; means for analyzing the preprocessed data and estimating the animal's emotions; means for generating sound based on the estimated emotions; and means for emitting the generated sound from a speaker attached to the animal's collar. This makes it possible to estimate the animal's emotions with high accuracy and in real time, and to express those emotions in sound.
[0535] A "camera" is a device used to capture images of an animal's behavior, gestures, and facial expressions as video data.
[0536] A "microphone" is a sound-collecting device used to capture animal sounds and ambient noises as audio data.
[0537] "Data preprocessing" is the process of removing noise from video and audio data acquired from cameras and microphones, and preparing it in a format that is easy to analyze.
[0538] A "server" is a central processing unit that receives data transmitted from terminals, performs analysis, and estimates emotions.
[0539] "Analysis" refers to the process of identifying and analyzing animal behavior and vocalizations based on pre-processed data.
[0540] "Emotion estimation" is the process of identifying an animal's emotions based on the results of an analysis.
[0541] "Voice generation" is the process of selecting appropriate voice text to express estimated emotions and generating it as voice data.
[0542] A "speaker" is an output device that plays back generated audio data as sound.
[0543] A "terminal" is a device that preprocesses data acquired from cameras and microphones and sends it to a server.
[0544] This invention is a system that uses a camera and microphone to acquire animal behavior, gestures, facial expressions, and vocalizations in real time, analyzes this data to estimate the animal's emotions, and expresses those estimated emotions in sound. This system is designed to make it easier for users to understand animal emotions. Its specific form is described below.
[0545] Hardware and software to be used
[0546] Camera: Used to continuously record the behavior and facial expressions of animals. Examples include the Logitec HD Pro Webcam.
[0547] Microphone: Used to record animal sounds and ambient noise. Examples include the Blue Yeti USB Mic.
[0548] Terminal: A small PC used for data preprocessing and communication with the server. For example, a Raspberry Pi can be used.
[0549] Server: A central processing unit that performs data analysis and sentiment estimation. Examples include AWS and Google Cloud Platform.
[0550] Generative AI models: AI / ML frameworks used for data analysis and sentiment estimation. Specific examples include TensorFlow and PyTorch.
[0551] Text-to-speech software: Generates speech based on estimated emotions. Examples include Amazon Polly and Google Text-to-Speech.
[0552] Speaker: An output device for playing back the generated sound, attached to the animal's collar.
[0553] Specific examples of actions
[0554] Example 1: When a cat is happy to be petted
[0555] 1. The device uses a camera to record the user's actions as they pet the cat, and a microphone to capture the cat's meows.
[0556] 2. The terminal preprocesses this data to remove unwanted noise.
[0557] 3. The terminal sends the pre-processed data to the server.
[0558] 4. The server analyzes the data and estimates whether the cat is "happy" based on its gestures and meows.
[0559] 5. The server generates the voice "Feels good!" which corresponds to the state of being "happy".
[0560] 6. The device sends the generated audio to the speaker and plays it.
[0561] 7. The user hears the cat say "Feels good!" from the cat's collar and understands that the cat is happy.
[0562] Example 2: When the cat looks anxious
[0563] 1. The device uses its camera and microphone to capture any actions or sounds the cat makes when it is startled by something.
[0564] 2. The terminal preprocesses this data.
[0565] 3. The terminal sends the pre-processed data to the server.
[0566] 4. The server analyzes the data and estimates whether the cat is "anxious" based on its behavior and meows.
[0567] 5. The server generates the voice message "What's wrong?" which corresponds to the state of being "anxious".
[0568] 6. The device sends the generated audio to the speaker and plays it.
[0569] 7. The user hears a voice saying "What's wrong?" from the cat's collar, understands that the cat is anxious, and takes action to soothe it.
[0570] Example of a prompt
[0571] "Analyze animal behavioral and auditory data to estimate the animal's emotions and generate appropriate voice-to-text. Specifically, if the animal is happy, generate the voice 'Feels good!', and if it is anxious, generate the voice 'What's wrong?'."
[0572] In this way, the present invention provides support for better communication between animals and their owners, and makes it possible to understand pets' emotions more accurately and in real time.
[0573] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0574] Step 1:
[0575] Data acquisition
[0576] The device uses a camera and microphone installed in the room to capture animal behavior and sounds in real time. The camera captures video of the animals, and the microphone records their audio. Specifically, the camera captures video data frame by frame, and the microphone continuously records audio.
[0577] Input: Animal behavior and sounds
[0578] Output: Raw video data and raw audio data
[0579] Step 2:
[0580] Data preprocessing
[0581] The device preprocesses the acquired video and audio data. Specifically, it removes background noise from the video data and filters the audio data to reduce unwanted noise. For video data, it masks parts other than animals, and for audio data, it applies filtering techniques to remove ambient noise.
[0582] Input: Raw video data and raw audio data
[0583] Output: Pre-processed video data and pre-processed audio data
[0584] Step 3:
[0585] Data transmission
[0586] The terminal sends pre-processed data to the server. This pre-processed video and audio data are transmitted to the server in real time using a high-speed network connection.
[0587] Input: Pre-processed video data and pre-processed audio data
[0588] Output: Data sent to the server
[0589] Step 4:
[0590] Data Analysis
[0591] The server analyzes the received pre-processed data. Specifically, it uses a generative AI model (e.g., TensorFlow or PyTorch) to identify animal behavior, gestures, and facial expressions from video data, and analyze animal sounds from audio data. Based on prompts, it obtains the analysis results.
[0592] Input: Pre-processed video data and pre-processed audio data
[0593] Output: Behavioral analysis results and vocalization analysis results
[0594] Step 5:
[0595] Emotion estimation
[0596] The server estimates the animal's emotions based on the results of data analysis. Specifically, it combines behavioral analysis results and vocalization analysis results to estimate the animal's current emotional state (e.g., "happy," "anxious").
[0597] Input: Behavioral analysis results and vocalization analysis results
[0598] Output: Estimated emotional state
[0599] Step 6:
[0600] Speech generation
[0601] The server generates speech based on estimated emotions. Specifically, it selects appropriate speech text corresponding to the emotional state and generates speech data using speech synthesis software (such as Amazon Polly or Google Text-to-Speech).
[0602] Input: Estimated emotional state
[0603] Output: Generated audio data
[0604] Step 7:
[0605] Audio transmission and playback
[0606] The device receives audio data from the server and transmits it to a speaker attached to the animal's collar, which then plays the audio. Specifically, the device transmits the audio data to the playback speaker via Bluetooth or Wi-Fi, and the speaker outputs the audio.
[0607] Input: Generated audio data
[0608] Output: Audio played through the speaker
[0609] Through these steps, users can understand the animals' emotions in real time, enabling better communication with them.
[0610] (Application Example 1)
[0611] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."
[0612] Conventional animal emotion estimation systems can estimate an animal's emotions, but they lack the ability to detect abnormal behavior and notify the user in real time. As a result, it is difficult for pet owners and caregivers to respond quickly to abnormal animal behavior. Furthermore, safety issues remain unresolved. This invention aims to improve animal safety and user peace of mind by detecting not only animal emotions but also abnormal behavior and notifying the user.
[0613] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0614] In this invention, the server includes means for acquiring animal behavior, gestures, facial expressions, and vocalizations from a camera and a microphone; means for preprocessing the data acquired from the camera and microphone; means for analyzing the preprocessed data and estimating the animal's emotions; means for generating sound based on the estimated emotions; means for emitting the generated sound from a speaker attached to the animal's collar; means for analyzing the data to detect abnormal animal behavior; and means for generating a warning based on the detected abnormal behavior and notifying the user. This makes it possible to detect both animal emotions and abnormal behavior in real time and notify the user.
[0615] A "camera" is a device used to capture images of an animal's behavior, gestures, and facial expressions.
[0616] A "microphone" is a sound acquisition device used to record animal sounds.
[0617] "Data preprocessing" refers to the process of removing noise and converting the data format of video and audio data acquired from cameras and microphones.
[0618] "Data analysis" refers to algorithmic processing that identifies animal behavior, gestures, and vocalizations from pre-processed video and audio data, and estimates the animal's emotions and abnormal behavior.
[0619] "Emotion estimation" is the process of inferring what emotional state an animal is in based on its behavior, gestures, and vocalizations.
[0620] "Voice generation" is the process of generating voice messages that convey emotions to the user based on estimated animal emotions.
[0621] A "speaker attached to an animal's collar" is an audio playback device installed on the collar worn by an animal, used to play generated audio messages.
[0622] "Detection of abnormal behavior" is the process of determining whether an animal is behaving in an unusual way based on data analysis.
[0623] "Notification to users" refers to the process of sending warning messages or notifications to users in real time based on detected abnormal behavior.
[0624] "Real-time" is a time standard that means processing or notifications are performed almost instantaneously or with very little delay.
[0625] The present invention is a system that uses a camera and microphone to acquire animal behavior, gestures, facial expressions, and vocalizations, analyzes this data to estimate the animal's emotions and abnormal behavior, and notifies the user as necessary. Specific embodiments for carrying out the present invention are described below.
[0626] System Overview
[0627] This system uses multiple cameras and microphones placed in a room or outdoors to capture animal behavior and vocalizations in real time, and transmits the data to a server for analysis. Based on the analysis results, it emits sounds from a speaker attached to the animal's collar, allowing the user to understand the animal's emotions. Furthermore, it has a function to notify the user of any abnormal behavior if it is detected.
[0628] System Configuration
[0629] 1. Camera and microphone
[0630] Multiple cameras and microphones, installed indoors or outdoors, capture animal behavior and sounds in real time. The cameras continuously record animal movements, while the microphones record sounds such as vocalizations.
[0631] 2. Preprocessing at the terminal
[0632] The device preprocesses the data acquired from the camera and microphone. Specifically, it removes background noise from the video and filters the audio data to reduce noise. This preprocessed data is then sent to the server for further analysis.
[0633] 3. Analysis on the server
[0634] The server analyzes pre-processed data sent from the terminal. It identifies animal behavior, gestures, and facial expressions from image data, and analyzes vocalizations from audio data. This allows it to estimate the animal's emotions and abnormal behavior. The server performs the analysis using machine learning algorithms such as TensorFlow.
[0635] 4. Speech generation
[0636] Based on the emotions of the animal estimated by the server, the appropriate voice text is selected and generated as speech. For example, if the emotion "happy" is estimated, the corresponding speech will be "Feels good!". In addition, if abnormal behavior is detected, a warning voice such as "Be careful!" will be generated.
[0637] 5. Audio playback via speaker
[0638] The generated audio is transmitted via a device to a speaker attached to the animal's collar. The audio emitted from the speaker allows the user to understand the animal's emotions and any abnormal behavior in real time.
[0639] Specific examples of actions
[0640] Example 1: When a dog senses a suspicious person and becomes alert.
[0641] 1. Device: The camera and microphone capture the dog's alert behavior and barking.
[0642] 2. Terminal: Preprocess this data to remove unwanted noise.
[0643] 3. Terminal: Sends pre-processed data to the server.
[0644] 4. Server: Analyzes data, estimates a "vigilant" state from the dog's movements and barks, and further detects abnormal behavior.
[0645] 5. Server: Generates the voice message "Be careful!" corresponding to the "Alert" state.
[0646] 6. Terminal: Sends the generated audio to the speaker and plays it.
[0647] 7. User: Hearing the "Be careful!" voice message from the dog's collar, the user understands that the dog is on alert and takes appropriate action.
[0648] Example 2: When the dog is playing safely
[0649] 1. Device: The camera and microphone capture the dog's playful behavior and barking.
[0650] 2. Terminal: Preprocess this data.
[0651] 3. Terminal: Sends pre-processed data to the server.
[0652] 4. Server: Analyzes data and estimates whether the dog is "enjoying" itself based on its movements and barks.
[0653] 5. Server: Generates the voice message "This is fun!" which corresponds to the "enjoying" state.
[0654] 6. Terminal: Sends the generated audio to the speaker and plays it.
[0655] 7. User: Confirm that the dog is happy by hearing the voice message "Happy!" coming from the dog's collar.
[0656] Examples of use and prompts for generative AI models
[0657] Usage example
[0658] If a security dog detects an intruder, it records the situation with a camera, analyzes any unusual barking sounds with a microphone, and emits a warning message saying, "An intruder has been detected!"
[0659] Example of a prompt:
[0660] basic information:
[0661] Animals used: Dogs
[0662] Usage scenario: Security guarding
[0663] detail:
[0664] Capture a real-time video stream with the camera.
[0665] The sounds are recorded in real time using a microphone.
[0666] Analyze the data using a TensorFlow model to detect anomalies.
[0667] If an anomaly is detected, the user will be notified with an audio alert.
[0668] In this way, the system detects not only the animals' emotions but also abnormal behavior in real time and notifies the user, thereby improving the safety of the animals and the user's sense of security.
[0669] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0670] Step 1:
[0671] Capture animal behavior and sounds from cameras and microphones.
[0672] Subject: terminal
[0673] Input: Animal movements and sounds
[0674] Output: Unprocessed video and audio data
[0675] Specific operation: The device acquires animal behavior and sounds in real time from multiple cameras and microphones installed indoors or outdoors. This includes capturing video and audio streams.
[0676] Step 2:
[0677] Preprocess the data acquired from the camera and microphone.
[0678] Subject: terminal
[0679] Input: Unprocessed video and audio data
[0680] Output: Pre-processed video and audio data
[0681] Specific operation: The terminal removes background noise from video data and filters audio data to reduce noise. This process uses OpenCV and audio filtering algorithms.
[0682] Step 3:
[0683] The pre-processed data is sent to the server.
[0684] Subject: terminal
[0685] Input: Pre-processed video and audio data
[0686] Output: Data sent to the server
[0687] Specific operation: The terminal sends pre-processed data to the server via the internet or local network. The data may be encrypted for security purposes during this process.
[0688] Step 4:
[0689] The server analyzes the data to estimate the animals' emotions and abnormal behaviors.
[0690] Subject: Server
[0691] Input: Pre-processed video and audio data
[0692] Output: Emotion and abnormal behavior estimation results
[0693] Specific operation: The server uses machine learning algorithms such as TensorFlow to analyze video and audio data and estimate emotions and abnormal behavior from animal behavior, gestures, and vocalizations. Specifically, it uses a trained model to classify emotional states such as "happy," "alert," and "anxious" and detect abnormal behavior.
[0694] Step 5:
[0695] It generates voices based on estimated emotions and abnormal behaviors.
[0696] Subject: Server
[0697] Input: Emotion and abnormal behavior estimation results
[0698] Output: Generated audio data
[0699] Specific operation: The server selects speech text corresponding to the estimated emotions and abnormal behavior, and converts the text into speech. This is done using speech synthesis software (such as the Google Text-to-Speech API).
[0700] Step 6:
[0701] The generated sound is emitted from a speaker attached to the animal's collar.
[0702] Subject: terminal
[0703] Input: Generated audio data
[0704] Output: Sound emitted from an animal collar
[0705] Specific operation: Audio data is generated and sent from the server to the terminal. The terminal then sends the audio data to a speaker attached to the animal's collar. The speaker plays the received audio data in real time.
[0706] Step 7:
[0707] The system notifies the user of abnormal behavior.
[0708] Subject: Server
[0709] Input: Abnormal behavior detection results
[0710] Output: Warning notification to the user
[0711] Specific operation: If abnormal behavior is detected, the server will send a warning message to the user's smartphone or other devices. This will be done using methods such as push notifications, email, or SMS. The user will receive the notification and be able to take prompt action.
[0712] This processing flow allows for the detection of not only animal emotions but also abnormal behavior in real time, and notifies the user, thereby improving animal safety and user peace of mind.
[0713] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0714] This invention is a system that combines a system for estimating animal emotions with an emotion engine that recognizes user emotions. Specific embodiments of this invention are described below.
[0715] System Overview
[0716] This system uses multiple cameras and microphones installed in a room to capture animal behavior, gestures, facial expressions, and vocalizations, and analyzes this data to estimate the animal's emotions. It also uses cameras and microphones to capture the user's gestures, facial expressions, and voice tone, recognizing the user's emotions. This recognized emotion generates appropriate responses to the animals, which are then provided to the user as voice or other feedback, thereby facilitating interaction.
[0717] System Configuration
[0718] 1. Camera and microphone:
[0719] Multiple cameras and microphones installed in the room capture the behavior, sounds, gestures, and facial expressions of both animals and users in real time. The cameras continuously capture high-resolution video, while the microphones record high-sensitivity audio data.
[0720] 2. Preprocessing at the terminal:
[0721] The device preprocesses data acquired from the camera and microphone. Video data is filtered to remove noise and extract areas of interest for both the animal and the user. Audio data is processed to remove background noise and convert it into a clear audio signal. Furthermore, preprocessing is performed to analyze the user's voice tone and language patterns.
[0722] 3. Analysis on the server:
[0723] The server analyzes the pre-processed data sent from the terminal. First, it applies a model to the video data to identify the actions, gestures, and facial expressions of the animals and the user. This allows it to detect things like whether the cat is being petted or whether the user is smiling.
[0724] 4. Estimating animal emotions:
[0725] The server applies a speech recognition model to estimate emotions from animal sounds and gestures. It analyzes the patterns and phonemes of the sounds to identify emotions such as "happy" or "anxious."
[0726] 5. User emotion recognition:
[0727] The server analyzes the user's gestures, facial expressions, and tone of voice to recognize their emotions. For example, if the user is smiling, it recognizes that they are "having fun," and if they are frowning, it recognizes that they are "anxious."
[0728] 6. Speech generation and interaction proposals:
[0729] The server generates appropriate voice messages based on the animal's emotions and the user's emotions. Furthermore, it suggests appropriate interactions between the user and the animal. For example, if the animal appears anxious and the user is also anxious, the system will make a voice suggestion such as, "Please gently stroke the cat to reassure it."
[0730] 7. Audio playback via speaker:
[0731] The generated audio is transmitted via the device to a speaker attached to the animal's collar. Feedback audio is also transmitted from the device to the user's speaker.
[0732] Specific examples of actions
[0733] Example 1: When a cat is happy to be petted
[0734] 1. Device: The camera records the user's actions as they pet the cat, and the microphone captures the cat's meows.
[0735] 2. Terminal: Preprocess this data to remove unwanted noise.
[0736] 3. Terminal: Sends pre-processed data to the server.
[0737] 4. Server: Analyzes data and estimates whether the cat is "happy" based on its gestures and meows.
[0738] 5. Server: On the other hand, it analyzes the user's facial expressions and tone of voice to recognize when the user is "enjoying" the experience.
[0739] 6. Server: Generates the voice "Feels good!" corresponding to the "happy" state, and simultaneously generates a voice message to the user saying "The cat is happy."
[0740] 7. Terminal: The generated audio is sent to the cat's collar and the user's speaker for playback.
[0741] 8. User: Hearing the voice saying "Feels good!" from the cat's collar, the user understands that the cat is happy and confirms that they themselves are also enjoying it.
[0742] Example 2: When the cat looks anxious
[0743] 1. Device: The camera and microphone capture the cat's startled movements and meows.
[0744] 2. Terminal: Preprocess this data.
[0745] 3. Terminal: Sends pre-processed data to the server.
[0746] 4. Server: Analyzes data and estimates whether the cat is "anxious" based on its behavior and meows.
[0747] 5. Server: On the other hand, it analyzes the user's facial expressions and tone of voice to recognize that the user is also "anxious."
[0748] 6. Server: Generates the voice message "What's wrong?" to indicate an anxious state, and simultaneously generates a voice message to the user saying, "Please gently stroke the cat to reassure it."
[0749] 7. Terminal: The generated audio is sent to the cat's collar and the user's speaker for playback.
[0750] 8. User: Hearing the sound from the speaker, the user understands that the cat is anxious and takes action to reassure the cat by gently stroking it.
[0751] In this way, the system can provide support for users and animals to understand each other's emotions and communicate smoothly.
[0752] The following describes the processing flow.
[0753] Step 1:
[0754] The device initializes the camera and microphone installed in the room and captures animal behavior, gestures, facial expressions, and vocalizations, as well as the user's gestures, facial expressions, and voice tone in real time. The camera continuously captures frame images, and the microphone records audio data.
[0755] Step 2:
[0756] The device preprocesses the acquired video and audio data. First, it removes noise from the video data and extracts the areas of interest for the animals and the user. Next, it filters the audio data to remove background noise and converts it into a clear audio signal.
[0757] Step 3:
[0758] The terminal packages pre-processed animal and user data and transmits it to a server via the internet. The data package includes video frames and filtered audio data.
[0759] Step 4:
[0760] The server receives the data package sent from the terminal and prepares it for analysis. It performs a consistency check on the received data to confirm that there are no missing or corrupted files.
[0761] Step 5:
[0762] The server analyzes the video data and applies models to identify animal behavior, gestures, and facial expressions, as well as user gestures and facial expressions. For example, it can detect whether an animal is being petted or whether the user is smiling.
[0763] Step 6:
[0764] The server analyzes audio data and applies a speech recognition model to estimate emotions from animal sounds. It also analyzes the tone and phonemes of the user's voice to identify their emotions.
[0765] Step 7:
[0766] The server integrates the analysis results of video and audio data to estimate the overall emotions of the animal and the user. For example, if the animal is showing signs of happiness, it is estimated to be "happy," and if the user is enjoying themselves, it is recognized as "enjoying themselves."
[0767] Step 8:
[0768] The server generates an appropriate voice message based on the estimated emotion. For example, if the server estimates that the animal is "happy," it will generate a voice message saying "Feels good!" and provide the user with feedback such as "The cat is happy."
[0769] Step 9:
[0770] The server sends the generated audio data to the terminal. This includes the audio message and associated metadata.
[0771] Step 10:
[0772] The device receives audio data from the server and transmits it to a speaker attached to the animal's collar, while simultaneously transmitting it to the user's speaker for playback.
[0773] Step 11:
[0774] Through voice messages emitted from the speaker, users can understand the emotions of animals and recognize how their own actions affect them. For example, they might hear the voice saying "Feels good!" from the animal's collar, confirm that the cat is happy, and become aware of their own enjoyment through the feedback "The cat is happy."
[0775] This allows users to understand the animals' emotions in real time and improve their relationship with them by interacting with them appropriately.
[0776] (Example 2)
[0777] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".
[0778] Conventional animal emotion estimation systems only estimate emotions based on animal behavior and vocalizations, and merely provide feedback to the animals. However, since the emotions of users living with animals directly influence the animals' behavior and state, a comprehensive analysis that includes the user's emotions is desirable. Furthermore, providing appropriate interaction suggestions based on the emotions of both the animal and the user can improve the relationship between the animal and the user. Against this backdrop, there is a need for a system that simultaneously analyzes the emotions of both animals and users and provides appropriate feedback and interaction suggestions.
[0779] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0780] In this invention, the server includes means for preprocessing data acquired from a camera and a microphone, means for analyzing the preprocessed data and estimating the emotions of the animal and the user, and means for generating speech based on the estimated emotions. This enables high-precision analysis of the emotions of the animal and the user, and allows for appropriate feedback and interaction suggestions based on the results.
[0781] A "camera" is an optical device used to capture the behavior, gestures, and facial expressions of animals and users.
[0782] A "microphone" is an audio device used to capture the voices of animals and users.
[0783] "Data preprocessing" refers to operations such as noise reduction and extraction of regions of interest on data acquired from cameras and microphones.
[0784] A "server" is an information processing device that analyzes pre-processed data, estimates the emotions of animals and users, and generates necessary feedback.
[0785] "Emotion estimation" is the process of inferring the emotional state of animals and users from acquired and pre-processed data.
[0786] "Voice generation" is the process of creating appropriate voice messages based on estimated emotions.
[0787] An "output device" is a device that emits the generated sound from an animal's collar or a speaker on the user's side.
[0788] "Interaction suggestion" is a process that proposes appropriate ways for animals and users to interact based on emotion estimation results.
[0789] This invention is a system that simultaneously analyzes the emotions of both animals and users, and provides appropriate feedback and interaction suggestions. Specific embodiments are described below.
[0790] First, multiple high-resolution cameras and high-sensitivity microphones installed in the room are used to capture the behavior, gestures, facial expressions, and voices of the animals and the user in real time. The cameras capture video at 30 frames per second, while the microphones collect audio data at a sampling rate of 20 kHz.
[0791] The terminal performs preprocessing on the acquired data. Specifically, it uses the open-source OpenCV library to denoise video data and extract regions of interest, and the librosa library to reduce background noise in audio data and convert it into a clear audio signal.
[0792] The pre-processed data is sent from the terminal to the server. The server uses the YOLOv3 model for image analysis and the Google Speech-to-Text API for speech analysis to analyze the received data. This analysis identifies the behavior, gestures, facial expressions, sounds, and voice tone of animals and users.
[0793] Next, the server estimates the emotions of both the animals and the users. For animal emotion estimation, a TensorFlow-based speech emotion recognition model is used, while for user emotion estimation, a combination of an image recognition model (e.g., OpenPose) and a speech emotion recognition model is used.
[0794] Based on the estimated emotions, the server generates an appropriate voice message. The Google Cloud Text-to-Speech API is used for voice generation. The generated voice message is sent to an output device attached to the animal's collar and to the user's output device (speaker) for playback.
[0795] As a concrete example, let's consider the case where a cat is enjoying being petted. In this case, the camera records the user's actions as they pet the cat, and the microphone captures the cat's meows. The device preprocesses this data and sends it to the server. The server analyzes the data, estimating the cat's "happy" state from its gestures and meows, while also analyzing the user's facial expressions and tone of voice to recognize that they are "enjoying" the experience. Based on this, the server generates the audio "That feels good!" and provides feedback to the user saying "The cat is happy." This audio is played from the animal's collar and the user's speaker.
[0796] As another example, consider a case where a cat appears anxious. In this case, the camera and microphone capture the cat's startled movements and meows, which the device preprocesses and sends to the server. The server estimates that the cat is "anxious" and similarly recognizes this from the user's facial expressions and tone of voice. The server generates a voice message saying "What's wrong?" and provides feedback to the user saying "Please gently stroke the cat to reassure it." This voice message is also played back from the output device.
[0797] Specific examples of prompt statements are as follows:
[0798] 1. "Explain how a system that uses a camera and microphone to analyze a cat's emotions provides feedback when the cat is happy while being petted."
[0799] 2. "Please describe the procedure for analyzing data from cats that appear surprised and anxious, and then suggesting appropriate interactions for the user."
[0800] As described above, this system analyzes the emotions of animals and users with high accuracy and improves the relationship between animals and users by providing appropriate feedback and interaction suggestions based on the results.
[0801] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0802] Step 1:
[0803] Data acquisition using camera and microphone
[0804] The device uses multiple high-resolution cameras and high-sensitivity microphones installed in the room to capture the behavior, gestures, facial expressions, and voices of animals and users in real time. Specifically, the cameras capture video at a rate of 30 frames per second, and the microphones collect audio data at a sampling rate of 20 kHz. This inputs data such as animal movements and sounds, and the user's facial expressions and voice tone. The output of this step is raw video and audio data.
[0805] Step 2:
[0806] Data preprocessing
[0807] The terminal performs preprocessing on the acquired raw video and audio data. Using the OpenCV library, it removes background noise from the video data and extracts the areas of interest for the animals and the user. It also uses the librosa library to reduce background noise from the audio data, converting it into a clear audio signal. Since the input video data contains noise, it undergoes filtering, and noise reduction is applied to the audio data. The output of this step is the preprocessed video and audio data.
[0808] Step 3:
[0809] Pre-processed data transmission to the server
[0810] The terminal sends pre-processed video and audio data to the server. The data is sent asynchronously using an HTTP POST request. As a specific example, compressed pre-processed data is sent to the server via a REST API. The input to this step is the pre-processed data, and the output is the data transfer to the server.
[0811] Step 4:
[0812] Server-based data analysis
[0813] The server analyzes the received preprocessed data. For image analysis, the YOLOv3 model is used to identify animal and user behavior, gestures, and facial expressions. For audio data analysis, the Google Speech-to-Text API is used to convert vocalizations and tone into text data for further analysis. The input for this step is preprocessed video and audio data, and the output is behavioral and emotion labels as analysis results. A concrete example of its operation is analyzing a cat's posture and vocalizations while being petted to identify that "the cat is happy."
[0814] Step 5:
[0815] Estimating animal emotions
[0816] The server estimates emotions from animal sounds and gestures. Using a TensorFlow-based speech emotion recognition model, it analyzes vocal patterns and phonemes to identify emotions such as "happy" or "anxious." The input for this step is the analyzed audio and video data, and the output is the animal's emotion label. A concrete example of its operation is analyzing a cat's high-pitched meow to estimate the emotion of "happiness."
[0817] Step 6:
[0818] User emotion recognition
[0819] The server analyzes the user's gestures, facial expressions, and voice tone to recognize their emotions. It uses a combination of an image recognition model (e.g., OpenPose) and a voice emotion recognition model. The input for this step is the analyzed video and audio data, and the output is the user's emotion label. A concrete example of its operation is detecting the user's smile and cheerful voice tone to identify a state of "enjoyment."
[0820] Step 7:
[0821] Speech generation and interaction proposals
[0822] The server generates an appropriate voice message based on the animal's emotions and the user's emotions. Using the Google Cloud Text-to-Speech API, the generated voice message is sent to an output device attached to the animal's collar and the user's output device (speaker) for playback. The input for this step is the emotion label, and the output is the generated voice message. A concrete example of its operation is the generation of a message such as "The cat is happy."
[0823] Step 8:
[0824] Audio playback and user feedback
[0825] The terminal transmits and plays the generated audio to an output device attached to the animal's collar and to the user's output device. The user listens to the audio from the speaker and takes appropriate action according to the animal's emotional state. The input for this step is the generated audio message, and the output is the playback of the audio. A concrete example of this operation would be the audio "Feels good!" being played from the cat's collar, and the message "The cat is happy" being played from the user's speaker.
[0826] (Application Example 2)
[0827] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."
[0828] In traditional pet shops, it was difficult for staff to accurately understand the emotions of customers and their pets, resulting in difficulties in recommending appropriate products and services. Understanding pets' emotions is particularly crucial for providing stress-free service, but effective methods for achieving this were lacking. Therefore, there is a growing need for a system that recognizes the emotions of customers and their pets in real time and provides optimal recommendations and feedback based on that understanding, in order to improve customer satisfaction.
[0829] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[0830] In this invention, the server includes means for acquiring animal behavior, gestures, facial expressions, and vocalizations from a camera and microphone; means for analyzing customer gestures, facial expressions, tone of voice, and content from the data to recognize customer emotions; and means for suggesting optimal products and services based on the emotions of the animals and customers. This makes it possible to recognize the emotions of customers and pets in real time within a pet shop and to provide optimal products and services based on that recognition.
[0831] A "camera" is a device that acquires visual information and converts it into digital images or video data.
[0832] A "microphone" is a device that captures sound and converts it into digital audio data.
[0833] "Data preprocessing" is the process of removing noise from acquired data and converting it into a format suitable for analysis.
[0834] "Means for estimating animal emotions" refers to algorithms or devices that identify and estimate an animal's emotional state based on its behavior, gestures, facial expressions, and vocalizations.
[0835] "Means for recognizing customer emotions" refers to an algorithm or device that identifies and recognizes a customer's emotional state based on their gestures, facial expressions, tone of voice, and content.
[0836] "Means of proposing optimal products and services" refers to algorithms or devices that select and propose appropriate products and services based on the emotional state of animals and customers.
[0837] "Means for generating speech" refers to algorithms or devices that create natural-sounding speech from text.
[0838] A "speaker" is a device that reproduces digital audio data as sound.
[0839] A "server" is a computing device or a system on a network that provides data reception, storage, and analysis functions.
[0840] The embodiments for carrying out this invention will be described in detail below. As an example of the embodiment of the invention, we will particularly illustrate its application to a system aimed at improving customer service in a pet shop.
[0841] System Configuration
[0842] This system recognizes the emotions of customers and their pets in real time and suggests the most suitable products and services based on that. The system includes a camera, microphone, terminal, server, and speaker.
[0843] Hardware configuration
[0844] 1. Camera
[0845] High-resolution cameras (e.g., Logitech C920) capture the behavior, gestures, and facial expressions of customers and pets inside the store.
[0846] 2. Mike
[0847] High-sensitivity microphones (e.g., Blue Yeti) can capture the voices of customers or pets.
[0848] 3. Terminal
[0849] Use a laptop or mobile device (e.g., iPhone) to preprocess data from the camera and microphone.
[0850] 4. Server
[0851] High-performance servers (e.g., AWS EC2) analyze pre-processed data to estimate the emotions of animals and customers.
[0852] 5. Speakers
[0853] Speakers installed in the store and speakers attached to pet collars play sounds generated by the system.
[0854] Software Configuration
[0855] 1. Preprocessing
[0856] The device performs noise reduction on data acquired from the camera and microphone, and detects faces and voices. Specifically, it uses OpenCV and the SpeechRecognition library.
[0857] 2. Data Analysis
[0858] The server uses the DeepFace library to perform facial recognition and estimate customer emotions in order to analyze pre-processed data sent from the terminal. It also uses Google's Speech-to-Text API for speech analysis.
[0859] 3. Emotion estimation
[0860] The server estimates the emotions of animals and customers based on the analyzed data. It identifies emotions such as "happy" or "anxious" from the animals' vocalizations and gestures, and identifies emotions such as "enjoying" or "anxious" from the customers' tone of voice and facial expressions.
[0861] 4. Suggestions and Feedback
[0862] Based on the emotions of animals and customers, the server suggests the most suitable products and services and generates voice messages. Voice generation uses gTTS (Google Text-to-Speech) and is played back using the Playsound library.
[0863] Specific example of processing
[0864] For example, if a customer visits a pet shop with their pet and is looking at products, the system will perform the following actions.
[0865] Example of a prompt
[0866] Camera image analysis prompt:
[0867] "Analyze the dominant emotion from this frame. Frame: {frame_data}"
[0868] Voice analysis prompt:
[0869] "Estimate the customer's emotions from their voice. Voice data: {audio_data}"
[0870] Based on camera footage and audio data, the system analyzes the emotions of the customer and their pet, understanding situations in real time, such as "the pet is interested in a toy" or "the customer is looking for new food." Based on this, the system generates voice messages such as "Your pet is interested in this toy. How about this toy?" and provides feedback to the customer through the speaker.
[0871] This improves the purchasing experience because customers can more easily understand their pet's condition and be offered appropriate products and services.
[0872] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0873] Step 1:
[0874] The device uses a high-resolution camera and a high-sensitivity microphone to acquire video and audio data of customers and pets in the store. The input consists of video data from the camera and audio data from the microphone, and this data is sent to the next step.
[0875] Step 2:
[0876] The video and audio data acquired by the device are preprocessed. Specifically, for video data, noise is removed using OpenCV, and faces and animal gestures are detected. For audio data, noise is removed using the SpeechRecognition library, and it is converted into clear audio data. The input is the video and audio data acquired in step 1, and the output is the preprocessed data.
[0877] Step 3:
[0878] The terminal sends pre-processed data to the server. The input is pre-processed data, which is sent from the terminal to the server via the network.
[0879] Step 4:
[0880] The server receives pre-processed data and analyzes the video data using the DeepFace library to estimate emotions from the customer's facial expressions. The input is video data sent from the terminal, and the output is the result of the customer's emotion recognition. The prompt message "Customer's face frame: {frame_data}" is used.
[0881] Step 5:
[0882] The server converts audio data into text using Google's Speech-to-Text API, and then analyzes the content and tone of voice to recognize the customer's emotions. The input is pre-processed audio data, and the output is the emotion recognition result based on speech recognition. The prompt message used is "Audio input data: {audio_data}".
[0883] Step 6:
[0884] The server analyzes animal video and audio data to determine their behavior, gestures, and vocalizations, and estimates their emotions. The input is video and audio data transmitted from a terminal, and the output is the estimated animal emotions. DeepFace and an audio analysis model are used.
[0885] Step 7:
[0886] The server generates messages suggesting optimal products and services based on animal emotion estimation results and customer emotion recognition results. gTTS is used to convert text to speech during this process. The input is the emotion recognition and emotion estimation results, and the output is the generated speech message.
[0887] Step 8:
[0888] The device sends the generated voice message to a speaker, which plays it back through the store's speakers and a speaker attached to the pet's collar. The input is the generated voice message, and the output is voice feedback to the customer and their pet.
[0889] Step 9:
[0890] The user listens to feedback from the speaker and selects appropriate products or services based on the system's suggestions. The input is voice feedback, and the output is the user's actions.
[0891] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[0892] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). An example of data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0893] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart glasses 214.
[0894] [Third Embodiment]
[0895] Figure 5 shows an example of the configuration of the data processing system 310 according to the third embodiment.
[0896] As shown in Figure 5, the data processing system 310 includes a data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.
[0897] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0898] The headset terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a display 343. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and display 343 are also connected to the bus 52.
[0899] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[0900] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).
[0901] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[0902] Figure 6 shows an example of the main functions of the data processing device 12 and the headset terminal 314. As shown in Figure 6, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[0903] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0904] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0905] In the headset terminal 314, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0906] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the headset terminal 314 will be referred to as the "terminal".
[0907] This invention is a system that uses a camera and microphone to acquire the behavior, gestures, facial expressions, and vocalizations of animals, analyzes this data to estimate the animals' emotions, and expresses the estimated emotions in sound. Specific embodiments for carrying out this invention are described below.
[0908] System Overview
[0909] This system uses multiple cameras and microphones placed in a room to capture animal behavior and vocalizations in real time, and sends the data to a server for analysis. Based on the analysis results, sounds are generated and emitted from speakers attached to the animals' collars, allowing users to understand the animals' emotions.
[0910] System Configuration
[0911] 1. Camera and microphone:
[0912] Multiple cameras and microphones installed in the room capture animal behavior and sounds in real time. The cameras continuously record animal movements, while the microphones record sounds such as vocalizations.
[0913] 2. Preprocessing at the terminal:
[0914] The device preprocesses the data acquired from the camera and microphone. Specifically, it removes background noise from the video and filters the audio data to reduce noise. This preprocessed data is then sent to the server for further analysis.
[0915] 3. Analysis on the server:
[0916] The server analyzes pre-processed data sent from the terminal. It identifies animal behavior, gestures, and facial expressions from image data, and analyzes animal sounds from audio data. This allows it to estimate the animal's emotions.
[0917] 4. Speech generation:
[0918] Based on the emotions of the animals estimated by the server, the appropriate voice text is selected and generated as speech. For example, if the emotion "happy" is estimated, the corresponding speech will be "Feels good!".
[0919] 5. Audio playback via speaker:
[0920] The generated audio is transmitted via a device to a speaker attached to the animal's collar. The audio emitted from the speaker allows the user to understand the animal's emotions in real time.
[0921] Specific examples of actions
[0922] Example 1: When a cat is happy to be petted
[0923] 1. Device: The camera records the user's actions as they pet the cat, and the microphone captures the cat's meows.
[0924] 2. Terminal: Preprocess this data to remove unwanted noise.
[0925] 3. Terminal: Sends pre-processed data to the server.
[0926] 4. Server: Analyzes data and estimates whether the cat is "happy" based on its gestures and meows.
[0927] 5. Server: Generates the voice "Feels good!" which corresponds to the "feeling happy" state.
[0928] 6. Terminal: Sends the generated audio to the speaker and plays it.
[0929] 7. User: Hearing the voice saying "Feels good!" from the cat's collar, the user understands that the cat is happy.
[0930] Example 2: When the cat looks anxious
[0931] 1. Device: The camera and microphone capture the cat's actions and meows when it is startled by something.
[0932] 2. Terminal: Preprocess this data.
[0933] 3. Terminal: Sends pre-processed data to the server.
[0934] 4. Server: Analyzes data and estimates whether the cat is "anxious" based on its behavior and meows.
[0935] 5. Server: Generates the voice message "What's wrong?" which corresponds to the state of "feeling anxious".
[0936] 6. Terminal: Sends the generated audio to the speaker and plays it.
[0937] 7. User: Hearing the voice saying "What's wrong?" from the cat's collar, the user understands that the cat is anxious and takes action to soothe it.
[0938] In this way, the system provides support for better communication between animals and their owners, and makes it possible to understand pets' emotions more accurately and in real time.
[0939] The following describes the processing flow.
[0940] Step 1:
[0941] The device initializes the camera and microphone installed in the room and captures the animal's behavior, gestures, facial expressions, and vocalizations in real time. Specifically, it continuously captures frame images from the camera and records audio data from the microphone.
[0942] Step 2:
[0943] The terminal preprocesses the acquired video and audio data. The video data is filtered to remove noise and extract the animal's areas of interest. The audio data is processed to remove background noise and convert it into a clear audio signal.
[0944] Step 3:
[0945] The terminal packages the pre-processed data and sends it to the server via the internet. The data package includes video frames and filtered audio data.
[0946] Step 4:
[0947] The server receives the data package sent from the terminal and prepares it for analysis. It performs a consistency check on the received data to confirm that there are no missing or corrupted files.
[0948] Step 5:
[0949] The server analyzes the video data and applies models to identify animal behavior, gestures, and facial expressions. For example, it can detect the posture a cat is in when being petted, or specific gestures.
[0950] Step 6:
[0951] The server analyzes the audio data and applies a speech recognition model to estimate emotions from animal sounds. It analyzes the patterns and phonemes of the sounds to identify emotions.
[0952] Step 7:
[0953] The server integrates the analysis results of video and audio data to estimate the animal's overall emotions. For example, if the animal shows signs of happiness, it will be estimated to be "happy."
[0954] Step 8:
[0955] The server generates an appropriate voice message based on the estimated emotion. For example, if the emotion is "happy," it will generate the voice message "Feels good!"
[0956] Step 9:
[0957] The server sends the generated audio data to the terminal. This includes the audio message and associated metadata.
[0958] Step 10:
[0959] The device receives audio data from the server and transmits it to a speaker attached to the animal's collar, playing the audio.
[0960] Step 11:
[0961] Users can understand animal emotions through voice messages emitted from the speaker. For example, by hearing the voice say "That feels good!", they can recognize that the cat is happy.
[0962] This allows users to understand the animals' emotions in real time and engage in appropriate dialogue and responses.
[0963] (Example 1)
[0964] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0965] Conventional pet emotion understanding systems have difficulty accurately estimating emotions from animal behavior and vocalizations, potentially leading owners to misunderstand their pets' feelings. Furthermore, the time required for real-time emotion estimation and feedback hinders smooth communication between pets and their owners. To address these issues, the present invention provides a system that accurately and in real-time estimates animal emotions and expresses those emotions verbally.
[0966] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0967] In this invention, the server includes means for acquiring animal behavior, gestures, facial expressions, and vocalizations from a camera and a microphone; means for preprocessing the data acquired from the camera and microphone; means for transmitting the preprocessed data to the server; means for analyzing the preprocessed data and estimating the animal's emotions; means for generating sound based on the estimated emotions; and means for emitting the generated sound from a speaker attached to the animal's collar. This makes it possible to estimate the animal's emotions with high accuracy and in real time, and to express those emotions in sound.
[0968] A "camera" is a device used to capture images of an animal's behavior, gestures, and facial expressions as video data.
[0969] A "microphone" is a sound-collecting device used to capture animal sounds and ambient noises as audio data.
[0970] "Data preprocessing" is the process of removing noise from video and audio data acquired from cameras and microphones, and preparing it in a format that is easy to analyze.
[0971] A "server" is a central processing unit that receives data transmitted from terminals, performs analysis, and estimates emotions.
[0972] "Analysis" refers to the process of identifying and analyzing animal behavior and vocalizations based on pre-processed data.
[0973] "Emotion estimation" is the process of identifying an animal's emotions based on the results of an analysis.
[0974] "Voice generation" is the process of selecting appropriate voice text to express estimated emotions and generating it as voice data.
[0975] A "speaker" is an output device that plays back generated audio data as sound.
[0976] A "terminal" is a device that preprocesses data acquired from cameras and microphones and sends it to a server.
[0977] This invention is a system that uses a camera and microphone to acquire animal behavior, gestures, facial expressions, and vocalizations in real time, analyzes this data to estimate the animal's emotions, and expresses those estimated emotions in sound. This system is designed to make it easier for users to understand animal emotions. Its specific form is described below.
[0978] Hardware and software to be used
[0979] Camera: Used to continuously record the behavior and facial expressions of animals. Examples include the Logitec HD Pro Webcam.
[0980] Microphone: Used to record animal sounds and ambient noise. Examples include the Blue Yeti USB Mic.
[0981] Terminal: A small PC used for data preprocessing and communication with the server. For example, a Raspberry Pi can be used.
[0982] Server: A central processing unit that performs data analysis and sentiment estimation. Examples include AWS and Google Cloud Platform.
[0983] Generative AI models: AI / ML frameworks used for data analysis and sentiment estimation. Specific examples include TensorFlow and PyTorch.
[0984] Text-to-speech software: Generates speech based on estimated emotions. Examples include Amazon Polly and Google Text-to-Speech.
[0985] Speaker: An output device for playing back the generated sound, attached to the animal's collar.
[0986] Specific examples of actions
[0987] Example 1: When a cat is happy to be petted
[0988] 1. The device uses a camera to record the user's actions as they pet the cat, and a microphone to capture the cat's meows.
[0989] 2. The terminal preprocesses this data to remove unwanted noise.
[0990] 3. The terminal sends the pre-processed data to the server.
[0991] 4. The server analyzes the data and estimates whether the cat is "happy" based on its gestures and meows.
[0992] 5. The server generates the voice "Feels good!" which corresponds to the state of being "happy".
[0993] 6. The device sends the generated audio to the speaker and plays it.
[0994] 7. The user hears the cat say "Feels good!" from the cat's collar and understands that the cat is happy.
[0995] Example 2: When the cat looks anxious
[0996] 1. The device uses its camera and microphone to capture any actions or sounds the cat makes when it is startled by something.
[0997] 2. The terminal preprocesses this data.
[0998] 3. The terminal sends the pre-processed data to the server.
[0999] 4. The server analyzes the data and estimates whether the cat is "anxious" based on its behavior and meows.
[1000] 5. The server generates the voice message "What's wrong?" which corresponds to the state of being "anxious".
[1001] 6. The device sends the generated audio to the speaker and plays it.
[1002] 7. The user hears a voice saying "What's wrong?" from the cat's collar, understands that the cat is anxious, and takes action to soothe it.
[1003] Example of a prompt
[1004] "Analyze animal behavioral and auditory data to estimate the animal's emotions and generate appropriate voice-to-text. Specifically, if the animal is happy, generate the voice 'Feels good!', and if it is anxious, generate the voice 'What's wrong?'."
[1005] In this way, the present invention provides support for better communication between animals and their owners, and makes it possible to understand pets' emotions more accurately and in real time.
[1006] The flow of the specific processing in Example 1 will be explained using Figure 11.
[1007] Step 1:
[1008] Data acquisition
[1009] The device uses a camera and microphone installed in the room to capture animal behavior and sounds in real time. The camera captures video of the animals, and the microphone records their audio. Specifically, the camera captures video data frame by frame, and the microphone continuously records audio.
[1010] Input: Animal behavior and sounds
[1011] Output: Raw video data and raw audio data
[1012] Step 2:
[1013] Data preprocessing
[1014] The device preprocesses the acquired video and audio data. Specifically, it removes background noise from the video data and filters the audio data to reduce unwanted noise. For video data, it masks parts other than animals, and for audio data, it applies filtering techniques to remove ambient noise.
[1015] Input: Raw video data and raw audio data
[1016] Output: Pre-processed video data and pre-processed audio data
[1017] Step 3:
[1018] Data transmission
[1019] The terminal sends pre-processed data to the server. This pre-processed video and audio data are transmitted to the server in real time using a high-speed network connection.
[1020] Input: Pre-processed video data and pre-processed audio data
[1021] Output: Data sent to the server
[1022] Step 4:
[1023] Data Analysis
[1024] The server analyzes the received pre-processed data. Specifically, it uses a generative AI model (e.g., TensorFlow or PyTorch) to identify animal behavior, gestures, and facial expressions from video data, and analyze animal sounds from audio data. Based on prompts, it obtains the analysis results.
[1025] Input: Pre-processed video data and pre-processed audio data
[1026] Output: Behavioral analysis results and vocalization analysis results
[1027] Step 5:
[1028] Emotion estimation
[1029] The server estimates the animal's emotions based on the results of data analysis. Specifically, it combines behavioral analysis results and vocalization analysis results to estimate the animal's current emotional state (e.g., "happy," "anxious").
[1030] Input: Behavioral analysis results and vocalization analysis results
[1031] Output: Estimated emotional state
[1032] Step 6:
[1033] Speech generation
[1034] The server generates speech based on estimated emotions. Specifically, it selects appropriate speech text corresponding to the emotional state and generates speech data using speech synthesis software (such as Amazon Polly or Google Text-to-Speech).
[1035] Input: Estimated emotional state
[1036] Output: Generated audio data
[1037] Step 7:
[1038] Audio transmission and playback
[1039] The device receives audio data from the server and transmits it to a speaker attached to the animal's collar, which then plays the audio. Specifically, the device transmits the audio data to the playback speaker via Bluetooth or Wi-Fi, and the speaker outputs the audio.
[1040] Input: Generated audio data
[1041] Output: Audio played through the speaker
[1042] Through these steps, users can understand the animals' emotions in real time, enabling better communication with them.
[1043] (Application Example 1)
[1044] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[1045] Conventional animal emotion estimation systems can estimate an animal's emotions, but they lack the ability to detect abnormal behavior and notify the user in real time. As a result, it is difficult for pet owners and caregivers to respond quickly to abnormal animal behavior. Furthermore, safety issues remain unresolved. This invention aims to improve animal safety and user peace of mind by detecting not only animal emotions but also abnormal behavior and notifying the user.
[1046] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[1047] In this invention, the server includes means for acquiring animal behavior, gestures, facial expressions, and vocalizations from a camera and a microphone; means for preprocessing the data acquired from the camera and microphone; means for analyzing the preprocessed data and estimating the animal's emotions; means for generating sound based on the estimated emotions; means for emitting the generated sound from a speaker attached to the animal's collar; means for analyzing the data to detect abnormal animal behavior; and means for generating a warning based on the detected abnormal behavior and notifying the user. This makes it possible to detect both animal emotions and abnormal behavior in real time and notify the user.
[1048] A "camera" is a device used to capture images of an animal's behavior, gestures, and facial expressions.
[1049] A "microphone" is a sound acquisition device used to record animal sounds.
[1050] "Data preprocessing" refers to the process of removing noise and converting the data format of video and audio data acquired from cameras and microphones.
[1051] "Data analysis" refers to algorithmic processing that identifies animal behavior, gestures, and vocalizations from pre-processed video and audio data, and estimates the animal's emotions and abnormal behavior.
[1052] "Emotion estimation" is the process of inferring what emotional state an animal is in based on its behavior, gestures, and vocalizations.
[1053] "Voice generation" is the process of generating voice messages that convey emotions to the user based on estimated animal emotions.
[1054] A "speaker attached to an animal's collar" is an audio playback device installed on the collar worn by an animal, used to play generated audio messages.
[1055] "Detection of abnormal behavior" is the process of determining whether an animal is behaving in an unusual way based on data analysis.
[1056] "Notification to users" refers to the process of sending warning messages or notifications to users in real time based on detected abnormal behavior.
[1057] "Real-time" is a time standard that means processing or notifications are performed almost instantaneously or with very little delay.
[1058] The present invention is a system that uses a camera and microphone to acquire animal behavior, gestures, facial expressions, and vocalizations, analyzes this data to estimate the animal's emotions and abnormal behavior, and notifies the user as necessary. Specific embodiments for carrying out the present invention are described below.
[1059] System Overview
[1060] This system uses multiple cameras and microphones placed in a room or outdoors to capture animal behavior and vocalizations in real time, and transmits the data to a server for analysis. Based on the analysis results, it emits sounds from a speaker attached to the animal's collar, allowing the user to understand the animal's emotions. Furthermore, it has a function to notify the user of any abnormal behavior if it is detected.
[1061] System Configuration
[1062] 1. Camera and microphone
[1063] Multiple cameras and microphones, installed indoors or outdoors, capture animal behavior and sounds in real time. The cameras continuously record animal movements, while the microphones record sounds such as vocalizations.
[1064] 2. Preprocessing at the terminal
[1065] The device preprocesses the data acquired from the camera and microphone. Specifically, it removes background noise from the video and filters the audio data to reduce noise. This preprocessed data is then sent to the server for further analysis.
[1066] 3. Analysis on the server
[1067] The server analyzes pre-processed data sent from the terminal. It identifies animal behavior, gestures, and facial expressions from image data, and analyzes vocalizations from audio data. This allows it to estimate the animal's emotions and abnormal behavior. The server performs the analysis using machine learning algorithms such as TensorFlow.
[1068] 4. Speech generation
[1069] Based on the emotions of the animal estimated by the server, the appropriate voice text is selected and generated as speech. For example, if the emotion "happy" is estimated, the corresponding speech will be "Feels good!". In addition, if abnormal behavior is detected, a warning voice such as "Be careful!" will be generated.
[1070] 5. Audio playback via speaker
[1071] The generated audio is transmitted via a device to a speaker attached to the animal's collar. The audio emitted from the speaker allows the user to understand the animal's emotions and any abnormal behavior in real time.
[1072] Specific examples of actions
[1073] Example 1: When a dog senses a suspicious person and becomes alert.
[1074] 1. Device: The camera and microphone capture the dog's alert behavior and barking.
[1075] 2. Terminal: Preprocess this data to remove unwanted noise.
[1076] 3. Terminal: Sends pre-processed data to the server.
[1077] 4. Server: Analyzes data, estimates a "vigilant" state from the dog's movements and barks, and further detects abnormal behavior.
[1078] 5. Server: Generates the voice message "Be careful!" corresponding to the "Alert" state.
[1079] 6. Terminal: Sends the generated audio to the speaker and plays it.
[1080] 7. User: Hearing the "Be careful!" voice message from the dog's collar, the user understands that the dog is on alert and takes appropriate action.
[1081] Example 2: When the dog is playing safely
[1082] 1. Device: The camera and microphone capture the dog's playful behavior and barking.
[1083] 2. Terminal: Preprocess this data.
[1084] 3. Terminal: Sends pre-processed data to the server.
[1085] 4. Server: Analyzes data and estimates whether the dog is "enjoying" itself based on its movements and barks.
[1086] 5. Server: Generates the voice message "This is fun!" which corresponds to the "enjoying" state.
[1087] 6. Terminal: Sends the generated audio to the speaker and plays it.
[1088] 7. User: Confirm that the dog is happy by hearing the voice message "Happy!" coming from the dog's collar.
[1089] Examples of use and prompts for generative AI models
[1090] Usage example
[1091] If a security dog detects an intruder, it records the situation with a camera, analyzes any unusual barking sounds with a microphone, and emits a warning message saying, "An intruder has been detected!"
[1092] Example of a prompt:
[1093] basic information:
[1094] Animals used: Dogs
[1095] Usage scenario: Security guarding
[1096] detail:
[1097] Capture a real-time video stream with the camera.
[1098] The sounds are recorded in real time using a microphone.
[1099] Analyze the data using a TensorFlow model to detect anomalies.
[1100] If an anomaly is detected, the user will be notified with an audio alert.
[1101] In this way, the system detects not only the animals' emotions but also abnormal behavior in real time and notifies the user, thereby improving the safety of the animals and the user's sense of security.
[1102] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[1103] Step 1:
[1104] Capture animal behavior and sounds from cameras and microphones.
[1105] Subject: terminal
[1106] Input: Animal movements and sounds
[1107] Output: Unprocessed video and audio data
[1108] Specific operation: The device acquires animal behavior and sounds in real time from multiple cameras and microphones installed indoors or outdoors. This includes capturing video and audio streams.
[1109] Step 2:
[1110] Preprocess the data acquired from the camera and microphone.
[1111] Subject: terminal
[1112] Input: Unprocessed video and audio data
[1113] Output: Pre-processed video and audio data
[1114] Specific operation: The terminal removes background noise from video data and filters audio data to reduce noise. This process uses OpenCV and audio filtering algorithms.
[1115] Step 3:
[1116] The pre-processed data is sent to the server.
[1117] Subject: terminal
[1118] Input: Pre-processed video and audio data
[1119] Output: Data sent to the server
[1120] Specific operation: The terminal sends pre-processed data to the server via the internet or local network. The data may be encrypted for security purposes during this process.
[1121] Step 4:
[1122] The server analyzes the data to estimate the animals' emotions and abnormal behaviors.
[1123] Subject: Server
[1124] Input: Pre-processed video and audio data
[1125] Output: Emotion and abnormal behavior estimation results
[1126] Specific operation: The server uses machine learning algorithms such as TensorFlow to analyze video and audio data and estimate emotions and abnormal behavior from animal behavior, gestures, and vocalizations. Specifically, it uses a trained model to classify emotional states such as "happy," "alert," and "anxious" and detect abnormal behavior.
[1127] Step 5:
[1128] It generates voices based on estimated emotions and abnormal behaviors.
[1129] Subject: Server
[1130] Input: Emotion and abnormal behavior estimation results
[1131] Output: Generated audio data
[1132] Specific operation: The server selects speech text corresponding to the estimated emotions and abnormal behavior, and converts the text into speech. This is done using speech synthesis software (such as the Google Text-to-Speech API).
[1133] Step 6:
[1134] The generated sound is emitted from a speaker attached to the animal's collar.
[1135] Subject: terminal
[1136] Input: Generated audio data
[1137] Output: Sound emitted from an animal collar
[1138] Specific operation: Audio data is generated and sent from the server to the terminal. The terminal then sends the audio data to a speaker attached to the animal's collar. The speaker plays the received audio data in real time.
[1139] Step 7:
[1140] The system notifies the user of abnormal behavior.
[1141] Subject: Server
[1142] Input: Abnormal behavior detection results
[1143] Output: Warning notification to the user
[1144] Specific operation: If abnormal behavior is detected, the server will send a warning message to the user's smartphone or other devices. This will be done using methods such as push notifications, email, or SMS. The user will receive the notification and be able to take prompt action.
[1145] This processing flow allows for the detection of not only animal emotions but also abnormal behavior in real time, and notifies the user, thereby improving animal safety and user peace of mind.
[1146] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[1147] This invention is a system that combines a system for estimating animal emotions with an emotion engine that recognizes user emotions. Specific embodiments of this invention are described below.
[1148] System Overview
[1149] This system uses multiple cameras and microphones installed in a room to capture animal behavior, gestures, facial expressions, and vocalizations, and analyzes this data to estimate the animal's emotions. It also uses cameras and microphones to capture the user's gestures, facial expressions, and voice tone, recognizing the user's emotions. This recognized emotion generates appropriate responses to the animals, which are then provided to the user as voice or other feedback, thereby facilitating interaction.
[1150] System Configuration
[1151] 1. Camera and microphone:
[1152] Multiple cameras and microphones installed in the room capture the behavior, sounds, gestures, and facial expressions of both animals and users in real time. The cameras continuously capture high-resolution video, while the microphones record high-sensitivity audio data.
[1153] 2. Preprocessing at the terminal:
[1154] The device preprocesses data acquired from the camera and microphone. Video data is filtered to remove noise and extract areas of interest for both the animal and the user. Audio data is processed to remove background noise and convert it into a clear audio signal. Furthermore, preprocessing is performed to analyze the user's voice tone and language patterns.
[1155] 3. Analysis on the server:
[1156] The server analyzes the pre-processed data sent from the terminal. First, it applies a model to the video data to identify the actions, gestures, and facial expressions of the animals and the user. This allows it to detect things like whether the cat is being petted or whether the user is smiling.
[1157] 4. Estimating animal emotions:
[1158] The server applies a speech recognition model to estimate emotions from animal sounds and gestures. It analyzes the patterns and phonemes of the sounds to identify emotions such as "happy" or "anxious."
[1159] 5. User emotion recognition:
[1160] The server analyzes the user's gestures, facial expressions, and tone of voice to recognize their emotions. For example, if the user is smiling, it recognizes that they are "having fun," and if they are frowning, it recognizes that they are "anxious."
[1161] 6. Speech generation and interaction proposals:
[1162] The server generates appropriate voice messages based on the animal's emotions and the user's emotions. Furthermore, it suggests appropriate interactions between the user and the animal. For example, if the animal appears anxious and the user is also anxious, the system will make a voice suggestion such as, "Please gently stroke the cat to reassure it."
[1163] 7. Audio playback via speaker:
[1164] The generated audio is transmitted via the device to a speaker attached to the animal's collar. Feedback audio is also transmitted from the device to the user's speaker.
[1165] Specific examples of actions
[1166] Example 1: When a cat is happy to be petted
[1167] 1. Device: The camera records the user's actions as they pet the cat, and the microphone captures the cat's meows.
[1168] 2. Terminal: Preprocess this data to remove unwanted noise.
[1169] 3. Terminal: Sends pre-processed data to the server.
[1170] 4. Server: Analyzes data and estimates whether the cat is "happy" based on its gestures and meows.
[1171] 5. Server: On the other hand, it analyzes the user's facial expressions and tone of voice to recognize when the user is "enjoying" the experience.
[1172] 6. Server: Generates the voice "Feels good!" corresponding to the "happy" state, and simultaneously generates a voice message to the user saying "The cat is happy."
[1173] 7. Terminal: The generated audio is sent to the cat's collar and the user's speaker for playback.
[1174] 8. User: Hearing the voice saying "Feels good!" from the cat's collar, the user understands that the cat is happy and confirms that they themselves are also enjoying it.
[1175] Example 2: When the cat looks anxious
[1176] 1. Device: The camera and microphone capture the cat's startled movements and meows.
[1177] 2. Terminal: Preprocess this data.
[1178] 3. Terminal: Sends pre-processed data to the server.
[1179] 4. Server: Analyzes data and estimates whether the cat is "anxious" based on its behavior and meows.
[1180] 5. Server: On the other hand, it analyzes the user's facial expressions and tone of voice to recognize that the user is also "anxious."
[1181] 6. Server: Generates the voice message "What's wrong?" to indicate an anxious state, and simultaneously generates a voice message to the user saying, "Please gently stroke the cat to reassure it."
[1182] 7. Terminal: The generated audio is sent to the cat's collar and the user's speaker for playback.
[1183] 8. User: Hearing the sound from the speaker, the user understands that the cat is anxious and takes action to reassure the cat by gently stroking it.
[1184] In this way, the system can provide support for users and animals to understand each other's emotions and communicate smoothly.
[1185] The following describes the processing flow.
[1186] Step 1:
[1187] The device initializes the camera and microphone installed in the room and captures animal behavior, gestures, facial expressions, and vocalizations, as well as the user's gestures, facial expressions, and voice tone in real time. The camera continuously captures frame images, and the microphone records audio data.
[1188] Step 2:
[1189] The device preprocesses the acquired video and audio data. First, it removes noise from the video data and extracts the areas of interest for the animals and the user. Next, it filters the audio data to remove background noise and converts it into a clear audio signal.
[1190] Step 3:
[1191] The terminal packages pre-processed animal and user data and transmits it to a server via the internet. The data package includes video frames and filtered audio data.
[1192] Step 4:
[1193] The server receives the data package sent from the terminal and prepares it for analysis. It performs a consistency check on the received data to confirm that there are no missing or corrupted files.
[1194] Step 5:
[1195] The server analyzes the video data and applies models to identify animal behavior, gestures, and facial expressions, as well as user gestures and facial expressions. For example, it can detect whether an animal is being petted or whether the user is smiling.
[1196] Step 6:
[1197] The server analyzes audio data and applies a speech recognition model to estimate emotions from animal sounds. It also analyzes the tone and phonemes of the user's voice to identify their emotions.
[1198] Step 7:
[1199] The server integrates the analysis results of video and audio data to estimate the overall emotions of the animal and the user. For example, if the animal is showing signs of happiness, it is estimated to be "happy," and if the user is enjoying themselves, it is recognized as "enjoying themselves."
[1200] Step 8:
[1201] The server generates an appropriate voice message based on the estimated emotion. For example, if the server estimates that the animal is "happy," it will generate a voice message saying "Feels good!" and provide the user with feedback such as "The cat is happy."
[1202] Step 9:
[1203] The server sends the generated audio data to the terminal. This includes the audio message and associated metadata.
[1204] Step 10:
[1205] The device receives audio data from the server and transmits it to a speaker attached to the animal's collar, while simultaneously transmitting it to the user's speaker for playback.
[1206] Step 11:
[1207] Through voice messages emitted from the speaker, users can understand the emotions of animals and recognize how their own actions affect them. For example, they might hear the voice saying "Feels good!" from the animal's collar, confirm that the cat is happy, and become aware of their own enjoyment through the feedback "The cat is happy."
[1208] This allows users to understand the animals' emotions in real time and improve their relationship with them by interacting with them appropriately.
[1209] (Example 2)
[1210] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[1211] Conventional animal emotion estimation systems only estimate emotions based on animal behavior and vocalizations, and merely provide feedback to the animals. However, since the emotions of users living with animals directly influence the animals' behavior and state, a comprehensive analysis that includes the user's emotions is desirable. Furthermore, providing appropriate interaction suggestions based on the emotions of both the animal and the user can improve the relationship between the animal and the user. Against this backdrop, there is a need for a system that simultaneously analyzes the emotions of both animals and users and provides appropriate feedback and interaction suggestions.
[1212] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[1213] In this invention, the server includes means for preprocessing data acquired from a camera and a microphone, means for analyzing the preprocessed data and estimating the emotions of the animal and the user, and means for generating speech based on the estimated emotions. This enables high-precision analysis of the emotions of the animal and the user, and allows for appropriate feedback and interaction suggestions based on the results.
[1214] A "camera" is an optical device used to capture the behavior, gestures, and facial expressions of animals and users.
[1215] A "microphone" is an audio device used to capture the voices of animals and users.
[1216] "Data preprocessing" refers to operations such as noise reduction and extraction of regions of interest on data acquired from cameras and microphones.
[1217] A "server" is an information processing device that analyzes pre-processed data, estimates the emotions of animals and users, and generates necessary feedback.
[1218] "Emotion estimation" is the process of inferring the emotional state of animals and users from acquired and pre-processed data.
[1219] "Voice generation" is the process of creating appropriate voice messages based on estimated emotions.
[1220] An "output device" is a device that emits the generated sound from an animal's collar or a speaker on the user's side.
[1221] "Interaction suggestion" is a process that proposes appropriate ways for animals and users to interact based on emotion estimation results.
[1222] This invention is a system that simultaneously analyzes the emotions of both animals and users, and provides appropriate feedback and interaction suggestions. Specific embodiments are described below.
[1223] First, multiple high-resolution cameras and high-sensitivity microphones installed in the room are used to capture the behavior, gestures, facial expressions, and voices of the animals and the user in real time. The cameras capture video at 30 frames per second, while the microphones collect audio data at a sampling rate of 20 kHz.
[1224] The terminal performs preprocessing on the acquired data. Specifically, it uses the open-source OpenCV library to denoise video data and extract regions of interest, and the librosa library to reduce background noise in audio data and convert it into a clear audio signal.
[1225] The pre-processed data is sent from the terminal to the server. The server uses the YOLOv3 model for image analysis and the Google Speech-to-Text API for speech analysis to analyze the received data. This analysis identifies the behavior, gestures, facial expressions, sounds, and voice tone of animals and users.
[1226] Next, the server estimates the emotions of both the animals and the users. For animal emotion estimation, a TensorFlow-based speech emotion recognition model is used, while for user emotion estimation, a combination of an image recognition model (e.g., OpenPose) and a speech emotion recognition model is used.
[1227] Based on the estimated emotions, the server generates an appropriate voice message. The Google Cloud Text-to-Speech API is used for voice generation. The generated voice message is sent to an output device attached to the animal's collar and to the user's output device (speaker) for playback.
[1228] As a concrete example, let's consider the case where a cat is enjoying being petted. In this case, the camera records the user's actions as they pet the cat, and the microphone captures the cat's meows. The device preprocesses this data and sends it to the server. The server analyzes the data, estimating the cat's "happy" state from its gestures and meows, while also analyzing the user's facial expressions and tone of voice to recognize that they are "enjoying" the experience. Based on this, the server generates the audio "That feels good!" and provides feedback to the user saying "The cat is happy." This audio is played from the animal's collar and the user's speaker.
[1229] As another example, consider a case where a cat appears anxious. In this case, the camera and microphone capture the cat's startled movements and meows, which the device preprocesses and sends to the server. The server estimates that the cat is "anxious" and similarly recognizes this from the user's facial expressions and tone of voice. The server generates a voice message saying "What's wrong?" and provides feedback to the user saying "Please gently stroke the cat to reassure it." This voice message is also played back from the output device.
[1230] Specific examples of prompt statements are as follows:
[1231] 1. "Explain how a system that uses a camera and microphone to analyze a cat's emotions provides feedback when the cat is happy while being petted."
[1232] 2. "Please describe the procedure for analyzing data from cats that appear surprised and anxious, and then suggesting appropriate interactions for the user."
[1233] As described above, this system analyzes the emotions of animals and users with high accuracy and improves the relationship between animals and users by providing appropriate feedback and interaction suggestions based on the results.
[1234] The flow of the specific processing in Example 2 will be explained using Figure 13.
[1235] Step 1:
[1236] Data acquisition using camera and microphone
[1237] The device uses multiple high-resolution cameras and high-sensitivity microphones installed in the room to capture the behavior, gestures, facial expressions, and voices of animals and users in real time. Specifically, the cameras capture video at a rate of 30 frames per second, and the microphones collect audio data at a sampling rate of 20 kHz. This inputs data such as animal movements and sounds, and the user's facial expressions and voice tone. The output of this step is raw video and audio data.
[1238] Step 2:
[1239] Data preprocessing
[1240] The terminal performs preprocessing on the acquired raw video and audio data. Using the OpenCV library, it removes background noise from the video data and extracts the areas of interest for the animals and the user. It also uses the librosa library to reduce background noise from the audio data, converting it into a clear audio signal. Since the input video data contains noise, it undergoes filtering, and noise reduction is applied to the audio data. The output of this step is the preprocessed video and audio data.
[1241] Step 3:
[1242] Pre-processed data transmission to the server
[1243] The terminal sends pre-processed video and audio data to the server. The data is sent asynchronously using an HTTP POST request. As a specific example, compressed pre-processed data is sent to the server via a REST API. The input to this step is the pre-processed data, and the output is the data transfer to the server.
[1244] Step 4:
[1245] Server-based data analysis
[1246] The server analyzes the received preprocessed data. For image analysis, the YOLOv3 model is used to identify animal and user behavior, gestures, and facial expressions. For audio data analysis, the Google Speech-to-Text API is used to convert vocalizations and tone into text data for further analysis. The input for this step is preprocessed video and audio data, and the output is behavioral and emotion labels as analysis results. A concrete example of its operation is analyzing a cat's posture and vocalizations while being petted to identify that "the cat is happy."
[1247] Step 5:
[1248] Estimating animal emotions
[1249] The server estimates emotions from animal sounds and gestures. Using a TensorFlow-based speech emotion recognition model, it analyzes vocal patterns and phonemes to identify emotions such as "happy" or "anxious." The input for this step is the analyzed audio and video data, and the output is the animal's emotion label. A concrete example of its operation is analyzing a cat's high-pitched meow to estimate the emotion of "happiness."
[1250] Step 6:
[1251] User emotion recognition
[1252] The server analyzes the user's gestures, facial expressions, and voice tone to recognize their emotions. It uses a combination of an image recognition model (e.g., OpenPose) and a voice emotion recognition model. The input for this step is the analyzed video and audio data, and the output is the user's emotion label. A concrete example of its operation is detecting the user's smile and cheerful voice tone to identify a state of "enjoyment."
[1253] Step 7:
[1254] Speech generation and interaction proposals
[1255] The server generates an appropriate voice message based on the animal's emotions and the user's emotions. Using the Google Cloud Text-to-Speech API, the generated voice message is sent to an output device attached to the animal's collar and the user's output device (speaker) for playback. The input for this step is the emotion label, and the output is the generated voice message. A concrete example of its operation is the generation of a message such as "The cat is happy."
[1256] Step 8:
[1257] Audio playback and user feedback
[1258] The terminal transmits and plays the generated audio to an output device attached to the animal's collar and to the user's output device. The user listens to the audio from the speaker and takes appropriate action according to the animal's emotional state. The input for this step is the generated audio message, and the output is the playback of the audio. A concrete example of this operation would be the audio "Feels good!" being played from the cat's collar, and the message "The cat is happy" being played from the user's speaker.
[1259] (Application Example 2)
[1260] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[1261] In traditional pet shops, it was difficult for staff to accurately understand the emotions of customers and their pets, resulting in difficulties in recommending appropriate products and services. Understanding pets' emotions is particularly crucial for providing stress-free service, but effective methods for achieving this were lacking. Therefore, there is a growing need for a system that recognizes the emotions of customers and their pets in real time and provides optimal recommendations and feedback based on that understanding, in order to improve customer satisfaction.
[1262] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[1263] In this invention, the server includes means for acquiring animal behavior, gestures, facial expressions, and vocalizations from a camera and microphone; means for analyzing customer gestures, facial expressions, tone of voice, and content from the data to recognize customer emotions; and means for suggesting optimal products and services based on the emotions of the animals and customers. This makes it possible to recognize the emotions of customers and pets in real time within a pet shop and to provide optimal products and services based on that recognition.
[1264] A "camera" is a device that acquires visual information and converts it into digital images or video data.
[1265] A "microphone" is a device that captures sound and converts it into digital audio data.
[1266] "Data preprocessing" is the process of removing noise from acquired data and converting it into a format suitable for analysis.
[1267] "Means for estimating animal emotions" refers to algorithms or devices that identify and estimate an animal's emotional state based on its behavior, gestures, facial expressions, and vocalizations.
[1268] "Means for recognizing customer emotions" refers to an algorithm or device that identifies and recognizes a customer's emotional state based on their gestures, facial expressions, tone of voice, and content.
[1269] "Means of proposing optimal products and services" refers to algorithms or devices that select and propose appropriate products and services based on the emotional state of animals and customers.
[1270] "Means for generating speech" refers to algorithms or devices that create natural-sounding speech from text.
[1271] A "speaker" is a device that reproduces digital audio data as sound.
[1272] A "server" is a computing device or a system on a network that provides data reception, storage, and analysis functions.
[1273] The embodiments for carrying out this invention will be described in detail below. As an example of the embodiment of the invention, we will particularly illustrate its application to a system aimed at improving customer service in a pet shop.
[1274] System Configuration
[1275] This system recognizes the emotions of customers and their pets in real time and suggests the most suitable products and services based on that. The system includes a camera, microphone, terminal, server, and speaker.
[1276] Hardware configuration
[1277] 1. Camera
[1278] High-resolution cameras (e.g., Logitech C920) capture the behavior, gestures, and facial expressions of customers and pets inside the store.
[1279] 2. Mike
[1280] High-sensitivity microphones (e.g., Blue Yeti) can capture the voices of customers or pets.
[1281] 3. Terminal
[1282] Use a laptop or mobile device (e.g., iPhone) to preprocess data from the camera and microphone.
[1283] 4. Server
[1284] High-performance servers (e.g., AWS EC2) analyze pre-processed data to estimate the emotions of animals and customers.
[1285] 5. Speakers
[1286] Speakers installed in the store and speakers attached to pet collars play sounds generated by the system.
[1287] Software Configuration
[1288] 1. Preprocessing
[1289] The device performs noise reduction on data acquired from the camera and microphone, and detects faces and voices. Specifically, it uses OpenCV and the SpeechRecognition library.
[1290] 2. Data Analysis
[1291] The server uses the DeepFace library to perform facial recognition and estimate customer emotions in order to analyze pre-processed data sent from the terminal. It also uses Google's Speech-to-Text API for speech analysis.
[1292] 3. Emotion estimation
[1293] The server estimates the emotions of animals and customers based on the analyzed data. It identifies emotions such as "happy" or "anxious" from the animals' vocalizations and gestures, and identifies emotions such as "enjoying" or "anxious" from the customers' tone of voice and facial expressions.
[1294] 4. Suggestions and Feedback
[1295] Based on the emotions of animals and customers, the server suggests the most suitable products and services and generates voice messages. Voice generation uses gTTS (Google Text-to-Speech) and is played back using the Playsound library.
[1296] Specific example of processing
[1297] For example, if a customer visits a pet shop with their pet and is looking at products, the system will perform the following actions.
[1298] Example of a prompt
[1299] Camera image analysis prompt:
[1300] "Analyze the dominant emotion from this frame. Frame: {frame_data}"
[1301] Voice analysis prompt:
[1302] "Estimate the customer's emotions from their voice. Voice data: {audio_data}"
[1303] Based on camera footage and audio data, the system analyzes the emotions of the customer and their pet, understanding situations in real time, such as "the pet is interested in a toy" or "the customer is looking for new food." Based on this, the system generates voice messages such as "Your pet is interested in this toy. How about this toy?" and provides feedback to the customer through the speaker.
[1304] This improves the purchasing experience because customers can more easily understand their pet's condition and be offered appropriate products and services.
[1305] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[1306] Step 1:
[1307] The device uses a high-resolution camera and a high-sensitivity microphone to acquire video and audio data of customers and pets in the store. The input consists of video data from the camera and audio data from the microphone, and this data is sent to the next step.
[1308] Step 2:
[1309] The video and audio data acquired by the device are preprocessed. Specifically, for video data, noise is removed using OpenCV, and faces and animal gestures are detected. For audio data, noise is removed using the SpeechRecognition library, and it is converted into clear audio data. The input is the video and audio data acquired in step 1, and the output is the preprocessed data.
[1310] Step 3:
[1311] The terminal sends pre-processed data to the server. The input is pre-processed data, which is sent from the terminal to the server via the network.
[1312] Step 4:
[1313] The server receives pre-processed data and analyzes the video data using the DeepFace library to estimate emotions from the customer's facial expressions. The input is video data sent from the terminal, and the output is the result of the customer's emotion recognition. The prompt message "Customer's face frame: {frame_data}" is used.
[1314] Step 5:
[1315] The server converts audio data into text using Google's Speech-to-Text API, and then analyzes the content and tone of voice to recognize the customer's emotions. The input is pre-processed audio data, and the output is the emotion recognition result based on speech recognition. The prompt message used is "Audio input data: {audio_data}".
[1316] Step 6:
[1317] The server analyzes animal video and audio data to determine their behavior, gestures, and vocalizations, and estimates their emotions. The input is video and audio data transmitted from a terminal, and the output is the estimated animal emotions. DeepFace and an audio analysis model are used.
[1318] Step 7:
[1319] The server generates messages suggesting optimal products and services based on animal emotion estimation results and customer emotion recognition results. gTTS is used to convert text to speech during this process. The input is the emotion recognition and emotion estimation results, and the output is the generated speech message.
[1320] Step 8:
[1321] The device sends the generated voice message to a speaker, which plays it back through the store's speakers and a speaker attached to the pet's collar. The input is the generated voice message, and the output is voice feedback to the customer and their pet.
[1322] Step 9:
[1323] The user listens to feedback from the speaker and selects appropriate products or services based on the system's suggestions. The input is voice feedback, and the output is the user's actions.
[1324] The specific processing unit 290 transmits the result of the specific processing to the headset terminal 314. In the headset terminal 314, the control unit 46A causes the speaker 240 and display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[1325] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). An example of data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[1326] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and specific processing may also be performed by the headset terminal 314.
[1327] [Fourth Embodiment]
[1328] Figure 7 shows an example of the configuration of the data processing system 410 according to the fourth embodiment.
[1329] As shown in Figure 7, the data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.
[1330] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[1331] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a controlled object 443. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and controlled object 443 are also connected to the bus 52.
[1332] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[1333] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).
[1334] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[1335] The controlled object 443 includes a display device, LEDs in the eyes, and motors that drive the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the robot 414's emotions can be expressed by controlling these motors. Furthermore, the robot 414's facial expressions can also be expressed by controlling the illumination state of the LEDs in its eyes.
[1336] Figure 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Figure 8, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[1337] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[1338] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[1339] In robot 414, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[1340] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[1341] This invention is a system that uses a camera and microphone to acquire the behavior, gestures, facial expressions, and vocalizations of animals, analyzes this data to estimate the animals' emotions, and expresses the estimated emotions in sound. Specific embodiments for carrying out this invention are described below.
[1342] System Overview
[1343] This system uses multiple cameras and microphones placed in a room to capture animal behavior and vocalizations in real time, and sends the data to a server for analysis. Based on the analysis results, sounds are generated and emitted from speakers attached to the animals' collars, allowing users to understand the animals' emotions.
[1344] System Configuration
[1345] 1. Camera and microphone:
[1346] Multiple cameras and microphones installed in the room capture animal behavior and sounds in real time. The cameras continuously record animal movements, while the microphones record sounds such as vocalizations.
[1347] 2. Preprocessing at the terminal:
[1348] The device preprocesses the data acquired from the camera and microphone. Specifically, it removes background noise from the video and filters the audio data to reduce noise. This preprocessed data is then sent to the server for further analysis.
[1349] 3. Analysis on the server:
[1350] The server analyzes pre-processed data sent from the terminal. It identifies animal behavior, gestures, and facial expressions from image data, and analyzes animal sounds from audio data. This allows it to estimate the animal's emotions.
[1351] 4. Speech generation:
[1352] Based on the emotions of the animals estimated by the server, the appropriate voice text is selected and generated as speech. For example, if the emotion "happy" is estimated, the corresponding speech will be "Feels good!".
[1353] 5. Audio playback via speaker:
[1354] The generated audio is transmitted via a device to a speaker attached to the animal's collar. The audio emitted from the speaker allows the user to understand the animal's emotions in real time.
[1355] Specific examples of actions
[1356] Example 1: When a cat is happy to be petted
[1357] 1. Device: The camera records the user's actions as they pet the cat, and the microphone captures the cat's meows.
[1358] 2. Terminal: Preprocess this data to remove unwanted noise.
[1359] 3. Terminal: Sends pre-processed data to the server.
[1360] 4. Server: Analyzes data and estimates whether the cat is "happy" based on its gestures and meows.
[1361] 5. Server: Generates the voice "Feels good!" which corresponds to the "feeling happy" state.
[1362] 6. Terminal: Sends the generated audio to the speaker and plays it.
[1363] 7. User: Hearing the voice saying "Feels good!" from the cat's collar, the user understands that the cat is happy.
[1364] Example 2: When the cat looks anxious
[1365] 1. Device: The camera and microphone capture the cat's actions and meows when it is startled by something.
[1366] 2. Terminal: Preprocess this data.
[1367] 3. Terminal: Sends pre-processed data to the server.
[1368] 4. Server: Analyzes data and estimates whether the cat is "anxious" based on its behavior and meows.
[1369] 5. Server: Generates the voice message "What's wrong?" which corresponds to the state of "feeling anxious".
[1370] 6. Terminal: Sends the generated audio to the speaker and plays it.
[1371] 7. User: Hearing the voice saying "What's wrong?" from the cat's collar, the user understands that the cat is anxious and takes action to soothe it.
[1372] In this way, the system provides support for better communication between animals and their owners, and makes it possible to understand pets' emotions more accurately and in real time.
[1373] The following describes the processing flow.
[1374] Step 1:
[1375] The device initializes the camera and microphone installed in the room and captures the animal's behavior, gestures, facial expressions, and vocalizations in real time. Specifically, it continuously captures frame images from the camera and records audio data from the microphone.
[1376] Step 2:
[1377] The terminal preprocesses the acquired video and audio data. A noise reduction filter is applied to the video data to extract the animal's areas of interest. Background noise is removed from the audio data, converting it into a clear audio signal.
[1378] Step 3:
[1379] The terminal packages the pre-processed data and sends it to the server via the internet. The data package includes video frames and filtered audio data.
[1380] Step 4:
[1381] The server receives the data package sent from the terminal and prepares it for analysis. It performs a consistency check on the received data to confirm that there are no missing or corrupted files.
[1382] Step 5:
[1383] The server analyzes the video data and applies models to identify animal behavior, gestures, and facial expressions. For example, it can detect the posture a cat is in when being petted, or specific gestures.
[1384] Step 6:
[1385] The server analyzes the audio data and applies a speech recognition model to estimate emotions from animal sounds. It analyzes the patterns and phonemes of the sounds to identify emotions.
[1386] Step 7:
[1387] The server integrates the analysis results of video and audio data to estimate the animal's overall emotions. For example, if the animal shows signs of happiness, it will be estimated to be "happy."
[1388] Step 8:
[1389] The server generates an appropriate voice message based on the estimated emotion. For example, if the emotion is "happy," it will generate the voice message "Feels good!"
[1390] Step 9:
[1391] The server sends the generated audio data to the terminal. This includes the audio message and associated metadata.
[1392] Step 10:
[1393] The device receives audio data from the server and transmits it to a speaker attached to the animal's collar, playing the audio.
[1394] Step 11:
[1395] Users can understand animal emotions through voice messages emitted from the speaker. For example, by hearing the voice say "That feels good!", they can recognize that the cat is happy.
[1396] This allows users to understand the animals' emotions in real time and engage in appropriate dialogue and responses.
[1397] (Example 1)
[1398] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[1399] Conventional pet emotion understanding systems have difficulty accurately estimating emotions from animal behavior and vocalizations, potentially leading owners to misunderstand their pets' feelings. Furthermore, the time required for real-time emotion estimation and feedback hinders smooth communication between pets and their owners. To address these issues, the present invention provides a system that accurately and in real-time estimates animal emotions and expresses those emotions verbally.
[1400] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[1401] In this invention, the server includes means for acquiring animal behavior, gestures, facial expressions, and vocalizations from a camera and a microphone; means for preprocessing the data acquired from the camera and microphone; means for transmitting the preprocessed data to the server; means for analyzing the preprocessed data and estimating the animal's emotions; means for generating sound based on the estimated emotions; and means for emitting the generated sound from a speaker attached to the animal's collar. This makes it possible to estimate the animal's emotions with high accuracy and in real time, and to express those emotions in sound.
[1402] A "camera" is a device used to capture images of an animal's behavior, gestures, and facial expressions as video data.
[1403] A "microphone" is a sound-collecting device used to capture animal sounds and ambient noises as audio data.
[1404] "Data preprocessing" is the process of removing noise from video and audio data acquired from cameras and microphones, and preparing it in a format that is easy to analyze.
[1405] A "server" is a central processing unit that receives data transmitted from terminals, performs analysis, and estimates emotions.
[1406] "Analysis" refers to the process of identifying and analyzing animal behavior and vocalizations based on pre-processed data.
[1407] "Emotion estimation" is the process of identifying an animal's emotions based on the results of an analysis.
[1408] "Voice generation" is the process of selecting appropriate voice text to express estimated emotions and generating it as voice data.
[1409] A "speaker" is an output device that plays back generated audio data as sound.
[1410] A "terminal" is a device that preprocesses data acquired from cameras and microphones and transmits it to a server.
[1411] This invention is a system that uses a camera and microphone to acquire animal behavior, gestures, facial expressions, and vocalizations in real time, analyzes this data to estimate the animal's emotions, and expresses those estimated emotions in sound. This system is designed to make it easier for users to understand animal emotions. Its specific form is described below.
[1412] Hardware and software to be used
[1413] Camera: Used to continuously record the behavior and facial expressions of animals. Examples include the Logitec HD Pro Webcam.
[1414] Microphone: Used to record animal sounds and ambient noise. Examples include the Blue Yeti USB Mic.
[1415] Terminal: A small PC used for data preprocessing and communication with the server. For example, a Raspberry Pi can be used.
[1416] Server: A central processing unit that performs data analysis and sentiment estimation. Examples include AWS and Google Cloud Platform.
[1417] Generative AI models: AI / ML frameworks used for data analysis and sentiment estimation. Specific examples include TensorFlow and PyTorch.
[1418] Text-to-speech software: Generates speech based on estimated emotions. Examples include Amazon Polly and Google Text-to-Speech.
[1419] Speaker: An output device for playing back the generated sound, attached to the animal's collar.
[1420] Specific examples of actions
[1421] Example 1: When a cat is happy to be petted
[1422] 1. The device uses a camera to record the user's actions as they pet the cat, and a microphone to capture the cat's meows.
[1423] 2. The terminal preprocesses this data to remove unwanted noise.
[1424] 3. The terminal sends the pre-processed data to the server.
[1425] 4. The server analyzes the data and estimates whether the cat is "happy" based on its gestures and meows.
[1426] 5. The server generates the voice "Feels good!" which corresponds to the state of being "happy".
[1427] 6. The device sends the generated audio to the speaker and plays it.
[1428] 7. The user hears the cat say "Feels good!" from the cat's collar and understands that the cat is happy.
[1429] Example 2: When the cat looks anxious
[1430] 1. The device uses its camera and microphone to capture any actions or sounds the cat makes when it is startled by something.
[1431] 2. The terminal preprocesses this data.
[1432] 3. The terminal sends the pre-processed data to the server.
[1433] 4. The server analyzes the data and estimates whether the cat is "anxious" based on its behavior and meows.
[1434] 5. The server generates the voice message "What's wrong?" which corresponds to the state of being "anxious".
[1435] 6. The device sends the generated audio to the speaker and plays it.
[1436] 7. The user hears a voice saying "What's wrong?" from the cat's collar, understands that the cat is anxious, and takes action to soothe it.
[1437] Example of a prompt
[1438] "Analyze animal behavioral and auditory data to estimate the animal's emotions and generate appropriate voice-to-text. Specifically, if the animal is happy, generate the voice 'Feels good!', and if it is anxious, generate the voice 'What's wrong?'."
[1439] In this way, the present invention provides support for better communication between animals and their owners, and makes it possible to understand pets' emotions more accurately and in real time.
[1440] The flow of the specific processing in Example 1 will be explained using Figure 11.
[1441] Step 1:
[1442] Data acquisition
[1443] The device uses a camera and microphone installed in the room to capture animal behavior and sounds in real time. The camera captures video of the animals, and the microphone records their audio. Specifically, the camera captures video data frame by frame, and the microphone continuously records audio.
[1444] Input: Animal behavior and sounds
[1445] Output: Raw video data and raw audio data
[1446] Step 2:
[1447] Data preprocessing
[1448] The device preprocesses the acquired video and audio data. Specifically, it removes background noise from the video data and filters the audio data to reduce unwanted noise. For video data, it masks parts other than animals, and for audio data, it applies filtering techniques to remove ambient noise.
[1449] Input: Raw video data and raw audio data
[1450] Output: Pre-processed video data and pre-processed audio data
[1451] Step 3:
[1452] Data transmission
[1453] The terminal sends pre-processed data to the server. This pre-processed video and audio data are transmitted to the server in real time using a high-speed network connection.
[1454] Input: Pre-processed video data and pre-processed audio data
[1455] Output: Data sent to the server
[1456] Step 4:
[1457] Data Analysis
[1458] The server analyzes the received pre-processed data. Specifically, it uses a generative AI model (e.g., TensorFlow or PyTorch) to identify animal behavior, gestures, and facial expressions from video data, and analyze animal sounds from audio data. Based on prompts, it obtains the analysis results.
[1459] Input: Pre-processed video data and pre-processed audio data
[1460] Output: Behavioral analysis results and vocalization analysis results
[1461] Step 5:
[1462] Emotion estimation
[1463] The server estimates the animal's emotions based on the results of data analysis. Specifically, it combines behavioral analysis results and vocalization analysis results to estimate the animal's current emotional state (e.g., "happy," "anxious").
[1464] Input: Behavioral analysis results and vocalization analysis results
[1465] Output: Estimated emotional state
[1466] Step 6:
[1467] Speech generation
[1468] The server generates speech based on estimated emotions. Specifically, it selects appropriate speech text corresponding to the emotional state and generates speech data using speech synthesis software (such as Amazon Polly or Google Text-to-Speech).
[1469] Input: Estimated emotional state
[1470] Output: Generated audio data
[1471] Step 7:
[1472] Audio transmission and playback
[1473] The device receives audio data from the server and transmits it to a speaker attached to the animal's collar, which then plays the audio. Specifically, the device transmits the audio data to the playback speaker via Bluetooth or Wi-Fi, and the speaker outputs the audio.
[1474] Input: Generated audio data
[1475] Output: Audio played through the speaker
[1476] Through these steps, users can understand the animals' emotions in real time, enabling better communication with them.
[1477] (Application Example 1)
[1478] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[1479] Conventional animal emotion estimation systems can estimate an animal's emotions, but they lack the ability to detect abnormal behavior and notify the user in real time. As a result, it is difficult for pet owners and caregivers to respond quickly to abnormal animal behavior. Furthermore, safety issues remain unresolved. This invention aims to improve animal safety and user peace of mind by detecting not only animal emotions but also abnormal behavior and notifying the user.
[1480] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[1481] In this invention, the server includes means for acquiring animal behavior, gestures, facial expressions, and vocalizations from a camera and a microphone; means for preprocessing the data acquired from the camera and microphone; means for analyzing the preprocessed data and estimating the animal's emotions; means for generating sound based on the estimated emotions; means for emitting the generated sound from a speaker attached to the animal's collar; means for analyzing the data to detect abnormal animal behavior; and means for generating a warning based on the detected abnormal behavior and notifying the user. This makes it possible to detect both animal emotions and abnormal behavior in real time and notify the user.
[1482] A "camera" is a device used to capture images of an animal's behavior, gestures, and facial expressions.
[1483] A "microphone" is a sound acquisition device used to record animal sounds.
[1484] "Data preprocessing" refers to the process of removing noise and converting the data format of video and audio data acquired from cameras and microphones.
[1485] "Data analysis" refers to algorithmic processing that identifies animal behavior, gestures, and vocalizations from pre-processed video and audio data, and estimates the animal's emotions and abnormal behavior.
[1486] "Emotion estimation" is the process of inferring what emotional state an animal is in based on its behavior, gestures, and vocalizations.
[1487] "Voice generation" is the process of generating voice messages that convey emotions to the user based on estimated animal emotions.
[1488] A "speaker attached to an animal's collar" is an audio playback device installed on the collar worn by an animal, used to play generated audio messages.
[1489] "Detection of abnormal behavior" is the process of determining whether an animal is behaving in an unusual way based on data analysis.
[1490] "Notification to users" refers to the process of sending warning messages or notifications to users in real time based on detected abnormal behavior.
[1491] "Real-time" is a time standard that means processing or notifications are performed almost instantaneously or with very little delay.
[1492] The present invention is a system that uses a camera and microphone to acquire animal behavior, gestures, facial expressions, and vocalizations, analyzes this data to estimate the animal's emotions and abnormal behavior, and notifies the user as necessary. Specific embodiments for carrying out the present invention are described below.
[1493] System Overview
[1494] This system uses multiple cameras and microphones placed in a room or outdoors to capture animal behavior and vocalizations in real time, and transmits the data to a server for analysis. Based on the analysis results, it emits sounds from a speaker attached to the animal's collar, allowing the user to understand the animal's emotions. Furthermore, it has a function to notify the user of any abnormal behavior if it is detected.
[1495] System Configuration
[1496] 1. Camera and microphone
[1497] Multiple cameras and microphones, installed indoors or outdoors, capture animal behavior and sounds in real time. The cameras continuously record animal movements, while the microphones record sounds such as vocalizations.
[1498] 2. Preprocessing at the terminal
[1499] The device preprocesses the data acquired from the camera and microphone. Specifically, it removes background noise from the video and filters the audio data to reduce noise. This preprocessed data is then sent to the server for further analysis.
[1500] 3. Analysis on the server
[1501] The server analyzes pre-processed data sent from the terminal. It identifies animal behavior, gestures, and facial expressions from image data, and analyzes vocalizations from audio data. This allows it to estimate the animal's emotions and abnormal behavior. The server performs the analysis using machine learning algorithms such as TensorFlow.
[1502] 4. Speech generation
[1503] Based on the emotions of the animal estimated by the server, the appropriate voice text is selected and generated as speech. For example, if the emotion "happy" is estimated, the corresponding speech will be "Feels good!". In addition, if abnormal behavior is detected, a warning voice such as "Be careful!" will be generated.
[1504] 5. Audio playback via speaker
[1505] The generated audio is transmitted via a device to a speaker attached to the animal's collar. The audio emitted from the speaker allows the user to understand the animal's emotions and any abnormal behavior in real time.
[1506] Specific examples of actions
[1507] Example 1: When a dog senses a suspicious person and becomes alert.
[1508] 1. Device: The camera and microphone capture the dog's alert behavior and barking.
[1509] 2. Terminal: Preprocess this data to remove unwanted noise.
[1510] 3. Terminal: Sends pre-processed data to the server.
[1511] 4. Server: Analyzes data, estimates a "vigilant" state from the dog's movements and barks, and further detects abnormal behavior.
[1512] 5. Server: Generates the voice message "Be careful!" corresponding to the "Alert" state.
[1513] 6. Terminal: Sends the generated audio to the speaker and plays it.
[1514] 7. User: Hearing the "Be careful!" voice message from the dog's collar, the user understands that the dog is on alert and takes appropriate action.
[1515] Example 2: When the dog is playing safely
[1516] 1. Device: The camera and microphone capture the dog's playful behavior and barking.
[1517] 2. Terminal: Preprocess this data.
[1518] 3. Terminal: Sends pre-processed data to the server.
[1519] 4. Server: Analyzes data and estimates whether the dog is "enjoying" itself based on its movements and barks.
[1520] 5. Server: Generates the voice message "This is fun!" which corresponds to the "enjoying" state.
[1521] 6. Terminal: Sends the generated audio to the speaker and plays it.
[1522] 7. User: Confirm that the dog is happy by hearing the voice message "Happy!" coming from the dog's collar.
[1523] Examples of use and prompts for generative AI models
[1524] Usage example
[1525] If a security dog detects an intruder, it records the situation with a camera, analyzes any unusual barking sounds with a microphone, and emits a warning message saying, "An intruder has been detected!"
[1526] Example of a prompt:
[1527] basic information:
[1528] Animals used: Dogs
[1529] Usage scenario: Security guarding
[1530] detail:
[1531] Capture a real-time video stream with the camera.
[1532] The sounds are recorded in real time using a microphone.
[1533] Analyze the data using a TensorFlow model to detect anomalies.
[1534] If an anomaly is detected, the user will be notified with an audio alert.
[1535] In this way, the system detects not only the animals' emotions but also abnormal behavior in real time and notifies the user, thereby improving the safety of the animals and the user's sense of security.
[1536] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[1537] Step 1:
[1538] Capture animal behavior and sounds from cameras and microphones.
[1539] Subject: terminal
[1540] Input: Animal movements and sounds
[1541] Output: Unprocessed video and audio data
[1542] Specific operation: The device acquires animal behavior and sounds in real time from multiple cameras and microphones installed indoors or outdoors. This includes capturing video and audio streams.
[1543] Step 2:
[1544] Preprocess the data acquired from the camera and microphone.
[1545] Subject: terminal
[1546] Input: Unprocessed video and audio data
[1547] Output: Pre-processed video and audio data
[1548] Specific operation: The terminal removes background noise from video data and filters audio data to reduce noise. This process uses OpenCV and audio filtering algorithms.
[1549] Step 3:
[1550] The pre-processed data is sent to the server.
[1551] Subject: terminal
[1552] Input: Pre-processed video and audio data
[1553] Output: Data sent to the server
[1554] Specific operation: The terminal sends pre-processed data to the server via the internet or local network. The data may be encrypted for security purposes during this process.
[1555] Step 4:
[1556] The server analyzes the data to estimate the animals' emotions and abnormal behaviors.
[1557] Subject: Server
[1558] Input: Pre-processed video and audio data
[1559] Output: Emotion and abnormal behavior estimation results
[1560] Specific operation: The server uses machine learning algorithms such as TensorFlow to analyze video and audio data and estimate emotions and abnormal behavior from animal behavior, gestures, and vocalizations. Specifically, it uses a trained model to classify emotional states such as "happy," "alert," and "anxious" and detect abnormal behavior.
[1561] Step 5:
[1562] It generates voices based on estimated emotions and abnormal behaviors.
[1563] Subject: Server
[1564] Input: Emotion and abnormal behavior estimation results
[1565] Output: Generated audio data
[1566] Specific operation: The server selects speech text corresponding to the estimated emotions and abnormal behavior, and converts the text into speech. This is done using speech synthesis software (such as the Google Text-to-Speech API).
[1567] Step 6:
[1568] The generated sound is emitted from a speaker attached to the animal's collar.
[1569] Subject: terminal
[1570] Input: Generated audio data
[1571] Output: Sound emitted from an animal collar
[1572] Specific operation: Audio data is generated and sent from the server to the terminal. The terminal then sends the audio data to a speaker attached to the animal's collar. The speaker plays the received audio data in real time.
[1573] Step 7:
[1574] The system notifies the user of abnormal behavior.
[1575] Subject: Server
[1576] Input: Abnormal behavior detection results
[1577] Output: Warning notification to the user
[1578] Specific operation: If abnormal behavior is detected, the server will send a warning message to the user's smartphone or other devices. This will be done using methods such as push notifications, email, or SMS. The user will receive the notification and be able to take prompt action.
[1579] This processing flow allows for the detection of not only animal emotions but also abnormal behavior in real time, and notifies the user, thereby improving animal safety and user peace of mind.
[1580] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[1581] This invention is a system that combines a system for estimating animal emotions with an emotion engine that recognizes user emotions. Specific embodiments of this invention are described below.
[1582] System Overview
[1583] This system uses multiple cameras and microphones installed in a room to capture animal behavior, gestures, facial expressions, and vocalizations, and analyzes this data to estimate the animal's emotions. It also uses cameras and microphones to capture the user's gestures, facial expressions, and voice tone, recognizing the user's emotions. This recognized emotion generates appropriate responses to the animals, which are then provided to the user as voice or other feedback, thereby facilitating interaction.
[1584] System Configuration
[1585] 1. Camera and microphone:
[1586] Multiple cameras and microphones installed in the room capture the behavior, sounds, gestures, and facial expressions of both animals and users in real time. The cameras continuously capture high-resolution video, while the microphones record high-sensitivity audio data.
[1587] 2. Preprocessing at the terminal:
[1588] The device preprocesses data acquired from the camera and microphone. Video data is filtered to remove noise and extract areas of interest for both the animal and the user. Audio data is processed to remove background noise and convert it into a clear audio signal. Furthermore, preprocessing is performed to analyze the user's voice tone and language patterns.
[1589] 3. Analysis on the server:
[1590] The server analyzes the pre-processed data sent from the terminal. First, it applies a model to the video data to identify the actions, gestures, and facial expressions of the animals and the user. This allows it to detect things like whether the cat is being petted or whether the user is smiling.
[1591] 4. Estimating animal emotions:
[1592] The server applies a speech recognition model to estimate emotions from animal sounds and gestures. It analyzes the patterns and phonemes of the sounds to identify emotions such as "happy" or "anxious."
[1593] 5. User emotion recognition:
[1594] The server analyzes the user's gestures, facial expressions, and tone of voice to recognize their emotions. For example, if the user is smiling, it recognizes that they are "having fun," and if they are frowning, it recognizes that they are "anxious."
[1595] 6. Speech generation and interaction proposals:
[1596] The server generates appropriate voice messages based on the animal's emotions and the user's emotions. Furthermore, it suggests appropriate interactions between the user and the animal. For example, if the animal appears anxious and the user is also anxious, the system will make a voice suggestion such as, "Please gently stroke the cat to reassure it."
[1597] 7. Audio playback via speaker:
[1598] The generated audio is transmitted via the device to a speaker attached to the animal's collar. Feedback audio is also transmitted from the device to the user's speaker.
[1599] Specific examples of actions
[1600] Example 1: When a cat is happy to be petted
[1601] 1. Device: The camera records the user's actions as they pet the cat, and the microphone captures the cat's meows.
[1602] 2. Terminal: Preprocess this data to remove unwanted noise.
[1603] 3. Terminal: Sends pre-processed data to the server.
[1604] 4. Server: Analyzes data and estimates whether the cat is "happy" based on its gestures and meows.
[1605] 5. Server: On the other hand, it analyzes the user's facial expressions and tone of voice to recognize when the user is "enjoying" the experience.
[1606] 6. Server: Generates the voice "Feels good!" corresponding to the "happy" state, and simultaneously generates a voice message to the user saying "The cat is happy."
[1607] 7. Terminal: The generated audio is sent to the cat's collar and the user's speaker for playback.
[1608] 8. User: Hearing the voice saying "Feels good!" from the cat's collar, the user understands that the cat is happy and confirms that they themselves are also enjoying it.
[1609] Example 2: When the cat looks anxious
[1610] 1. Device: The camera and microphone capture the cat's startled movements and meows.
[1611] 2. Terminal: Preprocess this data.
[1612] 3. Terminal: Sends pre-processed data to the server.
[1613] 4. Server: Analyzes data and estimates whether the cat is "anxious" based on its behavior and meows.
[1614] 5. Server: On the other hand, it analyzes the user's facial expressions and tone of voice to recognize that the user is also "anxious."
[1615] 6. Server: Generates the voice message "What's wrong?" to indicate an anxious state, and simultaneously generates a voice message to the user saying, "Please gently stroke the cat to reassure it."
[1616] 7. Terminal: The generated audio is sent to the cat's collar and the user's speaker for playback.
[1617] 8. User: Hearing the sound from the speaker, the user understands that the cat is anxious and takes action to reassure the cat by gently stroking it.
[1618] In this way, the system can provide support for users and animals to understand each other's emotions and communicate smoothly.
[1619] The following describes the processing flow.
[1620] Step 1:
[1621] The device initializes the camera and microphone installed in the room and captures animal behavior, gestures, facial expressions, and vocalizations, as well as the user's gestures, facial expressions, and voice tone in real time. The camera continuously captures frame images, and the microphone records audio data.
[1622] Step 2:
[1623] The device preprocesses the acquired video and audio data. First, it removes noise from the video data and extracts the areas of interest for the animals and the user. Next, it filters the audio data to remove background noise and converts it into a clear audio signal.
[1624] Step 3:
[1625] The terminal packages pre-processed animal and user data and transmits it to a server via the internet. The data package includes video frames and filtered audio data.
[1626] Step 4:
[1627] The server receives the data package sent from the terminal and prepares it for analysis. It performs a consistency check on the received data to confirm that there are no missing or corrupted files.
[1628] Step 5:
[1629] The server analyzes the video data and applies models to identify animal behavior, gestures, and facial expressions, as well as user gestures and facial expressions. For example, it can detect whether an animal is being petted or whether the user is smiling.
[1630] Step 6:
[1631] The server analyzes audio data and applies a speech recognition model to estimate emotions from animal sounds. It also analyzes the tone and phonemes of the user's voice to identify their emotions.
[1632] Step 7:
[1633] The server integrates the analysis results of video and audio data to estimate the overall emotions of the animal and the user. For example, if the animal is showing signs of happiness, it is estimated to be "happy," and if the user is enjoying themselves, it is recognized as "enjoying themselves."
[1634] Step 8:
[1635] The server generates an appropriate voice message based on the estimated emotion. For example, if the server estimates that the animal is "happy," it will generate a voice message saying "Feels good!" and provide the user with feedback such as "The cat is happy."
[1636] Step 9:
[1637] The server sends the generated audio data to the terminal. This includes the audio message and associated metadata.
[1638] Step 10:
[1639] The device receives audio data from the server and transmits it to a speaker attached to the animal's collar, while simultaneously transmitting it to the user's speaker for playback.
[1640] Step 11:
[1641] Through voice messages emitted from the speaker, users can understand the emotions of animals and recognize how their own actions affect them. For example, they might hear the voice saying "Feels good!" from the animal's collar, confirm that the cat is happy, and become aware of their own enjoyment through the feedback "The cat is happy."
[1642] This allows users to understand the animals' emotions in real time and improve their relationship with them by interacting with them appropriately.
[1643] (Example 2)
[1644] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[1645] Conventional animal emotion estimation systems only estimate emotions based on animal behavior and vocalizations, and merely provide feedback to the animals. However, since the emotions of users living with animals directly influence the animals' behavior and state, a comprehensive analysis that includes the user's emotions is desirable. Furthermore, providing appropriate interaction suggestions based on the emotions of both the animal and the user can improve the relationship between the animal and the user. Against this backdrop, there is a need for a system that simultaneously analyzes the emotions of both animals and users and provides appropriate feedback and interaction suggestions.
[1646] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[1647] In this invention, the server includes means for preprocessing data acquired from a camera and a microphone, means for analyzing the preprocessed data and estimating the emotions of the animal and the user, and means for generating speech based on the estimated emotions. This enables high-precision analysis of the emotions of the animal and the user, and allows for appropriate feedback and interaction suggestions based on the results.
[1648] A "camera" is an optical device used to capture the behavior, gestures, and facial expressions of animals and users.
[1649] A "microphone" is an audio device used to capture the voices of animals and users.
[1650] "Data preprocessing" refers to operations such as noise reduction and extraction of regions of interest on data acquired from cameras and microphones.
[1651] A "server" is an information processing device that analyzes pre-processed data, estimates the emotions of animals and users, and generates necessary feedback.
[1652] "Emotion estimation" is the process of inferring the emotional state of animals and users from acquired and pre-processed data.
[1653] "Voice generation" is the process of creating appropriate voice messages based on estimated emotions.
[1654] An "output device" is a device that emits the generated sound from an animal's collar or a speaker on the user's side.
[1655] "Interaction suggestion" is a process that proposes appropriate ways for animals and users to interact based on emotion estimation results.
[1656] This invention is a system that simultaneously analyzes the emotions of both animals and users, and provides appropriate feedback and interaction suggestions. Specific embodiments are described below.
[1657] First, multiple high-resolution cameras and high-sensitivity microphones installed in the room are used to capture the behavior, gestures, facial expressions, and voices of the animals and the user in real time. The cameras capture video at 30 frames per second, while the microphones collect audio data at a sampling rate of 20 kHz.
[1658] The terminal performs preprocessing on the acquired data. Specifically, it uses the open-source OpenCV library to denoise video data and extract regions of interest, and the librosa library to reduce background noise in audio data and convert it into a clear audio signal.
[1659] The pre-processed data is sent from the terminal to the server. The server uses the YOLOv3 model for image analysis and the Google Speech-to-Text API for speech analysis to analyze the received data. This analysis identifies the behavior, gestures, facial expressions, sounds, and voice tone of animals and users.
[1660] Next, the server estimates the emotions of both the animals and the users. For animal emotion estimation, a TensorFlow-based speech emotion recognition model is used, while for user emotion estimation, a combination of an image recognition model (e.g., OpenPose) and a speech emotion recognition model is used.
[1661] Based on the estimated emotions, the server generates an appropriate voice message. The Google Cloud Text-to-Speech API is used for voice generation. The generated voice message is sent to an output device attached to the animal's collar and to the user's output device (speaker) for playback.
[1662] As a concrete example, let's consider the case where a cat is enjoying being petted. In this case, the camera records the user's actions as they pet the cat, and the microphone captures the cat's meows. The device preprocesses this data and sends it to the server. The server analyzes the data, estimating the cat's "happy" state from its gestures and meows, while also analyzing the user's facial expressions and tone of voice to recognize that they are "enjoying" the experience. Based on this, the server generates the audio "That feels good!" and provides feedback to the user saying "The cat is happy." This audio is played from the animal's collar and the user's speaker.
[1663] As another example, consider a case where a cat appears anxious. In this case, the camera and microphone capture the cat's startled movements and meows, which the device preprocesses and sends to the server. The server estimates that the cat is "anxious" and similarly recognizes this from the user's facial expressions and tone of voice. The server generates a voice message saying "What's wrong?" and provides feedback to the user saying "Please gently stroke the cat to reassure it." This voice message is also played back from the output device.
[1664] Specific examples of prompt statements are as follows:
[1665] 1. "Explain how a system that uses a camera and microphone to analyze a cat's emotions provides feedback when the cat is happy while being petted."
[1666] 2. "Please describe the procedure for analyzing data from cats that appear surprised and anxious, and then suggesting appropriate interactions for the user."
[1667] As described above, this system analyzes the emotions of animals and users with high accuracy and improves the relationship between animals and users by providing appropriate feedback and interaction suggestions based on the results.
[1668] The flow of the specific processing in Example 2 will be explained using Figure 13.
[1669] Step 1:
[1670] Data acquisition using camera and microphone
[1671] The device uses multiple high-resolution cameras and high-sensitivity microphones installed in the room to capture the behavior, gestures, facial expressions, and voices of animals and users in real time. Specifically, the cameras capture video at a rate of 30 frames per second, and the microphones collect audio data at a sampling rate of 20 kHz. This inputs data such as animal movements and sounds, and the user's facial expressions and voice tone. The output of this step is raw video and audio data.
[1672] Step 2:
[1673] Data preprocessing
[1674] The terminal performs preprocessing on the acquired raw video and audio data. Using the OpenCV library, it removes background noise from the video data and extracts the areas of interest for the animals and the user. It also uses the librosa library to reduce background noise from the audio data, converting it into a clear audio signal. Since the input video data contains noise, it undergoes filtering, and noise reduction is applied to the audio data. The output of this step is the preprocessed video and audio data.
[1675] Step 3:
[1676] Pre-processed data transmission to the server
[1677] The terminal sends pre-processed video and audio data to the server. The data is sent asynchronously using an HTTP POST request. As a specific example, compressed pre-processed data is sent to the server via a REST API. The input to this step is the pre-processed data, and the output is the data transfer to the server.
[1678] Step 4:
[1679] Server-based data analysis
[1680] The server analyzes the received preprocessed data. For image analysis, the YOLOv3 model is used to identify animal and user behavior, gestures, and facial expressions. For audio data analysis, the Google Speech-to-Text API is used to convert vocalizations and tone into text data for further analysis. The input for this step is preprocessed video and audio data, and the output is behavioral and emotion labels as analysis results. A concrete example of its operation is analyzing a cat's posture and vocalizations while being petted to identify that "the cat is happy."
[1681] Step 5:
[1682] Estimating animal emotions
[1683] The server estimates emotions from animal sounds and gestures. Using a TensorFlow-based speech emotion recognition model, it analyzes vocal patterns and phonemes to identify emotions such as "happy" or "anxious." The input for this step is the analyzed audio and video data, and the output is the animal's emotion label. A concrete example of its operation is analyzing a cat's high-pitched meow to estimate the emotion of "happiness."
[1684] Step 6:
[1685] User emotion recognition
[1686] The server analyzes the user's gestures, facial expressions, and voice tone to recognize their emotions. It uses a combination of an image recognition model (e.g., OpenPose) and a voice emotion recognition model. The input for this step is the analyzed video and audio data, and the output is the user's emotion label. A concrete example of its operation is detecting the user's smile and cheerful voice tone to identify a state of "enjoyment."
[1687] Step 7:
[1688] Speech generation and interaction proposals
[1689] The server generates an appropriate voice message based on the animal's emotions and the user's emotions. Using the Google Cloud Text-to-Speech API, the generated voice message is sent to an output device attached to the animal's collar and the user's output device (speaker) for playback. The input for this step is the emotion label, and the output is the generated voice message. A concrete example of its operation is the generation of a message such as "The cat is happy."
[1690] Step 8:
[1691] Audio playback and user feedback
[1692] The terminal transmits and plays the generated audio to an output device attached to the animal's collar and to the user's output device. The user listens to the audio from the speaker and takes appropriate action according to the animal's emotional state. The input for this step is the generated audio message, and the output is the playback of the audio. A concrete example of this operation would be the audio "Feels good!" being played from the cat's collar, and the message "The cat is happy" being played from the user's speaker.
[1693] (Application Example 2)
[1694] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[1695] In traditional pet shops, it was difficult for staff to accurately understand the emotions of customers and their pets, resulting in difficulties in recommending appropriate products and services. Understanding pets' emotions is particularly crucial for providing stress-free service, but effective methods for achieving this were lacking. Therefore, there is a growing need for a system that recognizes the emotions of customers and their pets in real time and provides optimal recommendations and feedback based on that understanding, in order to improve customer satisfaction.
[1696] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[1697] In this invention, the server includes means for acquiring animal behavior, gestures, facial expressions, and vocalizations from a camera and microphone; means for analyzing customer gestures, facial expressions, tone of voice, and content from the data to recognize customer emotions; and means for suggesting optimal products and services based on the emotions of the animals and customers. This makes it possible to recognize the emotions of customers and pets in real time within a pet shop and to provide optimal products and services based on that recognition.
[1698] A "camera" is a device that acquires visual information and converts it into digital images or video data.
[1699] A "microphone" is a device that captures sound and converts it into digital audio data.
[1700] "Data preprocessing" is the process of removing noise from acquired data and converting it into a format suitable for analysis.
[1701] "Means for estimating animal emotions" refers to algorithms or devices that identify and estimate an animal's emotional state based on its behavior, gestures, facial expressions, and vocalizations.
[1702] "Means for recognizing customer emotions" refers to an algorithm or device that identifies and recognizes a customer's emotional state based on their gestures, facial expressions, tone of voice, and content.
[1703] "Means of proposing optimal products and services" refers to algorithms or devices that select and propose appropriate products and services based on the emotional state of animals and customers.
[1704] "Means for generating speech" refers to algorithms or devices that create natural-sounding speech from text.
[1705] A "speaker" is a device that reproduces digital audio data as sound.
[1706] A "server" is a computing device or a system on a network that provides data reception, storage, and analysis functions.
[1707] The embodiments for carrying out this invention will be described in detail below. As an example of the embodiment of the invention, we will particularly illustrate its application to a system aimed at improving customer service in a pet shop.
[1708] System Configuration
[1709] This system recognizes the emotions of customers and their pets in real time and suggests the most suitable products and services based on that. The system includes a camera, microphone, terminal, server, and speaker.
[1710] Hardware configuration
[1711] 1. Camera
[1712] High-resolution cameras (e.g., Logitech C920) capture the behavior, gestures, and facial expressions of customers and pets inside the store.
[1713] 2. Mike
[1714] High-sensitivity microphones (e.g., Blue Yeti) can capture the voices of customers or pets.
[1715] 3. Terminal
[1716] Use a laptop or mobile device (e.g., iPhone) to preprocess data from the camera and microphone.
[1717] 4. Server
[1718] High-performance servers (e.g., AWS EC2) analyze pre-processed data to estimate the emotions of animals and customers.
[1719] 5. Speakers
[1720] Speakers installed in the store and speakers attached to pet collars play sounds generated by the system.
[1721] Software Configuration
[1722] 1. Preprocessing
[1723] The device performs noise reduction on data acquired from the camera and microphone, and detects faces and voices. Specifically, it uses OpenCV and the SpeechRecognition library.
[1724] 2. Data Analysis
[1725] The server uses the DeepFace library to perform facial recognition and estimate customer emotions in order to analyze pre-processed data sent from the terminal. It also uses Google's Speech-to-Text API for speech analysis.
[1726] 3. Emotion estimation
[1727] The server estimates the emotions of animals and customers based on the analyzed data. It identifies emotions such as "happy" or "anxious" from the animals' vocalizations and gestures, and identifies emotions such as "enjoying" or "anxious" from the customers' tone of voice and facial expressions.
[1728] 4. Suggestions and Feedback
[1729] Based on the emotions of animals and customers, the server suggests the most suitable products and services and generates voice messages. Voice generation uses gTTS (Google Text-to-Speech) and is played back using the Playsound library.
[1730] Specific example of processing
[1731] For example, if a customer visits a pet shop with their pet and is looking at products, the system will perform the following actions.
[1732] Example of a prompt
[1733] Camera image analysis prompt:
[1734] "Analyze the dominant emotion from this frame. Frame: {frame_data}"
[1735] Voice analysis prompt:
[1736] "Estimate the customer's emotions from their voice. Voice data: {audio_data}"
[1737] Based on camera footage and audio data, the system analyzes the emotions of the customer and their pet, understanding situations in real time, such as "the pet is interested in a toy" or "the customer is looking for new food." Based on this, the system generates voice messages such as "Your pet is interested in this toy. How about this toy?" and provides feedback to the customer through the speaker.
[1738] This improves the purchasing experience because customers can more easily understand their pet's condition and be offered appropriate products and services.
[1739] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[1740] Step 1:
[1741] The device uses a high-resolution camera and a high-sensitivity microphone to acquire video and audio data of customers and pets in the store. The input consists of video data from the camera and audio data from the microphone, and this data is sent to the next step.
[1742] Step 2:
[1743] The video and audio data acquired by the device are preprocessed. Specifically, for video data, noise is removed using OpenCV, and faces and animal gestures are detected. For audio data, noise is removed using the SpeechRecognition library, and it is converted into clear audio data. The input is the video and audio data acquired in step 1, and the output is the preprocessed data.
[1744] Step 3:
[1745] The terminal sends pre-processed data to the server. The input is pre-processed data, which is sent from the terminal to the server via the network.
[1746] Step 4:
[1747] The server receives pre-processed data and analyzes the video data using the DeepFace library to estimate emotions from the customer's facial expressions. The input is video data sent from the terminal, and the output is the result of the customer's emotion recognition. The prompt message "Customer's face frame: {frame_data}" is used.
[1748] Step 5:
[1749] The server converts audio data into text using Google's Speech-to-Text API, and then analyzes the content and tone of voice to recognize the customer's emotions. The input is pre-processed audio data, and the output is the emotion recognition result based on speech recognition. The prompt message used is "Audio input data: {audio_data}".
[1750] Step 6:
[1751] The server analyzes animal video and audio data to determine their behavior, gestures, and vocalizations, and estimates their emotions. The input is video and audio data transmitted from a terminal, and the output is the estimated animal emotions. DeepFace and an audio analysis model are used.
[1752] Step 7:
[1753] The server generates messages suggesting optimal products and services based on animal emotion estimation results and customer emotion recognition results. gTTS is used to convert text to speech during this process. The input is the emotion recognition and emotion estimation results, and the output is the generated speech message.
[1754] Step 8:
[1755] The device sends the generated voice message to a speaker, which plays it back through the store's speakers and a speaker attached to the pet's collar. The input is the generated voice message, and the output is voice feedback to the customer and their pet.
[1756] Step 9:
[1757] The user listens to feedback from the speaker and selects appropriate products or services based on the system's suggestions. The input is voice feedback, and the output is the user's actions.
[1758] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the controlled object 443 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[1759] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). An example of data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[1760] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the robot 414.
[1761] Furthermore, the emotion identification model 59, acting as an emotion engine, may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to a specific mapping, which is an emotion map (see Figure 9). Similarly, the emotion identification model 59 may also determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.
[1762] Figure 9 shows an emotion map 400 in which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. The closer to the center of the concentric circles, the more primitive the emotions are located. Further out of the concentric circles, emotions representing states and actions arising from mental states are located. Emotion is a concept that includes feelings and mental states. On the left side of the concentric circles, emotions that are generally generated from reactions occurring in the brain are located. On the right side of the concentric circles, emotions that are generally induced by situational judgment are located. Above and below the concentric circles, emotions that are generally generated from reactions occurring in the brain and induced by situational judgment are located. In addition, the emotion of "pleasure" is located on the upper side of the concentric circles, and the emotion of "displeasure" is located on the lower side. Thus, in the emotion map 400, multiple emotions are mapped based on the structure in which emotions arise, and emotions that are likely to occur simultaneously are mapped close together.
[1763] These emotions are distributed at the 3 o'clock position on the Emotion Map 400, and usually fluctuate between feelings of security and anxiety. In the right half of the Emotion Map 400, situational awareness takes precedence over internal feelings, resulting in a calm impression.
[1764] The inside of the Emotion Map 400 represents inner thoughts, while the outside represents actions. Therefore, the further you go from the outside of the Emotion Map 400, the more visible (expressed in actions) your emotions become.
[1765] Here, human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. Similarly, in robots, cars, motorcycles, etc., emotions can be created based on various balances, such as posture and battery level. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. The emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on a system for analyzing brain physiological signals of speech emotion recognition and emotion, Tokushima University, doctoral dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map contains emotions belonging to a region called "response," where sensation is dominant. The right half of the emotion map contains emotions belonging to a region called "situation," where situational awareness is dominant.
[1766] The emotion map defines two emotions that promote learning. One is the emotion around the middle of the negative "repentance" and "reflection" on the situation side. In other words, it is when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is the emotion around the positive "desire" on the reaction side. In other words, it is when the robot has positive feelings such as "I want more" or "I want to know more."
[1767] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values representing each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple training data sets, which are combinations of user input and emotion values representing each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions located close together have similar values, as shown in the emotion map 900 in Figure 10. Figure 10 shows an example where multiple emotions such as "reassured," "calm," and "confident" have similar emotion values.
[1768] The above description primarily focuses on the functions of the data processing device 12 in relation to this disclosure. However, the system related to this disclosure is not necessarily implemented on a server. The system related to this disclosure may be implemented as a general information processing system. This disclosure may be implemented, for example, as a software program that runs on a personal computer or as an application that runs on a smartphone. The method related to this disclosure may be provided to users in SaaS (Software as a Service) format.
[1769] In the above embodiment, an example was given in which a specific process is performed by a single computer 22. However, the technology of this disclosure is not limited thereto, and a distributed processing of the specific process may be performed by multiple computers, including computer 22. For example, a data generation model 58 may be provided in an external device of the data processing device 12, and the external device may generate data according to the input data.
[1770] In the above embodiment, an example was given in which the specific processing program 56 is stored in the storage 32, but the technology of this disclosure is not limited thereto. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-temporary storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-temporary storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes specific processing according to the specific processing program 56.
[1771] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.
[1772] Furthermore, it is not necessary to store the entirety of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store the entirety of the specific processing program 56 in the storage 32; it is acceptable to store only a portion of the specific processing program 56.
[1773] The following types of processors can be used as hardware resources to perform specific processing. Examples of processors include a CPU, a general-purpose processor that functions as a hardware resource to perform specific processing by executing software, i.e., a program. Other examples of processors include dedicated electrical circuits, such as FPGAs (Field-Programmable Gate Arrays), PLDs (Programmable Logic Devices), or ASICs (Application Specific Integrated Circuits), which have circuit configurations specifically designed to perform specific processing. All of these processors have built-in or connected memory, and all of them perform specific processing by using memory.
[1774] The hardware resource that performs a specific process may consist of one of these various processors, or it may consist of a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Alternatively, the hardware resource that performs a specific process may consist of a single processor.
[1775] Examples of configurations using a single processor include, firstly, a configuration in which one or more CPUs and software are combined to form a single processor, and this processor functions as a hardware resource that performs a specific process. Secondly, there is a configuration using a processor that realizes the functions of the entire system, including multiple hardware resources that perform a specific process, on a single IC chip, as exemplified by SoCs (System-on-a-chip). In this way, a specific process is realized using one or more of the above types of processors as hardware resources.
[1776] Furthermore, the hardware structure of these various processors can more specifically utilize electrical circuits that combine circuit elements such as semiconductor devices. Also, the specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps can be deleted, new steps added, or the processing order rearranged, as long as it does not deviate from the main purpose.
[1777] The descriptions and illustrations presented above are detailed explanations of the technical aspects of this disclosure and are merely examples of the technical aspects. For example, the above descriptions of the structure, function, operation, and effect are examples of the structure, function, operation, and effect of the technical aspects of this disclosure. Therefore, it goes without saying that you may delete unnecessary parts, add new elements, or replace elements in the descriptions and illustrations presented above, as long as you do not deviate from the essence of the technical aspects of this disclosure. Furthermore, in order to avoid confusion and facilitate understanding of the technical aspects of this disclosure, explanations of common technical knowledge and the like that do not require special explanation to enable the implementation of the technical aspects of this disclosure have been omitted from the descriptions and illustrations presented above.
[1778] All documents, patent applications, and technical standards described herein are incorporated by reference to the same extent as if each individual document, patent application, and technical standard were specifically and individually noted as being incorporated by reference.
[1779] The following is further disclosed regarding the embodiments described above.
[1780] (Claim 1)
[1781] A means of acquiring animal behavior, gestures, facial expressions, and vocalizations from a camera and microphone,
[1782] means for preprocessing data acquired from the camera and microphone,
[1783] A means of analyzing pre-processed data to estimate animal emotions,
[1784] A means of generating voice based on estimated emotions,
[1785] A means of emitting the generated sound from a speaker attached to the animal's collar,
[1786] A system that includes this.
[1787] (Claim 2)
[1788] The system according to claim 1, further comprising means for transmitting data acquired from the camera and microphone to a server.
[1789] (Claim 3)
[1790] The system according to claim 1, further comprising means for estimating the emotions of an animal using the server.
[1791] "Example 1"
[1792] (Claim 1)
[1793] A means of acquiring animal behavior, gestures, facial expressions, and vocalizations from a camera and microphone,
[1794] means for preprocessing data acquired from the camera and microphone,
[1795] A means for sending pre-processed data to a server,
[1796] A means of analyzing pre-processed data to estimate animal emotions,
[1797] A means of generating voice based on estimated emotions,
[1798] A means of emitting the generated sound from a speaker attached to the animal's collar,
[1799] A system that includes this.
[1800] (Claim 2)
[1801] The system according to claim 1, further comprising means for transmitting audio generated by the server to an animal's speaker via a terminal.
[1802] (Claim 3)
[1803] The system according to claim 1, further comprising means for playing the generated sound from a speaker attached to the animal's collar.
[1804] "Application Example 1"
[1805] (Claim 1)
[1806] A means of acquiring animal behavior, gestures, facial expressions, and vocalizations from a camera and microphone,
[1807] means for preprocessing data acquired from the camera and microphone,
[1808] A means of analyzing pre-processed data to estimate animal emotions,
[1809] A means of generating voice based on estimated emotions,
[1810] A means of emitting the generated sound from a speaker attached to the animal's collar,
[1811] A means of analyzing data to detect abnormal animal behavior,
[1812] A means of generating a warning based on detected abnormal behavior and notifying the user of it,
[1813] A system that includes this.
[1814] (Claim 2)
[1815] The system according to claim 1, further comprising means for transmitting data acquired from the camera and microphone to a server.
[1816] (Claim 3)
[1817] The system according to claim 1, further comprising means for estimating the emotions and abnormal behavior of animals using the server.
[1818] "Example 2 of combining an emotion engine"
[1819] (Claim 1)
[1820] A means of acquiring the behavior, gestures, facial expressions, and voices of animals and users from a camera and microphone,
[1821] means for preprocessing data acquired from the camera and microphone,
[1822] A means for sending pre-processed data to a server,
[1823] A means for analyzing pre-processed data and estimating the emotions of animals and users,
[1824] A means of generating voice based on estimated emotions,
[1825] A means for emitting the generated sound from an output device attached to an animal's collar and from a user's output device,
[1826] A system that includes this.
[1827] (Claim 2)
[1828] The system according to claim 1, further comprising means for analyzing the emotions of animals and users using the server.
[1829] (Claim 3)
[1830] The system according to claim 1, further comprising means for the server to propose interactions based on the emotions of animals and users.
[1831] "Application example 2 when combining with an emotional engine"
[1832] (Claim 1)
[1833] A means of acquiring animal behavior, gestures, facial expressions, and vocalizations from a camera and microphone,
[1834] means for preprocessing data acquired from the camera and microphone,
[1835] A means of analyzing pre-processed data to estimate animal emotions,
[1836] A means of recognizing the customer's emotions by analyzing the customer's gestures, facial expressions, tone of voice, and content from the aforementioned data,
[1837] A means of proposing the most suitable products and services based on the emotions of animals and customers,
[1838] Means for generating voices based on estimated animal and recognized customer emotions,
[1839] A method for emitting the generated sound from speakers inside the pet shop,
[1840] A system that includes this.
[1841] (Claim 2)
[1842] The system according to claim 1, further comprising means for transmitting data acquired from the camera and microphone to a server.
[1843] (Claim 3)
[1844] The system according to claim 1, further comprising means for estimating and recognizing the emotions of animals and customers via the server. [Explanation of Symbols]
[1845] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Devices 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robots< / url:> < / url:> < / url:> < / url:>
Claims
1. A means of acquiring animal behavior, gestures, facial expressions, and vocalizations from a camera and microphone, means for preprocessing data acquired from the camera and microphone, A means of analyzing pre-processed data to estimate animal emotions, A means of generating voice based on estimated emotions, A means of emitting the generated sound from a speaker attached to the animal's collar, A system that includes this.
2. The system according to claim 1, further comprising means for transmitting data acquired from the camera and microphone to a server.
3. The system according to claim 1, further comprising means for estimating the emotions of an animal using the server.
Citation Information
Patent Citations
Persona chatbot control method and system
JP2022180282A