System
A system using a terminal, server, and AI model translates pets' movements and sounds into natural language, addressing the challenge of understanding pets' emotions and desires, enhancing communication and care.
Patent Information
- Application Number
- JP2024121559
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-07-26
- Publication Date
- 2026-02-05
AI Technical Summary
Existing systems struggle to accurately understand pets' emotions and desires, making it difficult for owners to provide appropriate care and communication with their pets.
A system comprising a terminal for collecting pet movements, facial expressions, and sounds, a server for data analysis using computer vision and voice recognition algorithms, and an AI model to translate these into natural language for user notification.
Enables owners to understand their pets' emotions and desires in real time, improving communication and enabling timely responses.
Smart Images

Figure 2026019811000001_ABST
Abstract
Description
[Technical Field]
[0001] The technology of the present disclosure relates to a system. [Background technology]
[0002] Patent document 1 discloses a persona chatbot control method performed by at least one processor, the method including the steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to a description of the chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance. [Prior art documents] [Patent documents]
[0003] [Patent Document 1] Japanese Patent Publication No. 2022-180282 Summary of the Invention [Problem to be solved by the invention]
[0004] One of the challenges of keeping pets is that it is difficult for owners to understand their pets' emotions and desires. There is also a need for improved health management and communication with pets. In particular, accurately analyzing the diverse sounds, movements, and facial expressions of pets and translating them into human language is a highly technical problem. There is a need to provide a system that can solve this problem and enable owners to hear the "real voice" of their pets. [Means for solving the problem]
[0005] The present invention provides a system including a terminal for collecting a pet's movements, facial expressions, and cries; a server for receiving data transmitted from the terminal; means for analyzing the received data and estimating the pet's emotions and desires, including an AI model; and means for translating the estimated emotions and desires into natural language and notifying the user. Furthermore, data protection is ensured by a configuration including a security protocol for securely transmitting the collected data. Furthermore, analysis accuracy is improved by utilizing computer vision algorithms and voice recognition algorithms to estimate the pet's emotions and desires. In this way, owners can understand their pet's specific emotions and desires in real time, thereby improving communication with their pets.
[0006] A "terminal" is a device used to collect pet movements, facial expressions, and sounds.
[0007] A "server" is a computer system that receives and analyzes data sent from a terminal.
[0008] "Reception" refers to the act of the server taking in data sent from the terminal.
[0009] "Analysis" refers to the process of processing the received data using an AI model to estimate the pet's emotions and desires.
[0010] An "AI model" is an algorithm that has been pre-trained using machine learning and deep learning to extract features from data and infer pet emotions and needs.
[0011] "Natural language" refers to a commonly used human language, in this case Japanese or English.
[0012] A "security protocol" is a technology that protects data during transmission and prevents unauthorized access or tampering.
[0013] A "computer vision algorithm" is an algorithm that analyzes video data to detect specific patterns or features.
[0014] A "voice recognition algorithm" is an algorithm that analyzes voice data to detect specific patterns or characteristics.
[0015] "Translation" is the act of converting the analysis results into natural language that the user can understand. [Brief explanation of the drawings]
[0016] [Figure 1] 1 is a conceptual diagram showing an example of the configuration of a data processing system according to a first embodiment. [Figure 2] 1 is a conceptual diagram showing an example of main functions of a data processing device and a smart device according to a first embodiment. [Figure 3] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a second embodiment. [Figure 4] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and smart glasses according to a second embodiment. [Figure 5] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a third embodiment. [Figure 6] FIG. 11 is a conceptual diagram showing an example of main functions of a data processing device and a headset-type terminal according to a third embodiment. [Figure 7] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a fourth embodiment. [Figure 8] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and a robot according to a fourth embodiment. [Figure 9] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 10] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 11] FIG. 3 is a sequence diagram showing a processing flow of the data processing system according to the first embodiment. [Figure 12]FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 1. [Figure 13] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system according to the second embodiment when an emotion engine is combined. [Figure 14] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 2 when an emotion engine is combined. DETAILED DESCRIPTION OF THE INVENTION
[0017] An example of an embodiment of a system according to the technology of the present disclosure will be described below with reference to the accompanying drawings.
[0018] First, the terms used in the following description will be explained.
[0019] In the following embodiments, a coded processor (hereinafter simply referred to as a "processor") may be a single arithmetic device or a combination of multiple arithmetic devices. Furthermore, a processor may be a single type of arithmetic device or a combination of multiple types of arithmetic devices. Examples of arithmetic devices include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), and an APU (Accelerated Processing Unit).
[0020] In the following embodiments, a coded RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a working memory by a processor.
[0021] In the following embodiments, the coded storage is one or more non-volatile storage devices that store various programs, various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), and magnetic tapes.
[0022] In the following embodiments, a communication I / F (Interface) with a symbol is an interface including a communication processor, an antenna, etc. The communication I / F controls communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), Bluetooth (registered trademark), etc.
[0023] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." In other words, "A and / or B" means that it may be only A, only B, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" is also applied when three or more things are expressed connected by "and / or."
[0024] [First embodiment]
[0025] FIG. 1 shows an example of the configuration of a data processing system 10 according to the first embodiment.
[0026] 1, a data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.
[0027] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0028] The smart device 14 includes a computer 36, a reception device 38, an output device 40, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The reception device 38, the output device 40, and the camera 42 are also connected to the bus 52.
[0029] The reception device 38 includes a touch panel 38A, a microphone 38B, and the like, and receives user input. The touch panel 38A detects contact with an indicator (for example, a pen or a finger) to receive user input by the touch of the indicator. The microphone 38B detects the user's voice to receive user input by voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.
[0030] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form of expression that the user 20 can perceive (for example, audio and / or text). The display 40A displays visible information such as text and images in accordance with instructions from the processor 46. The speaker 40B outputs audio in accordance with instructions from the processor 46. The camera 42 is a compact digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.
[0031] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54.
[0032] FIG. 2 shows an example of the main functions of the data processing device 12 and the smart device 14.
[0033] 2, in the data processing device 12, a specific process is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific process is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0034] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0035] In the smart device 14, the processor 46 performs the reception output process. The storage 50 stores a reception output program 60. The reception output program 60 is used in conjunction with the specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0036] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0037] System configuration overview
[0038] The present invention is a system that includes a "terminal" for collecting pet movements, facial expressions, and cries, a "server" for receiving data sent from the terminal, an "AI model" for analyzing the received data and inferring the pet's emotions and desires, and a "notification system" that translates the inferred emotions and desires into natural language and notifies the user.
[0039] Data collection details
[0040] The "terminal" is equipped with a camera and microphone, and collects the pet's movements, facial expressions, and sounds in real time. The collected data is divided into data packets at regular intervals and sent from the "terminal" to the "server." A security protocol is used to ensure the data is secure.
[0041] Data reception and analysis
[0042] The "server" receives the data packets sent from the terminal. The received data is first decoded and then processed separately into video data and audio data.
[0043] The "server" analyzes the video data using computer vision algorithms to identify the pet's facial expressions and movements (for example, tail movements). It also analyzes the audio data using speech recognition algorithms to identify the pattern, pitch, and frequency of the pet's cries. This allows the data's features to be extracted.
[0044] Estimating emotions and desires
[0045] The extracted features are input into an "AI model," which uses a pre-trained machine learning algorithm to infer the pet's emotions and desires. For example, if a pet frequently wags its tail and makes a high-pitched meow, it can be inferred to have a desire to play.
[0046] Natural language translation and notification
[0047] The estimated emotions and desires are translated into natural language. Specifically, the output of the AI model is converted into messages such as "I'm hungry" or "Play with the ball!". This translated message is then resent from the "server" to the "device," which then notifies the user.
[0048] Specific example explanation
[0049] Example 1: System behavior when a dog wants to play
[0050] 1. The "terminal" captures the dog's movements with a camera and collects the dog's tail wagging vigorously and high-pitched barking.
[0051] 2. The "terminal" sends the collected data to the "server."
[0052] 3. The "server" receives the data and identifies tail movements from the video data and high-pitched cries from the audio data.
[0053] 4. The “server” inputs these features into the “AI model.”
[0054] 5. The AI model uses these characteristics to infer that the dog wants to play.
[0055] 6. The "server" translates the result "I want to play" into natural language and makes it "Play with the ball!"
[0056] 7. The "terminal" notifies the user with the message "Play with the ball!"
[0057] Example 2: System behavior when the cat is hungry
[0058] 1. The "terminal" collects the cat's meowing (a low, sustained sound) and its movements around the dish.
[0059] 2. The "terminal" sends the collected data to the "server."
[0060] 3. The "server" receives the data and identifies the location from the video data (walking around the dish) and the sound from the audio data (a low, sustained sound).
[0061] 4. The “server” inputs these features into the “AI model.”
[0062] 5. The "AI model" uses these characteristics to infer that the cat is hungry.
[0063] 6. The "server" translates the result "hungry" into natural language and sets it as "I'm hungry."
[0064] 7. The "terminal" notifies the user with the message "I'm hungry."
[0065] In this way, the present invention can be implemented as a system that allows users to understand their pet's specific emotions and desires in real time and take appropriate action.
[0066] The processing flow will be explained below.
[0067] Step 1:
[0068] The device uses a camera and microphone to collect the pet's movements, facial expressions, and sounds in real time. The collected data is then separated into video data and audio data, and each is split into separate data packets.
[0069] Step 2:
[0070] The device sends the collected data packets to a server, where a security protocol is used to ensure the data is protected.
[0071] Step 3:
[0072] The server receives the data packets sent from the terminal. The received data is first decoded and then separated into video data and audio data for processing.
[0073] Step 4:
[0074] The server runs the video data through computer vision algorithms to analyze your pet's facial expressions and movements, specifically identifying tail movements, facial expressions, and overall body movements.
[0075] Step 5:
[0076] The server runs the audio data through a speech recognition algorithm to identify the pattern, pitch, and frequency of your pet's cries, including the pitch of the cries and repetitive patterns of cries.
[0077] Step 6:
[0078] The server integrates the results of the video and audio analysis to extract features to estimate the pet's emotions and desires, such as high-pitched meowing and tail wagging.
[0079] Step 7:
[0080] The server inputs the extracted features into a pre-trained AI model, which uses patterns learned from past data to infer the pet's emotions and desires.
[0081] Step 8:
[0082] The server receives the inferences output by the AI model and maps them to a list of predefined emotions and desires, such as "hunger," "want to play," "anxiety," and "relaxation."
[0083] Step 9:
[0084] The server translates the estimated emotions and desires into natural language. For example, "hungry" becomes "I'm hungry" and "I want to play" becomes "Play with my ball!"
[0085] Step 10:
[0086] The server then sends the translated message to the user's device, again using security protocols to protect the data.
[0087] Step 11:
[0088] The terminal receives the translated message sent from the server, and the received message is formatted to be displayed on the user interface.
[0089] Step 12:
[0090] The device will display messages to the user that indicate the pet's emotions and needs, such as "I'm hungry!" or "Play with my ball!", displayed as pop-ups or notification banners.
[0091] Step 13:
[0092] The user checks the messages displayed on the device and takes action according to the pet's request, such as feeding the pet or playing with a ball.
[0093] This series of processes allows users to understand their pet's specific emotions and desires in real time and respond appropriately.
[0094] Example 1
[0095] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0096] Conventional pet management systems have had problems with accurately monitoring pet behavior and conditions, and owners are unable to properly understand their pets' emotions and needs, which increases stress for pets and burdens on owners. Furthermore, insufficient coordination in data collection, analysis, and notification makes it difficult to understand pets' conditions in real time.
[0097] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0098] In this invention, the server includes means for decoding received data and processing the data separately into video data and audio data, means for analyzing the video data using a computer vision algorithm and the audio data using a speech recognition algorithm, and means for inputting the extracted features into an AI model to estimate emotions and desires. This makes it possible to instantly estimate emotions and desires from the pet's movements and cries, and notify the owner in natural language.
[0099] The "terminal" is a device equipped with a camera and microphone to collect the movements, facial expressions, and sounds of your pet.
[0100] A "server" is a computer system that receives, decrypts, and analyzes data sent from a terminal.
[0101] A "data packet" is a unit of information that divides collected data and is sent from a terminal to a server.
[0102] "Decoding" is the process of restoring encoded data to its original form.
[0103] "Video data" refers to data that visually records the movements and expressions of pets collected by a camera.
[0104] "Audio data" refers to data that records the sounds of pets as collected by a microphone.
[0105] A "computer vision algorithm" is a technology for analyzing video data and recognizing objects and actions.
[0106] A "voice recognition algorithm" is a technology for analyzing voice data and recognizing sounds and patterns.
[0107] "Features" are important information extracted from video and audio data for estimating a pet's emotions and desires.
[0108] The "AI model" is a system that is pre-trained using a machine learning algorithm and estimates a pet's emotions and desires based on input features.
[0109] "Natural language" refers to language that humans use on a daily basis, and is a language that expresses estimated emotions and desires in a way that is easy for users to understand.
[0110] A "security protocol" is a set of rules and procedures for securely communicating data.
[0111] "Notification" is the act of informing the user of the results of estimated emotions and desires.
[0112] MODE FOR CARRYING OUT THE INVENTION
[0113] The present invention is a system that includes a terminal that collects pet movements, facial expressions, and cries, a server that receives and analyzes the data, and a system that notifies the user of the estimated emotions and desires. Specifically, by using a pet monitoring device, a data analysis system, and a communication means, it is possible to grasp the pet's condition in real time and take appropriate measures.
[0114] System configuration overview
[0115] The present invention includes a "terminal" for collecting pet movements, facial expressions, and cries, a "server" for receiving data sent from the terminal, an "AI model" for analyzing the received data and inferring the pet's emotions and desires, and a "notification system" for translating the inferred emotions and desires into natural language and notifying the user.
[0116] Data collection details
[0117] The "terminal" is equipped with a camera and microphone, and collects the pet's movements, facial expressions, and cries in real time. For example, the camera records the pet's tail movements and eye expressions as video data, while the microphone collects the pet's cries as audio data. The collected data is divided into data packets at regular intervals and sent from the "terminal" to the "server." A security protocol is used for data transmission to ensure the safety of the data.
[0118] Data reception and analysis
[0119] The "server" receives data packets sent from the device. The received data is first decoded and then processed separately into video data and audio data. The "server" analyzes the video data using a computer vision algorithm (e.g., OpenCV) to identify the pet's facial expressions and movements. It also analyzes the audio data using a speech recognition algorithm (e.g., a TensorFlow or PyTorch-based model) to identify the pattern, pitch, and frequency of the pet's cries. This allows the data features to be extracted.
[0120] Estimating emotions and desires
[0121] The extracted features are input into an "AI model." The AI model uses a pre-trained machine learning algorithm (e.g., a generative AI model) to infer the pet's emotions and desires. For example, if the pet frequently wags its tail and makes a high-pitched meow, it can be inferred that the pet wants to play.
[0122] Natural language translation and notification
[0123] The estimated emotions and desires are translated into natural language. Specifically, the results output by the AI model are converted into messages such as "I'm hungry" or "Play with the ball!". These translated messages are then resent from the "server" to the "device," which then notifies the user. For example, the user can be notified by a push notification or voice alert on a smartphone app.
[0124] Specific example explanation
[0125] Example 1: System behavior when a dog wants to play
[0126] 1. The device captures the dog's movements with a camera and collects the dog's vigorously wagging tail and high-pitched barking sounds.
[0127] 2. The device sends the collected data to the server.
[0128] 3. The server receives the data and identifies tail movements from the video data and high-pitched cries from the audio data.
[0129] 4. The server inputs these features into the AI model.
[0130] 5. The AI model uses these characteristics to infer that the dog wants to play.
[0131] 6. The server translates the result "I want to play" into natural language and makes it "Play with the ball!"
[0132] 7. The device notifies the user with the message "Play with the ball!"
[0133] Example 2: System behavior when the cat is hungry
[0134] 1. The device collects the cat's meowing (a low, sustained sound) and its movements around the dish.
[0135] 2. The device sends the collected data to the server.
[0136] 3. The server receives the data and identifies the location from the video data (walking around the dish) and the sound from the audio data (a low, sustained sound).
[0137] 4. The server inputs these features into the AI model.
[0138] 5. The AI model uses these characteristics to infer that the cat is hungry.
[0139] 6. The server translates the result "hungry" into natural language and gives it the value "I'm hungry."
[0140] 7. The device notifies the user with the message "I'm hungry."
[0141] Example prompts to input to the generative AI model
[0142] 1. Example prompt for identifying a dog's desire to play: "When a dog wags its tail frequently and makes a high-pitched whine, it means it wants to play. Express this in natural language."
[0143] 2. Example prompt for identifying a cat's hunger state: "When a cat meows in a low, sustained tone and paces around its food bowl, it is likely hungry. Please describe this in natural language."
[0144] In this way, the present invention can be implemented as a system that allows users to understand their pet's specific emotions and desires in real time and take appropriate action.
[0145] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0146] Program processing steps
[0147] Step 1: Collect data from the device
[0148] The device uses a camera and microphone to collect your pet's movements, facial expressions, and cries in real time.
[0149] Input: Real-time pet movements, facial expressions, and sounds
[0150] Output: Collected video and audio data
[0151] Specific behavior:
[0152] The camera records your pet's tail movements and facial expressions as video data, while the microphone captures your pet's barks as audio data, such as when your dog wags its tail or barks in a high-pitched voice.
[0153] Step 2: Send data from the device to the server
[0154] The device divides the collected data into data packets and sends them to the server using a security protocol.
[0155] Input: Collected video and audio data
[0156] Output: Data packets sent from the device to the server
[0157] Specific behavior:
[0158] The collected data is divided into packets of a certain size, encrypted using the TLS protocol, and sent to the server.
[0159] Step 3: Server receives and decrypts data
[0160] The server receives the data packets sent from the terminal, decodes them, and separates them into video data and audio data.
[0161] Input: Data packets sent from the device
[0162] Output: Decoded video and audio data
[0163] Specific behavior:
[0164] The received data packets are decrypted using the TLS protocol and separated into collected video and audio data, for example, tail movements recorded by a camera or barks recorded by a microphone.
[0165] Step 4: Video data analysis by the server
[0166] The server analyzes the video data using computer vision algorithms to identify the pet's facial expressions and movements.
[0167] Input: Decoded video data
[0168] Output: Features of pet's facial expressions and movements
[0169] Specific behavior:
[0170] Using OpenCV, features such as tail wagging and eye movement are extracted from video data. For example, if a cat is wagging its tail vigorously, that movement can be identified.
[0171] Step 5: Server analyzes the audio data
[0172] The server analyzes the audio data using a voice recognition algorithm to identify the pattern, pitch and frequency of your pet's cries.
[0173] Input: Decoded audio data
[0174] Output: Features of pet sounds
[0175] Specific behavior:
[0176] Use TensorFlow or PyTorch to identify high-pitched meows and patterns in audio data, such as detecting the low, sustained meow of a cat.
[0177] Step 6: Feature extraction by the server
[0178] The server integrates the features extracted from the video and audio data and inputs them into the AI model.
[0179] Input: Features of pet's facial expressions and movements, features of pet's cries
[0180] Output: Integrated features
[0181] Specific behavior:
[0182] The identified tail movements from the video data and the call patterns from the audio data are integrated into a single dataset.
[0183] Step 7: Server inputs AI model
[0184] The server inputs the extracted features into an AI model to estimate emotions and desires.
[0185] Input: Integrated features
[0186] Output: Estimated pet's emotions and desires
[0187] Specific behavior:
[0188] The integrated features are input into a generative AI model to obtain an inference. For example, if a dog is wagging its tail and making a high-pitched bark, it may infer that it wants to play.
[0189] Step 8: Server-based translation into natural language
[0190] The server translates the estimated emotions and desires into natural language.
[0191] Input: Inference results from the AI model
[0192] Output: A message expressed in natural language
[0193] Specific behavior:
[0194] The result "I want to play" obtained from the AI model is converted into the message "Play with the ball!"
[0195] Step 9: Sending translation results from the server to the device
[0196] The server retransmits the translated message to the terminal.
[0197] Input: A message expressed in natural language
[0198] Output: Messages sent to the terminal
[0199] Specific behavior:
[0200] Using security protocols, send a translated message to the device, for example, "Play with the ball!"
[0201] Step 10: Notify the user via the device
[0202] The terminal receives the message sent from the server and notifies the user of the message.
[0203] Input: The message sent from the server
[0204] Output: The notification the user receives
[0205] Specific behavior:
[0206] The device uses push notifications and audio alerts on a smartphone app to notify users with messages like, "Play with the ball!"
[0207] Step 11: User Action
[0208] The user checks the notification from the device and responds to the pet.
[0209] Input: Notifications from your device
[0210] Output: Specific responses to pets
[0211] Specific behavior:
[0212] The user checks the notification and takes action based on the message, such as playing with their pet. For example, if they receive the message "Play with a ball!", they will get a ball and start playing with their dog.
[0213] The above is a description of the specific processing steps in the system program.
[0214] (Application example 1)
[0215] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0216] In recent years, there has been a demand for technology that can understand pet emotions and desires and enable owners to provide appropriate care. In particular, there is a need for a system that can grasp what pets want in real time and purchase appropriate products based on that information. However, current technology has difficulty accurately estimating pet emotions and desires and making product purchase suggestions based on those estimations. Therefore, the challenge is to provide a system that can analyze pet emotions and desires in real time and automatically make product purchase suggestions.
[0217] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0218] In this invention, the server includes a terminal for collecting the movements, expressions, and cries of the pet, a means for receiving data transmitted from the terminal, a means including an AI model for analyzing the received data and estimating the pet's emotions and desires, a means for translating the estimated emotions and desires into natural language and notifying the user, and a means for automatically generating product purchase suggestions based on the estimated emotions and desires and notifying the user. This makes it possible to grasp the pet's emotions and desires in real time and automatically suggest appropriate products.
[0219] The "terminal" is a device that collects your pet's movements, facial expressions, and sounds in real time, and includes a camera and microphone.
[0220] The "server" is a computer that receives data sent from the terminal, analyzes it, and infers the pet's emotions and desires.
[0221] An "AI model" is an analytical model that uses machine learning algorithms to estimate a pet's emotions and desires.
[0222] "Means of translating into natural language" refers to the function of converting the inferred results output by the AI model into words that are easy for users to understand, such as "I'm hungry" or "I want to play."
[0223] "Means for generating product purchase suggestions" refers to a function that selects appropriate products (e.g., pet food or toys) based on estimated emotions and desires and notifies the user of the suggestions.
[0224] System Overview
[0225] This invention is a system that collects and analyzes pet movements, facial expressions, and cries in real time to estimate the pet's emotions and desires and make appropriate product purchase suggestions to the user. It is composed of the following components.
[0226] Key Components
[0227] 1. Device:
[0228] The terminal is a device equipped with a camera and microphone that captures your pet's movements, facial expressions, and sounds in real time, and the data is then sent to a server using a security protocol.
[0229] 2. Server:
[0230] The server receives the data sent from the device and analyzes it to estimate the pet's emotions and needs. The analysis is performed using an AI model (machine learning algorithm). The server also translates the analysis results into natural language that is easy for the user to understand, and then generates appropriate product purchase suggestions.
[0231] 3. AI model:
[0232] The AI model, which uses the collected data to infer your pet's emotions and needs, includes computer vision and speech recognition algorithms.
[0233] 4. User Notification System:
[0234] The system notifies the user of the estimated emotions and desires, along with product purchase suggestions based on those emotions and desires. The user receives these notifications via their smartphone.
[0235] Software and Hardware Configuration
[0236] Hardware:
[0237] Camera: Capture videos and photos of your pet.
[0238] Microphone: Collects pet sounds.
[0239] Server: A computer that performs data analysis.
[0240] software:
[0241] OpenCV: Used for preprocessing video data.
[0242] Keras: Used to load and use AI models.
[0243] sounddevice: Used to collect audio data.
[0244] Requests: Used to send requests to the notification service.
[0245] How it works
[0246] The server receives the video and audio data sent from the device at runtime and pre-processes the received data appropriately. The video data is resized and normalized using OpenCV before being input to the AI model. The audio data is similarly pre-processed using the sounddevice and then input to the AI model.
[0247] The AI model uses this preprocessed data to infer the pet's emotions and desires. The inference results are translated into natural language and converted into user-friendly expressions such as "I'm hungry" or "I want to play." Based on this translation, the model then generates appropriate product purchase suggestions (e.g., pet food or toys).
[0248] Specific examples of implementation
[0249] Example 1: When a dog wants to play
[0250] 1. The device captures the dog's movements with a camera and collects the dog's vigorously wagging tail and high-pitched barking sounds.
[0251] 2. The collected data is sent to the server and separated into video data and audio data.
[0252] 3. The video and audio data are preprocessed and then input into the AI model.
[0253] 4. The AI model analyzes the data and infers that the dog wants to play.
[0254] 5. The server generates and sends the notification "Play with the ball!" and the suggestion "Would you like to buy a toy?" to the user.
[0255] Example prompt sentence:
[0256] "Write a Python program that monitors your pet's behavior and generates a notification saying 'Would you like to buy pet food?' if it detects that your pet is hungry."
[0257] In this way, a system can be provided for understanding a pet's emotions and desires in real time.
[0258] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[0259] Step 1:
[0260] The device uses a camera and microphone to collect the pet's movements, facial expressions, and cries in real time. Specifically, the camera captures the pet's movements as video, and the microphone records the pet's cries as audio data. These data are collected. The input is real-time video and audio data, and the output is the collected video and audio data.
[0261] Step 2:
[0262] The terminal divides the collected video and audio data into data packets at regular intervals and sends them to the server using a secure protocol. Specifically, the data is encrypted and sent using the HTTPS protocol. The input is the collected video and audio data, and the output is the encrypted data packets.
[0263] Step 3:
[0264] The server receives the data packets sent from the terminal and decrypts them. First, it receives the data, decrypts it, and returns it to the original video and audio data. The input is the encrypted data packet, and the output is the decrypted video and audio data.
[0265] Step 4:
[0266] The server analyzes the decoded data using computer vision and speech recognition algorithms. Specifically, it uses OpenCV to resize and normalize the video data. The audio data is preprocessed to extract audio features. The input is the decoded video and audio data, and the output is the preprocessed video and audio data.
[0267] Step 5:
[0268] The server inputs the preprocessed video and audio data into an AI model to estimate the pet's emotions and desires. Specifically, the data is input into a machine learning model trained using Keras to obtain an inference result. The input is the preprocessed video and audio data, and the output is the inference result of the pet's emotions and desires.
[0269] Step 6:
[0270] The server translates the inferences obtained from the AI model into natural language. For example, it converts them into messages such as "I want to play" or "I'm hungry." The input is the inferences from the AI model, and the output is a message in natural language.
[0271] Step 7:
[0272] The server generates appropriate product purchase suggestions based on the translated message, such as "Would you like to buy a toy?" or "Would you like to buy pet food?" The input is the translated message, and the output is the product purchase suggestion message.
[0273] Step 8:
[0274] The terminal receives the product purchase suggestion message sent from the server and notifies the user. The user can confirm the suggestion through a notification on their smartphone. The input is the product purchase suggestion message, and the output is a notification to the user.
[0275] Furthermore, an emotion engine that estimates the user's emotion may be combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59 and perform identification processing using the user's emotion.
[0276] System configuration overview
[0277] The present invention is a system that includes a "terminal" for collecting the movements, facial expressions, and cries of a pet, a "server" for receiving data sent from the terminal, an "AI model" for analyzing the received data and inferring the pet's emotions and desires, a "notification system" for translating the inferred emotions and desires into natural language and notifying the user, and an "emotion engine" for recognizing the user's emotions.
[0278] Data collection details
[0279] The "terminal" is equipped with a camera and microphone, and collects the pet's movements, facial expressions, and sounds in real time. The collected data is divided into data packets at regular intervals and sent from the "terminal" to the "server." A security protocol is used to ensure the data is secure.
[0280] Data reception and analysis
[0281] The "server" receives the data packets sent from the terminal. The received data is first decoded and then processed separately into video data and audio data.
[0282] The "server" analyzes the video data using computer vision algorithms to identify the pet's facial expressions and movements (for example, tail movements). It also analyzes the audio data using speech recognition algorithms to identify the pattern, pitch, and frequency of the pet's cries. This allows the data's features to be extracted.
[0283] Estimating emotions and desires
[0284] The extracted features are input into an "AI model," which uses a pre-trained machine learning algorithm to infer the pet's emotions and desires. For example, if the pet wags its tail and makes a high-pitched meow, it can be inferred to be wanting to play.
[0285] Natural language translation and notification
[0286] The estimated emotions and desires are translated into natural language. Specifically, the output of the AI model is converted into messages such as "I'm hungry" or "Play with the ball!" These translated messages are then resent from the "server" to the "device" and notified to the user via the "device."
[0287] Recognizing user emotions and proposing countermeasures
[0288] In addition, the user's voice and facial expressions are also collected from the "terminal" and sent to the "server." The "emotion engine" analyzes this data to recognize the user's emotions. Using voice recognition algorithms and facial expression analysis algorithms, it identifies emotions from the user's voice tone and facial expressions.
[0289] Proposing solutions based on user sentiment
[0290] The server then uses the user's emotional information to suggest ways to help the pet. For example, if the user is tired, the server suggests creating a relaxing atmosphere for the pet.
[0291] Specific example explanation
[0292] Example 1: System behavior when a dog wants to play
[0293] 1. The "terminal" captures the dog's movements with a camera and collects the dog's tail wagging vigorously and high-pitched barking.
[0294] 2. The "terminal" sends the collected data to the "server."
[0295] 3. The "server" receives the data and identifies tail movements from the video data and high-pitched cries from the audio data.
[0296] 4. The “server” inputs these features into the “AI model.”
[0297] 5. The AI model uses these characteristics to infer that the dog wants to play.
[0298] 6. The "server" translates the result "I want to play" into natural language and makes it "Play with the ball!"
[0299] 7. The "terminal" notifies the user with the message "Play with the ball!"
[0300] 8. The "terminal" also collects the user's facial expressions and voice and sends them to the "server."
[0301] 9. The "Emotion Engine" recognizes the user's emotion as "Relaxed."
[0302] 10. The "server" checks the user's relaxation state and sends a message suggesting "Relax with your pet."
[0303] 11. The Terminal displays the suggestion message to the user.
[0304] Example 2: System behavior when the cat is hungry
[0305] 1. The "terminal" collects the cat's meowing (a low, sustained sound) and its movements around the dish.
[0306] 2. The "terminal" sends the collected data to the "server."
[0307] 3. The "server" receives the data and identifies the location from the video data (walking around the dish) and the sound from the audio data (a low, sustained sound).
[0308] 4. The “server” inputs these features into the “AI model.”
[0309] 5. The "AI model" uses these characteristics to infer that the cat is hungry.
[0310] 6. The "server" translates the result "hungry" into natural language and sets it as "I'm hungry."
[0311] 7. The "terminal" notifies the user with the message "I'm hungry."
[0312] 8. The "terminal" collects the user's facial expressions and voice and sends them to the "server."
[0313] 9. The "emotion engine" recognizes the user's emotion as "tired."
[0314] 10. The "server" checks the user's tiredness and sends a message suggesting that they take their time to feed their pet.
[0315] 11. The Terminal displays the suggestion message to the user.
[0316] In this way, the present invention can be implemented as a system that allows the user to understand the specific emotions and desires of their pet in real time and take appropriate action according to their own emotions and state.
[0317] The processing flow will be explained below.
[0318] Step 1:
[0319] The device uses a camera and microphone to collect your pet's movements, facial expressions, and sounds in real time, and the collected data is divided into video data and audio data, each of which is then split into separate data packets.
[0320] Step 2:
[0321] The device sends the collected data packets to the server, using a security protocol to ensure the data is protected.
[0322] Step 3:
[0323] The server receives the data packets sent from the terminal. The received data is first decoded and then separated into video data and audio data for processing.
[0324] Step 4:
[0325] The server runs the video data through computer vision algorithms to analyze your pet's facial expressions and movements, specifically identifying tail movements, facial expressions, and overall body movements.
[0326] Step 5:
[0327] The server runs the audio data through a speech recognition algorithm to identify the pattern, pitch, and frequency of your pet's cries, including the pitch and repetition of the cries.
[0328] Step 6:
[0329] The server combines the results of the video and audio analysis to extract features to estimate the pet's emotions and desires, such as high-pitched meows and tail wagging.
[0330] Step 7:
[0331] The server inputs the extracted features into a pre-trained AI model, which uses patterns learned from past data to infer the pet's emotions and desires.
[0332] Step 8:
[0333] The server receives the inferences output by the AI model and maps them to a list of predefined emotions and desires, such as "hunger," "want to play," "anxiety," and "relaxation."
[0334] Step 9:
[0335] The server translates the estimated emotions and desires into natural language, for example, translating "hungry" into "I'm hungry" and "I want to play" into "Play with my ball!"
[0336] Step 10:
[0337] The server then sends the translated message to the user's device, again using security protocols to protect the data.
[0338] Step 11:
[0339] The terminal receives the translated message sent from the server, and the received message is formatted to be displayed on the user interface.
[0340] Step 12:
[0341] The device will display messages to the user that indicate the pet's emotions and needs, such as "I'm hungry!" or "Play with my ball!", displayed as pop-ups or notification banners.
[0342] Step 13:
[0343] The device also uses a camera and microphone to collect the user's voice and facial expressions, which are then split into data packets and sent to the server.
[0344] Step 14:
[0345] The server receives the user's data packets sent from the terminal, decodes the received data, and prepares them for processing by separating them into video data and audio data.
[0346] Step 15:
[0347] The server runs the video data through a facial expression analysis algorithm to analyze the user's facial expressions, and the audio data through a voice recognition algorithm to analyze the user's tone and words.
[0348] Step 16:
[0349] The server integrates the analysis results of the facial expression data and voice data and extracts features to identify the user's emotions.
[0350] Step 17:
[0351] The server inputs the extracted features into an emotion engine to estimate the user's emotion, such as "tired," "relaxed," or "excited."
[0352] Step 18:
[0353] The server generates a message suggesting how to handle pets based on the user's emotions. For example, if the user is tired, it suggests, "Take your time and don't rush, take your time to feed your pet."
[0354] Step 19:
[0355] The server sends a proposal message to the user's device, which is formatted for display through the user interface.
[0356] Step 20:
[0357] The device will notify the user with suggestion messages, such as "Take your time to feed your pet and don't rush," via a pop-up or notification banner.
[0358] In this way, the present invention can be implemented as a system that allows the user to understand the specific emotions and desires of their pet in real time and take appropriate action according to their own emotions and state.
[0359] Example 2
[0360] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0361] It is difficult to understand the emotions and desires of both pets and users in real time and propose appropriate responses. Conventional technologies lack systems that can accurately estimate a pet's emotions and desires, translate them into natural language, and notify the user. Furthermore, there is no system that can recognize the user's emotions and propose appropriate responses based on them. This leads to a decline in the quality of pet care and a lack of communication between pets and users.
[0362] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[0363] In this invention, the server includes means for receiving data transmitted from the terminal, decoding and classifying the data, means for analyzing video data using a computer vision algorithm, and means for analyzing audio data using a speech recognition algorithm. This allows features of the pet's movements and cries to be extracted and input into an AI model to estimate the pet's emotions and needs. Furthermore, the estimation results are translated into natural language and notified to the user, allowing for an accurate understanding of the pet's situation. Furthermore, the server collects and analyzes the user's voice and facial expressions, and proposes appropriate countermeasures based on the user's emotions, thereby improving communication between the pet and the user.
[0364] The "terminal" is a device equipped with a camera and microphone to collect your pet's movements, facial expressions, and sounds.
[0365] A "server" is a computer system that receives data sent from a terminal and decodes, classifies, and analyzes the data.
[0366] A "data packet" is a unit for dividing and transmitting data collected by a terminal at regular intervals.
[0367] A "security protocol" is a communication protocol used to ensure security when transmitting data.
[0368] "Video data" refers to image data captured by a camera showing the movements and expressions of a pet.
[0369] "Audio data" is sound data recorded with a microphone of a pet's cries.
[0370] A "computer vision algorithm" is an image processing technology used to analyze specific movements and facial expressions from video data.
[0371] A "voice recognition algorithm" is a technology for analyzing specific sound patterns and characteristics from audio data.
[0372] "Features" are important data points extracted from the analyzed data that indicate the pet's emotions and behavior.
[0373] An "AI model" is a computational model that uses machine learning algorithms to estimate a pet's emotions and desires.
[0374] "Natural language translation" is the process of converting the results estimated by an AI model into a language that a user can understand.
[0375] "Notification" refers to the act of conveying a translated message to the user, and is primarily done via the terminal.
[0376] "User data" is data that collects information such as the user's voice and facial expressions.
[0377] An "emotion engine" is a system for analyzing user data to identify a user's emotions.
[0378] "Countermeasure suggestion" is a process of suggesting actions that the user should take toward their pet based on the identified user emotions.
[0379] The present invention is a system that analyzes the movements, facial expressions, and cries of a pet, estimates the feelings and desires of the pet, and notifies the user of the results. Detailed embodiments of the system will be described below.
[0380] System configuration
[0381] The system includes a "terminal" that collects the pet's movements, facial expressions, and cries, a "server" that receives the data sent from the terminal, an "AI model" that analyzes the data and infers the pet's emotions and desires, and a "notification system" that notifies the user. It also includes an "emotion engine" that recognizes the user's emotions and suggests appropriate countermeasures.
[0382] Data collection
[0383] The device is equipped with a camera and microphone to collect pet movements, facial expressions, and sounds in real time. This data is divided into data packets at regular intervals (e.g., every second) and sent to a server using a security protocol (e.g., SSL / TLS).
[0384] Data reception and decryption
[0385] The server receives the data packets sent from the terminal, decrypts them according to the security protocol, and then classifies the received data into video data and audio data.
[0386] Data analysis
[0387] The server analyzes the video data using a computer vision algorithm (e.g., OpenCV) to identify the pet's movements and facial expressions (e.g., tail wagging and ear movement). It also analyzes the audio data using a speech recognition algorithm (e.g., LibriSpeech) to identify the pattern, pitch, and frequency of the pet's cries. This allows the extraction of features.
[0388] Estimating emotions and desires
[0389] The server inputs the extracted features into an AI model (e.g., TensorFlow or PyTorch). The AI model is pre-trained with data on pets' emotions and desires, and uses the input features to infer the pet's emotions and desires. For example, if the pet wags its tail and makes a high-pitched bark, it can infer that the pet wants to play.
[0390] Natural language translation and notification
[0391] The server translates the results output by the AI model into natural language. For example, the predicted result "I want to play" is converted into the message "Play with the ball!" The translated message is resent from the server to the device and notified to the user via the device.
[0392] Recognizing user emotions and proposing countermeasures
[0393] The device collects the user's voice and facial expressions using a camera and microphone. The collected user data is sent to a server where it is analyzed by an emotion engine. Using a voice recognition algorithm and an expression analysis algorithm, emotions are identified from the user's tone of voice and facial expressions. Based on this emotional information, the server suggests appropriate countermeasures. For example, if the user is tired, a message suggesting, "Let's create a relaxing atmosphere" is sent.
[0394] Specific examples
[0395] Example 1: System behavior when a dog wants to play
[0396] 1. The device captures the dog's movements with a camera and collects the dog's vigorously wagging tail and high-pitched barking sounds.
[0397] 2. The device sends the collected data to the server.
[0398] 3. The server receives the data and identifies tail movements from the video data and high-pitched cries from the audio data.
[0399] 4. The server inputs these features into the AI model.
[0400] 5. The AI model uses these characteristics to infer that the dog wants to play.
[0401] 6. The server translates the result "I want to play" into natural language and makes it "Play with the ball!"
[0402] 7. The device notifies the user with the message "Play with the ball!"
[0403] 8. The device collects the user's facial expressions and voice and sends them to the server.
[0404] 9. The emotion engine recognizes the user's emotion as "relaxed."
[0405] 10. The server checks the user's relaxation state and sends a message suggesting "Relax with your pet."
[0406] 11. The terminal displays the suggestion message to the user.
[0407] Example 2: System behavior when the cat is hungry
[0408] 1. The device collects the cat's meowing (a low, sustained sound) and its movements around the dish.
[0409] 2. The device sends the collected data to the server.
[0410] 3. The server receives the data and identifies the location from the video data (walking around the dish) and the sound from the audio data (a low, sustained sound).
[0411] 4. The server inputs these features into the AI model.
[0412] 5. The AI model uses these characteristics to infer that the cat is hungry.
[0413] 6. The server translates the result "hungry" into natural language and gives it the value "I'm hungry."
[0414] 7. The device notifies the user with the message "I'm hungry."
[0415] 8. The device collects the user's facial expressions and voice and sends them to the server.
[0416] 9. The emotion engine recognizes the user's emotion as "tired."
[0417] 10. The server checks the user's tiredness and sends a message suggesting that they take their time to feed their pet.
[0418] 11. The terminal displays the suggestion message to the user.
[0419] Example prompts for generative AI models
[0420] "We want you to infer what the dog wants from the dog's behavior and bark data collected by the device."
[0421] "I want you to analyze what the cat wants at the moment based on its meows and behavioral patterns."
[0422] Thus, the present invention provides a system that can analyze the emotions and desires of pets and users in real time and encourage users to respond appropriately, thereby improving the quality of pet care and facilitating communication between pets and users.
[0423] The flow of the identification process in the second embodiment will be described with reference to FIG.
[0424] Step 1:
[0425] Your pet's movements are collected on the device using a camera and microphone.
[0426] Input: Video and audio data captured through a camera and microphone.
[0427] Specific behavior: The device records your pet's movements, facial expressions, and sounds in real time, such as capturing a dog's tail wagging or a cat's meowing.
[0428] Output: Collected video and audio data.
[0429] Step 2:
[0430] Data collected by the terminal is divided into data packets and sent to a server using a security protocol.
[0431] Input: Collected video and audio data.
[0432] Specific operation: Data is packetized every second and transmitted using security protocols such as SSL / TLS.
[0433] Output: Data packets protected by security protocols.
[0434] Step 3:
[0435] The server receives and decodes the data packets sent from the terminal.
[0436] Input: Data packets protected by security protocols.
[0437] Specific operation: The data packets are decrypted using the SSL / TLS protocol and restored to the original video and audio data.
[0438] Output: Decoded video and audio data.
[0439] Step 4:
[0440] The server classifies the decoded data into video data and audio data.
[0441] Input: Decoded video and audio data.
[0442] Specific operation: Data is classified into video and audio categories based on data format and metadata.
[0443] Output: Classified video and audio data.
[0444] Step 5:
[0445] The server analyzes the video data using computer vision algorithms.
[0446] Input: Classified video data.
[0447] Specific operation: Using computer vision algorithms such as OpenCV, the system analyzes pet movements and facial expressions. For example, it detects the wagging of a dog's tail or the movement of a cat's ears.
[0448] Output: Features extracted from video data.
[0449] Step 6:
[0450] The server analyzes the voice data using a voice recognition algorithm.
[0451] Input: Classified audio data.
[0452] What it does: Uses speech recognition algorithms such as LibriSpeech to identify the pattern, pitch, and frequency of your pet's cries.
[0453] Output: Features extracted from the audio data.
[0454] Step 7:
[0455] The server inputs the features extracted from the video and audio data into an AI model to estimate the pet's emotions and desires.
[0456] Input: Features extracted from video and audio data.
[0457] Specific behavior: Using machine learning algorithms such as TensorFlow and PyTorch, the system infers a pet's emotions and desires from its behavior and sounds. For example, if a pet wags its tail and makes a high-pitched meow, it infers that it wants to play.
[0458] Output: Estimated pet's emotions and desires.
[0459] Step 8:
[0460] The server translates the estimated emotions and desires into natural language and notifies the user.
[0461] Input: Estimated pet emotions and desires.
[0462] Specific operation: The inference results are converted into messages such as "Play with the ball!" or "I'm hungry" and sent to the user's device.
[0463] Output: Notification message translated into natural language.
[0464] Step 9:
[0465] The device collects the user's voice and facial expressions and sends them to the server.
[0466] Input: User's audio and video data.
[0467] Specific operation: The camera and microphone are used to record the user's real-time voice and facial expressions, which are then divided into data packets and sent to the server.
[0468] Output: Collected user audio and video data.
[0469] Step 10:
[0470] The server analyzes the collected user data using voice recognition algorithms and facial expression analysis algorithms to identify the user's emotions.
[0471] Input: Collected user audio and video data.
[0472] Specific behavior: Identify the user's emotions by analyzing voice tone and facial expressions. For example, if the user's voice tone is low and they sound tired, it will recognize them as "tired."
[0473] Output: Identified user sentiment.
[0474] Step 11:
[0475] The server suggests measures to take for the pet based on the user's feelings.
[0476] Input: Identified user sentiment.
[0477] Specific behavior: Generate and send messages to the user's device suggesting things like "Create a relaxing atmosphere" and "Take your time to feed your pet without rushing."
[0478] Output: A message with the proposed workaround.
[0479] (Application example 2)
[0480] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0481] Currently, there are several systems on the market that monitor pets' movements and cries to estimate their emotions, but these systems cannot provide clear and specific suggestions to users. Furthermore, they lack a mechanism for providing appropriate content based on pets' emotions and desires. Therefore, there is a need for a system that can help users choose the best course of action for their pets.
[0482] The identification process by the identification processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes a terminal for collecting the movements, expressions, and cries of the pet, a means for receiving data transmitted from the terminal, a means including an AI model for analyzing the received data and estimating the pet's emotions and desires, a means for translating the estimated emotions and desires into natural language and notifying the user, and a means for selecting appropriate content based on the estimated emotions and desires of the pet and delivering it to the user terminal. This allows the user to understand the pet's emotions and desires in real time and obtain optimal content in a timely manner.
[0483] "Pet movements, expressions, and sounds" refers to the physical movements, facial expressions, and sounds made by pets.
[0484] The "terminal" is a device equipped with a camera and microphone that collects your pet's movements, facial expressions, and cries.
[0485] A "server" is a computer system that receives, analyzes, and processes data sent from a terminal.
[0486] An "AI model" is an artificial intelligence that uses machine learning algorithms to analyze data and estimate a pet's emotions and desires.
[0487] "Estimation" refers to predicting a pet's emotions and desires from the data obtained.
[0488] "Translation into natural language" means converting the inferred results output by the AI model into a message that is easy for the user to understand.
[0489] "Notifying the user" means conveying messages or information to the user through the terminal.
[0490] A "security protocol" is a communication protocol for maintaining the security of data when it is transmitted from a terminal to a server.
[0491] A "computer vision algorithm" is an algorithm that analyzes video data and identifies pet movements and facial expressions.
[0492] A "voice recognition algorithm" is an algorithm that analyzes voice data and identifies the sounds of pets.
[0493] "Selecting appropriate content" means selecting videos, music, information, etc. that are suitable for your pet based on the pet's estimated emotions and desires.
[0494] "Delivery" means transmitting selected content to a user terminal and having it played or displayed.
[0495] System configuration overview
[0496] This invention provides a system that collects pet movements, facial expressions, and cries, estimates the pet's emotions and desires based on the collected data, and delivers appropriate content to the user. The system mainly consists of the following components:
[0497] A "terminal" for collecting your pet's movements, facial expressions, and cries.
[0498] The "server" receives and analyzes data collected from the terminal.
[0499] An "AI model" that analyzes the received data and estimates the pet's emotions and desires.
[0500] A "notification system" that selects appropriate content based on estimated emotions and desires and delivers it to the user's device.
[0501] Data collection details
[0502] The "terminal" is equipped with a camera and microphone to collect the pet's movements, facial expressions, and sounds in real time, and the collected data is divided into data packets at regular intervals and sent to the "server" using a security protocol.
[0503] Data reception and analysis
[0504] The server receives and decodes the data packets sent from the device. It then processes the video and audio data separately. The video data uses computer vision algorithms to identify the pet's facial expressions and movements, while the audio data uses speech recognition algorithms to identify the pattern, pitch, and frequency of the pet's cries.
[0505] Estimating emotions and desires
[0506] The server inputs the analyzed features into an AI model to estimate the pet's emotions and desires. The AI model uses a trained machine learning algorithm, and can infer, for example, that a pet's desire to play is indicated by a wagging tail and high-pitched meow.
[0507] Natural language translation and notification
[0508] The estimated emotions and desires are translated into natural language and converted into specific messages (e.g., "I'm hungry" or "Play with a ball!"). The translated messages are then sent from the server to the device and notified to the user.
[0509] Recognizing user emotions and proposing countermeasures
[0510] The device also collects the user's voice and facial expressions and sends them to the server. The emotion engine analyzes this data and recognizes the user's emotions. For example, it uses voice recognition and facial expression analysis algorithms to identify emotions from the user's tone and facial expressions.
[0511] Proposing solutions based on user sentiment
[0512] The server generates a message suggesting appropriate actions for the pet based on the user's emotional information. For example, if the user is tired, the server might suggest, "Take your time to feed your pet and don't rush." The suggested message is displayed on the user's device, encouraging the user to take appropriate action.
[0513] Hardware and software used
[0514] The system uses the following hardware and software:
[0515] Hardware: Camera, microphone, smartphone, smart glasses, head-mounted display.
[0516] Software: OpenCV (used to collect video data from the camera), TensorFlow (used to analyze pet and user emotions), requests (used to communicate data with the server), virtual voice analysis module and emotion recognition module.
[0517] Specific example explanation
[0518] For example, if it is estimated that the pet wants to play, it will be treated as "playful" and the video content "play_video.mp4" will be selected and played. If the user is relaxed, the message "Enjoy a relaxing session with your pet" will be displayed, suggesting a better way to spend time with the pet.
[0519] Prompt Sentence Examples
[0520] "Suggest content suitable for pets who want to play."
[0521] "Suggest actions that are appropriate for when the user is relaxed."
[0522] In this way, a system is realized that can analyze emotions and desires based on real-time data from pets and users, and deliver optimal responses and content.
[0523] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[0524] Step 1:
[0525] The device uses a camera and microphone to collect the pet's movements, facial expressions, and sounds in real time. It receives camera video and audio data as input and divides it into data packets at regular intervals. This captures multiple features, such as the pet's movements, facial expressions, and sounds, which can then be sent to subsequent analysis steps.
[0526] Step 2:
[0527] The data packets collected by the terminal are sent to the server using a security protocol. The data packets are received as input, and the data is encrypted using a security protocol (e.g., SSL / TLS) and securely sent to the server. This ensures the confidentiality of the data.
[0528] Step 3:
[0529] The server receives data packets sent from the terminal. It receives encrypted data packets as input and first decrypts them. As output, it obtains decrypted video and audio data. This allows the server to store data in an analyzable format.
[0530] Step 4:
[0531] The server analyzes the video data to identify the pet's movements and facial expressions. It receives the decoded video data as input and analyzes it using a computer vision algorithm (e.g., OpenCV). The output is identified behavior and facial expression data, such as tail movements and facial expressions. This allows the pet's movement and facial expression features to be extracted.
[0532] Step 5:
[0533] The server analyzes the audio data to identify the pattern, pitch, and frequency of the pet's cries. It receives the decoded audio data as input and analyzes it using a speech recognition algorithm (e.g., a speech analysis module). The output is an analysis result that identifies the pattern, pitch, and frequency of the pet's cries. This allows the characteristics of the pet's cries to be extracted.
[0534] Step 6:
[0535] The server inputs the analyzed features into an AI model to estimate the pet's emotions and desires. The analysis results of movements, facial expressions, and cries are received as input and input into an AI model (for example, a model built with TensorFlow). The output is the pet's estimated emotions and desires (for example, "I want to play" or "I'm hungry"). This clarifies the pet's internal state.
[0536] Step 7:
[0537] The server translates the estimated emotions and desires into natural language and generates a message. It receives the estimated results from the AI model as input and translates them using a natural language processing engine. The output is a message that is easy for the user to understand, such as "I'm hungry" or "Play with my ball!" This communicates the pet's emotions and desires to the user.
[0538] Step 8:
[0539] The server selects appropriate content based on the estimated emotions and desires. It receives a message translated into natural language as input and uses an algorithm to select appropriate content (e.g., video, music). Specific content files (e.g., "play_video.mp4", "relaxing_music.mp3") are obtained as output. This allows content that matches the pet's state to be selected.
[0540] Step 9:
[0541] The server delivers the selected content to the user terminal. The server receives the selected content file as input and transmits it to the user terminal via the Internet. As output, the server delivers the content to be played or displayed on the user terminal. This allows the user to receive content that is appropriate for the condition of their pet.
[0542] Step 10:
[0543] The device collects the user's facial expressions and voice and sends them to the server. The device receives camera footage and voice data as input, divides it into data packets, and sends them to the server. This provides data for analyzing the user's emotions.
[0544] Step 11:
[0545] The server analyzes the user's facial expressions and voice to recognize the user's emotions. It receives the user's facial expressions and voice data as input and analyzes them using an emotion recognition algorithm. The output identifies the user's emotions (e.g., "relaxed" or "tired"). This clarifies the user's state.
[0546] Step 12:
[0547] The server generates a message suggesting how to respond to a pet based on the user's emotional information. It receives the user's emotion recognition results as input and generates the suggested message using a natural language processing engine. The output is specific suggested actions such as "Relax with your pet" or "Take your time to feed your pet without rushing." This provides guidance for the user to take appropriate action.
[0548] The above is a description of the specific processing steps of the system for implementing the present invention.
[0549] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0550] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0551] In the above embodiment, an example in which the specific process is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific process may be performed by the smart device 14.
[0552] [Second embodiment]
[0553] FIG. 3 shows an example of the configuration of a data processing system 210 according to the second embodiment.
[0554] 3, the data processing system 210 includes the data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.
[0555] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0556] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, and the camera 42 are also connected to the bus 52.
[0557] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[0558] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[0559] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[0560] Fig. 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Fig. 4, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[0561] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0562] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0563] In the smart glasses 214, the reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0564] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal."
[0565] System configuration overview
[0566] The present invention is a system that includes a "terminal" for collecting pet movements, facial expressions, and cries, a "server" for receiving data sent from the terminal, an "AI model" for analyzing the received data and inferring the pet's emotions and desires, and a "notification system" that translates the inferred emotions and desires into natural language and notifies the user.
[0567] Data collection details
[0568] The "terminal" is equipped with a camera and microphone, and collects the pet's movements, facial expressions, and sounds in real time. The collected data is divided into data packets at regular intervals and sent from the "terminal" to the "server." A security protocol is used to ensure the data is secure.
[0569] Data reception and analysis
[0570] The "server" receives the data packets sent from the terminal. The received data is first decoded and then processed separately into video data and audio data.
[0571] The "server" analyzes the video data using computer vision algorithms to identify the pet's facial expressions and movements (for example, tail movements). It also analyzes the audio data using speech recognition algorithms to identify the pattern, pitch, and frequency of the pet's cries. This allows the data's features to be extracted.
[0572] Estimating emotions and desires
[0573] The extracted features are input into an "AI model," which uses a pre-trained machine learning algorithm to infer the pet's emotions and desires. For example, if a pet frequently wags its tail and makes a high-pitched meow, it can be inferred to have a desire to play.
[0574] Natural language translation and notification
[0575] The estimated emotions and desires are translated into natural language. Specifically, the output of the AI model is converted into messages such as "I'm hungry" or "Play with the ball!". This translated message is then resent from the "server" to the "device," which then notifies the user.
[0576] Specific example explanation
[0577] Example 1: System behavior when a dog wants to play
[0578] 1. The "terminal" captures the dog's movements with a camera and collects the dog's tail wagging vigorously and high-pitched barking.
[0579] 2. The "terminal" sends the collected data to the "server."
[0580] 3. The "server" receives the data and identifies tail movements from the video data and high-pitched cries from the audio data.
[0581] 4. The “server” inputs these features into the “AI model.”
[0582] 5. The AI model uses these characteristics to infer that the dog wants to play.
[0583] 6. The "server" translates the result "I want to play" into natural language and makes it "Play with the ball!"
[0584] 7. The "terminal" notifies the user with the message "Play with the ball!"
[0585] Example 2: System behavior when the cat is hungry
[0586] 1. The "terminal" collects the cat's meowing (a low, sustained sound) and its movements around the dish.
[0587] 2. The "terminal" sends the collected data to the "server."
[0588] 3. The "server" receives the data and identifies the location from the video data (walking around the dish) and the sound from the audio data (a low, sustained sound).
[0589] 4. The “server” inputs these features into the “AI model.”
[0590] 5. The "AI model" uses these characteristics to infer that the cat is hungry.
[0591] 6. The "server" translates the result "hungry" into natural language and sets it as "I'm hungry."
[0592] 7. The "terminal" notifies the user with the message "I'm hungry."
[0593] In this way, the present invention can be implemented as a system that allows users to understand their pet's specific emotions and desires in real time and take appropriate action.
[0594] The processing flow will be explained below.
[0595] Step 1:
[0596] The device uses a camera and microphone to collect the pet's movements, facial expressions, and sounds in real time. The collected data is then separated into video data and audio data, and each is split into separate data packets.
[0597] Step 2:
[0598] The device sends the collected data packets to a server, where a security protocol is used to ensure the data is protected.
[0599] Step 3:
[0600] The server receives the data packets sent from the terminal. The received data is first decoded and then separated into video data and audio data for processing.
[0601] Step 4:
[0602] The server runs the video data through computer vision algorithms to analyze your pet's facial expressions and movements, specifically identifying tail movements, facial expressions, and overall body movements.
[0603] Step 5:
[0604] The server runs the audio data through a speech recognition algorithm to identify the pattern, pitch, and frequency of your pet's cries, including the pitch of the cries and the repetition of the cries.
[0605] Step 6:
[0606] The server combines the results of the video and audio analysis to extract features to estimate the pet's emotions and desires, such as high-pitched meows and tail wagging.
[0607] Step 7:
[0608] The server inputs the extracted features into a pre-trained AI model, which uses patterns learned from past data to infer the pet's emotions and desires.
[0609] Step 8:
[0610] The server receives the inferences output by the AI model and maps them to a list of predefined emotions and desires, such as "hunger," "want to play," "anxiety," and "relaxation."
[0611] Step 9:
[0612] The server translates the estimated emotions and desires into natural language. For example, "hungry" becomes "I'm hungry" and "I want to play" becomes "Play with my ball!"
[0613] Step 10:
[0614] The server then sends the translated message to the user's device, again using security protocols to protect the data.
[0615] Step 11:
[0616] The terminal receives the translated message sent from the server, and the received message is formatted to be displayed on the user interface.
[0617] Step 12:
[0618] The device will display messages to the user that indicate the pet's emotions and needs, such as "I'm hungry!" or "Play with my ball!", displayed as pop-ups or notification banners.
[0619] Step 13:
[0620] The user checks the messages displayed on the device and takes action according to the pet's request, such as feeding the pet or playing with a ball.
[0621] This series of processes allows users to understand their pet's specific emotions and desires in real time and respond appropriately.
[0622] Example 1
[0623] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0624] Conventional pet management systems have had problems with accurately monitoring pet behavior and conditions, and owners are unable to properly understand their pets' emotions and needs, which increases stress for pets and burdens on owners. Furthermore, insufficient coordination in data collection, analysis, and notification makes it difficult to understand pets' conditions in real time.
[0625] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0626] In this invention, the server includes means for decoding received data and processing the data separately into video data and audio data, means for analyzing the video data using a computer vision algorithm and the audio data using a speech recognition algorithm, and means for inputting the extracted features into an AI model to estimate emotions and desires. This makes it possible to instantly estimate emotions and desires from the pet's movements and cries, and notify the owner in natural language.
[0627] The "terminal" is a device equipped with a camera and microphone to collect the movements, facial expressions, and sounds of your pet.
[0628] A "server" is a computer system that receives, decrypts, and analyzes data sent from a terminal.
[0629] A "data packet" is a unit of information that divides collected data and is sent from a terminal to a server.
[0630] "Decoding" is the process of restoring encoded data to its original form.
[0631] "Video data" refers to data that visually records the movements and expressions of pets collected by a camera.
[0632] "Audio data" refers to data that records the sounds of pets as collected by a microphone.
[0633] A "computer vision algorithm" is a technology for analyzing video data and recognizing objects and actions.
[0634] A "voice recognition algorithm" is a technology for analyzing voice data and recognizing sounds and patterns.
[0635] "Features" are important information extracted from video and audio data for estimating a pet's emotions and desires.
[0636] The "AI model" is a system that is pre-trained using a machine learning algorithm and estimates a pet's emotions and desires based on input features.
[0637] "Natural language" refers to language that humans use on a daily basis, and is a language that expresses estimated emotions and desires in a way that is easy for users to understand.
[0638] A "security protocol" is a set of rules and procedures for securely communicating data.
[0639] "Notification" is the act of informing the user of the results of estimated emotions and desires.
[0640] MODE FOR CARRYING OUT THE INVENTION
[0641] The present invention is a system that includes a terminal that collects pet movements, facial expressions, and cries, a server that receives and analyzes the data, and a system that notifies the user of the estimated emotions and desires. Specifically, by using a pet monitoring device, a data analysis system, and a communication means, it is possible to grasp the pet's condition in real time and take appropriate measures.
[0642] System configuration overview
[0643] The present invention includes a "terminal" for collecting pet movements, facial expressions, and cries, a "server" for receiving data sent from the terminal, an "AI model" for analyzing the received data and inferring the pet's emotions and desires, and a "notification system" for translating the inferred emotions and desires into natural language and notifying the user.
[0644] Data collection details
[0645] The "terminal" is equipped with a camera and microphone, and collects the pet's movements, facial expressions, and cries in real time. For example, the camera records the pet's tail movements and eye expressions as video data, while the microphone collects the pet's cries as audio data. The collected data is divided into data packets at regular intervals and sent from the "terminal" to the "server." A security protocol is used for data transmission to ensure the safety of the data.
[0646] Data reception and analysis
[0647] The "server" receives data packets sent from the device. The received data is first decoded and then processed separately into video data and audio data. The "server" analyzes the video data using a computer vision algorithm (e.g., OpenCV) to identify the pet's facial expressions and movements. It also analyzes the audio data using a speech recognition algorithm (e.g., a TensorFlow or PyTorch-based model) to identify the pattern, pitch, and frequency of the pet's cries. This allows the data features to be extracted.
[0648] Estimating emotions and desires
[0649] The extracted features are input into an "AI model." The AI model uses a pre-trained machine learning algorithm (e.g., a generative AI model) to infer the pet's emotions and desires. For example, if the pet frequently wags its tail and makes a high-pitched meow, it can be inferred that the pet wants to play.
[0650] Natural language translation and notification
[0651] The estimated emotions and desires are translated into natural language. Specifically, the results output by the AI model are converted into messages such as "I'm hungry" or "Play with the ball!". These translated messages are then resent from the "server" to the "device," which then notifies the user. For example, the user can be notified by a push notification or voice alert on a smartphone app.
[0652] Specific example explanation
[0653] Example 1: System behavior when a dog wants to play
[0654] 1. The device captures the dog's movements with a camera and collects the dog's vigorously wagging tail and high-pitched barking sounds.
[0655] 2. The device sends the collected data to the server.
[0656] 3. The server receives the data and identifies tail movements from the video data and high-pitched cries from the audio data.
[0657] 4. The server inputs these features into the AI model.
[0658] 5. The AI model uses these characteristics to infer that the dog wants to play.
[0659] 6. The server translates the result "I want to play" into natural language and makes it "Play with the ball!"
[0660] 7. The device notifies the user with the message "Play with the ball!"
[0661] Example 2: System behavior when the cat is hungry
[0662] 1. The device collects the cat's meowing (a low, sustained sound) and its movements around the dish.
[0663] 2. The device sends the collected data to the server.
[0664] 3. The server receives the data and identifies the location from the video data (walking around the dish) and the sound from the audio data (a low, sustained sound).
[0665] 4. The server inputs these features into the AI model.
[0666] 5. The AI model uses these characteristics to infer that the cat is hungry.
[0667] 6. The server translates the result "hungry" into natural language and gives it the value "I'm hungry."
[0668] 7. The device notifies the user with the message "I'm hungry."
[0669] Example prompts to input to the generative AI model
[0670] 1. Example prompt for identifying a dog's desire to play: "When a dog wags its tail frequently and makes a high-pitched whine, it means it wants to play. Express this in natural language."
[0671] 2. Example prompt for identifying a cat's hunger state: "When a cat meows in a low, sustained tone and paces around its food bowl, it is likely hungry. Please describe this in natural language."
[0672] In this way, the present invention can be implemented as a system that allows users to understand their pet's specific emotions and desires in real time and take appropriate action.
[0673] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0674] Program processing steps
[0675] Step 1: Collect data from the device
[0676] The device uses a camera and microphone to collect your pet's movements, facial expressions, and cries in real time.
[0677] Input: Real-time pet movements, facial expressions, and sounds
[0678] Output: Collected video and audio data
[0679] Specific behavior:
[0680] The camera records your pet's tail movements and facial expressions as video data, while the microphone captures your pet's barks as audio data, such as when your dog wags its tail or barks in a high-pitched voice.
[0681] Step 2: Send data from the device to the server
[0682] The device divides the collected data into data packets and sends them to the server using a security protocol.
[0683] Input: Collected video and audio data
[0684] Output: Data packets sent from the device to the server
[0685] Specific behavior:
[0686] The collected data is divided into packets of a certain size, encrypted using the TLS protocol, and sent to the server.
[0687] Step 3: Server receives and decrypts data
[0688] The server receives the data packets sent from the terminal, decodes them, and separates them into video data and audio data.
[0689] Input: Data packets sent from the device
[0690] Output: Decoded video and audio data
[0691] Specific behavior:
[0692] The received data packets are decrypted using the TLS protocol and separated into collected video and audio data, for example, tail movements recorded by a camera or barks recorded by a microphone.
[0693] Step 4: Video data analysis by the server
[0694] The server analyzes the video data using computer vision algorithms to identify the pet's facial expressions and movements.
[0695] Input: Decoded video data
[0696] Output: Features of pet's facial expressions and movements
[0697] Specific behavior:
[0698] Using OpenCV, features such as tail wagging and eye movement are extracted from video data. For example, if a cat is wagging its tail vigorously, that movement can be identified.
[0699] Step 5: Server analyzes the audio data
[0700] The server analyzes the audio data using a voice recognition algorithm to identify the pattern, pitch and frequency of your pet's cries.
[0701] Input: Decoded audio data
[0702] Output: Features of pet sounds
[0703] Specific behavior:
[0704] Use TensorFlow or PyTorch to identify high-pitched meows and patterns in audio data, such as detecting the low, sustained meow of a cat.
[0705] Step 6: Feature extraction by the server
[0706] The server integrates the features extracted from the video and audio data and inputs them into the AI model.
[0707] Input: Features of pet's facial expressions and movements, features of pet's cries
[0708] Output: Integrated features
[0709] Specific behavior:
[0710] The identified tail movements from the video data and the call patterns from the audio data are integrated into a single dataset.
[0711] Step 7: Server inputs AI model
[0712] The server inputs the extracted features into an AI model to estimate emotions and desires.
[0713] Input: Integrated features
[0714] Output: Estimated pet's emotions and desires
[0715] Specific behavior:
[0716] The integrated features are input into a generative AI model to obtain an inference. For example, if a dog is wagging its tail and making a high-pitched bark, it may infer that it wants to play.
[0717] Step 8: Server-based translation into natural language
[0718] The server translates the estimated emotions and desires into natural language.
[0719] Input: Inference results from the AI model
[0720] Output: A message expressed in natural language
[0721] Specific behavior:
[0722] The result "I want to play" obtained from the AI model is converted into the message "Play with the ball!"
[0723] Step 9: Sending translation results from the server to the device
[0724] The server retransmits the translated message to the terminal.
[0725] Input: A message expressed in natural language
[0726] Output: Messages sent to the terminal
[0727] Specific behavior:
[0728] Using security protocols, send a translated message to the device, for example, "Play with the ball!"
[0729] Step 10: Notify the user via the device
[0730] The terminal receives the message sent from the server and notifies the user of the message.
[0731] Input: The message sent from the server
[0732] Output: The notification the user receives
[0733] Specific behavior:
[0734] The device uses push notifications and audio alerts on a smartphone app to notify users with messages like, "Play with the ball!"
[0735] Step 11: User Action
[0736] The user checks the notification from the device and responds to the pet.
[0737] Input: Notifications from your device
[0738] Output: Specific responses to pets
[0739] Specific behavior:
[0740] The user checks the notification and takes action based on the message, such as playing with their pet. For example, if they receive the message "Play with a ball!", they will get a ball and start playing with their dog.
[0741] The above is a description of the specific processing steps in the system program.
[0742] (Application example 1)
[0743] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0744] In recent years, there has been a demand for technology that can understand pet emotions and desires and enable owners to provide appropriate care. In particular, there is a need for a system that can grasp what pets want in real time and purchase appropriate products based on that information. However, current technology has difficulty accurately estimating pet emotions and desires and making product purchase suggestions based on those estimations. Therefore, the challenge is to provide a system that can analyze pet emotions and desires in real time and automatically make product purchase suggestions.
[0745] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0746] In this invention, the server includes a terminal for collecting the movements, expressions, and cries of the pet, a means for receiving data transmitted from the terminal, a means including an AI model for analyzing the received data and estimating the pet's emotions and desires, a means for translating the estimated emotions and desires into natural language and notifying the user, and a means for automatically generating product purchase suggestions based on the estimated emotions and desires and notifying the user. This makes it possible to grasp the pet's emotions and desires in real time and automatically suggest appropriate products.
[0747] The "terminal" is a device that collects your pet's movements, facial expressions, and sounds in real time, and includes a camera and microphone.
[0748] The "server" is a computer that receives data sent from the terminal, analyzes it, and infers the pet's emotions and desires.
[0749] An "AI model" is an analytical model that uses machine learning algorithms to estimate a pet's emotions and desires.
[0750] "Means of translating into natural language" refers to the function of converting the inferred results output by the AI model into words that are easy for users to understand, such as "I'm hungry" or "I want to play."
[0751] "Means for generating product purchase suggestions" refers to a function that selects appropriate products (e.g., pet food or toys) based on estimated emotions and desires and notifies the user of the suggestions.
[0752] System Overview
[0753] This invention is a system that collects and analyzes pet movements, facial expressions, and cries in real time to estimate the pet's emotions and desires and make appropriate product purchase suggestions to the user. It is composed of the following components.
[0754] Key Components
[0755] 1. Device:
[0756] The terminal is a device equipped with a camera and microphone that captures your pet's movements, facial expressions, and sounds in real time, and the data is then sent to a server using a security protocol.
[0757] 2. Server:
[0758] The server receives the data sent from the device and analyzes it to infer the pet's emotions and needs. The analysis is performed using an AI model (machine learning algorithm). The server also translates the analysis results into natural language that is easy for the user to understand, and then generates appropriate product purchase suggestions.
[0759] 3. AI model:
[0760] The AI model, which uses the collected data to infer your pet's emotions and needs, includes computer vision and speech recognition algorithms.
[0761] 4. User Notification System:
[0762] The system notifies the user of the estimated emotions and desires, along with product purchase suggestions based on those emotions and desires. The user receives these notifications via their smartphone.
[0763] Software and Hardware Configuration
[0764] Hardware:
[0765] Camera: Capture videos and photos of your pet.
[0766] Microphone: Collects pet sounds.
[0767] Server: A computer that performs data analysis.
[0768] software:
[0769] OpenCV: Used for preprocessing video data.
[0770] Keras: Used to load and use AI models.
[0771] sounddevice: Used to collect audio data.
[0772] Requests: Used to send requests to the notification service.
[0773] How it works
[0774] The server receives the video and audio data sent from the device at runtime and pre-processes the received data appropriately. The video data is resized and normalized using OpenCV before being input to the AI model. The audio data is similarly pre-processed using the sounddevice and then input to the AI model.
[0775] The AI model uses this preprocessed data to infer the pet's emotions and desires. The inference results are translated into natural language and converted into user-friendly expressions such as "I'm hungry" or "I want to play." Based on this translation, the model then generates appropriate product purchase suggestions (e.g., pet food or toys).
[0776] Specific examples of implementation
[0777] Example 1: When a dog wants to play
[0778] 1. The device captures the dog's movements with a camera and collects the dog's vigorously wagging tail and high-pitched barking sounds.
[0779] 2. The collected data is sent to the server and separated into video data and audio data.
[0780] 3. The video and audio data are preprocessed and then input into the AI model.
[0781] 4. The AI model analyzes the data and infers that the dog wants to play.
[0782] 5. The server generates and sends the notification "Play with the ball!" and the suggestion "Would you like to buy a toy?" to the user.
[0783] Example prompt sentence:
[0784] "Write a Python program that monitors your pet's behavior and generates a notification saying 'Would you like to buy pet food?' if it detects that your pet is hungry."
[0785] In this way, a system can be provided for understanding a pet's emotions and desires in real time.
[0786] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[0787] Step 1:
[0788] The device uses a camera and microphone to collect the pet's movements, facial expressions, and cries in real time. Specifically, the camera captures the pet's movements as video, and the microphone records the pet's cries as audio data. These data are collected. The input is real-time video and audio data, and the output is the collected video and audio data.
[0789] Step 2:
[0790] The terminal divides the collected video and audio data into data packets at regular intervals and sends them to the server using a secure protocol. Specifically, the data is encrypted and sent using the HTTPS protocol. The input is the collected video and audio data, and the output is the encrypted data packets.
[0791] Step 3:
[0792] The server receives the data packets sent from the terminal and decrypts them. First, it receives the data, decrypts it, and returns it to the original video and audio data. The input is the encrypted data packet, and the output is the decrypted video and audio data.
[0793] Step 4:
[0794] The server analyzes the decoded data using computer vision and speech recognition algorithms. Specifically, it uses OpenCV to resize and normalize the video data. The audio data is preprocessed to extract audio features. The input is the decoded video and audio data, and the output is the preprocessed video and audio data.
[0795] Step 5:
[0796] The server inputs the preprocessed video and audio data into an AI model to estimate the pet's emotions and desires. Specifically, the data is input into a machine learning model trained using Keras to obtain an inference result. The input is the preprocessed video and audio data, and the output is the inference result of the pet's emotions and desires.
[0797] Step 6:
[0798] The server translates the inferences obtained from the AI model into natural language. For example, it converts them into messages such as "I want to play" or "I'm hungry." The input is the inferences from the AI model, and the output is a message in natural language.
[0799] Step 7:
[0800] The server generates appropriate product purchase suggestions based on the translated message, such as "Would you like to buy a toy?" or "Would you like to buy pet food?" The input is the translated message, and the output is the product purchase suggestion message.
[0801] Step 8:
[0802] The terminal receives the product purchase suggestion message sent from the server and notifies the user. The user can confirm the suggestion through a notification on their smartphone. The input is the product purchase suggestion message, and the output is a notification to the user.
[0803] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[0804] System configuration overview
[0805] The present invention is a system that includes a "terminal" for collecting the movements, facial expressions, and cries of a pet, a "server" for receiving data sent from the terminal, an "AI model" for analyzing the received data and inferring the pet's emotions and desires, a "notification system" for translating the inferred emotions and desires into natural language and notifying the user, and an "emotion engine" for recognizing the user's emotions.
[0806] Data collection details
[0807] The "terminal" is equipped with a camera and microphone, and collects the pet's movements, facial expressions, and sounds in real time. The collected data is divided into data packets at regular intervals and sent from the "terminal" to the "server." A security protocol is used to ensure the data is secure.
[0808] Data reception and analysis
[0809] The "server" receives the data packets sent from the terminal. The received data is first decoded and then processed separately into video data and audio data.
[0810] The "server" analyzes the video data using computer vision algorithms to identify the pet's facial expressions and movements (for example, tail movements). It also analyzes the audio data using speech recognition algorithms to identify the pattern, pitch, and frequency of the pet's cries. This allows the data's features to be extracted.
[0811] Estimating emotions and desires
[0812] The extracted features are input into an "AI model," which uses a pre-trained machine learning algorithm to infer the pet's emotions and desires. For example, if the pet wags its tail and makes a high-pitched meow, it can be inferred to be wanting to play.
[0813] Natural language translation and notification
[0814] The estimated emotions and desires are translated into natural language. Specifically, the output of the AI model is converted into messages such as "I'm hungry" or "Play with the ball!" These translated messages are then resent from the "server" to the "device" and notified to the user via the "device."
[0815] Recognizing user emotions and proposing countermeasures
[0816] In addition, the user's voice and facial expressions are also collected from the "terminal" and sent to the "server." The "emotion engine" analyzes this data to recognize the user's emotions. Using voice recognition algorithms and facial expression analysis algorithms, it identifies emotions from the user's voice tone and facial expressions.
[0817] Proposing solutions based on user sentiment
[0818] The server then uses the user's emotional information to suggest ways to help the pet. For example, if the user is tired, the server suggests creating a relaxing atmosphere for the pet.
[0819] Specific example explanation
[0820] Example 1: System behavior when a dog wants to play
[0821] 1. The "terminal" captures the dog's movements with a camera and collects the dog's tail wagging vigorously and high-pitched barking.
[0822] 2. The "terminal" sends the collected data to the "server."
[0823] 3. The "server" receives the data and identifies tail movements from the video data and high-pitched cries from the audio data.
[0824] 4. The “server” inputs these features into the “AI model.”
[0825] 5. The AI model uses these characteristics to infer that the dog wants to play.
[0826] 6. The "server" translates the result "I want to play" into natural language and makes it "Play with the ball!"
[0827] 7. The "terminal" notifies the user with the message "Play with the ball!"
[0828] 8. The "terminal" also collects the user's facial expressions and voice and sends them to the "server."
[0829] 9. The "Emotion Engine" recognizes the user's emotion as "Relaxed."
[0830] 10. The "server" checks the user's relaxation state and sends a message suggesting "Relax with your pet."
[0831] 11. The Terminal displays the suggestion message to the user.
[0832] Example 2: System behavior when the cat is hungry
[0833] 1. The "terminal" collects the cat's meowing (a low, sustained sound) and its movements around the dish.
[0834] 2. The "terminal" sends the collected data to the "server."
[0835] 3. The "server" receives the data and identifies the location from the video data (walking around the dish) and the sound from the audio data (a low, sustained sound).
[0836] 4. The “server” inputs these features into the “AI model.”
[0837] 5. The "AI model" uses these characteristics to infer that the cat is hungry.
[0838] 6. The "server" translates the result "hungry" into natural language and sets it as "I'm hungry."
[0839] 7. The "terminal" notifies the user with the message "I'm hungry."
[0840] 8. The "terminal" collects the user's facial expressions and voice and sends them to the "server."
[0841] 9. The "emotion engine" recognizes the user's emotion as "tired."
[0842] 10. The "server" checks the user's tiredness and sends a message suggesting that they take their time to feed their pet.
[0843] 11. The Terminal displays the suggestion message to the user.
[0844] In this way, the present invention can be implemented as a system that allows the user to understand the specific emotions and desires of their pet in real time and take appropriate action according to their own emotions and state.
[0845] The processing flow will be explained below.
[0846] Step 1:
[0847] The device uses a camera and microphone to collect your pet's movements, facial expressions, and sounds in real time, and the collected data is divided into video data and audio data, each of which is then split into separate data packets.
[0848] Step 2:
[0849] The device sends the collected data packets to the server, using a security protocol to ensure the data is protected.
[0850] Step 3:
[0851] The server receives the data packets sent from the terminal. The received data is first decoded and then separated into video data and audio data for processing.
[0852] Step 4:
[0853] The server runs the video data through computer vision algorithms to analyze your pet's facial expressions and movements, specifically identifying tail movements, facial expressions, and overall body movements.
[0854] Step 5:
[0855] The server runs the audio data through a speech recognition algorithm to identify the pattern, pitch, and frequency of your pet's cries, including the pitch and repetition of the cries.
[0856] Step 6:
[0857] The server combines the results of the video and audio analysis to extract features to estimate the pet's emotions and desires, such as high-pitched meows and tail wagging.
[0858] Step 7:
[0859] The server inputs the extracted features into a pre-trained AI model, which uses patterns learned from past data to infer the pet's emotions and desires.
[0860] Step 8:
[0861] The server receives the inferences output by the AI model and maps them to a list of predefined emotions and desires, such as "hunger," "want to play," "anxiety," and "relaxation."
[0862] Step 9:
[0863] The server translates the estimated emotions and desires into natural language, for example, translating "hungry" into "I'm hungry" and "I want to play" into "Play with my ball!"
[0864] Step 10:
[0865] The server then sends the translated message to the user's device, again using security protocols to protect the data.
[0866] Step 11:
[0867] The terminal receives the translated message sent from the server, and the received message is formatted to be displayed on the user interface.
[0868] Step 12:
[0869] The device will display messages to the user that indicate the pet's emotions and needs, such as "I'm hungry!" or "Play with my ball!", displayed as pop-ups or notification banners.
[0870] Step 13:
[0871] The device also uses a camera and microphone to collect the user's voice and facial expressions, which are then split into data packets and sent to the server.
[0872] Step 14:
[0873] The server receives the user's data packets sent from the terminal, decodes the received data, and prepares them for processing by separating them into video data and audio data.
[0874] Step 15:
[0875] The server runs the video data through a facial expression analysis algorithm to analyze the user's facial expressions, and the audio data through a voice recognition algorithm to analyze the user's tone and words.
[0876] Step 16:
[0877] The server integrates the analysis results of the facial expression data and voice data and extracts features to identify the user's emotions.
[0878] Step 17:
[0879] The server inputs the extracted features into an emotion engine to estimate the user's emotion, such as "tired," "relaxed," or "excited."
[0880] Step 18:
[0881] The server generates a message suggesting how to handle pets based on the user's emotions. For example, if the user is tired, it suggests, "Take your time and don't rush, take your time to feed your pet."
[0882] Step 19:
[0883] The server sends a proposal message to the user's device, which is formatted for display through the user interface.
[0884] Step 20:
[0885] The device will notify the user with suggestion messages, such as "Take your time to feed your pet and don't rush," via a pop-up or notification banner.
[0886] In this way, the present invention can be implemented as a system that allows the user to understand the specific emotions and desires of their pet in real time and take appropriate action according to their own emotions and state.
[0887] Example 2
[0888] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0889] It is difficult to understand the emotions and desires of both pets and users in real time and propose appropriate responses. Conventional technologies lack systems that can accurately estimate a pet's emotions and desires, translate them into natural language, and notify the user. Furthermore, there is no system that can recognize the user's emotions and propose appropriate responses based on them. This leads to a decline in the quality of pet care and a lack of communication between pets and users.
[0890] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[0891] In this invention, the server includes means for receiving data transmitted from the terminal, decoding and classifying the data, means for analyzing video data using a computer vision algorithm, and means for analyzing audio data using a speech recognition algorithm. This allows features of the pet's movements and cries to be extracted and input into an AI model to estimate the pet's emotions and needs. Furthermore, the estimation results are translated into natural language and notified to the user, allowing for an accurate understanding of the pet's situation. Furthermore, the server collects and analyzes the user's voice and facial expressions, and proposes appropriate countermeasures based on the user's emotions, thereby improving communication between the pet and the user.
[0892] The "terminal" is a device equipped with a camera and microphone to collect your pet's movements, facial expressions, and sounds.
[0893] A "server" is a computer system that receives data sent from a terminal and decodes, classifies, and analyzes the data.
[0894] A "data packet" is a unit for dividing and transmitting data collected by a terminal at regular intervals.
[0895] A "security protocol" is a communication protocol used to ensure security when transmitting data.
[0896] "Video data" refers to image data captured by a camera showing the movements and expressions of a pet.
[0897] "Audio data" is sound data recorded with a microphone of a pet's cries.
[0898] A "computer vision algorithm" is an image processing technology used to analyze specific movements and facial expressions from video data.
[0899] A "voice recognition algorithm" is a technology for analyzing specific sound patterns and characteristics from audio data.
[0900] "Features" are important data points extracted from the analyzed data that indicate the pet's emotions and behavior.
[0901] An "AI model" is a computational model that uses machine learning algorithms to estimate a pet's emotions and desires.
[0902] "Natural language translation" is the process of converting the results estimated by an AI model into a language that a user can understand.
[0903] "Notification" refers to the act of conveying a translated message to the user, and is primarily done via the terminal.
[0904] "User data" is data that collects information such as the user's voice and facial expressions.
[0905] An "emotion engine" is a system for analyzing user data to identify a user's emotions.
[0906] "Countermeasure suggestion" is a process of suggesting actions that the user should take toward their pet based on the identified user emotions.
[0907] The present invention is a system that analyzes the movements, facial expressions, and cries of a pet, estimates the feelings and desires of the pet, and notifies the user of the results. Detailed embodiments of the system will be described below.
[0908] System configuration
[0909] The system includes a "terminal" that collects the pet's movements, facial expressions, and cries, a "server" that receives the data sent from the terminal, an "AI model" that analyzes the data and infers the pet's emotions and desires, and a "notification system" that notifies the user. It also includes an "emotion engine" that recognizes the user's emotions and suggests appropriate countermeasures.
[0910] Data collection
[0911] The device is equipped with a camera and microphone to collect pet movements, facial expressions, and sounds in real time. This data is divided into data packets at regular intervals (e.g., every second) and sent to a server using a security protocol (e.g., SSL / TLS).
[0912] Data reception and decryption
[0913] The server receives the data packets sent from the terminal, decrypts them according to the security protocol, and then classifies the received data into video data and audio data.
[0914] Data analysis
[0915] The server analyzes the video data using a computer vision algorithm (e.g., OpenCV) to identify the pet's movements and facial expressions (e.g., tail wagging and ear movement). It also analyzes the audio data using a speech recognition algorithm (e.g., LibriSpeech) to identify the pattern, pitch, and frequency of the pet's cries. This allows the extraction of features.
[0916] Estimating emotions and desires
[0917] The server inputs the extracted features into an AI model (e.g., TensorFlow or PyTorch). The AI model is pre-trained with data on pets' emotions and desires, and uses the input features to infer the pet's emotions and desires. For example, if the pet wags its tail and makes a high-pitched bark, it can infer that the pet wants to play.
[0918] Natural language translation and notification
[0919] The server translates the results output by the AI model into natural language. For example, the predicted result "I want to play" is converted into the message "Play with the ball!" The translated message is resent from the server to the device and notified to the user via the device.
[0920] Recognizing user emotions and proposing countermeasures
[0921] The device collects the user's voice and facial expressions using a camera and microphone. The collected user data is sent to a server where it is analyzed by an emotion engine. Using a voice recognition algorithm and an expression analysis algorithm, emotions are identified from the user's tone of voice and facial expressions. Based on this emotional information, the server suggests appropriate countermeasures. For example, if the user is tired, a message suggesting, "Let's create a relaxing atmosphere" is sent.
[0922] Specific examples
[0923] Example 1: System behavior when a dog wants to play
[0924] 1. The device captures the dog's movements with a camera and collects the dog's vigorously wagging tail and high-pitched barking sounds.
[0925] 2. The device sends the collected data to the server.
[0926] 3. The server receives the data and identifies tail movements from the video data and high-pitched cries from the audio data.
[0927] 4. The server inputs these features into the AI model.
[0928] 5. The AI model uses these characteristics to infer that the dog wants to play.
[0929] 6. The server translates the result "I want to play" into natural language and makes it "Play with the ball!"
[0930] 7. The device notifies the user with the message "Play with the ball!"
[0931] 8. The device collects the user's facial expressions and voice and sends them to the server.
[0932] 9. The emotion engine recognizes the user's emotion as "relaxed."
[0933] 10. The server checks the user's relaxation state and sends a message suggesting "Relax with your pet."
[0934] 11. The terminal displays the suggestion message to the user.
[0935] Example 2: System behavior when the cat is hungry
[0936] 1. The device collects the cat's meowing (a low, sustained sound) and its movements around the dish.
[0937] 2. The device sends the collected data to the server.
[0938] 3. The server receives the data and identifies the location from the video data (walking around the dish) and the sound from the audio data (a low, sustained sound).
[0939] 4. The server inputs these features into the AI model.
[0940] 5. The AI model uses these characteristics to infer that the cat is hungry.
[0941] 6. The server translates the result "hungry" into natural language and gives it the value "I'm hungry."
[0942] 7. The device notifies the user with the message "I'm hungry."
[0943] 8. The device collects the user's facial expressions and voice and sends them to the server.
[0944] 9. The emotion engine recognizes the user's emotion as "tired."
[0945] 10. The server checks the user's tiredness and sends a message suggesting that they take their time to feed their pet.
[0946] 11. The terminal displays the suggestion message to the user.
[0947] Example prompts for generative AI models
[0948] "We want you to infer what the dog wants from the dog's behavior and bark data collected by the device."
[0949] "I want you to analyze what the cat wants at the moment based on its meows and behavioral patterns."
[0950] Thus, the present invention provides a system that can analyze the emotions and desires of pets and users in real time and encourage users to respond appropriately, thereby improving the quality of pet care and facilitating communication between pets and users.
[0951] The flow of the identification process in the second embodiment will be described with reference to FIG.
[0952] Step 1:
[0953] Your pet's movements are collected on the device using a camera and microphone.
[0954] Input: Video and audio data captured through a camera and microphone.
[0955] Specific behavior: The device records your pet's movements, facial expressions, and sounds in real time, such as capturing a dog's tail wagging or a cat's meowing.
[0956] Output: Collected video and audio data.
[0957] Step 2:
[0958] Data collected by the terminal is divided into data packets and sent to a server using a security protocol.
[0959] Input: Collected video and audio data.
[0960] Specific operation: Data is packetized every second and transmitted using security protocols such as SSL / TLS.
[0961] Output: Data packets protected by security protocols.
[0962] Step 3:
[0963] The server receives and decodes the data packets sent from the terminal.
[0964] Input: Data packets protected by security protocols.
[0965] Specific operation: The data packets are decrypted using the SSL / TLS protocol and restored to the original video and audio data.
[0966] Output: Decoded video and audio data.
[0967] Step 4:
[0968] The server classifies the decoded data into video data and audio data.
[0969] Input: Decoded video and audio data.
[0970] Specific operation: Data is classified into video and audio categories based on data format and metadata.
[0971] Output: Classified video and audio data.
[0972] Step 5:
[0973] The server analyzes the video data using computer vision algorithms.
[0974] Input: Classified video data.
[0975] Specific operation: Using computer vision algorithms such as OpenCV, the system analyzes pet movements and facial expressions. For example, it detects the wagging of a dog's tail or the movement of a cat's ears.
[0976] Output: Features extracted from video data.
[0977] Step 6:
[0978] The server analyzes the voice data using a voice recognition algorithm.
[0979] Input: Classified audio data.
[0980] What it does: Uses speech recognition algorithms such as LibriSpeech to identify the pattern, pitch, and frequency of your pet's cries.
[0981] Output: Features extracted from the audio data.
[0982] Step 7:
[0983] The server inputs the features extracted from the video and audio data into an AI model to estimate the pet's emotions and desires.
[0984] Input: Features extracted from video and audio data.
[0985] Specific behavior: Using machine learning algorithms such as TensorFlow and PyTorch, the system infers a pet's emotions and desires from its behavior and sounds. For example, if a pet wags its tail and makes a high-pitched meow, it infers that it wants to play.
[0986] Output: Estimated pet's emotions and desires.
[0987] Step 8:
[0988] The server translates the estimated emotions and desires into natural language and notifies the user.
[0989] Input: Estimated pet emotions and desires.
[0990] Specific operation: The inference results are converted into messages such as "Play with the ball!" or "I'm hungry" and sent to the user's device.
[0991] Output: Notification message translated into natural language.
[0992] Step 9:
[0993] The device collects the user's voice and facial expressions and sends them to the server.
[0994] Input: User's audio and video data.
[0995] Specific operation: The camera and microphone are used to record the user's real-time voice and facial expressions, which are then divided into data packets and sent to the server.
[0996] Output: Collected user audio and video data.
[0997] Step 10:
[0998] The server analyzes the collected user data using voice recognition algorithms and facial expression analysis algorithms to identify the user's emotions.
[0999] Input: Collected user audio and video data.
[1000] Specific behavior: Identify the user's emotions by analyzing voice tone and facial expressions. For example, if the user's voice tone is low and they sound tired, it will recognize them as "tired."
[1001] Output: Identified user sentiment.
[1002] Step 11:
[1003] The server suggests measures to take for the pet based on the user's feelings.
[1004] Input: Identified user sentiment.
[1005] Specific behavior: Generate and send messages to the user's device suggesting things like "Create a relaxing atmosphere" and "Take your time to feed your pet without rushing."
[1006] Output: A message with the proposed workaround.
[1007] (Application example 2)
[1008] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[1009] Currently, there are several systems on the market that monitor pets' movements and cries to estimate their emotions, but these systems cannot provide clear and specific suggestions to users. Furthermore, they lack a mechanism for providing appropriate content based on pets' emotions and desires. Therefore, there is a need for a system that can help users choose the best course of action for their pets.
[1010] The identification process by the identification processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes a terminal for collecting the movements, expressions, and cries of the pet, a means for receiving data transmitted from the terminal, a means including an AI model for analyzing the received data and estimating the pet's emotions and desires, a means for translating the estimated emotions and desires into natural language and notifying the user, and a means for selecting appropriate content based on the estimated emotions and desires of the pet and delivering it to the user terminal. This allows the user to understand the pet's emotions and desires in real time and obtain optimal content in a timely manner.
[1011] "Pet movements, expressions, and sounds" refers to the physical movements, facial expressions, and sounds made by pets.
[1012] The "terminal" is a device equipped with a camera and microphone that collects your pet's movements, facial expressions, and cries.
[1013] A "server" is a computer system that receives, analyzes, and processes data sent from a terminal.
[1014] An "AI model" is an artificial intelligence that uses machine learning algorithms to analyze data and estimate a pet's emotions and desires.
[1015] "Estimation" refers to predicting a pet's emotions and desires from the data obtained.
[1016] "Translation into natural language" means converting the inferred results output by the AI model into a message that is easy for the user to understand.
[1017] "Notifying the user" means conveying messages or information to the user through the terminal.
[1018] A "security protocol" is a communication protocol for maintaining the security of data when it is transmitted from a terminal to a server.
[1019] A "computer vision algorithm" is an algorithm that analyzes video data and identifies pet movements and facial expressions.
[1020] A "voice recognition algorithm" is an algorithm that analyzes voice data and identifies the sounds of pets.
[1021] "Selecting appropriate content" means selecting videos, music, information, etc. that are suitable for your pet based on the pet's estimated emotions and desires.
[1022] "Delivery" means transmitting selected content to a user terminal and having it played or displayed.
[1023] System configuration overview
[1024] This invention provides a system that collects pet movements, facial expressions, and cries, estimates the pet's emotions and desires based on the collected data, and delivers appropriate content to the user. The system mainly consists of the following components:
[1025] A "terminal" for collecting your pet's movements, facial expressions, and cries.
[1026] The "server" receives and analyzes data collected from the terminal.
[1027] An "AI model" that analyzes the received data and estimates the pet's emotions and desires.
[1028] A "notification system" that selects appropriate content based on estimated emotions and desires and delivers it to the user's device.
[1029] Data collection details
[1030] The "terminal" is equipped with a camera and microphone to collect the pet's movements, facial expressions, and sounds in real time, and the collected data is divided into data packets at regular intervals and sent to the "server" using a security protocol.
[1031] Data reception and analysis
[1032] The server receives and decodes the data packets sent from the device. It then processes the video and audio data separately. The video data uses computer vision algorithms to identify the pet's facial expressions and movements, while the audio data uses speech recognition algorithms to identify the pattern, pitch, and frequency of the pet's cries.
[1033] Estimating emotions and desires
[1034] The server inputs the analyzed features into an AI model to estimate the pet's emotions and desires. The AI model uses a trained machine learning algorithm, and can infer, for example, that a pet's desire to play is indicated by a wagging tail and high-pitched meow.
[1035] Natural language translation and notification
[1036] The estimated emotions and desires are translated into natural language and converted into specific messages (e.g., "I'm hungry" or "Play with a ball!"). The translated messages are then sent from the server to the device and notified to the user.
[1037] Recognizing user emotions and proposing countermeasures
[1038] The device also collects the user's voice and facial expressions and sends them to the server. The emotion engine analyzes this data and recognizes the user's emotions. For example, it uses voice recognition and facial expression analysis algorithms to identify emotions from the user's tone and facial expressions.
[1039] Proposing solutions based on user sentiment
[1040] The server generates a message suggesting appropriate actions for the pet based on the user's emotional information. For example, if the user is tired, the server might suggest, "Take your time to feed your pet and don't rush." The suggested message is displayed on the user's device, encouraging the user to take appropriate action.
[1041] Hardware and software used
[1042] The system uses the following hardware and software:
[1043] Hardware: Camera, microphone, smartphone, smart glasses, head-mounted display.
[1044] Software: OpenCV (used to collect video data from the camera), TensorFlow (used to analyze pet and user emotions), requests (used to communicate data with the server), virtual voice analysis module and emotion recognition module.
[1045] Specific example explanation
[1046] For example, if it is estimated that the pet wants to play, it will be treated as "playful" and the video content "play_video.mp4" will be selected and played. If the user is relaxed, the message "Enjoy a relaxing session with your pet" will be displayed, suggesting a better way to spend time with the pet.
[1047] Prompt Sentence Examples
[1048] "Suggest content suitable for pets who want to play."
[1049] "Suggest actions that are appropriate for when the user is relaxed."
[1050] In this way, a system is realized that can analyze emotions and desires based on real-time data from pets and users, and deliver optimal responses and content.
[1051] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[1052] Step 1:
[1053] The device uses a camera and microphone to collect the pet's movements, facial expressions, and sounds in real time. It receives camera video and audio data as input and divides it into data packets at regular intervals. This captures multiple features, such as the pet's movements, facial expressions, and sounds, which can then be sent to subsequent analysis steps.
[1054] Step 2:
[1055] The data packets collected by the terminal are sent to the server using a security protocol. The data packets are received as input, and the data is encrypted using a security protocol (e.g., SSL / TLS) and securely sent to the server. This ensures the confidentiality of the data.
[1056] Step 3:
[1057] The server receives data packets sent from the terminal. It receives encrypted data packets as input and first decrypts them. As output, it obtains decrypted video and audio data. This allows the server to store data in an analyzable format.
[1058] Step 4:
[1059] The server analyzes the video data to identify the pet's movements and facial expressions. It receives the decoded video data as input and analyzes it using a computer vision algorithm (e.g., OpenCV). The output is identified behavior and facial expression data, such as tail movements and facial expressions. This allows the pet's movement and facial expression features to be extracted.
[1060] Step 5:
[1061] The server analyzes the audio data to identify the pattern, pitch, and frequency of the pet's cries. It receives the decoded audio data as input and analyzes it using a speech recognition algorithm (e.g., a speech analysis module). The output is an analysis result that identifies the pattern, pitch, and frequency of the pet's cries. This allows the characteristics of the pet's cries to be extracted.
[1062] Step 6:
[1063] The server inputs the analyzed features into an AI model to estimate the pet's emotions and desires. The analysis results of movements, facial expressions, and cries are received as input and input into an AI model (for example, a model built with TensorFlow). The output is the pet's estimated emotions and desires (for example, "I want to play" or "I'm hungry"). This clarifies the pet's internal state.
[1064] Step 7:
[1065] The server translates the estimated emotions and desires into natural language and generates a message. It receives the estimated results from the AI model as input and translates them using a natural language processing engine. The output is a message that is easy for the user to understand, such as "I'm hungry" or "Play with my ball!" This communicates the pet's emotions and desires to the user.
[1066] Step 8:
[1067] The server selects appropriate content based on the estimated emotions and desires. It receives a message translated into natural language as input and uses an algorithm to select appropriate content (e.g., video, music). Specific content files (e.g., "play_video.mp4", "relaxing_music.mp3") are obtained as output. This allows content that matches the pet's state to be selected.
[1068] Step 9:
[1069] The server delivers the selected content to the user terminal. The server receives the selected content file as input and transmits it to the user terminal via the Internet. As output, the server delivers the content to be played or displayed on the user terminal. This allows the user to receive content that is appropriate for the condition of their pet.
[1070] Step 10:
[1071] The device collects the user's facial expressions and voice and sends them to the server. The device receives camera footage and voice data as input, divides it into data packets, and sends them to the server. This provides data for analyzing the user's emotions.
[1072] Step 11:
[1073] The server analyzes the user's facial expressions and voice to recognize the user's emotions. It receives the user's facial expressions and voice data as input and analyzes them using an emotion recognition algorithm. The output identifies the user's emotions (e.g., "relaxed" or "tired"). This clarifies the user's state.
[1074] Step 12:
[1075] The server generates a message suggesting how to respond to a pet based on the user's emotional information. It receives the user's emotion recognition results as input and generates the suggested message using a natural language processing engine. The output is specific suggested actions such as "Relax with your pet" or "Take your time to feed your pet without rushing." This provides guidance for the user to take appropriate action.
[1076] The above is a description of the specific processing steps of the system for implementing the present invention.
[1077] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[1078] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[1079] In the above embodiment, an example in which the specific processing is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the smart glasses 214.
[1080] [Third embodiment]
[1081] FIG. 5 shows an example of the configuration of a data processing system 310 according to the third embodiment.
[1082] 5, the data processing system 310 includes the data processing device 12 and a headset type terminal 314. An example of the data processing device 12 is a server.
[1083] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[1084] The headset type terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a display 343. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the display 343 are also connected to the bus 52.
[1085] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[1086] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[1087] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[1088] Fig. 6 shows an example of the main functions of the data processing device 12 and the headset type terminal 314. As shown in Fig. 6, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[1089] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[1090] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[1091] In the headset type terminal 314, a reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[1092] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the headset type terminal 314 will be referred to as the "terminal."
[1093] System configuration overview
[1094] The present invention is a system that includes a "terminal" for collecting pet movements, facial expressions, and cries, a "server" for receiving data sent from the terminal, an "AI model" for analyzing the received data and inferring the pet's emotions and desires, and a "notification system" that translates the inferred emotions and desires into natural language and notifies the user.
[1095] Data collection details
[1096] The "terminal" is equipped with a camera and microphone, and collects the pet's movements, facial expressions, and sounds in real time. The collected data is divided into data packets at regular intervals and sent from the "terminal" to the "server." A security protocol is used to ensure the data is secure.
[1097] Data reception and analysis
[1098] The "server" receives the data packets sent from the terminal. The received data is first decoded and then processed separately into video data and audio data.
[1099] The "server" analyzes the video data using computer vision algorithms to identify the pet's facial expressions and movements (for example, tail movements). It also analyzes the audio data using speech recognition algorithms to identify the pattern, pitch, and frequency of the pet's cries. This allows the data's features to be extracted.
[1100] Estimating emotions and desires
[1101] The extracted features are input into an "AI model," which uses a pre-trained machine learning algorithm to infer the pet's emotions and desires. For example, if a pet frequently wags its tail and makes a high-pitched meow, it can be inferred to have a desire to play.
[1102] Natural language translation and notification
[1103] The estimated emotions and desires are translated into natural language. Specifically, the output of the AI model is converted into messages such as "I'm hungry" or "Play with the ball!". This translated message is then resent from the "server" to the "device," which then notifies the user.
[1104] Specific example explanation
[1105] Example 1: System behavior when a dog wants to play
[1106] 1. The "terminal" captures the dog's movements with a camera and collects the dog's tail wagging vigorously and high-pitched barking.
[1107] 2. The "terminal" sends the collected data to the "server."
[1108] 3. The "server" receives the data and identifies tail movements from the video data and high-pitched cries from the audio data.
[1109] 4. The “server” inputs these features into the “AI model.”
[1110] 5. The AI model uses these characteristics to infer that the dog wants to play.
[1111] 6. The "server" translates the result "I want to play" into natural language and makes it "Play with the ball!"
[1112] 7. The "terminal" notifies the user with the message "Play with the ball!"
[1113] Example 2: System behavior when the cat is hungry
[1114] 1. The "terminal" collects the cat's meowing (a low, sustained sound) and its movements around the dish.
[1115] 2. The "terminal" sends the collected data to the "server."
[1116] 3. The "server" receives the data and identifies the location from the video data (walking around the dish) and the sound from the audio data (a low, sustained sound).
[1117] 4. The “server” inputs these features into the “AI model.”
[1118] 5. The "AI model" uses these characteristics to infer that the cat is hungry.
[1119] 6. The "server" translates the result "hungry" into natural language and sets it as "I'm hungry."
[1120] 7. The "terminal" notifies the user with the message "I'm hungry."
[1121] In this way, the present invention can be implemented as a system that allows users to understand their pet's specific emotions and desires in real time and take appropriate action.
[1122] The processing flow will be explained below.
[1123] Step 1:
[1124] The device uses a camera and microphone to collect the pet's movements, facial expressions, and sounds in real time. The collected data is then separated into video data and audio data, and each is split into separate data packets.
[1125] Step 2:
[1126] The device sends the collected data packets to a server, where a security protocol is used to ensure the data is protected.
[1127] Step 3:
[1128] The server receives the data packets sent from the terminal. The received data is first decoded and then separated into video data and audio data for processing.
[1129] Step 4:
[1130] The server runs the video data through computer vision algorithms to analyze your pet's facial expressions and movements, specifically identifying tail movements, facial expressions, and overall body movements.
[1131] Step 5:
[1132] The server runs the audio data through a speech recognition algorithm to identify the pattern, pitch, and frequency of your pet's cries, including the pitch of the cries and the repetition of the cries.
[1133] Step 6:
[1134] The server combines the results of the video and audio analysis to extract features to estimate the pet's emotions and desires, such as high-pitched meows and tail wagging.
[1135] Step 7:
[1136] The server inputs the extracted features into a pre-trained AI model, which uses patterns learned from past data to infer the pet's emotions and desires.
[1137] Step 8:
[1138] The server receives the inferences output by the AI model and maps them to a list of predefined emotions and desires, such as "hunger," "want to play," "anxiety," and "relaxation."
[1139] Step 9:
[1140] The server translates the estimated emotions and desires into natural language. For example, "hungry" becomes "I'm hungry" and "I want to play" becomes "Play with my ball!"
[1141] Step 10:
[1142] The server then sends the translated message to the user's device, again using security protocols to protect the data.
[1143] Step 11:
[1144] The terminal receives the translated message sent from the server, and the received message is formatted to be displayed on the user interface.
[1145] Step 12:
[1146] The device will display messages to the user that indicate the pet's emotions and needs, such as "I'm hungry!" or "Play with my ball!", displayed as pop-ups or notification banners.
[1147] Step 13:
[1148] The user checks the messages displayed on the device and takes action according to the pet's request, such as feeding the pet or playing with a ball.
[1149] This series of processes allows users to understand their pet's specific emotions and desires in real time and respond appropriately.
[1150] Example 1
[1151] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1152] Conventional pet management systems have had problems with accurately monitoring pet behavior and conditions, and owners are unable to properly understand their pets' emotions and needs, which increases stress for pets and burdens on owners. Furthermore, insufficient coordination in data collection, analysis, and notification makes it difficult to understand pets' conditions in real time.
[1153] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[1154] In this invention, the server includes means for decoding received data and processing the data separately into video data and audio data, means for analyzing the video data using a computer vision algorithm and the audio data using a speech recognition algorithm, and means for inputting the extracted features into an AI model to estimate emotions and desires. This makes it possible to instantly estimate emotions and desires from the pet's movements and cries, and notify the owner in natural language.
[1155] The "terminal" is a device equipped with a camera and microphone to collect the movements, facial expressions, and sounds of your pet.
[1156] A "server" is a computer system that receives, decrypts, and analyzes data sent from a terminal.
[1157] A "data packet" is a unit of information that divides collected data and is sent from a terminal to a server.
[1158] "Decoding" is the process of restoring encoded data to its original form.
[1159] "Video data" refers to data that visually records the movements and expressions of pets collected by a camera.
[1160] "Audio data" refers to data that records the sounds of pets as collected by a microphone.
[1161] A "computer vision algorithm" is a technology for analyzing video data and recognizing objects and actions.
[1162] A "voice recognition algorithm" is a technology for analyzing voice data and recognizing sounds and patterns.
[1163] "Features" are important information extracted from video and audio data for estimating a pet's emotions and desires.
[1164] The "AI model" is a system that is pre-trained using a machine learning algorithm and estimates a pet's emotions and desires based on input features.
[1165] "Natural language" refers to language that humans use on a daily basis, and is a language that expresses estimated emotions and desires in a way that is easy for users to understand.
[1166] A "security protocol" is a set of rules and procedures for securely communicating data.
[1167] "Notification" is the act of informing the user of the results of estimated emotions and desires.
[1168] MODE FOR CARRYING OUT THE INVENTION
[1169] The present invention is a system that includes a terminal that collects pet movements, facial expressions, and cries, a server that receives and analyzes the data, and a system that notifies the user of the estimated emotions and desires. Specifically, by using a pet monitoring device, a data analysis system, and a communication means, it is possible to grasp the pet's condition in real time and take appropriate measures.
[1170] System configuration overview
[1171] The present invention includes a "terminal" for collecting pet movements, facial expressions, and cries, a "server" for receiving data sent from the terminal, an "AI model" for analyzing the received data and inferring the pet's emotions and desires, and a "notification system" for translating the inferred emotions and desires into natural language and notifying the user.
[1172] Data collection details
[1173] The "terminal" is equipped with a camera and microphone, and collects the pet's movements, facial expressions, and cries in real time. For example, the camera records the pet's tail movements and eye expressions as video data, while the microphone collects the pet's cries as audio data. The collected data is divided into data packets at regular intervals and sent from the "terminal" to the "server." A security protocol is used for data transmission to ensure the safety of the data.
[1174] Data reception and analysis
[1175] The "server" receives data packets sent from the device. The received data is first decoded and then processed separately into video data and audio data. The "server" analyzes the video data using a computer vision algorithm (e.g., OpenCV) to identify the pet's facial expressions and movements. It also analyzes the audio data using a speech recognition algorithm (e.g., a TensorFlow or PyTorch-based model) to identify the pattern, pitch, and frequency of the pet's cries. This allows the data features to be extracted.
[1176] Estimating emotions and desires
[1177] The extracted features are input into an "AI model." The AI model uses a pre-trained machine learning algorithm (e.g., a generative AI model) to infer the pet's emotions and desires. For example, if the pet frequently wags its tail and makes a high-pitched meow, it can be inferred that the pet wants to play.
[1178] Natural language translation and notification
[1179] The estimated emotions and desires are translated into natural language. Specifically, the results output by the AI model are converted into messages such as "I'm hungry" or "Play with the ball!". These translated messages are then resent from the "server" to the "device," which then notifies the user. For example, the user can be notified by a push notification or voice alert on a smartphone app.
[1180] Specific example explanation
[1181] Example 1: System behavior when a dog wants to play
[1182] 1. The device captures the dog's movements with a camera and collects the dog's vigorously wagging tail and high-pitched barking sounds.
[1183] 2. The device sends the collected data to the server.
[1184] 3. The server receives the data and identifies tail movements from the video data and high-pitched cries from the audio data.
[1185] 4. The server inputs these features into the AI model.
[1186] 5. The AI model uses these characteristics to infer that the dog wants to play.
[1187] 6. The server translates the result "I want to play" into natural language and makes it "Play with the ball!"
[1188] 7. The device notifies the user with the message "Play with the ball!"
[1189] Example 2: System behavior when the cat is hungry
[1190] 1. The device collects the cat's meowing (a low, sustained sound) and its movements around the dish.
[1191] 2. The device sends the collected data to the server.
[1192] 3. The server receives the data and identifies the location from the video data (walking around the dish) and the sound from the audio data (a low, sustained sound).
[1193] 4. The server inputs these features into the AI model.
[1194] 5. The AI model uses these characteristics to infer that the cat is hungry.
[1195] 6. The server translates the result "hungry" into natural language and gives it the value "I'm hungry."
[1196] 7. The device notifies the user with the message "I'm hungry."
[1197] Example prompts to input to the generative AI model
[1198] 1. Example prompt for identifying a dog's desire to play: "When a dog wags its tail frequently and makes a high-pitched whine, it means it wants to play. Express this in natural language."
[1199] 2. Example prompt for identifying a cat's hunger state: "When a cat meows in a low, sustained tone and paces around its food bowl, it is likely hungry. Please describe this in natural language."
[1200] In this way, the present invention can be implemented as a system that allows users to understand their pet's specific emotions and desires in real time and take appropriate action.
[1201] The flow of the identification process in the first embodiment will be described with reference to FIG.
[1202] Program processing steps
[1203] Step 1: Collect data from the device
[1204] The device uses a camera and microphone to collect your pet's movements, facial expressions, and cries in real time.
[1205] Input: Real-time pet movements, facial expressions, and sounds
[1206] Output: Collected video and audio data
[1207] Specific behavior:
[1208] The camera records your pet's tail movements and facial expressions as video data, while the microphone captures your pet's barks as audio data, such as when your dog wags its tail or barks in a high-pitched voice.
[1209] Step 2: Send data from the device to the server
[1210] The device divides the collected data into data packets and sends them to the server using a security protocol.
[1211] Input: Collected video and audio data
[1212] Output: Data packets sent from the device to the server
[1213] Specific behavior:
[1214] The collected data is divided into packets of a certain size, encrypted using the TLS protocol, and sent to the server.
[1215] Step 3: Server receives and decrypts data
[1216] The server receives the data packets sent from the terminal, decodes them, and separates them into video data and audio data.
[1217] Input: Data packets sent from the device
[1218] Output: Decoded video and audio data
[1219] Specific behavior:
[1220] The received data packets are decrypted using the TLS protocol and separated into collected video and audio data, for example, tail movements recorded by a camera or barks recorded by a microphone.
[1221] Step 4: Video data analysis by the server
[1222] The server analyzes the video data using computer vision algorithms to identify the pet's facial expressions and movements.
[1223] Input: Decoded video data
[1224] Output: Features of pet's facial expressions and movements
[1225] Specific behavior:
[1226] Using OpenCV, features such as tail wagging and eye movement are extracted from video data. For example, if a cat is wagging its tail vigorously, that movement can be identified.
[1227] Step 5: Server analyzes the audio data
[1228] The server analyzes the audio data using a voice recognition algorithm to identify the pattern, pitch and frequency of your pet's cries.
[1229] Input: Decoded audio data
[1230] Output: Features of pet sounds
[1231] Specific behavior:
[1232] Use TensorFlow or PyTorch to identify high-pitched meows and patterns in audio data, such as detecting the low, sustained meow of a cat.
[1233] Step 6: Feature extraction by the server
[1234] The server integrates the features extracted from the video and audio data and inputs them into the AI model.
[1235] Input: Features of pet's facial expressions and movements, features of pet's cries
[1236] Output: Integrated features
[1237] Specific behavior:
[1238] The identified tail movements from the video data and the call patterns from the audio data are integrated into a single dataset.
[1239] Step 7: Server inputs AI model
[1240] The server inputs the extracted features into an AI model to estimate emotions and desires.
[1241] Input: Integrated features
[1242] Output: Estimated pet's emotions and desires
[1243] Specific behavior:
[1244] The integrated features are input into a generative AI model to obtain an inference. For example, if a dog is wagging its tail and making a high-pitched bark, it may infer that it wants to play.
[1245] Step 8: Server-based translation into natural language
[1246] The server translates the estimated emotions and desires into natural language.
[1247] Input: Inference results from the AI model
[1248] Output: A message expressed in natural language
[1249] Specific behavior:
[1250] The result "I want to play" obtained from the AI model is converted into the message "Play with the ball!"
[1251] Step 9: Sending translation results from the server to the device
[1252] The server retransmits the translated message to the terminal.
[1253] Input: A message expressed in natural language
[1254] Output: Messages sent to the terminal
[1255] Specific behavior:
[1256] Using security protocols, send a translated message to the device, for example, "Play with the ball!"
[1257] Step 10: Notify the user via the device
[1258] The terminal receives the message sent from the server and notifies the user of the message.
[1259] Input: The message sent from the server
[1260] Output: The notification the user receives
[1261] Specific behavior:
[1262] The device uses push notifications and audio alerts on a smartphone app to notify users with messages like, "Play with the ball!"
[1263] Step 11: User Action
[1264] The user checks the notification from the device and responds to the pet.
[1265] Input: Notifications from your device
[1266] Output: Specific responses to pets
[1267] Specific behavior:
[1268] The user checks the notification and takes action based on the message, such as playing with their pet. For example, if they receive the message "Play with a ball!", they will get a ball and start playing with their dog.
[1269] The above is a description of the specific processing steps in the system program.
[1270] (Application example 1)
[1271] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1272] In recent years, there has been a demand for technology that can understand pet emotions and desires and enable owners to provide appropriate care. In particular, there is a need for a system that can grasp what pets want in real time and purchase appropriate products based on that information. However, current technology has difficulty accurately estimating pet emotions and desires and making product purchase suggestions based on those estimations. Therefore, the challenge is to provide a system that can analyze pet emotions and desires in real time and automatically make product purchase suggestions.
[1273] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[1274] In this invention, the server includes a terminal for collecting the movements, expressions, and cries of the pet, a means for receiving data transmitted from the terminal, a means including an AI model for analyzing the received data and estimating the pet's emotions and desires, a means for translating the estimated emotions and desires into natural language and notifying the user, and a means for automatically generating product purchase suggestions based on the estimated emotions and desires and notifying the user. This makes it possible to grasp the pet's emotions and desires in real time and automatically suggest appropriate products.
[1275] The "terminal" is a device that collects your pet's movements, facial expressions, and sounds in real time, and includes a camera and microphone.
[1276] The "server" is a computer that receives data sent from the terminal, analyzes it, and infers the pet's emotions and desires.
[1277] An "AI model" is an analytical model that uses machine learning algorithms to estimate a pet's emotions and desires.
[1278] "Means of translating into natural language" refers to the function of converting the inferred results output by the AI model into words that are easy for users to understand, such as "I'm hungry" or "I want to play."
[1279] "Means for generating product purchase suggestions" refers to a function that selects appropriate products (e.g., pet food or toys) based on estimated emotions and desires and notifies the user of the suggestions.
[1280] System Overview
[1281] This invention is a system that collects and analyzes pet movements, facial expressions, and cries in real time to estimate the pet's emotions and desires and make appropriate product purchase suggestions to the user. It is composed of the following components.
[1282] Key Components
[1283] 1. Device:
[1284] The terminal is a device equipped with a camera and microphone that captures your pet's movements, facial expressions, and sounds in real time, and the data is then sent to a server using a security protocol.
[1285] 2. Server:
[1286] The server receives the data sent from the device and analyzes it to infer the pet's emotions and needs. The analysis is performed using an AI model (machine learning algorithm). The server also translates the analysis results into natural language that is easy for the user to understand, and then generates appropriate product purchase suggestions.
[1287] 3. AI model:
[1288] The AI model, which uses the collected data to infer your pet's emotions and needs, includes computer vision and speech recognition algorithms.
[1289] 4. User Notification System:
[1290] The system notifies the user of the estimated emotions and desires, along with product purchase suggestions based on those emotions and desires. The user receives these notifications via their smartphone.
[1291] Software and Hardware Configuration
[1292] Hardware:
[1293] Camera: Capture videos and photos of your pet.
[1294] Microphone: Collects pet sounds.
[1295] Server: A computer that performs data analysis.
[1296] software:
[1297] OpenCV: Used for preprocessing video data.
[1298] Keras: Used to load and use AI models.
[1299] sounddevice: Used to collect audio data.
[1300] Requests: Used to send requests to the notification service.
[1301] How it works
[1302] The server receives the video and audio data sent from the device at runtime and pre-processes the received data appropriately. The video data is resized and normalized using OpenCV before being input to the AI model. The audio data is similarly pre-processed using the sounddevice and then input to the AI model.
[1303] The AI model uses this preprocessed data to infer the pet's emotions and desires. The inference results are translated into natural language and converted into user-friendly expressions such as "I'm hungry" or "I want to play." Based on this translation, the model then generates appropriate product purchase suggestions (e.g., pet food or toys).
[1304] Specific examples of implementation
[1305] Example 1: When a dog wants to play
[1306] 1. The device captures the dog's movements with a camera and collects the dog's vigorously wagging tail and high-pitched barking sounds.
[1307] 2. The collected data is sent to the server and separated into video data and audio data.
[1308] 3. The video and audio data are preprocessed and then input into the AI model.
[1309] 4. The AI model analyzes the data and infers that the dog wants to play.
[1310] 5. The server generates and sends the notification "Play with the ball!" and the suggestion "Would you like to buy a toy?" to the user.
[1311] Example prompt sentence:
[1312] "Write a Python program that monitors your pet's behavior and generates a notification saying 'Would you like to buy pet food?' if it detects that your pet is hungry."
[1313] In this way, a system can be provided for understanding a pet's emotions and desires in real time.
[1314] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[1315] Step 1:
[1316] The device uses a camera and microphone to collect the pet's movements, facial expressions, and cries in real time. Specifically, the camera captures the pet's movements as video, and the microphone records the pet's cries as audio data. These data are collected. The input is real-time video and audio data, and the output is the collected video and audio data.
[1317] Step 2:
[1318] The terminal divides the collected video and audio data into data packets at regular intervals and sends them to the server using a secure protocol. Specifically, the data is encrypted and sent using the HTTPS protocol. The input is the collected video and audio data, and the output is the encrypted data packets.
[1319] Step 3:
[1320] The server receives the data packets sent from the terminal and decrypts them. First, it receives the data, decrypts it, and returns it to the original video and audio data. The input is the encrypted data packet, and the output is the decrypted video and audio data.
[1321] Step 4:
[1322] The server analyzes the decoded data using computer vision and speech recognition algorithms. Specifically, it uses OpenCV to resize and normalize the video data. The audio data is preprocessed to extract audio features. The input is the decoded video and audio data, and the output is the preprocessed video and audio data.
[1323] Step 5:
[1324] The server inputs the preprocessed video and audio data into an AI model to estimate the pet's emotions and desires. Specifically, the data is input into a machine learning model trained using Keras to obtain an inference result. The input is the preprocessed video and audio data, and the output is the inference result of the pet's emotions and desires.
[1325] Step 6:
[1326] The server translates the inferences obtained from the AI model into natural language. For example, it converts them into messages such as "I want to play" or "I'm hungry." The input is the inferences from the AI model, and the output is a message in natural language.
[1327] Step 7:
[1328] The server generates appropriate product purchase suggestions based on the translated message, such as "Would you like to buy a toy?" or "Would you like to buy pet food?" The input is the translated message, and the output is the product purchase suggestion message.
[1329] Step 8:
[1330] The terminal receives the product purchase suggestion message sent from the server and notifies the user. The user can confirm the suggestion through a notification on their smartphone. The input is the product purchase suggestion message, and the output is a notification to the user.
[1331] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[1332] System configuration overview
[1333] The present invention is a system that includes a "terminal" for collecting the movements, facial expressions, and cries of a pet, a "server" for receiving data sent from the terminal, an "AI model" for analyzing the received data and inferring the pet's emotions and desires, a "notification system" for translating the inferred emotions and desires into natural language and notifying the user, and an "emotion engine" for recognizing the user's emotions.
[1334] Data collection details
[1335] The "terminal" is equipped with a camera and microphone, and collects the pet's movements, facial expressions, and sounds in real time. The collected data is divided into data packets at regular intervals and sent from the "terminal" to the "server." A security protocol is used to ensure the data is secure.
[1336] Data reception and analysis
[1337] The "server" receives the data packets sent from the terminal. The received data is first decoded and then processed separately into video data and audio data.
[1338] The "server" analyzes the video data using computer vision algorithms to identify the pet's facial expressions and movements (for example, tail movements). It also analyzes the audio data using speech recognition algorithms to identify the pattern, pitch, and frequency of the pet's cries. This allows the data's features to be extracted.
[1339] Estimating emotions and desires
[1340] The extracted features are input into an "AI model," which uses a pre-trained machine learning algorithm to infer the pet's emotions and desires. For example, if the pet wags its tail and makes a high-pitched meow, it can be inferred to be wanting to play.
[1341] Natural language translation and notification
[1342] The estimated emotions and desires are translated into natural language. Specifically, the output of the AI model is converted into messages such as "I'm hungry" or "Play with the ball!" These translated messages are then resent from the "server" to the "device" and notified to the user via the "device."
[1343] Recognizing user emotions and proposing countermeasures
[1344] In addition, the user's voice and facial expressions are also collected from the "terminal" and sent to the "server." The "emotion engine" analyzes this data to recognize the user's emotions. Using voice recognition algorithms and facial expression analysis algorithms, it identifies emotions from the user's voice tone and facial expressions.
[1345] Proposing solutions based on user sentiment
[1346] The server then uses the user's emotional information to suggest ways to help the pet. For example, if the user is tired, the server suggests creating a relaxing atmosphere for the pet.
[1347] Specific example explanation
[1348] Example 1: System behavior when a dog wants to play
[1349] 1. The "terminal" captures the dog's movements with a camera and collects the dog's tail wagging vigorously and high-pitched barking.
[1350] 2. The "terminal" sends the collected data to the "server."
[1351] 3. The "server" receives the data and identifies tail movements from the video data and high-pitched cries from the audio data.
[1352] 4. The “server” inputs these features into the “AI model.”
[1353] 5. The AI model uses these characteristics to infer that the dog wants to play.
[1354] 6. The "server" translates the result "I want to play" into natural language and makes it "Play with the ball!"
[1355] 7. The "terminal" notifies the user with the message "Play with the ball!"
[1356] 8. The "terminal" also collects the user's facial expressions and voice and sends them to the "server."
[1357] 9. The "Emotion Engine" recognizes the user's emotion as "Relaxed."
[1358] 10. The "server" checks the user's relaxation state and sends a message suggesting "Relax with your pet."
[1359] 11. The Terminal displays the suggestion message to the user.
[1360] Example 2: System behavior when the cat is hungry
[1361] 1. The "terminal" collects the cat's meowing (a low, sustained sound) and its movements around the dish.
[1362] 2. The "terminal" sends the collected data to the "server."
[1363] 3. The "server" receives the data and identifies the location from the video data (walking around the dish) and the sound from the audio data (a low, sustained sound).
[1364] 4. The “server” inputs these features into the “AI model.”
[1365] 5. The "AI model" uses these characteristics to infer that the cat is hungry.
[1366] 6. The "server" translates the result "hungry" into natural language and sets it as "I'm hungry."
[1367] 7. The "terminal" notifies the user with the message "I'm hungry."
[1368] 8. The "terminal" collects the user's facial expressions and voice and sends them to the "server."
[1369] 9. The "emotion engine" recognizes the user's emotion as "tired."
[1370] 10. The "server" checks the user's tiredness and sends a message suggesting that they take their time to feed their pet.
[1371] 11. The Terminal displays the suggestion message to the user.
[1372] In this way, the present invention can be implemented as a system that allows the user to understand the specific emotions and desires of their pet in real time and take appropriate action according to their own emotions and state.
[1373] The processing flow will be explained below.
[1374] Step 1:
[1375] The device uses a camera and microphone to collect your pet's movements, facial expressions, and sounds in real time, and the collected data is divided into video data and audio data, each of which is then split into separate data packets.
[1376] Step 2:
[1377] The device sends the collected data packets to the server, using a security protocol to ensure the data is protected.
[1378] Step 3:
[1379] The server receives the data packets sent from the terminal. The received data is first decoded and then separated into video data and audio data for processing.
[1380] Step 4:
[1381] The server runs the video data through computer vision algorithms to analyze your pet's facial expressions and movements, specifically identifying tail movements, facial expressions, and overall body movements.
[1382] Step 5:
[1383] The server runs the audio data through a speech recognition algorithm to identify the pattern, pitch, and frequency of your pet's cries, including the pitch and repetition of the cries.
[1384] Step 6:
[1385] The server combines the results of the video and audio analysis to extract features to estimate the pet's emotions and desires, such as high-pitched meows and tail wagging.
[1386] Step 7:
[1387] The server inputs the extracted features into a pre-trained AI model, which uses patterns learned from past data to infer the pet's emotions and desires.
[1388] Step 8:
[1389] The server receives the inferences output by the AI model and maps them to a list of predefined emotions and desires, such as "hunger," "want to play," "anxiety," and "relaxation."
[1390] Step 9:
[1391] The server translates the estimated emotions and desires into natural language, for example, translating "hungry" into "I'm hungry" and "I want to play" into "Play with my ball!"
[1392] Step 10:
[1393] The server then sends the translated message to the user's device, again using security protocols to protect the data.
[1394] Step 11:
[1395] The terminal receives the translated message sent from the server, and the received message is formatted to be displayed on the user interface.
[1396] Step 12:
[1397] The device will display messages to the user that indicate the pet's emotions and needs, such as "I'm hungry!" or "Play with my ball!", displayed as pop-ups or notification banners.
[1398] Step 13:
[1399] The device also uses a camera and microphone to collect the user's voice and facial expressions, which are then split into data packets and sent to the server.
[1400] Step 14:
[1401] The server receives the user's data packets sent from the terminal, decodes the received data, and prepares them for processing by separating them into video data and audio data.
[1402] Step 15:
[1403] The server runs the video data through a facial expression analysis algorithm to analyze the user's facial expressions, and the audio data through a voice recognition algorithm to analyze the user's tone and words.
[1404] Step 16:
[1405] The server integrates the analysis results of the facial expression data and voice data and extracts features to identify the user's emotions.
[1406] Step 17:
[1407] The server inputs the extracted features into an emotion engine to estimate the user's emotion, such as "tired," "relaxed," or "excited."
[1408] Step 18:
[1409] The server generates a message suggesting how to handle pets based on the user's emotions. For example, if the user is tired, it suggests, "Take your time and don't rush, take your time to feed your pet."
[1410] Step 19:
[1411] The server sends a proposal message to the user's device, which is formatted for display through the user interface.
[1412] Step 20:
[1413] The device will notify the user with suggestion messages, such as "Take your time to feed your pet and don't rush," via a pop-up or notification banner.
[1414] In this way, the present invention can be implemented as a system that allows the user to understand the specific emotions and desires of their pet in real time and take appropriate action according to their own emotions and state.
[1415] Example 2
[1416] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1417] It is difficult to understand the emotions and desires of both pets and users in real time and propose appropriate responses. Conventional technologies lack systems that can accurately estimate a pet's emotions and desires, translate them into natural language, and notify the user. Furthermore, there is no system that can recognize the user's emotions and propose appropriate responses based on them. This leads to a decline in the quality of pet care and a lack of communication between pets and users.
[1418] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[1419] In this invention, the server includes means for receiving data transmitted from the terminal, decoding and classifying the data, means for analyzing video data using a computer vision algorithm, and means for analyzing audio data using a speech recognition algorithm. This allows features of the pet's movements and cries to be extracted and input into an AI model to estimate the pet's emotions and needs. Furthermore, the estimation results are translated into natural language and notified to the user, allowing for an accurate understanding of the pet's situation. Furthermore, the server collects and analyzes the user's voice and facial expressions, and proposes appropriate countermeasures based on the user's emotions, thereby improving communication between the pet and the user.
[1420] The "terminal" is a device equipped with a camera and microphone to collect your pet's movements, facial expressions, and sounds.
[1421] A "server" is a computer system that receives data sent from a terminal and decodes, classifies, and analyzes the data.
[1422] A "data packet" is a unit for dividing and transmitting data collected by a terminal at regular intervals.
[1423] A "security protocol" is a communication protocol used to ensure security when transmitting data.
[1424] "Video data" refers to image data captured by a camera showing the movements and expressions of a pet.
[1425] "Audio data" is sound data recorded with a microphone of a pet's cries.
[1426] A "computer vision algorithm" is an image processing technology used to analyze specific movements and facial expressions from video data.
[1427] A "voice recognition algorithm" is a technology for analyzing specific sound patterns and characteristics from audio data.
[1428] "Features" are important data points extracted from the analyzed data that indicate the pet's emotions and behavior.
[1429] An "AI model" is a computational model that uses machine learning algorithms to estimate a pet's emotions and desires.
[1430] "Natural language translation" is the process of converting the results estimated by an AI model into a language that a user can understand.
[1431] "Notification" refers to the act of conveying a translated message to the user, and is primarily done via the terminal.
[1432] "User data" is data that collects information such as the user's voice and facial expressions.
[1433] An "emotion engine" is a system for analyzing user data to identify a user's emotions.
[1434] "Countermeasure suggestion" is a process of suggesting actions that the user should take toward their pet based on the identified user emotions.
[1435] The present invention is a system that analyzes the movements, facial expressions, and cries of a pet, estimates the feelings and desires of the pet, and notifies the user of the results. Detailed embodiments of the system will be described below.
[1436] System configuration
[1437] The system includes a "terminal" that collects the pet's movements, facial expressions, and cries, a "server" that receives the data sent from the terminal, an "AI model" that analyzes the data and infers the pet's emotions and desires, and a "notification system" that notifies the user. It also includes an "emotion engine" that recognizes the user's emotions and suggests appropriate countermeasures.
[1438] Data collection
[1439] The device is equipped with a camera and microphone to collect pet movements, facial expressions, and sounds in real time. This data is divided into data packets at regular intervals (e.g., every second) and sent to a server using a security protocol (e.g., SSL / TLS).
[1440] Data reception and decryption
[1441] The server receives the data packets sent from the terminal, decrypts them according to the security protocol, and then classifies the received data into video data and audio data.
[1442] Data analysis
[1443] The server analyzes the video data using a computer vision algorithm (e.g., OpenCV) to identify the pet's movements and facial expressions (e.g., tail wagging and ear movement). It also analyzes the audio data using a speech recognition algorithm (e.g., LibriSpeech) to identify the pattern, pitch, and frequency of the pet's cries. This allows the extraction of features.
[1444] Estimating emotions and desires
[1445] The server inputs the extracted features into an AI model (e.g., TensorFlow or PyTorch). The AI model is pre-trained with data on pets' emotions and desires, and uses the input features to infer the pet's emotions and desires. For example, if the pet wags its tail and makes a high-pitched bark, it can infer that the pet wants to play.
[1446] Natural language translation and notification
[1447] The server translates the results output by the AI model into natural language. For example, the predicted result "I want to play" is converted into the message "Play with the ball!" The translated message is resent from the server to the device and notified to the user via the device.
[1448] Recognizing user emotions and proposing countermeasures
[1449] The device collects the user's voice and facial expressions using a camera and microphone. The collected user data is sent to a server where it is analyzed by an emotion engine. Using a voice recognition algorithm and an expression analysis algorithm, emotions are identified from the user's tone of voice and facial expressions. Based on this emotional information, the server suggests appropriate countermeasures. For example, if the user is tired, a message suggesting, "Let's create a relaxing atmosphere" is sent.
[1450] Specific examples
[1451] Example 1: System behavior when a dog wants to play
[1452] 1. The device captures the dog's movements with a camera and collects the dog's vigorously wagging tail and high-pitched barking sounds.
[1453] 2. The device sends the collected data to the server.
[1454] 3. The server receives the data and identifies tail movements from the video data and high-pitched cries from the audio data.
[1455] 4. The server inputs these features into the AI model.
[1456] 5. The AI model uses these characteristics to infer that the dog wants to play.
[1457] 6. The server translates the result "I want to play" into natural language and makes it "Play with the ball!"
[1458] 7. The device notifies the user with the message "Play with the ball!"
[1459] 8. The device collects the user's facial expressions and voice and sends them to the server.
[1460] 9. The emotion engine recognizes the user's emotion as "relaxed."
[1461] 10. The server checks the user's relaxation state and sends a message suggesting "Relax with your pet."
[1462] 11. The terminal displays the suggestion message to the user.
[1463] Example 2: System behavior when the cat is hungry
[1464] 1. The device collects the cat's meowing (a low, sustained sound) and its movements around the dish.
[1465] 2. The device sends the collected data to the server.
[1466] 3. The server receives the data and identifies the location from the video data (walking around the dish) and the sound from the audio data (a low, sustained sound).
[1467] 4. The server inputs these features into the AI model.
[1468] 5. The AI model uses these characteristics to infer that the cat is hungry.
[1469] 6. The server translates the result "hungry" into natural language and gives it the value "I'm hungry."
[1470] 7. The device notifies the user with the message "I'm hungry."
[1471] 8. The device collects the user's facial expressions and voice and sends them to the server.
[1472] 9. The emotion engine recognizes the user's emotion as "tired."
[1473] 10. The server checks the user's tiredness and sends a message suggesting that they take their time to feed their pet.
[1474] 11. The terminal displays the suggestion message to the user.
[1475] Example prompts for generative AI models
[1476] "We want you to infer what the dog wants from the dog's behavior and bark data collected by the device."
[1477] "I want you to analyze what the cat wants at the moment based on its meows and behavioral patterns."
[1478] Thus, the present invention provides a system that can analyze the emotions and desires of pets and users in real time and encourage users to respond appropriately, thereby improving the quality of pet care and facilitating communication between pets and users.
[1479] The flow of the identification process in the second embodiment will be described with reference to FIG.
[1480] Step 1:
[1481] Your pet's movements are collected on the device using a camera and microphone.
[1482] Input: Video and audio data captured through a camera and microphone.
[1483] Specific behavior: The device records your pet's movements, facial expressions, and sounds in real time, such as capturing a dog's tail wagging or a cat's meowing.
[1484] Output: Collected video and audio data.
[1485] Step 2:
[1486] Data collected by the terminal is divided into data packets and sent to a server using a security protocol.
[1487] Input: Collected video and audio data.
[1488] Specific operation: Data is packetized every second and transmitted using security protocols such as SSL / TLS.
[1489] Output: Data packets protected by security protocols.
[1490] Step 3:
[1491] The server receives and decodes the data packets sent from the terminal.
[1492] Input: Data packets protected by security protocols.
[1493] Specific operation: The data packets are decrypted using the SSL / TLS protocol and restored to the original video and audio data.
[1494] Output: Decoded video and audio data.
[1495] Step 4:
[1496] The server classifies the decoded data into video data and audio data.
[1497] Input: Decoded video and audio data.
[1498] Specific operation: Data is classified into video and audio categories based on data format and metadata.
[1499] Output: Classified video and audio data.
[1500] Step 5:
[1501] The server analyzes the video data using computer vision algorithms.
[1502] Input: Classified video data.
[1503] Specific operation: Using computer vision algorithms such as OpenCV, the system analyzes pet movements and facial expressions. For example, it detects the wagging of a dog's tail or the movement of a cat's ears.
[1504] Output: Features extracted from video data.
[1505] Step 6:
[1506] The server analyzes the voice data using a voice recognition algorithm.
[1507] Input: Classified audio data.
[1508] What it does: Uses speech recognition algorithms such as LibriSpeech to identify the pattern, pitch, and frequency of your pet's cries.
[1509] Output: Features extracted from the audio data.
[1510] Step 7:
[1511] The server inputs the features extracted from the video and audio data into an AI model to estimate the pet's emotions and desires.
[1512] Input: Features extracted from video and audio data.
[1513] Specific behavior: Using machine learning algorithms such as TensorFlow and PyTorch, the system infers a pet's emotions and desires from its behavior and sounds. For example, if a pet wags its tail and makes a high-pitched meow, it infers that it wants to play.
[1514] Output: Estimated pet's emotions and desires.
[1515] Step 8:
[1516] The server translates the estimated emotions and desires into natural language and notifies the user.
[1517] Input: Estimated pet emotions and desires.
[1518] Specific operation: The inference results are converted into messages such as "Play with the ball!" or "I'm hungry" and sent to the user's device.
[1519] Output: Notification message translated into natural language.
[1520] Step 9:
[1521] The device collects the user's voice and facial expressions and sends them to the server.
[1522] Input: User's audio and video data.
[1523] Specific operation: The camera and microphone are used to record the user's real-time voice and facial expressions, which are then divided into data packets and sent to the server.
[1524] Output: Collected user audio and video data.
[1525] Step 10:
[1526] The server analyzes the collected user data using voice recognition algorithms and facial expression analysis algorithms to identify the user's emotions.
[1527] Input: Collected user audio and video data.
[1528] Specific behavior: Identify the user's emotions by analyzing voice tone and facial expressions. For example, if the user's voice tone is low and they sound tired, it will recognize them as "tired."
[1529] Output: Identified user sentiment.
[1530] Step 11:
[1531] The server suggests measures to take for the pet based on the user's feelings.
[1532] Input: Identified user sentiment.
[1533] Specific behavior: Generate and send messages to the user's device suggesting things like "Create a relaxing atmosphere" and "Take your time to feed your pet without rushing."
[1534] Output: A message with the proposed workaround.
[1535] (Application example 2)
[1536] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1537] Currently, there are several systems on the market that monitor pets' movements and cries to estimate their emotions, but these systems cannot provide clear and specific suggestions to users. Furthermore, they lack a mechanism for providing appropriate content based on pets' emotions and desires. Therefore, there is a need for a system that can help users choose the best course of action for their pets.
[1538] The identification process by the identification processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes a terminal for collecting the movements, expressions, and cries of the pet, a means for receiving data transmitted from the terminal, a means including an AI model for analyzing the received data and estimating the pet's emotions and desires, a means for translating the estimated emotions and desires into natural language and notifying the user, and a means for selecting appropriate content based on the estimated emotions and desires of the pet and delivering it to the user terminal. This allows the user to understand the pet's emotions and desires in real time and obtain optimal content in a timely manner.
[1539] "Pet movements, expressions, and sounds" refers to the physical movements, facial expressions, and sounds made by pets.
[1540] The "terminal" is a device equipped with a camera and microphone that collects your pet's movements, facial expressions, and cries.
[1541] A "server" is a computer system that receives, analyzes, and processes data sent from a terminal.
[1542] An "AI model" is an artificial intelligence that uses machine learning algorithms to analyze data and estimate a pet's emotions and desires.
[1543] "Estimation" refers to predicting a pet's emotions and desires from the data obtained.
[1544] "Translation into natural language" means converting the inferred results output by the AI model into a message that is easy for the user to understand.
[1545] "Notifying the user" means conveying messages or information to the user through the terminal.
[1546] A "security protocol" is a communication protocol for maintaining the security of data when it is transmitted from a terminal to a server.
[1547] A "computer vision algorithm" is an algorithm that analyzes video data and identifies pet movements and facial expressions.
[1548] A "voice recognition algorithm" is an algorithm that analyzes voice data and identifies the sounds of pets.
[1549] "Selecting appropriate content" means selecting videos, music, information, etc. that are suitable for your pet based on the pet's estimated emotions and desires.
[1550] "Delivery" means transmitting selected content to a user terminal and having it played or displayed.
[1551] System configuration overview
[1552] This invention provides a system that collects pet movements, facial expressions, and cries, estimates the pet's emotions and desires based on the collected data, and delivers appropriate content to the user. The system mainly consists of the following components:
[1553] A "terminal" for collecting your pet's movements, facial expressions, and cries.
[1554] The "server" receives and analyzes data collected from the terminal.
[1555] An "AI model" that analyzes the received data and estimates the pet's emotions and desires.
[1556] A "notification system" that selects appropriate content based on estimated emotions and desires and delivers it to the user's device.
[1557] Data collection details
[1558] The "terminal" is equipped with a camera and microphone to collect the pet's movements, facial expressions, and sounds in real time, and the collected data is divided into data packets at regular intervals and sent to the "server" using a security protocol.
[1559] Data reception and analysis
[1560] The server receives and decodes the data packets sent from the device. It then processes the video and audio data separately. The video data uses computer vision algorithms to identify the pet's facial expressions and movements, while the audio data uses speech recognition algorithms to identify the pattern, pitch, and frequency of the pet's cries.
[1561] Estimating emotions and desires
[1562] The server inputs the analyzed features into an AI model to estimate the pet's emotions and desires. The AI model uses a trained machine learning algorithm, and can infer, for example, that a pet's desire to play is indicated by a wagging tail and high-pitched meow.
[1563] Natural language translation and notification
[1564] The estimated emotions and desires are translated into natural language and converted into specific messages (e.g., "I'm hungry" or "Play with a ball!"). The translated messages are then sent from the server to the device and notified to the user.
[1565] Recognizing user emotions and proposing countermeasures
[1566] The device also collects the user's voice and facial expressions and sends them to the server. The emotion engine analyzes this data and recognizes the user's emotions. For example, it uses voice recognition and facial expression analysis algorithms to identify emotions from the user's tone and facial expressions.
[1567] Proposing solutions based on user sentiment
[1568] The server generates a message suggesting appropriate actions for the pet based on the user's emotional information. For example, if the user is tired, the server might suggest, "Take your time to feed your pet and don't rush." The suggested message is displayed on the user's device, encouraging the user to take appropriate action.
[1569] Hardware and software used
[1570] The system uses the following hardware and software:
[1571] Hardware: Camera, microphone, smartphone, smart glasses, head-mounted display.
[1572] Software: OpenCV (used to collect video data from the camera), TensorFlow (used to analyze pet and user emotions), requests (used to communicate data with the server), virtual voice analysis module and emotion recognition module.
[1573] Specific example explanation
[1574] For example, if it is estimated that the pet wants to play, it will be treated as "playful" and the video content "play_video.mp4" will be selected and played. If the user is relaxed, the message "Enjoy a relaxing session with your pet" will be displayed, suggesting a better way to spend time with the pet.
[1575] Prompt Sentence Examples
[1576] "Suggest content suitable for pets who want to play."
[1577] "Suggest actions that are appropriate for when the user is relaxed."
[1578] In this way, a system is realized that can analyze emotions and desires based on real-time data from pets and users, and deliver optimal responses and content.
[1579] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[1580] Step 1:
[1581] The device uses a camera and microphone to collect the pet's movements, facial expressions, and sounds in real time. It receives camera video and audio data as input and divides it into data packets at regular intervals. This captures multiple features, such as the pet's movements, facial expressions, and sounds, which can then be sent to subsequent analysis steps.
[1582] Step 2:
[1583] The terminal sends the collected data packets to the server using a security protocol. The terminal receives the data packets as input, encrypts the data using a security protocol (e.g., SSL / TLS), and securely transmits the data to the server. This ensures the confidentiality of the data.
[1584] Step 3:
[1585] The server receives data packets sent from the terminal. It receives encrypted data packets as input and first decrypts them. As output, it obtains decrypted video and audio data. This allows the data to be stored in an analyzable format on the server.
[1586] Step 4:
[1587] The server analyzes the video data to identify the pet's movements and facial expressions. It receives the decoded video data as input and analyzes it using a computer vision algorithm (e.g., OpenCV). The output is identified behavior and facial expression data, such as tail movements and facial expressions. This allows the pet's movement and facial expression features to be extracted.
[1588] Step 5:
[1589] The server analyzes the audio data to identify the pattern, pitch, and frequency of the pet's cries. It receives the decoded audio data as input and analyzes it using a speech recognition algorithm (e.g., a speech analysis module). The output is an analysis result that identifies the pattern, pitch, and frequency of the pet's cries. This allows the characteristics of the pet's cries to be extracted.
[1590] Step 6:
[1591] The server inputs the analyzed features into an AI model to estimate the pet's emotions and desires. The analysis results of movements, facial expressions, and cries are received as input and input into an AI model (for example, a model built with TensorFlow). The output is the pet's estimated emotions and desires (for example, "I want to play" or "I'm hungry"). This clarifies the pet's internal state.
[1592] Step 7:
[1593] The server translates the estimated emotions and desires into natural language and generates a message. It receives the estimated results from the AI model as input and translates them using a natural language processing engine. The output is a message that is easy for the user to understand, such as "I'm hungry" or "Play with my ball!" This communicates the pet's emotions and desires to the user.
[1594] Step 8:
[1595] The server selects appropriate content based on the estimated emotions and desires. It receives a message translated into natural language as input and uses an algorithm to select appropriate content (e.g., video, music). Specific content files (e.g., "play_video.mp4", "relaxing_music.mp3") are obtained as output. This allows content that matches the pet's state to be selected.
[1596] Step 9:
[1597] The server delivers the selected content to the user terminal. The server receives the selected content file as input and transmits it to the user terminal via the Internet. As output, the server delivers the content to be played or displayed on the user terminal. This allows the user to receive content that is appropriate for the condition of their pet.
[1598] Step 10:
[1599] The device collects the user's facial expressions and voice and sends them to the server. It receives camera footage and voice data as input, divides them into data packets, and sends them to the server. This provides data for analyzing the user's emotions.
[1600] Step 11:
[1601] The server analyzes the user's facial expressions and voice to recognize the user's emotions. It receives the user's facial expressions and voice data as input and analyzes them using an emotion recognition algorithm. The output identifies the user's emotions (e.g., "relaxed" or "tired"). This clarifies the user's state.
[1602] Step 12:
[1603] The server generates a message suggesting how to respond to a pet based on the user's emotional information. It receives the user's emotion recognition results as input and generates the suggested message using a natural language processing engine. The output is specific suggested actions such as "Relax with your pet" or "Take your time to feed your pet without rushing." This provides guidance for the user to take appropriate action.
[1604] The above is a description of the specific processing steps of the system for implementing the present invention.
[1605] The specific processing unit 290 transmits the result of the specific processing to the headset type terminal 314. In the headset type terminal 314, the control unit 46A causes the speaker 240 and the display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[1606] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[1607] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the headset type terminal 314.
[1608] [Fourth embodiment]
[1609] FIG. 7 shows an example of the configuration of a data processing system 410 according to the fourth embodiment.
[1610] 7, a data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.
[1611] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[1612] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a control target 443. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the control target 443 are also connected to the bus 52.
[1613] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[1614] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[1615] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[1616] The control object 443 includes a display device, LEDs in the eyes, and motors for driving the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the emotions of the robot 414 can be expressed by controlling these motors. In addition, the facial expressions of the robot 414 can also be expressed by controlling the light emission state of the LEDs in the eyes of the robot 414.
[1617] Fig. 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Fig. 8, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[1618] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[1619] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[1620] In the robot 414, the processor 46 performs the reception output process. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[1621] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1622] System configuration overview
[1623] The present invention is a system that includes a "terminal" for collecting pet movements, facial expressions, and cries, a "server" for receiving data sent from the terminal, an "AI model" for analyzing the received data and inferring the pet's emotions and desires, and a "notification system" that translates the inferred emotions and desires into natural language and notifies the user.
[1624] Data collection details
[1625] The "terminal" is equipped with a camera and microphone, and collects the pet's movements, facial expressions, and sounds in real time. The collected data is divided into data packets at regular intervals and sent from the "terminal" to the "server." A security protocol is used to ensure the data is secure.
[1626] Data reception and analysis
[1627] The "server" receives the data packets sent from the terminal. The received data is first decoded and then processed separately into video data and audio data.
[1628] The "server" analyzes the video data using computer vision algorithms to identify the pet's facial expressions and movements (for example, tail movements). It also analyzes the audio data using speech recognition algorithms to identify the pattern, pitch, and frequency of the pet's cries. This allows the data's features to be extracted.
[1629] Estimating emotions and desires
[1630] The extracted features are input into an "AI model," which uses a pre-trained machine learning algorithm to infer the pet's emotions and desires. For example, if a pet frequently wags its tail and makes a high-pitched meow, it can be inferred to have a desire to play.
[1631] Natural language translation and notification
[1632] The estimated emotions and desires are translated into natural language. Specifically, the output of the AI model is converted into messages such as "I'm hungry" or "Play with the ball!". This translated message is then resent from the "server" to the "device," which then notifies the user.
[1633] Specific example explanation
[1634] Example 1: System behavior when a dog wants to play
[1635] 1. The "terminal" captures the dog's movements with a camera and collects the dog's tail wagging vigorously and high-pitched barking.
[1636] 2. The "terminal" sends the collected data to the "server."
[1637] 3. The "server" receives the data and identifies tail movements from the video data and high-pitched cries from the audio data.
[1638] 4. The “server” inputs these features into the “AI model.”
[1639] 5. The AI model uses these characteristics to infer that the dog wants to play.
[1640] 6. The "server" translates the result "I want to play" into natural language and makes it "Play with the ball!"
[1641] 7. The "terminal" notifies the user with the message "Play with the ball!"
[1642] Example 2: System behavior when the cat is hungry
[1643] 1. The "terminal" collects the cat's meowing (a low, sustained sound) and its movements around the dish.
[1644] 2. The "terminal" sends the collected data to the "server."
[1645] 3. The "server" receives the data and identifies the location from the video data (walking around the dish) and the sound from the audio data (a low, sustained sound).
[1646] 4. The “server” inputs these features into the “AI model.”
[1647] 5. The "AI model" uses these characteristics to infer that the cat is hungry.
[1648] 6. The "server" translates the result "hungry" into natural language and sets it as "I'm hungry."
[1649] 7. The "terminal" notifies the user with the message "I'm hungry."
[1650] In this way, the present invention can be implemented as a system that allows users to understand their pet's specific emotions and desires in real time and take appropriate action.
[1651] The processing flow will be explained below.
[1652] Step 1:
[1653] The device uses a camera and microphone to collect the pet's movements, facial expressions, and sounds in real time. The collected data is then separated into video data and audio data, and each is split into separate data packets.
[1654] Step 2:
[1655] The device sends the collected data packets to a server, where a security protocol is used to ensure the data is protected.
[1656] Step 3:
[1657] The server receives the data packets sent from the terminal. The received data is first decoded and then separated into video data and audio data for processing.
[1658] Step 4:
[1659] The server runs the video data through computer vision algorithms to analyze your pet's facial expressions and movements, specifically identifying tail movements, facial expressions, and overall body movements.
[1660] Step 5:
[1661] The server runs the audio data through a speech recognition algorithm to identify the pattern, pitch, and frequency of your pet's cries, including the pitch of the cries and the repetition of the cries.
[1662] Step 6:
[1663] The server combines the results of the video and audio analysis to extract features to estimate the pet's emotions and desires, such as high-pitched meows and tail wagging.
[1664] Step 7:
[1665] The server inputs the extracted features into a pre-trained AI model, which uses patterns learned from past data to infer the pet's emotions and desires.
[1666] Step 8:
[1667] The server receives the inferences output by the AI model and maps them to a list of predefined emotions and desires, such as "hunger," "want to play," "anxiety," and "relaxation."
[1668] Step 9:
[1669] The server translates the estimated emotions and desires into natural language. For example, "hungry" becomes "I'm hungry" and "I want to play" becomes "Play with my ball!"
[1670] Step 10:
[1671] The server then sends the translated message to the user's device, again using security protocols to protect the data.
[1672] Step 11:
[1673] The terminal receives the translated message sent from the server, and the received message is formatted to be displayed on the user interface.
[1674] Step 12:
[1675] The device will display messages to the user that indicate the pet's emotions and needs, such as "I'm hungry!" or "Play with my ball!", displayed as pop-ups or notification banners.
[1676] Step 13:
[1677] The user checks the messages displayed on the device and takes action according to the pet's request, such as feeding the pet or playing with a ball.
[1678] This series of processes allows users to understand their pet's specific emotions and desires in real time and respond appropriately.
[1679] Example 1
[1680] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1681] Conventional pet management systems have had problems with accurately monitoring pet behavior and conditions, and owners are unable to properly understand their pets' emotions and needs, which increases stress for pets and burdens on owners. Furthermore, insufficient coordination in data collection, analysis, and notification makes it difficult to understand pets' conditions in real time.
[1682] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[1683] In this invention, the server includes means for decoding received data and processing the data separately into video data and audio data, means for analyzing the video data using a computer vision algorithm and the audio data using a speech recognition algorithm, and means for inputting the extracted features into an AI model to estimate emotions and desires. This makes it possible to instantly estimate emotions and desires from the pet's movements and cries, and notify the owner in natural language.
[1684] The "terminal" is a device equipped with a camera and microphone to collect the movements, facial expressions, and sounds of your pet.
[1685] A "server" is a computer system that receives, decrypts, and analyzes data sent from a terminal.
[1686] A "data packet" is a unit of information that divides collected data and is sent from a terminal to a server.
[1687] "Decoding" is the process of restoring encoded data to its original form.
[1688] "Video data" refers to data that visually records the movements and expressions of pets collected by a camera.
[1689] "Audio data" refers to data that records the sounds of pets as collected by a microphone.
[1690] A "computer vision algorithm" is a technology for analyzing video data and recognizing objects and actions.
[1691] A "voice recognition algorithm" is a technology for analyzing voice data and recognizing sounds and patterns.
[1692] "Features" are important information extracted from video and audio data for estimating a pet's emotions and desires.
[1693] The "AI model" is a system that is pre-trained using a machine learning algorithm and estimates a pet's emotions and desires based on input features.
[1694] "Natural language" refers to language that humans use on a daily basis, and is a language that expresses estimated emotions and desires in a way that is easy for users to understand.
[1695] A "security protocol" is a set of rules and procedures for securely communicating data.
[1696] "Notification" is the act of informing the user of the results of estimated emotions and desires.
[1697] MODE FOR CARRYING OUT THE INVENTION
[1698] The present invention is a system that includes a terminal that collects pet movements, facial expressions, and cries, a server that receives and analyzes the data, and a system that notifies the user of the estimated emotions and desires. Specifically, by using a pet monitoring device, a data analysis system, and a communication means, it is possible to grasp the pet's condition in real time and take appropriate measures.
[1699] System configuration overview
[1700] The present invention includes a "terminal" for collecting pet movements, facial expressions, and cries, a "server" for receiving data sent from the terminal, an "AI model" for analyzing the received data and inferring the pet's emotions and desires, and a "notification system" for translating the inferred emotions and desires into natural language and notifying the user.
[1701] Data collection details
[1702] The "terminal" is equipped with a camera and microphone, and collects the pet's movements, facial expressions, and cries in real time. For example, the camera records the pet's tail movements and eye expressions as video data, while the microphone collects the pet's cries as audio data. The collected data is divided into data packets at regular intervals and sent from the "terminal" to the "server." A security protocol is used for data transmission to ensure the safety of the data.
[1703] Data reception and analysis
[1704] The "server" receives data packets sent from the device. The received data is first decoded and then processed separately into video data and audio data. The "server" analyzes the video data using a computer vision algorithm (e.g., OpenCV) to identify the pet's facial expressions and movements. It also analyzes the audio data using a speech recognition algorithm (e.g., a TensorFlow or PyTorch-based model) to identify the pattern, pitch, and frequency of the pet's cries. This allows the data features to be extracted.
[1705] Estimating emotions and desires
[1706] The extracted features are input into an "AI model." The AI model uses a pre-trained machine learning algorithm (e.g., a generative AI model) to infer the pet's emotions and desires. For example, if the pet frequently wags its tail and makes a high-pitched meow, it can be inferred that the pet wants to play.
[1707] Natural language translation and notification
[1708] The estimated emotions and desires are translated into natural language. Specifically, the results output by the AI model are converted into messages such as "I'm hungry" or "Play with the ball!". These translated messages are then resent from the "server" to the "device," which then notifies the user. For example, the user can be notified by a push notification or voice alert on a smartphone app.
[1709] Specific example explanation
[1710] Example 1: System behavior when a dog wants to play
[1711] 1. The device captures the dog's movements with a camera and collects the dog's vigorously wagging tail and high-pitched barking sounds.
[1712] 2. The device sends the collected data to the server.
[1713] 3. The server receives the data and identifies tail movements from the video data and high-pitched cries from the audio data.
[1714] 4. The server inputs these features into the AI model.
[1715] 5. The AI model uses these characteristics to infer that the dog wants to play.
[1716] 6. The server translates the result "I want to play" into natural language and makes it "Play with the ball!"
[1717] 7. The device notifies the user with the message "Play with the ball!"
[1718] Example 2: System behavior when the cat is hungry
[1719] 1. The device collects the cat's meowing (a low, sustained sound) and its movements around the dish.
[1720] 2. The device sends the collected data to the server.
[1721] 3. The server receives the data and identifies the location from the video data (walking around the dish) and the sound from the audio data (a low, sustained sound).
[1722] 4. The server inputs these features into the AI model.
[1723] 5. The AI model uses these characteristics to infer that the cat is hungry.
[1724] 6. The server translates the result "hungry" into natural language and gives it the value "I'm hungry."
[1725] 7. The device notifies the user with the message "I'm hungry."
[1726] Example prompts to input to the generative AI model
[1727] 1. Example prompt for identifying a dog's desire to play: "When a dog wags its tail frequently and makes a high-pitched whine, it means it wants to play. Express this in natural language."
[1728] 2. Example prompt for identifying a cat's hunger state: "When a cat meows in a low, sustained tone and paces around its food bowl, it is likely hungry. Please describe this in natural language."
[1729] In this way, the present invention can be implemented as a system that allows users to understand their pet's specific emotions and desires in real time and take appropriate action.
[1730] The flow of the identification process in the first embodiment will be described with reference to FIG.
[1731] Program processing steps
[1732] Step 1: Collect data from the device
[1733] The device uses a camera and microphone to collect your pet's movements, facial expressions, and cries in real time.
[1734] Input: Real-time pet movements, facial expressions, and sounds
[1735] Output: Collected video and audio data
[1736] Specific behavior:
[1737] The camera records your pet's tail movements and facial expressions as video data, while the microphone captures your pet's barks as audio data, such as when your dog wags its tail or barks in a high-pitched voice.
[1738] Step 2: Send data from the device to the server
[1739] The device divides the collected data into data packets and sends them to the server using a security protocol.
[1740] Input: Collected video and audio data
[1741] Output: Data packets sent from the device to the server
[1742] Specific behavior:
[1743] The collected data is divided into packets of a certain size, encrypted using the TLS protocol, and sent to the server.
[1744] Step 3: Server receives and decrypts data
[1745] The server receives the data packets sent from the terminal, decodes them, and separates them into video data and audio data.
[1746] Input: Data packets sent from the device
[1747] Output: Decoded video and audio data
[1748] Specific behavior:
[1749] The received data packets are decrypted using the TLS protocol and separated into collected video and audio data, for example, tail movements recorded by a camera or barks recorded by a microphone.
[1750] Step 4: Video data analysis by the server
[1751] The server analyzes the video data using computer vision algorithms to identify the pet's facial expressions and movements.
[1752] Input: Decoded video data
[1753] Output: Features of pet's facial expressions and movements
[1754] Specific behavior:
[1755] Using OpenCV, features such as tail wagging and eye movement are extracted from video data. For example, if a cat is wagging its tail vigorously, that movement can be identified.
[1756] Step 5: Server analyzes the audio data
[1757] The server analyzes the audio data using a voice recognition algorithm to identify the pattern, pitch and frequency of your pet's cries.
[1758] Input: Decoded audio data
[1759] Output: Features of pet sounds
[1760] Specific behavior:
[1761] Use TensorFlow or PyTorch to identify high-pitched meows and patterns in audio data, such as detecting the low, sustained meow of a cat.
[1762] Step 6: Feature extraction by the server
[1763] The server integrates the features extracted from the video and audio data and inputs them into the AI model.
[1764] Input: Features of pet's facial expressions and movements, features of pet's cries
[1765] Output: Integrated features
[1766] Specific behavior:
[1767] The identified tail movements from the video data and the call patterns from the audio data are integrated into a single dataset.
[1768] Step 7: Server inputs AI model
[1769] The server inputs the extracted features into an AI model to estimate emotions and desires.
[1770] Input: Integrated features
[1771] Output: Estimated pet's emotions and desires
[1772] Specific behavior:
[1773] The integrated features are input into a generative AI model to obtain an inference. For example, if a dog is wagging its tail and making a high-pitched bark, it may infer that it wants to play.
[1774] Step 8: Server-based translation into natural language
[1775] The server translates the estimated emotions and desires into natural language.
[1776] Input: Inference results from the AI model
[1777] Output: A message expressed in natural language
[1778] Specific behavior:
[1779] The result "I want to play" obtained from the AI model is converted into the message "Play with the ball!"
[1780] Step 9: Sending translation results from the server to the device
[1781] The server retransmits the translated message to the terminal.
[1782] Input: A message expressed in natural language
[1783] Output: Messages sent to the terminal
[1784] Specific behavior:
[1785] Using security protocols, send a translated message to the device, for example, "Play with the ball!"
[1786] Step 10: Notify the user via the device
[1787] The terminal receives the message sent from the server and notifies the user of the message.
[1788] Input: The message sent from the server
[1789] Output: The notification the user receives
[1790] Specific behavior:
[1791] The device uses push notifications and audio alerts on a smartphone app to notify users with messages like, "Play with the ball!"
[1792] Step 11: User Action
[1793] The user checks the notification from the device and responds to the pet.
[1794] Input: Notifications from your device
[1795] Output: Specific responses to pets
[1796] Specific behavior:
[1797] The user checks the notification and takes action based on the message, such as playing with their pet. For example, if they receive the message "Play with a ball!", they will get a ball and start playing with their dog.
[1798] The above is a description of the specific processing steps in the system program.
[1799] (Application example 1)
[1800] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1801] In recent years, there has been a demand for technology that can understand pet emotions and desires and enable owners to provide appropriate care. In particular, there is a need for a system that can grasp what pets want in real time and purchase appropriate products based on that information. However, current technology has difficulty accurately estimating pet emotions and desires and making product purchase suggestions based on those estimations. Therefore, the challenge is to provide a system that can analyze pet emotions and desires in real time and automatically make product purchase suggestions.
[1802] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[1803] In this invention, the server includes a terminal for collecting the movements, expressions, and cries of the pet, a means for receiving data transmitted from the terminal, a means including an AI model for analyzing the received data and estimating the pet's emotions and desires, a means for translating the estimated emotions and desires into natural language and notifying the user, and a means for automatically generating product purchase suggestions based on the estimated emotions and desires and notifying the user. This makes it possible to grasp the pet's emotions and desires in real time and automatically suggest appropriate products.
[1804] The "terminal" is a device that collects your pet's movements, facial expressions, and sounds in real time, and includes a camera and microphone.
[1805] The "server" is a computer that receives data sent from the terminal, analyzes it, and infers the pet's emotions and desires.
[1806] An "AI model" is an analytical model that uses machine learning algorithms to estimate a pet's emotions and desires.
[1807] "Means of translating into natural language" refers to the function of converting the inferred results output by the AI model into words that are easy for users to understand, such as "I'm hungry" or "I want to play."
[1808] "Means for generating product purchase suggestions" refers to a function that selects appropriate products (e.g., pet food or toys) based on estimated emotions and desires and notifies the user of the suggestions.
[1809] System Overview
[1810] This invention is a system that collects and analyzes pet movements, facial expressions, and cries in real time to estimate the pet's emotions and desires and make appropriate product purchase suggestions to the user. It is composed of the following components.
[1811] Key Components
[1812] 1. Device:
[1813] The terminal is a device equipped with a camera and microphone that captures your pet's movements, facial expressions, and sounds in real time, and the data is then sent to a server using a security protocol.
[1814] 2. Server:
[1815] The server receives the data sent from the device and analyzes it to infer the pet's emotions and needs. The analysis is performed using an AI model (machine learning algorithm). The server also translates the analysis results into natural language that is easy for the user to understand, and then generates appropriate product purchase suggestions.
[1816] 3. AI model:
[1817] The AI model, which uses the collected data to infer your pet's emotions and needs, includes computer vision and speech recognition algorithms.
[1818] 4. User Notification System:
[1819] The system notifies the user of the estimated emotions and desires, along with product purchase suggestions based on those emotions and desires. The user receives these notifications via their smartphone.
[1820] Software and Hardware Configuration
[1821] Hardware:
[1822] Camera: Capture videos and photos of your pet.
[1823] Microphone: Collects pet sounds.
[1824] Server: A computer that performs data analysis.
[1825] software:
[1826] OpenCV: Used for preprocessing video data.
[1827] Keras: Used to load and use AI models.
[1828] sounddevice: Used to collect audio data.
[1829] Requests: Used to send requests to the notification service.
[1830] How it works
[1831] The server receives the video and audio data sent from the device at runtime and pre-processes the received data appropriately. The video data is resized and normalized using OpenCV before being input to the AI model. The audio data is similarly pre-processed using the sounddevice and then input to the AI model.
[1832] The AI model uses this preprocessed data to infer the pet's emotions and desires. The inference results are translated into natural language and converted into user-friendly expressions such as "I'm hungry" or "I want to play." Based on this translation, the model then generates appropriate product purchase suggestions (e.g., pet food or toys).
[1833] Specific examples of implementation
[1834] Example 1: When a dog wants to play
[1835] 1. The device captures the dog's movements with a camera and collects the dog's vigorously wagging tail and high-pitched barking sounds.
[1836] 2. The collected data is sent to the server and separated into video data and audio data.
[1837] 3. The video and audio data are preprocessed and then input into the AI model.
[1838] 4. The AI model analyzes the data and infers that the dog wants to play.
[1839] 5. The server generates and sends the notification "Play with the ball!" and the suggestion "Would you like to buy a toy?" to the user.
[1840] Example prompt sentence:
[1841] "Write a Python program that monitors your pet's behavior and generates a notification saying 'Would you like to buy pet food?' if it detects that your pet is hungry."
[1842] In this way, a system can be provided for understanding a pet's emotions and desires in real time.
[1843] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[1844] Step 1:
[1845] The device uses a camera and microphone to collect the pet's movements, facial expressions, and cries in real time. Specifically, the camera captures the pet's movements as video, and the microphone records the pet's cries as audio data. These data are collected. The input is real-time video and audio data, and the output is the collected video and audio data.
[1846] Step 2:
[1847] The terminal divides the collected video and audio data into data packets at regular intervals and sends them to the server using a secure protocol. Specifically, the data is encrypted and sent using the HTTPS protocol. The input is the collected video and audio data, and the output is the encrypted data packets.
[1848] Step 3:
[1849] The server receives the data packets sent from the terminal and decrypts them. First, it receives the data, decrypts it, and returns it to the original video and audio data. The input is the encrypted data packet, and the output is the decrypted video and audio data.
[1850] Step 4:
[1851] The server analyzes the decoded data using computer vision and speech recognition algorithms. Specifically, it uses OpenCV to resize and normalize the video data. The audio data is preprocessed to extract audio features. The input is the decoded video and audio data, and the output is the preprocessed video and audio data.
[1852] Step 5:
[1853] The server inputs the preprocessed video and audio data into an AI model to estimate the pet's emotions and desires. Specifically, the data is input into a machine learning model trained using Keras to obtain an inference result. The input is the preprocessed video and audio data, and the output is the inference result of the pet's emotions and desires.
[1854] Step 6:
[1855] The server translates the inferences obtained from the AI model into natural language. For example, it converts them into messages such as "I want to play" or "I'm hungry." The input is the inferences from the AI model, and the output is a message in natural language.
[1856] Step 7:
[1857] The server generates appropriate product purchase suggestions based on the translated message, such as "Would you like to buy a toy?" or "Would you like to buy pet food?" The input is the translated message, and the output is the product purchase suggestion message.
[1858] Step 8:
[1859] The terminal receives the product purchase suggestion message sent from the server and notifies the user. The user can confirm the suggestion through a notification on their smartphone. The input is the product purchase suggestion message, and the output is a notification to the user.
[1860] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[1861] System configuration overview
[1862] The present invention is a system that includes a "terminal" for collecting the movements, facial expressions, and cries of a pet, a "server" for receiving data sent from the terminal, an "AI model" for analyzing the received data and inferring the pet's emotions and desires, a "notification system" for translating the inferred emotions and desires into natural language and notifying the user, and an "emotion engine" for recognizing the user's emotions.
[1863] Data collection details
[1864] The "terminal" is equipped with a camera and microphone, and collects the pet's movements, facial expressions, and sounds in real time. The collected data is divided into data packets at regular intervals and sent from the "terminal" to the "server." A security protocol is used to ensure the data is secure.
[1865] Data reception and analysis
[1866] The "server" receives the data packets sent from the terminal. The received data is first decoded and then processed separately into video data and audio data.
[1867] The "server" analyzes the video data using computer vision algorithms to identify the pet's facial expressions and movements (for example, tail movements). It also analyzes the audio data using speech recognition algorithms to identify the pattern, pitch, and frequency of the pet's cries. This allows the data's features to be extracted.
[1868] Estimating emotions and desires
[1869] The extracted features are input into an "AI model," which uses a pre-trained machine learning algorithm to infer the pet's emotions and desires. For example, if the pet wags its tail and makes a high-pitched meow, it can be inferred to be wanting to play.
[1870] Natural language translation and notification
[1871] The estimated emotions and desires are translated into natural language. Specifically, the output of the AI model is converted into messages such as "I'm hungry" or "Play with the ball!" These translated messages are then resent from the "server" to the "device" and notified to the user via the "device."
[1872] Recognizing user emotions and proposing countermeasures
[1873] In addition, the user's voice and facial expressions are also collected from the "terminal" and sent to the "server." The "emotion engine" analyzes this data to recognize the user's emotions. Using voice recognition algorithms and facial expression analysis algorithms, it identifies emotions from the user's voice tone and facial expressions.
[1874] Proposing solutions based on user sentiment
[1875] The server then uses the user's emotional information to suggest ways to help the pet. For example, if the user is tired, the server suggests creating a relaxing atmosphere for the pet.
[1876] Specific example explanation
[1877] Example 1: System behavior when a dog wants to play
[1878] 1. The "terminal" captures the dog's movements with a camera and collects the dog's tail wagging vigorously and high-pitched barking.
[1879] 2. The "terminal" sends the collected data to the "server."
[1880] 3. The "server" receives the data and identifies tail movements from the video data and high-pitched cries from the audio data.
[1881] 4. The “server” inputs these features into the “AI model.”
[1882] 5. The AI model uses these characteristics to infer that the dog wants to play.
[1883] 6. The "server" translates the result "I want to play" into natural language and makes it "Play with the ball!"
[1884] 7. The "terminal" notifies the user with the message "Play with the ball!"
[1885] 8. The "terminal" also collects the user's facial expressions and voice and sends them to the "server."
[1886] 9. The "Emotion Engine" recognizes the user's emotion as "Relaxed."
[1887] 10. The "server" checks the user's relaxation state and sends a message suggesting "Relax with your pet."
[1888] 11. The Terminal displays the suggestion message to the user.
[1889] Example 2: System behavior when the cat is hungry
[1890] 1. The "terminal" collects the cat's meowing (a low, sustained sound) and its movements around the dish.
[1891] 2. The "terminal" sends the collected data to the "server."
[1892] 3. The "server" receives the data and identifies the location from the video data (walking around the dish) and the sound from the audio data (a low, sustained sound).
[1893] 4. The “server” inputs these features into the “AI model.”
[1894] 5. The "AI model" uses these characteristics to infer that the cat is hungry.
[1895] 6. The "server" translates the result "hungry" into natural language and sets it as "I'm hungry."
[1896] 7. The "terminal" notifies the user with the message "I'm hungry."
[1897] 8. The "terminal" collects the user's facial expressions and voice and sends them to the "server."
[1898] 9. The "emotion engine" recognizes the user's emotion as "tired."
[1899] 10. The "server" checks the user's tiredness and sends a message suggesting that they take their time to feed their pet.
[1900] 11. The Terminal displays the suggestion message to the user.
[1901] In this way, the present invention can be implemented as a system that allows the user to understand the specific emotions and desires of their pet in real time and take appropriate action according to their own emotions and state.
[1902] The processing flow will be explained below.
[1903] Step 1:
[1904] The device uses a camera and microphone to collect your pet's movements, facial expressions, and sounds in real time, and the collected data is divided into video data and audio data, each of which is then split into separate data packets.
[1905] Step 2:
[1906] The device sends the collected data packets to the server, using a security protocol to ensure the data is protected.
[1907] Step 3:
[1908] The server receives the data packets sent from the terminal. The received data is first decoded and then separated into video data and audio data for processing.
[1909] Step 4:
[1910] The server runs the video data through computer vision algorithms to analyze your pet's facial expressions and movements, specifically identifying tail movements, facial expressions, and overall body movements.
[1911] Step 5:
[1912] The server runs the audio data through a speech recognition algorithm to identify the pattern, pitch, and frequency of your pet's cries, including the pitch and repetition of the cries.
[1913] Step 6:
[1914] The server combines the results of the video and audio analysis to extract features to estimate the pet's emotions and desires, such as high-pitched meows and tail wagging.
[1915] Step 7:
[1916] The server inputs the extracted features into a pre-trained AI model, which uses patterns learned from past data to infer the pet's emotions and desires.
[1917] Step 8:
[1918] The server receives the inferences output by the AI model and maps them to a list of predefined emotions and desires, such as "hunger," "want to play," "anxiety," and "relaxation."
[1919] Step 9:
[1920] The server translates the estimated emotions and desires into natural language, for example, translating "hungry" into "I'm hungry" and "I want to play" into "Play with my ball!"
[1921] Step 10:
[1922] The server then sends the translated message to the user's device, again using security protocols to protect the data.
[1923] Step 11:
[1924] The terminal receives the translated message sent from the server, and the received message is formatted to be displayed on the user interface.
[1925] Step 12:
[1926] The device will display messages to the user that indicate the pet's emotions and needs, such as "I'm hungry!" or "Play with my ball!", displayed as pop-ups or notification banners.
[1927] Step 13:
[1928] The device also uses a camera and microphone to collect the user's voice and facial expressions, which are then split into data packets and sent to the server.
[1929] Step 14:
[1930] The server receives the user's data packets sent from the terminal, decodes the received data, and prepares them for processing by separating them into video data and audio data.
[1931] Step 15:
[1932] The server runs the video data through a facial expression analysis algorithm to analyze the user's facial expressions, and the audio data through a voice recognition algorithm to analyze the user's tone and words.
[1933] Step 16:
[1934] The server integrates the analysis results of the facial expression data and voice data and extracts features to identify the user's emotions.
[1935] Step 17:
[1936] The server inputs the extracted features into an emotion engine to estimate the user's emotion, such as "tired," "relaxed," or "excited."
[1937] Step 18:
[1938] The server generates a message suggesting how to handle pets based on the user's emotions. For example, if the user is tired, it suggests, "Take your time and don't rush, take your time to feed your pet."
[1939] Step 19:
[1940] The server sends a proposal message to the user's device, which is formatted for display through the user interface.
[1941] Step 20:
[1942] The device will notify the user with suggestion messages, such as "Take your time to feed your pet and don't rush," via a pop-up or notification banner.
[1943] In this way, the present invention can be implemented as a system that allows the user to understand the specific emotions and desires of their pet in real time and take appropriate action according to their own emotions and state.
[1944] Example 2
[1945] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1946] It is difficult to understand the emotions and desires of both pets and users in real time and propose appropriate responses. Conventional technologies lack systems that can accurately estimate a pet's emotions and desires, translate them into natural language, and notify the user. Furthermore, there is no system that can recognize the user's emotions and propose appropriate responses based on them. This leads to a decline in the quality of pet care and a lack of communication between pets and users.
[1947] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[1948] In this invention, the server includes means for receiving data transmitted from the terminal, decoding and classifying the data, means for analyzing video data using a computer vision algorithm, and means for analyzing audio data using a speech recognition algorithm. This allows features of the pet's movements and cries to be extracted and input into an AI model to estimate the pet's emotions and needs. Furthermore, the estimation results are translated into natural language and notified to the user, allowing for an accurate understanding of the pet's situation. Furthermore, the server collects and analyzes the user's voice and facial expressions, and proposes appropriate countermeasures based on the user's emotions, thereby improving communication between the pet and the user.
[1949] The "terminal" is a device equipped with a camera and microphone to collect your pet's movements, facial expressions, and sounds.
[1950] A "server" is a computer system that receives data sent from a terminal and decodes, classifies, and analyzes the data.
[1951] A "data packet" is a unit for dividing and transmitting data collected by a terminal at regular intervals.
[1952] A "security protocol" is a communication protocol used to ensure security when transmitting data.
[1953] "Video data" refers to image data captured by a camera showing the movements and expressions of a pet.
[1954] "Audio data" is sound data recorded with a microphone of a pet's cries.
[1955] A "computer vision algorithm" is an image processing technology used to analyze specific movements and facial expressions from video data.
[1956] A "voice recognition algorithm" is a technology for analyzing specific sound patterns and characteristics from audio data.
[1957] "Features" are important data points extracted from the analyzed data that indicate the pet's emotions and behavior.
[1958] An "AI model" is a computational model that uses machine learning algorithms to estimate a pet's emotions and desires.
[1959] "Natural language translation" is the process of converting the results estimated by an AI model into a language that a user can understand.
[1960] "Notification" refers to the act of conveying a translated message to the user, and is primarily done via the terminal.
[1961] "User data" is data that collects information such as the user's voice and facial expressions.
[1962] An "emotion engine" is a system for analyzing user data to identify a user's emotions.
[1963] "Countermeasure suggestion" is a process of suggesting actions that the user should take toward their pet based on the identified user emotions.
[1964] The present invention is a system that analyzes the movements, facial expressions, and cries of a pet, estimates the feelings and desires of the pet, and notifies the user of the results. Detailed embodiments of the system will be described below.
[1965] System configuration
[1966] The system includes a "terminal" that collects the pet's movements, facial expressions, and cries, a "server" that receives the data sent from the terminal, an "AI model" that analyzes the data and infers the pet's emotions and desires, and a "notification system" that notifies the user. It also includes an "emotion engine" that recognizes the user's emotions and suggests appropriate countermeasures.
[1967] Data collection
[1968] The device is equipped with a camera and microphone to collect pet movements, facial expressions, and sounds in real time. This data is divided into data packets at regular intervals (e.g., every second) and sent to a server using a security protocol (e.g., SSL / TLS).
[1969] Data reception and decryption
[1970] The server receives the data packets sent from the terminal, decrypts them according to the security protocol, and then classifies the received data into video data and audio data.
[1971] Data analysis
[1972] The server analyzes the video data using a computer vision algorithm (e.g., OpenCV) to identify the pet's movements and facial expressions (e.g., tail wagging and ear movement). It also analyzes the audio data using a speech recognition algorithm (e.g., LibriSpeech) to identify the pattern, pitch, and frequency of the pet's cries. This allows the extraction of features.
[1973] Estimating emotions and desires
[1974] The server inputs the extracted features into an AI model (e.g., TensorFlow or PyTorch). The AI model is pre-trained with data on pets' emotions and desires, and uses the input features to infer the pet's emotions and desires. For example, if the pet wags its tail and makes a high-pitched bark, it can infer that the pet wants to play.
[1975] Natural language translation and notification
[1976] The server translates the results output by the AI model into natural language. For example, the predicted result "I want to play" is converted into the message "Play with the ball!" The translated message is resent from the server to the device and notified to the user via the device.
[1977] Recognizing user emotions and proposing countermeasures
[1978] The device collects the user's voice and facial expressions using a camera and microphone. The collected user data is sent to a server where it is analyzed by an emotion engine. Using a voice recognition algorithm and an expression analysis algorithm, emotions are identified from the user's tone of voice and facial expressions. Based on this emotional information, the server suggests appropriate countermeasures. For example, if the user is tired, a message suggesting, "Let's create a relaxing atmosphere" is sent.
[1979] Specific examples
[1980] Example 1: System behavior when a dog wants to play
[1981] 1. The device captures the dog's movements with a camera and collects the dog's vigorously wagging tail and high-pitched barking sounds.
[1982] 2. The device sends the collected data to the server.
[1983] 3. The server receives the data and identifies tail movements from the video data and high-pitched cries from the audio data.
[1984] 4. The server inputs these features into the AI model.
[1985] 5. The AI model uses these characteristics to infer that the dog wants to play.
[1986] 6. The server translates the result "I want to play" into natural language and makes it "Play with the ball!"
[1987] 7. The device notifies the user with the message "Play with the ball!"
[1988] 8. The device collects the user's facial expressions and voice and sends them to the server.
[1989] 9. The emotion engine recognizes the user's emotion as "relaxed."
[1990] 10. The server checks the user's relaxation state and sends a message suggesting "Relax with your pet."
[1991] 11. The terminal displays the suggestion message to the user.
[1992] Example 2: System behavior when the cat is hungry
[1993] 1. The device collects the cat's meowing (a low, sustained sound) and its movements around the dish.
[1994] 2. The device sends the collected data to the server.
[1995] 3. The server receives the data and identifies the location from the video data (walking around the dish) and the sound from the audio data (a low, sustained sound).
[1996] 4. The server inputs these features into the AI model.
[1997] 5. The AI model uses these characteristics to infer that the cat is hungry.
[1998] 6. The server translates the result "hungry" into natural language and gives it the value "I'm hungry."
[1999] 7. The device notifies the user with the message "I'm hungry."
[2000] 8. The device collects the user's facial expressions and voice and sends them to the server.
[2001] 9. The emotion engine recognizes the user's emotion as "tired."
[2002] 10. The server checks the user's tiredness and sends a message suggesting that they take their time to feed their pet.
[2003] 11. The terminal displays the suggestion message to the user.
[2004] Example prompts for generative AI models
[2005] "We want you to infer what the dog wants from the dog's behavior and bark data collected by the device."
[2006] "I want you to analyze what the cat wants at the moment based on its meows and behavioral patterns."
[2007] Thus, the present invention provides a system that can analyze the emotions and desires of pets and users in real time and encourage users to respond appropriately, thereby improving the quality of pet care and facilitating communication between pets and users.
[2008] The flow of the identification process in the second embodiment will be described with reference to FIG.
[2009] Step 1:
[2010] Your pet's movements are collected on the device using a camera and microphone.
[2011] Input: Video and audio data captured through a camera and microphone.
[2012] Specific behavior: The device records your pet's movements, facial expressions, and sounds in real time, such as capturing a dog's tail wagging or a cat's meowing.
[2013] Output: Collected video and audio data.
[2014] Step 2:
[2015] Data collected by the terminal is divided into data packets and sent to a server using a security protocol.
[2016] Input: Collected video and audio data.
[2017] Specific operation: Data is packetized every second and transmitted using security protocols such as SSL / TLS.
[2018] Output: Data packets protected by security protocols.
[2019] Step 3:
[2020] The server receives and decodes the data packets sent from the terminal.
[2021] Input: Data packets protected by security protocols.
[2022] Specific operation: The data packets are decrypted using the SSL / TLS protocol and restored to the original video and audio data.
[2023] Output: Decoded video and audio data.
[2024] Step 4:
[2025] The server classifies the decoded data into video data and audio data.
[2026] Input: Decoded video and audio data.
[2027] Specific operation: Data is classified into video and audio categories based on data format and metadata.
[2028] Output: Classified video and audio data.
[2029] Step 5:
[2030] The server analyzes the video data using computer vision algorithms.
[2031] Input: Classified video data.
[2032] Specific operation: Using computer vision algorithms such as OpenCV, the system analyzes pet movements and facial expressions. For example, it detects the wagging of a dog's tail or the movement of a cat's ears.
[2033] Output: Features extracted from video data.
[2034] Step 6:
[2035] The server analyzes the voice data using a voice recognition algorithm.
[2036] Input: Classified audio data.
[2037] What it does: Uses speech recognition algorithms such as LibriSpeech to identify the pattern, pitch, and frequency of your pet's cries.
[2038] Output: Features extracted from the audio data.
[2039] Step 7:
[2040] The server inputs the features extracted from the video and audio data into an AI model to estimate the pet's emotions and desires.
[2041] Input: Features extracted from video and audio data.
[2042] Specific behavior: Using machine learning algorithms such as TensorFlow and PyTorch, the system infers a pet's emotions and desires from its behavior and sounds. For example, if a pet wags its tail and makes a high-pitched meow, it infers that it wants to play.
[2043] Output: Estimated pet's emotions and desires.
[2044] Step 8:
[2045] The server translates the estimated emotions and desires into natural language and notifies the user.
[2046] Input: Estimated pet emotions and desires.
[2047] Specific operation: The inference results are converted into messages such as "Play with the ball!" or "I'm hungry" and sent to the user's device.
[2048] Output: Notification message translated into natural language.
[2049] Step 9:
[2050] The device collects the user's voice and facial expressions and sends them to the server.
[2051] Input: User's audio and video data.
[2052] Specific operation: The camera and microphone are used to record the user's real-time voice and facial expressions, which are then divided into data packets and sent to the server.
[2053] Output: Collected user audio and video data.
[2054] Step 10:
[2055] The server analyzes the collected user data using voice recognition algorithms and facial expression analysis algorithms to identify the user's emotions.
[2056] Input: Collected user audio and video data.
[2057] Specific behavior: Identify the user's emotions by analyzing voice tone and facial expressions. For example, if the user's voice tone is low and they sound tired, it will recognize them as "tired."
[2058] Output: Identified user sentiment.
[2059] Step 11:
[2060] The server suggests measures to take for the pet based on the user's feelings.
[2061] Input: Identified user sentiment.
[2062] Specific behavior: Generate and send messages to the user's device suggesting things like "Create a relaxing atmosphere" and "Take your time to feed your pet without rushing."
[2063] Output: A message with the proposed workaround.
[2064] (Application example 2)
[2065] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[2066] Currently, there are several systems on the market that monitor pets' movements and cries to estimate their emotions, but these systems cannot provide clear and specific suggestions to users. Furthermore, they lack a mechanism for providing appropriate content based on pets' emotions and desires. Therefore, there is a need for a system that can help users choose the best course of action for their pets.
[2067] The identification process by the identification processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes a terminal for collecting the movements, expressions, and cries of the pet, a means for receiving data transmitted from the terminal, a means including an AI model for analyzing the received data and estimating the pet's emotions and desires, a means for translating the estimated emotions and desires into natural language and notifying the user, and a means for selecting appropriate content based on the estimated emotions and desires of the pet and delivering it to the user terminal. This allows the user to understand the pet's emotions and desires in real time and obtain optimal content in a timely manner.
[2068] "Pet movements, expressions, and sounds" refers to the physical movements, facial expressions, and sounds made by pets.
[2069] The "terminal" is a device equipped with a camera and microphone that collects your pet's movements, facial expressions, and cries.
[2070] A "server" is a computer system that receives, analyzes, and processes data sent from a terminal.
[2071] An "AI model" is an artificial intelligence that uses machine learning algorithms to analyze data and estimate a pet's emotions and desires.
[2072] "Estimation" refers to predicting a pet's emotions and desires from the data obtained.
[2073] "Translation into natural language" means converting the inferred results output by the AI model into a message that is easy for the user to understand.
[2074] "Notifying the user" means conveying messages or information to the user through the terminal.
[2075] A "security protocol" is a communication protocol for maintaining the security of data when it is transmitted from a terminal to a server.
[2076] A "computer vision algorithm" is an algorithm that analyzes video data and identifies pet movements and facial expressions.
[2077] A "voice recognition algorithm" is an algorithm that analyzes voice data and identifies the sounds of pets.
[2078] "Selecting appropriate content" means selecting videos, music, information, etc. that are suitable for your pet based on the pet's estimated emotions and desires.
[2079] "Delivery" means transmitting selected content to a user terminal and having it played or displayed.
[2080] System configuration overview
[2081] This invention provides a system that collects pet movements, facial expressions, and cries, estimates the pet's emotions and desires based on the collected data, and delivers appropriate content to the user. The system mainly consists of the following components:
[2082] A "terminal" for collecting your pet's movements, facial expressions, and cries.
[2083] The "server" receives and analyzes data collected from the terminal.
[2084] An "AI model" that analyzes the received data and estimates the pet's emotions and desires.
[2085] A "notification system" that selects appropriate content based on estimated emotions and desires and delivers it to the user's device.
[2086] Data collection details
[2087] The "terminal" is equipped with a camera and microphone to collect the pet's movements, facial expressions, and sounds in real time, and the collected data is divided into data packets at regular intervals and sent to the "server" using a security protocol.
[2088] Data reception and analysis
[2089] The server receives and decodes the data packets sent from the device. It then processes the video and audio data separately. The video data uses computer vision algorithms to identify the pet's facial expressions and movements, while the audio data uses speech recognition algorithms to identify the pattern, pitch, and frequency of the pet's cries.
[2090] Estimating emotions and desires
[2091] The server inputs the analyzed features into an AI model to estimate the pet's emotions and desires. The AI model uses a trained machine learning algorithm, and can infer, for example, that a pet's desire to play is indicated by a wagging tail and high-pitched meow.
[2092] Natural language translation and notification
[2093] The estimated emotions and desires are translated into natural language and converted into specific messages (e.g., "I'm hungry" or "Play with a ball!"). The translated messages are then sent from the server to the device and notified to the user.
[2094] Recognizing user emotions and proposing countermeasures
[2095] The device also collects the user's voice and facial expressions and sends them to the server. The emotion engine analyzes this data and recognizes the user's emotions. For example, it uses voice recognition and facial expression analysis algorithms to identify emotions from the user's tone and facial expressions.
[2096] Proposing solutions based on user sentiment
[2097] The server generates a message suggesting appropriate actions for the pet based on the user's emotional information. For example, if the user is tired, the server might suggest, "Take your time to feed your pet and don't rush." The suggested message is displayed on the user's device, encouraging the user to take appropriate action.
[2098] Hardware and software used
[2099] The system uses the following hardware and software:
[2100] Hardware: Camera, microphone, smartphone, smart glasses, head-mounted display.
[2101] Software: OpenCV (used to collect video data from the camera), TensorFlow (used to analyze pet and user emotions), requests (used to communicate data with the server), virtual voice analysis module and emotion recognition module.
[2102] Specific example explanation
[2103] For example, if it is estimated that the pet wants to play, it will be treated as "playful" and the video content "play_video.mp4" will be selected and played. If the user is relaxed, the message "Enjoy a relaxing session with your pet" will be displayed, suggesting a better way to spend time with the pet.
[2104] Prompt Sentence Examples
[2105] "Suggest content suitable for pets who want to play."
[2106] "Suggest actions that are appropriate for when the user is relaxed."
[2107] In this way, a system is realized that can analyze emotions and desires based on real-time data from pets and users, and deliver optimal responses and content.
[2108] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[2109] Step 1:
[2110] The device uses a camera and microphone to collect the pet's movements, facial expressions, and sounds in real time. It receives camera video and audio data as input and divides it into data packets at regular intervals. This captures multiple features, such as the pet's movements, facial expressions, and sounds, which can then be sent to subsequent analysis steps.
[2111] Step 2:
[2112] The terminal sends the collected data packets to the server using a security protocol. The terminal receives the data packets as input, encrypts the data using a security protocol (e.g., SSL / TLS), and securely transmits the data to the server. This ensures the confidentiality of the data.
[2113] Step 3:
[2114] The server receives data packets sent from the terminal. It receives encrypted data packets as input and first decrypts them. As output, it obtains decrypted video and audio data. This allows the data to be stored in an analyzable format on the server.
[2115] Step 4:
[2116] The server analyzes the video data to identify the pet's movements and facial expressions. It receives the decoded video data as input and analyzes it using a computer vision algorithm (e.g., OpenCV). The output is identified behavior and facial expression data, such as tail movements and facial expressions. This allows the pet's movement and facial expression features to be extracted.
[2117] Step 5:
[2118] The server analyzes the audio data to identify the pattern, pitch, and frequency of the pet's cries. It receives the decoded audio data as input and analyzes it using a speech recognition algorithm (e.g., a speech analysis module). The output is an analysis result that identifies the pattern, pitch, and frequency of the pet's cries. This allows the characteristics of the pet's cries to be extracted.
[2119] Step 6:
[2120] The server inputs the analyzed features into an AI model to estimate the pet's emotions and desires. The analysis results of movements, facial expressions, and cries are received as input and input into an AI model (for example, a model built with TensorFlow). The output is the pet's estimated emotions and desires (for example, "I want to play" or "I'm hungry"). This clarifies the pet's internal state.
[2121] Step 7:
[2122] The server translates the estimated emotions and desires into natural language and generates a message. It receives the estimated results from the AI model as input and translates them using a natural language processing engine. The output is a message that is easy for the user to understand, such as "I'm hungry" or "Play with my ball!" This communicates the pet's emotions and desires to the user.
[2123] Step 8:
[2124] The server selects appropriate content based on the estimated emotions and desires. It receives a message translated into natural language as input and uses an algorithm to select appropriate content (e.g., video, music). Specific content files (e.g., "play_video.mp4", "relaxing_music.mp3") are obtained as output. This allows content that matches the pet's state to be selected.
[2125] Step 9:
[2126] The server delivers the selected content to the user terminal. The server receives the selected content file as input and transmits it to the user terminal via the Internet. As output, the server delivers the content to be played or displayed on the user terminal. This allows the user to receive content that is appropriate for the condition of their pet.
[2127] Step 10:
[2128] The device collects the user's facial expressions and voice and sends them to the server. It receives camera footage and voice data as input, divides them into data packets, and sends them to the server. This provides data for analyzing the user's emotions.
[2129] Step 11:
[2130] The server analyzes the user's facial expressions and voice to recognize the user's emotions. It receives the user's facial expressions and voice data as input and analyzes them using an emotion recognition algorithm. The output identifies the user's emotions (e.g., "relaxed" or "tired"). This clarifies the user's state.
[2131] Step 12:
[2132] The server generates a message suggesting how to respond to a pet based on the user's emotional information. It receives the user's emotion recognition results as input and generates the suggested message using a natural language processing engine. The output is specific suggested actions such as "Relax with your pet" or "Take your time to feed your pet without rushing." This provides guidance for the user to take appropriate action.
[2133] The above is a description of the specific processing steps of the system for implementing the present invention.
[2134] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the control target 443 to output the result of the specific processing. The microphone 238 acquires voice indicating a user input regarding the result of the specific processing. The control unit 46A transmits voice data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the voice data.
[2135] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[2136] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the robot 414.
[2137] The emotion identification model 59 as an emotion engine may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to an emotion map (see FIG. 9), which is a specific mapping. Similarly, the emotion id...
Claims
1. A device to collect your pet's movements, expressions, and cries, a server that receives data transmitted from the terminal; means for analyzing the received data and including an AI model for inferring the pet's feelings and desires; A means for translating the estimated emotions and desires into natural language and notifying the user of the translation; A system including:
2. 10. The system of claim 1, including a security protocol for securely transmitting collected data.
3. 10. The system of claim 1, further comprising means for utilizing computer vision and speech recognition algorithms to infer the pet's emotions and desires.
Citation Information
Patent Citations
Persona chatbot control method and system
JP2022180282A